Skip to content

Commit d8508cb

Browse files
author
Ronald Tse
committed
teacher measured at greedy: PER 1.25% — same-protocol client-tier gap is +1.60pp
The beam pathology inflated the teacher too (4.43 beam-4 → 1.25 greedy, n=1219, EM 95.16%). The Thai client tier's true shrink cost is +1.60pp greedy-to-greedy, not the +7.63pp the beam-vs-beam comparison suggested. Frontier figure, RESULTS table, metadata, metrics chain all updated.
1 parent 3d497b8 commit d8508cb

4 files changed

Lines changed: 19 additions & 7 deletions

File tree

‎docs/RESULTS.md‎

Lines changed: 9 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -85,10 +85,15 @@ better than the beam-4 harness numbers.** Re-measured on the shipped
8585
int4 zip through the Python runtime (the exact ONNX KV decode users
8686
get), true Levenshtein, full 1,219-sentence set:
8787

88-
| Decode | PER | Exact match |
89-
|---|---|---|
90-
| beam-4 (published, torch harness) | 12.06% | 87.94% |
91-
| **greedy (runtime protocol)** | **2.85%** | **88.93%** |
88+
| Decode | Teacher PER | Student PER | Student EM |
89+
|---|---|---|---|
90+
| beam-4 (published, torch harness) | 4.43% | 12.06% | 87.94% |
91+
| **greedy (runtime protocol)** | **1.25%** | **2.85%** | **88.93%** |
92+
93+
The teacher is also affected by the beam pathology (4.43 beam-4 → 1.25
94+
greedy, measured 2026-08-25 through the same harness at num_beams=1):
95+
same-protocol, the client tier's true shrink cost is **+1.60pp**, not
96+
the +7.63pp the beam-vs-beam comparison suggested.
9297

9398
The beam-4 numbers are inflated by length-normalized beam preferring
9499
long garbage on this model's flat per-token distributions (top-1

‎docs/paper-assets/frontier.png‎

1.36 KB
Loading

‎models/metrics-sources.yaml‎

Lines changed: 4 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -69,13 +69,14 @@ heb-diac-small-1.0:
6969
- {row: "Student (ByT5-small, gate)", column: DER, as: der_student_greedy}
7070
tha-g2p-small-1.0:
7171
repo: interscript/interscript-ml
72-
ref: docs/greedy-correction
72+
ref: docs/paper-ci
7373
path: docs/RESULTS.md
7474
anchor: tha-g2p-small-10-thai-g2p-client-tier-2026-08-22
7575
protocol: "greedy decode via the Python runtime (the shipped ONNX KV path), corpus-level PER, true Levenshtein; 1,219 held-out Kaikki Thai test sentences"
7676
tables:
77-
- {row: "beam-4 (published, torch harness)", column: PER, as: per_student_beam4}
78-
- {row: "greedy (runtime protocol)", column: PER, as: per_student}
77+
- {row: "greedy (runtime protocol)", column: Teacher PER, as: per_teacher_greedy}
78+
- {row: "beam-4 (published, torch harness)", column: Student PER, as: per_student_beam4}
79+
- {row: "greedy (runtime protocol)", column: Student PER, as: per_student}
7980
ara-diac-small-1.0:
8081
repo: interscript/interscript-ml
8182
ref: release/ara-diac-small-1.0

‎models/tha-g2p-small/tha-g2p-small-1.0.metadata.yaml‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -17,6 +17,12 @@ trained_from: 'sequence-level KD from the B-K/umt5-thai-g2p-v2-0.5k teacher (4.4
1717
this is the smallest rung that does not collapse (docs/RESULTS.md frontier table,
1818
~300MB int8).'
1919
metrics:
20+
- name: per_teacher_greedy
21+
value: 1.25
22+
protocol: greedy decode (num_beams=1) through the same harness; 1,219 Kaikki Thai
23+
test sentences; exact match 95.16%; measured 2026-08-25 — the published beam-4
24+
teacher figure 4.43 was itself inflated by the decode pathology
25+
source: interscript/interscript-ml docs/RESULTS.md#tha-g2p-small-1.0
2026
- name: per_student_beam4
2127
value: 12.06
2228
protocol: beam-4, corpus-level PER, same harness as the teacher (src/gpu/ modal_distill.py::evaluate_per,

0 commit comments

Comments
 (0)