Skip to content

Commit 5dfa990

Browse files
author
Ronald Tse
committed
release: tha-g2p-base-1.0 (gated ByT5-base student, +4.76pp PER)
Sequence-level KD from B-K/umt5 teacher (4.43% PER): 48,757 beam-4 labels, gate student 9.19% (+4.76 <= +5pp). Distill runner hardened along the way: resumable beam-4 labeling with OOM split-retry, token-budget batching, byte-length truncation in the training collate (the OOM saga's root cause), gradient checkpointing, degenerate-label filtering, teacher recovery script with secryst-recipe stages.
1 parent 53a837a commit 5dfa990

5 files changed

Lines changed: 269 additions & 73 deletions

File tree

Lines changed: 22 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,22 @@
1+
# tha-g2p-base-1.0
2+
3+
Thai grapheme-to-phoneme (IPA). Client-tier ByT5-base (580M) student
4+
distilled from the B-K/umt5-thai-g2p teacher via sequence-level KD:
5+
48,757 beam-4 teacher-generated labels over the Kaikki + epitran-Wikipedia
6+
corpus (deduplicated, degenerate outputs filtered).
7+
8+
Gate: student 9.19% PER vs teacher 4.43% on the same harness
9+
(1,219 Kaikki test sentences, beam-4, corpus-level PER) — +4.76pp,
10+
inside the +5pp distillation budget (docs/DISTILL-SOURCE-PROMPT.md).
11+
12+
Note on the teacher: the secryst 2.32%-PER umt5 artifacts are
13+
unrecoverable (transformers 5.15 save drops the untied umt5 lm_head;
14+
the volume's epitran corpus is tone-less). The 2.32% tier re-enters
15+
this pipeline when secryst ships repaired artifacts; this model
16+
distills the best verified teacher available (4.43%).
17+
18+
```python
19+
from interscript_ml import Model
20+
model = Model.load("tha-g2p-base-1.0")
21+
model.translate("สวัสดี")
22+
```
Lines changed: 32 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,32 @@
1+
format: imf-v1
2+
id: tha-g2p-base-1.0
3+
task: g2p
4+
source_script: Thai
5+
target: IPA
6+
tokenizer: bytes
7+
opset: 14
8+
decoder: kv
9+
precision: fp32
10+
license: BSD-3-Clause
11+
trained_from: >-
12+
sequence-level KD from the B-K/umt5-thai-g2p-v2-0.5k teacher (4.43%
13+
PER on this harness; secryst's saved umt5 artifacts are unusable —
14+
transformers 5.15 dropped the untied lm_head — and the volume's
15+
epitran corpus is tone-less, so the published 2.32% tier is
16+
unrecoverable until secryst regenerates it); 48,757 beam-4
17+
teacher-generated labels; ByT5-base init google/byt5-base; checkpoint
18+
secryst-checkpoints:/secryst_thai_g2p_distill_small/run-004/best
19+
metrics:
20+
- name: per_teacher
21+
value: 4.43
22+
protocol: >-
23+
beam-4, corpus-level PER (total_ed/total_gold over chars of
24+
joined-piece decode); 1,219 Kaikki Thai test sentences;
25+
B-K/umt5-thai-g2p-v2-0.5k teacher
26+
source: interscript/ml-models src/gpu/modal_distill.py::evaluate_per
27+
- name: per_student
28+
value: 9.19
29+
protocol: >-
30+
beam-4, corpus-level PER, same harness as the teacher; gate
31+
+4.76pp <= +5pp (docs/DISTILL-SOURCE-PROMPT.md); exact match 90.81%
32+
source: interscript/ml-models release tha-g2p-base-1.0

0 commit comments

Comments
 (0)