Skip to content

Commit 1140883

Browse files
author
Ronald Tse
committed
feat: ara-diac-small-2.0 retargeted to run-006 — E4 landed at 4.8218
The E4 arm completed as pre-registered (gate ≤6.26, prediction 4.3-5.0): r7 canonical teacher labels + Muon on the vanilla ByT5-small scores 4.8218 full-set (teacher reproduces 2.289), superseding the run-005 rung (5.2945) this branch originally targeted — as its own metadata anticipated. 42% error reduction vs the shipped 1.0 at identical architecture and artifact size. Strict teacher+0.5pp gate still missed (+2.53pp), disclosed.
1 parent ada4e93 commit 1140883

5 files changed

Lines changed: 53 additions & 28 deletions

File tree

‎docs/EXPERIMENTS.md‎

Lines changed: 5 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -152,7 +152,11 @@ All rows passed the CER parity gate at release. Readings:
152152

153153
## E4 — ara-diac-small-2.0 candidate (run-006-r7-muon)
154154

155-
- **Status:** queued (registered 2026-08-28 before launch).
155+
- **Status:** COMPLETE (2026-08-29). **PASSED — 4.8218** full-set
156+
windowed DER (gate ≤ 6.26; registered prediction 4.3–5.0; teacher r7
157+
reproduces 2.289 vs documented 2.2864). −3.44pp / 42% relative vs the
158+
shipped 1.0 at identical architecture and size; matches the PKM arm's
159+
4.829 without memory layers. Released as ara-diac-small-2.0.
156160
- **Hypothesis:** the two measured wins compound — the r7 canonical
157161
teacher (better labels; ID 2.2864 vs 2.5793) plus the E3-adopted
158162
Muon optimizer (−2.727pp on r6 labels) — moving the client rung far

‎docs/RESULTS.md‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -315,6 +315,25 @@ gap) — roughly additive, slightly sub-additive on memory. Residual
315315
~2.2pp is domain coverage. Optimization is the dominant recoverable
316316
term of the distillation gap at the ByT5-small rung.
317317

318+
### run-006-r7-muon — E4: the 2.0 release candidate (2026-08-29)
319+
320+
The two measured wins compounded on the vanilla architecture: r7
321+
canonical teacher (fresh greedy labels) + Muon optimizer, same
322+
corpus/limits/seed family. Pre-registered E4 gate ≤ 6.26 (prediction
323+
4.3–5.0):
324+
325+
| Model | DER-CE (full 1,200) |
326+
|---|---|
327+
| Teacher (r7, in-run) | 2.2890% |
328+
| **ByT5-small, r7 labels + Muon (run-006)** | **4.8218%** |
329+
| ByT5-small, r6 labels + Muon (run-005) | 5.2945% |
330+
| ByT5-small, r6 labels + AdamW (run-002, shipped 1.0) | 8.2590% |
331+
332+
**4.8218 — gate passed; −3.44pp / 42% relative vs the shipped 1.0** at
333+
identical architecture and artifact size. Matches the PKM arm's 4.829
334+
without the memory layers. Release: ara-diac-small-2.0 (run-006
335+
checkpoint; strict teacher+0.5pp still missed at +2.53pp, disclosed).
336+
318337
Leaderboard context (SadeedDiac-25, Misraj evaluator, zero-skip,
319338
harakat-projected DER-CE): the teacher tier (r6, 580M) at 2.5793
320339
(reproduced at 2.5815, 2026-08-26) is the best dedicated model measured

‎models/ara-diac-small/ara-diac-small-2.0.README.md‎

Lines changed: 10 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -1,17 +1,18 @@
11
# ara-diac-small-2.0
22

33
Arabic diacritization (haraqat restoration). Client-tier ByT5-small
4-
student, identical corpus/labels/epochs to ara-diac-small-1.0 with one
5-
variable changed: **the Muon optimizer** (E3 factorial). That single
6-
change closes 2.96pp of the 5.68pp teacher-student gap:
4+
student — the two measured wins of the campaign compounded on the same
5+
architecture and artifact size as 1.0: the **r7 canonical teacher**
6+
(2.2864; fresh greedy labels) and the **Muon optimizer** (E3-adopted).
77

8-
- 1.0 (AdamW): 8.26 full-set windowed DER-CE
9-
- **2.0 (Muon): 5.29** (teacher r6: 2.58; in-run reproduction 2.60)
8+
- 1.0 (r6 labels, AdamW): 8.26 full-set windowed DER-CE
9+
- 2.0 (r7 labels, Muon): **4.82** (teacher r7 in-run: 2.289)
1010

11-
The strict teacher+0.5pp gate is still missed; the residual decomposes
12-
as ~0.70pp capacity + ~2.25pp domain coverage (E2/E3 factorial), and an
13-
r7-teacher re-distillation is in flight. Identical IMF v1 contract:
14-
dynamic fetch, sha256-verified, KV decode, margins JSON alongside.
11+
A 42% error reduction, pre-registered as E4 (gate ≤ 6.26; prediction
12+
4.3–5.0 — landed at 4.82). The strict teacher+0.5pp gate is still
13+
missed (+2.53pp, disclosed); the E2/E3 factorial attributes the
14+
residual to domain coverage. Identical IMF v1 contract: dynamic fetch,
15+
sha256-verified, KV decode, margins JSON alongside.
1516

1617
```python
1718
from interscript_ml import Model

‎models/ara-diac-small/ara-diac-small-2.0.metadata.yaml‎

Lines changed: 18 additions & 17 deletions
Original file line numberDiff line numberDiff line change
@@ -8,25 +8,26 @@ opset: 14
88
decoder: kv
99
precision: fp32
1010
license: BSD-3-Clause
11-
trained_from: 'sequence-level KD from the r6 teacher (rababa_arabic_byt5/run-006-morph/best,
12-
2.5793 windowed DER-CE full protocol): the same 29,322 greedy teacher labels as
13-
ara-diac-small-1.0, with the Muon optimizer (E3 factorial, vanilla student arm).
14-
Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-005-muon/best. The
15-
optimizer alone closes 2.96pp of the 5.68pp teacher-student gap measured for the
16-
1.0 (AdamW) release: 8.259 -> 5.2945 full-set windowed DER-CE (teacher reproduces
17-
2.5997 in-run). Still misses the strict teacher+0.5pp gate — the residual decomposes
18-
as ~0.70pp capacity + ~2.25pp domain coverage per the E2/E3 factorial; miss disclosed.
19-
An r7-teacher re-distillation is in flight and expected to supersede as a later rung.'
11+
trained_from: 'sequence-level KD from the r7 canonical teacher (rababa_arabic_byt5/run-007-news/best,
12+
2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units
13+
corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small
14+
(E4, pre-registered gate <= 6.26). Checkpoint
15+
rababa-checkpoints:/rababa_arabic_distill_small/run-006-r7-muon/best. The two
16+
measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher reproduces
17+
2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0 release at the
18+
same architecture and artifact size. Still misses the strict teacher+0.5pp gate
19+
(+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
20+
coverage.'
2021
metrics:
2122
- name: der_teacher_fullset
22-
value: 2.5997
23+
value: 2.289
2324
protocol: windowed DER-CE (1400-byte windows, word-boundary split, greedy,
2425
haraqat-projected, Misraj evaluator); full 1,200-paragraph SadeedDiac-25;
25-
in-run reproduction of the documented 2.5815/2.5793
26-
source: interscript/interscript-ml docs/RESULTS.md#run-005-muon
26+
in-run reproduction of the documented 2.2864 (r7 canonical teacher)
27+
source: interscript/interscript-ml docs/RESULTS.md#run-006-r7-muon
2728
- name: der_student_fullset
28-
value: 5.2945
29-
protocol: same full-set harness; +2.69pp over the in-run teacher — the Muon
30-
arm of the E3 2x2 factorial (vanilla student, optimizer as the only variable
31-
vs the 8.259 AdamW release)
32-
source: interscript/interscript-ml docs/RESULTS.md#run-005-muon
29+
value: 4.8218
30+
protocol: same full-set harness; E4 (r7 teacher labels + Muon, vanilla
31+
ByT5-small) vs the 8.259 AdamW/r6-labels 1.0 release — a 42% error
32+
reduction at identical architecture and artifact size
33+
source: interscript/interscript-ml docs/RESULTS.md#run-006-r7-muon

‎src/gpu/modal_export.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -116,7 +116,7 @@
116116
},
117117
"ara-diac-small-2": {
118118
"volume": "/volumes/rababa-checkpoints",
119-
"checkpoint": "rababa_arabic_distill_small/run-005-muon/best",
119+
"checkpoint": "rababa_arabic_distill_small/run-006-r7-muon/best",
120120
"metadata": "models/ara-diac-small/ara-diac-small-2.0.metadata.yaml",
121121
"readme": "models/ara-diac-small/ara-diac-small-2.0.README.md",
122122
"test_volume": "/datasets/rababa",

0 commit comments

Comments
 (0)