From cb5e49e38e2aeba4664f94cd1b41744ff53d15ea Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Wed, 7 Oct 2026 18:09:58 +0800 Subject: [PATCH] =?UTF-8?q?docs:=20TODO.sota=2017-19=20=E2=80=94=20registe?= =?UTF-8?q?r=20closure=20arms,=20parity=20runs,=20lexicon=20loader?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- TODO.sota/00-README.md | 3 +++ TODO.sota/17-arabic-register-closure.md | 21 +++++++++++++++++++++ TODO.sota/18-parity-runs.md | 7 +++++++ TODO.sota/19-lexicon-loader.md | 6 ++++++ 4 files changed, 37 insertions(+) create mode 100644 TODO.sota/17-arabic-register-closure.md create mode 100644 TODO.sota/18-parity-runs.md create mode 100644 TODO.sota/19-lexicon-loader.md diff --git a/TODO.sota/00-README.md b/TODO.sota/00-README.md index 9cf8a9a..8519322 100644 --- a/TODO.sota/00-README.md +++ b/TODO.sota/00-README.md @@ -54,6 +54,9 @@ hf jobs run --flavor a100-large -d \ | 14 | 14-ts-plane-port.md | P2 | — | DONE (ts#99/#100/#101, CI leg models#265) | | 15 | 15-api-edge.md | P2 | 14 | DONE (api#29: edge-first /v1/infer, kind in index) | | 16 | 16-parking-lot.md | — | — | parked (not now) | +| 17 | 17-arabic-register-closure.md | P1 | 11 | r8a RUNNING; r8b launching | +| 18 | 18-parity-runs.md | P1 | 10 | pending | +| 19 | 19-lexicon-loader.md | P2 | 13 | pending | ## Schedule diff --git a/TODO.sota/17-arabic-register-closure.md b/TODO.sota/17-arabic-register-closure.md new file mode 100644 index 0000000..55d86e6 --- /dev/null +++ b/TODO.sota/17-arabic-register-closure.md @@ -0,0 +1,21 @@ +# WO17 — Arabic register closure (the WikiNews gap attack) + +**Why:** WO01 proved the 17.38-vs-2.70 WikiNews-2024 gap is register, +not vocabulary (OOV 1.61%). QCRI's own README confirms their news model +is trained on 5M words of Wikipedia SILVER labels from their BiLSTM — +the same noisy-teacher technique as our r8. Two arms decide which +teacher's silver closes our gap: + +- **r8a (self)**: r7 pseudo-labels arwiki (RUNNING, label-v4; keep_frac + filter at train time). Our teacher: 2.2864 on the harder + expert-reviewed SadeedDiac-25. +- **r8b (QCRI silver)**: their published Wikipedia_20240420.diac.jsonl + (95MB, ~5M words, machine-labeled by their BiLSTM ~3% WER). Published + dataset — same standing as Tashkeela/Nakdimon in our stack; NOT an LLM + (no-LLM rule intact). Provenance + license recorded in the dataset + repo. + +**Dual-surface gate (both arms, in-job):** SadeedDiac-25 DER (hold the +2.2864 line) AND WikiNews-2024 multiref WER/DER (move 17.38/11.83). +A model that wins OOD but collapses ID is a regression, not a win. +Ship the arm that dominates both; if neither does, negative verdict. diff --git a/TODO.sota/18-parity-runs.md b/TODO.sota/18-parity-runs.md new file mode 100644 index 0000000..835caf5 --- /dev/null +++ b/TODO.sota/18-parity-runs.md @@ -0,0 +1,7 @@ +# WO18 — parity runs for the shipped artifacts + +1. Attach golden/heb-diac-plane-2.0.jsonl to the golden-v2 release. +2. Dispatch neural-parity (3 legs: ruby, python, ts) for + heb-diac-plane-2.0 — first model validated across ALL runtimes at + ship time. +3. Green runs recorded in RESULTS.md. diff --git a/TODO.sota/19-lexicon-loader.md b/TODO.sota/19-lexicon-loader.md new file mode 100644 index 0000000..12336a2 --- /dev/null +++ b/TODO.sota/19-lexicon-loader.md @@ -0,0 +1,6 @@ +# WO19 — lexicon fetch helper (client convenience) + +`thai_hybrid.fetch_lexicon()`: resolve the canonical release asset +(tha-lexicon-kaikki-1.0) with sha256 verification into a Lexicon — +same trust contract as model artifacts (no unverified data paths). +TDD; py first (ts/ruby follow on demand).