Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions TODO.sota/00-README.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,9 @@ hf jobs run --flavor a100-large -d \
| 14 | 14-ts-plane-port.md | P2 | — | DONE (ts#99/#100/#101, CI leg models#265) |
| 15 | 15-api-edge.md | P2 | 14 | DONE (api#29: edge-first /v1/infer, kind in index) |
| 16 | 16-parking-lot.md | — | — | parked (not now) |
| 17 | 17-arabic-register-closure.md | P1 | 11 | r8a RUNNING; r8b launching |
| 18 | 18-parity-runs.md | P1 | 10 | pending |
| 19 | 19-lexicon-loader.md | P2 | 13 | pending |

## Schedule

Expand Down
21 changes: 21 additions & 0 deletions TODO.sota/17-arabic-register-closure.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
# WO17 — Arabic register closure (the WikiNews gap attack)

**Why:** WO01 proved the 17.38-vs-2.70 WikiNews-2024 gap is register,
not vocabulary (OOV 1.61%). QCRI's own README confirms their news model
is trained on 5M words of Wikipedia SILVER labels from their BiLSTM —
the same noisy-teacher technique as our r8. Two arms decide which
teacher's silver closes our gap:

- **r8a (self)**: r7 pseudo-labels arwiki (RUNNING, label-v4; keep_frac
filter at train time). Our teacher: 2.2864 on the harder
expert-reviewed SadeedDiac-25.
- **r8b (QCRI silver)**: their published Wikipedia_20240420.diac.jsonl
(95MB, ~5M words, machine-labeled by their BiLSTM ~3% WER). Published
dataset — same standing as Tashkeela/Nakdimon in our stack; NOT an LLM
(no-LLM rule intact). Provenance + license recorded in the dataset
repo.

**Dual-surface gate (both arms, in-job):** SadeedDiac-25 DER (hold the
2.2864 line) AND WikiNews-2024 multiref WER/DER (move 17.38/11.83).
A model that wins OOD but collapses ID is a regression, not a win.
Ship the arm that dominates both; if neither does, negative verdict.
7 changes: 7 additions & 0 deletions TODO.sota/18-parity-runs.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# WO18 — parity runs for the shipped artifacts

1. Attach golden/heb-diac-plane-2.0.jsonl to the golden-v2 release.
2. Dispatch neural-parity (3 legs: ruby, python, ts) for
heb-diac-plane-2.0 — first model validated across ALL runtimes at
ship time.
3. Green runs recorded in RESULTS.md.
6 changes: 6 additions & 0 deletions TODO.sota/19-lexicon-loader.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
# WO19 — lexicon fetch helper (client convenience)

`thai_hybrid.fetch_lexicon()`: resolve the canonical release asset
(tha-lexicon-kaikki-1.0) with sha256 verification into a Lexicon —
same trust contract as model artifacts (no unverified data paths).
TDD; py first (ts/ruby follow on demand).
Loading