@@ -6,12 +6,12 @@ v1.0, 2026-09-05
66
77== Introduction
88
9- Two weeks ago the phonological layer shipped with an honest gap: the
10- Arabic client student scored 8.26 on the full benchmark — a disclosed
11- miss. This week that gap closed the only way gaps should close: every
12- rung measured on the full set, confidence intervals on every
13- separation, and the two causal hypotheses that remained both tested to
14- a verdict. One verdict was negative. That is the point .
9+ Two weeks ago the phonological layer shipped with a disclosed gap:
10+ the Arabic client model scored 8.26 on the full benchmark. This week
11+ every rung was measured on the full test set, with confidence
12+ intervals on each separation, and the two remaining hypotheses were
13+ both tested. One test came back negative, and that result is published
14+ here with the rest .
1515
1616== The frontier, bracketed
1717
@@ -76,8 +76,7 @@ Every leaderboard number we publish can now be re-derived by anyone:
7676
7777The tool reproduces our published verdicts exactly — it re-derives this
7878week's 4.8231 run from its raw predictions and the public benchmark,
79- intervals included. Protocol-matched comparison should be a command,
80- not a promise.
79+ intervals included. The comparison is a command, not a claim.
8180
8281== Where this leaves the stack
8382
0 commit comments