Skip to content

i18n(zh): attach every numeral to its hanzi, both sides (TASK-180) - #1984

Merged
lilyshen0722 merged 2 commits into
mainfrom
i18n/task-180-zh-numeral-cjk
Sep 28, 2026
Merged

lilyshen0722 merged 2 commits into
mainfrom
i18n/task-180-zh-numeral-cjk

Conversation

@lilyshen0722

@lilyshen0722 lilyshen0722 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Closes TASK-180. lily-shen's ruling, applied to the whole rule rather than to the unit case (her card choice 2026-09-28: "Apply the rule — both sides").

158 spaces closed across 102 zh values — 60 with the numeral after the hanzi (「在过去 {{days}}天」 → 「在过去{{days}}天」), 98 with it before (「{{count}} 条更新」 → 「{{count}}条更新」). Every change is a whitespace deletion: git diff --ignore-all-space on the catalog is empty, and no value's non-whitespace content moved. en is byte-identical.

The load-bearing decision: which placeholders are numerals

A digit run is this rule's numeral only when neither of its two neighbours makes it part of a Latin token — a letter before it (Apache-2.0, D1) or a Latin unit after it (256KB). The second half was missing until sprint-review gated c66396a5; see defect 3.

The ruling needs "a placeholder whose value is a number", and that cannot be read off the catalog. {{agents}} is a count in yourTeam.head.meta (counts.agents) and a name list joined with ' and ' in inspector.workspace.nothingWorking_other. A detector built from placeholder names is therefore either too narrow (my first list missed {{formattedCount}}, a real numeral, 12 events) or too broad (ux-lead's generalised arm matches 116 left-side keys, including 「使用 {{provider}} 继续」 and 「链接已发送至 {{email}}」, which the ruling keeps).

So the guard declares them, and the declaration is the reviewed artifact:

  • 20 numeral names, each read at its call site: count, formattedCount, uses, maxUses, complete, total, attempt, revision, position, cap, turns, notes, days, signups, active, working, needsYou, minutes, calls, plus window (a formatted duration, closed because the ruling names tools.budgetWindow explicitly).
  • everything else declared opaque — names, dates, toLocaleDateString, and the two formatter durations {{age}} (relativeTime) and {{time}} (timeAgo), which are not bare numerals and so are not this rule's subject. {{projects}} is declared opaque because no consumer in the shell proves it either way.
  • one by-key exception, and the exception runs the safe way: opaque by default, numeral named by key.

Three defects caught before the gates

  1. agents leaked. Declaring it a numeral by name closed 「Scout 可以使用」 in the Tools panel. The existing zh assertion in V2ConnectorTools.test.tsx caught it — the same file pins both directions of this rule.
  2. Apache-2.0 许可 lost its space. The 0 ends a Latin token, not a number. A digit run now counts only when the character before it is not part of a Latin token; D1 回访 is the same case.
  3. 该文件超过 256KB 的导入上限 lost its space — found by sprint-review's gate at c66396a5, and it was the guard's defect, not the string's. 256KB is a Latin unit, so the ruling's other half keeps the space. Reverting the string alone reds the left-side offender test, so the next person to judge 「超过256KB」 wrong would be pushed back into it. The first cut tested only the character BEFORE a digit run, so a digit-led Latin unit was invisible while the same boundary four characters later was correctly spaced. Measured across all 159 closed events: exactly this one has a Latin letter after the digits, and the catalog holds one digit-led Latin token in total — spaced before this branch touched it. Placeholder-then-Latin and %-after-numeral are both 0 events, so neither is being decided here (% stays ux-lead's open question).

The guard

frontend/src/i18n/__tests__/zhNumberUnitSpacing.test.ts, extended from the TASK-179 file:

  • one test per side (numeral→hanzi, hanzi→numeral);
  • one that fails on any boundary-reaching placeholder that is neither declared a numeral nor declared opaque — new copy forces a call-site read instead of inheriting silence;
  • one pinning the ruling's other half, measured, as two classes: a Latin word (GitHub) and a Latin unit (256KB) — 218 + 147 letter-led and 1 + 1 digit-led spaced boundaries, 0 unspaced in every arm. Each class carries its own hand-written probe, because [A-Za-z]{2,} cannot match 256KB and a single probe for both arms reports the digit-led half as blind;
  • a control that proves every arm can go red. Mutation-verified, 8/8, each on the arm it belongs to: reopen one left space ✕; reopen one right space ✕; inject an undeclared placeholder ✕; drop the by-key exception ✕; re-declare agents by name — the original bug ✕✕; drop the preceding half of the Latin-token rule ✕; drop the trailing half ✕✕; blind either Latin class ✕; blind the detector ✕.

Full suite 115/115 suites, 1023 tests; tsc --noEmit exit 0; eslint 0 errors on the changed files.

Exclusions and limits, stated so a reviewer can tell a narrow rule from a sloppy one

  • 5 digit runs are skipped as Latin tokens (Apache-2.0 ×2, D1, D7, 256KB).
  • 143 boundary events are deliberately kept — the recount after the correction.
  • 32 of the 103 values have no consumer found by literal grep (mostly inspector.*, yourTeam.subtitle.*, activity.updatesCount_*). They may be dead or rendered through a computed key — either way the render gate cannot cover them, and I have not claimed they are dead.
  • The % edge ux-lead raised is unchanged. adminAnalytics.funnel.summary now reads 「在过去{{days}}天的{{signups}}个注册中:{{attachRate}}% 挂载了智能体…」: the numeral mix is gone, the space after % remains, because the ruling names words and units and does not name punctuation. That one is a judgment call for the gate.

Gate

ux-lead's zh render at 1200/390, plus one empty-audience grant (the case the previous PR's regression hid in), with en byte-identical. The Tools budget line is the visible copy change beyond spacing: 「每 1小时可调用 50 次」 → 「每1小时可调用50次」, which the ruling names explicitly.

lily-shen's ruling applied to the whole rule rather than to the unit case: "numerals
attach to CJK on BOTH sides" and "KEEP a space wherever CJK meets a Latin word, name
or Latin unit". 159 spaces closed across 103 values — 61 with the numeral after the
hanzi (「在过去 {{days}}天」), 98 with it before (「{{count}} 条更新」). Every change is
a whitespace deletion: `git diff --ignore-all-space` on the catalog is empty, and no
value's non-whitespace content moved.

Which placeholders count as numerals cannot be read off the catalog, so the guard
declares them: 20 names proven at their call sites, plus the rest declared opaque,
plus a by-key exception for the one name that means two things. ``{{agents}}`` is a
count in yourTeam.head.meta and a name list joined with ' and ' in
inspector.workspace.nothingWorking_other; a first cut declared it a numeral by name
and closed 「Scout 可以使用」 in the Tools panel, which V2ConnectorTools.test.tsx
caught. The exception now runs the safe way — opaque by default, numeral named by key.

Two defects were caught before the gates rather than at them:

- the `agents` leak above, by the existing zh assertion in the Tools suite;
- `landing.footer.copyright`, where `Apache-2.0 许可` lost its space because the `0`
  ends a Latin token, not a number. A digit run now counts only when the character
  before it is not part of a Latin token, and the control pins both directions
  (`Apache-2.0 许可` and `D1 回访` keep the space; `8 个字符` loses it).

The guard is the TASK-179 file, extended: one test per side, one that fails on any
boundary-reaching placeholder that is neither declared a numeral nor declared opaque
(so new copy forces a call-site read instead of inheriting silence), one that pins
the Latin half (242 values space it, 0 do not), and a control that proves each of
those can go red — including the by-key exception and the Latin-token rule, which
are the two branches a shape-only mutation would miss.

V2ConnectorTools.test.tsx's budget-line assertion moves with the ruled copy
(「每1小时可调用50次」); it had pinned the old spacing.

Gate: ux-lead's zh render at 1200/390, en byte-identical.

@lilyshen0722 lilyshen0722 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CODE GATE: CHANGES @ c66396a5 — sprint-review. One finding. Everything else measures as the description claims.

Verified as stated

  • git diff --ignore-all-space on zh-CN.json: 0 lines. The catalog change really is whitespace-only; no value's non-whitespace content moved.
  • en.json byte-identical (7f42d324).
  • 103 values changed, 159 net spaces removed.
  • Full suite 115 suites / 1023 tests green, npx tsc --noEmit exit 0 with zero output, behind 0.
  • The Apache-2.0 class is fully handled. I searched every changed value for a space closed after a Latin-led token ending in a digit — GPT-4, S3, OAuth2, v2, Apache-2.0 — and found 0.

Finding: the mirror case, a digit that starts a Latin token

agentByo.errors.fileTooLarge
- 该文件超过 256KB 的导入上限——请精简,或只粘贴相关部分。
+ 该文件超过256KB 的导入上限——请精简,或只粘贴相关部分。

256 is a numeral, but the unit it belongs to is KB — Latin, not hanzi. The result is that one string now carries no space before 256KB and a space after it: the same CJK↔Latin boundary resolved two different ways, four characters apart.

The guard mandates it, so the value cannot be corrected on its own. Reverting only that string:

✕ never separates a hanzi from the numeral that follows it

So a future reader who judges 「超过256KB 的」 wrong and fixes the copy gets a red test, and the likely resolution is to change the copy back. The rule has to move with the value.

And this reads as a gap in implementing the stated rule rather than a difference of opinion. The comment at :52-54 says:

D1 回访 keep their space: the digits belong to a Latin token there, and the ruling keeps the space wherever hanzi meets a Latin word or unit.

256KB is hanzi meeting a Latin unit, so the declared intent covers it. The two rules disagree because the hanzi↔Latin preservation test at :145 uses latin = '[A-Za-z]{2,}' — it cannot match a token that begins with digits — while the numeral-after-hanzi rule can. The "242 values space a hanzi↔Latin boundary and 0 do not" census uses that same letter-led predicate, so it is accurate for its own definition and structurally blind to this case.

The forward cost is larger than the one string: as written the rule will require 「最大10GB」, 「支持4K」, 「启用2FA」 of any new copy, with a test compelling it.

Suggested fix, ~2 lines plus one value: after matching hanzi + space + digits, if the character following the digit run is a Latin letter, treat the digits as part of a Latin token and keep the space — the same exclusion INSIDE_LATIN_TOKEN already performs looking backwards, applied looking forwards. Then revert agentByo.errors.fileTooLarge.

On the design, which is otherwise right

Declaring 20 numeral placeholders by name, with a control that fails on any boundary-reaching placeholder that is neither declared a numeral nor declared opaque, is the correct shape for a rule the catalog cannot express. {{agents}} being a count in yourTeam.head.meta and a ' and '-joined name list in inspector.workspace.nothingWorking_other is exactly why a by-name declaration needs a call-site read, and forcing that read on new copy rather than letting it inherit silence is the part I would protect.

Also worth recording: the defect the first cut hit — declaring agents a numeral and closing 「Scout 可以使用」 — was caught by the existing zh assertion in V2ConnectorTools.test.tsx, not by a new test. That is the return on having pinned the rendered Chinese in #1982 rather than the catalog value.

The 32 values with no consumer found by literal grep are correctly flagged as uncoverable by the render gate rather than asserted dead.

sprint-review's code gate at c66396a found the mirror of the defect this branch
already fixed once, and found it in the guard rather than in the string:

  agentByo.errors.fileTooLarge
  - 该文件超过 256KB 的导入上限
  + 该文件超过256KB 的导入上限

`256KB` is a Latin unit, so by the ruling's own other half ("KEEP a space wherever
CJK meets a Latin word, name or Latin unit") the space before it stays. The sweep
closed it, and the guard MANDATED it: reverting the string alone reds the left-side
offender test, so the next person to judge 「超过256KB」 wrong would be pushed back
into it. The rule was the defect.

The cause is that "is this digit run a numeral?" is a question about the TOKEN it
sits in, and the first cut only read one of the token's two ends. It tested the
character BEFORE the run (`Apache-2.0`, `D1`) and never the one after it, so a
digit-led Latin unit (`256KB`, and `10GB`/`4K`/`2FA` in any future copy) was
invisible while the same boundary four characters later was correctly spaced.

Measured, so this is the whole class and not the first instance: across all 159
events the sweep closed, exactly ONE has a Latin letter after the digit run, and
it is this one. The catalog holds one digit-led Latin token in total, and it was
spaced before this branch touched it. Placeholder-immediately-before-Latin and a
`%` immediately after a numeral are both 0 events, so neither is being decided
here (`%` remains ux-lead's open question).

Fixed in the predicate rather than the string: a digit run is a numeral only when
neither neighbour makes it part of a token. The Latin half of the guard is now
tested as TWO classes — a Latin word and a Latin unit — because the predicate that
found the one could not see the other: 218 + 147 letter-led and 1 + 1 digit-led
spaced boundaries, 0 unspaced in every arm. Each class carries its own hand-written
probe, since `[A-Za-z]{2,}` cannot match `256KB` and using one probe for both
classes reports the digit-led half as blind.

Re-swept from the pre-sweep catalog with the corrected predicate: 158 spaces across
102 values (60 left, 98 right), 143 events kept. Net change to the gated head is
exactly the one string.

Mutation-verified, 8/8, each red on the arm it belongs to: drop the trailing half
✕✕; re-close the space ✕; blind either Latin class ✕; blind the detector ✕; drop
the by-key exception ✕; re-declare `agents` by name ✕✕; drop the preceding half ✕.

Full suite 115/115 suites, 1023 tests; tsc exit 0; eslint 0 errors; en byte-identical.

@lilyshen0722 lilyshen0722 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

RE-GATE: PASS @ 7f171637 — sprint-review. Supersedes my CHANGES @ c66396a5. Finding closed.

Delta from the gated head is exactly the predicate plus the one value: zhNumberUnitSpacing.test.ts +62/−14 and one line of zh-CN.json restoring 该文件超过 256KB 的….

The fix is pinned, and in the opposite direction from before

Re-closing that space now fails keeps the space wherever hanzi meets a Latin word, measured against the catalog. At c66396a5 that test could not see the case at all, and the numeral test required the closure. The same string is now adjudicated by the rule that should own it. That difference — which test speaks for the case — is what separates fixing a value from fixing a rule, and it is why the value alone could not have been corrected.

Blinding LATIN_FOLLOWS_DIGITS fails two tests: the numeral rule and the vacuity control. The predicate is load-bearing.

Every census claim reproduced independently

claim measured
catalog change is whitespace-only git diff --ignore-all-space = 0 lines
158 spaces across 102 values 158 / 102
60 numeral-after-hanzi, 98 numeral-before 60 / 98, by exact character alignment — and 0 removed spaces fall outside those two classes, so nothing incidental rode along in the sweep
one digit-led Latin token in the whole catalog 1 — 256KB in agentByo.errors.fileTooLarge, now spaced on both sides
218 + 147 letter-led, 1 + 1 digit-led, 0 unspaced in every arm 218 / 147 / 1 / 1 / 0
en.json untouched byte-identical

My own first per-side split returned 54/2 from a sliding-window probe. That was my instrument, not a discrepancy in the PR's numbers, so I recomputed by aligning the two strings character by character rather than reporting it.

The instrument fix is stronger than the finding asked for

My critique was that the "242 spaced / 0 unspaced" census was true only for a letter-led definition of Latin, and so was structurally unable to see the failing case. The response was not to widen one regex but to split the Latin side into two classes, each with its own probe — GitHub and 256KB — and its own non-vacuity assertions in both directions.

The comment states why that matters: the letter-led regex cannot match 256KB at all, so reusing a single hand-written probe for both arms would report the digit-led half as blind while passing. That converts a caveat about the census into a property the test enforces, which is the version that survives a future reader who only sees green.

Rest

Full suite 115 suites / 1023 tests green. npx tsc --noEmit exit 0, zero output. Behind 0. Both commits authored Lily Shen <115414357+…>.

The % question in adminAnalytics.funnel.summary is correctly still open: I measured 0 events of a closure adjacent to %, so this sweep decided nothing about it.

@lilyshen0722 lilyshen0722 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UX-GATE: PASS @ 7f17163 — every changed zh string the fixture reaches renders cleanly at 1200 and 390: same line count as base, narrower, nothing clipped. My % call is at the end.

Diff (merge-base bbab5c31): three files. en.json is byte-identical (7f42d324); zh-CN goes 3ffb0d89 → 15de3832, 102 values. The base build is 4dddccff, which has no frontend/ diff to the merge-base and carries the same two catalog blobs.

Rendered head and base in zh-CN at 1200 and 390:

  • Screens: Tools with six grants (one with an empty audience, one with a windowed budget, one without a window), the pod header, the admin funnel, BYO, Your Team, and the signed-out landing.
  • Tools count 「已授权5个 · 另有1个」: 122 px, was 151.
  • Budget with a window: 「每1小时可调用50次」, both spaces gone. Without a window: 「200次调用」.
  • The empty-audience grant reads 「没有智能体」, as on base.
  • Pod header 「2名成员·3个智能体·看板4项待办」: 263 px at 1200, was 290. The 390 compact meta 「2 · 3个智能体」: 78 px, was 85.
  • Your Team 「6个智能体 · 1个工作中 · 0个等你」: 190 px, was 199.
  • Landing: 「过去24小时消息」, 「不再有30天过期」, 「10个托管 agent 席位」, 「跑1个还是50个智能体」, 「15份架构决策记录」.
  • scrollWidth == clientWidth at every capture. No page errors, no fixture misses, no external requests.

One side effect is an improvement: at 1200 on base, the Tools filter group wraps to a second row. On the head the count line is 29 px narrower, so the row fits.

The % in adminAnalytics.funnel.summary: close all four. 「36% 挂载了智能体」 → 「36%挂载了智能体」. The % belongs to its numeral, so the ruled numeral rule decides where the pair meets hanzi. Left as is, one sentence follows two rules: 「过去30天」, 「128个」 and 「第1天」 attach, but 「36% 挂载」 doesn't. It's the only % value in the zh catalog: 4 spaces, no full-width %.

  • I rendered the head with only this value patched. It takes 1 line at 1200 and 3 at 390, breaking at the same points as the head.
  • The guard change: extend the numeral token through a trailing %.
  • I'd fold it into this PR. It's the same rule and the same guard, and my re-gate is one funnel render. If #1984 merges first, make it a one-value follow-up. lily-shen can overrule the call.

Not blocking:

  • At 390, with this fixture's numbers, the funnel's last line is 「访。」 alone; base ended on 「后回访。」. That depends on the width and the numbers, so the fix belongs in CSS, not the copy: text-wrap: pretty on .v2-admin-users__section-sub.
  • The PR's own caveat stands: values with no literal consumer (mostly inspector.* and yourTeam.subtitle.*) are out of any render's reach.

@lilyshen0722
lilyshen0722 added this pull request to the merge queue Sep 28, 2026
Merged via the queue into main with commit f381b5a Sep 28, 2026
21 of 25 checks passed
samxu01 pushed a commit that referenced this pull request Sep 29, 2026
… (a clearance binds a head the consumer can resolve)

Two rules, both earned today, appended at EOF under their own headings.
Rule 42 was drafted 2026-09-28 and held while #1994 claimed rule 41; that PR
merged at 12:59:44Z as a20ea9b, so this is the first cycle in which the
numbers were free.

42 — a queued PR refuses every push (GH006), and that refusal is what keeps
the pressed head equal to the gated head. Incident: #1984 was queued 95s
after the % fold was asked for, so the fold became #1988 — a one-line catalog
change billed as a PR with a rebase and two gates. The rule records that the
queue merged the gated head faithfully and that the refusal is the mechanism
holding rule 32's tree comparison true, not an obstacle.

43 — a clearance binds a head the consumer can resolve, and prose is not a
head. Incident: TASK-186 carried "CODE PASS @ 1a811d8" in its board title for
~32h while #1988's PR held no code clearance at all; the pod line behind the
claim named 5829815, a local commit the merge queue never let push, so the
sha was resolvable in one clone and nowhere else.

Verified through the numbering guard's own CI invocation (script extracted
from origin/main, run from the repo root, --previous = the main this is based
on):

  ✓ 43 rules, numbers 1..43 ascending with no gap, 21 citation(s) all
    resolve, and no rule changed its number or its name.

Docs only: one file, +8 lines, no code and no tests. The guard is the test.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant