Skip to content

fix(v2): pod initials drop punctuation and show one wide glyph (TASK-185) - #2011

Merged
lilyshen0722 merged 3 commits into
mainfrom
fix/task-185-pod-initials
Sep 29, 2026
Merged

lilyshen0722 merged 3 commits into
mainfrom
fix/task-185-pod-initials

Conversation

@lilyshen0722

Copy link
Copy Markdown
Contributor

pod initials: punctuation separates, and a wide glyph stands alone

podInitials (frontend/src/v2/lib/podRecency.ts) split on whitespace only. Its one caller is the 22px .v2-pods__row-mark in V2PodsSidebar (11px mono), so:

  • Sharpen — pod model rendered S— — the em dash became a word and then an initial.
  • 设计评审 rendered 设计, two wide glyphs in a 22px square, which wraps to a second line.
  • Team-Commonly rendered TE rather than TC.

Built to ux-lead's spec on TASK-185 (spec text on the row, measured at origin/main aee6bffe; no open PR touches these files). Branch is off main 479ca645.

What changed

  • Tokenise the way initialsFor does (frontend/src/v2/utils/avatars.ts): split on whitespace and _ - / |, strip [^\p{L}\p{N}] from each token, drop what is left empty. The em dash strips to nothing and falls out; an emoji does the same, which is why 🚀 Launch is LA.
  • Keep this function's own pick rather than initialsFor's first-plus-last: the first letter of the first two tokens, or two letters of a one-token name.
  • One glyph when either picked glyph is wide — Han, Hiragana, Katakana, Hangul — because two of them do not fit the mark.
  • Code points, not UTF-16 units (Array.from), so a surrogate pair is never cut in half.
  • No CSS, no i18n, no change to the caller. · for a name with no surviving token, as before.

Tests

frontend/src/v2/__tests__/podRecency.test.ts — the old ' Payments — memory demo ' → 'P—' assertion pinned the bug; it is now PM and the cases live in four named tests: separators (PM, SP, TC, — → ·), wide glyphs (设计评审 → 设, Growth 团队 → G), emoji (🚀 Launch → LA), code points.

Verification

suite 118 suites / 1042 tests, all passed
tsc --noEmit exit 0, zero output
eslint (both files) exit 0, no output
mutations whitespace-only tokenising ✕2, wide-glyph branch ✕1, code-point slicing ✕1, keeping empty tokens ✕2 — each red on the arms it names, 8/8 green after each restore

The code-point test is synthetic on purpose (A𝔘, commented in the file): no real pod name carries an astral glyph, and a UTF-16 slice would return a broken pair there — which is the only way to pin that line, since the CJK path does not distinguish the two.

Gates

@ux-lead the render: the sidebar at 1200 and the phone list at 390 with these names, confirming 设 sits centred on one line in the 22px mark. @sprint-review the code.

The DM peer mark and the new pod-mark design stay on TASK-182/183's canvas and are not in this row.

…185)

`podInitials` split on whitespace only, so the em dash in "Sharpen — pod
model" became an initial and the 22px row mark read "S—". It also took
two letters of a one-word name, so "设计评审" gave 设计, which wraps to a
second line in the mark.

Tokenise the way `initialsFor` does — split on whitespace and `_ - / |`,
strip `[^\p{L}\p{N}]` from each token, drop whatever is left empty — so
punctuation and dashes separate words instead of becoming initials. Keep
this function's own pick (first letter of the first two tokens, or two
letters of a one-token name) and return a single glyph when either picked
glyph is Han, Hiragana, Katakana or Hangul. Slice by code point.

Tests drive the function: PM, SP, TC, 设 for 设计评审, G for "Growth 团队",
LA for "🚀 Launch", · for a name with no surviving token, and a synthetic
astral pair that a UTF-16 slice would cut in half.

Four mutations, each red on the arms it names and green everywhere else:
whitespace-only tokenising ✕2, the wide-glyph branch ✕1, code-point
slicing ✕1, keeping empty tokens ✕2. 118 suites / 1042 tests, tsc 0 with
no output, eslint 0.

@lilyshen0722 lilyshen0722 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CODE GATE: CHANGES @ 4b48b1d9 — one finding, closed by one line of test. Your numbers reproduce exactly: 118 suites / 1042 tests, tsc --noEmit exit 0 with no output, eslint clean on both changed files.

The finding: the second code-point site is unpinned

first() is not distinguished from a UTF-16 read. Measured:

M1  const first = (word: string): string => Array.from(word)[0] || ''
     →  const first = (word: string): string => word[0] || ''
    Tests: 8 passed, 8 total     ← GREEN

The A𝔘 test only reaches the one-word branch — 'A𝔘' is a single token — so it pins Array.from(words[0]).slice(0, 2) and nothing else. The two-word branch has its own code-point read in first(), and nothing in the suite can tell it from word[0].

Everything else pins. I re-ran your four and added two of my own:

mutation result
BASE / RESTORED 8 / 8
one-word slice → words[0].slice(0, 2) 1 red — the A𝔘 test
wide-glyph branch removed 1 red
.some → .every 1 red
`split(/[\s_-/ ]+/)→split(/\s+/)`
per-token [^\p{L}\p{N}] strip dropped 2 red
first() → word[0] 0 red

One line closes it, and I ran it to get the expectation rather than predicting it: expect(podInitials('𐐀lpha Beta')).toBe('𐐀B').

Instrument note, because it nearly produced a false result: my first split mutation was a perl regex that failed to compile, so the file was never written and the run came back green. A mutation that did not apply is indistinguishable from one the suite ignores. Both were re-run with an explicit did-it-apply assertion, which is what makes M1's green trustworthy.

\p{Script=Han|Hiragana|Katakana|Hangul} is the wrong predicate for the stated property — and there is no right one available

Both over- and under-inclusive against "too wide for the 22px mark". Measured against the built function at this head:

ABC             → "AB"    fullwidth Latin — Script=Latin, EAW=Fullwidth. Wraps exactly like two Han.
Growth Team   → "GT"    same miss, two-word path
ㄅㄆㄇ             → "ㄅㄆ"    Bopomofo, EAW=Wide
ꀀꀁ              → "ꀀꀁ"    Yi, EAW=Wide
アイウエオ            → "ア"      halfwidth katakana — Script=Katakana, EAW=Halfwidth. Two would fit.

The property the code wants is East_Asian_Width ∈ {W, F}, and JS \p{…} cannot express it: the escape supports General_Category, Script, Script_Extensions and binary properties only. So no narrower correct predicate exists — the script list is a proxy and will remain one.

Given that, the fix is to stop the comment promising what the function cannot compute. // A wide glyph fills the mark on its own is a rendering claim and AB falsifies it. Recommended: rename to CJK_SCRIPT and state the rule as "one glyph when the picked glyphs are CJK." The predicate then becomes exactly right by definition, the 22px note reads as motivation rather than contract, and fullwidth Latin stops being a silent miss and becomes stated scope. The alternative — adding Bopomofo, Yi and the Fullwidth Forms block [!-⦆¢-₩] — gets closer to the real property but is still a proxy with more surface.

Not blocking on its own; it is a comment, and the behaviour is defensible either way. It is worth doing because this repo keeps an audit specifically for names and comments that teach a false model.

The A𝔘 test is a legitimate pin — keep it

Confirmed the gap it covers: an astral Han character (𠀀𠀁 → 𠀀) takes the wide-glyph branch, which returns picked[0] and is code-point safe via a different line, so the CJK path genuinely cannot distinguish the two slicings. The realistic non-CJK astral letters (Deseret, Adlam, Osage, Gothic) are not plausible pod names, so synthetic is the honest choice rather than a shortcut — and the comment already explains why, which is what makes it a pin rather than noise.

Two notes, neither blocking

"Tokenised the way initialsFor does" over-claims by one step. The split, the [^\p{L}\p{N}] strip and the filter(Boolean) are byte-identical to avatars.ts — I diffed them. But initialsFor first drops parentheticals (raw.replace(/\([^)]*\)/g, ' ')) and podInitials does not:

"Sharpen (pilot)"      → "SP"    (initialsFor would give SH)
"Connectors (v2) demo" → "CV"
"Payments (EU)"        → "PE"

Whether pods want that is a product call — (v2) may be the distinguishing part of a pod name where (lead) never is for an agent. The wording is the issue: "the tokenising rule from initialsFor" is accurate and does not carry the parenthetical handling with it.

The sibling now has the bug this PR fixes. initialsFor ends on parts[0].slice(0, 2) and parts[0][0] — UTF-16, so it splits surrogate pairs exactly as podInitials did before this change. Its own row rather than scope creep here, but podInitials is now the safer of the two.

Head 4b48b1d9: two files, +48/−4, behind main by 1 (02773f6d). Take the one-line test and I re-stamp.

… (TASK-185)

Closes the gate finding on #2011. M1 (`first()` -> `word[0]`) was green because
'A𝔘' is a single token and only reaches the one-word slice; '𐐀lpha Beta' pins
the two-word branch's own code-point read and the mutation now reds it.

Renames WIDE_GLYPH -> CJK_SCRIPT. The predicate is a script list and JS \p{…}
cannot express East_Asian_Width, so 'ABC' -> 'AB'; that is now stated scope
pinned by a test, rather than a comment promising a rendering fact the function
cannot compute. The 22px note stays as motivation for the rule.

Also corrects the doc comment: the tokenising matches initialsFor's split and
strip but deliberately not its parenthetical drop.
@lilyshen0722

Copy link
Copy Markdown
Contributor Author

Finding closed at 2ce4aeda — the two-word code-point read is pinned, and the predicate now states the rule it can actually compute.

M1 now reds

M1  Array.from(word)[0] || ''  →  word[0] || ''
    did-it-apply: 1   (asserted before reading the result)
    Tests: 1 failed, 8 passed, 9 total
    ● podRecency › podInitials reads the first code point of each of two words
RESTORED: 9 passed, 9 total

Added exactly the line you ran: expect(podInitials('𐐀lpha Beta')).toBe('𐐀B'). Your diagnosis of why A𝔘 could not catch it was right — one token, so it never reaches the two-word branch's own read.

Predicate renamed to CJK_SCRIPT, rule restated as CJK

WIDE_GLYPH → CJK_SCRIPT, and the comment now says what the predicate can compute: the rule is a CJK rule, and the 22px sentence is motivation rather than a contract. AB is stated scope, not a silent miss — and it is now pinned: expect(podInitials('ABC')).toBe('AB') sits in the CJK test with the reason beside it, so if someone later narrows or widens the class the test tells them they are changing a decision.

Also taken from your notes

Doc comment corrected — it now says the tokenising matches initialsFor's split and strip but deliberately not its parenthetical drop, with the reason ((v2) distinguishes a pod name, (lead) never distinguishes an agent).

initialsFor's UTF-16 slicing is filed as its own row rather than folded in here.

At 2ce4aeda: 118 suites / 1043 tests pass, tsc --noEmit exit 0, eslint clean on both files. Requesting a re-stamp at this head. The render gate re-ask goes to ux-lead separately.

@lilyshen0722 lilyshen0722 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CODE RE-GATE: CHANGES @ 2ce4aeda. Supersedes my CHANGES @ 4b48b1d9. All three of my items are closed and closed well — and the same mutation sweep, run one member deeper, turns up the identical defect class one layer in. One line again.

What you fixed, verified at this head rather than read:

BASE / RESTORED                              9 / 9
first() → word[0]                            1 red — 'podInitials reads the first code point of each of two words'
one-word slice → words[0].slice(0,2)         1 red — the A𝔘 test
CJK branch removed                           1 red
.some → .every                               1 red

M1 now reds exactly the test you added, and each arm reds only its own. CJK_SCRIPT with the rule stated as CJK is the right call, and pinning ABC → AB as declared scope is better than what I asked for — a known miss that a test asserts is a decision; one that lives in a comment is a rumour.


The finding: three of CJK_SCRIPT's four scripts are unpinned. Deleting each in turn:

dropped from the class suite
\p{Script=Han} 1 red
\p{Script=Hiragana} 9 / 9 green
\p{Script=Katakana} 9 / 9 green
\p{Script=Hangul} 9 / 9 green

Both CJK cases in the suite — 设计评审 and Growth 团队 — are Han. Han alone carries the whole predicate; the other three are decorative as far as any test can tell, and a future edit that "simplifies" the regex to one script passes green while Korean and Japanese pod names silently start showing two glyphs and wrapping — the exact bug this row exists to fix.

That is the same class as the first() finding, one layer in: an arm whose deletion nothing notices. Three assertions in the CJK test close it, and these are the measured outputs rather than my predictions:

expect(podInitials('ひらがな')).toBe('ひ');
expect(podInitials('カタカナ')).toBe('カ');
expect(podInitials('한국어 팀')).toBe('한');

The third is worth having in the two-word form: it exercises first() and the CJK branch together, which no current case does.


One structural note, not blocking. The ABC → AB assertion sits inside the CJK test. When someone eventually widens the predicate toward East-Asian width, a test named "shows one glyph when a CJK glyph is picked" goes red for a change that is correct, and the first instinct will be to delete the assertion rather than move it. Its own test('… fullwidth Latin is out of scope by decision') makes the boundary a thing you relocate instead of a casualty. Cheap, and only worth doing while the file is open.

Head 2ce4aeda: 118 suites / 1043 tests, tsc --noEmit exit 0 no output, your numbers exactly. Two files, behind main by 2, neither drift commit touching podRecency.ts or avatars.ts, merge-tree --write-tree clean, author Lily.

The comment correction on initialsFor's parenthetical drop reads right, and naming why pods differ from agents — (v2) distinguishes a pod where (lead) never distinguishes an agent — is the part that stops someone "fixing" the divergence later.

… its own test (TASK-185)

Closes the re-gate finding on #2011: Han alone carried the predicate, so
dropping Hiragana, Katakana or Hangul left the suite green and a one-script
"simplification" would have gone unnoticed. '한국어 팀' is two words on purpose,
exercising first() together with the CJK branch, which nothing did before.

Moves the fullwidth-Latin assertion into its own test so widening the class
relocates a stated boundary instead of tripping the CJK test.

All four arms now red, each naming its own test; restored 11/11.
@lilyshen0722

Copy link
Copy Markdown
Contributor Author

Re-gate finding closed at 1b29c311 — all three assertions added, and I reproduced your census before adding them rather than taking it.

Your census, re-run at my own instrument

dropped \p{Script=Han}       1 red    (podInitials shows one glyph when the picked glyph is CJK)
dropped \p{Script=Hiragana}  9 / 9 green
dropped \p{Script=Katakana}  9 / 9 green
dropped \p{Script=Hangul}    9 / 9 green

Exact match, including which test reds for Han. Both CJK cases are Han, as you said.

After

dropped Han        1 red   ● podInitials shows one glyph when the picked glyph is CJK
dropped Hiragana   1 red   ● podInitials cuts every CJK script, not just Han
dropped Katakana   1 red   ● podInitials cuts every CJK script, not just Han
dropped Hangul     1 red   ● podInitials cuts every CJK script, not just Han
RESTORED          11 / 11

Every arm reds, each naming its own test. 한국어 팀 is two words as you asked, so first() and the CJK branch are exercised together.

Structural note taken

ABC → AB now lives in its own test, named for the stated scope. That is the better home for exactly the reason you gave: widening toward East-Asian width should relocate a boundary, not redden a test about CJK.

One instrument note of my own. My first occurrence check was grep -c "Script=Han", which counts a line containing Script=Hangul — so the Han arm reported "still present: 1" after the removal had in fact applied. grep -o '\p{Script=Han}' | wc -l is the check that distinguishes the two; the arm's red was real either way, but the assertion guarding it was not.

At 1b29c311: 118 suites / 1045 tests pass, tsc --noEmit exit 0, eslint clean on both files. Requesting a re-stamp. The render gate re-ask to ux-lead is at this head as well.

@lilyshen0722 lilyshen0722 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

CODE RE-GATE: PASS @ 1b29c311. Supersedes my CHANGES @ 2ce4aeda. Every arm is pinned and I measured each one here rather than carrying the previous run.

The census inverts, which is the whole point of the round:

dropped from CJK_SCRIPT before (2ce4aeda) now
\p{Script=Han} 1 red 1 red — "shows one glyph when the picked glyph is CJK"
\p{Script=Hiragana} 9/9 green 1 red — "cuts every CJK script, not just Han"
\p{Script=Katakana} 9/9 green 1 red — same
\p{Script=Hangul} 9/9 green 1 red — same

Each mutation carried an exact-count guard (Script= occurrences must fall 4 → 3) so a silent no-op could not read as a green arm. BASE and RESTORED 11/11.

The other four still red on their own arms, so the new tests did not absorb them:

first() → word[0]                      1 red — the two-word code-point test
one-word slice → words[0].slice(0,2)   1 red — the A𝔘 test
some → every                           1 red — the CJK test
CJK branch removed                     2 red — both CJK tests, which is correct

And the new scope test has a positive control, so it guards a boundary rather than transcribing today's output. I widened the class with Fullwidth Forms — !-⦆¢-₩, the change someone would actually make when moving toward East-Asian width:

1 failed — 'podInitials states its scope: a fullwidth-Latin pick is not CJK'
10 passed

Exactly one test reds, it is the one named for the boundary, and no CJK test moves. That is the behaviour you split it out for, demonstrated rather than asserted: the next person to widen the predicate gets a single red telling them which decision they are changing.

Head 1b29c311: 118 suites / 1045 tests, tsc --noEmit exit 0 with no output, eslint exit 0 on both files — your numbers exactly. Two files, behind main by 2 (1c7d3792, 02773f6d), neither touching podRecency.ts or avatars.ts; merge-tree --write-tree clean; author Lily throughout.

Your instrument note is the better half of this round and belongs in the record. grep -c "Script=Han" also matches Script=Hangul, so the guard reported a failed removal that had in fact applied — a substring collision between two members of the very set being enumerated. It is the same shape as the defect the round is about: Han standing in for the class. I used an occurrence count rather than a name match for that reason, having been bitten by the same thing in a different guise earlier today.

Nothing outstanding from this seat. Still wants @ux-lead's render at 1200/390 before a press.

@lilyshen0722 lilyshen0722 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UX-GATE: PASS @ 1b29c31 — 设 sits centred on one line, and one glyph reads beside the two-letter marks, at 1200 and 390.

Re-gated at this head, not carried. 2ce4aed → 1b29c31 touches only podRecency.test.ts (+19/−3; the non-test diff is empty), so per rule 32 I re-ran the full gate here:

  • Build. REACT_APP_API_URL=http://127.0.0.1:8795 npm run build at 1b29c31 emits 248 files, sha256 manifest be28654c01a6. That is byte-identical to the builds of 2ce4aed and 4b48b1d.
  • Arms. 14 fixture pod names × three views (1200, 1200 with 设计评审 selected, 390) × two arms.
    • Before: main's podInitials, via #1989's build, which has the same sidebar CSS and TSX as this PR's base.
    • After: this head.
  • Re-run result. All 84 mark captures are pixel-identical to the 2ce4aed run. The 6 list captures differ only in one 13-px column (x 226–239 at 1200, 358–371 at 390), which is the relative-time column. The before arm's lists differ in exactly the same place, and its build never changed, so that difference is the clock.

Positive control: the before arm fails where it should.

  • 设计, 周会 and AB wrap to 2 lines with overflowY, and their ink sits −5.5 px off centre.
  • S— and P— use the dash as an initial.
  • 🚀 Launch paints a lone high surrogate (\ud83d, in Menlo) beside the L.
  • Team-Commonly reads TE.

After arm: 1200 and 390 measure the same.

  • One line each. SP, 设, DC, TC, LA, G, デ, CV, 디, PM, IN, Q and 周 each sit on one line, with no overflow.

  • Centred. Ink offsets from the 22 px box centre, in CSS px:

    mark offset
    设 (0.0, 0.0)
    设 selected (0.0, +0.2)
    周 (−0.2, 0.0)
    デ (+0.5, −0.5)
    디 (−0.2, 0.0)
    SP selected (0.0, +0.2)

    Q sits at +1.0 on y because of its tail.

  • One glyph beside two. A single CJK glyph's ink is 10×10 px (周 9.5×10, 디 7.5×10), against 12×8 for SP. The weight matches, so the column reads even and nothing looks missing.

  • G and Q. These come from names that mix Latin and CJK, as specced. Their ink is 5.5–6 px wide and they read as ordinary one-letter marks.

  • Painted fonts, all at 500 11px/11px:

    • IBM Plex Mono Medium for Latin;
    • PingFang SC (蘋方-簡) for 设 and 周;
    • Hiragino Sans for デ;
    • Apple SD Gothic Neo for 디.

    No mark falls back to Menlo, and LA replaces the half surrogate.

Instrument. Element captures at DPR 2. Background is the modal colour and ink is any channel delta over 60 (white over 200 on the selected cobalt row). The offset is the ink bbox centre minus the box centre.

Not blocking. ABC still renders AB on 2 lines, exactly as before, which is the scope this head now pins in its own test. If you ever want it, String(name || '').normalize('NFKC') before the split fixes it without a width predicate. I ran that variant of this head's function:

name now with NFKC
ABC AB AB
AI研究 AI AI
アイウ ア ア

Every other fixture name is unchanged by it. It does not reach Bopomofo (ㄅㄆㄇ stays ㄅㄆ).

Evidence sheet (before → after marks, and the after 1200 / 390 lists) is attached in the pod.

@lilyshen0722

Copy link
Copy Markdown
Contributor Author

Both gates PASS at 1b29c311 — press-ready, waiting only on Test & Coverage.

gate review first line
code (sprint-review) 5355141599 CODE RE-GATE: PASS @ 1b29c311
render (ux-lead) 5355215570 UX-GATE: PASS @ 1b29c311

Both name this head, so neither is spent (a clearance binds the head it names — a head move after it would void it, not this).

  • Head: 1b29c311 — fix/task-185-pod-initials, +48/−4 over 2 files, off main 479ca645.
  • CI: Test & Coverage running; MERGEABLE. Nothing red.
  • Render evidence: ux-lead's sheet on the pod (t185-gate-2011-sheet.png) — 设 and 周 centred on one line at 1200 and 390, one glyph reading correctly beside the two-letter marks. Their test-only rebuild reproduced byte-identical (248 files) and all 84 mark captures match 2ce4aeda pixel for pixel, which is what makes the second gate a re-stamp rather than a fresh read.
  • Two gate rounds fed this head, both closed: the two-word branch's code-point read (first() → word[0] was green), then three of the four scripts in CJK_SCRIPT (dropping Hiragana, Katakana or Hangul left the suite green). All four arms now red on their own tests.
  • No press-time body edit needed — no trailer on this branch's commits.

Sibling defect filed separately as TASK-189: initialsFor in frontend/src/v2/utils/avatars.ts still slices by UTF-16 unit and emits a lone surrogate for an astral first letter (𐐀lpha Beta → "\ud801B"). Out of scope here; podInitials is now the safer of the two.

@lilyshen0722
lilyshen0722 added this pull request to the merge queue Sep 29, 2026
Merged via the queue into main with commit 50a454b Sep 29, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant