Skip to content

test(mmc-trajectory): ship the regression guards that were only on one machine - #16

Open
avatasia wants to merge 15 commits into
MiniMax-AI:mainfrom
avatasia:add-mmc-trajectory-tests
Open

avatasia wants to merge 15 commits into
MiniMax-AI:mainfrom
avatasia:add-mmc-trajectory-tests

Conversation

@avatasia

@avatasia avatasia commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

会话轨迹 — 让每一轮轮询不再重建整场会话

本 PR 在 ab717e3(检视面板跟随筛选 + sticky 重算收敛到单一出口)之上,继续处理 L0 数据层。

0. 上一轮的状态

  • Stage 1(把 L2 在文档里收成连续的)已否决,理由见 D1:查询条件是聚合,UI 与请求构造本来就该拆开。
  • D8 / D9(检视面板无响应、sticky 偏移重算只挂一根线)已修,docs/LAYERS.md 有记录。
  • 守卫 26 条,变异 34 条全捕。

1. 问题:轮询重建整场会话

一个已结束的会话不花钱——ETag 由 stat 事实派生,会话没变就是 304,什么都不重建。

花钱的是还在增长的会话,而那恰恰是读者主动盯着看的场景:每一轮轮询都从头重建一遍。

本机实测(23 代会话,payload.session.sizeBytes = 99,986,501):

23 个 generation 文件里 22 个是 snapshot 运行时轮转代时写一次,之后在压缩之间完全不动
只有 messages.jsonl 在追加 97,526 字节 = 0.10%
trajectoryCacheKey() 把每代 byteLength 拼进键 会话一增长 → 整场缓存 miss → 重读重解析全部 95 MB
其中与上一轮逐字节相同 99.90%
线上字节(5 分钟) 1.31 GB 换 118,875 字节真实新增 = 11,006×,99.99% 重发

三层各自独立地表现出同一个比例:服务端输入 99.90% 冗余、线上字节 99.99% 冗余、页面解析 ~99.99% 冗余。会话在忙碌时每 3–6 秒变一次,轮询每 2.5 秒问一次——几乎每次轮询都撞上全量重发。


2. 修法一:按代缓存解析结果

loadParsedRows() 用 path + size + mtime + limit 做键,返回任何没有变动的 generation 文件的行。

真实文件上直接测(22 个冻结 snapshot,95.3 MB):

冷读(读 + 解析全部)   306 ms
命中缓存                 2 ms
省下                  303 ms  (99%)
解析吞吐              311 MB/s

按字节预算而不是条数——条数与体积无关,一个 generation 文件可能 40 KB 也可能 10 MB。

缓存键的每个成分都对应一个它要防的失败:

  • 用 filePath 不用 fileName —— 每个会话都有一个 messages.jsonl,basename 键会让两个会话撞车。夹具专门造了这个碰撞(同尺寸、同 mtime、不同对话),断言互不可见。
  • 用 stat 的 size + mtimeMs,不用 catalog 里的 byteLength —— 代码自己在 server.mjs:1889-1894 注释了那个值是会过期的快照,而 active 文件每轮都在它之外继续增长。
  • 含 limit —— readHead 按它截断,而"是否被截断"是这条目持有的事实之一。
  • stat 与 readHead 之间的竞态是安全的 —— 条目会键在一个它并不匹配的大小上,下一轮轮询重新 stat → size 变了 → 键不同 → 重读。只会多一次读,永远不会供出陈旧解析。

3. 修法二:原文改成按需构建

buildRawPayload() 对每一条记录执行 JSON.stringify(message) —— 18,337 次、约 100 MB —— 只为回答两个问题:序列化后是否超过 MAX_RAW_BYTES(65536)、是否含 .minimax 需要脱敏。那个字符串用完就扔:splitRawRecords() 在最终序列化之前把消息摘掉,而 /raw 带走的是消息对象,不是这个字符串。

于是每轮重建都要把 100 MB 消息体完整序列化一遍、量一下、再扔。而 raw 只有打开某条记录的原文页、复制、交接时才读得到。

离线复测(同语料、同代码路径):

之前 之后
buildTrajectoryPayload 142 ms 48 ms
折算到真实 95.4 MB lineage 214 ms 72 ms

记录现在只带一个定位符(它来自 rows 的第几条),/raw 被请求时才调 buildRawPayload()——一次一条。


4. 口径变更(契约级)

「原文过大,已省略」从 bulk payload 迁到 /raw 响应。

/api/trajectory/raw?index=N 现在返回 { index, raw, omitted }。

按本项目规矩逐个列读点,并在 tests/verify-raw-lazy.mjs 里逐个加断言:

# 位置 现在读 改后读
1 index.html:2867 交接元信息 record.rawOmitted rawText === null(交接本来就先取 rawSourceFor,copySelected 同样)
2 index.html:5989 rawSourceFor record.rawOmitted → 短路 null 无条件取 /raw,省略由服务端给出
3 index.html:5998-5999 原文页提示 record.rawOmitted /raw 的 omitted
4 index.html:6007-6008 原文页内容 record.rawOmitted → — 同上
5 index.html:6059 copySelected source === null → stringify(record) 不变

五个读点没有一个在台账行上,全部在读者打开某条记录之后才发生。

读者可见的变化只有一处:原文页对「过大」的记录会多发一次 /raw 请求(之前零次)。这是一条用户主动打开才走的路径,用一次请求换每轮少序列化 100 MB。


5. 守卫 26 → 30,变异 34 → 58,全捕

verify-parsed-rows.mjs   53 断言
mutate-parsed-rows.py    12 变异

verify-raw-lazy.mjs      39 断言
mutate-raw-lazy.py       15 变异

verify-parsed-rows.mjs 的核心:跑两次真实路径(冷读 / 命中),两边都 JSON.stringify 成 body,要求逐字节相同,不是"看起来一样"。读错来源或键错字段的版本在这里会算出错误答案而不是抛异常。

verify-raw-lazy.mjs 里的 64 KB 上限没有丢,而是跟着定位符走到它现在该在的地方被断言。

变异替守卫抓到的三个真缺陷(不是补断言凑数):

  • Number(record.rawRow) 把 null 当成第 0 行(Number(null) === 0),/raw 会答出另一条记录的消息。改成拒绝而不是强转。
  • 去掉"把定位符从记录上摘掉"的那行 delete,rawRow 就会随每条记录发到页面上。此前没有任何断言说它不在。
  • mutate-pr15-fixes.py 里有一条变异锚点被本次改动冲掉了。它报的是 SKIP 不是 pass,但死掉的变异看起来像绿的,已重新锚定到承载同一判定的行。

6. 对 D10 的更正(这条我自己写错了)

D10(按代缓存)记的是「读+解析 ≈ 1,000 ms」。这个数是错的,实测 306 ms。

错在哪:那个数是用 1,138ms 重建 − 139ms stringify 减出来的,把没有查过的 buildTrajectoryPayload 和 buildRawPayload 一起算进了「读+解析」里。

所以诚实的账:

D10 里承诺的 实测
按代缓存收益 ~1,000 ms / 3.5× ~300 ms / 约 21%
常驻堆 130–290 MB(外推) 116 MB(实测,95.3 MB 源 × 1.22×)

顺带修了 D10 自己发的代码里一个缺陷:预算代理写成 3×,于是 95.3 MB × 3 = 286 MB 计费压过 256 MB 预算,LRU 把放得下的条目也赶走了。实测保留率 1.22×,代理改为 2×。预算的好坏完全取决于这个数。

端到端轮询中位数没有明显变化(1,405 → 1,434 ms,样本区间 348–2,916 ms)。原因有二,如实说明:300 ms 埋在这个噪声带里,而且浏览器标签一直在开、每 2.5 秒轮询同一个服务端,我的探针和页面在抢同一个事件循环——这两组数都不足以支撑结论。


7. 验证

node tests/run.mjs                          30/30
python tests/mutate-parsed-rows.py          all 12 caught
python tests/mutate-raw-lazy.py             all 15 caught
python tests/mutate-pr15-fixes.py           all 9 caught
npm run validate                            plugins/avatasia/mmc-trajectory: OK

浏览器实测(新运行时 http://127.0.0.1:55344/dashboard):

  • payload 里 raw 字段 0 个、rawOmitted 字段 0 个(18,611 条记录)
  • /raw 抽 40 条:39 条带回原文、1 条报「过大」、0 条说不清的
  • 点行 → 检视面板填出该记录;点「原文」页 → 取回并渲染出真实消息 JSON
  • console 无 error

8. 还没做的

  • 增量协议(?since=,只发新增尾部)。它比这一轮两项加起来更彻底:服务端不重建、页面不解析 36 MB、不传 36 MB,1,400 ms → ~10 ms 量级,而且不需要那 116 MB 常驻。这是这个项目真正该还的设计债,按 AGENTS.md 的规矩要先在 LAYERS.md 里定口径再动手。
  • 「显示全部」后的两百万像素 DOM:state.limit = Infinity 会一次构建 143,568 个 DOM 节点,8 万次点击测得中位143,568 节点路径是 state.limit += PAGE_STEP 的二十几倍,同样实测:这条已修(分批建行 + 进度,D12,见下)。剩下的杠杆是给「显示全部」封顶或走增量协议,两者都改承诺,未动。

9. 本轮新增:D12 分批建行 / D13 检视不再改窗口

两个缺陷,都是同一次会话上复现的,根因都是结构性的,不是少写一行。

A —「显示全部」把两万行塞进一次同步循环

实测(本机,150 轮 / 18,924 条 / 视口 735×900):

分批建行(25 轮/帧)总耗时      3,956 ms
  每批建行 JS   中位 141 ms / 最大 151 ms
  帧间(浏览器对整篇文档的 layout + paint)  ≈ 617 ms/帧

卡的不是建行 JS,是每帧对已经两百万像素高的文档做的 layout。 六批 JS 加起来约 870 ms,其余 3,086 ms 是 5 个帧间隙。所以把批调小只会让帧次变多、对不断变长的文档重复 layout——治不了。

修法不改承诺,只改节奏:一帧最多建 25 个轮组就交还主线程;窗口条在期间报「正在显示 N / M 条」;每轮 renderLedger() 先 ledgerBuildToken += 1,在途分批下次醒来发现 token 变了就地退出、不碰 DOM;轮询 tick 发现 ledgerBuilding() 为真就记住这一拍,由 flushDeferredPoll() 在构建结束后补跑一次——3,956 ms > 2,500 ms,不这么做读者会看着计数一次次归零。

「显示全部」仍然是全部。

实测并否决的替代方案(记在 LAYERS.md D12,因为它是下一个会被问到的问题):content-visibility: auto 是真实的 4.0×——3,956 ms → 979 ms,行数一致,帧间 layout 一并消失。但它需要给被跳过的元素填一个 contain-intrinsic-size,而 150 个轮组的真实高度跨度 11 倍(p25 1,011 / 中位 2,667 / 最大 29,754 px)。填任何单一值都会让滚动高度大幅失真,auto 关键字只能记住已渲染过的那两个。这是拿「控件如实描述页面」换速度,不采用。

B — 检视面板描述着一条已经不在表里的记录

在「显示全部」下选中一条早期记录,点「显示最新 200 条」,记录离开窗口,而面板还在描述它。

revealSelection() 判定在切片之前,且会 state.limit = Math.min(needed, entries.length) 主动把窗口撑大到装下那条选中:

点下「显示最新 200 条」   已显示 200 条
一秒钟之后               已显示 15,224 条

读者自己点的按钮,在点击之后把窗口条上的数字改了。 两处错:窗口成了检视面板能写的东西(state.limit 因此有第二个写手),以及判据错了——「读者能看见什么」是窗口,不是筛选结果。

现在 revealSelection(entries) 是纯谓词:只写 state.selectedId,不再碰 state.limit;调用点挪到切片之后,参数从 matched 改成 windowed。窗口收窄到不含选中记录时取消选中,不撑大窗口——读者的窗口选择优先于更早那次点击,且与筛选轴既有规则一致。

口径变更与读点

tests/verify-selection.mjs 整份重写(14 断言)。这是口径变更不是重构:该文件原本正守着「窗口会被拉到选中所在处」(断言 limit === 8)和「判定发生在切片之前」——两者现在都反了,改为断言新承诺,并新增一条:revealSelection 源码里不允许出现 limit 这个词。

守卫

文件 状态
tests/verify-ledger-batch.mjs 新增,20 断言
tests/mutate-ledger-batch.py 新增,19 变异全捕
tests/verify-selection.mjs 重写,14 断言
tests/mutate-selection.py 9 变异全捕,其中 2 条是改写不是删除
tests/mutate-sticky.py / mutate-turn-breakdown.py / client-check.mjs 重新锚定并收紧

两条变异值得单独说:原来的「窗口不会被拉到选中」守的正是这次删掉的那行 state.limit 写入,它现在改为弄坏替换它的新规则;原来的「判定发生在切片之后」守的是相反顺序,现在改为弄坏「判定用的是窗口而不是筛选结果」。两者都是改写,没有删除。

client-check.mjs 里窗口条的「全场 N 轮」现在钉在两条分支各自上——它被写了两次,只断言「字符串出现过」会被另一条分支兜住。

验证

node tests/run.mjs                          32/32(本地与 fork 均是),--real 也是 32/32
python tests/mutate-ledger-batch.py         all 19 caught
python tests/mutate-selection.py            all 9 caught
python tests/mutate-sticky.py               all 14 caught
python tests/mutate-turn-breakdown.py       all 12 caught
npm run validate                            plugins/avatasia/mmc-trajectory: OK / repository: OK

浏览器实测(served 字节数 == 磁盘字节数 302,384):

  • 进度读数在浏览器里被当场抓到:正在显示 15,813 / 19,150 条,随后落定为 已显示 19,157 条
  • 选中 #1 后切到最新 200:检视回到 尚未选择记录,表里没有任何行带 aria-current,计数精确停在 已显示 200 条 · 更早的 18,964 条未加载——没有像以前那样悄悄撑到 18,964
  • console 无 error

这一轮没做的

给「显示全部」封顶、或走增量协议,两者都改变页面已经承诺的东西,所以没有塞进一个性能修复里。

The PR MiniMax-AI#15 review asked whether the rule-name guard could be reproduced, and noted that the
script behind it is not in the repository. It was not: the mutation controls and verifiers lived
in a scratch directory on the machine that wrote the client, with absolute paths in them.

Fifteen of them now live in tests/, and they run on a clean checkout with nothing but Node and
Python:

    node tests/run.mjs

Every path resolves from the directory's own location, and client-check.mjs writes its extracted
copy of the inline script to a temporary directory rather than into the checkout.

Why half of them are mutate-*.py: a guard that cannot fail is not a guard, because it reads as
coverage. Each one breaks exactly one rule in the shipped source, requires the guard that owns
that rule to fail *and* to name the right check, then restores the file and requires green again.
Two of the six report "all N mutations caught"; the rest print each mutation and the check that
caught it.

Four were left out rather than carried:

  mutate-filter-guards.mjs, mutate-qa-lens.py
      Guard the layer buttons and the turn lens. Commit 91464cd removed the layer presets in
      favour of the type chips; turns-btn, setLayer, LAYER_KINDS and applyLayerToKinds now occur
      zero times in index.html. The rule they guard is gone.

  mutate-steer.py, mutate-window-default.py
      The rules are still live -- verify-window-widen.mjs still covers the second one -- but the
      mutation anchors no longer match the source. These need re-anchoring, which is a separate
      change; shipping them red would only teach reviewers to ignore the suite.

tests/README.md records the same reasoning, and states plainly that the one-shot patch scripts
that write the package are not here and must not be run: they have no dry-run and no backup, they
write straight back to index.html / server.mjs / the READMEs, and re-running one edits text that
has since moved on.

15/15 pass. The four payload files are byte-identical to main -- this commit adds only tests/.
avatasia and others added 14 commits October 11, 2026 15:40
verify-parsers.mjs and verify-prompt-fix.mjs read this machine's ~/.minimax.
On a runner without one the suite drops to 13/15: verify-parsers crashes and
verify-prompt-fix fails two checks, so the guards would have gone red the
moment CI was allowed to run.

They now read a committed synthetic store under tests/fixtures/minimax-home.
The fixture is distilled from real transcripts rather than hand-written:
make-fixtures.mjs keeps the structure the assertions depend on -- roles,
content-part types, stopReason (the reply/narration split turns on it),
toolCallId pairing, usage numbers, and the injection markers
splitUserMessage() branches on -- and replaces every string payload.

Three things keep that honest:

- a leak assertion on the rendered bytes refuses to write if a path, user
  name or store reference survived synthesis (--selftest proves it trips);
- a coverage gate refuses to write unless all eight required shapes are
  present, so a sparse sample cannot pass by accident;
- toolCall id / toolResult toolCallId are memoised per original value, so
  corpus pairing survives -- regenerating each occurrence independently
  silently broke it and the pairing assertion would have failed.

`node tests/run.mjs --real` still runs the two guards against the local
~/.minimax, for checking a change against transcripts nobody curated.

Also corrects verify-prompt-fix.mjs, which passed resolveSessionsRoot().home
to createRedactor -- that property does not exist on the returned object, so
the store root was never one of the redacted secrets.

15/15 on this machine, 15/15 on a runner with no ~/.minimax, 15/15 with
--real. The four payload files are byte-identical to main; this adds only
tests/. npm run validate reports plugins/avatasia/mmc-trajectory: OK.
…of its own

The suite went 15 -> 22. The six guards that were only runnable against a booted service on a
developer machine now read a committed payload instead, and the two mutation controls whose
anchors had stopped matching work again.

Seven guards wanted one /api/trajectory body and got it by fetching it from a port read out of
.probe/port.txt. The route is buildTrajectoryPayload(...) plus a few session.* fields, so the
same body falls out of the fixture store offline -- make-payload.mjs writes it, with --check to
catch drift between the two fixtures.

Distilling that store properly needed the whole lineage. A compacted session keeps its old
messages in snapshots/g<gen>--ctx_*.jsonl, and the shapes the guards need -- an empty reply
above all -- live in those older generations. Reading only the active file found none of them and
verify-anomaly failed for want of something to find. The distiller now walks the lineage and the
coverage gate requires all ten shapes, including one that produced no segment at all: server.mjs
only pushes a text:null reply when segmentOrder is empty, so an assistant message that called a
tool but said nothing is not a silent reply.

Two of these were not what they looked like:

- mutate-steer.py was not stale. server.mjs is CRLF and the anchor joined its lines with \n, so
  a multi-line anchor matched zero times while every single-line anchor in the same file matched
  fine. Anchors are now built in the file's own line ending.
- mutate-window-default.py was stale, but not where it looked: the rule is intact and rePinWindow
  had gained a time-range early-return *between* the two lines the anchor spanned.

verify-block-order.mjs was wrong about the rule it names. It broke a run only on a non-assistant
record, so two assistant messages in one turn were ranked against each other and three
toolCall -> thinking pairs were reported that never occurred inside any message. The rule is
per-message; splitting runs at the message boundary finds more runs, not fewer (16 -> 19, 0
violations).

mutate-anomaly.py does not carry a case that nulls the page's liveTurn, even though that looks
like the obvious thing to break. verify-anomaly.mjs re-derives the live turn itself instead of
calling the page's function, so the mutation is invisible to every assertion in the file -- tried,
and the probe passed. Shipping it would ship a mutation known to be unobservable. The gap is
named in tests/README.md instead.

verify-time-range.mjs and mutate-time-range.py are present but skipped by run.mjs: their
assertions are calibrated against a full multi-generation session (stats.messages > 1000, a
checkpoint nested inside a turn) and reproducing that means committing a multi-megabyte payload.
Relaxing the numbers would stop exercising the timezone arithmetic the file exists for.

22/22 here, 22/22 on a runner with no ~/.minimax, 22/22 with --real. The four payload files are
byte-identical to main; this adds and changes only tests/. npm run validate reports
plugins/avatasia/mmc-trajectory: OK.
The 异常 row was the last field of the search form, which put it above the KPI
cards and the type axis — four cards away from the control it reads as part of.
It now sits directly under 全部/用户/追加, because both axes answer the same
question: which records the ledger shows.

It is a second filter system rather than another field of the first. Ticking a
rule switches the type axis, the search box, the time range and the overview
brush off; a new 全部数据 button — the state the page opens in — switches them
back on with their values intact. The inactive controls are `inert`, not merely
dimmed: a control that takes a click and then changes nothing on screen is the
defect this change is about, and reintroducing it one line away would be worse
than the original. The session and generation pickers are deliberately left
working — they choose which session is on screen, not which records of it.

The counts followed. anomalyCounts() walked passesBaseFilter while the ledger
filtered the same way, so a button promising 283 rows sat above a list of
1,494 — reproducible on a 15k-record session, and the arithmetic was right in
both places, which is why nothing caught it. It now walks the records anomaly
mode actually shows, so the number on a button is what ticking that button
puts on screen. A rule scoped to a turn is counted once per turn and published
in 轮: 无答复 was advertising 1,244 "unanswered" rows in a session that was
broken 13 times, none of which carried a sign of being unanswered. The unit
comes from the rule table, so a rename or a rescope cannot leave it lying.

verify-anomaly.mjs 116 → 158 assertions. The two that required the count to
follow the other axis are inverted rather than deleted — they were correct
while the ledger filtered the same way, and they are the reason the mismatch
had a written spec. Added: the unit follows the rule table, a turn rule is
counted per turn, the axes are alternatives rather than a stack in both the
ledger and the overview strip, the way back exists and is not a toggle, the
inactive axis is inert, and the position of the axis on the page.

mutate-anomaly.py 36 → 45 mutations, one per new invariant. Two cases were
repointed and one deleted: they encoded the stacked-axis behaviour that has
just been removed, and the anchor of a third had moved with the markup.

Also fixes a dead class — `.anomaly-btn.is-empty` was toggled on every render
and had no rule anywhere in the stylesheet.

22/22 locally, 22/22 on a clean runner, 22/22 with --real.
The button added in the previous commit as 全部数据 was neither of those things.
It is 重置: not a reverse selection — nothing is being inverted — but the row
returning to the state it was in on arrival, where the table is filtered by the
row above.

That distinction has a consequence worth spelling out. In 异常 mode the row above
filters nothing, so a 重置 that restored what that row used to hold would be
restoring a condition the page is not applying. It resets it instead — and so
does ticking a rule, on the way in. Otherwise the axis keeps drawing 「用户」
while the table is showing tool failures: the same defect as a button whose
number does not match its list, one control over.

Ticking a rule therefore puts the row above back to its default — every type, no
search, no range, no brush. `resetBaseFilters()` sets the state *and* calls
`syncKindButtons()`, because a write without the sync is what leaves the axis
drawing what was there before. That combination has its own mutation control.

重置 also drops its count. 「怎么多会留下多少」 is a question about a rule, and a
control that keeps whatever the row above says has no answer to it; printing a
number there only invites adding it up against the numbers beside it.

verify-anomaly.mjs 158 → 172 assertions. mutate-anomaly.py 45 → 50 mutations,
five new controls for the reset. One of them survived the first run, because the
assertion that catches it was in client-check.mjs and this runner does not
execute that file; the assertion moved to the file whose runner can, rather than
being dropped. client-check's "state.kinds is written from exactly three
controls" becomes four, for the reason its own comment demands — change the check
in the same commit that adds the control.

22/22 locally, 22/22 on a runner with no ~/.minimax, 22/22 with --real.
多选是一个模式开关,不是第五条规则:它在规则按钮组之外,所以不会被
syncAnomalyMode 的 data-filter-scope 打成 inert——那会把它唯一一条退路掐掉。

单选语义下点第二条是替换而非叠加,重复点同一条才关掉;多选下点第二条才叠加,
点已选中的那条只摘掉它。关掉多选时保留最后勾选的那条(读者刚才在看的那条),
而不是表格第一条。重置把规则和多选一起清回默认,并复位 checkbox。

守卫:verify-anomaly 172 -> 196,mutate-anomaly 50 -> 62。
- 两条旧断言(ticking a second keeps table order / unticking removes only that
  one)编码的是 toggle 无条件叠加的旧语义,本轮按需求推翻重写:单选写到单选
  块,原来的并集语义挪到多选块重断言,不是删掉。
- 新增覆盖:关多选保留最后勾选那条、多选开关不改变已选集合、重置同时关掉多选
  并复位 checkbox、每次 render 从 state 重绘 checkbox、checkbox 不得带
  data-filter-scope、change 而非 click 绑定。
- 修了 FAILED_NAMES 用 split(':')[0] 取名会把带冒号的断言名截断的 bug——它
  让变异跑器把「已抓住」读成「没抓住」。名字改为单独收集,可读输出不变。

浏览器实测(大会话 126 轮):单选替换、并集叠加(343 = 200 已显示 + 143 未
加载,与 324+19 有 1 条重叠吻合)、关多选保留最后勾选那条、重置回到默认且
checkbox 视觉复位。
黑盒走了一遍按钮矩阵,发现三处缺陷——都是「控件说了什么」而不是「控件算了什么」,
读代码读不出来,因为算的那部分是对的。

1. 部分并集时,提示把「合计」指向「全部异常」。实测:多选下勾 工具失败+上下文压缩,
   提示写「合计见「全部异常」」,而「全部异常」是四条规则的全集 1,707 条,当前所选的合计
   约 354 行。按提示去找合计,拿到的是另一个选择的数。
   改为在 anomalyCounts() 同一次遍历里数出所选规则的并集,提示直接给这个数。

2. 多选已勾但未选规则时,「重置」自称「已是默认状态」且 aria-pressed=true。实测证据:
   多选可见、规则全未选,重置仍读「已是默认状态」。但默认是单选,而且多选会改变之后每次
   点击的语义(替换 vs 叠加),重置是唯一能把模式拨回去的控件——告诉读者已经在默认态,
   他就没有理由去按它。

3. 单选下按「全部异常」会用多选词汇描述状态:提示写「已叠加 4 条规则」,而多选是关的,
   读者并没有叠加。改为「全部 4 条规则同时生效」。

顺带把 counts 改成单次快照:render() 原来为提示和按钮各走一遍,而会话在轮询间隙还在长,
两遍可能给出不同数字。现在一次 anomalyCounts() 同时喂两者——提示和按钮不可能各答各的。

守卫:verify-anomaly 214 -> 234,mutate-anomaly 62 -> 70。
- 补上了此前根本没测的部分并集:�ll.any 是四条规则的全集,逐条规则的检查又只选一条,
  所以多选唯一能产出的那种选择此前一条断言都没有。现在对 6 个规则对逐一断言
  counts.selected == visibleEntries().length,并用与页面无关的代码独立算出期望值。
- 两条新变异第一次跑就存活:一条是我把变异写成了空操作(sel.length === 0 || 在有选中时
  等价于原式),换成了「去掉 break」——只在两条规则重叠于一条记录时才出错的那种;另一条
  指向了不存在的断言。两条都是补守卫,不是删变异。
- 每条规则对的断言名带规则对,人读着好用,但变异跑器按名字精确匹配,哪一对命中取决于夹具。
  因此另加族级断言(名字稳定)作为变异的归属,逐对的保留作诊断。

新增 tests/METHODOLOGY.md:验证方法论。每条规则都对应一次真实代价,包括本轮踩的——
读父容器而不是过滤子元素、期望值只能来自冻结夹具、空操作变异、断言名不能带冒号、
散文要跑出来断言而不是看源码形状、以及工具驱不动的事要写明驱不动。

浏览器实测:多选两条 → 提示「合计 354 条」,账本 200 已显示 + 154 未加载 = 354;
单选按全部异常 → 「全部 4 条规则同时生效 · 合计 1,716 条」,账本 200 + 1,516 = 1,716;
多选未选规则 → 重置改为「点击回到默认状态」。

未验证:多选 checkbox 的键盘操作。Shift+Tab 在该 harness 里移不动焦点(焦点环仍在重置上,
Space 无效),所以该结论只由静态守卫(绑定 change 而非 click,配变异对照)加原生行为支撑,
不做端到端声称。
滚动时看不到异常筛选:类型轴自己 claim 过 position: sticky,异常轴没有,于是
筛选过的会话滚下去之后,页面上没有任何一行还在说明「现在被筛过什么」。

- `.filter-sticky` 把类型轴和异常轴包成一个钉在 top: 0 的不透明块。它们回答
  同一个问题(账本显示哪些记录),分开钉就会出现「类型轴还在、异常轴滚走了」
  的半截状态,比完全不钉更糟。
- 下方三层的偏移不再写死。原来的 0–32 / 32–83 只在「总览条上方只有一行筛选轴」
  的布局里成立,而异常轴会随窄屏换行、提示行的出现和消失改变高度,任何常量都会
  悄悄变成重叠。`syncStickyOffsets()` 在 `render()` 末尾量三个块,写进
  `--sticky-band-h` / `--sticky-overview-h` / `--sticky-window-h`,样式表只读变量。
- 轮头从 `top: 0; z-index: 2` 移到实测偏移下方。它原先压在总览刻度和类型轴按钮
  上,两处都读不出来;z-index 6 的筛选带盖住它只是把重叠藏起来,不是修好。

守卫:`verify-sticky.mjs`(20 条断言)把 syncStickyOffsets() 拿 137 / 65 / 57 的
stub 高度跑一遍——文件里没有任何常量等于这三个数,所以「写数字而不是量」的版本
过不去,把 band 量成 strip 的版本也过不去。`mutate-sticky.py` 18 条变异逐条要求
它以正确的名字失败,其中三条是补了守卫才有的:样式表读到的变量集合必须等于测量
写出的集合(单向都不够)、被测量的块必须真的绑定到 markup 里存在的 id(stub 永远
喂得起函数,抓不到这种静默回退)、以及任何钉在账本之上的层都必须有不透明背景
(账本行是 `background: transparent`)。

测试 22 → 24 个守卫,全绿。
这个页面的筛选逻辑是一次加一个功能长出来的,代价能逐条指到代码:

- 「看哪些记录」这一个职责住在文档位置 3 和 5 两处(工具条 / 筛选带),
  中间隔着整块 KPI 区。上一轮的 sticky 修复是这笔债的利息不是本金。
- state 里两句注释互相矛盾:`kinds` 写「类型选项是唯一筛选」,`rangeFrom`
  写「区间是 KPI 跟随的唯一筛选」。两句各自描述一个更早的页面。
- kpiStats() 只跟区间,不跟类型轴、不跟搜索、不跟异常轴——汇总层和筛选层
  共用了同一套视觉语言。
- 异常轴被实现成「第二套系统」而不是筛选层的一个字段,于是互斥只能散落在
  两个渲染函数里手写。上一轮修的三处缺陷没有一处是算错,全是「两个筛选器
  各自推理自己」的症状。
- window-bar 在列表面板里,写的却是「渲染多少」这一层的事。

docs/LAYERS.md 给出 L0–L6 的归属表、六条不可违反的规则,以及实现之前必须先
答的四个问题:归层、关联、布局、口径。最后一问是「联动异常」的主要来源——
加了一个能力改了一个数字,别的地方还在按旧口径说话。

本 PR 不改页面代码,Stage 1–3 在文档里分阶段列出。
按 AGENTS.md 的强约束,Stage 1 在实现之前先把设计冻结成可判定的断言。

docs/STAGE1-SPEC.md:
- 四问在本阶段的答案。L2 容器装四组控件(搜索 / 区间 / 类型轴 / 异常轴),
  L1 的会话与上下文代留在外面,不动 Stage 2 的模型,一个数字都不改。
- 口径读点清单:KPI↔区间这条口径共有 4 个页面读点、4 个守卫读点。其中
  verify-time-range.mjs 在 run.mjs 的 SKIP 集合里默认不跑,所以这条口径在
  默认套件里目前没有保护——而 Stage 1 恰好要动 KPI 的位置,所以必须补一条
  能默认跑的断言,不能依赖那个跳过文件。
- 10 条断言,每条都写了「必须拒绝的错误实现」和配套变异。其中 V7(把区间塞进
  类型轴)是这一阶段最容易犯的错。
- 4.1 节记了搬家会引出、原 scope 没写的后果,逐条对过代码:
  F1 .toolbar-row 承载区间那行的 flex,移出表单后静默失配;
  F2 form#toolbar 的 submit handler 只为拦搜索框的隐式提交,搬走后表单里只剩
     两个 select,role="search" 成了假话;
  F3 异常模式的 inert 靠后代选择器 [data-filter-scope],搬家后仍生效,异常轴
     本身不带该属性(否则会把自己 inert 掉);
  F4 applyFilterScope() 名字说 L2、只写 L4 的口径行——记进 LAYERS.md 的 D7,
     本阶段不改名,留给 Stage 2。

跨家族只读查漏试了 codex / claude / codebuddy,三个都被认证挡住(模型不被
ChatGPT 账号支持 / OAuth 过期 / 需要 /login),三次都设了只读机制且比对过文件
哈希,未被改动。所以这份 spec 是「还没查过漏的」,已在 7.1 节写明解锁方式。

24/24 守卫全绿,validate OK。
用 codex / gpt-5.6-sol(OpenAI,与本会话的模型不同家族)做只读查漏。
Windows 沙箱在 exec 下初始化失败,第一次交白卷——它明确拒绝捏造文件:行,
且哈希确认未改动文件。改为把材料内联进 stdin 后跑通。
只有 gpt-6.1-sol 被 ChatGPT 账号拒绝,gpt-5.6-sol 等可用。

它找对的(已接受并改进 spec):
- V1 的变异是空操作:把 KPI 接在四组之后时四组仍连续,V1 照样通过。
  这是我给守卫定的规矩(变异存活就补守卫)在自己身上失效了。改成把 KPI
  插在搜索与区间之间,并让 V2 用「接在四组之后」的互不等价变异。
- V1/V2 共用同一变异,等于其中一条没被测到。已拆开。
- V10 只覆盖一种丢标记方式。拆成 V10a/b/c + V11。
- 第 4 问的口径表漏了一个读点:#kpi-scope 不只是三句措辞,还发布
  stats.messages 与 stats.turnsWithPrompt 两个数字。
- 同一份表里我把写 #kpi-scope 的函数写成了 updateKpiScope()——那个函数
  不存在,真正的写入者是 applyFilterScope(),也就是 D7 说「名字与职责不符」
  的那个。我在自己的文档里把它的名字写错了一次。
- 「保留 form 因为 select 需要 label 关联」是错的:for/id 关联与是否在
  form 内无关。改为直接删掉失去对象的 submit handler。
- verify-sticky.mjs「只改指向」不足:现有子树断言只要求类型轴与异常轴同在
  band 内,扩到四组必须重写成员与排他性,并显式排除 L1 与 L4。
- I1 用下标比大小发现不了中间的插入;改为解析容器子节点的标签序列。
- I4 的不透明判断太粗;限定为已定义的 --mcode-* token。
- 焦点顺序变化(会话→代→搜索→区间→KPI 变成 KPI 插进筛选流程中间)不能列为
  「不验证」——DOM 顺序正是本阶段直接改的东西。新增 F5 并撤销原排除。
- aria-controls 只指向 overview、而同一筛选也控制账本;L2 容器也没有可访问
  名称。新增 F6。
- 窄屏下四组全 sticky 可能占满视口。新增 F7 与 B1–B4 四条浏览器补测。

它指错的(已驳回):
- 「目前没有保护」过度表述。spec 第 94 行原文是「在默认套件里目前没有保护」,
  已经限定范围;它引的第 88 行是表格行,那行里没有这句话。
- verify-sticky.mjs:79 的行号落在 value() 这个辅助函数上,不是它要谈的
  子树成员断言。

断言 10 条 -> 15 条(V10 拆三条 + 新增 V11/V12),后果 4 条 -> 7 条,
浏览器补测 4 条。spec 第 8 节留了完整对照表,含驳回项与理由。

24/24 守卫全绿,validate OK。spec 仍未冻结,需再查漏一轮。
Selecting a record and then changing any filter left the inspector describing a
row that was no longer on screen. Two functions decided what 'selected' meant
and they read different lists: the rows were built from visibleEntries(), while
revealSelection() and selectedRecord() scanned state.flat — the whole session,
unfiltered. Worse, the only caller of revealSelection() was applyPayload(), and
only when the payload signature changed, so no local filter change ever reached
it. The ledger read 已显示 2 条, the inspector still named the old record, and no
row anywhere carried the selection mark; 复制 JSON and 交给 AI 分析这条 read the
same stale record and stayed enabled.

revealSelection() is now handed the very array the rows are built from and is
called from renderLedger(), the one exit every path that changes what the table
shows goes through — nine of them call it directly instead of calling render().
It reports whether it dropped the selection, and the exit redraws the panel when
it did: clearing the id alone is not enough, because those nine callers never
redraw it, so the search box left the panel naming a record that had just been
dropped. selectedRecord() stays a plain lookup and documents why that is safe.

The sticky offsets had the same shape of problem. syncStickyOffsets() had exactly
one call site — the end of render() — while nine writers change what the table
shows without going through render(), so the invariant lived in a convention
rather than in the structure. There was also no resize hook at all, and on a
session whose payload has stopped growing every poll is a 304 and render() never
runs, so a window resize that wrapped the type or 异常 row left the pinned strips
at the old width indefinitely. The measurement now runs at the same renderLedger()
exit, and resize re-measures through a requestAnimationFrame so a window drag
does not measure on every event.

docs/LAYERS.md D1 is withdrawn. Treating the filter controls being split across
two regions as structural debt was wrong: query conditions are an aggregation, and
the frontend UI being separate from the request construction is the design.
docs/STAGE1-SPEC.md is marked not-adopted with the reason. D8 and D9 record what
was actually wrong, and both are closed.

Guards 24 -> 26. verify-selection.mjs runs revealSelection() against a state
carrying both the session and the filtered list, so a version that reads the
unfiltered one again answers wrongly instead of throwing; mutate-selection.py
breaks nine of its rules. verify-sticky.mjs gains five assertions for the new
exit and the resize hook; mutate-sticky.py gains seven mutations. One of the nine
mutations survived the first run — it moved the call to the line above the slice,
which still raised state.limit in time, so it changed the text and not the page;
the mutation was rewritten to actually move it past the slice.
A finished session costs nothing: the ETag is derived from stat facts, so an
unchanged session is a 304 and nothing is rebuilt. The cost is the session that
is still growing, which is the one a reader is watching on purpose — and there
every poll rebuilt everything from scratch.

Measured on a 23-generation session (payload.session.sizeBytes = 99,986,501):

  * 22 of the 23 generation files are snapshots, written once when the runtime
    rotates generations. Between compactions they do not move at all. Only
    messages.jsonl is appended to — 97,526 bytes, 0.10% of the lineage.
  * trajectoryCacheKey() carries each generation's byteLength, so a growing
    session misses the whole-session cache every poll and re-reads and
    re-parses all 95 MB, 99.90% of it byte-identical to the previous poll.
  * On the wire: 1.31 GB delivered over five minutes to carry 118,875 bytes of
    real growth. 11,006x amplification, 99.99% re-sent.

Two fixes, both on the read path.

Per-generation parsed rows. loadParsedRows() keys a cache on
path+size+mtime+limit and hands back the rows of any generation file that has
not moved. Cold 306 ms, warm 2 ms over the 22 frozen files, measured directly
against the real files. Bounded by bytes rather than entries, because entries
say nothing about volume — one generation file is 40 KB and the next is 10 MB.

  * The key names the path, not the basename: every session has a
    messages.jsonl, so a basename key lets two of them collide. The fixture
    builds exactly that collision — same size, same mtime, different
    conversations — and asserts neither sees the other's rows.
  * The key carries the size and the mtime rather than the catalog's recorded
    byteLength, which this file already documents as a snapshot that goes
    stale and that the active file grows past between polls.
  * The key carries the limit, because readHead truncates at it and the
    truncation verdict is part of what the entry holds.
  * A file that changes between the stat and the read lands on a key that
    describes a size it does not match; the next poll stats again and misses.
    The race can only cost a redundant read.

The untouched message, on demand. buildRawPayload() ran JSON.stringify on every
record to measure it against a 64 KB cap — about 100 MB per rebuild — to
produce a string that was thrown away, because splitRawRecords() strips the
message off the bulk response before it is ever serialized and /raw serves the
message object rather than the string. A row now carries the position of its
message instead, and /raw builds it for the one record that asked.

That moves 「原文过大,已省略」 out of the bulk payload and into the /raw
response. It is a contract change, so every reader is listed in LAYERS.md D11
and asserted by name: none of the five is in the ledger — all of them are on
the path a reader takes after opening a record. The only visible difference is
that the raw tab now spends one request to learn a message was too large
instead of knowing it in advance.

Guards 26 -> 30, mutations 34 -> 58, all caught.

  verify-parsed-rows.mjs   53 assertions. A hit is compared against a cold read
                           by serializing both bodies and requiring them equal,
                           not merely similar. A version that read the wrong
                           source, or keyed on the wrong fields, produces the
                           wrong answer here rather than looking plausible.
  verify-raw-lazy.mjs      39 assertions. The 64 KB cap is asserted through the
                           path that now owns it, rather than dropped: the
                           locator is followed to its row and buildRawPayload()
                           is run there.

Three things the mutations caught in this change and the guards did not have:

  * A locator written as Number(record.rawRow) accepted null as row zero —
    Number(null) is 0 — so /raw would have answered with a different record's
    message. Now rejected rather than coerced.
  * Removing the delete that takes the locator off the row let rawRow ship on
    every record. Nothing asserted its absence.
  * mutate-pr15-fixes.py had a mutation whose anchor this change erased. It was
    reporting SKIP, not a pass, but a dead mutation reads like a green one; it is
    re-anchored to the line that now carries the same decision.

Correction to LAYERS.md D10, which this change's measurement disproved: the
read-and-parse was ~306 ms, not the ~1,000 ms written there. That number was
obtained by subtracting the serializer from the rebuild total, which folded
buildTrajectoryPayload and buildRawPayload into "read and parse". D10's
payoff is therefore ~300 ms of a ~1,400 ms rebuild, and the resident heap is a
measured 116 MB rather than the 130-290 MB estimated there.
Two defects, both reported from the same session, both with a cause that was
structural rather than a missing line.

**A — 「显示全部」 froze the page.** Building 18,924 rows inside one
synchronous loop took 3,956 ms of main thread with nothing on screen saying
so. Measured on this machine (150 turns, 18,924 records, 735x900):

    batched build (25 turns/frame)   3,956 ms total
      row construction    median 141 ms / max 151 ms per batch
      between frames      ~617 ms  <- layout + paint of the whole document

The blocking work was not the row construction. It was each frame's layout
pass over a document that was already two million pixels tall, so making the
batches smaller does not help — it multiplies the passes.

The fix keeps the promise and changes the pacing: one rAF builds at most 25
turn groups, then hands the thread back, and the window bar reports
「正在显示 N / M 条」 while it goes. Every pass through renderLedger() bumps
ledgerBuildToken first, so a build still in flight retires itself at its next
step instead of appending into a list that has been cleared underneath it.
The poll tick checks ledgerBuilding() first and remembers the tick — at 3,956
ms against a 2,500 ms interval, letting every tick through would restart the
build before it ever finished and the reader would watch the count reset for
ever. flushDeferredPoll() runs the held tick once the build ends.

Nothing about what the window shows has changed. 「显示全部」 still means every
record.

**Rejected, and recorded in LAYERS.md D12 because it is the obvious next
question:** `content-visibility: auto` on `.turn-group` is a real 4.0x —
3,956 ms to 979 ms, same row count, and the per-frame layout cost disappears
with it. It needs one `contain-intrinsic-size` for elements it skips, and the
150 turn groups span 1,011 px (p25) to 29,754 px (max), median 2,667 px. Any
single value makes the scroll height wrong by a large factor, and `auto` only
remembers sizes of groups that have already rendered. That is trading a
control that describes the page accurately for speed, so it is not shipped.

**B — the inspector survived the record it was describing.** Select an early
record in 「显示全部」, press 「显示最新 200 条」, and the row leaves the table —
but the panel went on naming it. revealSelection() was judged before the
window was sliced and raised state.limit to reach the selection, so the reader's
own click changed the number on the bar a second later:

    press 「显示最新 200 条」   已显示 200 条
    one second later           已显示 15,224 条

Two things wrong with that: the window became something the inspector could
write, which gave state.limit a second writer, and the test was the wrong one —
「what the reader can see」 is the window, not the filter result.

revealSelection() is now a predicate over the rows that will actually be drawn.
It writes state.selectedId and nothing else, and it runs after the slice. A
selection below the window is dropped rather than revealed: the reader's
explicit window choice outranks an earlier click, and it matches what the
filter axes already do.

**Guards**

- tests/verify-selection.mjs — rewritten, 14 assertions. This is a 口径 change,
  not a refactor: the file previously pinned 「the window is pulled down to a
  selection below it」 (asserting limit === 8) and that the judgement ran
  *before* the slice. Both are now the opposite, and the new promise is asserted
  instead, including that the function's source may not mention `limit` at all.
- tests/verify-ledger-batch.mjs — new, 20 assertions over the four things that
  have to hold together: the frame is really handed back; a new exit retires the
  build in flight; progress goes through updateWindowBar, the one writer of that
  count; a poll defers without swallowing the tick.
- tests/mutate-ledger-batch.py — new, 19 mutations, all caught by name.
- tests/mutate-selection.py — two mutations rewritten rather than deleted. The
  old 「the window is never pulled down to the selection」 was guarding the very
  state.limit write this change removes; it now breaks the rule that replaced
  it. The old 「the selection is judged after the window is sliced」 guarded the
  opposite ordering, so it now breaks the new one.
- mutate-sticky.py, mutate-turn-breakdown.py, client-check.mjs — re-anchored
  and tightened where the refactor moved text. The window-bar total is now
  pinned on both of its branches, since it is written twice.

**Verified**

- `node tests/run.mjs` — 32/32 locally and in the fork; `--real` also 32/32
- mutate-selection 9/9, mutate-ledger-batch 19/19, mutate-sticky 14/14,
  mutate-turn-breakdown 12/12
- `npm run validate` — mmc-trajectory OK, repository OK
- Browser, served bytes == disk bytes (302,384): the progress readout observed
  live mid-build at 「正在显示 15,813 / 19,150 条」, then settling to
  「已显示 19,157 条」; with row MiniMax-AI#1 selected and the window narrowed, the
  inspector returned to 「尚未选择记录」, no row carried aria-current, and the
  count stayed at exactly 「已显示 200 条 · 更早的 18,964 条未加载」 instead of
  silently widening. Console clean.

The remaining lever on the A side is making the two-million-pixel number
smaller — capping 「显示全部」, or the incremental protocol. Both change what the
page promises, so neither was done inside a performance fix.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant