fix: handle read index overshoot - #246
Conversation
kibertoad
left a comment
There was a problem hiding this comment.
The core fix looks correct. Atomics.waitAsync registers atomically, so a READ_INDEX change either resolves the promise with 'ok' or makes the call return synchronously with 'not-equal'; routing both back to the caller lets waitForRead re-sample WRITE_INDEX and closes the overshoot loop from #245. Full suite passes on the branch (the test/error-flush.test.js flakiness I saw also reproduces on main, so it is not from this PR).
Three comments inline, one of which I think leaves a live variant of the same bug.
|
Addressed all review threads in a011ff3 and added coverage for fallback-timeout changes, synchronous I also retried the failed CI jobs three times. The remaining failures are pre-existing Node 26 flakes outside this change: Windows repeatedly times out in |
Outside production the server logs through pino-pretty in a thread-stream worker, which stays referenced until its READY handshake sees the read index reach a write index it snapshotted earlier. thread-stream 4.2.0 compares with ===, so when logging continues during startup the read index can jump past the snapshot, or be reset under it, and the handshake never completes. The worker then holds the event loop open forever. That is the image-dimension regression's hang (#6529): its retry warnings land while the worker boots, it prints its success line, and then sits until the runner kills it at 30 seconds. It was never a keep-alive socket; the only live handle at hang time is the worker's MessagePort, with the stream stuck at ready=false and read == write. Apply upstream's fix (pinojs/thread-stream#246 and #251) as a pnpm patch until a release after 4.2.0 ships it. Measured on this machine, the spec alone: 11/30 runs hung unpatched, 0/40 patched. The regression runner also names every file that did not pass, so a timeout no longer shows up only as "1 failed".
Restore
wait()'snot-equalresult when a notified value skips the expected snapshot, allowingwaitForRead()to resample instead of waiting forever. This also covers values that cycle before notification and error sentinels.The regression was introduced by #178, whose
Atomics.waitAsyncrefactor removed the historicalnot-equalresult. #198 later moved the affected logic intowaitForRead()without correcting it.Fixes #245