Repository navigation
Remove a coordinated-retry park's lock wake callback when the park ends - #905
Merged
Merged
Conversation
A park that lost a conflict registered a wake callback on the holder's LockTracker and never removed it when the park ended by timeout, env teardown, or close. Behind a holder that never releases, every re-park added another inert callback for as long as the lock was held. addWakeCallback() now returns a WakeRegistration: a weak handle to a lazily allocated callback list plus a list iterator. The park's ParkTimeoutRegistry entry owns it and cancels it before calling the TSFN (and on destruction for every other drain). The handle never references the tracker, so cancelling after the tracker is freed is a no-op and no path takes the VT writerMutex_ that wake() already runs under. wake() marks the list drained before detaching its callbacks, so a cancel that loses that race leaves the detached callback to run, as before. lockWakeCallbackCount() reports the process-wide registered total for tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TCdH23J2tPQgQq4PrWGjEh Dispatch-Task: rocksdb-js-wake-callback-deregistration
Address pre-push review round 1: - wake() constructed a local std::list to swap the callbacks into, which allocates a sentinel on MSVC; a throw there, under writerMutex_ after a commit has landed, has no recovery. wake() now marks the list drained and runs the callbacks in place. - A WakeRegistration held its list weakly, so moving one after wake() freed the list copied a dangling iterator. It now holds the list, which outlives every live registration (only a registration's own cancel() erases its node), and resets its iterator when cancelled or moved from. - The timeout fixture observes each registration within an 800 ms window of a 1 s park timeout, instead of 200 ms of 250 ms. - AGENTS.md invariant 12: fire(id) also takes the wake list's leaf mutex. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TCdH23J2tPQgQq4PrWGjEh Dispatch-Task: rocksdb-js-wake-callback-deregistration
Contributor
There was a problem hiding this comment.
Code Review
This pull request introduces a mechanism to cancel lock wake registrations when a coordinated-retry park ends due to timeout, environment exiting, or database closing. It adds a LockTracker::WakeRegistration class to manage the lifetime of wake callbacks, ensuring they are properly removed from the LockTracker's wake list when cancelled, preventing memory growth issues. Additionally, it introduces diagnostic functions and comprehensive tests to verify that registrations are cleaned up correctly under various scenarios. There are no review comments, so I have no feedback to provide.
kriszyp
marked this pull request as ready for review
October 6, 2026 16:47
Resolve the root DESIGN.md index conflict by keeping both new entries. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TCdH23J2tPQgQq4PrWGjEh Dispatch-Task: rocksdb-js-wake-callback-deregistration
Contributor
📊 Benchmark Resultsget-sync.bench.tsgetSync() > random keys - small key size (100 records)
getSync() > sequential keys - small key size (100 records)
ranges.bench.tsgetRange() > small range (100 records, 50 range)
realistic-load.bench.tsRealistic write load with workers > write variable records with transaction log
transaction-log.bench.tsTransaction log > read 100 iterators while write log with 100 byte records
Transaction log > read one entry from random position from log with 1000 100 byte records
worker-put-sync.bench.tsputSync() > random keys - small key size (100 records, 10 workers)
worker-transaction-log.bench.tsTransaction log with workers > write log with 100 byte records
Results from commit be1ab81 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
⊙ Problem
A
coordinatedRetrycommit that loses a conflict parks on the holder's verification-table (VT) lock by registering a wake callback on that lock'sLockTracker.LockTrackeronly hadaddWakeCallback()andwake(), so when a park ended any other way than the holder releasing (the bounded park timeout from fix(transaction): bound the coordinated-retry park with a descriptor-owned timeout, its worker exiting, or its database closing), its callback stayed registered until the holder eventually released. Behind a holder that never releases (Unsettled commit promise wedges all writes on a thread indefinitely, the case the timeout exists for), every re-park, about one per waiting transaction every 5 s, added another inert closure for as long as the lock was held. AGENTS.md invariant 12 recorded this as a known, deferred gap.💡 Solution
LockTracker::addWakeCallback()now returns a move-onlyWakeRegistration; cancelling or destroying it removes the callback in O(1). Each park'sParkTimeoutRegistryentry owns its registration and cancels it as the park ends, so a lock's registered callbacks are exactly the parks still waiting on it.wake()behaves as before for live waiters: registration order, each once, invoked with no lock of its own held, and a registration afterwake()returns empty so the caller resolves inline.WakeList(ashared_ptrplus astd::listiterator), never the tracker. The park drops its tracker reference right after registering, and dropping one takes the globalwriterMutex_thatwake()already runs under, so a tracker reference cannot be held across the park.writerMutex_, and only contended locks get a list.wake()marks the list drained and runs the callbacks in place. It does not allocate, because a throw there (underwriterMutex_, after a commit has landed) has no recovery. A cancel that loses the race towake()leaves its node alone and the callback still runs, which is why the park's closure keeps its weak references and exactly-once gate.ParkTimeoutRegistry::resolve()cancels before calling the TSFN, so JavaScript that observes a park's result never sees it still registered. Env teardown and shutdown cancel by destroying the entry.lockWakeCallbackCount()diagnostic (internal, not exported fromindex.ts) follows thetransactionLogMapCountprecedent.⚖️ Alternatives
The planning review returned
Framing-Verdict: better-alternative-exists; I adopted its alternative: astd::listwith iterator tokens instead of the plannedstd::mapkeyed by registration id. It has O(1) removal under the registry mutex, the same one allocation per registration, and no id counter.fire()on the real-wake path, which already holdswriterMutex_, andunrefTracker()takes it again (self-deadlock); avoiding that needs a lock-free "never the last reference" decrement that breaks the VT rule that every tracker free is serialized bywriterMutex_.writerMutex_, and tracker addresses are reused, so identity needs the 14-bit generation, which wraps every 16K installs.addWakeCallback(): rejected. It leaves the last batch registered on a lock nobody re-parks on (it fails "zero after N timed-out parks"), and scans every live waiter on each add.🔧 Changes
src/binding/core/verification_table.h/.cpp:LockTrackergainsWakeList(mutex,drainedflag,std::listof callbacks) and the move-onlyWakeRegistration(listshared_ptr+ iterator; move-assignment cancels the registration it replaces).addWakeCallback()creates the list on first use and returns a registration with the strong exception guarantee.cancel()erases its node unless the list is drained and reports whether it did.wake()moves the list out of the tracker, marks it drained, and runs the callbacks in place. A process-wide atomic counts registered callbacks.src/binding/database/db_descriptor.h/.cpp: eachParkTimeoutentry owns itsWakeRegistration, and the class comment now gives the remaining reason parks are keyed by id.resolve()cancels before calling the TSFN;releaseByEnv()andshutdown()cancel by destroying the entry. NewattachWakeRegistration()(declaration) hands a registration to a park that is still pending, or cancels it if the park already ended.src/binding/transaction/transaction.cpp:completeCommitWork()attaches the returned registration to its park, and falls back to the existing inline resolve when the lock was already released or registration throws. The comment on the closure'sweak_ptrcaptures now gives the remaining reason they are needed (a callbackwake()already claimed can outlive its park).src/binding/binding.cpp,src/load-binding.ts: thelockWakeCallbackCount()diagnostic (native function, export, TypeScript binding).AGENTS.mdinvariant 12 (known-gap paragraph replaced,fire(id)'s locks and the weak-capture reason updated),src/binding/core/DESIGN.md(new "Wake registrations" section, including the no-allocation rule forwake()) and the rootDESIGN.mdindex line.✅ Verification
Route: new child-process integration tests, plus native unit tests for the list itself.
pnpm test:native: 319/319 before the merge. TheLockTrackerWakesuite cover cancel-then-wake, order and exactly-once delivery, wake with no waiters and repeated wake, cancel racingwake(), cancel after the tracker is freed, release of captured state on destruction, and move-assignment. The concurrent test repeated 30× clean; an instrumented run showed every round reaching all three outcomes (≈2000 cancelled before the wake, 15–46 claimed by it, ≈2000 added after it).test/lock-tracker.test.ts: 16/16 underROCKSDB_JS_COMMIT_THREADunset,0, and2. The four new scenarios each first observe a registration, then require zero while the holder still holds:timeout(3 parks against one held lock),wake(release well under the 5 s default timeout),worker-exit(Node only),foreign-close(a one-slot VT puts the park on another database's tracker; closing the waiting database).timeout,worker-exitandforeign-closefail with 1 callback left registered;wakestill passes, as it should.pnpm test(full, Node): 1139 passed, 10 skipped, 0 failed at 47781d1; after mergingmain(which brought in Make optimistic commit lock buckets and validation policy configurable and its concurrent commit threads), 1155 passed, 10 skipped, 0 failed, andpnpm test:native323/323.pnpm check: clean.Test files:
test/native/verification_table_test.cc(theLockTrackerWakesuite, concurrent test),test/fixtures/fork-park-wake-registration.mts(timeout scenario, foreign-close scenario),test/workers/park-wake-registration-worker.mts(the parked worker),test/lib/park.ts(shared holder/park setup; it callspopulateVersionfirst because the process-wide VT is created on the first version lookup and writes before that lock nothing), andtest/lock-tracker.test.ts(runs the scenarios; the existing park-timeout test now shares its child-process helper).🤖 Generated with Claude Code
https://claude.ai/code/session_01TCdH23J2tPQgQq4PrWGjEh
— Claude Opus 5.5
Related PRs: #897 independent (now in main and merged into this branch; it shares transaction.cpp and db_descriptor.*, not the park code), #842 independent, #767 independent, #900 independent, #902 independent (reached main through #897), #742 independent (already in base), #901 independent (already in base), #890 independent (already in base)
Complexity: complicated
Dispatch: task
rocksdb-js-wake-callback-deregistration· queued by unknown · ran by claude/opus/xhigh · worker kzyp-xps-1Review-Coverage: authored=claude; ran=cursor-composer,gemini,codex; adjudicated=domain; declined=cursor-grok,cursor-kimi,cursor-muse; rounds=3; full=1 @ 2b9befd
Review-Attention: study ~15m (critical: transaction.cpp; decisions: do-less-alternative, strong-list-ownership, cancel-inside-resolve, diagnostic-surface) @ 2b9befd