Skip to content

fix: lock the actor instance by primary key - #53

Merged
cardmagic merged 1 commit into
mainfrom
fix/enqueue-create-race-deadlock
Sep 21, 2026
Merged

cardmagic merged 1 commit into
mainfrom
fix/enqueue-create-race-deadlock

Conversation

@cardmagic

Copy link
Copy Markdown
Owner

Why

Concurrent callers that create the same actor from inside their own transaction
deadlock on MySQL.

mysqlSql rewrites the portable ON CONFLICT (...) DO NOTHING into
INSERT IGNORE (src/database/mysql.ts:237). When the row already exists,
INSERT IGNORE leaves a shared lock on the identity index.
enqueueInTransaction then read that same row FOR UPDATE, so each caller held
shared and each waited for exclusive on one index record.

*** (1) HOLDS THE LOCK(S):
  index actor_type of table `..._instances`  lock mode S
*** (1) WAITING FOR THIS LOCK TO BE GRANTED:
  index actor_type ... lock_mode X locks rec but not gap waiting

Six of eight concurrent callers failed with ER_LOCK_DEADLOCK.

Why it was not visible

enqueue retries ER_LOCK_DEADLOCK up to eight times (repository.ts:280),
so the top level path absorbs it. That retry does not wrap
enqueueInTransaction, and every nested caller enqueues inside a transaction
that already wrote: the effect executor (:804), the retry path (:1137),
effect recovery (:1523, :1571), reminders (:1762), and the recovery
coordinator (:2025). A deadlock there aborts the caller's whole transaction.

The top level path deadlocks too. It only looks healthy because it pays a
deadlock and a backoff on every cold start burst.

What changed

  • Read the instance without a lock first, on every database, then lock it by
    primary key rather than by actor type and actor id.
  • After an ignored insert on MySQL, read the winning row in shared mode and
    select only its id.

Selecting only the id matters. A shared read that selects every column also
locks the clustered record, which does not remove the upgrade deadlock, it
moves it onto the primary key:

index PRIMARY of table `..._instances`  lock mode S
index PRIMARY ... lock_mode X locks rec but not gap waiting

An index-only read stays inside the identity index and never touches the
clustered record, so the later primary-key lock has nothing to upgrade.

PostgreSQL and SQLite take no shared lock. PostgreSQL defaults to read
committed, so a plain read already sees the winning row, and a share lock there
creates the upgrade deadlock it prevents on MySQL.

Tests

Both are new and follow the failing-test-first order.

  • test/mysql.test.ts drives eight concurrent nested callers at one new actor.
    It fails on main with six ER_LOCK_DEADLOCK results and passes here, four
    runs out of four.
  • test/postgresql.test.ts carries the same race. PostgreSQL never had this
    bug, so that test guards the fix rather than the original defect: it fails
    when the MySQL shared read is applied to PostgreSQL as well, which is how the
    adapter split was chosen rather than assumed.

Both assert one instance row and sequences 1 through 8.

Validation

Gate Result
default suite 404 passed, 30 skipped
MySQL 8.4 suite 39 passed, 7 skipped
PostgreSQL 18 suite 50 passed
format:check pass
check pass
build pass
test:package pass
test:recovery pass

Compatibility

No migration and no API change. FOR SHARE needs MySQL 8.0, which is already
the tested minimum in src/doctor.ts and both CI versions. The removed lock
option on the private findInstance had no remaining caller.

Release

This branch also cuts 0.15.2, matching solid_objects 0.15.2 for Ruby, which
carries the same fix.

Concurrent callers that created the same actor from inside their own
transaction deadlocked on MySQL. The portable conflict clause becomes
INSERT IGNORE, which leaves a shared lock on the identity index when the
row already exists. The mailbox then read the same row FOR UPDATE, so
two callers each held shared and each waited for exclusive on one index
record. Six of eight concurrent callers failed with ER_LOCK_DEADLOCK.

Top level enqueue hid this, because it retries ER_LOCK_DEADLOCK eight
times. Nested callers get no retry, and the effect executor, the
reminder scheduler, the retry path, and effect recovery all enqueue
inside a transaction that already wrote.

Read the instance without a lock first, on every database, then lock it
by primary key. After an ignored insert on MySQL, read the winning row
in shared mode and select only its id, so the read stays inside the
identity index and never locks the clustered record. A shared read that
selects every column locks the clustered record too, which only moves
the same upgrade deadlock onto the primary key.

PostgreSQL and SQLite take no shared lock. A share lock on PostgreSQL
creates the upgrade deadlock it prevents on MySQL, and its regression
test fails when that lock is applied there.

Validate with the default suite, the MySQL suite, the PostgreSQL suite,
format:check, check, build, test:package, and test:recovery.
@greptile-apps

greptile-apps Bot commented Sep 21, 2026

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with the database-specific locking behavior covered by focused MySQL and PostgreSQL concurrency tests.

Summary

This PR changes instance acquisition so enqueue first resolves an actor without locking, performs a MySQL-specific shared identity-index read after a contested insert, and then acquires the update lock by primary key.

  • Adds MySQL and PostgreSQL concurrency coverage for eight transactional callers creating the same actor.
  • Documents the concurrency guarantee and releases version 0.15.2.
  • Keeps the shared-lock behavior isolated to MySQL while retaining primary-key update locking on MySQL and PostgreSQL.

Diagram

sequenceDiagram
  participant E as enqueueInTransaction
  participant I as Identity index
  participant P as Primary row

  E->>I: SELECT id by actor type and actor id
  alt Instance is absent
    E->>I: INSERT instance ON CONFLICT DO NOTHING
    alt MySQL
      E->>I: SELECT id FOR SHARE
    else PostgreSQL or SQLite
      E->>I: SELECT id
    end
  end
  E->>P: SELECT row by id FOR UPDATE
  E->>P: Allocate sequence and update instance
  E->>P: Insert queued message
Loading

Reviews (1) · Last reviewed commit: "fix: lock the actor instance by primary ..."

@cardmagic
cardmagic merged commit a35954d into main Sep 21, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant