Skip to content

turbine: end the open transaction before a partition drop - #440

Merged
turbolytics merged 2 commits into
mainfrom
fix/drop-conflict-on-lost
Oct 5, 2026
Merged

turbolytics merged 2 commits into
mainfrom
fix/drop-conflict-on-lost

Conversation

@turbolytics

Copy link
Copy Markdown
Owner

Closes #437.

The bug

When a partition_owned worker's group session failed, it crashed while dropping the partitions it had lost:

error dropping revoked partitions: dropping partitions 3, 4, 5 of usage.events from usage_window:
  Invalid Argument: TransactionContext Error: Conflict on tuple deletion!
dropping revoked partitions failed; the pipeline is stopping

This was the other half of the network blip in #436: one worker stalled silently, the other crashed.

Why

The window table is written from two DuckDB connections: the pipeline's (batch inserts, the partition drop) and the window manager's (the pass that publishes closed minutes and deletes their rows). Each has its own transaction, and DuckDB's concurrency is optimistic.

The drop ran inside whatever transaction the pipeline's connection already had open. Between batches that's an idle tick's progress write, which opens a transaction and takes its snapshot before the drop takes the pass lock. If a pass then deleted the same rows and committed, the drop's DELETE hit tuples already gone, and DuckDB refused it. The pass lock couldn't help: it excludes passes from now on, not one that committed before the snapshot.

Nine statements across two connections reproduce it exactly.

The fix

dropPartitions commits the pipeline's open transaction before deleting, so the drop begins a transaction whose snapshot is after every pass the lock excludes. The progress write is rewritten at every commit, so committing it early loses nothing.

Test

TestTurbine_PartitionDropSucceedsAfterAPassDeletedTheRows drives a real Turbine with a DuckDB dropper over two connections in exactly that order, then loses the partitions.

  • On main: the run stops with the crash log's error, word for word.
  • With the fix: the run continues.

internal/core and internal/cli/run pass under -race.

Not reproduced end to end

Unlike #436, this one needs a window pass to land in a ~300ms gap after an idle tick, so it's timing-dependent in the stack. The two-connection test is the proof.

The drop ran inside whatever transaction the pipeline's connection had
open. Between batches that is an idle tick's progress write, begun
before the drop took the pass lock. Its snapshot could predate a window
pass that had since deleted closed buckets' rows on the manager's own
connection and committed; DuckDB then refused the drop's delete of those
rows with "Conflict on tuple deletion", and the pipeline stopped over it
(#437). The pass lock could not help: it excludes passes from now on,
not the one that committed before the snapshot was taken.

dropPartitions now commits the open transaction first, so the drop
begins one whose snapshot is after every pass the lock excludes. The
progress write is rewritten at every commit, so nothing is lost by
committing it early.

Test: TestTurbine_PartitionDropSucceedsAfterAPassDeletedTheRows drives
a real Turbine with a DuckDB dropper over two connections in exactly
that order. On main the run stops with the crash log's error; with the
fix it continues. internal/core and internal/cli/run pass under -race.

Closes #437.
The test left the loop running and let t.Cleanup close its connection
and database under it: a use after free in the ADBC driver, which
segfaulted on CI's Linux runner and happened to pass on a Mac. The loop
is now cancelled and waited for before anything closes.
@turbolytics
turbolytics merged commit 81bf1bd into main Oct 5, 2026
7 checks passed
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 5, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

partition_owned: losing a group session crashes the worker while dropping its partitions (DuckDB tuple-deletion conflict)

1 participant