Skip to content

Commit d7e3e0e

Browse files
cardmagicclaude
andcommitted
docs: describe retry and redrive
The roadmap still named the gaps this branch closes: no retry for a dead effect or broadcast, one row at a time, and no record of who pressed what. It now states what exists and what still does not, which is the dashboard surface for the new scopes. Operations gains the API, the idempotency rule, the batching, the authorization resource names, and the audit row. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent 1e9d9cc commit d7e3e0e

4 files changed

Lines changed: 121 additions & 9 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 22 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,28 @@
22

33
## Unreleased
44

5+
- Retry a dead effect or broadcast. `SolidObjects.dead_letters` keeps its
6+
message meaning and answers `effects` and `broadcasts`, so the kind rides on
7+
the receiver. `retry` returns a dead row to pending with a zero attempt count
8+
and no claim, reuses the stable id so a deduplicating handler sees the same
9+
key, and acts only on a dead row, so a second press cannot double-enqueue. A
10+
dead transmit effect replays rather than stay lost.
11+
- Redrive a whole scope. `redrive` opens a durable task, returns at once, and is
12+
idempotent over its scope and filters, which a dashboard button needs. A
13+
unique index on the active scope enforces that in the database, so two
14+
processes that start the same redrive share one task. The supervisor advances
15+
one bounded batch per pass, so a redrive never holds a transaction longer than
16+
one batch. `SolidObjects.redrives` reads tasks back, and `task.cancel` stops
17+
one and leaves the rows it already moved.
18+
- Record who pressed what. Every retry and every redrive transition writes one
19+
row to `solid_objects_administration_events`. The identity comes from the
20+
authorization context through a new `administration_identity` hook.
21+
- Add `redrive_batch_size`, which defaults to 100, and `redrive_batch_pause`,
22+
which defaults to 0.05 seconds.
23+
- Add two tables, `solid_objects_administration_events` and
24+
`solid_objects_redrives`. Run `bin/rails solid_objects:install:migrations` and
25+
migrate.
26+
527
- Select a wake-up adapter automatically. `config.wake_up_adapter` now takes a
628
name or an adapter, as `config.cache_store` and
729
`config.active_job.queue_adapter` do, and defaults to `:automatic`. Selection

‎docs/dashboard.md‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -149,6 +149,16 @@ actor class that no longer exists, a full mailbox, a payload over the cap. The
149149
dashboard renders the dead letter again with the reason and a 422 status,
150150
rather than failing the request.
151151

152+
Dead effects and broadcasts have the same API, which the dashboard does not yet
153+
surface. `SolidObjects.dead_letters.effects` and
154+
`SolidObjects.dead_letters.broadcasts` read and retry their own kind, and
155+
`redrive` moves a whole scope as a durable task. See
156+
[Operations](operations.md) for both.
157+
158+
Every retry and every redrive transition writes one row to
159+
`solid_objects_administration_events`, holding the action, the kind, the
160+
subject, and the identity that asked for it.
161+
152162
**Pause an instance** sets `paused_at`, and the activation manager stops
153163
claiming that identity. Two consequences matter:
154164

‎docs/operations.md‎

Lines changed: 77 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -282,6 +282,83 @@ interval, and current interval. The polling-only warning is also emitted as
282282
adapters should return `true` for a notification and `false` for a timeout; an
283283
older adapter that returns `nil` remains compatible and keeps the fast cadence.
284284

285+
## Dead letters, retry, and redrive
286+
287+
A message that exhausts its attempts becomes a dead letter. An effect or a
288+
broadcast that exhausts its attempts stays in its own table with
289+
`status = 'dead'`. All three are read and retried through one receiver, which
290+
carries the kind:
291+
292+
```ruby
293+
SolidObjects.dead_letters.all(authorization_context: current_admin)
294+
SolidObjects.dead_letters.retry(dead_letter_id, authorization_context: current_admin)
295+
296+
SolidObjects.dead_letters.effects.all(authorization_context: current_admin)
297+
SolidObjects.dead_letters.effects.retry(effect_id, authorization_context: current_admin)
298+
SolidObjects.dead_letters.broadcasts.retry(broadcast_id, authorization_context: current_admin)
299+
```
300+
301+
An effect or broadcast retry returns the row to pending with a zero attempt
302+
count, no claim, and immediate availability. It keeps the stable id, so a
303+
handler that deduplicates on `effect_id` still sees the same key. An effect is
304+
at-least-once by contract, so a retried effect can run twice.
305+
306+
Retry acts only on a dead row. A row that is pending, processing, or completed
307+
comes back unchanged, so pressing a button twice cannot double-enqueue and
308+
cannot take a row away from a worker that holds it.
309+
310+
An incident produces dead rows in the hundreds, so a scope also answers
311+
`redrive`:
312+
313+
```ruby
314+
task = SolidObjects.dead_letters.effects.redrive(
315+
actor_type: "payments",
316+
failed_after: 6.hours.ago,
317+
limit: 5_000,
318+
authorization_context: current_admin
319+
)
320+
321+
task.id # => "redrive_..."
322+
task.status # => "running"
323+
task.moved # => 412
324+
task.remaining # => 4_588
325+
326+
task.cancel(authorization_context: current_admin)
327+
```
328+
329+
`redrive` returns at once. The task is durable, and the supervisor advances one
330+
bounded batch per pass, so a redrive of thousands of rows never holds a
331+
transaction longer than one batch. `redrive_batch_size` defaults to 100 and
332+
`redrive_batch_pause` to 0.05 seconds.
333+
334+
A redrive is idempotent over its scope and its filters. Starting the same one
335+
while it runs returns the running task rather than a second one, which a
336+
dashboard button an operator can press twice needs. A different scope or a
337+
different filter starts its own task, and the same scope can be redriven again
338+
once the first task finishes.
339+
340+
Read tasks back with `SolidObjects.redrives`:
341+
342+
```ruby
343+
SolidObjects.redrives.find(task.id, authorization_context: current_admin)
344+
SolidObjects.redrives.all(status: :running, authorization_context: current_admin)
345+
```
346+
347+
A running task reports what is left to move rather than a stored estimate,
348+
because rows die and are retried while it runs.
349+
350+
Retry, redrive, and cancel each go through `authorize_administration` under
351+
their own resource name: `dead_letters`, `effect_dead_letters`,
352+
`broadcast_dead_letters`, and `redrives`. Every retry and every task transition
353+
writes one row to `solid_objects_administration_events`, holding the action, the
354+
kind, the subject, the identity, and when it happened. The identity comes from
355+
`administration_identity`, which receives the authorization context the caller
356+
passed and defaults to its `to_s`. A refused caller writes nothing.
357+
358+
Automatic redrive on a schedule is deliberately absent. A dead row means a
359+
person decided something, and these APIs give that person an alternative to an
360+
`UPDATE` against a runtime table.
361+
285362
## Graceful shutdown
286363

287364
The supervisor requests shutdown, stops new claims, lets active loops return,

‎docs/roadmap.md‎

Lines changed: 12 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -178,15 +178,18 @@
178178
or turn off. Every route declares its
179179
own administration policy and a route declared without one raises at load
180180
time, so the deny-by-default posture is enforced by construction rather than
181-
by remembering to add a check. It changes only two things: an idempotent dead
182-
letter retry and instance pause/resume. What does not exist is audit records
183-
of who pressed what, and bulk-safe tools: retry is one dead letter at a time,
184-
because `DeadLetterManager` exposes no bulk operation. Pause is an operator
185-
brake and not a stop, since a pass already in flight finishes its turn and a
186-
synchronous caller waiting on a paused instance times out. Retry also only
187-
exists for message dead letters: a dead effect or broadcast has no retry
188-
API, which matters for transmit effects because a dead one is a lost
189-
replay until an operator returns its row to pending. The page cost was
181+
by remembering to add a check. It changes only three things: an idempotent dead
182+
letter retry, a redrive, and instance pause/resume. Retry covers all three
183+
kinds. `SolidObjects.dead_letters` keeps its message meaning and answers
184+
`effects` and `broadcasts`, so a dead effect or broadcast returns to pending
185+
through an API rather than an operator's `UPDATE`, and a dead transmit effect
186+
is no longer a lost replay. `redrive` moves a whole scope as a durable task
187+
that is idempotent over its filters, cancellable, and advanced in bounded
188+
batches by the supervisor. Every retry and task transition writes one row to
189+
`solid_objects_administration_events`, so who pressed what is recorded. Pause
190+
is an operator brake and not a stop, since a pass already in flight finishes
191+
its turn and a synchronous caller waiting on a paused instance times out. The
192+
dashboard does not yet surface the scopes or redrive; the API does. The page cost was
190193
reasoned about rather than measured: the summary bar issues a fixed set of
191194
indexed aggregate queries per page, which is why `HEAD /` exists for uptime
192195
monitors, but no dashboard latency has been benchmarked against a large

0 commit comments

Comments
 (0)