Provide environment information
- Trigger.dev Cloud, project
proj_puwcdoxtxcywuaznmabj, prod env
@trigger.dev/sdk 4.5.14, Node 22
- Deployed versions:
20260922.2, 20260927.69, 20260929.11
Describe the bug
Follow-up to #4971, which was closed on 2026-09-28 with no comment. The bug is still happening, most recently on 2026-09-29 at 02:14 UTC.
A task awaiting batchTriggerAndWait is never resumed, even though every child run is COMPLETED and GET /api/v1/batches/{id} returns status: "COMPLETED" with all run ids listed. This is shape 2 from #4971. The parent just sits there (no attempts, durationMs: 0, finishedAt: null) until something cancels it. We now run a watchdog that cancels a parent if it hasn't resumed 10 minutes after its last child finished. Every run below was canceled by that watchdog, not resumed by Trigger.
Parent run (scheduled-metrics-orchestration) |
Version |
Batch |
runCount |
Last child finishedAt |
Batch updatedAt / status |
Resumed? |
run_06gchv0ng30rdpg14g75f9gg01 |
20260922.2 |
batch_06gchv1gfi8jslsjv421mdcg01 |
496 |
2026-09-22T12:02:15Z |
12:02:15Z / COMPLETED |
No, canceled 12:15Z |
run_06ge9aup7qou37gk70528bhb01 |
20260927.69 |
batch_06ge9d3c2s0u2sf4fo8njd0a01 |
519 |
2026-09-27T21:15:14Z |
21:15:14Z / COMPLETED |
No, canceled 21:25Z |
run_06gelqcci5mtosn4hkkachci01 |
20260929.11 |
batch_06gelresnakoaisb5uabtkje01 |
57 |
2026-09-29T02:14:10Z |
02:14:10Z / COMPLETED |
No, canceled ~02:25Z |
In each wave, only one parent gets stuck. Its siblings (25 to 29 other runs of the same task in the same wave, running the same code) resume within seconds of their batches completing. The 2026-09-29 batch had only 57 items, so this doesn't look like it depends on batch size.
The stuck parents are themselves children of a batchTriggerAndWait from the cron task social-metrics-collection (for example, root run run_06gelou4op6gt2avuuvrns0501).
Reproduction
We can't reproduce it on demand. It happens intermittently in production, roughly once every few days out of a cron that runs every 5 minutes. The pattern is a nested await childTask.batchTriggerAndWait(items): a cron triggers 20 to 30 parents, and each parent triggers 50 to 600 children. Every item has an idempotency key and a small JSON payload. The parent has maxDuration: 300 and the child has maxDuration: 600, and each is on its own queue with a concurrency limit.
Expected behavior
When a batch is COMPLETED, the parent's waitpoint should be completed and the parent should resume. If that can't happen, the parent should fail with an error instead of waiting forever.
Additional information
We can share more run ids or give you dashboard access to the project if that helps.
Provide environment information
proj_puwcdoxtxcywuaznmabj, prod env@trigger.dev/sdk4.5.14, Node 2220260922.2,20260927.69,20260929.11Describe the bug
Follow-up to #4971, which was closed on 2026-09-28 with no comment. The bug is still happening, most recently on 2026-09-29 at 02:14 UTC.
A task awaiting
batchTriggerAndWaitis never resumed, even though every child run isCOMPLETEDandGET /api/v1/batches/{id}returnsstatus: "COMPLETED"with all run ids listed. This is shape 2 from #4971. The parent just sits there (no attempts,durationMs: 0,finishedAt: null) until something cancels it. We now run a watchdog that cancels a parent if it hasn't resumed 10 minutes after its last child finished. Every run below was canceled by that watchdog, not resumed by Trigger.scheduled-metrics-orchestration)finishedAtupdatedAt/ statusrun_06gchv0ng30rdpg14g75f9gg01batch_06gchv1gfi8jslsjv421mdcg01COMPLETEDrun_06ge9aup7qou37gk70528bhb01batch_06ge9d3c2s0u2sf4fo8njd0a01COMPLETEDrun_06gelqcci5mtosn4hkkachci01batch_06gelresnakoaisb5uabtkje01COMPLETEDIn each wave, only one parent gets stuck. Its siblings (25 to 29 other runs of the same task in the same wave, running the same code) resume within seconds of their batches completing. The 2026-09-29 batch had only 57 items, so this doesn't look like it depends on batch size.
The stuck parents are themselves children of a
batchTriggerAndWaitfrom the cron tasksocial-metrics-collection(for example, root runrun_06gelou4op6gt2avuuvrns0501).Reproduction
We can't reproduce it on demand. It happens intermittently in production, roughly once every few days out of a cron that runs every 5 minutes. The pattern is a nested
await childTask.batchTriggerAndWait(items): a cron triggers 20 to 30 parents, and each parent triggers 50 to 600 children. Every item has an idempotency key and a small JSON payload. The parent hasmaxDuration: 300and the child hasmaxDuration: 600, and each is on its own queue with a concurrency limit.Expected behavior
When a batch is
COMPLETED, the parent's waitpoint should be completed and the parent should resume. If that can't happen, the parent should fail with an error instead of waiting forever.Additional information
We can share more run ids or give you dashboard access to the project if that helps.