Skip to content

batchTriggerAndWait parent never resumes after its batch is COMPLETED (follow-up to #4971) #4988

Description

@devin-ai-integration

Provide environment information

  • Trigger.dev Cloud, project proj_puwcdoxtxcywuaznmabj, prod env
  • @trigger.dev/sdk 4.5.14, Node 22
  • Deployed versions: 20260922.2, 20260927.69, 20260929.11

Describe the bug

Follow-up to #4971, which was closed on 2026-09-28 with no comment. The bug is still happening, most recently on 2026-09-29 at 02:14 UTC.

A task awaiting batchTriggerAndWait is never resumed, even though every child run is COMPLETED and GET /api/v1/batches/{id} returns status: "COMPLETED" with all run ids listed. This is shape 2 from #4971. The parent just sits there (no attempts, durationMs: 0, finishedAt: null) until something cancels it. We now run a watchdog that cancels a parent if it hasn't resumed 10 minutes after its last child finished. Every run below was canceled by that watchdog, not resumed by Trigger.

Parent run (scheduled-metrics-orchestration) Version Batch runCount Last child finishedAt Batch updatedAt / status Resumed?
run_06gchv0ng30rdpg14g75f9gg01 20260922.2 batch_06gchv1gfi8jslsjv421mdcg01 496 2026-09-22T12:02:15Z 12:02:15Z / COMPLETED No, canceled 12:15Z
run_06ge9aup7qou37gk70528bhb01 20260927.69 batch_06ge9d3c2s0u2sf4fo8njd0a01 519 2026-09-27T21:15:14Z 21:15:14Z / COMPLETED No, canceled 21:25Z
run_06gelqcci5mtosn4hkkachci01 20260929.11 batch_06gelresnakoaisb5uabtkje01 57 2026-09-29T02:14:10Z 02:14:10Z / COMPLETED No, canceled ~02:25Z

In each wave, only one parent gets stuck. Its siblings (25 to 29 other runs of the same task in the same wave, running the same code) resume within seconds of their batches completing. The 2026-09-29 batch had only 57 items, so this doesn't look like it depends on batch size.

The stuck parents are themselves children of a batchTriggerAndWait from the cron task social-metrics-collection (for example, root run run_06gelou4op6gt2avuuvrns0501).

Reproduction

We can't reproduce it on demand. It happens intermittently in production, roughly once every few days out of a cron that runs every 5 minutes. The pattern is a nested await childTask.batchTriggerAndWait(items): a cron triggers 20 to 30 parents, and each parent triggers 50 to 600 children. Every item has an idempotency key and a small JSON payload. The parent has maxDuration: 300 and the child has maxDuration: 600, and each is on its own queue with a concurrency limit.

Expected behavior

When a batch is COMPLETED, the parent's waitpoint should be completed and the parent should resume. If that can't happen, the parent should fail with an error instead of waiting forever.

Additional information

We can share more run ids or give you dashboard access to the project if that helps.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions