Skip to content

Broker spawn treats Relaycast database_overloaded preregistration as fatal after one attempt #1715

Description

@khaliqgant

Problem

Agent Relay v11.10.4 cannot start a new broker worker when Relaycast agent registration returns database_overloaded. The local broker is healthy and five existing agents remain deliverable, but POST /api/spawn fails before process creation.

Clean reproduction

Environment: running chief-broker, Agent Relay v11.10.4, headless task-exit OpenCode worker, no Daytona sandbox.

Observed across independent paths on 2026-09-09:

  • Direct agent-relay node agent spawn timed out; broker roster confirmed no process.
  • Direct broker POST /api/spawn returned failure: failed to pre-register worker; HTTP 503; code database_overloaded; attempts: 1.
  • Three further bounded broker retries returned the same typed 503; a fourth returned 500 internal_error.
  • Eight explicit agent-relay agent register attempts also failed; no worker token or process was created.
  • Relaycast incident evidence is tracked at GET /v1/agents latency varies 5x on an identical payload (5.3s-28.7s), and intermittently fails relaycast#389.

Contract mismatch

The v11.10.4 source documents retry_agent_registration as up to three retries for transient errors. The HTTP spawn path also has a RetryableExhausted branch intended to continue spawning without preregistration while emitting a warning. Actual database_overloaded responses report attempts: 1 and take the fatal branch, so a transient control-plane write outage prevents even a headless local task from starting.

Expected

  • Classify database_overloaded and equivalent Retry-After 503 registration failures as retryable.
  • Honor Retry-After within a bounded total deadline and perform the documented retry count.
  • When retries exhaust and the requested task can safely run without Relay messaging, continue the local spawn with an explicit preregistration warning, as the source path already intends.
  • Never report success unless the worker process is actually present; preserve clean task-exit cleanup.
  • Add a regression covering node agent.register failure followed by HTTP registration database_overloaded and proving the fallback branch is reached rather than fatal after one attempt.

This blocks using cheap OpenCode workers to supervise the current cleanroom/Fleet qualification campaign.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions