This file is automatically loaded by odek when running inside this repository. It provides context about the project's architecture, conventions, and how to update/maintain it.
- Package:
odek(Go module:github.com/BackendStack21/odek) - What it is: Minimal Go autonomous agent runtime — ReAct (Reasoning + Acting) loop with zero frameworks (stdlib + a few focused packages).
- Binary:
odek— single static binary, <15 MB, instant startup. - Config: Five-layer priority:
~/.odek/secrets.env→~/.odek/config.json→./odek.json→ODEK_*env vars → CLI flags. - Extension contract:
odek-extension/v1(docs/EXTENSIONS.md) — MCP server limits, artifact refs, runtime events, external session refs, execution budgets. - Releases: tag-driven (
git tag vX.Y.Z && git push --tagsbuilds binaries + release). See GitHub Releases.
odek.go Compatibility facade for the public Go API (type aliases + forwarding functions)
cmd/odek/
main.go CLI entry point, flag parsing, commands, sandbox setup, system prompt,
--events-jsonl/--external-ref/budget flag wiring, init config templates
dispatch.go CLI subcommand dispatch (+ budget.Error → exit code 4 mapping)
shell.go Built-in shell tool (local or docker exec; danger-gated; optional timeout_seconds)
serve.go Web UI server (HTTP + WebSocket; @-resource completion; protocol v2: delta
streaming, ping/pong, WS cancel, session_switch)
serve_api.go REST management API (/api/health, sessions search/pagination/pin/export,
memory facts + episode promote/discard + consolidate, skills + promote, tools,
profiles, sanitized config view, MCP listing, shutdown)
serve_runs.go Headless REST runs (POST /api/prompt → run registry, remote approval
bridge, cancel) + events ring (/api/events), usage stats (/api/usage),
WS connection registry (/api/connections, kick)
serve_jobs.go Serve-side job registry API (bg jobs, /api/jobs) + wake-on-complete wiring
serve_supervision.go WebUI task supervision: sub-agent state frames, cancel, safe recovery
serve_artifacts.go Artifact registry + delivery over the serve surface
serve_workspace.go WebUI workspace/session-state access endpoints
serve_diagnostics.go Diagnostics endpoints (/api/diagnostics)
serve_retention.go Session/artifact retention enforcement on the serve surface
serve_bg_sandbox.go Sandbox follow-up handling for background jobs (pidfile kill)
repl.go Interactive REPL with multi-turn session support
repl_editor.go Terminal raw-mode input editor
telegram.go Telegram bot command — wires odek agent into Telegram poller
telegram_voice.go Telegram voice messages: STT transcription of audio notes
telegram_vision.go Telegram vision: photo/image input routing
subagent.go Sub-agent command (--goal, --context, --task) + flag parsing/limits
subagent_tool.go delegate_tasks built-in tool (sub-agent spawning)
subagent_key.go FD-based API key handoff (parent → sub-agent, never via env)
subagent_registry.go Parent-side registry of in-flight sub-agent tasks (cancel, status)
subagent_artifacts.go Sub-agent result-artifact protocol (files auto-delivered to the parent)
subagent_artifact_registry.go Collation-time artifact-id registry backing artifact_read
subagent_accounting.go Sub-agent token/cost accounting rollup into the parent session
artifact_read_tool.go Parent-only artifact_read tool — resolves artifact ids to validated
content; the model supplies only the id, never a filesystem path
browser_tool.go Built-in browser tool (HTTP fetch + headless navigation)
file_tool.go Built-in file tools (read_file, write_file, search_files, patch, glob, file_info)
perf_tools.go Native utilities (math_eval, diff, json_query, tree, checksum, head_tail, base64)
http_request_tool.go Single-URL HTTP status/size checks with SSRF and redirect guards
introspect.go config_view + list_tools tools and shared sanitized view builders
(structural sanitization: secrets never enter the view map)
profiles_tool.go list_subagent_profiles tool (capability profile discovery)
speak_tool.go speak tool (TTS via configured speech provider)
speech_provider.go Speech provider resolution (provider-backed TTS/STT)
proactive.go Proactive engagement presentation: return-after-break summary injection
and follow-up suggestions (presentation-only, never persisted)
iteration_progress.go Iteration-progress signal plumbing for clients
native_outcomes.go Native-tool outcome recording (feeds plan-check evidence)
bg_tools.go Background command tools (bg_*) + per-surface runtime
bg_wake.go Wake-on-complete: system-initiated turns when jobs finish (serve surface)
bg_telegram.go Telegram background-job integration
bg_telegram_wake.go Wake-on-complete for idle Telegram chats (coalesced, rate-bounded)
wsapprover.go WebSocket interactive approval relay (with friction + class-trust gates)
refs.go @-resource reference resolution (files, sessions)
untrusted.go <untrusted·content_<nonce>> wrapper + per-call ingest recorder
unreadscan.go Pre-exec injection-scan enrichment for unread-script approvals
audit.go Per-turn audit + `odek audit` subcommand (divergence heuristic)
sandbox_file.go Sandbox-aware file-tool bridge
sandbox_cleanup.go Sandbox teardown/cleanup helpers
ssrf_guard.go URL / SSRF validation helpers
vision_tool.go Vision / image-input tool
vision_provider.go Provider-backed vision backend (reuses the providers.<id> registry)
transcribe_tool.go Whisper.cpp audio transcription (local mode)
web_search_tool.go Web search tool
session_search_tool.go Session search tool
mcp.go MCP server implementation (stdio transport)
mcp_approval.go Per-tool MCP server approval UI and persistence (key hashes limits/artifact_roots)
project_sandbox_approval.go Project-level sandbox config approval gate
skill_promote.go `odek skill promote` — clear NeedsReview on a tainted skill
schedule.go `odek schedule` command and scheduler wiring
schedule_telegram.go Telegram surface for schedules
memory_cmd.go `odek memory` command
cleanup.go `odek cleanup [--dry-run]` one-shot storage sweep + janitor wiring (telegram/serve/schedule daemon)
upgrade.go `odek upgrade [--check]` self-upgrade from GitHub Releases (checksums.txt SHA-256 verified)
external_ref.go --external-ref flag parsing (run + continue) → session.ExternalRef
native_tool_context.go Invocation-scoped contexts for independent native tool calls
toolctx.go Tool-call context plumbing
logs.go / operational_logging.go / surface_logging.go Structured operational logging per surface
security_report_validation_test.go Regression bar for every documented mitigation
*_test.go Unit + E2E tests covering all tools
internal/
agent/ Public runtime implementation and its package-local tests
llmclient/ Adapter over go-llm-sdk (DTO mapping, temperature polarity, SimpleCall, SideCall)
loop/ ReAct engine: observe → think → parallel-act → repeat. plan.go + plan_*.go —
built-in plan tool and check lifecycle; verify.go — final-answer verification
pass; completion_checks.go — engine-outcome check recording (never claims
from tool output); effects.go + reconcile.go — effect evidence and
post-batch reconciliation; signal.go — SignalEvent observability
(context_trimmed, tool_recovery, tool_running heartbeat); budget enforcement
(budget.Checker) + odek.event/v1 emission.
tool/ Thread-safe tool registry, clarify.go, send_message.go
danger/ Command/URL classification + bypass-resistant tokenizer. Approver interface +
TTYApprover with friction mode (interactive approval system lives here).
bgproc/ Session-scoped background process manager (bg_* tools): bounded output rings,
spawn-time danger classification parity, group-signal stop
diagnostics/ Runtime diagnostics helpers
embedding/ Embedding + featurization for semantic session search
eval/ Plan-tool / behavioral evaluation harness
guard/ Content-scope injection scanner (guard.ScanContentWithScope) + PIGuard sidecar integration
logquery/ Structured query over runtime logs
memory/ MemoryManager (facts, buffer, episodes, merge, scan). EpisodeProvenance — tainted episodes never auto-replayed.
session/ Session store (CRUD, trim, cleanup, compact JSON). AuditStore + divergence heuristic.
ExternalRef — opaque operator-supplied refs, never dereferenced.
artifact/ odek.artifact-ref/v1 + odek.tool-result/v1: fail-closed ref validation
(roots/symlinks/sha256/size) and model-safe rendering (metadata only, no paths/content).
events/ odek.event/v1 runtime event stream: Event, non-blocking panic-isolated Emitter
(args hashed, redact applied), JSONLSink (0600, no symlinks, flush per event).
budget/ Hard execution budgets: Limits (+ model_prices resolution), typed Error, per-run Checker.
maintenance/ Storage-maintenance janitor (session/audit/plan retention, log rotation, media sweep).
Config: `maintenance` section (operator-only).
skills/ Skill system (types, loader, triggers, import, cache). SkillProvenance gate.
config/ Config file loading, env vars, secrets.env, priority merge, limits clamp (project may only lower budgets)
telegram/ Telegram bot: bot.go, poller.go, handler.go, commands.go, session.go, health.go, plan.go, media_path.go
render/ Terminal output and narrator support
narrate/ LLM-powered emoji-rich progress messages
redact/ Secret redaction (20+ patterns)
mcp/ MCP server handler (stdio JSON-RPC: tools/list, tools/call)
mcpclient/ MCP client (connect to external MCP servers); per-server limits (timeout/response bytes/result
chars/artifact roots) + odek-extension/v1 contract (contract.go), artifact-ref enforcement in CallTool
sandbox/ Docker sandbox lifecycle
flock/ Advisory file-locking helpers
fsatomic/ Atomic file-write helpers
pathutil/ Path helpers
resource/ @-resource resolver (files, sessions) with size/symlink hardening
runtimelog/ Runtime log store (replaces the retired ~/.odek/serve.log) with retention
schedule/ Cron-style scheduler: cronexpr, store, scheduler (global-only surface)
transport/ Shared HTTP transport with connection pooling
ws/ RFC 6455 WebSocket framing
docs/ Documentation (CLI, API, CONFIG, MCP, EXTENSIONS, MEMORY, PLANNING, TELEGRAM, SECURITY, etc.)
ReAct cycle: observe → think → act → repeat.
- LLM returns tool calls or a final answer.
- Parallel tool execution — independent tool calls run concurrently (
max_tool_parallel, default: 4). - Plan tool — built-in
plantool (docs/PLANNING.md) for multi-step work: steps, optional acceptance checks, revise/check_replace lifecycle. Sub-agents never own or mutate parent plans. Check recording (completion_checks.go) consumes engine outcomes only — a check passes when its declared tool actually ran and succeeded, never from claims inside tool output; a failed or unclassifiable effect invalidates earlier checks in the same batch. - Batch approval gate — multiple risky tools shown in one prompt.
classifyToolCallclassifies each individualshellcommand, eachpatch/write_filetarget, and thebrowsertool; shows full commands; withholds blanketSetTrustAllwhen unclassifiable tools (incl. MCP tools, classifiedunknown) remain. - Tool-failure recovery — retry transient errors, skip permanent failures, continue without crashing. Stall detection: 3 consecutive identical tool calls inject a corrective hint + fire
tool_recovery— a hint, never aborts the run. - Context-limit protection — graduated trimming near the context window: old large tool results replaced with markers (4 most recent kept intact), then oldest turn groups dropped atomically (tool messages stay grouped with their parent). The protected head (system prompt, memory block, compaction digest, original task) is never dropped. Token estimator counts tool schemas + reasoning content; safety margin self-tightens when provider-reported tokens exceed estimates (
margin_calibratedsignal).trimToSurvivalhandles provider context-length errors. Rolling compaction (on by default;compaction: false/ODEK_COMPACTION=false/--no-compaction) sketches dropped groups extractively into a rolling digest immediately, then a thinking-off side call replaces that sketch with a model digest on a later iteration if it arrives. - Final-answer verification pass (
internal/loop/verify.go, optional) — after a candidate final answer, a side-call checks the answer against the run's evidence and can trigger corrective cycles. Config:verifysection (enabled,modehint|strict|off, optional cheapermodel,max_cycles≤ 3,max_tokens); default off,hintwhen enabled. Sub-agents are opted in only via an explicitsubagent.verifysection —delegate_tasksnever enables it implicitly, so N children cannot multiply verify cost. - Interaction modes — engaging (narrated), enhance (persistent), verbose (raw), off.
- Max 90 iterations by default. On iteration-budget exhaustion the engine makes one final tool-less LLM call for a partial-progress summary (30s bound), returned marked
[Iteration budget reached — partial summary]. - Post-response async processing — episode extraction and extended-memory extraction run in background goroutines;
Agent.Closedrains them (~15s bound). Pending memory episodes can be promoted or discarded from the WebUI and REST API. - Per-turn session persistence —
SetMessagesPersistCallbackfires after each completed step with a fresh snapshot; CLI/REPL/serve/Telegram wire it toStore.SaveNoIndex(atomic, skips remote vector indexing). Interrupted runs resume viaodek continuefrom the last completed step; error paths persist partial history minus dangling tool calls. The final save per turn still updates the semantic index once. - Storage maintenance janitor —
maintenance.Startsweeps~/.odek(retention, log rotation, media sweep) insideodek telegram/serve/schedule daemon;odek cleanup [--dry-run]runs it on demand. Session files are trimmed at write time when they would exceedMaxSessionFileBytes. Operator-only config. See docs/MAINTENANCE.md. - Artifact-aware file search —
search_filesskipnode_modules,vendor,.git,__pycache__,.venv, etc. - Semantic session search —
session_searchtool: go-vector RandomProjections + k-NN, two-tier (vector index → exhaustive fallback). - Background commands —
bg_start/bg_list/bg_status/bg_output/bg_stoptools over a process-scoped, session-keyed process manager (internal/bgproc): session-scoped job lifetime, bounded in-memory output rings, spawn-time danger classification parity withshell, group-signal stop (sandbox mode via the pidfile follow-up). Wake-on-complete (default on;background.wake_on_complete, coalesced bywake_coalesce_ms, bounded bymax_wakes_per_hour): when a job finishes while its session is idle, a system-initiated turn lets the model fetchbg_outputand report unprompted. Config:backgroundsection (docs/CONFIG.md).
- MCP per-server limits —
timeout_seconds(30s/3600s cap),max_response_bytes(10 MiB/64 MiB ceiling),max_result_chars(200k/1M cap, structured truncation notice),artifact_roots(empty ⇒ refs rejected). Resolved per client; approval keys hash all four fields. Per-serverenabledtoggle (nil/true = enabled, back-compat). - Artifact references — MCP tools return
odek.tool-result/v1envelopes withfile://refs instead of bulk content; validated fail-closed ininternal/artifact; model sees metadata only. The parent reads full content via the parent-onlyartifact_readtool (id-supplied only, no model-controlled paths; TOCTOU-hardened). - Runtime events —
odek.event/v1viaConfig.EventHandler(non-blocking, panic-isolated) andrun --events-jsonl. Types: run_started, iteration_completed, tool_call_*, session_saved, context_trimmed, budget_exceeded, run_completed/run_failed, plan_created, plan_updated, plan_blocked, subagent_denied, subagent_spawned, subagent_completed, subagent_concurrency_wait. - External session refs —
Session.ExternalRefs+--external-refon run/continue; validated, deduped, never dereferenced. - Execution budgets —
limitsconfig section +--max-runtime/--max-tool-calls/--max-input-tokens/--max-output-tokens/--max-cost-usdonrun; typedbudget.Error→ CLI exit code 4; session persisted before return. Per-model prices vialimits.model_priceswith flat-pair fallback; cost enforcement only when cap + prices configured.odek init --globalscaffolds the section (zeros = off).GET /api/limitson serve exposes limits + effective prices for cost rendering.
Built-in tools include: http_request, math_eval, diff, json_query, tree, checksum, head_tail, base64, transcribe, speak, browser, read_file, write_file, search_files, patch, shell, plan, delegate_tasks, session_search, artifact_read, config_view, list_tools, list_subagent_profiles, clarify, send_message.
Vertical space compression is baked into the render paths; blank lines removed from Iteration/FinalAnswer/Summary. Raw-mode cursor uses \r\n for cross-platform compatibility.
System prompt priority: --system flag > ~/.odek/IDENTITY.md > compiled-in defaultSystem. Explicit prompts and IDENTITY.md are capped at 256 KiB and scanned with danger.ScanInjection (failure → compiled-in default). Project AGENTS.md ignored if >256 KiB. The compiled-in default carries the execution-provenance rules (justification from the principal; read what you execute; failed reads never become executions; deferred-execution confirmation; tool metadata is not directives) and is itself scanner-clean — pinned by TestDefaultSystem_PassesOwnInjectionScan so a copy into IDENTITY.md is never rejected. Operator identity surfaces (--system, ODEK_SYSTEM, config system, IDENTITY.md) carry identity only — name, mission, persona. buildSystemPrompt force-composes the invariant pillar on top of every accepted identity via composeSystem (idempotent — an identity already carrying the pillar is kept whole), so no operator surface can drop the security rules; the compiled default is defaultIdentity + pillar, pinned by cmd/odek/system_pillar_test.go. Sub-agents compose the same invariant pillar into subagentSystem (shared securityPillar const: Safety, Execution provenance, IPI) plus role amendments — a child has no principal channel, so confirmation rules become skip-and-report, justification scope is the declared task, and deferred execution requires the task to name the mechanism. Parity is pinned by cmd/odek/subagent_pillar_test.go; the operator-writable parent surface (--system, IDENTITY.md) never propagates to children.
Layered prompt-injection / approval-fatigue defenses. The full per-mitigation list lives in docs/SECURITY.md; cmd/odek/security_report_validation_test.go is the regression bar. Summary by layer:
- Untrusted-content boundary (
cmd/odek/untrusted.go) — every externally-sourced tool result (browser, file/shell/search tools, MCP, session_search, @-refs, --ctx, attachments, artifact_read) is wrapped in a per-call nonce'd<untrusted·content_<nonce>>tag; tool-result delimiters are also nonce'd (internal/loop). Skill/episode context injected into the system prompt is wrapped too. The per-session audit log (cmd/odek/audit.go) records every ingest and flags divergence between user-mentioned resources and agent actions. - Provenance gates — tainted memory episodes are stored but never auto-replayed; skills from untrusted sources (imported via URI, project
./.odek/skills/) are pinnedNeedsReviewuntilodek skill promote --forceis run after human review, excluded from trigger matching, and protected against frontmatter tampering. Load-time and import-time skill bodies go through the injection scanner (guard.ScanContentWithScope).odekself-invocation via shell issystem_writeso the agent can't reach its own trust mutations. - Danger classifier (
internal/danger/classifier.go) — bypass-resistant normalization ($IFS, command substitution, wrappers, backslashes, basenames); covers awk/sed/editor escapes, pipe-fed xargs composition, root-level mutation targets, git data-loss verbs,ghas network egress,git -c/config code exec, find/rsync destructive flags, env dumps, shell operand/redirect path classification (writes to shell rc files, ~/.ssh, ~/.odek escalate to system_write). Trust anchors under~/.odekare write-protected from generic file tools. The read ledger is fingerprinted (WasReadFresh: post-read mutation re-fires the unread-script gate), and unread-script approvals carry a pre-exec injection-scan enrichment incl. single-layer base64/hex decode (cmd/odek/unreadscan.go, scan never populates the ledger). - Plan-check honesty — plan acceptance checks record outcomes from actual engine observation only; failed or unclassifiable tool effects invalidate earlier checks in the same batch, so a check can never be satisfied by a claim inside tool output.
- Approval friction — TTY/WS/Telegram approvers engage friction after 3 same-class approvals in 60s (type
approve, pause, trust shortcut hidden);destructive/blocked/unknownnever get trust shortcuts. TTY prompts are process-wide serialized. - Sub-agent caps —
delegate_taskscarries trust_level + max_risk enforced via the sub-agent's DangerousConfig; MCP tools withheld from untrusted sub-agents; API keys handed off via unlinked-tempfile FD, never env. Sub-agent results over ~2000 chars are delivered as registry-backed artifacts the parent reads viaartifact_read. - MCP hardening — subprocess env sanitization (secret-pattern stripping), tool-name/description/inputSchema validation + injection scans, per-tool approval for every server (keys hash command/args/env + schema hash + description text + all four limit fields), per-server limits with absolute ceilings, per-server
enabledtoggle, artifact-ref fail-closed validation. - Config trust split —
./odek.jsonis untrusted: sensitive sections (provider, providers, llm, base_url, api_key, system, dangerous, memory, telegram, web_search, embedding, sessions, skills.dirs, verify, profiles, guard, schedules) ignored with warnings; sandbox knobs gated behind explicit operator approval (incl. implicitDockerfile.odekbuilds, content-hash keyed); project limits may only lower global budgets, project prices rejected outright. Global config/secrets permission-checked; config files size-capped. - Serve / network surface — per-instance CSRF token on
/wsand all/api/*, loopback Host checks, local-origin requirement for mutations, per-session auth tokens + rate limiting, clickjacking headers, WS message-size caps. SSRF dial guard (DNS-rebinding-safe, internal-IP refusal, proxy refusal) on browser/http_request/web_search. WebUI done-frame stats are markup-escape hardened; WebUI task supervision renders sub-agent state without trusting frame content. - Budgets, events, refs (v1.24.0) — budget clamp merge (see above); event stream carries SHA-256 arg hashes + sizes only (never raw args), redact applied, JSONL sink 0600/no-symlink/fsync-per-event, drop-on-full dispatch; external refs validated and never dereferenced.
- Resource bounds — pervasive size caps (shell output 1 MiB/stream, perf-tool files 10 MiB, session files 32 MiB, skill files 1 MiB, browser snapshots/history/elements, tree width, search results, write_file content, patch expansion) to keep hostile input from OOMing the process.
- Telegram — chat-scoped sessions/plans/media, callback binding to originating user, outbound media allowlist + approval, secret-subtree rejection, singleton flock, 0600 logs.
- Redaction —
internal/redact(20+ patterns: provider keys, cloud creds, PEM, JWT, DB URLs, …) applied to sessions, logs, and the event stream.
sec_findings.md at the repository root is the running security audit log. It is
intentionally listed in .gitignore so that audit output and in-progress
findings are not committed to the repository by default. Do not commit this
file in pull requests unless you explicitly intend to publish a finalized
audit snapshot.
CLI, REPL, Web UI, Telegram bot — all in a single binary.
# All unit tests
go test ./... -count=1
# Race detector
go test -race ./... -count=1
# E2E tests (builds odek binary, tests real subprocess spawning)
ODEK_E2E=true go test -v -count=1 ./cmd/odek/ -run "TestE2E_"
# MCP E2E tests (builds fakeserver from source at test time)
ODEK_E2E=true go test -v -count=1 ./cmd/odek/ -run "TestMCPClientE2E_"
# Sandbox integration tests (requires Docker)
go test -v -count=1 ./cmd/odek/ -run "TestSandbox"
# PIGuard sidecar E2E (local only — too heavy for CI; provisions the
# docker stack, runs the env-gated test, tears down. Use --linux on macOS
# for full socket-mode coverage.)
docker/piguard-e2e.sh
# Fuzz soaks
go test -fuzz=FuzzExtractJSON -fuzztime=30s ./internal/memory/extended/
go test -fuzz=FuzzParseSkillContent -fuzztime=30s ./internal/skills/
go test -fuzz=FuzzSessionLoad -fuzztime=30s ./internal/session/
go test -fuzz=FuzzParseEnvelope -fuzztime=30s ./internal/artifact/ # + FuzzValidateRef
go test -fuzz=FuzzEventJSON -fuzztime=30s ./internal/events/
go test -fuzz=FuzzExternalRefValidate -fuzztime=30s ./internal/session/
go test -fuzz=FuzzParseExternalRefFlag -fuzztime=30s ./cmd/odek/
go test -fuzz=FuzzClampProjectLimits -fuzztime=30s ./internal/config/CI also runs golangci-lint (staticcheck) and govulncheck on every push/PR — run both locally before pushing.
Note: MCP client E2E tests build the fakeserver from internal/mcpclient/testdata/main.go at test time (the extension mock is env-gated via FAKE_ARTIFACT_MODE=1). macOS temp dirs are classified as LocalWrite (not SystemWrite), and the Docker availability check verifies daemon reachability (5s timeout) before running sandbox tests.
Agent workflow notes:
- Scope test runs. Prefer
go test -count=1 <changed packages>over./...; run-raceonly on changed packages. A fullgo test -race ./...takes several minutes (race-instrumented build + 5–10× runtime slowdown). - Sandbox test runs use
make test-sandbox. Inside the odek sandbox (Dockerfile.odekimage),/tmpis mounted noexec and the workspace mount does not enforce write bits, so tests must run viamake test-sandbox— it splits TempDir by need (exec-capableGOTMPDIR/TMPDIRfor the cmd/odek mocks and the mcpclient fakeserver; the permission-enforcing default for flock/maintenance). Host runs keepmake test/make test-cmd. - Comments name behavior, not plan tickets. Do not put
Slice C,M1.2,P0-1,WP6,H-9, or similar milestone/review IDs in comments or product docs. Describe what the code does. Runtime sequence labels (Phase 1: execute tools) are fine. - The
shelltool is fully buffered. Nothing is shown or returned until the command exits, and the default timeout is 30 minutes — a long command looks "stuck" even though it is running (atool_runningheartbeat signal fires every 60s). Settimeout_secondsexplicitly for known long-running commands (builds, test suites) so a genuinely stuck command fails fast, and passgo test -timeoutfor test runs.