Last reviewed: 2026-08-31 Review trigger: Backlog priorities or shipped/beta status claims change. Scheduling gate (2026-08-26): DR-006 remains authoritative. EFX-001–003 are in-progress P0
governance-gapremediations on existing seams; no product horizon, distributed effect subsystem, or external adoption work is reopened. Advisor note 2026-08-31 (hypothesis, not authority): H4 extend only if dogfood scheduled for organic events else revert (0 organic ≠ promotion, demo synthetic ≠ C1); EFX promote only on live proof; M4 dogfood is only DR-006 lane that can generate organic events/friction/BG-001; H2/H5/H6/WDH-002/EFX-FUTURE remain Hold — see.omx/artifacts/claude-you-are-an-external-advisor-for-teaagent-a-harness-first-own-2026-08-31T06-14-02-585Z.md.
Prioritized by impact order: security and production risk → core platform capabilities → developer experience and ecosystem.
Tags classify why work exists. Scheduling is governed by DR-006. DR-001 evidence intake was met on 2026-06-22; completing that intake does not authorize new work.
| Tag | Meaning |
|---|---|
friction-driven |
Owner friction log evidence or closed hypothesis promoted after validation |
governance-gap |
Security/audit/approval correctness; not UX speculation |
legacy-competitive |
Pre–2026-06-13 competitive positioning or survey machinery |
owner-override |
Explicit owner direction record or dated override rationale |
harness-migration |
Event spine, thin runner, TASK-006 arc |
| Item | Tag | Scheduling |
|---|---|---|
| CP-4 OpenCode Gap Watch | legacy-competitive |
Monitoring only; escalation → feature sprint blocked without friction evidence |
| CP-6 Community Presence | legacy-competitive |
Hold — external acquisition non-goal |
| M4 cloud/background/control-plane cockpit | legacy-competitive |
Hold except DR-006 carve-out — background lifecycle + operator cockpit eligible under owner-override co-maintainer dogfood; cloud/SaaS/multi-tenant GTM held (T4). Owner 2026-07-22: carve-out remains held with no scheduled dogfood |
| RBAC enforce flip (ADR 0031, 2026-09-12) | governance-gap |
Hold — ADR-0031 review due 2026-09-12 (decision packet evaluation: promote/extend/revert; no premature enforce flip) |
| TASK-006 RunEvent taxonomy + M0 | harness-migration |
Done — ADR-0032 run-event taxonomy; teaagent/runner/_events.py spine + audit dual-write (M0–M7 in work-log) |
| TASK-001 constitution repositioning | owner-override |
Done (Human Review 2026-07-22) — owner ratified harness-first positioning; no README/product-contract changes required |
| TASK-002 docs tiering | harness-migration |
Done — tier column in docs/generated/docs-inventory.md + aging dashboard; check-docs-inventory pre-commit regen |
| TASK-003 test typing pass | harness-migration |
Done — all 586 test files typed (contract/behavior/adversarial/lifecycle) per scripts/audit_test_quality.py |
| TASK-004 flagship tests off deprecated approval | governance-gap |
Done (2026-07-01) — G-P2-2 removed call-id preapproval; flagship uses --approve-scoped; see work-log/task-004-blocked-2026-06-13.md resolution |
| TASK-007 friction log bootstrap | friction-driven |
Met (2026-06-22) — 5/5 owner evidence entries (F2/F3/F6/F7/F8); DR-001 7b satisfied |
| Phase 4–6 Beta (consensus, sandbox, control plane) | owner-override |
Shipped; freeze new platform surface per T3/T4 |
| Competitive refresh release gate | legacy-competitive |
Resolved (DR-006 T5): split gate — CI blocks on main; manual survey quarterly |
| Seven-control-loops SCL follow-on work | legacy-competitive |
Listed implementation rows are historical Complete; follow-on scheduling is held without friction, governance-gap proof, or owner override |
| Community-pain-point CPP follow-on work | legacy-competitive |
Listed implementation rows are historical Complete; follow-on scheduling is held without friction, governance-gap proof, or owner override |
| Consensus validation (ADR-0029) | governance-gap |
Deleted/quarantined (2026-07-22) — owner-ratified Option D executed; no production importers; feature intent and git recovery path preserved in ADR-0029 disposition spec |
| EFX-001 ambiguous mutating-tool dispatch | governance-gap |
In progress (2026-08-25) — runner commits pending_effect before execute; unmatched non-idempotent starts become OUTCOME_UNKNOWN and refuse blind redispatch (tests/test_efx001_interrupted_dispatch.py, tests/acceptance/test_efx_durable_effect_flow.py). Live GitHub/browser/provider proof still required for Complete |
| EFX-002 effect classification and approval escalation | governance-gap |
In progress (2026-08-25) — GitHub/browser mutators are local external_effect; Prompt/read-only/workspace-write fail closed; MCP hints cannot relax policy (tests/test_efx002_effect_classification.py, tests/acceptance/test_efx_durable_effect_flow.py). Live GitHub/browser/provider proof still required for Complete |
| EFX-003 one-time approval identity and consumption | governance-gap |
In progress (2026-08-25) — omitted call IDs include payload digest; one-time grants bind digest and consume at assert_allowed (tests/test_efx003_one_time_approval.py, tests/acceptance/test_efx_durable_effect_flow.py). Live GitHub/browser/provider proof still required for Complete |
| EFX-FUTURE generic external-effect settlement | legacy-competitive / owner-override |
Hold — ADR-0042 remains binding; exactly-once, generic ledger/outbox/reconciliation, fencing, actor supervision, and distributed leases need a new owner promise and sink-enforced evidence |
Items below were deferred at baseline and have since been implemented in-repo.
| Item | Sprint | Files |
|---|---|---|
Key ring CLI support (--oauth-key-ring-file, --oauth-active-kid, fail-closed validation) |
P0-r1 | cli/_mcp_parsers.py, cli/_handlers/_mcp.py |
External state checkpoint store (InMemoryCheckpointStore, SQLiteCheckpointStore, --checkpoint-store CLI) |
P0-r1 | teaagent/checkpoint.py, runner/_core.py, chat_agent.py |
DPoP replay TTL configurable (dpop_replay_ttl + --oauth-dpop-replay-ttl CLI) |
P0-r1 | oauth21/_server.py, cli/_mcp_parsers.py |
OAuth key rotation overlap window (OAuthKeyRing.rotate, key_for_validation, --oauth-rotation-window) |
P0-r2 | oauth21/_store.py, oauth21/_server.py, oauth21/_resource.py |
Code Mode sandbox profile matrix (SandboxProfile enum, default_sandbox, validate_runtime_support) |
P0-r2 | code_mode/_types.py, code_mode/__init__.py |
LLM-as-Judge scoring (JudgeScore, run_eval_with_judge, make_llm_judge_fn) |
P1-r1 | teaagent/eval.py |
AgentCard + InMemoryAgentRegistry + agent card CLI |
P1-r1 | teaagent/agentcard.py, cli/_agent_parsers.py |
Managed runtime interface (ManagedRuntimeAdapter Protocol, ManagedAgentRunner, provider stubs) |
P1-r2 | teaagent/managed_runtime.py |
LATENCY conformance tier (p50/p95 sampling, threshold check, latency_samples/latency_threshold_ms params) |
P1-r2 | llm_conformance/_types.py, llm_conformance/_runner.py |
| SQLiteAgentRegistry + A2ADispatcher (persistent A2A registry, in-process routing) | P1-r2 | teaagent/agentcard.py |
Extended conformance tiers: STREAMING, STRUCTURED_OUTPUT |
P2-r1 | llm_conformance/_types.py, llm_conformance/_runner.py |
OpenAPI 3.1 schema auto-generation from ToolRegistry + workspace openapi CLI |
P2-r1 | teaagent/openapi.py, cli/_misc_parsers.py |
Web audit viewer (AuditViewerServer, HTML/JSON routes, audit serve CLI) |
P2-r2 | teaagent/audit_viewer.py, cli/_handlers/_audit.py |
Schema migration framework (SchemaMigration, SQLiteMigrationStore, MigrationRunner, doctor migration CLI) |
P2-r2 | teaagent/schema_migration.py, cli/_handlers/_doctor.py |
Code Mode kernel sandbox hardening (seccomp_profile, apparmor_profile, selinux_label, oci_runtime on ContainerCodeModeBackend; IsolateCodeModeBackend gVisor wrapper; sandbox_profile_selected/sandbox_violation audit events; profile+audit_logger params on execute_code_mode) |
P0-r3 | code_mode/_types.py, code_mode/_container.py, code_mode/_isolate.py, code_mode/__init__.py |
Cross-host OAuthStore persistence (PostgreSQLOAuthStore with DELETE…RETURNING atomic consume; RedisOAuthStore with Lua-script atomic consume, NX nonce/code saves, configurable key prefix) |
P0-r3 | oauth21/_pg_store.py, oauth21/_redis_store.py, oauth21/__init__.py |
Managed runtime audit events (managed_task_started/completed/failed on ManagedAgentRunner.run; audit_logger+run_id params); tool-context forwarding for Anthropic and OpenAI runtimes |
P1-r3 | teaagent/managed_runtime.py |
Extended conformance tiers: TOOL_CALLING (invokes get_current_time tool, checks tool_calls); SAFETY (API-level block + text refusal taxonomy); LLMToolDefinition, LLMToolCall, SafetyCategory, LLMSafetyBlock types; tool wiring in all three adapters |
P1-r3 | llm/_types.py, llm/_adapters.py, llm/__init__.py, llm_conformance/_types.py, llm_conformance/_runner.py |
A2A HTTP discovery + wire protocol: A2ADiscoveryServer (serves /.well-known/agent.json, handles POST /a2a/task); A2AClient (fetch_card, delegate); FederatedAgentRegistry (TTL-cached pulls from remote endpoints) |
P1-r3 | teaagent/agentcard.py |
GraphQLite production deployment (GraphQLitePersistentStore, GraphQLiteProductionConfig, index strategy, migration integration, graphqlite migrate CLI, production deployment guide) |
P2-r3 | teaagent/graphqlite_production.py, docs/graphqlite-production.md, cli/_handlers/_misc.py, cli/_misc_parsers.py |
Hosted doc site infrastructure (pdoc dependency, scripts/build_docs.py build script, class-level docstrings on core modules) |
P2-r3 | pyproject.toml, scripts/build_docs.py, teaagent/tools.py, teaagent/runner/_core.py, teaagent/budget.py, teaagent/policy.py, teaagent/memory.py |
ANP bidirectional adapter governed federation (ANPGovernedService, audit correlation, approval/budget invariants) |
P1 | teaagent/anp_adapter.py, tests/acceptance/test_anp_adapter_flow.py, docs/adr/0007-anp-adapter-boundary.md |
| Daily-use ergonomics (init, providerless CLI, recipes, sessions, background/attach, model capabilities, KPI) | P1 | teaagent/ergonomics/, teaagent/recipes/, scripts/measure_time_to_first_run.py, examples/ergonomics/ |
mtime read-before-write concurrent modification guard (workspace_write_file rejects overwrites on mtime mismatch; workspace_read_file returns mtime; backward compatible) |
P0 | teaagent/workspace_tools/_files.py, tests/acceptance/test_mtime_read_before_write_flow.py |
Protected paths default deny rules (.git/* and .teaagent/* blocked by default in FilePolicy; load_file_policy(include_protected_dirs=) gating) |
P0 | teaagent/file_policy.py, tests/acceptance/test_protected_paths_flow.py |
| LSP code analysis acceptance (7 tests covering tool registration, tree-sitter relations, candidate path extraction, config enablement) | P0 | tests/acceptance/test_code_analysis_lsp_flow.py |
Declarative sub-agent definitions with Markdown frontmatter (.md file support matching Claude Code .claude/agents/*.md convention; isolation/background/disallowed_tools/effort fields on SubagentDef) |
P1 | teaagent/subagents/_types.py, teaagent/subagents/_loader.py, tests/acceptance/test_subagent_definitions_flow.py |
| Context compaction latency SLO (traffic-light zoning boundary tests; compaction preserves recent observations; latency < 100ms SLO) | P1 | teaagent/context.py, tests/acceptance/test_context_compaction_slo_flow.py |
| Hook lifecycle acceptance elevation (16 tests: PreToolUse veto, PostToolUse chaining, permission_check_hook deny/allow/patterns, all 8 Claude Code events, registry enabled flag) | P1 | tests/acceptance/test_hook_lifecycle_flow.py |
| Architecture comparison matrix documenting TeaAgent vs Claude Code/Codex/OpenCode feature coverage | P2 | docs/architecture.md |
| Use-case matrix updated with LSP, sub-agent orchestration, mtime guard, and protected paths entries | P2 | docs/use-cases.md, docs/use-case-matrix.md |
| 5-Loop Governance Hardening (Tool, Coding, Audit, Memory, Swarm) | P0 | teaagent/tools.py, teaagent/governance/tool_lint.py, teaagent/selftest.py, teaagent/plan_mode.py, teaagent/workflow_engine.py, teaagent/audit.py, teaagent/memory/failure_card.py, teaagent/swarm.py, teaagent/tournament/, docs/architecture.md |
Note: All P0 items below have been implemented with acceptance tests. See docs/specs/ for technical specifications.
Status: ✅ Implemented
Journey: User pastes an issue, gets ambiguity score, plan artifact, safe command, and acceptance checklist.
Acceptance test: test_issue_to_plan_acceptance_flow.py
Dependencies: Plan mode, issue parsing, ambiguity detection, plan generation
Size: Large (2-3 weeks)
Reference: docs/plans/agent-ecosystem-acceptance-roadmap-2026-05-31.md
Files: teaagent/issue_intake.py, tests/test_issue_intake.py, docs/specs/issue-to-plan-intake-spec-2026-05-31.md
Status: ✅ Implemented
Journey: User can compare two plan revisions before execution and bind run to the accepted plan hash.
Acceptance test: test_plan_review_revision_flow.py
Dependencies: Plan versioning, plan diff, run-to-plan binding, hash verification
Size: Large (2-3 weeks)
Reference: docs/plans/agent-ecosystem-acceptance-roadmap-2026-05-31.md
Files: teaagent/plan_storage.py, tests/test_plan_storage.py, docs/specs/plan-review-revision-spec-2026-05-31.md
Status: ✅ Implemented
Journey: Failed/partial run suggests resume, undo, inspect audit, or retry with safer mode.
Acceptance test: test_guided_recovery_flow.py
Dependencies: Failure classification, recovery strategy selection, undo/resume integration
Size: Medium (1-2 weeks)
Reference: docs/plans/agent-ecosystem-acceptance-roadmap-2026-05-31.md
Files: teaagent/guided_recovery.py, tests/test_guided_recovery.py, tests/acceptance/test_guided_recovery_flow.py, docs/specs/guided-recovery-spec-2026-05-31.md
Status: ✅ Implemented
Journey: Completed run emits changed files, commands, tests, approvals, costs, failures, and rollback path.
Acceptance test: test_run_evidence_summary_flow.py
Dependencies: Audit event extraction, evidence bundle serialization, sensitive value redaction
Size: Medium (1-2 weeks)
Reference: docs/plans/agent-ecosystem-acceptance-roadmap-2026-05-31.md
Files: teaagent/run_evidence.py, tests/test_run_evidence.py, tests/acceptance/test_run_evidence_summary_flow.py, docs/specs/run-evidence-bundle-spec-2026-06-01.md
Status: ✅ Implemented
Journey: CLI, TUI, and dashboard expose the same core run, approval, cost, and warning state.
Acceptance test: test_daily_cockpit_parity_flow.py
Dependencies: Cockpit state model, serialization, run store integration
Size: Medium (1-2 weeks)
Reference: docs/plans/agent-ecosystem-acceptance-roadmap-2026-05-31.md
Files: teaagent/cockpit.py, tests/test_cockpit.py, tests/acceptance/test_daily_cockpit_parity_flow.py
Note: Items extracted from strategic reference documents (competitive positioning, UX improvements, future roadmap).
Status: ✅ Complete
Journey: README already leads with governance-first positioning (permission gates, audit trail, undo, cost cap, model-agnostic)
Acceptance criteria: README answers all three persona questions in first screen, governance table in top half, time to comprehension < 30 seconds
Size: Small (3 days)
Reference: docs/plans/competitive-positioning-plan-2026-05-31.md
Status: ✅ Complete
Journey: Security whitepaper exists documenting governance model, control catalog, NIST mapping, data handling, deployment isolation, incident response
Acceptance criteria: Document exists and reviewed by security background, every control has traceable code path, honest "known limitations" section
Size: Medium (1 week)
Reference: docs/plans/competitive-positioning-plan-2026-05-31.md
Status: ✅ Complete
Journey: Surface git commit-per-change workflow prominently (stage in branch, review diff, commit/discard, structured commit messages with run metadata)
Acceptance criteria: teaagent show --diff works without git knowledge, commit message includes run metadata, teaagent commit is idempotent
Size: Medium (1 week)
Reference: docs/plans/competitive-positioning-plan-2026-05-31.md
Status: ✅ Complete
Journey: Emit structured summary at end of every run (tools called, files changed, cost, budget remaining, audit log location, undo command)
Acceptance criteria: Run summary always emitted at session end, --no-summary flag suppresses it, summary includes files changed/cost/undo command
Size: Small (3 days)
Reference: docs/plans/ux-improvement-roadmap-2026-05-31.md
Status: ✅ Complete
Journey: Emit budget warnings at 50%, 80%, 90% (with prompt), 100% (offer read-only mode)
Acceptance criteria: Tests trigger at 50%/80%/90%/100%, TUI shows persistent budget status line, warning in post-run summary
Size: Small (3 days)
Reference: docs/plans/ux-improvement-roadmap-2026-05-31.md
Status: ✅ Complete
Journey: Add --preview flag to undo command, add --last shortcut, show unified diff of what will be reverted
Acceptance criteria: teaagent undo --last --preview shows readable diff without executing, teaagent undo --last reverts all writes, post-run summary includes undo command
Size: Small (3 days)
Reference: docs/plans/ux-improvement-roadmap-2026-05-31.md
Status: ✅ Complete
Journey: Implement guided init flow (provider selection, permission mode, config write) that completes in < 2 minutes
Acceptance criteria: Init completes in < 2 minutes with guided prompts, after init teaagent "hello" produces useful response, init is documented first step in README
Size: Medium (1 week)
Reference: docs/plans/ux-improvement-roadmap-2026-05-31.md
Status: ✅ Complete
Journey: On first run (no .teaagent/), show orientation message explaining governance features (approval gates, audit log, undo, budget cap)
Acceptance criteria: Orientation message shown on first run, explains key governance features, includes next steps
Size: Small (2 days)
Reference: docs/plans/ux-improvement-roadmap-2026-05-31.md
Status: ✅ Complete
Journey: Implement DecisionLog stored in .teaagent/decisions.md, inject into system prompt, CLI commands to list/add decisions
Acceptance criteria: teaagent memory decisions list shows all decisions, teaagent memory decisions add appends entry, system prompt includes 10 most recent decisions
Size: Medium (1 week)
Reference: docs/plans/ux-improvement-roadmap-2026-05-31.md
Status: ✅ Complete
Journey: Track context window usage, warn at 60% with /compact suggestion, implement one-command summary-and-continue
Acceptance criteria: Warning triggers at configurable threshold (default 60%), /compact summarizes conversation and replaces with compressed version
Size: Medium (1 week)
Reference: docs/plans/ux-improvement-roadmap-2026-05-31.md
Status: ✅ Complete
Journey: Write structured scratchpad on session end (goal, progress, open questions, next step), offer resumption on next session start
Acceptance criteria: Scratchpad written on clean exit/ctrl+c/crash, offered for resumption on next session start, teaagent sessions list shows recent sessions
Size: Medium (1 week)
Reference: docs/plans/ux-improvement-roadmap-2026-05-31.md
Note: All P0 security and reliability items from comprehensive audit have been completed.
Status: ✅ Complete
Journey: Removed incorrect "encrypted at rest" claim from audit.py docstring, updated threat-model.md with L3 plaintext row
Acceptance criteria: grep returns no "encrypted at rest" in audit.py, threat-model.md has L3 plaintext row, no other file claims L3 encrypts
Size: Small (1 day)
Reference: docs/plans/comprehensive-plan-all-aspects-2026-05-31.md
Status: ✅ Complete
Journey: Added docstring warning, added trusted_only field that raises ValueError when False, updated security-spec.md with trust boundary table
Acceptance criteria: ChildProcessCodeModeBackend(trusted_only=False) raises ValueError, docstring change auditable via git log
Size: Small (1 day)
Reference: docs/plans/comprehensive-plan-all-aspects-2026-05-31.md
Status: ✅ Complete
Journey: Replaced new_loop pattern in approval_manager.py and policy.py with proper async loop handling to prevent "Event loop is closed" errors
Acceptance criteria: No asyncio.set_event_loop(new_loop) pattern in approval paths, tests verify no closed loop errors
Size: Small (2 days)
Reference: docs/plans/comprehensive-plan-all-aspects-2026-05-31.md
Status: ✅ Complete
Journey: Converted shell=True to argv-only path in workspace_run_shell for security
Acceptance criteria: shell=True removed, argv-only path used, tests pass
Size: Small (2 days)
Reference: docs/plans/remediation-roadmap.md
Status: ✅ Complete
Journey: Completed ShellObfuscationTests coverage for adversarial matrix
Acceptance criteria: ShellObfuscationTests pass, adversarial matrix complete
Size: Medium (3 days)
Reference: docs/plans/remediation-roadmap.md
Status: ✅ Complete
Journey: Required token when TEAAGENT_STRICT_LOCAL=1 for MCP loopback
Acceptance criteria: MCP loopback requires token when strict local mode enabled, tests pass
Size: Small (2 days)
Reference: docs/plans/remediation-roadmap.md
Status: ✅ Complete
Journey: Auto-load relay-tokens.json or env for vote relay auth
Acceptance criteria: surface_auth.default_relay_token_file implemented, tests pass
Size: Small (2 days)
Reference: docs/plans/remediation-roadmap.md
Status: ✅ Complete
Journey: Implemented fail-closed when TEAAGENT_PLUGINS_STRICT=1
Acceptance criteria: Plugin fails closed when strict mode enabled, tests pass
Size: Small (2 days)
Reference: docs/plans/remediation-roadmap.md
Status: ✅ Complete
Journey: Documented plaintext bearer token files as known limitation with operational guidance
Acceptance criteria: Bearer token files documented as known limitation with chmod 600 guidance
Size: Small (documentation only)
Reference: docs/plans/remediation-roadmap.md
Status: ✅ Complete Journey: Verified audit verify raises exceptions properly Acceptance criteria: Audit verify raises exceptions properly, tests pass Size: Small (verification only) Reference: `docs/plans/remediation-roadmap.md``
The following plans are marked as complete and are retained for reference:
- daily-driver-backlog-2026-06-01.md: ✅ COMPLETE - All 6 tickets implemented (chat REPL fixes, cost accounting, compaction, shared controller, TUI scrollback)
- governance-hardening.md: ✅ SHIPPED - All tranches shipped (tool lint, plan gate, audit, failure cards, MCP trust, approval queue, swarm, Phase 4-6 Beta)
- daily-driver-execution-readiness-2026-06-01.md: ✅ COMPLETE - All items marked as completed
- daily-driver-hardening-plan-2026-06-01.md: ✅ COMPLETE - All items marked as completed
Status: 📋 Ongoing monitoring (automation script created)
Journey: Monitor OpenCode GitHub releases monthly for governance-related features, escalate if approval gates added
Acceptance criteria: Monthly checks performed, escalation trigger defined (50+ votes on governance issues)
Size: Ongoing
Reference: docs/plans/competitive-positioning-plan-2026-05-31.md
Process document: docs/processes/opencode-gap-watch.md
Automation script: scripts/opencode_gap_watch.py
Status: ✅ Complete
Journey: Create model capability matrix document mapping features to models (extended thinking, tool use, streaming, multi-sig, cost tracking)
Acceptance criteria: Matrix exists and updated with provider adapter changes, README links to matrix
Size: Small (2 days)
Reference: docs/plans/competitive-positioning-plan-2026-05-31.md
Status: On hold — external acquisition and community-growth activity are not current goals.
Journey: Historical proposal only; monitoring or public outreach requires a new owner-ratified positioning decision.
Acceptance criteria: No activity is scheduled from this row while harness-first and DR-006 remain unchanged.
Size: Not scheduled
Reference: docs/plans/competitive-positioning-plan-2026-05-31.md (historical)
Process document: docs/processes/community-presence.md (inactive reference)
Automation script: scripts/community_presence_monitor.py (must not imply scheduled work)
For long-term planning, positioning, and UX research, see:
docs/plans/competitive-positioning-plan-2026-05-31.md- README rewrite, security whitepaper, demo scriptsdocs/plans/ux-improvement-roadmap-2026-05-31.md- Post-run summary, budget warnings, undo UX, memory/contextdocs/plans/future-roadmap-risk-usability-backlog-2026-05-31.md- Large strategic backlog with horizons
These are reference documents for strategic decisions. Extract actionable items to this backlog as needed.
The listed Phase 4–6 code surfaces and cited focused tests exist at Beta maturity. This does not complete canonical M4–M6 milestones, prove owner dogfood, authorize new platform surface, or establish hosted, multi-tenant, or production readiness.
Phase 4 consensus and Phase 5 sandbox code surfaces have CLI, unit-test, and acceptance evidence. Follow-on work remains governed by roadmap holds, DR-006, and the applicable ADR decision dates; code existence is not milestone acceptance.
| Task | Priority | Status | Evidence |
|---|---|---|---|
| Consensus data structures | P1 | Shipped | teaagent/consensus/ engine, tests/test_consensus_engine_history.py |
| Peer identity management | P1 | Shipped | PeerRegistry, CLI teaagent consensus peers |
| Voting mechanism | P1 | Shipped | teaagent/consensus/voting.py, tests/test_consensus_cli.py |
| Consensus engine | P0 | Shipped | ConsensusEngine, attestation + pre-approval |
| Swarm integration | P0 | Shipped | SwarmManager + pre-approval gate |
| CLI commands | P2 | Shipped | teaagent/cli/_handlers/_consensus.py |
| E2E acceptance | P2 | Shipped | tests/acceptance/test_consensus_flow.py |
| Task | Priority | Status | Evidence |
|---|---|---|---|
| Docker resource limits | P1 | Shipped | prepare_subagent_isolation --cpus / --memory |
| WASM runtime wrapper | P0 | Shipped | teaagent/wasm_runtime.py, tests/test_wasm_runtime.py |
| Skill execution routing | P1 | Shipped | teaagent/skill_router.py, isolation=auto on subagents |
| Resource monitoring | P2 | Shipped | teaagent/resource_monitor.py, teaagent sandbox monitor |
| CLI commands | P2 | Shipped | teaagent/cli/_handlers/_sandbox.py |
| E2E acceptance | P2 | Shipped | tests/acceptance/test_sandbox_enhancement_flow.py |
| Task | Priority | Status | Description |
|---|---|---|---|
| Async consensus UX | P2 | Shipped | Non-blocking vote collection via async_vote_collection + poll |
| WASM skill execution | P1 | Shipped | teaagent/skill_executor.py, teaagent sandbox execute |
| Docker dogfood benchmarks | P2 | Shipped | Optional CI docker-smoke job (tests/test_phase6_docker.py) |
| Task | Priority | Status | Evidence |
|---|---|---|---|
| SkillWriter publish + review | P1 | Shipped | teaagent/skill_writer.py, tests/test_phase6_skill_writer.py |
| Docker resource monitor + abort | P1 | Shipped | teaagent/docker_sandbox.py, tests/test_phase6_docker.py |
| Prompt tournament scoring | P2 | Shipped | teaagent/swarm.py, tests/test_phase6_swarm_score.py |
| Control plane HTTP + dashboard | P2 | Shipped | teaagent/control_plane_api.py, tests/test_phase6_control_plane.py |
| Control plane CLI | P2 | Shipped | teaagent control-plane serve |
| JIT approval server | P2 | Shipped | teaagent/jit_approval_server.py, tests/test_phase6_jit_server.py |
| External peer voting UX | P2 | Shipped | teaagent consensus wait, votes-import, import_votes_batch |
| Comparator tournament from subagent results | P2 | Shipped | teaagent/tournament/benchmark.py, tests/test_phase6_remaining_features.py |
| Control plane swarm dogfood | P2 | Shipped | teaagent/control_plane_bridge.py, SwarmManager(control_plane_state=…) |
| WASM skill contract CLI | P2 | Shipped | teaagent/wasm_skill.py, teaagent sandbox wasm-contract |
| SSH vote relay (production peers) | P1 | Shipped | teaagent/vote_relay.py, teaagent consensus relay |
| WASM skill CI template workflow | P2 | Shipped | .github/workflows/wasm-skill-build.yml, docs/wasm-skill-ci.md |
| Multi-tenant control plane | P2 | Shipped | teaagent/control_plane_tenant.py, X-TeaAgent-Tenant |
| Relay mTLS + API tokens | P1 | Shipped | teaagent/surface_auth.py, teaagent/tls_server.py, docs/http-surface-auth.md |
| Per-tenant authZ + reverse proxy templates | P2 | Shipped | templates/reverse-proxy/, control plane bearer enforcement |
All items from the 2026-05-22 ergonomics tranche are shipped. Future work is strategic (hosted surfaces, repo-map quality eval) under ongoing maintenance below.
| Item | Notes |
|---|---|
| Docs/provider architecture drift guard | scripts/validate_docs_consistency.py now checks README/USAGE/architecture counts against PROVIDER_CONFIGS, unique credential env vars (including shared CLOUDFLARE_API_TOKEN / OPENCODEZEN_API_KEY), and survey doc freshness; docs/architecture.md updated to 13 providers. |
| Mode and safety comparison matrix | docs/USAGE.md matrix maps permission modes, Plan/Auto/Code lanes, shell mutation, approvals, audit, and rollback; validator enforces every PermissionMode value and required topics. |
| Multi-surface launch recipes | docs/USAGE.md “Choose Your Surface” table with CLI/TUI/VS Code/MCP/ACP/A2A/ANP/managed-runtime recipes; validate_surface_recipes + test_surface_launch_recipes_flow.py smoke local commands. |
| Subagent lineage and isolation hardening | Child runs record lineage metadata; isolation: shared, worktree, or container snapshot; worktree_path/container_path in lineage; test_subagent_worktree_isolation_flow.py, test_subagent_container_isolation_flow.py. |
| Repo-map / context pack for coding runs | build_context_pack on agent preflight; candidate files, memories, optional LSP symbols, hybrid/knowledge/GraphQLite read-only hits; test_context_pack_read_only_flow.py. |
| Plugin/skill compatibility catalog | docs/plugin-skill-catalog.md with fixture-backed validate_plugin_skill_catalog. |
| Competitive use-case dashboard refresh | Matrix/HTML include survey review date and open P1/P2 gap counts via build_use_case_matrix.py + render_use_case_dashboard.py. |
| Periodic mainstream-agent refresh cadence | docs/release-checklist.md requires survey refresh before minor releases or protocol ADRs. |
| Everyday usage onboarding | README “Daily Use in 5 Commands”, USAGE daily section, cli recipe table, TUI daily → preflight → ask → resume flow. |
Competitive-refresh feature work is shipped. Before minor releases or protocol ADRs, run:
python3 scripts/refresh_competitive_docs.pyThen manually refresh scripts/refresh_agent_readme_survey.md when upstream signals change (see docs/release-checklist.md).
| Cadence | What to keep in sync |
|---|---|
| Survey | Last reviewed date, source table, docs/use-cases.md differentiators |
| Docs drift | validate_docs_consistency.py (providers, mode matrix, surface recipes, catalog) |
| Coverage artifacts | docs/acceptance.md count, docs/use-case-matrix.md, docs/use-case-matrix.html via refresh_competitive_docs.py --check (CI/review) and refresh_competitive_docs.py (intentional regeneration) |
| Extension surfaces | docs/plugin-skill-catalog.md, docs/USAGE.md, docs/cli.md when hooks, isolation modes, or preflight/context APIs change |
| New index backends | teaagent/context_pack.py read-only graph evidence when adding search stores |
| Strategic P2 research | hosted/cloud surface docs, background sessions, desktop/client-server packaging, repo-map quality evaluation |
| Ergonomics KPI | python3 scripts/measure_time_to_first_run.py --write docs/ergonomics-kpi.json before regenerating the use-case matrix |