Skip to content

feat(agent-profiles): one launch pipeline for every conversation start - #5154

Closed
simonrosenberg wants to merge 7 commits into
mainfrom
agent-profile-unified-launch
Closed

simonrosenberg wants to merge 7 commits into
mainfrom
agent-profile-unified-launch

Conversation

@simonrosenberg

@simonrosenberg simonrosenberg commented Sep 17, 2026 •

Copy link
Copy Markdown
Member

HUMAN:

Filing the SDK half of #5141: the default profile and named profiles have to build the same agent from one pipeline, so new profile fields stop needing an implementation per launch path. Canvas and cloud adoption follow separately.


AGENT:

Why

Launching the default Agent Profile and launching a named one built different agents from the same stored settings, because the launch had several code paths and each set agent fields its own way: a Pydantic validator for agent_settings, the agent_profile_id branch of conversation_service, a second copy of that branch in the Docker runtime's prepare_start, and a separate dry-run for materialize. Every profile field had to be implemented, and kept correct, once per path — #3967, #4014, #4016 and #4542 were that cost paid one field at a time.

This is step 1–3 of #5141 (the SDK steps). Canvas and cloud adoption are the follow-ups listed below.

Summary

  • One function builds every launch. openhands.sdk.profiles.prepare_agent_launch(source, *, catalog, runtime, additions, profile_origin, build_agent) resolves a profile's llm_profile_ref / mcp_server_refs / disabled_skills and owns the launch-time fields: tool defaults and browser injection, forced streaming, skill catalog and ACP skill sourcing, project-skill loading, suffix + additions, current_datetime, load_memory, and the secret scope. Runtime-dependent answers come in as an explicit AgentLaunchRuntime, so the Docker runtime can pass its container's answer instead of the host's. conversation_service, Docker mediation and materialize all call it through the new agent_server/agent_launch.py; _resolve_agent_from_profile, _with_load_memory and _apply_acp_skill_sourcing are gone, as is the duplicated profile branch in prepare_start.
  • materialize is the same call with side effects off (build_agent=False, which skips create_agent() — the only step that can refresh a subscription LLM's credentials over the network). Its verdict and resolved_settings now come from the launch itself.
  • New request shapes. StartConversationRequest accepts an inline agent_profile draft (resolved exactly like a stored one, never saved), and agent_settings is deprecated (deprecated: true in OpenAPI, removal target v1.55.0). During the window the server converts it into an inline profile with the payload's own LLM/MCP/skills as its catalog, so it takes the same pipeline instead of the validator shortcut; fields no profile models (critic_api_key, user_message_suffix, agent_context.secrets, acp_isolate_data_dir, …) are carried through, not dropped.
  • AgentLaunchAdditions.llm_profile_ref gives the chat LLM picker a per-launch override, recorded in LaunchedAgentProfile (which also gained inline). Additions stay additive: they carry no tools, MCP servers, skills or secrets, and a test asserts a scoped profile's tools/MCP/skills/secret scope are byte-identical with and without them.
  • One structured failure. A dangling LLM profile, MCP server or meta-profile ref raises UnresolvedProfileReferences and the start endpoint returns 422 with {"code": "unresolved_profile_references", "message", "dangling_llm_profile_ref", "dangling_mcp_server_refs", "dangling_meta_profile_ref"} — no silent fallback to a different path.
  • Meta-profile routing now reaches a profile launch. OpenHandsAgentProfile gains enable_classify_and_switch_llm_tool (behavior, like the enable_switch_llm_tool beside it) and meta_profile_ref (a reference resolved against the meta-profile store, like llm_profile_ref), and the launch hydrates the meta-profile plus the LLM profiles it routes to. See the note below for why hydration, not just the name.

REST API contract changes

Compared with base OpenAPI 3311ba9eec50 for public /api/** paths.

--- base public OpenAPI
+++ head public OpenAPI
@@ -771,0 +772,15 @@
+schema ACPAgentProfile property acp_args optional schema=anyOf=[type="array" items=type="string",type="null"]
+schema ACPAgentProfile property acp_command optional schema=anyOf=[type="string",type="null"]
+schema ACPAgentProfile property acp_model optional schema=anyOf=[type="string",type="null"]
+schema ACPAgentProfile property acp_prompt_timeout optional schema=type="number" default=1800.0 exclusiveMinimum=0.0
+schema ACPAgentProfile property acp_server optional schema=type="string" enum=["claude-code","codex","gemini-cli","kimi-code","pi","opencode","custom"] default="claude-code"
+schema ACPAgentProfile property acp_session_mode optional schema=anyOf=[type="string",type="null"]
+schema ACPAgentProfile property acp_startup_timeout optional schema=type="number" default=90.0 exclusiveMinimum=0.0
+schema ACPAgentProfile property agent_kind optional schema=type="string" const="acp" default="acp"
+schema ACPAgentProfile property id optional schema=type="string" format="uuid"
+schema ACPAgentProfile property mcp_server_refs optional schema=anyOf=[type="array" items=type="string",type="null"]
+schema ACPAgentProfile property name required schema=type="string" minLength=1
+schema ACPAgentProfile property revision optional schema=type="integer" default=0 minimum=0.0
+schema ACPAgentProfile property schema_version optional schema=type="integer" default=2 minimum=1.0
+schema ACPAgentProfile property secret_refs optional schema=anyOf=[type="array" items=type="string",type="null"]
+schema ACPAgentProfile type="object" additionalProperties=false
@@ -960,0 +976 @@
+schema AgentLaunchAdditions property llm_profile_ref optional schema=anyOf=[type="string" minLength=1,type="null"]
@@ -1922,0 +1939,9 @@
+schema LLMSummarizingCondenserSettings property condenser_kind optional schema=type="string" const="llm_summarizing" default="llm_summarizing"
+schema LLMSummarizingCondenserSettings property enabled optional schema=type="boolean" default=true
+schema LLMSummarizingCondenserSettings property hard_context_reset_context_scaling optional schema=type="number" default=0.8 exclusiveMinimum=0.0 exclusiveMaximum=1.0
+schema LLMSummarizingCondenserSettings property hard_context_reset_max_retries optional schema=type="integer" default=5 exclusiveMinimum=0.0
+schema LLMSummarizingCondenserSettings property keep_first optional schema=type="integer" default=2 minimum=0.0
+schema LLMSummarizingCondenserSettings property max_size optional schema=type="integer" default=240 minimum=20.0
+schema LLMSummarizingCondenserSettings property max_tokens optional schema=anyOf=[type="integer" exclusiveMinimum=0.0,type="null"]
+schema LLMSummarizingCondenserSettings property minimum_progress optional schema=type="number" default=0.1 exclusiveMinimum=0.0 exclusiveMaximum=1.0
+schema LLMSummarizingCondenserSettings type="object"
@@ -1923,0 +1949,2 @@
+schema LaunchedAgentProfile property inline optional schema=type="boolean" default=false
+schema LaunchedAgentProfile property llm_profile_ref optional schema=anyOf=[type="string",type="null"]
@@ -2230,0 +2258,3 @@
+schema NoOpCondenserSettings property condenser_kind optional schema=type="string" const="no_op" default="no_op"
+schema NoOpCondenserSettings property enabled optional schema=type="boolean" default=true
+schema NoOpCondenserSettings type="object"
@@ -2242,0 +2273,20 @@
+schema OpenHandsAgentProfile property agent optional schema=type="string" default="CodeActAgent"
+schema OpenHandsAgentProfile property agent_kind optional schema=type="string" const="openhands" default="openhands"
+schema OpenHandsAgentProfile property condenser optional schema=oneOf=[LLMSummarizingCondenserSettings,NoOpCondenserSettings]
+schema OpenHandsAgentProfile property disabled_skills optional schema=type="array" items=type="string"
+schema OpenHandsAgentProfile property enable_classify_and_switch_llm_tool optional schema=type="boolean" default=false
+schema OpenHandsAgentProfile property enable_sub_agents optional schema=type="boolean" default=false
+schema OpenHandsAgentProfile property enable_switch_llm_tool optional schema=type="boolean" default=true
+schema OpenHandsAgentProfile property id optional schema=type="string" format="uuid"
+schema OpenHandsAgentProfile property llm_profile_ref required schema=type="string" minLength=1
+schema OpenHandsAgentProfile property mcp_server_refs optional schema=anyOf=[type="array" items=type="string",type="null"]
+schema OpenHandsAgentProfile property meta_profile_ref optional schema=anyOf=[type="string",type="null"]
+schema OpenHandsAgentProfile property name required schema=type="string" minLength=1
+schema OpenHandsAgentProfile property revision optional schema=type="integer" default=0 minimum=0.0
+schema OpenHandsAgentProfile property schema_version optional schema=type="integer" default=2 minimum=1.0
+schema OpenHandsAgentProfile property secret_refs optional schema=anyOf=[type="array" items=type="string",type="null"]
+schema OpenHandsAgentProfile property system_message_suffix optional schema=anyOf=[type="string",type="null"]
+schema OpenHandsAgentProfile property tool_concurrency_limit optional schema=type="integer" default=1 minimum=1.0
+schema OpenHandsAgentProfile property tools optional schema=anyOf=[type="array" items=Tool-Input,type="null"]
+schema OpenHandsAgentProfile property verification optional schema=ProfileVerificationSettings
+schema OpenHandsAgentProfile type="object" additionalProperties=false
@@ -2349,0 +2400,8 @@
+schema ProfileVerificationSettings property critic_enabled optional schema=type="boolean" default=false
+schema ProfileVerificationSettings property critic_mode optional schema=type="string" enum=["finish_and_message","all_actions"] default="finish_and_message"
+schema ProfileVerificationSettings property critic_model_name optional schema=anyOf=[type="string",type="null"]
+schema ProfileVerificationSettings property critic_server_url optional schema=anyOf=[type="string",type="null"]
+schema ProfileVerificationSettings property critic_threshold optional schema=type="number" default=0.6 minimum=0.0 maximum=1.0
+schema ProfileVerificationSettings property enable_iterative_refinement optional schema=type="boolean" default=false
+schema ProfileVerificationSettings property max_refinement_iterations optional schema=type="integer" default=3 minimum=1.0
+schema ProfileVerificationSettings type="object"
@@ -2551,0 +2610 @@
+schema StartConversationRequest property agent_profile optional schema=anyOf=[oneOf=[OpenHandsAgentProfile,ACPAgentProfile],type="null"]

Issue Number

Fixes #5141

How to Test

Unit tests (the meta-profile wiring is covered by 5 cases in tests/sdk/profiles/test_launch.py and test_meta_profile_routing_reaches_a_profile_launch in the parity file):

uv run pytest tests/sdk/profiles tests/agent_server/test_agent_launch_parity.py \
  tests/agent_server/test_agent_profile_conv_start.py \
  tests/agent_server/test_agent_launch_additions.py \
  tests/agent_server/test_acp_skill_sourcing.py \
  tests/agent_server/docker_runtime tests/agent_server/test_conversation_router.py -q

End-to-end against a real agent-server (this is the interesting one — it reproduces the table in #5141 without canvas):

uv run python .pr/launch_parity_e2e.py

It boots python -m openhands.agent_server on a temp persistence dir, stores two identically-configured profiles (default and default-copy), launches a conversation through agent_profile_id for each, through an inline agent_profile draft, and through the deprecated agent_settings, then diffs the agents the server actually built against the materialize preview of the same profile. Output in .pr/launch_parity_e2e_output.txt:

legacy current_datetime sent 2020-01-01T00:00, launched with 2026-09-17T14:08:55.977969-04:00
default-copy (agent_profile_id): MATCH
inline (agent_profile): MATCH
legacy (agent_settings): MATCH
materialize (default): MATCH
dangling refs -> HTTP 422 {"detail":{"code":"unresolved_profile_references","message":"LLM profile 'gone' not found; MCP server(s) not configured: nope","dangling_llm_profile_ref":"gone","dangling_mcp_server_refs":["nope"]}}
RESULT: PASS

Compared field by field: llm (whole dump), tools, MCP keys, skills, suffix, disabled_skills, load_project_skills, load_memory, condenser, critic, concurrency, switch-LLM and whether a timestamp is present. The one deliberate exception is the skill catalog on the agent_settings path: that payload carries the client's own catalog (canvas assembles one today), which is exactly what the canvas follow-up removes.

Type

  • Bug fix
  • Feature
  • Refactor
  • Breaking change
  • Docs / chore

Notes

Rebased onto main on 2026-09-28 (was 11 days behind; 76 commits). The rebase was conflict-free, and both suites pass locally with the meta-profile wiring included: tests/agent_server 2252 passed, tests/sdk 6538 passed, ruff + pyright clean. Run them separately — tests/agent_server and tests/sdk in one pytest process produce three unrelated ordering failures (CI runs them as separate jobs).

Two interactions found while rebasing; the first is fixed in this PR:

  1. Add Pareto prompt meta-profile routing #4287 (Pareto meta-profile routing, merged 09-24) had re-opened this exact divergence for four new fields — now wired through the profile. It added enable_classify_and_switch_llm_tool, active_meta_profile, meta_profile and meta_profile_llms to OpenHandsAgentSettings and not to OpenHandsAgentProfile, and _build_openhands_settings composes from an allow-list of profile fields. Measured on this branch before the fix:

    field agent_settings launch stored-profile launch
    enable_classify_and_switch_llm_tool True False
    active_meta_profile 'pareto' None

    The legacy path survived only because it passes base_settings; a profile launch silently lost route_task_to_model. The split follows this issue's own rule — behavior on the profile, shared resources global — so the profile carries the toggle plus a meta_profile_ref, and the meta-profile store stays global like the MCP registry and the LLM profiles.

    Why the launch hydrates rather than just passing the name: the routing tool reads the store by name and only falls back to the inline meta_profile blob when the store cannot resolve it. A conversation container's HOME is a per-conversation runtime dir with no meta-profile store, so a name alone would have left Docker routing failing at tool-call time. prepare_agent_launch therefore loads the meta-profile and the LLM profiles it routes to; a caller that hydrated them itself (a cloud control plane, through base_settings) keeps its own copies. A dangling meta_profile_ref fails the launch only when the tool is enabled — an inert ref is never resolved, so it cannot fail a launch it has no effect on.

    Known limit: on the deprecated agent_settings path the routing targets are not hydrated, because that catalog's LLM loader only knows the payload's single inline LLM. It is {} there today as well, so nothing regresses, and a local runtime resolves them from its own store.

  2. fix(agent-server): enforce profile secret scope on runtime-launched conversations #5193's secret-scope gap is adjacent but not closed here. apply_launch filters request.secrets only when the source is a profile; a raw agent bound to a profile through OH_RUNTIME_LAUNCHED_PROFILE (the in-container launch) is filtered on resume, not at create. I corrected the docstring that overclaimed this and left the fix to fix(agent-server): enforce profile secret scope on runtime-launched conversations #5193, where it belongs.

Merge-order note: this PR supersedes agent_server/profile_launch.py from #5151 — agent_launch.py subsumes gather_profile_launch_inputs, and probes the container's browser availability rather than the host's. #5151 is already conflicting with main independently (via #4287). Several open PRs touch functions this one rewrites (#5315, #4717, #5193, #4961, #5199, #5318, #4410, #5126); this PR itself merges cleanly with main.

Behavior changes reviewers should weigh:

  • A dangling LLM ref at start is now 422, not 404, and the 422 detail gained code / dangling_llm_profile_ref (the old message / dangling_mcp_server_refs keys are unchanged). Canvas never reads that status — it pre-checks the profile list and downgrades to agent_settings — and that rule is what the follow-up deletes.
  • agent_settings is no longer converted in the validator, so StartConversationRequest(agent_settings=...).agent is None until the server resolves it, and an invalid payload is rejected by the start endpoint (422) rather than at parse time. The field also lost exclude=True so it round-trips over the wire. An explicit "agent": null alongside another source no longer crashes.
  • The agent_settings path now gets the launch-owned fields too, because it goes through the same pipeline: streaming forced on, load_project_skills=True, browser added when the payload's tools is null and the runtime has it, and a fresh current_datetime instead of the saved one. Canvas already sends stream: true, load_project_skills: true and an explicit tool list, so its payload is unaffected; a hand-rolled REST client that relied on those staying off would see the change.
  • The Docker runtime now resolves twice: once before the container starts (build_agent=False, so a dangling ref still fails fast without paying for a container) and once after, with the container's own runtime answer (GET /server_info for browser availability, openhands_managed skills). Previously the host's answers were used for a container agent, which meant an ACP profile in Docker got no managed skills.
  • /server_info advertises unified_agent_launch_v1.

Not fixed here, found while testing: a condenser's max_tokens inheritance keys off model_fields_set, so a stored profile (loaded from JSON, every field "set") does not inherit the LLM's token limit while an in-memory one does. It is consistent across today's product paths (both read persisted JSON) and predates this PR, so I left it alone rather than widen the diff.

Follow-ups, per #5141:

🤖 Generated with Claude Code


🐳 Agent Server images for this PR — GHCR package, pull/run commands, and all pushed tags (click to expand)

• GHCR package: https://github.com/OpenHands/agent-sdk/pkgs/container/agent-server

Variants & Base Images

Variant Architectures Base Image Docs / Tags
java amd64, arm64 eclipse-temurin:17-jdk Link
python-slim amd64, arm64 python-node-runtime Link
python-minimal amd64, arm64 python-node-runtime Link
python amd64, arm64 python-node-runtime Link
golang amd64, arm64 golang:1.21-bookworm Link

Pull (multi-arch manifest)

# Each variant is a multi-arch manifest supporting both amd64 and arm64
docker pull ghcr.io/openhands/agent-server:dbd6b27-python

Run

docker run -it --rm \
  -p 8000:8000 \
  --name agent-server-dbd6b27-python \
  ghcr.io/openhands/agent-server:dbd6b27-python

All tags pushed for this build

ghcr.io/openhands/agent-server:dbd6b27-golang-amd64
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-golang-amd64
ghcr.io/openhands/agent-server:agent-profile-unified-launch-golang-amd64
ghcr.io/openhands/agent-server:dbd6b27-golang_tag_1.21-bookworm-amd64
ghcr.io/openhands/agent-server:dbd6b27-golang-arm64
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-golang-arm64
ghcr.io/openhands/agent-server:agent-profile-unified-launch-golang-arm64
ghcr.io/openhands/agent-server:dbd6b27-golang_tag_1.21-bookworm-arm64
ghcr.io/openhands/agent-server:dbd6b27-java-amd64
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-java-amd64
ghcr.io/openhands/agent-server:agent-profile-unified-launch-java-amd64
ghcr.io/openhands/agent-server:dbd6b27-eclipse-temurin_tag_17-jdk-amd64
ghcr.io/openhands/agent-server:dbd6b27-java-arm64
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-java-arm64
ghcr.io/openhands/agent-server:agent-profile-unified-launch-java-arm64
ghcr.io/openhands/agent-server:dbd6b27-eclipse-temurin_tag_17-jdk-arm64
ghcr.io/openhands/agent-server:dbd6b27-python-amd64
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-python-amd64
ghcr.io/openhands/agent-server:agent-profile-unified-launch-python-amd64
ghcr.io/openhands/agent-server:dbd6b27-python-node-runtime-amd64
ghcr.io/openhands/agent-server:dbd6b27-python-arm64
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-python-arm64
ghcr.io/openhands/agent-server:agent-profile-unified-launch-python-arm64
ghcr.io/openhands/agent-server:dbd6b27-python-node-runtime-arm64
ghcr.io/openhands/agent-server:dbd6b27-python-minimal-amd64
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-python-minimal-amd64
ghcr.io/openhands/agent-server:agent-profile-unified-launch-python-minimal-amd64
ghcr.io/openhands/agent-server:dbd6b27-python-node-runtime-minimal-amd64
ghcr.io/openhands/agent-server:dbd6b27-python-minimal-arm64
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-python-minimal-arm64
ghcr.io/openhands/agent-server:agent-profile-unified-launch-python-minimal-arm64
ghcr.io/openhands/agent-server:dbd6b27-python-node-runtime-minimal-arm64
ghcr.io/openhands/agent-server:dbd6b27-python-slim-amd64
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-python-slim-amd64
ghcr.io/openhands/agent-server:agent-profile-unified-launch-python-slim-amd64
ghcr.io/openhands/agent-server:dbd6b27-python-node-runtime-slim-amd64
ghcr.io/openhands/agent-server:dbd6b27-python-slim-arm64
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-python-slim-arm64
ghcr.io/openhands/agent-server:agent-profile-unified-launch-python-slim-arm64
ghcr.io/openhands/agent-server:dbd6b27-python-node-runtime-slim-arm64
ghcr.io/openhands/agent-server:dbd6b27-golang
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-golang
ghcr.io/openhands/agent-server:agent-profile-unified-launch-golang
ghcr.io/openhands/agent-server:dbd6b27-golang_tag_1.21-bookworm
ghcr.io/openhands/agent-server:dbd6b27-java
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-java
ghcr.io/openhands/agent-server:agent-profile-unified-launch-java
ghcr.io/openhands/agent-server:dbd6b27-eclipse-temurin_tag_17-jdk
ghcr.io/openhands/agent-server:dbd6b27-python-minimal
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-python-minimal
ghcr.io/openhands/agent-server:agent-profile-unified-launch-python-minimal
ghcr.io/openhands/agent-server:dbd6b27-python-node-runtime-minimal
ghcr.io/openhands/agent-server:dbd6b27-python-slim
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-python-slim
ghcr.io/openhands/agent-server:agent-profile-unified-launch-python-slim
ghcr.io/openhands/agent-server:dbd6b27-python-node-runtime-slim
ghcr.io/openhands/agent-server:dbd6b27-python
ghcr.io/openhands/agent-server:dbd6b272796f8e1c707d418eac6db571c6e324ae-python
ghcr.io/openhands/agent-server:agent-profile-unified-launch-python
ghcr.io/openhands/agent-server:dbd6b27-python-node-runtime

About Multi-Architecture Support

  • Each variant tag (e.g., dbd6b27-python) is a multi-arch manifest supporting both amd64 and arm64
  • Docker automatically pulls the correct architecture for your platform
  • Individual architecture tags (e.g., dbd6b27-python-amd64) are also available if needed

@github-actions

Copy link
Copy Markdown
Contributor

📁 PR Artifacts Notice

This PR contains a .pr/ directory with temporary PR-specific documents. The directory will be automatically removed when the PR is approved.

@github-actions

github-actions Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

REST API breakage checks (OpenAPI) — ✅ PASSED

Result: ✅ PASSED

Action log

@github-actions

github-actions Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Coverage

Coverage Report •
FileStmtsMissCoverMissing
openhands-agent-server/openhands/agent_server
   agent_launch.py901484%54–55, 61, 67, 96–97, 100–101, 116–117, 157–158, 192–193
   agent_profiles_router.py2191693%170, 174, 178, 197, 200, 336, 340, 452, 492–493, 552–554, 565–567
   config.py130199%447
   conversation_router.py2691594%190, 316, 398, 444, 504, 674–677, 689–692, 732, 770
   conversation_service.py119812590%188–189, 198, 225–226, 230–231, 236, 432–433, 494, 597, 604–605, 706, 792, 843–844, 851, 883–884, 900, 931, 935, 947, 967, 979–982, 988–989, 998, 1000, 1066, 1076, 1101, 1107–1108, 1112–1113, 1121, 1148, 1154, 1248, 1254, 1259, 1265, 1273–1274, 1283–1286, 1295, 1307, 1315, 1361, 1367–1368, 1371–1373, 1400, 1452, 1501–1502, 1506, 1555–1556, 1632, 1687–1689, 1691–1692, 1695–1696, 1733, 1807–1808, 1840, 1843, 1850–1852, 1855–1856, 1860–1862, 1865–1866, 1870–1872, 1875–1876, 1905, 1914, 1957, 1967–1969, 2029, 2032, 2059, 2069, 2074–2077, 2091, 2102, 2114–2115, 2147, 2242, 2299, 2357, 2372–2373, 2751, 2804, 2807
   server_details_router.py61297%27–28
openhands-agent-server/openhands/agent_server/docker_runtime
   mediation.py67790%74–77, 101–102, 134
   routers.py20310150%49, 56–62, 98–100, 113–118, 120–121, 125–127, 129–130, 133–136, 138–147, 149–151, 154–156, 160–162, 168–173, 176–179, 181–184, 186, 189, 200, 209, 219, 225–227, 237–239, 248, 251, 295, 302–303, 346–348, 369, 371–383, 386–388, 394–395, 406, 415
openhands-sdk/openhands/sdk/conversation
   message_request.py9189%19
   request.py931089%308, 328, 335–336, 339–340, 354, 360, 372, 379
openhands-sdk/openhands/sdk/profiles
   agent_profile.py112794%347, 358, 361, 428, 434, 439, 467
   resolver.py273499%286, 335, 484, 672
openhands-sdk/openhands/sdk/settings
   model.py8306992%314, 332, 541, 558, 568–571, 574, 587, 591, 597, 607, 613, 618, 734, 737–738, 744–747, 752, 756, 771, 774–776, 824, 829–830, 835, 856, 868, 917, 926, 973, 1161, 1186, 1188, 1190, 1192, 1194, 1196, 1198, 1200, 1202, 1204, 1534, 1536, 1838, 1858, 1995, 2124, 2163, 2318–2320, 2322, 2408, 2418, 2420, 2425, 2443, 2456, 2458, 2460, 2462, 2469
TOTAL45405819982% 

@simonrosenberg
simonrosenberg marked this pull request as ready for review September 17, 2026 18:19
@all-hands-bot

Copy link
Copy Markdown
Collaborator

🤖 OpenHands is reviewing this PR.

Head commit: 97f40abbf1634e6239b09069b40844b42ccadca6
View the conversation: https://oss-agent-canvas.ngrok.dev/conversations/f993aa67-a43b-4257-a8c0-ae63d9370c21

This comment was posted by an AI agent (OpenHands).

@all-hands-bot all-hands-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This review was created by an AI agent (OpenHands) on behalf of the repository maintainers.

Summary

This PR unifies every conversation-start path (stored profile, inline profile, raw agent, deprecated agent_settings, materialize preview, and the Docker runtime) behind a single prepare_agent_launch function. I reviewed the resolver, the agent-server launch glue, the Docker mediation double-resolve, the request model changes, and the test suite.

No material bugs found. The design is clean and achieves what it set out to do: each profile field is now implemented once instead of once-per-launch-path.

What holds up under scrutiny

  • Exception hierarchy is correct. UnresolvedProfileReferences -> AgentLaunchError -> ValueError, and both routers catch AgentLaunchError before ValueError, so the structured 422 detail (code/dangling_llm_profile_ref/dangling_mcp_server_refs) is preserved rather than collapsed to a plain string. ProfileNotFound stays a separate Exception -> 404.
  • Docker double-resolve is sound. prepare_start runs with build_agent=False (and browser_available=False) so a dangling ref fails before any container is provisioned; finish_start re-resolves the same LaunchSource/catalog with the container's real browser_available from GET /server_info (usable_tools field -- verified correct). Skill discovery happens once in the catalog and is reused, not re-run. Container cleanup (registry.stop) is wired on every failure branch.
  • Secret scoping is enforced server-side in apply_launch (plan.allowed_secrets filters request.secrets), and agent_settings_launch_source correctly sets allowed_secrets=None (unrestricted) to match the legacy path. The additions-cannot-widen-scope invariant is asserted by a real test.
  • Backward compatibility is preserved. LaunchedAgentProfile.inline/llm_profile_ref and AgentLaunchAdditions.llm_profile_ref are additive fields with defaults; LaunchedAgentProfile has no extra="forbid", so old persisted conversations load. agent_settings is deprecated with deprecated_in=1.50.0 -> removed_in=1.55.0 (5 minor releases, meeting the policy) and is converted to an inline profile through the same pipeline rather than dropped.
  • Tests exercise real code paths, not mock wiring: test_launch.py covers runtime pieces, additions, provenance, dangling refs, and the deprecated agent_settings round-trip (verifying non-profile fields like critic_api_key, user_message_suffix, agent_context.secrets survive). The parity test diffs every launch path field-by-field.

Eval / benchmark risk -- flagging for a human maintainer

This PR changes agent launch behavior in ways that could plausibly move benchmark numbers, and there is no eval-monitor link or maintainer eval confirmation in the PR description or comments:

  • The agent_settings path now goes through the unified pipeline, so it gains forced stream=True, load_project_skills=True, browser injection (when the runtime has it and tools is null), and a fresh current_datetime instead of the saved timestamp. Canvas already sends these explicitly so it's unaffected, but a hand-rolled REST client relying on the old defaults would see a change.
  • ACP skill sourcing in Docker switched from the host's answer to the container's (openhands_managed), so an ACP profile in Docker now gets managed skills where it previously got none.

Per the repo's review policy I'm leaving a COMMENT rather than approving. Recommend a maintainer run lightweight evals (or confirm Canvas-only impact) before merging.

Minor note (non-blocking)

warn_deprecated(..., deprecated_in="1.50.0") is called from the agent-server while the current SDK version is 1.49.1. _should_warn compares current >= deprecated_in, so the runtime warning won't actually fire until the SDK ships as 1.50.0 -- which is presumably the release this PR targets, so this is consistent, just worth knowing the warning is effectively inert until then.

Risk assessment: MEDIUM -- no correctness/security issues, but behavior changes on the agent_settings/ACP-in-Docker paths that warrant eval confirmation.

Verdict: Worth merging after a maintainer confirms no eval regression.


Improve this review? If any feedback above seems incorrect or irrelevant to this repository, you can teach the reviewer to do better:

  1. Add a .agents/skills/custom-codereview-guide.md file to your branch (or edit it if one already exists) with the /codereview trigger and the context the reviewer is missing. See the customization docs for the required frontmatter format.
  2. Re-request a review - the reviewer reads guidelines from the PR branch, so your changes take effect immediately.
  3. When your PR is merged, the guideline file goes through normal code review by repository maintainers.

Resolve with AI? Install the iterate skill in your agent and run /iterate to automatically drive this PR through CI, review, and QA until it's merge-ready.

Was this review helpful? React with thumbs up or thumbs down to give feedback.

@simonrosenberg

Copy link
Copy Markdown
Member Author

Thanks — on the eval-risk flag, here is what I can evidence from this branch so a maintainer has the facts to decide. I have not run evals.

Who actually sees the agent_settings behavior changes. Canvas's payload builder (buildConfiguredOpenHandsAgentSettings / buildAgentContext on OpenHands/OpenHands@main) already sends llm.stream = true, load_project_skills: true, load_user_skills: true and an explicit tools list from getAgentTools. So of the four changes, three are no-ops for canvas and the fourth is the fix itself: current_datetime is now computed at launch instead of being the value saved with settings. Verified live — the e2e script sent 2020-01-01T00:00 and the launched agent came back with the launch timestamp (.pr/launch_parity_e2e_output.txt).

Cloud is not on this path. The enterprise app server builds a concrete agent (create_kwargs = {'agent': agent, ...} in live_status_app_conversation_service.py) and never sets the request's agent_settings, so neither the pipeline change nor the removed validator conversion reaches it.

ACP-in-Docker. Previously the host's acp_skill_sourcing decided what a container's ACP agent got, so with the default host config (native) an ACP profile in Docker launched with no managed skills while the container image itself sets OH_ACP_SKILL_SOURCING=openhands_managed. That is the drift this PR removes by passing the container's answer. It does change the ACP prompt in Docker, and it is the one change I would point an eval at if you want one.

Unchanged for stored-profile launches (the path evals exercise): the parity e2e diffs llm, tools, MCP keys, skills, suffix, disabled_skills, load_project_skills, load_memory, condenser, critic, concurrency and switch-LLM between two identically-configured profiles, an inline draft, the legacy payload and the materialize preview — all MATCH.

On the warn_deprecated note: agreed and intentional — 1.50.0 is the next minor, so the warning arms exactly when the deprecation takes effect.

simonrosenberg and others added 4 commits September 28, 2026 09:44
Collapse the several code paths that built a launch agent into
prepare_agent_launch(), the single SDK function that resolves an Agent
Profile's references and applies the runtime-dependent and per-launch
pieces. conversation_service, the Docker runtime's mediation and the
materialize preview all call it, so a profile named `default` and a named
one build the same agent, and a preview can no longer disagree with a
launch.

Adds an inline `agent_profile` draft and a per-launch `llm_profile_ref`
override, deprecates `agent_settings` (converted into an inline profile so
it takes the same pipeline), and returns one structured error for dangling
LLM/MCP references.

Fixes #5141

Co-authored-by: openhands <openhands@all-hands.dev>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: openhands <openhands@all-hands.dev>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…agent

Co-authored-by: openhands <openhands@all-hands.dev>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: openhands <openhands@all-hands.dev>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@simonrosenberg
simonrosenberg force-pushed the agent-profile-unified-launch branch from 97f40ab to 5786443 Compare September 28, 2026 07:54
simonrosenberg and others added 2 commits September 28, 2026 09:56
Co-authored-by: openhands <openhands@all-hands.dev>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#4287 added the routing settings to OpenHandsAgentSettings only, so
route_task_to_model reached an agent_settings launch and not a profile
launch. The profile now carries enable_classify_and_switch_llm_tool and
meta_profile_ref, and the launch hydrates the meta-profile plus the LLMs
it routes to, so a runtime without the store on disk can still route.

Co-authored-by: openhands <openhands@all-hands.dev>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@simonrosenberg

Copy link
Copy Markdown
Member Author

Code review (xhigh) — head 72d8960

Focus: no unintended agent-behavior changes, and keep the code simple. CI is green, but none of the items below are covered by tests. Items marked (repro) were reproduced locally.

Meta-profile routing

  1. (repro) Routing target LLMs are copied, with decrypted keys, into ClassifyAndSwitchLLMTool's untyped Tool.params — openhands-sdk/openhands/sdk/profiles/resolver.py:403. Because params is dict[str, Any], the keys don't round-trip: in Docker they arrive as Fernet ciphertext, and after a resume they come back as ciphertext or **********. The tool prefers _meta_profile_llms over switch_profile(), so route_task_to_model switches to an LLM that fails auth. That also defeats the Docker motivation for this change.
  2. (repro) A meta-profile missing from the store now fails the whole launch with 422, including on the deprecated agent_settings path — resolver.py:505. Before, the tool resolved it lazily and fell back to the inline meta_profile blob. A corrupt meta-profile file also became a 422 at start instead of an error at tool-call time.
  3. (repro) The dry run reports valid=True when only meta_profile_ref dangles — resolver.py:865. The except UnresolvedProfileReferences branch records only MCP and LLM refs, so materialize says the profile is launchable while the real launch returns 422.
  4. Deleting a meta-profile doesn't update agent profiles that reference it — meta_profiles_router.py:200. The seeded default profile (with meta_profile_ref copied from active_meta_profile) then returns 422 on every launch. /activate also stops affecting profile launches.
  5. (repro) resolve_agent_profile raises ProfileNotFound("LLM profile None not found") when only the meta-profile dangles — resolver.py:797.

Other behavior changes

  1. (repro) Explicit opt-outs in agent_settings are overridden — resolver.py:381. Setting agent_context.current_datetime=null or load_project_skills=false now gives dt=now and proj=True, so the system prompt changes and repo skills get loaded.
  2. Default tool order changed — resolver.py:617. default_tool_specs(enable_browser=...) puts browser_tool_set before task_tool_set. Before, the browser was appended after create_agent's defaults.
  3. (repro) agent_settings lost exclude=True but is still a raw dict[str, Any] — conversation/request.py:275. model_dump_json() now emits a plaintext llm.api_key.
  4. (repro) A malformed inline agent_profile returns 500 instead of 422 — request.py:333. validate_agent_profile raises TypeError inside a mode="before" validator, and Pydantic doesn't convert that into a ValidationError.
  5. MetaProfileStore() is built on every launch and materialize, even without a meta_profile_ref — agent_launch.py:63. It calls mkdir(parents=True), so an unwritable HOME gives a 500 for every launch.
  6. The browser-usability probe now runs on the event loop for every conversation start — conversation_service.py:1512, and the same in materialize. launch_runtime(...) is evaluated before asyncio.to_thread.
  7. Docker: create_agent() (subscription credential refresh) and secret materialization moved to after container creation — docker_runtime/routers.py:160. Their failures now cost a container cycle and surface as a generic 502.

Simplicity

  1. Resolution runs twice. The dry run loads and decrypts the LLM twice, and computes the MCP filter and skills twice — resolver.py:841. The Docker path runs the whole resolution twice.
  2. Launch policy is duplicated in four places. Materialize re-implements profile_catalog and the MetaProfileStore construction, and has its own copy of the "docker ⇒ openhands_managed" rule — agent_profiles_router.py:570. The explicit-null filter exists in both mediation.py:96 and _normalize_agent_source, and ACPSkillSourcing is defined twice.
  3. Comment conventions. apply_launch's docstring (agent_launch.py:171) is multi-paragraph, describes non-local behavior and cites fix(agent-server): enforce profile secret scope on runtime-launched conversations #5193. The same applies to resolver.py:399-400 and the _resolve_meta_profile docstring.

Minor: inline launches record the draft's random default UUID as agent_profile_id. _error_text drops the field location from 422s. The dry run drops a dangling MCP ref when the LLM load raises. The ACP skill-sourcing default differs between prepare_agent_launch (native) and the dry run (openhands_managed).

Checked and found safe: secrets under a cipher context, import cycles, containers with OH_ACP_SKILL_SOURCING=openhands_managed, the enterprise launch path, and container cleanup on a failed start.

- Pass meta_profile_ref to route_task_to_model by name instead of hydrating
  the meta-profile and its target LLMs (with decrypted keys) into the tool's
  untyped params at launch. The tool resolves it at call time again, so a
  missing meta-profile no longer fails the launch, and the meta-profile store
  is no longer created on every launch or materialize.
- Keep explicit agent_settings opt-outs (current_datetime=null,
  load_project_skills=false); only unset values take the launch defaults.
- Append the browser after the default tools, restoring the old tool order.
- Serialize agent_settings through the settings model so secrets are masked
  unless expose_secrets/cipher context is given.
- Turn a malformed inline agent_profile into a validation error (422) instead
  of an uncaught TypeError (500).
- Run the browser probe in the worker thread, and skip it for raw agents.
- Surface Docker build failures after the container starts as a 422 launch
  error with the real cause instead of a generic 502.
- Dry run reuses the launch's resolved LLM instead of loading it twice.
- Share ACPSkillSourcing and the server's skill-sourcing rule; drop the
  duplicate explicit-null filter in prepare_start; trim docstrings.

Co-authored-by: openhands <openhands@all-hands.dev>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@simonrosenberg

Copy link
Copy Markdown
Member Author

Fixes for the review above — dbd6b27

Meta-profile routing (1–5, 10). The launch no longer resolves meta_profile_ref. It passes the ref to route_task_to_model as active_meta_profile by name, and the tool resolves it at call time, as it did before this PR. The deprecated agent_settings path keeps the payload's own inline meta_profile and meta_profile_llms.

  • No decrypted target LLMs in Tool.params anymore, so resume and Docker can't pick up ciphertext keys (1).
  • A missing or corrupt meta-profile no longer fails a launch or shows a misleading preview verdict (2, 3). Deleting a meta-profile no longer breaks the profiles that reference it (4).
  • resolve_agent_profile can't report "LLM profile None" anymore (5).
  • No MetaProfileStore is created per launch or materialize (10).
  • MetaProfileLoader, AgentLaunchCatalog.meta_profile_store and dangling_meta_profile_ref are removed, net −50 lines.
  • Known limit: routing in a Docker container still needs the store there, same as before this PR.

Other behavior changes
6. On the agent_settings path, an explicit load_project_skills=false and current_datetime=null are kept. Only unset values take the launch defaults, and a saved timestamp is still refreshed.
7. The browser is appended after the default tools again, so the order is …, task_tool_set, browser_tool_set.
8. agent_settings serializes through the settings model, so secrets are masked unless expose_secrets or a cipher context is passed. RemoteConversation's exposed dump is unchanged.
9. A malformed inline agent_profile now raises a ValidationError (422) instead of a TypeError (500).
11. The browser probe runs in the worker thread, and raw-agent starts skip it. Materialize also builds its runtime off the event loop.
12. In Docker, a failure in finish_start becomes an AgentLaunchError: the container is stopped and the client gets a 422 with the real cause instead of a generic 502. The extra container cycle remains.

Simplicity
13. The dry run takes llm_profile_resolved and llm_api_key_set from the launch's own resolved LLM and loads it once. It reloads only in the rare case where MCP dangles. Transient store errors now come back as a launch error ("Could not load LLM profile …") on both paths. Docker still resolves twice, but meta hydration was the costly part, and that's gone.
14. ACPSkillSourcing is defined once (config imports it from the SDK). The "docker ⇒ openhands_managed" rule is one helper, server_acp_skill_sourcing. The duplicate explicit-null filter in prepare_start is removed.
15. The docstrings are trimmed: the apply_launch one is now a single line, and the cloud and #5193 notes are gone.

Tests. I added regression tests for 6, 7, 8, 9 and 12, and replaced the meta-profile tests to cover the name-only contract. I negative-controlled each new test: with its fix reverted it fails. Results:

  • tests/agent_server: 2257 passed.
  • tests/sdk: 6538 passed, plus 1 failure in test_truncate::test_maybe_truncate_hash_based_filename. It also fails on main under -n 8 and passes serially.
  • .pr/launch_parity_e2e.py: PASS.
  • ruff and pyright: clean.

Not changed:

  • The minor "inline launches record the draft's UUID" item.
  • The declared agent_settings behavior changes: forced streaming, and the browser added when tools is null.

🤖 Generated with Claude Code

@neubig
neubig requested review from all-hands-bot and removed request for all-hands-bot September 30, 2026 08:06

@all-hands-bot all-hands-bot left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This review was posted by an AI agent (OpenHands).

Summary

I reviewed the current head dbd6b272 (25 files, +3074/-1280): the unified prepare_agent_launch pipeline, the agent_server/agent_launch.py adapter, the Docker double-resolve in mediation.py/routers.py, the StartConversationRequest changes, and the new/updated tests.

The code itself holds up. I ran the focused suites locally against this head and they pass:

  • tests/sdk/profiles — 181 passed
  • tests/agent_server/test_agent_launch_parity.py, test_agent_launch_additions.py, test_acp_skill_sourcing.py — 25 passed
  • tests/agent_server/test_agent_profile_conv_start.py, test_conversation_router.py, test_conversation_service.py, tests/agent_server/docker_runtime — 317 passed

CI for dbd6b272 is green (52 checks, all success). I also verified by hand that the deprecated agent_settings path carries through fields no profile models (critic_api_key, user_message_suffix, agent_context.secrets, current_datetime=null, load_project_skills=false), that a scoped profile's secret filter is applied in apply_launch, and that the inline-profile validator now rejects a malformed draft at parse time (422) instead of a 500.

I am leaving a COMMENT rather than an approval for one reason: the PR description on this head describes a meta_profile_ref implementation that the head no longer contains.

Finding — the PR description contradicts the shipped contract, and hides a real Docker routing gap

The current body says:

prepare_agent_launch therefore loads the meta-profile and the LLM profiles it routes to; a caller that hydrated them itself (a cloud control plane, through base_settings) keeps its own copies.

and lists the 422 detail as:

{"code": "unresolved_profile_references", "message", "dangling_llm_profile_ref", "dangling_mcp_server_refs", "dangling_meta_profile_ref"}

Neither is true on dbd6b272. The final commit dbd6b272 ("address xhigh review of the unified launch") removed the hydration path and the dangling_meta_profile_ref key entirely. On this head:

  • OpenHandsAgentSettings.active_meta_profile is set to profile.meta_profile_ref (name only); meta_profile and meta_profile_llms are left unset.
  • UnresolvedProfileReferences.to_detail() emits only dangling_llm_profile_ref and dangling_mcp_server_refs; the meta_profile_ref parameter no longer exists (grep confirms zero remaining references to dangling_meta_profile_ref in the tree).
  • AgentLaunchCatalog.meta_profile_store and MetaProfileLoader are gone.

I verified the consequence for the case the body claims is fixed. A profile launch inside a Docker conversation container produces a route_task_to_model spec with params == {'active_meta_profile': 'pareto'} only, and the container's HOME/OH_PERSISTENCE_DIR is a per-conversation runtime dir (/var/openhands/.openhands, mounted from <runtime_dir>/persistence) that carries no meta-profiles/ directory and is never populated with the host's store. Reconstructing the tool the way the container does (ClassifyAndSwitchLLMTool.create(**spec["params"])) then fails at call time with:

FileNotFoundError: Active meta-profile 'pareto' could not be resolved from the
store or inline configuration. Available meta-profiles: none

So the routing tool a profile launch wires up cannot route inside Docker, which is exactly the failure the description says hydration was added to prevent. No test exercises a routing profile through the Docker prepare_start/finish_start path, so this is silent.

To be clear: this is not a regression against main — the base had no profile-level routing fields at all, and the author's 2026-09-28 comment discloses the Docker limitation as known. But the durable PR description still advertises hydration as the fix and still lists a 422 key that does not exist, and the REST detail schema is a public contract (the OpenAPI summary in the body is itself the artifact other clients read). A maintainer approving on that description would be approving a different change than the one at dbd6b272.

Please either (a) correct the description: drop "hydrates the meta-profile plus the LLM profiles it routes to" and dangling_meta_profile_ref, and state the Docker limitation plainly, or (b) if Docker routing is meant to work with a profile-carried meta_profile_ref, restore a store-independent path (e.g. hydrate into meta_profile/meta_profile_llms as agent_settings already does), with a test that drives it through the container launch.

Other notes (non-blocking)

  • warn_deprecated(..., deprecated_in="1.50.0") will not fire until the SDK ships 1.50.0, which is presumably the target release; the removal target 1.55.0 is similarly inert on this branch. Fine, just worth knowing.
  • The behavior changes the author enumerated for the agent_settings path (forced stream=True, load_project_skills=True, fresh current_datetime, browser injection when tools is null) are real and intentional; I confirmed each against the base. Worth a maintainer eyeball for eval movement, as the earlier review flagged.

🔄 CHANGES REQUESTED

@simonrosenberg

Copy link
Copy Markdown
Member Author

I'm closing this in favour of #5398, which replaces the design.

A review of the head commit (dbd6b27) found that this PR moves the launch paths behind one prepare_agent_launch, but nothing stops the next path or field from bypassing it. That is the underlying cause of #5141. The OpenAI-compatible gateway already bypasses it, and the #4287 meta-profile fields had to be added to the profile by hand. Fixing that needs a different structure, not more patches here.

Bugs at head:

  • Meta-profile routing fails outside the host. resolver.py:381 passes only the meta-profile name as active_meta_profile. It never loads the meta-profile or the LLMs it routes to. In a Docker container or cloud sandbox with no meta-profile store, the agent gets a route_task_to_model tool that fails on every call. The PR description's claim that the meta-profile and its LLMs are loaded at launch, and that a dangling ref returns dangling_meta_profile_ref, doesn't match the code.
  • Re-posting a running Docker conversation can stop it. docker_runtime/routers.py:155-177 runs finish_start and the /server_info probe after registry.get_or_create, and every failure path calls registry.stop(). A transient secret-lookup failure or probe timeout while re-posting a running conversation kills its container. Before this PR, that case returned a 422 before the container was touched.
  • On a Docker server, materialize and the launch disagree. materialize checks browser availability on the host, while the launch uses the container's answer. For a profile with tools: null, the preview and the real launch can differ. The parity test uses only explicit tools, so it misses this.

Design problems:

  • agent_settings is already a fully resolved settings object. Converting it into a fake profile and then restoring the lost fields through base_settings requires base and no-base branches in every builder.
  • For Docker, the host resolves the profile twice, then asks the container over HTTP for runtime facts the container already has.
  • Every non-FileNotFoundError LLM-store error hides the other dangling refs and becomes a 422, including transient lock timeouts.
  • StartConversationRequest(agent_settings=…).agent becomes None, and the new agent_settings serializer revalidates on every dump.

What #5398 does instead: it splits the launch into resolve and finalize. resolve runs where the stores live and returns a settings object with no references left in it. finalize runs in the process that runs the agent. agent_settings goes straight to finalize with no conversion, and Docker needs neither the probe nor a second resolution. Three tests enforce the split:

  • _start_event_service accepts only the value finalize returns.
  • An architecture test fails on any agent built outside the launch module.
  • A field-ownership test fails on any settings field the profile doesn't account for.

Some pieces here can be reused, and #5398 lists them: the SendMessageRequest move, UnresolvedProfileReferences, the inline agent_profile field, and .pr/launch_parity_e2e.py. I'm keeping the branch for that reason.

#5151 and #5315 refer to this PR. They should track #5398 instead.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Agent Profile] The default profile and named profiles build different agents — collapse launch into one pipeline

2 participants