Skip to content

fix(agent): keep typed hosted failures and honor transcript-autoload suppression; un-ignore harness e2e tests - #6879

Merged
senamakel merged 21 commits into
tinyhumansai:mainfrom
senamakel:fu-bug-harness
Oct 1, 2026
Merged

senamakel merged 21 commits into
tinyhumansai:mainfrom
senamakel:fu-bug-harness

Conversation

@senamakel

@senamakel senamakel commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Summary

Root-causes the ignored and failing tests in agent_harness_e2e and agent_turn_overrides_e2e. Two are real product bugs, one is a fixture that described a stream no provider produces, and the rest are tests that predate intended product changes.

Closes #6375
Closes #6377

Root causes and fixes

  1. streaming_tool_call_accumulation (Hosted TinyAgents stream regressions lose turn continuation and typed failures #6375): the 3x dispatch was the fixture, not the harness. ScriptedProvider.stream replayed the same ToolCallDelta fragments before every response, including the text-only final answer. The hosted harness deliberately treats streamed tool fragments as authoritative and rebuilds the terminal response's tool calls from them (invoke_model_streaming_once, agent_loop/model_call.rs). So the final "stream final" response also dispatched echo_tool: three identical successful batches, then the repeat guard stopped the turn. The fixture now streams fragments only for a response that carries tool calls. The test and its provider-level SSE sibling (provider_sse_tool_args_accumulation) had been dropped from main by a merge resolution, so they are restored here, with no #[ignore].
  2. model_call_ceiling_... (Hosted TinyAgents stream regressions lose turn continuation and typed failures #6375): the hosted boundary dropped the failure kind. The host drained invoke_agent_stream, whose terminal Failed item is sanitized to one string (hosted agent invocation failed) for every failure, then wrapped it as TinyAgentsError::Model. A per-call timeout was therefore indistinguishable from a provider error and classified as inference. The intended contract is turn_timeout (is_turn_timeout_error, timeout_bound_tag). The root fix is upstream, tinyagents#261: a public invoke_agent_streaming returning the typed HostedError. The host now uses it for both surfaces and converts with the existing From<HostedError> for TinyAgentsError, so the kind survives. The test's < 8s bound is replaced because a per-call ceiling is now a retryable CallTimeout that the turn policy retries on its 5-attempt schedule (Upstream 429 on the selected model kills the turn after one retry — no backoff, no fallback, no user-visible retry #6413), so the turn ends after about 50s. It now pins turn_timeout, at least 2 upstream calls, and < 90s.
  3. Orchestrator tests: stale, not a regression. perf(orchestrator): cut the first-turn prompt from 13.3k to 5.3k tokens #6787 (f6bffc219f, "defer rarely-used tools and pack setup_skills hand-off") deliberately moved setup_skills into the skills pack, deferred the four mcp_registry_* tools off the orchestrator's wire, and made a packed hand-off stop closing its pack to the orchestrator. Updated to that documented intent:
    • orchestrator_hands_skill_installs_to_skill_setup_directly becomes ..._through_the_skills_pack: the hand-off is called as use_skill {skills, setup_skills}.
    • orchestrator_cannot_install_a_skill_through_the_raw_registry_tool becomes orchestrator_raw_skill_install_through_use_skill_needs_approval: the guard is now the approval gate, and a denied install writes nothing.
    • orchestrator_advertises_direct_mcp_tools becomes orchestrator_defers_its_mcp_registry_tools_without_a_hand_off.
  4. Turn override E2E fixtures need hosted root authority #6377, two more causes.
    • suppress_transcript_autoload is a real product bug. turn() ran the thread-bound explicit resume (runtime.resume) before the lifecycle hook applied begin_turn_resume, so the override suppressed nothing for a thread-bound session. The flows builder and others pass a thread id together with this override. The resume mode is now decided before the explicit resume (OpenHumanSessionHost::turn_resume_mode). The test's old control (a different thread id replays another thread's transcript) described the removed latest-by-agent-name lookup. It now pins that another thread never gets the transcript, that the same thread resumes (the control), and that the override makes the same-thread turn start clean.
    • turn_overrides_apply_to_exactly_one_turn_and_then_reset is a fixture issue. ModelProfile::default() has no native tool calling, so the hosted harness renders the belt into the prompt and sends an empty request.tools. The test now uses ScriptedModel::native_tools.

Pin

vendor/tinyagents -> 4583fd645c2426a9484df429493c9d16dd8bc4fa (branch fix-hosted-stream-typed-failure, tinyagents#261, based on the previous pin b88a2728). Merge #261 first; the gitlink then needs a repoint to the merged main SHA.

Tests

  • RED before: streaming_tool_call_accumulation (repeat-guard text instead of "stream final"), model_call_ceiling (error_type=inference), the three orchestrator tests, suppress_transcript_autoload... (prior marker present in the suppressed prompt), as run on this base.
  • GREEN after, with scripts/test-rust-with-mock.sh:
    • --test agent_harness_e2e: 27 passed, 0 ignored
    • --test agent_turn_overrides_e2e: 4 passed, 0 ignored
  • cargo test -p openhuman --lib --features <product>: 9033 passed, 0 failed, 39 ignored.
  • cargo check --tests -p openhuman -p openhuman-cli --features <product> clean.
  • tinyagents: cargo test -p tinyagents-harness --lib runtime:: (47 passed) and clippy -D warnings clean.
  • CI scripts:
    • check-gated-test-allowlist and check-feature-forwarding: OK.
    • check-submodule-monotonic: OK.
    • check-agent-runtime-boundary: only the 2 pre-existing stale openhuman_backend_model baseline entries.
    • rust:layout: only the pre-existing failures in all_tests.rs and openhuman_backend_model*.rs. The runtime_session.rs pin is lowered to 1966.
    • check-ignored-tests: tests baseline 22 -> 17. openhuman-core is 39 against a baseline of 38 on main already, which is not from this PR.

Notes

  • The pre-push hook reformatted unrelated files. Those were reverted from this PR and the push used --no-verify.
  • On the hosted path the harness sanitizes the timeout message, so the Sentry per_model_call vs run_remaining tag (timeout_bound_tag) can no longer see which bound fired and reports unclassified_timeout. The user-facing class is correct. Restoring the tag needs the kind or bound carried on HostedError; it is not done here.

Summary by CodeRabbit

  • Bug Fixes
    • Transcript resumption now stays within the bound thread, and the suppress-autoload override prevents replaying that thread’s transcript.
    • Hosted turns preserve more specific error types when tool-call streaming is enabled or disabled.
  • Improvements
    • Streamed tool-call argument updates are covered across scripted and OpenAI-compatible providers, including arguments delivered in fragments.
    • Skill hand-off and registry tool access now have broader end-to-end coverage.

senamakel and others added 21 commits October 1, 2026 11:07
Add two end-to-end tests for streaming tool-call argument accumulation. The first test drives the engine and UI delta forwarding through a ScriptedProvider, asserting that the tool receives the full assembled argument set from ModelResponse.message.tool_calls while the progress channel carries the individual ToolCallArgsDelta fragments. The second test exercises the real provider HTTP and SSE-parse path against an in-test upstream, verifying that the provider's accumulation buffer correctly reassembles function.arguments fragments split across SSE chunks at awkward byte offsets.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ion test

The `#[ignore]` attribute was removed from the `streaming_tool_call_accumulation` test function, allowing it to run as part of the normal test suite instead of requiring the `--ignored` flag.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added eprintln! calls in the ScriptedProvider test helper to log the remaining response count and incoming message count during streaming tests, making it easier to diagnose test failures by observing the sequence of calls at runtime.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Remove debug print statements and conditionally replay stream events only when the response contains tool calls. Without this guard, the test harness would emit tool-call fragments ahead of a text-only final answer, which no real provider produces and which causes the agent loop to re-dispatch the call on every iteration until the repeat guard stops the turn.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ol call test

The comment describing why the test filters to iteration 1 was misleading, as it suggested iteration 2 also emits the same delta sequence. The ScriptedProvider actually replays stream events only for the response carrying the tool call, so only iteration 1 emits deltas and the filter simply pins that iteration.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The `is_turn_timeout_error` function now also matches the `exceeded its per-model-call ceiling` phrase, so that a wedged model call that exhausts its retries is classified as a turn timeout rather than falling through to the generic inference bucket. A test case and the corresponding end-to-end test are updated to reflect this fix for issue tinyhumansai#6375.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The per-model-call ceiling error is now handled as a retryable `CallTimeout` rather than a turn timeout, so the anchor string and its associated test case have been removed to prevent misclassification.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
… use_skill

The end-to-end test for skill install hand-offs was updated to reflect that `setup_skills` is now a member of the `skills` tool pack rather than a directly advertised tool. The helper function was renamed and extended to accept a pack parameter, and the assertions now check that the orchestrator offers the pack through `use_skill` instead of advertising the hand-off tool directly. The test function and its inner async counterpart were also renamed to match the new behaviour.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…_skill

Update the orchestrator e2e test to reflect that since tinyhumansai#6787, `setup_skills` is a member of the `skills` pack and a packed hand-off no longer closes its pack to the orchestrator. The raw `skill_registry_install` is now reachable through `use_skill`, but guarded by the approval gate, so the test now denies the approval prompt instead of approving it and verifies the install is refused.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Rename the test that verifies the orchestrator defers its MCP registry tools without a hand-off, and update its doc comment to clarify that the tools are not advertised on the orchestrator's wire due to issue tinyhumansai#6787.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Replace the manual stream-draining logic with a call to `invoke_agent_streaming` and unify error handling for both streaming and non-streaming paths. Previously, the streaming branch consumed the agent stream item by item and collapsed all failures into a sanitized string, losing the original error kind. The new code keeps the typed `HostedError` through the turn error, allowing `web_errors` to classify timeouts, limits, and provider failures correctly.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Two end-to-end tests for turn-level overrides were previously ignored because the hosted root session did not carry tool-calling capability, causing the tests to fail on their own control assertions. A new `ScriptedModel::native_tools` constructor now creates a model profile that declares native tool calling, which lets the hosted harness send the toolbelt schema upstream. The `suppress_transcript_autoload` and `turn_overrides_reset` tests are un-ignored and updated to use this constructor, restoring coverage for the scenarios they were designed to validate.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…d-scoped session resume

The test for `suppress_transcript_autoload` is updated to reflect that `turn()` now resumes by durable session identity rather than by agent name, making the lookup thread-scoped. The test is restructured into three clear assertions: a different thread never replays another thread's transcript, the same thread resumes its own transcript as a control, and the suppression override makes the same-thread turn start clean.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The baseline count of ignored tests was reduced from 22 to 17, reflecting a decrease in the number of tests that are expected to be skipped during CI runs.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test for per-call ceiling behaviour now accounts for the retry schedule introduced in tinyhumansai#6413, where a wedged call is retried up to five times with exponential backoff rather than ending immediately after the first ceiling cut. The assertion on elapsed time is relaxed from 8 seconds to 90 seconds, and a new assertion verifies that the upstream was called at least twice, confirming the ceiling triggered retries.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a runtime session attempts to transition to a new state, the code now checks whether a session host is present before proceeding. This prevents a panic that occurred when the host was unexpectedly absent, ensuring graceful handling of edge cases in session lifecycle management.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a runtime session is not found in the session host, the system now returns an appropriate error instead of panicking or proceeding with an invalid state. This ensures graceful failure and clearer diagnostics for missing session scenarios.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Replace the `transcript_autoload_suppressed` method with `turn_resume_mode` that returns a `ResumeMode` enum, enabling more nuanced control over how session history is resumed. This change addresses issue tinyhumansai#6377 where the previous boolean approach could not distinguish between suppressing autoload and other resume strategies needed for thread-bound sessions.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The line-count exception for the runtime_session.rs file in the OpenHuman Rust layout check was reduced from 1979 to 1966 to reflect the removal of generic session state code that was moved to the tinyagents-runtime crate, keeping the allowance exact for the remaining composition logic.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformat several long string literals and function call arguments to stay within the project's line-length limit, and reorder `#[cfg(test)] mod` declarations in a handful of files so that the module attributes appear immediately before their corresponding module name. No behaviour changes are introduced.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformatted several multi-line string literals and error messages to fit on single lines, and reordered `#[cfg(test)]` module declarations to follow a consistent pattern. These changes are purely cosmetic with no behavioral impact.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@tinysweeper

tinysweeper Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

⚠️ Review failed for 7a77a6dfabc2. the review of #6879 did not finish within 900s

@coderabbitai

coderabbitai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 5dc8da8f-476f-4d2e-9996-7252f6b4fb92

📥 Commits

Reviewing files that changed from the base of the PR and between 18470c8 and 7a77a6d.

📒 Files selected for processing (8)
  • crates/openhuman-core/src/agent/session_host/runtime_session.rs
  • crates/openhuman-core/src/agent/session_host/types.rs
  • crates/openhuman-core/src/agent/tinyagents/turn_runner.rs
  • scripts/ci/check-openhuman-rust-layout.mjs
  • scripts/ci/ignored-test-baseline.json
  • tests/agent_harness_e2e.rs
  • tests/agent_turn_overrides_e2e.rs
  • vendor/tinyagents

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 6 remain after this review.


📝 Walkthrough

Walkthrough

Hosted turns now select transcript resume modes through a dedicated method and preserve typed invocation errors. End-to-end tests cover transcript overrides, streamed tool arguments, retry-aware timeouts, packed skill hand-offs, and MCP tool advertisement.

Changes

Hosted turns and harness behavior

Layer / File(s) Summary
Transcript resume selection and override tests
crates/openhuman-core/src/agent/session_host/{runtime_session.rs,types.rs}, tests/agent_turn_overrides_e2e.rs, scripts/ci/ignored-test-baseline.json, scripts/ci/check-openhuman-rust-layout.mjs
Resume-mode selection now accounts for transcript-autoload suppression, session binding, and runtime history. Tests cover different-thread isolation, same-thread resumption, suppression, and one-shot overrides. The layout pin and ignored-test baseline were updated.
Hosted invocation and streaming coverage
crates/openhuman-core/src/agent/tinyagents/turn_runner.rs, vendor/tinyagents, tests/agent_harness_e2e.rs
Hosted turns use streaming or non-streaming invocation APIs and retain typed hosted errors. Tests cover streamed tool-argument accumulation, provider SSE handling, and retry-aware timeout behavior.
Packed skill hand-offs and MCP advertisement
tests/agent_harness_e2e.rs
Tests check packed hand-offs through use_skill, denied raw skill installation, and deferral of MCP registry tools from the orchestrator’s advertised tools.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Bug fix · Severity of issue fixed: Medium

Suggested reviewers: m3ga-mind

Merge Risk: ⚪ Minimal · up to 7a77a

The changes preserve hosted error classifications and honor transcript suppression, with restored regression coverage. No actionable blocker remains; merge after normal checks.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 7a77a

The change strengthens conversation-history suppression and preserves failure categories. No new security weakness was established, but interruption handling and shared-session persistence remain only partially verified.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The security-relevant surface is persisted conversation content entering a turn's prompt, together with failure information and commit-driven downstream consumers. The inspected host change reduces transcript-loading reachability under suppression; it does not demonstrate an expansion of tool authority. Maximum tenant and deployment exposure is not established by the available evidence.

Trust Boundaries and Controls

  • observed — A bound thread uses a scoped session identity, and the suppression check now precedes transcript restoration and adoption of recorded tool declarations. This addresses the prior control-ordering defect. Session identity selection itself is not evidence that externally supplied thread identities are authorized.

Resilience and Maintainability Implications

  • inferred — Receipt-gated host side effects and sequential override-reset assertions support failure containment and one-turn control ownership. They do not resolve whether the dependency preserves prior durable history after a suppressed turn, commits partial state on failure or interruption, serializes concurrent hosts, or preserves sanitized error payloads during typed conversion.
🚥 Pre-merge checks | ✅ 3 | ❌ 2

❌ Failed checks (2 warnings)

Check name Status Explanation Resolution
Out of Scope Changes check ⚠️ Warning The PR also changes behavior and tests that are not required by #6375. suppress_transcript_autoload handling and the restored agent_turn_overrides_e2e coverage target the closed, historical #6377 … Remove the unrelated transcript-autoload and turn-override behavior and tests from this PR, or link them to an active issue with matching coding requirements. Move the skill-pack, registry-install, MCP, and orchestrator tool-pack updates to…
Docstring Coverage ⚠️ Warning Docstring coverage is 43.40% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 53 functions across 6 files. (2 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely summarizes the main changes: preserving typed hosted failures, honoring transcript-autoload suppression, and restoring E2E tests.
Linked Issues check ✅ Passed The PR meets the active coding requirements in #6375. The hosted turn path now uses invoke_agent_streaming when progress streaming is enabled and preserves typed HostedError values. The restored t…
Full details: Out of Scope Changes check

Explanation

The PR also changes behavior and tests that are not required by #6375. suppress_transcript_autoload handling and the restored agent_turn_overrides_e2e coverage target the closed, historical #6377 issue. The skill hand-off, raw registry install, MCP deferred-tool, and orchestrator tool-pack test updates address separate tool-pack behavior. These changes are not needed to restore the three hosted regressions in #6375.

Resolution

Remove the unrelated transcript-autoload and turn-override behavior and tests from this PR, or link them to an active issue with matching coding requirements. Move the skill-pack, registry-install, MCP, and orchestrator tool-pack updates to a separate scoped PR or provide an active linked issue that requires them.

Full details: Docstring Coverage

Explanation

Docstring coverage is 43.40% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 53 functions across 6 files. (2 skipped: 2 unsupported.)

  • Fix all pre-merge checks with AI
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


A rabbit watches arguments stream,
Four small fragments form one beam.
A thread resumes, or starts anew,
A skill pack takes the offered route.
Typed errors hop along the way,
Then carrots close the review day.

Comment @coderabbitai help to get the list of available commands.

✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR

@senamakel
senamakel merged commit b34cc6e into tinyhumansai:main Oct 1, 2026
30 of 33 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Turn override E2E fixtures need hosted root authority Hosted TinyAgents stream regressions lose turn continuation and typed failures

1 participant