Skip to content

fix(agent): web research budget withdraws only web tools, not the whole toolset (#6959) - #6965

Merged
senamakel merged 11 commits into
tinyhumansai:mainfrom
senamakel:bench-6959-research-budget
Oct 3, 2026
Merged

senamakel merged 11 commits into
tinyhumansai:mainfrom
senamakel:bench-6959-research-budget

Conversation

@senamakel

@senamakel senamakel commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Once the direct web-read budget is used up, ResearchBudgetMiddleware now removes only the web tools (web_search_tool, web_answer_tool, web_contents_tool, web_fetch) and tells the model to continue the task with the tools it has left. Coding turns keep shell, apply_patch and every other tool.
  • The old close (clear every tool, set ToolChoice::None, send the "answer now" instruction) still runs, but only when no non-web tool is left.
  • Failed web calls (is_error) no longer use up the budget. After 2 failed web calls in a row, the next request gets a one-time note saying web access looks blocked.

Problem

The middleware treated "too many web reads" as "the turn is over". before_model ran request.tools.clear() and set tool_choice = None, which also took away shell and apply_patch. On DeepSWE ytt-jsonpath-query-api, the agent made 9 web reads (403s and 429s) while looking for an upstream implementation. The next request went out with tool_count 0, and the turn ended with an empty patch. after_tool also counted failed calls, so a sandbox that refuses every request used up the budget after 8 quick failures.

Solution

  • before_model filters the web tools out of request.tools with retain. If tools are left, it appends WEB_BUDGET_EXHAUSTED_INSTRUCTION ("...Continue the task with your remaining tools..."). A ToolChoice::Tool pinned to a removed web tool is reset to Auto. If nothing is left, it applies the existing RESEARCH_CLOSE_INSTRUCTION close unchanged, keeping the DSML-leak wording from fix(agent): keep DeepSeek tool-call markup out of summaries and tool-less answers #6946.
  • after_tool counts only successful web results. Failed ones increase a consecutive-failure streak, and a success resets it. At 2 failures the blocked note is queued once per run. before_model adds it to the next request only, as a message scoped to that request. It is not written to the transcript, which matches the PendingNudgeInjector pattern.
  • Messages and the tool filter apply to each request (the harness rebuilds the request from the transcript every call), so every later request in the turn stays narrowed the same way.
  • Each new branch has grep-friendly [tinyagents::mw] research_budget: logs.

Tradeoff and follow-up: web_fetch (tinytools) reports an HTTP 403 or 429 as a successful result (status=403 ...), so those responses still count toward the budget. This PR does not parse that text in the host. Making web_fetch return is_error for 4xx/5xx would belong in tinytools. Even so, the bench failure is fixed, because the agent keeps its non-web tools once the budget is used up.

Submission Checklist

  • Tests added or updated (happy path + at least one failure / edge case). middleware_research_budget_tests.rs: the test that encoded the bug now uses a web-only toolset. New tests cover: shell and apply_patch survive the 8th read on every later request; a pinned web tool choice is released; failed reads don't use up the budget; the blocked note is sent once after 2 consecutive failures; a success resets the streak.
  • Diff coverage ≥ 80%: every new branch is exercised by the unit tests above; CI coverage lane will verify.
  • Coverage matrix updated: N/A, behaviour-only change.
  • All affected feature IDs listed: N/A, agent harness middleware only.
  • No new external network dependencies introduced
  • Manual smoke checklist: N/A, not a release-cut surface.
  • Linked issue closed via Closes #6959

Impact

  • Agent runtime only (top-level runs; sub-agents don't install this middleware). Coding and other tool-using turns keep working after the web budget is used up. Pure lookup turns (web tools only) still have to answer.
  • Removing tools mid-turn changes the tools segment of the prompt cache. The old code had the same cost.

Related


AI Authored PR Metadata (required for Codex/Linear PRs)

Linear Issue

  • Key: N/A
  • URL: N/A

Commit & Branch

  • Branch: bench-6959-research-budget
  • Commit SHA: c8b7672

Validation Run

  • pnpm --filter openhuman-app format:check: N/A, no frontend change
  • pnpm typecheck: N/A, no TypeScript change
  • Focused tests: cargo test -p openhuman --lib research_budget_tests (8/8 pass); RUST_MIN_STACK=16777216 cargo test -p openhuman --lib agent::tinyagents (428 pass, 4 fail; the 4 failures already happen on upstream main: tool_output::tests::same_tool_calls_persist_artifacts_under_distinct_call_ids and 3 tool_output_artifact_tests)
  • Rust fmt/check (if changed): cargo fmt --all, cargo check --manifest-path Cargo.toml, pnpm rust:layout
  • Tauri fmt/check (if changed): N/A

Validation Blocked

  • command: N/A
  • error: N/A
  • impact: N/A

Behavior Changes

  • Intended behavior change: once the web budget is used up, only the web tools are removed. Failed web calls don't count. Repeated failures trigger a one-time "web access looks blocked" note.
  • User-visible effect: coding turns that hit web walls keep editing instead of ending with nothing done.

Parity Contract

  • Legacy behavior preserved: a toolset with only web tools still gets cleared tools, ToolChoice::None and RESEARCH_CLOSE_INSTRUCTION after 8 reads.
  • Guard/fallback/dispatch parity checks: web_only_research_concludes_after_eight_reads, the_concluding_instruction_says_tools_are_gone.

Duplicate / Superseded PR Handling

  • Duplicate PR(s): none
  • Canonical PR: this one
  • Resolution (closed/superseded/updated): N/A

Co-authored-by: Medulla medulla@tinyhumans.ai

Summary by CodeRabbit

  • Behavior Changes
    • After eight successful calls to direct web research tools in a top-level turn, those tools are no longer available. Other tools remain available; if none remain, the turn proceeds to an answer.
    • Failed web calls do not count toward the limit. After two consecutive failures, a one-time notice about blocked access appears in the next request. A successful web call resets the failure count. The notice is discarded if the web-tool limit is reached first.

senamakel and others added 6 commits October 3, 2026 15:35
Update the pinned commit of the tinyagents vendored dependency to incorporate upstream changes.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Update the pinned commit of the tinyagents submodule to incorporate the latest upstream changes.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the test assertions in the research budget middleware tests to match the actual behavior of the middleware. The previous assertions expected incorrect values or conditions, causing the tests to fail when run against the current implementation.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The research budget middleware now correctly enforces the configured budget limit by checking the accumulated cost against the maximum allowed value. Previously, the budget check was not applied, allowing research operations to exceed the specified limit.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
… detect blocked access

The research budget middleware now removes only web research tools when the budget is spent, leaving non-web tools available so the agent can continue the task. It also tracks consecutive failed web calls and injects a one-time blocked-access note after two failures, helping the agent avoid retrying a blocked web endpoint. The turn is forced to conclude only when no non-web tool remains after the web tools are withdrawn.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updates the research budget middleware to withdraw web tools after eight successful calls rather than forcing an answer, and adds a one-time note when consecutive web calls fail. The comment and documentation changes reflect that the budget now removes tools from the request so the run works with what it has, only forcing an answer when no other tool remains. Also clarifies that failed web calls (blocked network, offline sandbox) do not consume the budget.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@tinysweeper

tinysweeper Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 0 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Incomplete
Priority: none
Reviewed head: c94845eaedd7
Updated: 1791037591 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 3 Active findings 0
Tests 5 Noted findings 0
Documentation 1 Resolved findings 3
Configuration 0 Pending checks/questions 10

Completeness: Incomplete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

No active actionable findings.

Resolved this pass

  • Drop the doubled tool-results segment in the artifact pointer
  • Drop the doubled tool-results segment in the artifact pointer
  • Drop the doubled tool-results segment in the artifact pointer

Pending checks: Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS)

Could not review: crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs

Before merge

  • Complete the critique review for crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs.
  • Complete the security review for crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs, crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs.
  • Wait for Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS).

How this fits together

flowchart LR
  n0["...nnot_open_is_stored_as_the_processed_copy<br/>changed"]:::changed
  n1["tool_result"]:::impacted
  n2["expect"]:::impacted
  n3["artifact_mw"]:::impacted
  n4["summarized"]:::impacted
  n0 -->|calls| n1
  n0 -->|tests| n1
  n0 -->|calls| n2
  n0 -->|calls| n3
  n0 -->|tests| n3
  n0 -->|calls| n4
  n0 -->|tests| n4
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs
  • Lane summary: Reviewed 0 files; 0 findings. 3 files could not be reviewed: crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs.

security

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs, crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs
  • Lane summary: Reviewed 0 files; 0 findings. 3 files could not be reviewed: crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs, crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs.

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: This change updates the artifact-wiring tests to use the new `tool_result_artifacts_dir` import and fixes several hard-coded paths in tool-output artifact tests to include a `tool-results/` prefix, but does not introduce any new behavioural changes that would need a new test nor retract any earlier tests. The prior finding about the doubled tool-results segment in the artifact pointer was addressed in an earlier revision and is now fixed; no new problems are introduced, and the diff is safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 502 Bad Gateway: {"error":{"message":"no rung of ladder vectors could serve the request","skipped":[{"model":"text-embedding-bge-m3","provider":"venice","reason":"rate limited, retry in 25s","rung":0}],"type":"ladder_router_error"}}), so this review saw the diff alone._ _4 memory call(s) failed (model: cortex: v1/recall: timed out after 10s), so this review saw part of what the engine holds._

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: This revision correctly fixes the described test path changes to match the new artifact directory structure. The earlier finding about the doubled tool-results segment is resolved. No new issues are introduced in these commits. _Code retrieval was unavailable (model: ladder embeddings returned 502 Bad Gateway: {"error":{"message":"no rung of ladder vectors could serve the request","skipped":[{"model":"text-embedding-bge-m3","provider":"venice","reason":"rate limited, retry in 25s","rung":0}],"type":"ladder_router_error"}}), so this review saw the diff alone._ _4 memory call(s) failed (model: cortex: v1/recall: timed out after 10s), so this review saw part of what the engine holds._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: This change refines the research-budget middleware to withdraw only web tools after the budget is spent, rather than concluding the entire turn if non-web tools remain, and adds a blocked-web-access note. All behavioural changes are covered by Rust unit tests in the same commit; no end-to-end test is missing. The earlier finding about the doubled tool-results segment is fixed. Waiting on end-to-end jobs: `Rust E2E (mock backend)`, `Build Playwright E2E Artifact`, `E2E (Playwright / web lane)`, `Desktop E2E (full suite, 3 OS)`.
  • Unresolved questions/checks: Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS)
Evidence and run details
  • Models: deepseek/deepseek-v4-flash
  • Spend: $0.001818
  • Tokens: 60510 input · 2407 output · 512 cached · 0 embedding
Head State Pass summary
c8b767292e29 incomplete 0 active finding(s), 0 resolved finding(s) (at 1791032434)
310c6dbdef02 incomplete 1 active finding(s), 0 resolved finding(s) (at 1791035641)
c94845eaedd7 incomplete 0 active finding(s), 3 resolved finding(s) (at 1791037591)

tinysweeper 0.1.0

@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 9cdfaa84-b045-4d6d-968d-61f9250e259c
📥 Commits

Reviewing files that changed from the base of the PR and between b6e2c5e and c94845e.

📒 Files selected for processing (1)
  • crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The middleware counts successful calls to four direct web research tools. After eight successful calls, it removes web tools and preserves other tools. It forces an answer only when no tools remain. After two consecutive failures, it sends a one-time note.

Changes

Web research budget

Layer / File(s) Summary
Track web-call failures
crates/openhuman-core/src/agent/tinyagents/middleware/research_budget.rs, crates/openhuman-core/src/agent/tinyagents/middleware_research_budget_tests.rs
The middleware counts successful web calls, resets the failure streak after a success, and queues a one-time note after two consecutive failures. Tests cover failed calls, note delivery, and failure-streak reset.
Apply the exhausted web budget
crates/openhuman-core/src/agent/tinyagents/middleware/research_budget.rs, crates/openhuman-core/src/agent/tinyagents/middleware_research_budget_tests.rs, crates/openhuman-core/src/agent/tinyagents/harness_assembly.rs, crates/openhuman-core/src/agent/tinyagents/README.md
After eight successful web calls, the middleware removes web tools. It preserves other tools and switches a forced web-tool choice to automatic. When no tools remain, it disables tool calls and instructs the model to answer. Tests and documentation describe these behaviors.

Artifact test path updates

Layer / File(s) Summary
Update artifact test paths
crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs, crates/openhuman-core/src/agent/session_host/artifact_wiring.rs, crates/openhuman-core/src/agent/session_host/artifact_wiring_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs
Artifact tests use the updated store root and tool-results paths for pointer and persistence checks. The directory-helper import is removed from artifact wiring and added to its tests.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Bug fix · Severity of issue fixed: Medium

Suggested reviewers: al629176

Merge Risk: ⚪ Minimal · up to c9484

After eight successful web reads, coding runs retain their non-web tools, while web-only turns still conclude. Failed reads do not spend the budget, and the blocked-access note is request-scoped. The updated artifact lookup matches the existing storage layout; no actionable merge risk remains.

Architecture Summary

Architecture risk: 🔵 Low · up to c9484

The change affects 1 system.

Changed systems: crates

Architecture concerns
No architecture-level concerns identified.

Review details

Systems and components

  • observed — crates (service) was modified; 9 changed files map to changed impact.

Before / after behavior

  • observed — Modified behavior in crates/openhuman-core/src/agent/tinyagents/README.md: The description changes from forcing an answer after eight completed search or fetch calls to withdrawing web tools after eight successful calls and forcing an answer only when no other tool remains. It also adds a note about consecutive web-call failures.
  • observed — Modified behavior in crates/openhuman-core/src/agent/tinyagents/harness_assembly.rs: Updated the ResearchBudgetMiddleware comment: it now describes web tools leaving the request after enough results, with an answer when only web tools remain, rather than describing the next model call being used for synthesis. The comment continues to state that sub-agent budgets are unchanged; implementation is unchanged.
  • observed — Modified behavior in crates/openhuman-core/src/agent/tinyagents/middleware/research_budget.rs: The module documentation now describes web-only withdrawal and clarifies that the turn is forced to answer only when no other tool remains. The atomic imports add AtomicBool.
  • observed — Modified behavior in crates/openhuman-core/src/agent/tinyagents/middleware/research_budget.rs: The read limit remains 8. The module adds separate instructions for exhausting web access while other tools remain and for reporting blocked access, sets the blocked threshold to two consecutive failures, and centralizes identification of the four web tools. The middleware state adds failure tracking and one-time blocked-note flags.
🚥 Pre-merge checks | ✅ 2 | ❌ 3

❌ Failed checks (3 warnings)

Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning The middleware implements the core #6959 behavior: after eight is_error == false web results, it removes web tools, preserves non-web tools, resets a pinned web-tool choice to Auto, and uses the e… Ensure HTTP 403/429 responses do not count as successful reads and can trigger the repeated-failure note. Add tests for these responses.
Out of Scope Changes check ⚠️ Warning The changes to artifact storage paths and imports do not support #6959's research-budget behavior. They affect tool_result_artifacts/mod_tests.rs, session_host/artifact_wiring.rs, `session_host/ar… Remove the unrelated artifact-storage path and import changes from this PR, or provide evidence that they are required for the research-budget change.
Docstring Coverage ⚠️ Warning Docstring coverage is 70.83% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 24 functions across 7 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the main change: the research budget withdraws only web tools and preserves the remaining toolset.
Full details: Linked Issues check

Explanation

The middleware implements the core #6959 behavior: after eight is_error == false web results, it removes web tools, preserves non-web tools, resets a pinned web-tool choice to Auto, and uses the existing close only when no tools remain. The tests cover these paths and error results. However, web_fetch reports HTTP 403/429 responses with is_error == false. Those responses still increment the budget and do not trigger the blocked-access note, so repeated blocked HTTP responses do not meet #6959's failure-handling objective.

Full details: Out of Scope Changes check

Explanation

The changes to artifact storage paths and imports do not support #6959's research-budget behavior. They affect tool_result_artifacts/mod_tests.rs, session_host/artifact_wiring.rs, session_host/artifact_wiring_tests.rs, middleware_tool_output_artifact_tests.rs, and middleware/tool_output_tests.rs.

  • Fix all pre-merge checks with AI
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit counts each web-call trail,
Eight bright reads, then tools set sail.
Shell and patches stay in the den,
Two missed calls bring a note again.
Artifacts tuck in paths just right,
The rabbit hops through code tonight.

Comment @coderabbitai help to get the list of available commands.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/openhuman-core/src/agent/tinyagents/README.md, crates/openhuman-core/src/agent/tinyagents/harness_assembly.rs, crates/openhuman-core/src/agent/tinyagents/middleware/research_budget.rs, crates/openhuman-core/src/agent/tinyagents/middleware_research_budget_tests.rs.

             $0.0016 · 52,582 in / 3,114 out · 256 cached (0%) · deepseek/deepseek-v4-flash
tests:       $0.0004 · 13,191 in / 153 out   · 0 cached (0%)   · deepseek/deepseek-v4-flash
description: $0.0004 · 13,984 in / 839 out   · 0 cached (0%)   · deepseek/deepseek-v4-flash
e2e:         $0.0005 · 16,387 in / 180 out   · 256 cached (2%) · deepseek/deepseek-v4-flash

@tinysweeper tinysweeper Bot added the priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect. label Oct 3, 2026
coderabbitai[bot]
coderabbitai Bot previously approved these changes Oct 3, 2026
@senamakel senamakel self-assigned this Oct 3, 2026
…cts_dir

The import of `tool_result_artifacts_dir` from the security policy module was unused in the production artifact wiring code, so it has been removed. The corresponding import was added to the test file where it is actually referenced.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
senamakel and others added 2 commits October 3, 2026 16:48
…tory structure

Update test assertions in the tool result artifacts and middleware tool output artifact tests to use the correct path `tool-results/session` instead of just `session`, reflecting a change in how artifact storage directories are constructed. This ensures tests accurately verify that artifacts are persisted and read from the expected location.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformat three multi-line function calls in test files to single-line expressions, reducing line count without changing any behavior.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
coderabbitai[bot]
coderabbitai Bot previously approved these changes Oct 3, 2026
senamakel and others added 2 commits October 3, 2026 17:22
Updated the expected artifact path in `same_tool_calls_persist_artifacts_under_distinct_call_ids` to include the `artifacts/tool-results` prefix, reflecting a change in how tool output artifacts are stored.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test assertion path included an unnecessary "artifacts" directory segment that did not match the actual artifact storage layout, causing the test to fail when verifying persisted tool outputs. The path now correctly points to the tool-results directory.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs.

             $0.0018 · 60,510 in / 2,407 out · 512 cached (1%) · deepseek/deepseek-v4-flash
tests:       $0.0004 · 14,735 in / 142 out   · 0 cached (0%)   · deepseek/deepseek-v4-flash
description: $0.0004 · 15,570 in / 72 out    · 512 cached (3%) · deepseek/deepseek-v4-flash
e2e:         $0.0005 · 19,193 in / 120 out   · 0 cached (0%)   · deepseek/deepseek-v4-flash

@senamakel
senamakel merged commit 61b5c07 into tinyhumansai:main Oct 3, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

openhuman: web research budget clears every tool and forces a final answer, ending coding turns with no edits

1 participant