Repository navigation
fix(agent): web research budget withdraws only web tools, not the whole toolset (#6959) - #6965
Conversation
Update the pinned commit of the tinyagents vendored dependency to incorporate upstream changes. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Update the pinned commit of the tinyagents submodule to incorporate the latest upstream changes. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the test assertions in the research budget middleware tests to match the actual behavior of the middleware. The previous assertions expected incorrect values or conditions, causing the tests to fail when run against the current implementation. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The research budget middleware now correctly enforces the configured budget limit by checking the accumulated cost against the maximum allowed value. Previously, the budget check was not applied, allowing research operations to exceed the specified limit. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
… detect blocked access The research budget middleware now removes only web research tools when the budget is spent, leaving non-web tools available so the agent can continue the task. It also tracks consecutive failed web calls and injects a one-time blocked-access note after two failures, helping the agent avoid retrying a blocked web endpoint. The turn is forced to conclude only when no non-web tool remains after the web tools are withdrawn. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updates the research budget middleware to withdraw web tools after eight successful calls rather than forcing an answer, and adds a one-time note when consecutive web calls fail. The comment and documentation changes reflect that the budget now removes tools from the request so the run works with what it has, only forcing an answer when no other tool remains. Also clarifies that failed web calls (blocked network, offline sandbox) do not consume the budget. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Tiny Sweeper reviewTiny Sweeper reviewed this change across 6 lane(s) and found 0 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below. State: Incomplete Review snapshot
Completeness: Incomplete What changedThe review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below. FeaturesNone identified with supported citations. TestsNo supported feature-to-test mapping was produced. Test execution is not inferred. FindingsNo active actionable findings. Resolved this pass
Pending checks: Rust E2E (mock backend), Build Playwright E2E Artifact, E2E (Playwright / web lane), Desktop E2E (full suite, 3 OS) Could not review: crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs Before merge
How this fits togetherflowchart LR
n0["...nnot_open_is_stored_as_the_processed_copy<br/>changed"]:::changed
n1["tool_result"]:::impacted
n2["expect"]:::impacted
n3["artifact_mw"]:::impacted
n4["summarized"]:::impacted
n0 -->|calls| n1
n0 -->|tests| n1
n0 -->|calls| n2
n0 -->|calls| n3
n0 -->|tests| n3
n0 -->|calls| n4
n0 -->|tests| n4
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe middleware counts successful calls to four direct web research tools. After eight successful calls, it removes web tools and preserves other tools. It forces an answer only when no tools remain. After two consecutive failures, it sends a one-time note. ChangesWeb research budget
Artifact test path updates
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Bug fix · Severity of issue fixed: Medium Suggested reviewers: Merge Risk: ⚪ Minimal · up to After eight successful web reads, coding runs retain their non-web tools, while web-only turns still conclude. Failed reads do not spend the budget, and the blocked-access note is request-scoped. The updated artifact lookup matches the existing storage layout; no actionable merge risk remains. Architecture SummaryArchitecture risk: 🔵 Low · up to The change affects 1 system. Changed systems: Architecture concerns Review detailsSystems and components
Before / after behavior
🚥 Pre-merge checks | ✅ 2 | ❌ 3❌ Failed checks (3 warnings)
✅ Passed checks (2 passed)
Full details: Linked Issues checkExplanation The middleware implements the core Full details: Out of Scope Changes checkExplanation The changes to artifact storage paths and imports do not support
A rabbit counts each web-call trail, Comment |
There was a problem hiding this comment.
tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/openhuman-core/src/agent/tinyagents/README.md, crates/openhuman-core/src/agent/tinyagents/harness_assembly.rs, crates/openhuman-core/src/agent/tinyagents/middleware/research_budget.rs, crates/openhuman-core/src/agent/tinyagents/middleware_research_budget_tests.rs.
$0.0016 · 52,582 in / 3,114 out · 256 cached (0%) · deepseek/deepseek-v4-flash
tests: $0.0004 · 13,191 in / 153 out · 0 cached (0%) · deepseek/deepseek-v4-flash
description: $0.0004 · 13,984 in / 839 out · 0 cached (0%) · deepseek/deepseek-v4-flash
e2e: $0.0005 · 16,387 in / 180 out · 256 cached (2%) · deepseek/deepseek-v4-flash
…cts_dir The import of `tool_result_artifacts_dir` from the security policy module was unused in the production artifact wiring code, so it has been removed. The corresponding import was added to the test file where it is actually referenced. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…tory structure Update test assertions in the tool result artifacts and middleware tool output artifact tests to use the correct path `tool-results/session` instead of just `session`, reflecting a change in how artifact storage directories are constructed. This ensures tests accurately verify that artifacts are persisted and read from the expected location. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformat three multi-line function calls in test files to single-line expressions, reducing line count without changing any behavior. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the expected artifact path in `same_tool_calls_persist_artifacts_under_distinct_call_ids` to include the `artifacts/tool-results` prefix, reflecting a change in how tool output artifacts are stored. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test assertion path included an unnecessary "artifacts" directory segment that did not match the actual artifact storage layout, causing the test to fail when verifying persisted tool outputs. The path now correctly points to the tool-results directory. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
There was a problem hiding this comment.
tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/openhuman-core/src/agent/harness/tool_result_artifacts/mod_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware/tool_output_tests.rs, crates/openhuman-core/src/agent/tinyagents/middleware_tool_output_artifact_tests.rs.
$0.0018 · 60,510 in / 2,407 out · 512 cached (1%) · deepseek/deepseek-v4-flash
tests: $0.0004 · 14,735 in / 142 out · 0 cached (0%) · deepseek/deepseek-v4-flash
description: $0.0004 · 15,570 in / 72 out · 512 cached (3%) · deepseek/deepseek-v4-flash
e2e: $0.0005 · 19,193 in / 120 out · 0 cached (0%) · deepseek/deepseek-v4-flash
Summary
ResearchBudgetMiddlewarenow removes only the web tools (web_search_tool,web_answer_tool,web_contents_tool,web_fetch) and tells the model to continue the task with the tools it has left. Coding turns keepshell,apply_patchand every other tool.ToolChoice::None, send the "answer now" instruction) still runs, but only when no non-web tool is left.is_error) no longer use up the budget. After 2 failed web calls in a row, the next request gets a one-time note saying web access looks blocked.Problem
The middleware treated "too many web reads" as "the turn is over".
before_modelranrequest.tools.clear()and settool_choice = None, which also took awayshellandapply_patch. On DeepSWEytt-jsonpath-query-api, the agent made 9 web reads (403s and 429s) while looking for an upstream implementation. The next request went out withtool_count 0, and the turn ended with an empty patch.after_toolalso counted failed calls, so a sandbox that refuses every request used up the budget after 8 quick failures.Solution
before_modelfilters the web tools out ofrequest.toolswithretain. If tools are left, it appendsWEB_BUDGET_EXHAUSTED_INSTRUCTION("...Continue the task with your remaining tools..."). AToolChoice::Toolpinned to a removed web tool is reset toAuto. If nothing is left, it applies the existingRESEARCH_CLOSE_INSTRUCTIONclose unchanged, keeping the DSML-leak wording from fix(agent): keep DeepSeek tool-call markup out of summaries and tool-less answers #6946.after_toolcounts only successful web results. Failed ones increase a consecutive-failure streak, and a success resets it. At 2 failures the blocked note is queued once per run.before_modeladds it to the next request only, as a message scoped to that request. It is not written to the transcript, which matches thePendingNudgeInjectorpattern.[tinyagents::mw] research_budget:logs.Tradeoff and follow-up:
web_fetch(tinytools) reports an HTTP 403 or 429 as a successful result (status=403 ...), so those responses still count toward the budget. This PR does not parse that text in the host. Makingweb_fetchreturnis_errorfor 4xx/5xx would belong in tinytools. Even so, the bench failure is fixed, because the agent keeps its non-web tools once the budget is used up.Submission Checklist
middleware_research_budget_tests.rs: the test that encoded the bug now uses a web-only toolset. New tests cover:shellandapply_patchsurvive the 8th read on every later request; a pinned web tool choice is released; failed reads don't use up the budget; the blocked note is sent once after 2 consecutive failures; a success resets the streak.Closes #6959Impact
Related
web_fetchmark HTTP error statuses asis_errorso 403/429 walls stop counting.AI Authored PR Metadata (required for Codex/Linear PRs)
Linear Issue
Commit & Branch
Validation Run
pnpm --filter openhuman-app format:check: N/A, no frontend changepnpm typecheck: N/A, no TypeScript changecargo test -p openhuman --lib research_budget_tests(8/8 pass);RUST_MIN_STACK=16777216 cargo test -p openhuman --lib agent::tinyagents(428 pass, 4 fail; the 4 failures already happen on upstream main:tool_output::tests::same_tool_calls_persist_artifacts_under_distinct_call_idsand 3tool_output_artifact_tests)cargo fmt --all,cargo check --manifest-path Cargo.toml,pnpm rust:layoutValidation Blocked
command:N/Aerror:N/Aimpact:N/ABehavior Changes
Parity Contract
ToolChoice::NoneandRESEARCH_CLOSE_INSTRUCTIONafter 8 reads.web_only_research_concludes_after_eight_reads,the_concluding_instruction_says_tools_are_gone.Duplicate / Superseded PR Handling
Co-authored-by: Medulla medulla@tinyhumans.ai
Summary by CodeRabbit