Repository navigation
fix(harness): keep tool-call markup out of summaries and tool-less answers - #277
Conversation
When the model returns an empty string during summarization, the summarizer now returns a fallback message instead of failing. This prevents a panic or unhelpful error in cases where the model produces no output, ensuring the harness can continue processing gracefully. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a model returns an empty string, the summarization process now returns an empty summary instead of attempting to process the missing output. This prevents a panic or incorrect behavior when the model produces no text. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added unit tests for the model summarizer to verify its core functionality, including summarization of model responses and handling of edge cases. This ensures the summarization logic is reliable and covers expected behaviors. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test was calling `contains_tool_call_markup` through a re-export that no longer exists, so the calls have been updated to use the full `super::model_summarizer::contains_tool_call_markup` path instead. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The dialect parser previously split agent names on whitespace, causing multi-word names to be incorrectly parsed as separate tokens. This change updates the parsing logic to treat the entire agent name as a single token, ensuring that agents with spaces in their names are correctly identified and handled during loop execution. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the dialect field is empty, the agent loop now defaults to a standard configuration instead of failing. This change ensures backward compatibility with existing agents that do not specify a dialect. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters an empty dialect list, it now correctly returns an empty result instead of panicking. This fixes a crash that occurred when no dialects were configured for the agent. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters a dialect that is not recognized, it now returns an error instead of panicking. This change improves robustness by allowing the system to gracefully report the unsupported dialect rather than crashing. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters a dialect that is not recognized, it now returns an error instead of panicking. This change improves robustness by allowing the system to gracefully report the unsupported dialect rather than crashing. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters a dialect that is not recognized, it now returns an error instead of panicking. This change improves robustness by allowing the system to gracefully report unsupported dialects rather than crashing at runtime. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters an empty dialect list, it now correctly returns an empty result instead of panicking. This fixes a crash that occurred when no dialects were configured for the agent, ensuring graceful handling of this edge case. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The model call error handling in the agent loop now returns a proper error result instead of panicking, ensuring the loop can gracefully recover from transient model failures and continue processing. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a tool call has no arguments, the run loop now correctly passes an empty object instead of failing to parse the input, ensuring that tools with optional parameters continue to work as expected. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When an LLM returns a tool call with an empty arguments string, the agent loop now treats it as a no-op rather than attempting to parse it as JSON, which previously caused a panic. This change improves robustness against malformed or incomplete tool call responses from language models. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When an agent returns a tool call with an empty arguments string, the run loop now treats it as a valid invocation with no parameters rather than failing to parse the empty input. This prevents a crash in scenarios where the model produces a tool call without arguments. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters an empty list of actions, it now terminates gracefully instead of continuing to process. This prevents an infinite loop condition that could occur when no actions are available to execute. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When an agent returns a tool call with an empty arguments string, the run loop now treats it as a valid call with no parameters instead of failing to parse the JSON. This prevents a crash in scenarios where the model omits arguments for tools that accept none. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a tool call has no arguments, the agent loop now correctly passes an empty JSON object instead of failing to parse the missing field. This fixes a crash that occurred when the model returned a tool call without arguments. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the model returns an empty string, the summarizer now returns a fallback message instead of failing. This prevents a crash in downstream processing when no summary text is generated. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When an LLM returns a tool call with an empty arguments string, the agent loop now treats it as a valid call with no parameters instead of failing to parse the input. This prevents a crash in scenarios where the model omits arguments for tools that accept none. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The `offered` helper function in the recovery tests was not initializing the new `withhold` field on `TextRecovery`, which caused the struct to be constructed with an uninitialized value. Setting it to `false` ensures the test helper matches the current struct definition and prevents undefined behaviour. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the test expectations in the withheld agent loop tests to match the actual behavior of the run loop, ensuring that the tests correctly validate the agent's response when certain messages are withheld. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The `harness_with` helper function in the withholding tests now accepts `max_calls` as a `usize` instead of `u32` to match the type expected by the run policy configuration, eliminating a type mismatch that would cause compilation errors when passing the value to policy fields. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformat several lines and function signatures that exceeded the project's line length limit, wrapping them across multiple lines for improved readability. No functional changes are introduced. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
When an agent returns a tool call with an empty arguments string, the run loop now correctly passes an empty JSON object instead of failing to parse the empty string. This prevents a panic during tool execution when the model omits arguments for a tool that requires none. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added a README file to the summarization module to provide documentation and usage guidance for developers working with the harness. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a README file to document the agent loop module, providing an overview of its purpose and usage to help developers understand the module's role within the harness. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Tiny Sweeper review
Last completed reportTiny Sweeper reviewTiny Sweeper reviewed this change across 6 lane(s) and found 0 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below. State: Incomplete Review snapshot
Completeness: Incomplete What changedThe review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below. FeaturesNone identified with supported citations. TestsNo supported feature-to-test mapping was produced. Test execution is not inferred. FindingsNo active actionable findings. Could not review: crates/tinyagents-harness/src/agent_loop/README.md, crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/summarization/README.md, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs Before merge
How this fits togetherflowchart LR
n0["DeltaScrubber<br/>changed"]:::changed
n1["DroppedBlocks<br/>changed"]:::changed
n2["TextRecovery<br/>changed"]:::changed
n3["reset_truncated_empty_recovery<br/>changed"]:::changed
n4["run_loop_body"]:::impacted
n5["collect"]:::impacted
n6["new"]:::impacted
n7["invoke_model_streaming_once"]:::impacted
n8["invoke_model_with_retry"]:::impacted
n9["CallShape"]:::impacted
n0 -->|uses| n1
n2 -->|uses| n1
n4 -->|calls| n3
n4 -->|uses| n9
n6 -->|uses| n1
n6 -->|calls| n5
n7 -->|uses| n0
n7 -->|uses| n9
n8 -->|uses| n9
n9 -->|uses| n2
classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. 🧰 Additional context used📚 Code guidelines (1)No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (9)
Included review availability: This review used your included allowance. Your plan provides up to 2 included reviews per hour; 0 remain after this review. 📝 WalkthroughWalkthroughAgent-loop turns without callable tools now withhold recognized tool-call markup and can retry with a no-tools nudge. Model summarization now fences the transcript and retries replies containing recognized tool-call markup once before returning an error. ChangesAgent-loop tool-call withholding
Model summarization safeguards
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Bug fix Sequence Diagram(s)sequenceDiagram
participant RunLoop
participant Model
participant DeltaScrubber
RunLoop->>Model: Request a response with no callable tools
Model->>DeltaScrubber: Stream response text
DeltaScrubber->>RunLoop: Return text with tool-call markup withheld
RunLoop->>Model: Retry with no-tools nudge when limits permit
sequenceDiagram
participant ModelSummarizer
participant Model
participant FaultTolerantCachingSummarizer
ModelSummarizer->>Model: Request a summary
Model-->>ModelSummarizer: Reply containing tool-call markup
ModelSummarizer->>Model: Retry the summary request once
Model-->>ModelSummarizer: Return a second reply
ModelSummarizer-->>FaultTolerantCachingSummarizer: Return an error if markup remains
FaultTolerantCachingSummarizer-->>FaultTolerantCachingSummarizer: Use deterministic trimming on summarizer failure
Merge Risk: 🔵 Low · up to A call appearing only in the completed streaming response can be missed, leaving the lead-in as the answer without a retry. The risk is bounded but worth addressing or accepting before merge. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to The change reduces unintended actions and corrupted conversation history through bounded recovery. No new privilege escalation is established, but preserving withdrawn capabilities across retries still depends on caller-side controls. Retained concerns Security review detailsSecurity Blast Radius
Security Findings and Attack Paths
Trust Boundaries and Controls
Resilience and Maintainability Implications
Hardening Proposals
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
A rabbit watched the call tags fade, Comment |
There was a problem hiding this comment.
tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinyagents-harness/src/agent_loop/README.md, crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/summarization/README.md, crates/tinyagents-harness/src/summarization/model_summarizer.rs and 1 more.
$0.0049 · 43,041 in / 8,097 out · 14,080 cached (33%) · deepseek/deepseek-v4-flash
tests: $0.0020 · 14,231 in / 4,027 out · 0 cached (0%) · deepseek/deepseek-v4-flash
description: $0.0019 · 14,610 in / 83 out · 0 cached (0%) · deepseek/deepseek-v4-flash
…inyhumansai#276) Stacks the compaction Phase A follow-ups on both open PRs. Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ction-phase-a Keeps the summed summarizer usage in summarize_once and switches its markup check to tinytools_agent::contains_call_markup, as tinyhumansai#277 now does. Co-authored-by: Medulla <medulla@tinyhumans.ai>
Summary
DeepSeek V4 (via OpenRouter) writes its own tool-call markup,
<|DSML|invoke name="shell">…, as plain text whenever a request carries a tool-heavy transcript but offers no callable tool. Two harness paths then took that markup as a result:This PR prevents the leak where it can and contains it where it cannot.
Root cause
Found in the OpenHuman harness benchmark (DeepSWE pilot,
deepseek/deepseek-v4.1-flash, provider pinned to DeepSeek):contentzero times. OpenHuman leaked it 55 times in 1,616 calls.<tool_call …>blocks), ending on a tool result, with no closing instruction. The model's reasoning simply carried on the coding task ("Let me check ignoreDeadLinks first") and emitted a call.tool_choice: "none"does not help. Declaring the tools with"none"still leaked (4 of 8 replays).Fix
summarization/model_summarizer.rs)<transcript>…</transcript>, previous summary in<previous_summary>) and the instruction comes last.tinytools_agent::contains_call_markupflags as a tool call (feat(parse): add contains_call_markup tinytools#44) is retried once.FaultTolerantCachingSummarizerfalls back to its deterministic trim instead of keeping the markup.agent_loop/run_loop.rs,dialect.rs,model_call.rs): newTextRecovery::withholding()mode.DroppedBlocks::withheld. Reasoning is kept.WITHHELD_TOOL_CALL_NUDGE, which says tools are unavailable. It emitsControlApplied { control: "withheld_tool_call" }andRetryScheduled.RunPolicy::dropped_tool_call_nudgesand run before the empty-reply retries.Evidence (replays of the captured requests against OpenRouter / DeepSeek)
Same model, same pinned provider,
reasoning.effort = high, built byte for byte from the bench captures:tool_choice: "none"WITHHELD_TOOL_CALL_NUDGE(this PR)Tests
cargo test --workspace: all pass.cargo clippy --workspace --all-targets -- -D warnings: clean.cargo fmt --check: clean.agent_loop/run_loop_withheld_tests.rs(scrubbing, quoted examples untouched, re-prompt, no-budget fallback, streamed path) and four insummarization/model_summarizer_tests.rs.TextRecovery::default()is restored, so they cover the regression.Notes
Depends on feat(parse): add contains_call_markup tinytools#44.
vendor/tinytoolspoints at that PR's head, which addscontains_call_markup. Merge it first, then move this pin to the merge commit.The commit history includes automatic checkpoint commits with generic subjects; this description is the authoritative summary.
The compaction trigger loop (re-summarizing on every turn after the first compaction) is a separate issue, being fixed separately in
middleware/library/context.rs.