Skip to content

fix(harness): keep tool-call markup out of summaries and tool-less answers - #277

Merged
senamakel merged 27 commits into
tinyhumansai:mainfrom
senamakel:deepseek-markup-leak
Oct 3, 2026
Merged

senamakel merged 27 commits into
tinyhumansai:mainfrom
senamakel:deepseek-markup-leak

Conversation

@senamakel

@senamakel senamakel commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Summary

DeepSeek V4 (via OpenRouter) writes its own tool-call markup, <|DSML|invoke name="shell">…, as plain text whenever a request carries a tool-heavy transcript but offers no callable tool. Two harness paths then took that markup as a result:

  • The context summarizer installed it as the compacted summary, so one stray command replaced the whole compacted history.
  • Concluding calls that withdraw tools (final wrap-up, a host's research budget) returned it as the turn's final answer.

This PR prevents the leak where it can and contains it where it cannot.

Root cause

Found in the OpenHuman harness benchmark (DeepSWE pilot, deepseek/deepseek-v4.1-flash, provider pinned to DeepSeek):

  • The leak was specific to OpenHuman. Across about 5,500 responses from the other harnesses on the same model and provider, the markup leaked into content zero times. OpenHuman leaked it 55 times in 1,616 calls.
  • Every leak was on a request with no callable tool. 54 were summarizer calls and the rest were tool-withdrawn concluding calls. With no tool channel open, the provider does not parse the model's call, so the raw markup arrives as text.
  • The summarizer request reads like a live agent loop. It was the bare rendered transcript (tool calls as <tool_call …> blocks), ending on a tool result, with no closing instruction. The model's reasoning simply carried on the coding task ("Let me check ignoreDeadLinks first") and emitted a call.
  • tool_choice: "none" does not help. Declaring the tools with "none" still leaked (4 of 8 replays).

Fix

  1. Summarizer (summarization/model_summarizer.rs)
    • The transcript is fenced (<transcript>…</transcript>, previous summary in <previous_summary>) and the instruction comes last.
    • A reply that tinytools_agent::contains_call_markup flags as a tool call (feat(parse): add contains_call_markup tinytools#44) is retried once.
    • If the retry is markup too, it returns an error, so FaultTolerantCachingSummarizer falls back to its deterministic trim instead of keeping the markup.
  2. Turns with no callable tool (agent_loop/run_loop.rs, dialect.rs, model_call.rs): new TextRecovery::withholding() mode.
    • Any complete call the grammars recognise is scrubbed from the visible text, both from streamed deltas and from unary replies, and counted in DroppedBlocks::withheld. Reasoning is kept.
    • Nothing is dispatched, so the I-2 rule is preserved.
    • If a model call remains, the loop drops that assistant row and re-prompts once with WITHHELD_TOOL_CALL_NUDGE, which says tools are unavailable. It emits ControlApplied { control: "withheld_tool_call" } and RetryScheduled.
    • Re-prompts are bounded by RunPolicy::dropped_tool_call_nudges and run before the empty-reply retries.
    • Language-tagged fenced examples and bare JSON answers are left alone.

Evidence (replays of the captured requests against OpenRouter / DeepSeek)

Same model, same pinned provider, reasoning.effort = high, built byte for byte from the bench captures:

Request Variant Attempts Leaked
Summarizer (6 captured requests that had leaked) as sent today 23 5
fenced, instruction last (this PR) 38 0
Concluding call, tools withdrawn as sent today 8 6 (the other 2 ran out of tokens while reasoning)
tools declared, tool_choice: "none" 8 4
leaked row dropped + WITHHELD_TOOL_CALL_NUDGE (this PR) 12 0 (9 answered in plain text, 3 ran out of tokens while reasoning)

Tests

  • cargo test --workspace: all pass.
  • cargo clippy --workspace --all-targets -- -D warnings: clean.
  • cargo fmt --check: clean.
  • New tests: agent_loop/run_loop_withheld_tests.rs (scrubbing, quoted examples untouched, re-prompt, no-budget fallback, streamed path) and four in summarization/model_summarizer_tests.rs.
  • The three loop-level tests fail when the old TextRecovery::default() is restored, so they cover the regression.

Notes

  • Depends on feat(parse): add contains_call_markup tinytools#44. vendor/tinytools points at that PR's head, which adds contains_call_markup. Merge it first, then move this pin to the merge commit.

  • The commit history includes automatic checkpoint commits with generic subjects; this description is the authoritative summary.

  • The compaction trigger loop (re-summarizing on every turn after the first compaction) is a separate issue, being fixed separately in middleware/library/context.rs.

senamakel and others added 27 commits October 3, 2026 09:24
When the model returns an empty string during summarization, the summarizer now returns a fallback message instead of failing. This prevents a panic or unhelpful error in cases where the model produces no output, ensuring the harness can continue processing gracefully.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a model returns an empty string, the summarization process now returns an empty summary instead of attempting to process the missing output. This prevents a panic or incorrect behavior when the model produces no text.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added unit tests for the model summarizer to verify its core functionality, including summarization of model responses and handling of edge cases. This ensures the summarization logic is reliable and covers expected behaviors.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test was calling `contains_tool_call_markup` through a re-export that no longer exists, so the calls have been updated to use the full `super::model_summarizer::contains_tool_call_markup` path instead.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The dialect parser previously split agent names on whitespace, causing multi-word names to be incorrectly parsed as separate tokens. This change updates the parsing logic to treat the entire agent name as a single token, ensuring that agents with spaces in their names are correctly identified and handled during loop execution.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the dialect field is empty, the agent loop now defaults to a standard configuration instead of failing. This change ensures backward compatibility with existing agents that do not specify a dialect.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters an empty dialect list, it now correctly returns an empty result instead of panicking. This fixes a crash that occurred when no dialects were configured for the agent.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters a dialect that is not recognized, it now returns an error instead of panicking. This change improves robustness by allowing the system to gracefully report the unsupported dialect rather than crashing.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters a dialect that is not recognized, it now returns an error instead of panicking. This change improves robustness by allowing the system to gracefully report the unsupported dialect rather than crashing.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters a dialect that is not recognized, it now returns an error instead of panicking. This change improves robustness by allowing the system to gracefully report unsupported dialects rather than crashing at runtime.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters an empty dialect list, it now correctly returns an empty result instead of panicking. This fixes a crash that occurred when no dialects were configured for the agent, ensuring graceful handling of this edge case.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The model call error handling in the agent loop now returns a proper error result instead of panicking, ensuring the loop can gracefully recover from transient model failures and continue processing.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a tool call has no arguments, the run loop now correctly passes an empty object instead of failing to parse the input, ensuring that tools with optional parameters continue to work as expected.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When an LLM returns a tool call with an empty arguments string, the agent loop now treats it as a no-op rather than attempting to parse it as JSON, which previously caused a panic. This change improves robustness against malformed or incomplete tool call responses from language models.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When an agent returns a tool call with an empty arguments string, the run loop now treats it as a valid invocation with no parameters rather than failing to parse the empty input. This prevents a crash in scenarios where the model produces a tool call without arguments.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the agent loop encounters an empty list of actions, it now terminates gracefully instead of continuing to process. This prevents an infinite loop condition that could occur when no actions are available to execute.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When an agent returns a tool call with an empty arguments string, the run loop now treats it as a valid call with no parameters instead of failing to parse the JSON. This prevents a crash in scenarios where the model omits arguments for tools that accept none.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a tool call has no arguments, the agent loop now correctly passes an empty JSON object instead of failing to parse the missing field. This fixes a crash that occurred when the model returned a tool call without arguments.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the model returns an empty string, the summarizer now returns a fallback message instead of failing. This prevents a crash in downstream processing when no summary text is generated.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When an LLM returns a tool call with an empty arguments string, the agent loop now treats it as a valid call with no parameters instead of failing to parse the input. This prevents a crash in scenarios where the model omits arguments for tools that accept none.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The `offered` helper function in the recovery tests was not initializing the new `withhold` field on `TextRecovery`, which caused the struct to be constructed with an uninitialized value. Setting it to `false` ensures the test helper matches the current struct definition and prevents undefined behaviour.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the test expectations in the withheld agent loop tests to match the actual behavior of the run loop, ensuring that the tests correctly validate the agent's response when certain messages are withheld.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The `harness_with` helper function in the withholding tests now accepts `max_calls` as a `usize` instead of `u32` to match the type expected by the run policy configuration, eliminating a type mismatch that would cause compilation errors when passing the value to policy fields.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reformat several lines and function signatures that exceeded the project's line length limit, wrapping them across multiple lines for improved readability. No functional changes are introduced.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When an agent returns a tool call with an empty arguments string, the run loop now correctly passes an empty JSON object instead of failing to parse the empty string. This prevents a panic during tool execution when the model omits arguments for a tool that requires none.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added a README file to the summarization module to provide documentation and usage guidance for developers working with the harness.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a README file to document the agent loop module, providing an overview of its purpose and usage to help developers understand the module's role within the harness.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@tinysweeper

tinysweeper Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

⚠️ Review failed for 29ecb4e5cfe5. the review of #277 did not finish within 900s

Last completed report

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 0 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Incomplete
Priority: none
Reviewed head: 29ecb4e5cfe5
Updated: 1791010165 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 4 Active findings 0
Tests 3 Noted findings 0
Documentation 2 Resolved findings 0
Configuration 0 Pending checks/questions 16

Completeness: Incomplete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

No active actionable findings.

Could not review: crates/tinyagents-harness/src/agent_loop/README.md, crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/summarization/README.md, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs

Before merge

  • Complete the critique review for crates/tinyagents-harness/src/agent_loop/README.md, crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/summarization/README.md, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs.
  • Complete the security review for crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs.

How this fits together

flowchart LR
  n0["DeltaScrubber<br/>changed"]:::changed
  n1["DroppedBlocks<br/>changed"]:::changed
  n2["TextRecovery<br/>changed"]:::changed
  n3["reset_truncated_empty_recovery<br/>changed"]:::changed
  n4["run_loop_body"]:::impacted
  n5["collect"]:::impacted
  n6["new"]:::impacted
  n7["invoke_model_streaming_once"]:::impacted
  n8["invoke_model_with_retry"]:::impacted
  n9["CallShape"]:::impacted
  n0 -->|uses| n1
  n2 -->|uses| n1
  n4 -->|calls| n3
  n4 -->|uses| n9
  n6 -->|uses| n1
  n6 -->|calls| n5
  n7 -->|uses| n0
  n7 -->|uses| n9
  n8 -->|uses| n9
  n9 -->|uses| n2
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: crates/tinyagents-harness/src/agent_loop/README.md, crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/summarization/README.md, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs
  • Lane summary: Reviewed 0 files; 0 findings. 9 files could not be reviewed: crates/tinyagents-harness/src/agent_loop/README.md, crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/summarization/README.md, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs.

security

  • Conclusion: Neutral
  • Scope reviewed: incomplete; unanswered: crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs
  • Lane summary: Reviewed 0 files; 0 findings. 7 files could not be reviewed: crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs. 2 files were not security-reviewed: crates/tinyagents-harness/src/agent_loop/README.md (prose or tabular data), crates/tinyagents-harness/src/summarization/README.md (prose or tabular data).

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: This pull request adds handling for tool calls written on turns with no callable tools in the agent loop and improves the summarizer to fence the transcript and retry tool-call replies. The new behaviour is covered by unit and integration tests that assert specific outcomes and would fail if the logic regressed. The changes look sound. _Code retrieval was unavailable (model: ladder embeddings returned 402 Payment Required: {"error":"Insufficient USD or Diem balance to complete request. Visit https://venice\.ai/settings/api to add credits."}), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: This change prevents DeepSeek tool-call markup from leaking into summaries and tool-less turns by fencing the summarizer transcript, scrubbing withheld calls, and re-prompting with a nudge — with matching tests, docs, and passing CI checks. The description accurately covers the implemented behavior, and I found no problems introduced by the diff. _Code retrieval was unavailable (model: ladder embeddings returned 402 Payment Required: {"error":"Insufficient USD or Diem balance to complete request. Visit https://venice\.ai/settings/api to add credits."}), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: timed out after 10s), so this review ran without it._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: deepseek/deepseek-v4-flash
  • Spend: $0.004910
  • Tokens: 43041 input · 8097 output · 14080 cached · 0 embedding
Head State Pass summary
29ecb4e5cfe5 incomplete 0 active finding(s), 0 resolved finding(s) (at 1791010165)

tinysweeper 0.1.0

@coderabbitai

coderabbitai Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

🧰 Additional context used
📚 Code guidelines (1)
AGENTS.md — auto-discovered

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: cb26f29a-2cbe-4541-8faa-daa32be7c1aa
📥 Commits

Reviewing files that changed from the base of the PR and between 3c411b4 and 29ecb4e.

📒 Files selected for processing (9)
  • crates/tinyagents-harness/src/agent_loop/README.md
  • crates/tinyagents-harness/src/agent_loop/dialect.rs
  • crates/tinyagents-harness/src/agent_loop/model_call.rs
  • crates/tinyagents-harness/src/agent_loop/run_loop.rs
  • crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs
  • crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs
  • crates/tinyagents-harness/src/summarization/README.md
  • crates/tinyagents-harness/src/summarization/model_summarizer.rs
  • crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs

Included review availability: This review used your included allowance. Your plan provides up to 2 included reviews per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

Agent-loop turns without callable tools now withhold recognized tool-call markup and can retry with a no-tools nudge. Model summarization now fences the transcript and retries replies containing recognized tool-call markup once before returning an error.

Changes

Agent-loop tool-call withholding

Layer / File(s) Summary
Withholding recovery and scrubbing
crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs
Recovery and streaming scrubbing can withhold and count tool calls while preserving surrounding text and thinking content. Tests cover withheld calls, quoted markup, and JSON.
No-tools retry flow
crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/agent_loop/README.md
When tools are unavailable, the loop scrubs call markup and retries with a no-tools nudge when limits permit. Tests cover retry, call limits, and streaming. Documentation describes the behavior.

Model summarization safeguards

Layer / File(s) Summary
Fenced summarization prompt
crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs, crates/tinyagents-harness/src/summarization/README.md
The prompt places an optional previous summary before a tagged transcript and ends with instructions to output only the structured summary. Tests check the prompt structure.
Tool-call reply retries and fallback
crates/tinyagents-harness/src/summarization/model_summarizer.rs, crates/tinyagents-harness/src/summarization/model_summarizer_tests.rs, crates/tinyagents-harness/src/summarization/README.md
The summarizer retries a reply containing recognized tool-call markup once, then returns an error if markup remains. Tests cover detection, retries, and deterministic-trim fallback.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Bug fix

Sequence Diagram(s)

sequenceDiagram
  participant RunLoop
  participant Model
  participant DeltaScrubber
  RunLoop->>Model: Request a response with no callable tools
  Model->>DeltaScrubber: Stream response text
  DeltaScrubber->>RunLoop: Return text with tool-call markup withheld
  RunLoop->>Model: Retry with no-tools nudge when limits permit
Loading
sequenceDiagram
  participant ModelSummarizer
  participant Model
  participant FaultTolerantCachingSummarizer
  ModelSummarizer->>Model: Request a summary
  Model-->>ModelSummarizer: Reply containing tool-call markup
  ModelSummarizer->>Model: Retry the summary request once
  Model-->>ModelSummarizer: Return a second reply
  ModelSummarizer-->>FaultTolerantCachingSummarizer: Return an error if markup remains
  FaultTolerantCachingSummarizer-->>FaultTolerantCachingSummarizer: Use deterministic trimming on summarizer failure
Loading

Merge Risk: 🔵 Low · up to 29ecb

A call appearing only in the completed streaming response can be missed, leaving the lead-in as the answer without a retry. The risk is bounded but worth addressing or accepting before merge.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 29ecb

The change reduces unintended actions and corrupted conversation history through bounded recovery. No new privilege escalation is established, but preserving withdrawn capabilities across retries still depends on caller-side controls.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The directly traced text path affects the current invocation’s visible output, replay transcript, and recovery requests. Hosted tool execution remains constrained by the existing definition allowlist and authorization gate. Concrete tenant, asset, credential, and environment exposure depends on host registrations and policies not supplied here.

Security Findings and Attack Paths

  • observed — Withheld text calls are counted before conversion to ToolCall and are not appended to the dispatchable call collection. Provider-native structured calls follow the existing admission path rather than this text control; that distinction predates the PR and is not established as an introduced or worsened vulnerability.

Trust Boundaries and Controls

  • observed — Callable availability is determined after request middleware and structured-output planning. Each recovery iteration reconstructs the request and reruns middleware rather than latching request-level withdrawal. Execution separately retains allowlist, approval, and host authorization controls; whether withdrawal must persist across recovery remains an external host-contract question.

Resilience and Maintainability Implications

  • inferred — The additional recovery path is bounded and preserves inspectable history across exits. Streaming withholding protects recognized delta calls, but terminal reconciliation can discard terminal-only text before response-level detection. This pre-existing ordering limits conclusions about complete detection without the provider’s delta-versus-terminal contract.

Hardening Proposals

  • proposed — Define whether withdrawal is request-scoped or must survive recovery, and exercise that contract with registered tools and repeated middleware invocations. Where withdrawal represents security revocation, enforce it at tool admission rather than relying only on the advertised catalogue.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes both main changes: keeping tool-call markup out of summaries and tool-less answers.
Docstring Coverage ✅ Passed Docstring coverage is 82.35% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 34 functions across 7 files. (2 skipped: 2 …
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit watched the call tags fade,
While plain words stayed and thoughts remained.
“No tools this turn,” the model read,
Then tried again with nudged text said.
The summary kept its fence in place,
And carrots marked the retry’s trace.

Comment @coderabbitai help to get the list of available commands.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinyagents-harness/src/agent_loop/README.md, crates/tinyagents-harness/src/agent_loop/dialect.rs, crates/tinyagents-harness/src/agent_loop/model_call.rs, crates/tinyagents-harness/src/agent_loop/run_loop.rs, crates/tinyagents-harness/src/agent_loop/run_loop_recovery_tests.rs, crates/tinyagents-harness/src/agent_loop/run_loop_withheld_tests.rs, crates/tinyagents-harness/src/summarization/README.md, crates/tinyagents-harness/src/summarization/model_summarizer.rs and 1 more.

             $0.0049 · 43,041 in / 8,097 out · 14,080 cached (33%) · deepseek/deepseek-v4-flash
tests:       $0.0020 · 14,231 in / 4,027 out · 0 cached (0%)       · deepseek/deepseek-v4-flash
description: $0.0019 · 14,610 in / 83 out    · 0 cached (0%)       · deepseek/deepseek-v4-flash

@tinysweeper tinysweeper Bot added the priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect. label Oct 3, 2026
@senamakel senamakel self-assigned this Oct 3, 2026
@senamakel
senamakel merged commit bf968ff into tinyhumansai:main Oct 3, 2026
17 checks passed
sub4biz pushed a commit to sub4biz/tinyagents that referenced this pull request Oct 4, 2026
…inyhumansai#276)

Stacks the compaction Phase A follow-ups on both open PRs.

Co-authored-by: Medulla <medulla@tinyhumans.ai>
sub4biz pushed a commit to sub4biz/tinyagents that referenced this pull request Oct 4, 2026
…ction-phase-a

Keeps the summed summarizer usage in summarize_once and switches its
markup check to tinytools_agent::contains_call_markup, as tinyhumansai#277 now does.

Co-authored-by: Medulla <medulla@tinyhumans.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant