Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions packages/agent-core/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -130,6 +130,15 @@ The `beforeToolCall` hook runs after `tool_execution_start` and validated argume

Tools, blocked `beforeToolCall` results, and `afterToolCall` overrides can return `terminate: true` to hint that the automatic follow-up LLM call should be skipped. The loop only stops early when every finalized tool result in that batch sets `terminate: true`. Mixed batches continue normally.

`transformToolResult` is an optional final content-only transform on `AgentOptions`
and `AgentLoopConfig`. It runs before terminal tool events and result messages for
all outcomes, including validation failures, denied calls and rejected truncated
calls. It runs after `afterToolCall` for executed tools; existing execution-hook
semantics are unchanged. Its returned content cannot override usage, metadata or
termination hints. A thrown error becomes a bounded error result while preserving
those fields. Applications can use this boundary to retain and limit output;
agent-core itself does not own storage or output-limit policy.

The `Agent` class accepts `shouldStopAfterTurn` in `AgentOptions`. Low-level loop callers can set the same hook in `AgentLoopConfig`:

```typescript
Expand Down
40 changes: 34 additions & 6 deletions packages/agent-core/src/agent-loop.ts
Original file line number Diff line number Diff line change
Expand Up @@ -242,7 +242,7 @@ async function runLoop(
// them all instead of executing potentially borked calls.
const executedToolBatch =
message.stopReason === "length"
? await failToolCallsFromTruncatedMessage(toolCalls, emit)
? await failToolCallsFromTruncatedMessage(toolCalls, config, signal, emit)
: await executeToolCalls(currentContext, message, config, signal, emit);
toolResults.push(...executedToolBatch.messages);
hasMoreToolCalls = !executedToolBatch.terminate;
Expand Down Expand Up @@ -417,6 +417,8 @@ async function streamAssistantResponse(
*/
async function failToolCallsFromTruncatedMessage(
toolCalls: AgentToolCall[],
config: AgentLoopConfig,
signal: AbortSignal | undefined,
emit: AgentEventSink,
): Promise<ExecutedToolCallBatch> {
const messages: ToolResultMessage[] = [];
Expand All @@ -434,7 +436,7 @@ async function failToolCallsFromTruncatedMessage(
),
isError: true,
};
await emitToolExecutionEnd(finalized, emit);
await emitToolExecutionEnd(finalized, config, signal, emit);
const toolResultMessage = createToolResultMessage(finalized);
await emitToolResultMessage(toolResultMessage, emit);
messages.push(toolResultMessage);
Expand Down Expand Up @@ -506,7 +508,7 @@ async function executeToolCallsSequential(
);
}

await emitToolExecutionEnd(finalized, emit);
await emitToolExecutionEnd(finalized, config, signal, emit);
const toolResultMessage = createToolResultMessage(finalized);
await emitToolResultMessage(toolResultMessage, emit);
finalizedCalls.push(finalized);
Expand Down Expand Up @@ -548,7 +550,7 @@ async function executeToolCallsParallel(
result: preparation.result,
isError: preparation.isError,
} satisfies FinalizedToolCallOutcome;
await emitToolExecutionEnd(finalized, emit);
await emitToolExecutionEnd(finalized, config, signal, emit);
finalizedCalls.push(finalized);
if (signal?.aborted) {
break;
Expand All @@ -566,7 +568,7 @@ async function executeToolCallsParallel(
config,
signal,
);
await emitToolExecutionEnd(finalized, emit);
await emitToolExecutionEnd(finalized, config, signal, emit);
return finalized;
});
if (signal?.aborted) {
Expand Down Expand Up @@ -801,7 +803,33 @@ function createErrorToolResult(message: string): AgentToolResult<any> {
};
}

async function emitToolExecutionEnd(finalized: FinalizedToolCallOutcome, emit: AgentEventSink): Promise<void> {
async function emitToolExecutionEnd(
finalized: FinalizedToolCallOutcome,
config: AgentLoopConfig,
signal: AbortSignal | undefined,
emit: AgentEventSink,
): Promise<void> {
if (config.transformToolResult) {
try {
const content = await config.transformToolResult(finalized.result.content ?? [], signal);
finalized.result = { ...finalized.result, content };
} catch (error) {
const reason = error instanceof Error ? error.message.slice(0, 512) : "Unknown output processing error";
finalized.result = {
...finalized.result,
content: [
{
type: "text",
text: `Tool output could not be prepared: ${reason}. The tool may already have run; check its effects before repeating a state-changing call.`,
},
...(Array.isArray(finalized.result.content)
? finalized.result.content.filter((part) => part?.type === "image")
: []),
],
};
finalized.isError = true;
}
}
await emit({
type: "tool_execution_end",
toolCallId: finalized.toolCall.id,
Expand Down
4 changes: 4 additions & 0 deletions packages/agent-core/src/agent.ts
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,7 @@ export interface AgentOptions {
onResponse?: SimpleStreamOptions["onResponse"];
beforeToolCall?: (context: BeforeToolCallContext, signal?: AbortSignal) => Promise<BeforeToolCallResult | undefined>;
afterToolCall?: (context: AfterToolCallContext, signal?: AbortSignal) => Promise<AfterToolCallResult | undefined>;
transformToolResult?: AgentLoopConfig["transformToolResult"];
shouldStopAfterTurn?: (context: ShouldStopAfterTurnContext, signal?: AbortSignal) => boolean | Promise<boolean>;
prepareNextTurn?: (
signal?: AbortSignal,
Expand Down Expand Up @@ -190,6 +191,7 @@ export class Agent {
context: AfterToolCallContext,
signal?: AbortSignal,
) => Promise<AfterToolCallResult | undefined>;
public transformToolResult?: AgentLoopConfig["transformToolResult"];
public shouldStopAfterTurn?: (
context: ShouldStopAfterTurnContext,
signal?: AbortSignal,
Expand Down Expand Up @@ -225,6 +227,7 @@ export class Agent {
this.onResponse = runtimeOptions.onResponse;
this.beforeToolCall = runtimeOptions.beforeToolCall;
this.afterToolCall = runtimeOptions.afterToolCall;
this.transformToolResult = runtimeOptions.transformToolResult;
this.shouldStopAfterTurn = runtimeOptions.shouldStopAfterTurn;
this.prepareNextTurn = runtimeOptions.prepareNextTurn;
this.prepareNextTurnWithContext = runtimeOptions.prepareNextTurnWithContext;
Expand Down Expand Up @@ -457,6 +460,7 @@ export class Agent {
toolExecution: this.toolExecution,
beforeToolCall: this.beforeToolCall,
afterToolCall: this.afterToolCall,
transformToolResult: this.transformToolResult,
shouldStopAfterTurn: shouldStopAfterTurn
? async (context) => await shouldStopAfterTurn(context, this.signal)
: undefined,
Expand Down
11 changes: 11 additions & 0 deletions packages/agent-core/src/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -303,6 +303,17 @@ export interface AgentLoopConfig extends SimpleStreamOptions {
* The hook receives the agent abort signal and is responsible for honoring it.
*/
afterToolCall?: (context: AfterToolCallContext, signal?: AbortSignal) => Promise<AfterToolCallResult | undefined>;

/**
* Final content-only transform for every tool outcome, including validation failures,
* denied calls and truncated-call rejections. Runs after execution hooks and before
* terminal events or result messages; it cannot override usage or termination policy.
* A transform failure becomes an error result while preserving the original metadata.
*/
transformToolResult?: (
content: AgentToolResult<unknown>["content"],
signal?: AbortSignal,
) => Promise<AgentToolResult<unknown>["content"]>;
}

/**
Expand Down
196 changes: 196 additions & 0 deletions packages/agent-core/test/tool-result-transform.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,196 @@
import { fauxAssistantMessage, fauxToolCall } from "@step-harness/providers";
import { registerFauxProvider, streamSimple } from "@step-harness/providers/compat";
import { Type } from "typebox";
import { afterEach, describe, expect, it, vi } from "vitest";
import { Agent } from "../src/agent.ts";
import type { AgentEvent, AgentTool, AgentToolResult, ToolExecutionMode } from "../src/types.ts";

const providers: ReturnType<typeof registerFauxProvider>[] = [];
afterEach(() => {
while (providers.length) providers.pop()?.unregister();
});

function setup(toolExecution: ToolExecutionMode, tools: AgentTool[]) {
const provider = registerFauxProvider();
providers.push(provider);
const agent = new Agent({
streamFn: streamSimple,
getApiKey: () => "test-key",
initialState: { model: provider.getModel(), tools },
toolExecution,
});
const events: AgentEvent[] = [];
agent.subscribe((event) => {
events.push(event);
});
return { provider, agent, events };
}

function tool(name: string): AgentTool {
return {
name,
label: name,
description: name,
parameters: Type.Object({ count: Type.Number() }),
execute: async () => {
if (name === "broken") throw new Error("execution failed");
return { content: [{ type: "text", text: "original" }], details: { kept: true } };
},
};
}

describe.each<ToolExecutionMode>(["sequential", "parallel"])("final tool result transform (%s)", (mode) => {
it("transforms executed, invalid, missing and denied results before publication", async () => {
const { provider, agent, events } = setup(mode, [tool("ok"), tool("invalid"), tool("denied"), tool("broken")]);
const after = vi.fn(async () => ({ content: [{ type: "text" as const, text: "after hook" }] }));
agent.beforeToolCall = async ({ toolCall }) =>
toolCall.name === "denied" ? { block: true, reason: "denied", terminate: true } : undefined;
agent.afterToolCall = after;
const transform = vi.fn(async (content: AgentToolResult<unknown>["content"]) => [
{ type: "text" as const, text: `final:${JSON.stringify(content)}` },
]);
agent.transformToolResult = transform;
provider.setResponses([
fauxAssistantMessage(
[
fauxToolCall("ok", { count: 1 }),
fauxToolCall("invalid", { count: "bad" }),
fauxToolCall("missing", {}),
fauxToolCall("denied", { count: 1 }),
fauxToolCall("broken", { count: 1 }),
],
{ stopReason: "toolUse" },
),
fauxAssistantMessage("done"),
]);
await agent.prompt("run tools");
expect(transform).toHaveBeenCalledTimes(5);
expect(after).toHaveBeenCalledTimes(2);
const ended = events.filter((event) => event.type === "tool_execution_end");
expect(ended).toHaveLength(5);
for (const event of ended)
expect(event.result.content[0]).toMatchObject({ text: expect.stringMatching(/^final:/) });
const results = agent.state.messages.filter((message) => message.role === "toolResult");
expect(results.map((message) => message.isError)).toEqual([false, true, true, true, true]);
for (const message of results)
expect(message.content[0]).toMatchObject({ text: expect.stringMatching(/^final:/) });
});

it("keeps a denied batch terminating and leaves the execution hook untouched", async () => {
const blocked = tool("denied");
const execute = vi.spyOn(blocked, "execute");
const { provider, agent } = setup(mode, [blocked]);
agent.beforeToolCall = async () => ({ block: true, reason: "policy says no", terminate: true });
const after = vi.fn();
agent.afterToolCall = after;
agent.transformToolResult = async () => [{ type: "text", text: "bounded denial" }];
provider.setResponses([fauxAssistantMessage(fauxToolCall("denied", { count: 1 }), { stopReason: "toolUse" })]);
await agent.prompt("denied");
expect(execute).not.toHaveBeenCalled();
expect(after).not.toHaveBeenCalled();
expect(agent.state.messages.at(-1)).toMatchObject({
role: "toolResult",
isError: true,
content: [{ type: "text", text: "bounded denial" }],
});
});

it("reports transform failure without dropping the termination hint", async () => {
const done: AgentTool = {
...tool("done"),
execute: async () => ({
content: [{ type: "text", text: "completed" }],
details: { kept: true },
terminate: true,
}),
};
const { provider, agent } = setup(mode, [done]);
agent.transformToolResult = async () => {
throw new Error("cannot retain output");
};
provider.setResponses([fauxAssistantMessage(fauxToolCall("done", { count: 1 }), { stopReason: "toolUse" })]);
await agent.prompt("run");
expect(agent.state.messages.at(-1)).toMatchObject({
role: "toolResult",
isError: true,
details: { kept: true },
content: [{ type: "text", text: expect.stringContaining("cannot retain output") }],
});
});
});

it("transforms tools rejected for a truncated assistant message", async () => {
const { provider, agent } = setup("parallel", [tool("ok")]);
agent.transformToolResult = async () => [{ type: "text", text: "bounded length rejection" }];
provider.setResponses([
fauxAssistantMessage(fauxToolCall("ok", { count: 1 }), { stopReason: "length" }),
fauxAssistantMessage("done"),
]);
await agent.prompt("run");
expect(agent.state.messages.find((message) => message.role === "toolResult")).toMatchObject({
content: [{ type: "text", text: "bounded length rejection" }],
isError: true,
});
});

it("honors the constructor transform while preserving usage and newly discovered tool names", async () => {
const provider = registerFauxProvider();
providers.push(provider);
const usage = {
input: 1,
output: 2,
cacheRead: 0,
cacheWrite: 0,
totalTokens: 3,
cost: { input: 0.01, output: 0.02, cacheRead: 0, cacheWrite: 0, total: 0.03 },
};
const discover: AgentTool = {
...tool("discover"),
execute: async () => ({
content: [{ type: "text", text: "large source" }],
details: { kept: true },
usage,
addedToolNames: ["new_tool"],
terminate: true,
}),
};
const agent = new Agent({
streamFn: streamSimple,
getApiKey: () => "test-key",
initialState: { model: provider.getModel(), tools: [discover] },
transformToolResult: async () => [{ type: "text", text: "retained preview" }],
});
provider.setResponses([fauxAssistantMessage(fauxToolCall("discover", { count: 1 }), { stopReason: "toolUse" })]);
await agent.prompt("discover");
expect(agent.state.messages.at(-1)).toMatchObject({
role: "toolResult",
content: [{ type: "text", text: "retained preview" }],
details: { kept: true },
usage,
addedToolNames: ["new_tool"],
isError: false,
});
});

it("keeps image blocks when final text processing fails", async () => {
const image = { type: "image" as const, data: "aGVsbG8=", mimeType: "image/png" };
const capture: AgentTool = {
...tool("capture"),
execute: async () => ({
content: [image, { type: "text", text: "x".repeat(60000) }],
details: {},
terminate: true,
}),
};
const { agent, provider } = setup("parallel", [capture]);
agent.transformToolResult = async () => {
throw new Error("storage unavailable");
};
provider.setResponses([fauxAssistantMessage(fauxToolCall("capture", { count: 1 }), { stopReason: "toolUse" })]);
await agent.prompt("capture");
const result = agent.state.messages.at(-1);
expect(result?.role).toBe("toolResult");
if (result?.role !== "toolResult") throw new Error("Expected tool result");
expect(result.isError).toBe(true);
expect(result.content.filter((part) => part.type === "image")).toEqual([image]);
});
35 changes: 35 additions & 0 deletions packages/coding-agent/docs/tool-output-retention.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# Tool output retention

Coding-agent bounds the combined text in each final tool result to 2,000 lines
and 50 KiB, including its truncation notice. This applies to built-in, MCP and
extension tools, text replaced by `tool_result` hooks, validation errors and
blocked calls. Accepted `message_end` tool-result replacements are checked again
before session persistence and model replay. Images are preserved separately. Tool details, usage, error
status and termination hints retain their existing contracts.

An oversized result keeps a prefix and a `Full output:` path. The complete text
from all text blocks, joined by newlines, is saved before the preview is
published. Individual oversized lines may leave no complete line in the
preview; the full file remains readable with the normal file tools. A producer's
`truncated` metadata does not disable this final bound. Results that exceed the
limits by at most 8 lines and 1 KiB are left unchanged: built-in tools truncate
to the same limits and then append a notice and, for bash, an exit status, and
that tail must not be cut again. If a producer already lost content before
returning, this file contains only what it returned.

Artifacts live in `tool-output` under the session directory. If the session has
no storage directory, the configured agent directory (or its default) is used.
Paths in tool results are absolute JSON-quoted strings, so spaces and line breaks
in directory names do not change the notice's structure. New artifact directories are private and
files are created with exclusive creation and owner-only permissions. Each
retention operation makes a bounded best-effort pass over expired owned files;
files older than seven days may be removed. Cleanup skips unrelated names,
directories and symlinks, and a cleanup failure does not fail the tool.

If complete output cannot be saved, the result reports an output-processing
error. It states that the tool may already have run and that effects should be
checked before repeating a state-changing call. Image blocks remain available even if text retention fails. Existing denial and
termination policy remains in force. Small results create no artifact and are unchanged.

This bounds final model-facing text, not streaming progress updates or arbitrary
typed `details` payloads. No new permission bypass or external upload is added.
Loading
Loading