Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 28 additions & 0 deletions packages/coding-agent/docs/compaction-overflow-recovery.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
# Compaction summary overflow recovery

Conversation-history summaries and split-turn prefix summaries can recover when the provider rejects the summary request for context overflow. The first request keeps the existing prompt, serialization, output-token limit, reasoning settings, and routing options. Summary requests continue to use `cacheRetention: "none"`.

After a classified overflow error, the next requests target 70%, 50%, then 35% of the **original rejected request's estimated input size**. The estimate uses the serialized text actually sent, including the system prompt, conversation tags, summarization instructions, custom focus, and any previous summary. Each reduction must produce a strictly smaller request within its target. The percentages are not compounded, and there are at most three reductions.

The reducer removes complete older message groups, oldest first. An assistant's tool calls and all their results form one group, including parallel results and messages interleaved before the batch completes. The retained source preserves:

- The previous summary and summarization/custom instructions verbatim.
- The latest real user request, including image input with blank or absent text, even if synthetic messages follow it.
- The newest assistant group and newest complete tool batch, plus trailing context.
- Checkpoint messages already present in the source, including compaction and branch summaries.

Retained groups use the existing serialization and tool-output truncation. Recovery does not further rewrite their contents. A note inside the conversation reports how many older groups were omitted. If the protected content cannot fit, no further measurable reduction is possible, or tool calls/results cannot be grouped safely, compaction fails explicitly instead of resending an unchanged rejected payload.

Only an error response classified as context overflow activates source reduction. Authentication and other permanent errors retain their failure behavior. Empty summaries, length-limited summaries, and successful responses with large reported input usage do not trigger input loss. Cancellation prevents subsequent requests. Output-integrity validation remains separate from overflow recovery.

Transient errors retain the configured retry policy and reporting callbacks. They retry the current request without reducing it, and the transient retry budget is shared across all overflow reductions for that summary. A summary therefore makes at most four requests plus the configured number of transient retries. A split compaction gives the history and prefix summaries their own budgets, as before. Branch summaries retain their existing behavior.

## Usage and failure contract

When a summary succeeds, its returned usage includes reported usage from the failed overflow/transient attempts as well as the successful response, including cache, reasoning, and cost fields. A successful split compaction combines the history and prefix totals.

`AgentSession` persists usage only with a successful compaction. This change does **not** add persistent accounting for a terminally failed or cancelled compaction, or for a successful history summary followed by a failed prefix summary. Those paths retain the existing failure contract; no usage entry is written for them.

## Regression coverage

`test/compaction-overflow-recovery.test.ts` drives the real summary functions through an injected stream function. It covers both summary paths, protected content and tool pairing, bounded reductions, no progress, cancellation, permanent errors, unchanged successful requests, transient retries, successful usage aggregation, and unchanged branch behavior. Existing summary-reasoning, serialization, transient-stream-drop, and truncated-summary regressions remain applicable.
144 changes: 122 additions & 22 deletions packages/coding-agent/src/core/compaction/compaction.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@
import type { AgentMessage, StreamFn, ThinkingLevel } from "@step-harness/agent-core";
import {
contentText,
isContextOverflow,
type RetryCallbacks,
type RetryPolicy,
retryAssistantCall,
Expand All @@ -22,6 +23,7 @@ import {
type SessionEntry,
sessionEntryToContextMessages,
} from "../session-manager.ts";
import { createSummarySourceReducer } from "./summary-overflow.ts";
import {
collectSkillInstructions,
formatSkillInstructions,
Expand Down Expand Up @@ -589,6 +591,9 @@ export function getSummarizationFailure(
label: string,
maxTokens?: number,
): string | undefined {
if (response.stopReason === "aborted") {
return `${label} aborted before producing a summary`;
}
if (response.stopReason === "error") {
return `${label} failed: ${response.errorMessage || "Unknown error"}`;
}
Expand Down Expand Up @@ -699,6 +704,93 @@ function buildSummarizationContext(promptText: string): Context {
};
}

/** Shares of the first rejected request's estimate, never compounded across retries. */
const SUMMARY_OVERFLOW_TARGETS = [0.7, 0.5, 0.35] as const;

/**
* Keep the initial summary payload intact, reducing only its source history after an
* explicit context-overflow error. The outer transient policy retains one retry budget
* across reductions; completeSummarization performs each individual request with the
* usual cache/routing behavior. Branch summaries keep their existing independent path.
*/
async function completeSummaryWithOverflowRecovery(
messages: AgentMessage[],
buildPrompt: (conversation: string) => string,
model: Model<any>,
options: SimpleStreamOptions,
label: string,
streamFn?: StreamFn,
retry?: RetryPolicy,
callbacks?: RetryCallbacks,
): Promise<AssistantMessage> {
let conversation = serializeConversation(convertToLlm(messages));
let context = buildSummarizationContext(buildPrompt(conversation));
const promptOverhead = buildPrompt("").length;
const estimateRequestTokens = (conversationChars: number) =>
Math.ceil(SUMMARIZATION_SYSTEM_PROMPT.length / 4) + Math.ceil((promptOverhead + conversationChars) / 4);
const originalTokens = estimateRequestTokens(conversation.length);
let rejectedTokens = originalTokens;
let reductions = 0;
let reduceSource: ReturnType<typeof createSummarySourceReducer>;
let usage: Usage | undefined;
const requestOptions = { ...options, sessionId: options.sessionId ?? uuidv7() };

// A local recovery/cancellation failure can end produce() while a transient retry is
// active. Close its reporting lifecycle just as retryAssistantCall does for responses.
let activeRetry: number | undefined;
const retryCallbacks: RetryCallbacks = {
...callbacks,
onRetryScheduled: async (...args) => {
activeRetry = args[0];
await callbacks?.onRetryScheduled?.(...args);
},
onRetryFinished: async (...args) => {
activeRetry = undefined;
await callbacks?.onRetryFinished?.(...args);
},
};

const produce = async (): Promise<AssistantMessage> => {
for (;;) {
if (options.signal?.aborted) throw new Error("Compaction cancelled");
const response = await completeSummarization(model, context, requestOptions, streamFn);
usage = usage ? combineUsage(usage, response.usage) : response.usage;
if (options.signal?.aborted) throw new Error("Compaction cancelled");
// Do not infer overflow from successful usage or empty/length-limited output.
if (response.stopReason !== "error" || !isContextOverflow(response)) return response;

const fraction = SUMMARY_OVERFLOW_TARGETS[reductions];
if (fraction === undefined) {
throw new Error(
`${label} failed: context overflow after ${reductions} reduced requests (70%, 50%, 35%): ${response.errorMessage}`,
);
}
reduceSource ??= createSummarySourceReducer(messages, estimateRequestTokens);
const reduced = reduceSource?.(Math.floor(originalTokens * fraction), rejectedTokens);
if (reduced === undefined) {
throw new Error(
`${label} failed: context overflow recovery cannot reduce the request further while preserving protected summaries, instructions, the latest user request and complete tool groups`,
{ cause: response.errorMessage },
);
}
reductions++;
conversation = reduced;
rejectedTokens = estimateRequestTokens(conversation.length);
context = buildSummarizationContext(buildPrompt(conversation));
}
};

try {
const response = await retryAssistantCall(produce, retry, options.signal, retryCallbacks);
return { ...response, usage: usage ?? response.usage };
} catch (error) {
if (activeRetry !== undefined) {
await callbacks?.onRetryFinished?.(false, activeRetry, error instanceof Error ? error.message : String(error));
}
throw error;
}
}

/** Generate or update a conversation summary and return its provider usage. */
export async function generateSummaryWithUsage(
currentMessages: AgentMessage[],
Expand All @@ -724,17 +816,14 @@ export async function generateSummaryWithUsage(
basePrompt = `${basePrompt}\n\nAdditional focus: ${customInstructions}`;
}

// Serialize conversation to text so model doesn't try to continue it
// Convert to LLM messages first (handles custom types like bashExecution, custom, etc.)
const llmMessages = convertToLlm(currentMessages);
const conversationText = serializeConversation(llmMessages);

// Build the prompt with conversation wrapped in tags
let promptText = `<conversation>\n${conversationText}\n</conversation>\n\n`;
if (previousSummary) {
promptText += `<previous-summary>\n${previousSummary}\n</previous-summary>\n\n`;
}
promptText += basePrompt;
// Only the conversation can shrink. Instructions and the previous summary remain verbatim.
const buildPrompt = (conversation: string): string => {
let promptText = `<conversation>\n${conversation}\n</conversation>\n\n`;
if (previousSummary) {
promptText += `<previous-summary>\n${previousSummary}\n</previous-summary>\n\n`;
}
return promptText + basePrompt;
};

const completionOptions = createSummarizationOptions(
model,
Expand All @@ -747,24 +836,31 @@ export async function generateSummaryWithUsage(
sessionId,
);

const response = await completeSummarization(
const response = await completeSummaryWithOverflowRecovery(
currentMessages,
buildPrompt,
model,
buildSummarizationContext(promptText),
completionOptions,
"Summarization",
streamFn,
retry,
callbacks,
);

const failure = getSummarizationFailure(response, "Summarization", maxTokens);
if (failure) {
throw new Error(failure);
const error = new Error(failure);
if (response.stopReason === "aborted") error.name = "AbortError";
throw error;
}
if (response.content.some((block) => block.type === "toolCall")) {
throw new Error("Summarization attempted to call a tool");
}

const textContent = contentText(response.content);
if (!textContent.trim()) {
throw new Error("Summarization failed: response contained no summary text");
}

return { text: textContent, usage: response.usage };
}
Expand Down Expand Up @@ -1040,29 +1136,33 @@ async function generateTurnPrefixSummary(
// Smaller output budget for turn-prefix summaries: the suffix of the turn
// is retained verbatim, so the summary only needs to describe the prefix.
const maxTokens = pickSummaryMaxTokens(model, reserveTokens, 0.5);
const llmMessages = convertToLlm(messages);
const conversationText = serializeConversation(llmMessages);
const promptText = `<conversation>\n${conversationText}\n</conversation>\n\n${TURN_PREFIX_SUMMARIZATION_PROMPT}`;

const response = await completeSummarization(
const response = await completeSummaryWithOverflowRecovery(
messages,
(conversation) => `<conversation>\n${conversation}\n</conversation>\n\n${TURN_PREFIX_SUMMARIZATION_PROMPT}`,
model,
buildSummarizationContext(promptText),
createSummarizationOptions(model, maxTokens, apiKey, headers, env, signal, thinkingLevel, sessionId),
"Turn prefix summarization",
streamFn,
retry,
callbacks,
);

const failure = getSummarizationFailure(response, "Turn prefix summarization", maxTokens);
if (failure) {
throw new Error(failure);
const error = new Error(failure);
if (response.stopReason === "aborted") error.name = "AbortError";
throw error;
}
if (response.content.some((block) => block.type === "toolCall")) {
throw new Error("Turn prefix summarization attempted to call a tool");
}
const textContent = contentText(response.content);
if (!textContent.trim()) {
throw new Error("Turn prefix summarization failed: response contained no summary text");
}

return {
text: contentText(response.content),
text: textContent,
usage: response.usage,
};
}
111 changes: 111 additions & 0 deletions packages/coding-agent/src/core/compaction/summary-overflow.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,111 @@
/** Source-history reduction used only after a summary request is rejected for context overflow. */
import type { AgentMessage } from "@step-harness/agent-core";
import { contentText } from "@step-harness/providers";
import { convertToLlm } from "../messages.ts";
import { serializeConversation } from "./utils.ts";

type SourceReducer = (targetTokens: number, rejectedTokens: number) => string | undefined;

/**
* Build atomic groups without cutting across outstanding tool calls. Interleaved messages
* stay with the batch until every result arrives. Ambiguous or incomplete tool histories
* cannot be reduced safely; the original request may still succeed without this fallback.
*/
function groupMessages(messages: AgentMessage[]): AgentMessage[][] | undefined {
const groups: AgentMessage[][] = [];
let group: AgentMessage[] = [];
const pending = new Set<string>();
const calls = new Set<string>();
for (const message of messages) {
group.push(message);
if (message.role === "assistant") {
for (const block of message.content) {
if (block.type !== "toolCall") continue;
if (calls.has(block.id)) return undefined;
calls.add(block.id);
pending.add(block.id);
}
} else if (message.role === "toolResult" && !pending.delete(message.toolCallId)) {
return undefined;
}
if (pending.size === 0) {
groups.push(group);
group = [];
calls.clear();
}
}
return pending.size === 0 ? groups : undefined;
}

/**
* Drop only older complete groups, preserving the latest real user request, the newest
* assistant and tool batch, and any trailing context. Serialization matches the initial
* request, including its existing tool-output truncation; retained content is not rewritten.
* The estimator includes fixed instructions/previous-summary overhead supplied by the caller.
*/
export function createSummarySourceReducer(
messages: AgentMessage[],
estimateRequestTokens: (conversationChars: number) => number,
): SourceReducer | undefined {
const groups = groupMessages(messages);
if (!groups?.length) return undefined;

let latestUser = -1;
let latestAssistant = -1;
let latestTools = -1;
for (let index = groups.length - 1; index >= 0; index--) {
for (const message of groups[index]) {
// An image can carry the latest request even when its text is blank.
if (
latestUser < 0 &&
message.role === "user" &&
(contentText(message.content, "").trim().length > 0 ||
(Array.isArray(message.content) && message.content.some((block) => block.type === "image")))
) {
latestUser = index;
}
if (message.role === "assistant") {
if (latestAssistant < 0) latestAssistant = index;
if (latestTools < 0 && message.content.some((block) => block.type === "toolCall")) latestTools = index;
}
}
}
const tailStart = latestAssistant >= 0 ? latestAssistant : groups.length - 1;
const sources = groups.map((group, index) => ({
text: serializeConversation(convertToLlm(group)),
protected:
index === latestUser ||
index === latestTools ||
index >= tailStart ||
group.some((message) => message.role === "compactionSummary" || message.role === "branchSummary"),
}));
let textChars = sources.reduce((sum, source) => sum + source.text.length, 0);
let textCount = sources.filter((source) => source.text.length > 0).length;
let nextGroup = 0;
let omitted = 0;

return (targetTokens, rejectedTokens) => {
while (nextGroup < sources.length) {
const source = sources[nextGroup++];
if (source.protected || !source.text) continue;
textChars -= source.text.length;
textCount--;
source.text = "";
omitted++;
const note = `[${omitted} older message groups omitted after context overflow]\n\n`;
const tokens = estimateRequestTokens(note.length + textChars + Math.max(0, textCount - 1) * 2);
// A large earlier removal may already be below the next target. Still require
// a strictly smaller payload: never submit the same rejected request again.
if (tokens <= targetTokens && tokens < rejectedTokens) {
return (
note +
sources
.map((entry) => entry.text)
.filter(Boolean)
.join("\n\n")
);
}
}
return undefined;
};
}
Loading
Loading