Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 20 additions & 1 deletion apps/cli/src/main.ts
Original file line number Diff line number Diff line change
Expand Up @@ -71,7 +71,7 @@ import {
updateGlobalMcpConfig,
withStepDefaults,
} from "@step-harness/coding-agent";
import { parseArgs, toPrintOutputMode } from "#args/index";
import { parseArgs, resolveAppMode, toPrintOutputMode } from "#args/index";
import { loadStepStartupConfig } from "#bootstrap/config";
import { createStepExtensionFactories } from "#bootstrap/extensions";
import { captureRawStdout, sdkStdioRequested } from "#bootstrap/stdout-capture";
Expand Down Expand Up @@ -554,6 +554,22 @@ try {
syncStepLoginProfileEndpoint(getStepAuthPath());
let shouldLaunchMain = true;
const parsedInteractiveArgs = parseArgs(compatibility?.args ?? stepCodeArgs);
if (
parsedInteractiveArgs.completionReview &&
!parsedInteractiveArgs.completionCheck &&
!parsedInteractiveArgs.help &&
!parsedInteractiveArgs.version
) {
throw new Error("--completion-review requires --completion-check git-committed");
}
if (
parsedInteractiveArgs.completionCheck &&
!parsedInteractiveArgs.help &&
!parsedInteractiveArgs.version &&
resolveAppMode(parsedInteractiveArgs, process.stdin.isTTY, process.stdout.isTTY) === "interactive"
) {
throw new Error("--completion-check requires print or JSON mode; use --print or --mode json.");
}
const interactiveStartup = isStepInteractiveLoginStartup({
stdinIsTTY: process.stdin.isTTY,
stdoutIsTTY: process.stdout.isTTY,
Expand Down Expand Up @@ -721,6 +737,9 @@ async function dispatchStepAppMode(prep: Extract<MainPreparation, { kind: "dispa
// print / json headless channels.
const exitCode = await runPrintMode(prep.runtimeHost, {
mode: toPrintOutputMode(prep.appMode),
completionCheck: prep.parsed.completionCheck,
completionCheckAttempts: prep.parsed.completionCheckAttempts,
completionReview: prep.parsed.completionReview,
messages: prep.parsed.messages,
initialMessage: prep.initialMessage,
initialImages: prep.initialImages,
Expand Down
165 changes: 165 additions & 0 deletions docs/compaction-integrity.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,165 @@
# Compaction integrity

Built-in compaction replaces earlier conversation context with a generated
summary. It must preserve the previous checkpoint when no new history needs
summarizing, and must fail before returning replacement context if a required
summary has no text. The coding-agent and agent-core compaction implementations
enforce the same rules.

## Preserve history during repeated split-turn compaction

A previous checkpoint can contain the only remaining copy of the original goal,
constraints, and verified results. When the next cut falls inside the first
retained turn, `messagesToSummarize` is empty but `turnPrefixMessages` is not.
The earlier summary still belongs in the next checkpoint.

For that case, `compact()` copies `previousSummary` verbatim as the history
portion and generates only the turn-prefix summary. The previous summary does
not need another model request. `No prior history.` is used only when there is
no previous summary. When new history exists, the existing update-summary request
continues to receive `previousSummary`.

## Reject empty generated summaries before assembly

Each history and turn-prefix generation extracts text blocks from the provider
response, then requires `text.trim().length > 0` before returning success. Empty
content arrays, empty text, whitespace-only text blocks, and thinking-only
responses fail this check. Thinking mixed with whitespace also fails. Accepted
text retains its original whitespace and formatting.

Validation happens before combining history and prefix text or appending file
operation metadata. In particular, none of the following can make a missing
generated summary valid:

- A preserved or newly generated history summary beside an empty turn prefix.
- `No prior history.`, the split-turn heading, or its separators.
- `<read-files>` and `<modified-files>` metadata.

Failure of either required generation fails the whole compaction. An empty
history response stops before requesting a turn-prefix summary. An empty prefix
response discards the newly generated history result as a candidate checkpoint.
No partial compaction result is returned.

The coding-agent helpers throw `Summarization failed: empty summary` or
`Turn prefix summarization failed: empty summary`. Agent-core returns a
`CompactionError` with code `summarization_failed` and the same message. Existing
session callers therefore keep their checkpoint, retained messages, and active
context when generation fails. Both manual and automatic built-in compaction
use these helpers.

Length-stop and provider-error diagnostics take precedence over the empty-text
check. Cancellation also remains a cancellation: coding-agent throws an
`AbortError` for an aborted summary response, and agent-core returns error code
`aborted`. Existing bounded retries for transient provider errors are unchanged;
an otherwise successful response with empty text fails without an added retry.

This check prevents missing summaries from being persisted. It does not judge
the factual quality of nonempty model text or validate extension-supplied
compaction results.

## Bound generated summary output

`reserveTokens` also sets the automatic trigger:
`contextTokens > contextWindow - reserveTokens`. Raising it must not raise a
summary request above the model's output limit or the existing 32000-token
summary ceiling. Both compaction implementations clamp the final candidate
budget after considering the reserve fraction and model budget. History and
history updates use fraction 0.8; split-turn prefixes use 0.5.

For a model declaring `contextWindow=1048576`, `maxTokens=65536`, a 196608-token
trigger uses `reserveTokens=851968`. History and prefix requests both have output
cap 32000. Previously their caps were 681574 and 425984. A positive lower model
limit is respected; when `maxTokens <= 0` denotes an unknown limit, the reserve
fraction remains the fallback, bounded by 32000.

Settings defaults and valid budgets are unchanged. In particular, a 65536-output
model keeps its 32000 summary cap with either the session's 16384 reserve or the
low-level helper's 24576 reserve. With an unknown model limit and reserve 24576,
history remains 19660 and prefix remains 12288. Ordinary model requests still use
their existing output budget; this change only bounds generated compaction
summaries. No new setting is introduced.

## Recognize compaction requests without relaxing the normal cap

Both implementations currently use the same system prompt: 581 UTF-8 bytes,
SHA-256 `7f4677db342c3991df3ed0ba729c514db1772ef7af4c39155d2a0d87d08b12bb`.
Hash decoded prompt text exactly: retain whitespace/newlines and do not add a
trailing newline. The text is `SUMMARIZATION_SYSTEM_PROMPT` in each compaction
`utils.ts`. The complete generated static instruction suffixes are pinned as:

| Kind | UTF-8 bytes | SHA-256 |
| --- | ---: | --- |
| History | 2702 | `4379f7f63f9fbd36f3f273e94b9566d78967e2e938e4b483957bf67266490572` |
| Prefix | 2948 | `0a355395dcd867cb08c3be3229d31475b3a468c51021489fcca4fab6c967cdb6` |
| Update | 3457 | `90546e9b55c76b8ac2570b98b0d2849cb3808f102097f52f643b00253edbb333` |

For the OpenAI Chat transport, positively identify a built-in compaction only
when all of these hold:

1. The top-level system message matches the pinned system prompt exactly.
2. The request has exactly a system message and one user message, both text-only,
with no `tools` or `tool_choice` field.
3. The entire user text matches the generated framing: `<conversation>\n...\n</conversation>\n\n`
followed by the exact pinned history or prefix instructions. An update also
has `<previous-summary>\n...\n</previous-summary>\n\n` before its exact update
instructions. History/update may append `\n\nAdditional focus: ...`; prefix
does not append it. The match must be unambiguous and anchored, not a search
for a phrase anywhere in the transcript.

The static suffix starts with `The messages above are a conversation to summarize.`,
`The messages above are the PREFIX of a single turn`, or
`The messages above are NEW conversation messages`, respectively. These starts
are labels for inspection; they are not sufficient classifiers on their own.
The complete suffix includes the eight-section format and detail rules.

Branch summaries share the system prompt but have different instructions and a
2048-token output budget. The system hash alone must not grant a compaction cap
allowance. Unknown, drifted, or ambiguous summary-shaped requests require review.
A normal request containing quoted summary instructions remains normal; neither
an observed 32000 cap nor a session ID establishes that a request is compaction.
The coding-agent main system prompt varies with runtime resources and configuration
and therefore has no single compaction-style static hash.

For the Harbor adaptive custom-model profile, the generated model overlay omits
`reasoning`, `thinkingLevelMap`, and `compat`, and the adapter emits no `--thinking`
flag. The model loader defaults `reasoning` to false. The default OpenAI Chat
serializer emits `max_completion_tokens`: 65536 for normal requests and 32000 for
recognized compaction requests with this model. Neither kind includes
`reasoning_effort` or `thinking`. Do not add `--thinking off` merely to test this
profile. Other integrity checks, including model identity and no-effort fields,
still apply to every request.

## Offline regression tests

The tests import the actual compaction and context modules and use the faux
provider. Coding-agent also exercises the real `AgentSession` and in-memory
`SessionManager`, checking that manual and automatic failures do not append a
checkpoint or change the active messages. Automatic cases also cover histories
without file metadata, where an empty generation previously produced either an
empty handoff or only the fixed split-turn boilerplate. No model service is used.

```sh
# From packages/coding-agent
pnpm exec vitest --run test/suite/regressions/compaction-integrity.test.ts

# From packages/agent-core
pnpm exec vitest --run test/harness/compaction-integrity.test.ts
```

Budget and adaptive wire regressions are also offline:

```sh
# From packages/agent-core
pnpm exec vitest --run test/harness/compaction-summary-budget.test.ts

# From packages/coding-agent
pnpm exec vitest --run test/compaction.test.ts -t 'summary output budget'
pnpm exec vitest --run test/suite/regressions/compaction-adaptive-wire.test.ts
```

The adaptive fixture was generated with the Harbor adapter's actual overlay
builder using a synthetic model ID and endpoint. The native test loads it through
the model registry, injects an offline HTTP fetch backed by the suite faux
provider, forces automatic history/prefix compaction, and verifies an explicit
history update. It pins the system and static instruction bytes, checks both
normal and summary wire caps, and supplies no thinking-level override.
121 changes: 121 additions & 0 deletions docs/completion-check.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,121 @@
# Print-mode completion check

The Step CLI can perform a bounded completion check in the existing session:

```sh
step --print --completion-check git-committed --completion-check-attempts 2 "Complete the task and commit the changes."
step --mode json --completion-check git-committed "Complete the task and commit the changes."
step --mode json --completion-check git-committed --completion-review "Complete the task and commit the changes."
```

The feature is off unless `--completion-check git-committed` is supplied. The
attempts option counts **additional prompts**, defaults to 2, and accepts integers
1 through 3. Both flags accept `--flag=value` syntax. Attempts without the check,
unsupported values, interactive mode, RPC, and SDK stdio are rejected. Piped
print mode is supported. Direct `runPrintMode` callers can supply the same
`completionCheck` and `completionCheckAttempts` options. The boolean
`--completion-review` flag takes no value, requires `--completion-check
git-committed`, and defaults off. Direct callers can set `completionReview: true`.

Before binding extensions or sending the first prompt, the check requires a Git
worktree with an existing HEAD commit and saves that HEAD. It then sends the
initial prompt, its images, and all additional user messages in their original
order. Once those prompts finish, completion requires all of these conditions:

- `starting-HEAD..HEAD` contains at least one commit. A preexisting commit or
moving HEAD backwards is insufficient.
- The committed tree differs from the starting HEAD's tree. An empty commit or
a change fully reverted before completion is insufficient. This tests delivery
of a change, not its correctness; the canonical verifier still owns correctness.
- The index and tracked worktree are clean, including submodule changes.
- No unignored untracked files remain. Ignored files do not block completion.
- The final assistant message has non-whitespace text and no pending tool calls.

Headless clients can set `STEP_CODING_AGENT_PLAN_DIR` to place generated
Markdown plans outside the worktree; see [plan file storage](step-configuration.md#plan-file-storage).
The conditions above still apply to all files remaining in the worktree.

If any condition is missing, a short status-only prompt asks the same session to
finish the task's required verification and commit work and give a final answer.
The original conversation and session ID remain in use. With review off, already
complete output costs no extra model calls. The follow-up budget applies to the whole invocation,
not separately to each user message. These prompts consume the original trial's
time budget; no trial timeout is extended or reset, and no new attempt is started.
The checker neither changes source files nor commits changes or runs hidden tests.

With `--completion-review`, the first eligible completion follow-up requests one
generic self-review even if the Git conditions and final text already pass. It
asks the same session and model to compare the work with the original visible
task, check public interfaces and types, boundary cases and the final diff, and
confirm relevant tests and checks ran after the last edit. It asks the assistant
to fix issues, commit remaining task changes, and provide a final answer while
preserving unrelated user changes and permission denials.

Any missing Git or final-text conditions are included in that same review prompt.
Review consumes **one of the existing follow-up slots**, with no extra budget or
new session. Later iterations use ordinary completion checks and feedback without
repeating the self-review. For example, a two-slot budget permits one combined
review/repair prompt and at most one further completion prompt. The review is an
optional experiment, not a grader; its benefit to task scores is unproven. It
introduces no hidden tests, external grading feedback, automatic commits, or
independent correctness judgment.

An explicit terminating tool denial, or an assistant error/abort observed during
this invocation, prevents further prompts from the checker, including when a
native retry subsequently succeeds. Pending user messages also stop at such a
terminal outcome when the check is enabled. Native provider retry policies are
unchanged. Explicit runtime session replacement continues to rebind listeners and
extensions, but the checker does not carry automatic feedback into another
session or working directory.
These rules also suppress review, including after an error recovered by native
retry. If cancellation or replacement occurs while waiting for stdout, a planned
review is not sent. Budget exhaustion and final-output exit codes are unchanged.

After the follow-up budget is exhausted, a valid final answer still returns exit
code **0** even if Git conditions remain unsatisfied. The canonical task verifier
owns the score; an ordinary failed task must not become an infrastructure error
that resamples the attempt. Missing/thinking-only final output returns **2** with
an explicit incomplete diagnostic. Existing terminal denials and final assistant
errors keep exit code **1**. Invalid configuration or failed Git preflight returns
**1** before a model call. If Git becomes unreadable after the model runs, the
checker stops adding prompts and reports that state; final text still returns 0,
and missing final text returns 2.

Text stdout contains only the last assistant answer. Diagnostics use stderr.
JSON mode keeps the ordinary session event stream, including the added user
prompts, and adds `completion_check` events. Each successful inspection includes
`check`, `attempt` (follow-ups already used, starting at 0), `maxAttempts`,
`hasNewCommit`, `hasCommittedChanges`, `trackedDirty`, `untrackedFiles`, `hasFinalText`, `status`
(`passed`, `follow_up`, or `exhausted`), and `willFollowUp`. A failed inspection
emits `status: "unavailable"` and `willFollowUp: false`. No filenames, file
contents, diffs, commit messages, or Git stderr appear in check feedback/events.

Only when review is enabled, completion events also contain
`review: { requested: boolean, sent: boolean }`. `requested` becomes true when the
first eligible inspection schedules review. `sent` becomes true only when the
original review text is emitted as a user message by that session; it does not
assert that the model completed a review or that the result is correct. A prompt
that fails preflight or is intercepted without delivery remains unsent.
Events scheduling feedback include `followUpKind: "review"` or `"completion"`.

When the review user message is delivered, JSON emits an additional
`completion_check` receipt with the same attempt and inspection fields and
`review.sent: true`. This preserves delivery evidence even if a later assistant
error or runtime replacement prevents another inspection. The receipt does not
run Git or consume another follow-up slot; count prompts by attempts and user
messages rather than the number of events. Subsequent inspections retain the
review state. With review disabled, these fields and the receipt are absent and
the existing event shape is unchanged.

Git is invoked directly with fixed argument arrays, no shell, and only a
validated starting object ID as a variable argument. Reads use `rev-parse`,
`rev-list --max-count=1`, `diff --quiet <starting-HEAD> HEAD --`, and NUL-delimited
porcelain `status --no-renames` with normal untracked-directory reporting. The
tree diff disables external diffs, text conversion, and rename detection. Only
its exit status is used: 0 means no committed changes, 1 means committed changes,
and any other code, timeout, or cancellation makes the check unavailable. Each command has a 5-second timeout,
64-KiB stdout/stderr limits, and SIGKILL termination; optional Git index/cache
writes and fsmonitor are disabled. Lazy fetching and interactive Git prompts are
disabled. Normal disposal and SIGINT/SIGTERM/SIGHUP cancel outstanding Git reads
and retain the existing runtime, detached-child, stdout-backpressure, and signal
cleanup paths.
Loading
Loading