Skip to content

v0.9.15: performance improvements, cli run until, memory fixes, xhigh thinking - #8654

Open
waleedlatif1 wants to merge 65 commits into
mainfrom
staging
Open

waleedlatif1 wants to merge 65 commits into
mainfrom
staging

Conversation

@waleedlatif1

@waleedlatif1 waleedlatif1 commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

icecrasher321 and others added 26 commits October 5, 2026 00:52
…he-top-options-compare (#8612)

* docs(library): update agentic-ai-coding-tools-what-they-are-and-how-the-top-options-compare

* Pi Babysit: address PR #8612 feedback

---------

Co-authored-by: Sim Pi Agent <pi@sim.ai>
…8613)

* feat(library): How do you build an AI bot for Discord without code?

* Pi Babysit: address PR #8613 feedback

---------

Co-authored-by: Sim Pi Agent <pi@sim.ai>
…ead of extra E2B round trips (#8621)

* perf(sandbox): grant the run_code session lease in the reconnect instead of extra E2B round trips

Reused Mothership workbench calls made five E2B control-plane requests:
list, connect, getInfo + setTimeout on acquire (the set always fired), and
getInfo on release. Connect now asks for max(5 min, remaining, lease), the
handle records the deadline it requested, and acquisition/release skip the
provider while that lower bound covers the request. Final deadlines are
unchanged; unrequested deadlines are still read back.

* test(sandbox): model connect as setting the deadline so a dropped preserve is caught
The simple Chat effort picker labeled medium as Low, high as Medium, and
xhigh as High. Derive its options from MOTHERSHIP_EFFORT_OPTIONS so each
label names the effort it sends: Medium, High, Extra High. Values, the
default (high), and the stored preference are unchanged.
…un --stop-after` (#8622)

* feat(workflows): stop a manual v2 run after a block, and `workflows run --stop-after`

A manual v2 run can now name `run.stopAfterBlockId`; the run stops once that
block completes and downstream blocks do not execute. Combined with a block
entry on the same block, it re-runs exactly one block against a prior run's
persisted upstream outputs, server-side:

  sim workflows run W --from-block X --source-run R --stop-after X --select-output X.result

Agents verifying an edit no longer re-run every upstream block (often a slow
LLM or API call) or toggle blocks off to skip them.

- Contract: optional `stopAfterBlockId` on the manual run selection.
- Application: both manual operations refuse a block missing from the saved
  workflow or nested in a loop/parallel (the engine would otherwise run to the
  end or stop after one iteration), before anything runs.
- Execute service: threads the trusted value to the sync and stream paths.
- CLI: `--stop-after <blockId>` implies --manual and rejects --async.
- E2E: test-workflow-stop-after-e2e.ts against a running app; the http-e2e job
  gains a Redis service because hosted billing admits runs through a Redis
  usage reservation.

* fix(workflows): refuse stop targets a run cannot reach, and run the E2E self-hosted

- The manual operations refuse a stop block the run cannot reach from its
  entry (an upstream block would let the run finish everything after the
  entry), and look blocks up as own properties.
- The executor fails a run whose stop block is absent from the workflow it
  executes, instead of running everything; this closes the window between
  validation and the executor's own draft load, for every caller.
- The CLI refuses an empty --stop-after rather than dropping it.
- CI: the stop-after E2E gets its own self-hosted app step; the SCIM suite
  asserts PostgreSQL rate-limit storage, so Redis is not added to that app.
  Fixture cleanup waits for run logs to finalize before deleting.

* fix(workflows): refuse a disabled stop block, or one reached only through one

The executor omits disabled blocks from its graph, so a disabled stop target,
or one whose only path runs through a disabled block, is never reached and
the run would finish everything after the entry.

* fix(workflows): the executor refuses a disabled stop block too

The serialized workflow keeps disabled blocks, but the DAG skips them, so a
disabled stop target would never be reached.

---------

Co-authored-by: Waleed Latif <waleed@sim.ai>
…pletion (#8620)

* fix(executor): stop retaining duplicate copies of loop block outputs

* test(executor): guard output sharing in block logs and loop aggregates

* fix(logs): size execution data the way JSON.stringify writes it

* fix(logs): count JSON string bytes without copying and unbox primitive wrappers

* fix(logs): measure execution data iteratively and apply toJSON on functions

* fix(executor): drop the full JSON clone of execution state at run completion

* fix(executor): normalize only live state for PII masking and walk serializability lazily

* fix(executor): snapshot array lengths and unbox wrappers in JSON walks

* fix(logs): read boxed boolean and bigint values the way JSON.stringify does
* feat(projects): add project identity and lifecycle foundation

* feat(projects): create projects with their initial environment

* docs(projects): record project files follow-up

* docs(projects): explain project and workspace creation flows

* fix(projects): stage activation after compatible writers deploy

* refactor(projects): prepare compatible writers for the SQL backfill

* fix(projects): clean up automatically created fixture Projects

* fix(workflows): guard restore against concurrent workspace archive

* fix(projects): close lifecycle races and surface rollout conflicts

* fix(workflows): return not found when import loses archive race
* fix(mcp): restrict MCP server destination changes to admins

* fix(mcp): compare exact MCP paths and guard concurrent URL changes

* fix(mcp): treat setting a URL on a URL-less server as a destination change

* fix(mcp): guard re-registration against concurrent URL changes

* fix(mcp): narrow re-registration URL before the guarded update

* chore(mcp): use absolute import in utils test
…ck (#8635)

* fix(executor): fail a stop-after run whose routing skips the stop block

A run with stopAfterBlockId only stopped when the stop block completed. When a
router, condition, or untaken error path routed the run away from it, the stop
never triggered and the run finished every other branch, reporting success as
if it had stopped there. A static check before the run cannot see this.

- The engine ends the run as soon as every path into the stop block has been
  deactivated, before any further block starts, and fails it with
  `Stop block "<name>" (<id>) was not reached: no path this run took leads to it`.
- Any run that ends without completing its stop block fails the same way: a
  stop block missing from the executed graph, or a Response block that ended
  the run first.
- A loop or parallel stop with nothing to run completes at its start sentinel,
  whose end sentinel never runs, so that exit now counts as reaching it.
- The v2 contract and the CLI `--stop-after` help describe the failure.
- E2E: a condition fixture checks the stop on the taken branch still stops
  there, a stop on the skipped branch fails the run before the other branch's
  slow block finishes, and the CLI exits non-zero.

* fix(executor): a skipped stop block fails a run another branch paused, and names a Response ending

- A run whose stop block was proven unreachable fails even when another branch
  paused, instead of returning a paused run that would resume past it.
- When a Response block ended the run first, the error says so rather than
  claiming no path leads to the stop block.
…le (#8631)

* fix(projects): restore workspace deletion and tighten Project lifecycle

- Archive a Project with its last active environment instead of refusing the
  workspace delete; account deletion follows the same rule, and the implicit
  archive is audited
- Gate Project APIs on a `projects` AppConfig flag (PROJECT_API_ENABLED fallback)
- Run Project reads in a read-only snapshot without locks; list Projects from the
  caller's grants with batched authorization
- Batch workflow archival, Project transfer and owner reassignment; move Project
  ownership on organization ownership transfer
- Reuse shared advisory-lock and text-array helpers; narrow admin-move conflict
  mapping to ProjectConflictError

* improvement(projects): unify environment archive and align with shared patterns

- Archive a workspace's workflows atomically with it through one
  archiveEnvironmentInTransaction shared by workspace delete and Project archive;
  the workspace row is locked before the sweep so concurrent creates are covered
- Scope the Project lock timeout to lock acquisition and map lock timeouts and
  deadlocks to a retryable conflict; backfill and multi-Project locks use the
  shared advisory-lock helpers in code-unit order
- Shared orchestrationFailureResponse for raw routes; contracts use the ID
  primitives and export only what is consumed; audit enums and mock in sync
- Project restrictions section matches its sibling settings rows
- Batch account-deletion Project loads/locks; skip inconsistent Projects in lists
- Harden the foundation integration suite (user-keyed cleanup, poll helper,
  pid-scoped waits, precise assertions)

* fix(projects): trim the requested organization id before validating it

* fix(projects): address review on Project locking, archive notifications and list policy cost

* fix(projects): keep archive retries from re-stamping MCP servers and isolate post-commit notifications
…m, restore Low (#8634)

* feat(mothership): keep each chat's reasoning effort, default to medium, restore Low

The simple picker offers Low / Medium / High / Extra High again, each sending
exactly that effort. New chats and chats never changed run at medium instead
of high. An effort the user picks is stored on the chat (copilot_chats.config)
through a new PUT /api/mothership/chats/[chatId]/effort and at turn admission,
and later turns of that chat keep it. The global last-used effort is no longer
persisted. The Sim Chat block defaults to medium.

* fix(mothership): keep the latest effort pick through refetches, failed saves and abandoned new chats

* fix(mothership): leave a deduplicated send's chat on the pick its first attempt stored

* fix(mothership): show a recovered chat's pick while its details load

* fix(mothership): hand a withdrawn first send's effort pick back to the new-chat composer

* fix(mothership): hand a withdrawn send's pick back only while its new-chat surface is open
… Tools Compared (#8638)

Co-authored-by: Sim Pi Agent <pi@sim.ai>
… Integrations (#8639)

Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
…gine-optimization-actually-mean (#8647)

Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
@waleedlatif1
waleedlatif1 requested a review from a team as a code owner October 6, 2026 02:21
@vercel

vercel Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

1 Skipped Deployment
Project Deployment Actions Updated
docs Skipped Skipped Oct 6, 2026 11:29pm UTC

Request Review

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 180 files

Confidence score: 3/5

  • In apps/sim/executor/execution/edge-manager.ts, an active incoming edge from a disconnected component can keep a stop-after target waiting for unrelated work, even though that edge’s source will never execute. Exclude unreachable sources from the early-failure check.
Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="apps/sim/executor/execution/edge-manager.ts">

<violation number="1" location="apps/sim/executor/execution/edge-manager.ts:133">
P2: An active incoming edge from a disconnected component keeps this true even though its source will never execute. A stop-after target in that component therefore cannot fail early and may wait for unrelated work to finish; account for source reachability from the execution frontier.</violation>
</file>

Re-trigger cubic

Comment thread apps/sim/lib/projects/membership.ts
Comment thread apps/sim/lib/execution/remote-sandbox/e2b.ts Outdated
Comment thread apps/sim/lib/workflows/lifecycle.ts Outdated
Comment thread apps/sim/lib/mcp/utils.ts Outdated
Comment thread apps/sim/hooks/queries/mothership-chats.ts Outdated
Comment thread apps/sim/lib/mcp/orchestration/server-lifecycle.ts
Comment thread apps/sim/app/api/workspaces/[id]/route.ts Outdated
Comment thread apps/sim/lib/workspaces/organization-workspaces.ts
Comment thread apps/sim/executor/execution/snapshot-serializer.ts Outdated
Comment thread apps/sim/lib/logs/execution/trace-spans/trace-spans.ts
@greptile-apps

greptile-apps Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 3/5

[High risk] Workflow execution engine and API contract changes.

The PR should not merge until stop-after runs handle already-launched branches and unrelated pauses consistently with the new API contract.

Findings

  1. P1 Other branches can run ▶
  2. P1 Pause overrides reached stop ▶

Summary

The PR adds Project identity and lifecycle operations, exposes manual stop-after workflow runs through the API and CLI, and updates chat effort, MCP configuration, execution performance, logging, and content indexing.

  • Stop-after runs need additional handling for concurrently launched branches and pauses.
  • Project creation, membership, and archival now share transactional workspace lifecycle paths.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Manual run with stop-after target] --> B[Execute ready branches]
  B --> C{Condition reaches target?}
  C -- No --> D[Fail run]
  C -- Yes --> E[Complete target and stop scheduling]
  B --> F[Other branch may already run or pause]
  F --> D
  F --> G[Paused result can override reached target]
Loading

Reviews (1) · Last reviewed commit: "fix(seo): make content pagination indexa..."

* docs(library): update sim-vs-dedicated-chatbot-builders

* Pi Babysit: address PR #8684 feedback

---------

Co-authored-by: Sim Pi Agent <pi@sim.ai>
…d ungated pickup windows (#8685)

* fix(desktop): scope the overdue sweep to desktop calls and keep ungated pickup windows

- The overdue listing and the per-call deadline read only consider desktop
  tool calls, and the settlement checks the call is one the desktop runs, so
  a bound run's workflow call or Sim-files VFS read is never failed with the
  desktop's not-started result.
- The cron scans only runs inside the inbox's horizon, oldest calls first,
  within its batch limit.
- Recording a decision clears the pickup deadline only for a call that was
  gated; a call that was never gated keeps the window it was offered with.

* fix(desktop): limit the overdue sweep over desktop calls only, and bound it by deadline

The sweep's tool-name filter admitted read, grep and glob calls on Sim's own
files, which the settlement then skipped, so enough of them could fill every
batch and starve real desktop calls. Queries now use the SQL form of
isDesktopToolCall (a desktop tool by name, or a VFS read of a granted local
folder) before their limit.

The sweep's horizon is now on the deadline that lapsed, not on the run's
start: Sim does not enforce a run's wall clock, so a long-lived run's
recently overdue call is still settled.

* fix(desktop): hand a device its calls in the order they were persisted

The inbox ordered calls by created_at, then tool_call_id. Calls of one turn
can be persisted in the same millisecond, and then the tie broke on random
ids: a device could run a click before the type the model emitted first.

Calls now carry persist_seq, a strictly increasing number assigned on insert
(pre-persist writes them in emission order), and the inbox and the overdue
sweep order by it. The migration adds the column without a default and then
sets the default, so existing rows are not rewritten; rows persisted before it
have no position and sort first, as the oldest.

* fix(desktop): reach every overdue desktop call, and declare the desktop tool names as const

A 24 h horizon on the lapsed deadline meant a call the backstop missed for a
day was never settled. The sweep starts from the few unsettled (pending or
running) calls in persistence order, so it needs no horizon to stay small.

The desktop tool names are a literal array declared as const, and the lookup
set is derived from it.

* test(mothership): give the hand-built tool call table the persistence sequence column
…finished turn (#8676)

* fix(mothership): re-read the transcript until the server has saved a finished turn

A tab finalizes a turn on the stream's `complete` event and reads the saved
transcript. The server sends that event before it persists the turn and clears
the chat's stream marker, so the read can land in the gap and return the
in-flight copy: the finished stream still listed as active and the answer under
its live id. That copy matches the optimistic one, so the history effect saw
nothing new, and the chat's `completed` event only marks a cached live stream
stale, so the tab kept it until a reload. On the local QA stack this happened
in roughly one turn in four: the finished answer had no Fork action, and the
cached stream marker kept a queued message from draining on a later return.

Finalize now re-reads with a short backoff while the server still lists the
stream it saw end. The first pass joins finalize's own read, so a turn saved in
time is still read once.

* fix(mothership): keep waiting for a slow save, and wait with a follow-up queued

- The re-read gave up after six reads (about 8s), so a slow save left the
  finished turn on its in-flight copy. It now keeps reading on a capped
  backoff for up to two minutes, and still stops as soon as the chat moves on
  (another send or another chat).
- It skipped the wait when a follow-up was queued, but a queued follow-up that
  does not go out at once (held for an edit) left the finished stream listed
  as running in the cache. It now runs for every finished turn; a follow-up
  that does go out ends it, since its send cancels the read and replaces the
  stream.

* improvement(mothership): pace the saved-turn re-read with the shared jittered backoff
…ed stream (#8675)

* fix(mothership): re-attach at once when the user returns to a recovered stream

Once the user left a running chat and came back, the return recovery owned the
stream for the rest of the turn, and every later online/visible/pageshow event
just joined it. So a network drop after that waited out whatever the recovery
was doing: a tail that went silent held the stream until the 45s idle timeout,
and a tail that failed slept out a reconnect backoff of up to 30s. The same
drop on a stream the send still owned re-attached immediately, because the
return signal supersedes the send's reader. On the local QA stack the stream
resumed 48s after the network came back.

A return signal now supersedes an in-flight recovery the same way, and the new
recovery re-attaches from the cursor.

* fix(mothership): finish a turn that ended while its reader was silent

A return recovery used the chat history's `activeStreamId` to decide what to
re-attach to. When the turn had ended while this surface was not listening (its
reader stalled, or a later return superseded the recovery that held it), the
history listed no running turn, so recovery returned without touching the
stream this surface still showed as running, and the chat stayed on Stop until
an idle timeout or backoff happened to run into the terminal state.

When the history lists no running turn but this surface is still sending,
recovery now resolves that stream: its terminal status replays the remaining
events and finalizes.

* fix(mothership): leave a send waiting for admission alone on return

A send shows as running before its POST is admitted, but the chat cannot list
it yet. Resolving that stream on a return event read it as ended and aborted
the POST. Recovery now resolves only a locally running stream whose POST was
admitted.

* fix(mothership): finish an admitted turn whose POST never answered on return

Recovery left a send alone while its POST had not answered, so a turn the server
admitted and finished, whose answer never reached the client, kept the chat on
Stop with nothing to clear it. The send now counts as admitted once the loaded
chat holds its message, and recovery resolves its stream as any other.
…dropping it (#8674)

* fix(mothership): keep a message the server never admitted instead of dropping it

Two sends were lost without an error:

- A send whose POST got no response (offline, Wi-Fi drop, waking a laptop)
  reconnected to the stream it would have opened. That stream does not exist,
  so the 404 read as "finished", the turn finalized as a success, and the
  refetched transcript no longer held the message. A queued follow-up was lost
  the same way, since it had already left the queue.
- A send refused with 409 because another turn held the chat (started in
  another tab, or one this surface lost track of) reconnected to that turn
  under the new message's bubble, then vanished when it finished.

Both now hand the message back under its id. An unreachable send is held in
the queue, so it is not redispatched into the same failure, and goes out when
the browser is back online (or when the user sends it). A send that found the
chat busy waits in the queue behind that turn, which the chat shows as
running, and goes out when it ends. Reusing the id keeps a retry deduplicated
if the server did admit the first attempt.

* fix(mothership): keep held and busy-refused sends exactly once across remounts

- A first message held offline on the new-chat page sat under that mount's
  queue key, which dies with the mount, so a reload or remount before the
  network returned stranded it. Held sends on a chatless surface now carry the
  surface they belong to, and the next chatless mount of that surface adopts
  them.
- Held sends are released for every chat when the browser comes back online,
  and on mount when it already is, so a send held in a chat the user is not
  viewing (or one whose `online` event fired with no surface mounted) still
  goes out.
- A send refused because the chat is busy is handed back only after the chat's
  running turn has been read, so the queue cannot redispatch it before that
  turn ends. A busy refusal that does not name the running turn no longer reads
  as a deduplicated send, which reconnected to a stream that never existed and
  lost the message.

* fix(mothership): send released held messages through the queue's own drain rules

Releasing held sends on mount kicked the queue dispatcher directly, which skips
the drain's guards. After a reload the chat history is not loaded yet, so a
follow-up queued behind a still-running turn went out at once and was refused as
busy. The release now only clears the hold; the drain effect, which waits for
the history and for the running turn to end, sends a released head, and now
also re-runs when the head's hold clears.

* fix(mothership): don't hold a send whose network returned while it was failing

An `online` event can fire while the failing POST is still pending, so the
release ran before the message was held and the message then waited for a
release that had already happened. A send now notes whether the browser came
back online while it was in flight, and if so goes back to the queue unheld,
for the drain to send under its usual rules.

* chore(mothership): name sign-out in desktop tool lease docs and say "not run" once

Follow-ups from the #8673 review: the lease docs now name sign-out
(`stopAllDesktopTools`) alongside the user's Stop as what cancels a desktop
tool, and the stale-observation message no longer says it was not run twice.

* fix(mothership): restore a withdrawn queued send even after the user switched chats

A queued send had already left the queue when its POST failed, and a dispatch
whose epoch changed meanwhile (the user switched chats) skipped restoring it,
so the message was lost. A withdrawn send was never admitted, so it now goes
back to its own chat's queue regardless of the epoch.

* fix(mothership): don't recreate a deleted chat's queue from a late restore

A withdrawn send restored after its dispatch outlived a chat switch could land
after the user deleted that chat, recreating a queue (and a message) for a
conversation that no longer exists. Clearing a chat's queue now leaves a
session tombstone that restores respect; a new enqueue for that key lifts it.
…n agreement (#8694)

The inbox and the overdue sweep filter calls with a SQL form of
isDesktopToolCall, and nothing tied the two together: a desktop tool added on
one side only would silently drop out of the inbox and the sweep, or let
Sim-only calls crowd them. One table of calls now runs through both against
real PostgreSQL, covering every named desktop tool, read and grep under,
exactly at and outside user-local, glob patterns, non-string and missing
arguments, a user-localX lookalike and non-desktop tools, and the test
requires them to agree on every case.
…ents-in-2026 (#8689)

* docs(library): update 6-best-ai-observability-tools-for-production-agents-in-2026

* Pi Babysit: address PR #8689 feedback

---------

Co-authored-by: Sim Pi Agent <pi@sim.ai>
* fix(plane): use a white integration background

* fix(plane): preserve distinct records in capped connector syncs

* fix(desktop): preserve legacy inbox timestamp precision
…an empty dev cache (#8686)

* fix(ci): start each HTTP end-to-end app from an empty Turbopack dev cache

The stop-after suite failed 12 of 175 CI runs (Oct 1-6), always in its own
`next dev` app, which restored the Turbopack dev cache (.next/dev) the
SCIM/version-compare app left in the same job. That app ran with other
NEXT_PUBLIC_* values, and `next dev` SIGKILLs its server 100ms after SIGTERM.
Restoring the cache:
- panicked Turbopack at startup (inner_of_upper_lost_follower), 6 runs;
- panicked mid-compile of the execute route ("socket connection was closed
  unexpectedly"), 4 runs;
- wedged that compile until the 300s request timeout ("The operation timed
  out."), 2 runs.
No run reached the executor: the cleanup check found no open execution log.

- Each app step removes .next/dev before starting, so no app restores another
  app's cache.
- A failing step prints the server log tail, not only a startup failure.
- The suite compiles the execute route with a refused request under its own
  300s budget and named check; every later request and CLI run is bounded at
  60s. A timed-out or dropped request names the route and elapsed time, lands
  in the report with a null status, and a CLI timeout is reported as one.

* fix(ci): start the desktop inbox app from an empty Turbopack dev cache too

* fix(e2e): record a status only for a fully read stop-after response

* fix(ci): stop each HTTP end-to-end app's whole process session before the next one starts

* fix(e2e): compile the desktop executor routes before timing executor calls
…is off (#8669)

* feat(mothership): let the worker choose models when the model picker is off

With the picker off, chats sent the hidden hosted default selection (GPT-6 Astra),
which the worker treats as an explicit pick for both the main agent and subagents.
No selection is sent now, so the worker routes every model role itself. A stale
client selection is still cleared at admission. Picker-on chats and the Sim Chat
block are unchanged.

* test(mothership): assert the omitted selection on the serialized worker request

* test(mothership): read the serialized request without a JSON round trip
…ts lean (#8625)

* perf(cli): load a manual run's draft once and keep embedded CLI results lean

- Manual runs reuse the draft state the use case loaded to pick and validate
  the trigger, instead of reading the draft tables twice more in the execute
  service and the execution core. The trigger is now chosen from the same
  snapshot that runs.
- Embedded (in-app agent) synchronous runs ask for file references without
  inline base64 unless --include-file-base64 is passed; --async never sends it.
- Embedded logs get leaves the workflow snapshot out unless
  --include-workflow-state is passed, via a new embeddedRequestDefault contract
  field. The installed CLI and public API defaults are unchanged.
- Add test-cli-run-latency-e2e.ts, which times both commands per segment
  against a running app and writes a JSON report; CI runs a short pass.

* fix(cli): cover followed runs, assert lean results, own CI step for the CLI E2E

- Embedded --follow runs also ask for file references only.
- The E2E asserts this checkout's embedded results carry no inline file bytes
  or workflow snapshot; a compared CLI build is measured, not asserted.
- Run the CLI E2E against its own self-hosted app: hosted billing admits runs
  through Redis, which the HTTP E2E job does not provision.

* fix(cli): remove the CLI E2E's execution logs and snapshots on cleanup

This branch was previously deployed

1 inactive deployment
Preview — f4855cdf Deployed Oct 6, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants