Repository navigation
v0.9.15: performance improvements, cli run until, memory fixes, xhigh thinking - #8654
Open
waleedlatif1 wants to merge 65 commits into
Open
waleedlatif1 wants to merge 65 commits into
waleedlatif1 wants to merge 65 commits into
Conversation
…ead of extra E2B round trips (#8621) * perf(sandbox): grant the run_code session lease in the reconnect instead of extra E2B round trips Reused Mothership workbench calls made five E2B control-plane requests: list, connect, getInfo + setTimeout on acquire (the set always fired), and getInfo on release. Connect now asks for max(5 min, remaining, lease), the handle records the deadline it requested, and acquisition/release skip the provider while that lower bound covers the request. Final deadlines are unchanged; unrequested deadlines are still read back. * test(sandbox): model connect as setting the deadline so a dropped preserve is caught
The simple Chat effort picker labeled medium as Low, high as Medium, and xhigh as High. Derive its options from MOTHERSHIP_EFFORT_OPTIONS so each label names the effort it sends: Medium, High, Extra High. Values, the default (high), and the stored preference are unchanged.
…un --stop-after` (#8622) * feat(workflows): stop a manual v2 run after a block, and `workflows run --stop-after` A manual v2 run can now name `run.stopAfterBlockId`; the run stops once that block completes and downstream blocks do not execute. Combined with a block entry on the same block, it re-runs exactly one block against a prior run's persisted upstream outputs, server-side: sim workflows run W --from-block X --source-run R --stop-after X --select-output X.result Agents verifying an edit no longer re-run every upstream block (often a slow LLM or API call) or toggle blocks off to skip them. - Contract: optional `stopAfterBlockId` on the manual run selection. - Application: both manual operations refuse a block missing from the saved workflow or nested in a loop/parallel (the engine would otherwise run to the end or stop after one iteration), before anything runs. - Execute service: threads the trusted value to the sync and stream paths. - CLI: `--stop-after <blockId>` implies --manual and rejects --async. - E2E: test-workflow-stop-after-e2e.ts against a running app; the http-e2e job gains a Redis service because hosted billing admits runs through a Redis usage reservation. * fix(workflows): refuse stop targets a run cannot reach, and run the E2E self-hosted - The manual operations refuse a stop block the run cannot reach from its entry (an upstream block would let the run finish everything after the entry), and look blocks up as own properties. - The executor fails a run whose stop block is absent from the workflow it executes, instead of running everything; this closes the window between validation and the executor's own draft load, for every caller. - The CLI refuses an empty --stop-after rather than dropping it. - CI: the stop-after E2E gets its own self-hosted app step; the SCIM suite asserts PostgreSQL rate-limit storage, so Redis is not added to that app. Fixture cleanup waits for run logs to finalize before deleting. * fix(workflows): refuse a disabled stop block, or one reached only through one The executor omits disabled blocks from its graph, so a disabled stop target, or one whose only path runs through a disabled block, is never reached and the run would finish everything after the entry. * fix(workflows): the executor refuses a disabled stop block too The serialized workflow keeps disabled blocks, but the DAG skips them, so a disabled stop target would never be reached. --------- Co-authored-by: Waleed Latif <waleed@sim.ai>
…pletion (#8620) * fix(executor): stop retaining duplicate copies of loop block outputs * test(executor): guard output sharing in block logs and loop aggregates * fix(logs): size execution data the way JSON.stringify writes it * fix(logs): count JSON string bytes without copying and unbox primitive wrappers * fix(logs): measure execution data iteratively and apply toJSON on functions * fix(executor): drop the full JSON clone of execution state at run completion * fix(executor): normalize only live state for PII masking and walk serializability lazily * fix(executor): snapshot array lengths and unbox wrappers in JSON walks * fix(logs): read boxed boolean and bigint values the way JSON.stringify does
* feat(projects): add project identity and lifecycle foundation * feat(projects): create projects with their initial environment * docs(projects): record project files follow-up * docs(projects): explain project and workspace creation flows * fix(projects): stage activation after compatible writers deploy * refactor(projects): prepare compatible writers for the SQL backfill * fix(projects): clean up automatically created fixture Projects * fix(workflows): guard restore against concurrent workspace archive * fix(projects): close lifecycle races and surface rollout conflicts * fix(workflows): return not found when import loses archive race
* fix(mcp): restrict MCP server destination changes to admins * fix(mcp): compare exact MCP paths and guard concurrent URL changes * fix(mcp): treat setting a URL on a URL-less server as a destination change * fix(mcp): guard re-registration against concurrent URL changes * fix(mcp): narrow re-registration URL before the guarded update * chore(mcp): use absolute import in utils test
…v2 provider discovery (#8632)
…ck (#8635) * fix(executor): fail a stop-after run whose routing skips the stop block A run with stopAfterBlockId only stopped when the stop block completed. When a router, condition, or untaken error path routed the run away from it, the stop never triggered and the run finished every other branch, reporting success as if it had stopped there. A static check before the run cannot see this. - The engine ends the run as soon as every path into the stop block has been deactivated, before any further block starts, and fails it with `Stop block "<name>" (<id>) was not reached: no path this run took leads to it`. - Any run that ends without completing its stop block fails the same way: a stop block missing from the executed graph, or a Response block that ended the run first. - A loop or parallel stop with nothing to run completes at its start sentinel, whose end sentinel never runs, so that exit now counts as reaching it. - The v2 contract and the CLI `--stop-after` help describe the failure. - E2E: a condition fixture checks the stop on the taken branch still stops there, a stop on the skipped branch fails the run before the other branch's slow block finishes, and the CLI exits non-zero. * fix(executor): a skipped stop block fails a run another branch paused, and names a Response ending - A run whose stop block was proven unreachable fails even when another branch paused, instead of returning a paused run that would resume past it. - When a Response block ended the run first, the error says so rather than claiming no path leads to the stop block.
…le (#8631) * fix(projects): restore workspace deletion and tighten Project lifecycle - Archive a Project with its last active environment instead of refusing the workspace delete; account deletion follows the same rule, and the implicit archive is audited - Gate Project APIs on a `projects` AppConfig flag (PROJECT_API_ENABLED fallback) - Run Project reads in a read-only snapshot without locks; list Projects from the caller's grants with batched authorization - Batch workflow archival, Project transfer and owner reassignment; move Project ownership on organization ownership transfer - Reuse shared advisory-lock and text-array helpers; narrow admin-move conflict mapping to ProjectConflictError * improvement(projects): unify environment archive and align with shared patterns - Archive a workspace's workflows atomically with it through one archiveEnvironmentInTransaction shared by workspace delete and Project archive; the workspace row is locked before the sweep so concurrent creates are covered - Scope the Project lock timeout to lock acquisition and map lock timeouts and deadlocks to a retryable conflict; backfill and multi-Project locks use the shared advisory-lock helpers in code-unit order - Shared orchestrationFailureResponse for raw routes; contracts use the ID primitives and export only what is consumed; audit enums and mock in sync - Project restrictions section matches its sibling settings rows - Batch account-deletion Project loads/locks; skip inconsistent Projects in lists - Harden the foundation integration suite (user-keyed cleanup, poll helper, pid-scoped waits, precise assertions) * fix(projects): trim the requested organization id before validating it * fix(projects): address review on Project locking, archive notifications and list policy cost * fix(projects): keep archive retries from re-stamping MCP servers and isolate post-commit notifications
…m, restore Low (#8634) * feat(mothership): keep each chat's reasoning effort, default to medium, restore Low The simple picker offers Low / Medium / High / Extra High again, each sending exactly that effort. New chats and chats never changed run at medium instead of high. An effort the user picks is stored on the chat (copilot_chats.config) through a new PUT /api/mothership/chats/[chatId]/effort and at turn admission, and later turns of that chat keep it. The global last-used effort is no longer persisted. The Sim Chat block defaults to medium. * fix(mothership): keep the latest effort pick through refetches, failed saves and abandoned new chats * fix(mothership): leave a deduplicated send's chat on the pick its first attempt stored * fix(mothership): show a recovered chat's pick while its details load * fix(mothership): hand a withdrawn first send's effort pick back to the new-chat composer * fix(mothership): hand a withdrawn send's pick back only while its new-chat surface is open
… Tools Compared (#8638) Co-authored-by: Sim Pi Agent <pi@sim.ai>
… Integrations (#8639) Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
…8642) Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
…gine-optimization-actually-mean (#8647) Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
Contributor
There was a problem hiding this comment.
1 issue found across 180 files
Confidence score: 3/5
- In
apps/sim/executor/execution/edge-manager.ts, an active incoming edge from a disconnected component can keep a stop-after target waiting for unrelated work, even though that edge’s source will never execute. Exclude unreachable sources from the early-failure check.
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="apps/sim/executor/execution/edge-manager.ts">
<violation number="1" location="apps/sim/executor/execution/edge-manager.ts:133">
P2: An active incoming edge from a disconnected component keeps this true even though its source will never execute. A stop-after target in that component therefore cannot fail early and may wait for unrelated work to finish; account for source reachability from the execution frontier.</violation>
</file>
Contributor
|
* docs(library): update sim-vs-dedicated-chatbot-builders * Pi Babysit: address PR #8684 feedback --------- Co-authored-by: Sim Pi Agent <pi@sim.ai>
Co-authored-by: Sim Pi Agent <pi@sim.ai>
…d ungated pickup windows (#8685) * fix(desktop): scope the overdue sweep to desktop calls and keep ungated pickup windows - The overdue listing and the per-call deadline read only consider desktop tool calls, and the settlement checks the call is one the desktop runs, so a bound run's workflow call or Sim-files VFS read is never failed with the desktop's not-started result. - The cron scans only runs inside the inbox's horizon, oldest calls first, within its batch limit. - Recording a decision clears the pickup deadline only for a call that was gated; a call that was never gated keeps the window it was offered with. * fix(desktop): limit the overdue sweep over desktop calls only, and bound it by deadline The sweep's tool-name filter admitted read, grep and glob calls on Sim's own files, which the settlement then skipped, so enough of them could fill every batch and starve real desktop calls. Queries now use the SQL form of isDesktopToolCall (a desktop tool by name, or a VFS read of a granted local folder) before their limit. The sweep's horizon is now on the deadline that lapsed, not on the run's start: Sim does not enforce a run's wall clock, so a long-lived run's recently overdue call is still settled. * fix(desktop): hand a device its calls in the order they were persisted The inbox ordered calls by created_at, then tool_call_id. Calls of one turn can be persisted in the same millisecond, and then the tie broke on random ids: a device could run a click before the type the model emitted first. Calls now carry persist_seq, a strictly increasing number assigned on insert (pre-persist writes them in emission order), and the inbox and the overdue sweep order by it. The migration adds the column without a default and then sets the default, so existing rows are not rewritten; rows persisted before it have no position and sort first, as the oldest. * fix(desktop): reach every overdue desktop call, and declare the desktop tool names as const A 24 h horizon on the lapsed deadline meant a call the backstop missed for a day was never settled. The sweep starts from the few unsettled (pending or running) calls in persistence order, so it needs no horizon to stay small. The desktop tool names are a literal array declared as const, and the lookup set is derived from it. * test(mothership): give the hand-built tool call table the persistence sequence column
…finished turn (#8676) * fix(mothership): re-read the transcript until the server has saved a finished turn A tab finalizes a turn on the stream's `complete` event and reads the saved transcript. The server sends that event before it persists the turn and clears the chat's stream marker, so the read can land in the gap and return the in-flight copy: the finished stream still listed as active and the answer under its live id. That copy matches the optimistic one, so the history effect saw nothing new, and the chat's `completed` event only marks a cached live stream stale, so the tab kept it until a reload. On the local QA stack this happened in roughly one turn in four: the finished answer had no Fork action, and the cached stream marker kept a queued message from draining on a later return. Finalize now re-reads with a short backoff while the server still lists the stream it saw end. The first pass joins finalize's own read, so a turn saved in time is still read once. * fix(mothership): keep waiting for a slow save, and wait with a follow-up queued - The re-read gave up after six reads (about 8s), so a slow save left the finished turn on its in-flight copy. It now keeps reading on a capped backoff for up to two minutes, and still stops as soon as the chat moves on (another send or another chat). - It skipped the wait when a follow-up was queued, but a queued follow-up that does not go out at once (held for an edit) left the finished stream listed as running in the cache. It now runs for every finished turn; a follow-up that does go out ends it, since its send cancels the read and replaces the stream. * improvement(mothership): pace the saved-turn re-read with the shared jittered backoff
…ed stream (#8675) * fix(mothership): re-attach at once when the user returns to a recovered stream Once the user left a running chat and came back, the return recovery owned the stream for the rest of the turn, and every later online/visible/pageshow event just joined it. So a network drop after that waited out whatever the recovery was doing: a tail that went silent held the stream until the 45s idle timeout, and a tail that failed slept out a reconnect backoff of up to 30s. The same drop on a stream the send still owned re-attached immediately, because the return signal supersedes the send's reader. On the local QA stack the stream resumed 48s after the network came back. A return signal now supersedes an in-flight recovery the same way, and the new recovery re-attaches from the cursor. * fix(mothership): finish a turn that ended while its reader was silent A return recovery used the chat history's `activeStreamId` to decide what to re-attach to. When the turn had ended while this surface was not listening (its reader stalled, or a later return superseded the recovery that held it), the history listed no running turn, so recovery returned without touching the stream this surface still showed as running, and the chat stayed on Stop until an idle timeout or backoff happened to run into the terminal state. When the history lists no running turn but this surface is still sending, recovery now resolves that stream: its terminal status replays the remaining events and finalizes. * fix(mothership): leave a send waiting for admission alone on return A send shows as running before its POST is admitted, but the chat cannot list it yet. Resolving that stream on a return event read it as ended and aborted the POST. Recovery now resolves only a locally running stream whose POST was admitted. * fix(mothership): finish an admitted turn whose POST never answered on return Recovery left a send alone while its POST had not answered, so a turn the server admitted and finished, whose answer never reached the client, kept the chat on Stop with nothing to clear it. The send now counts as admitted once the loaded chat holds its message, and recovery resolves its stream as any other.
…dropping it (#8674) * fix(mothership): keep a message the server never admitted instead of dropping it Two sends were lost without an error: - A send whose POST got no response (offline, Wi-Fi drop, waking a laptop) reconnected to the stream it would have opened. That stream does not exist, so the 404 read as "finished", the turn finalized as a success, and the refetched transcript no longer held the message. A queued follow-up was lost the same way, since it had already left the queue. - A send refused with 409 because another turn held the chat (started in another tab, or one this surface lost track of) reconnected to that turn under the new message's bubble, then vanished when it finished. Both now hand the message back under its id. An unreachable send is held in the queue, so it is not redispatched into the same failure, and goes out when the browser is back online (or when the user sends it). A send that found the chat busy waits in the queue behind that turn, which the chat shows as running, and goes out when it ends. Reusing the id keeps a retry deduplicated if the server did admit the first attempt. * fix(mothership): keep held and busy-refused sends exactly once across remounts - A first message held offline on the new-chat page sat under that mount's queue key, which dies with the mount, so a reload or remount before the network returned stranded it. Held sends on a chatless surface now carry the surface they belong to, and the next chatless mount of that surface adopts them. - Held sends are released for every chat when the browser comes back online, and on mount when it already is, so a send held in a chat the user is not viewing (or one whose `online` event fired with no surface mounted) still goes out. - A send refused because the chat is busy is handed back only after the chat's running turn has been read, so the queue cannot redispatch it before that turn ends. A busy refusal that does not name the running turn no longer reads as a deduplicated send, which reconnected to a stream that never existed and lost the message. * fix(mothership): send released held messages through the queue's own drain rules Releasing held sends on mount kicked the queue dispatcher directly, which skips the drain's guards. After a reload the chat history is not loaded yet, so a follow-up queued behind a still-running turn went out at once and was refused as busy. The release now only clears the hold; the drain effect, which waits for the history and for the running turn to end, sends a released head, and now also re-runs when the head's hold clears. * fix(mothership): don't hold a send whose network returned while it was failing An `online` event can fire while the failing POST is still pending, so the release ran before the message was held and the message then waited for a release that had already happened. A send now notes whether the browser came back online while it was in flight, and if so goes back to the queue unheld, for the drain to send under its usual rules. * chore(mothership): name sign-out in desktop tool lease docs and say "not run" once Follow-ups from the #8673 review: the lease docs now name sign-out (`stopAllDesktopTools`) alongside the user's Stop as what cancels a desktop tool, and the stale-observation message no longer says it was not run twice. * fix(mothership): restore a withdrawn queued send even after the user switched chats A queued send had already left the queue when its POST failed, and a dispatch whose epoch changed meanwhile (the user switched chats) skipped restoring it, so the message was lost. A withdrawn send was never admitted, so it now goes back to its own chat's queue regardless of the epoch. * fix(mothership): don't recreate a deleted chat's queue from a late restore A withdrawn send restored after its dispatch outlived a chat switch could land after the user deleted that chat, recreating a queue (and a message) for a conversation that no longer exists. Clearing a chat's queue now leaves a session tombstone that restores respect; a new enqueue for that key lifts it.
…n agreement (#8694) The inbox and the overdue sweep filter calls with a SQL form of isDesktopToolCall, and nothing tied the two together: a desktop tool added on one side only would silently drop out of the inbox and the sweep, or let Sim-only calls crowd them. One table of calls now runs through both against real PostgreSQL, covering every named desktop tool, read and grep under, exactly at and outside user-local, glob patterns, non-string and missing arguments, a user-localX lookalike and non-desktop tools, and the test requires them to agree on every case.
* fix(plane): use a white integration background * fix(plane): preserve distinct records in capped connector syncs * fix(desktop): preserve legacy inbox timestamp precision
…an empty dev cache (#8686) * fix(ci): start each HTTP end-to-end app from an empty Turbopack dev cache The stop-after suite failed 12 of 175 CI runs (Oct 1-6), always in its own `next dev` app, which restored the Turbopack dev cache (.next/dev) the SCIM/version-compare app left in the same job. That app ran with other NEXT_PUBLIC_* values, and `next dev` SIGKILLs its server 100ms after SIGTERM. Restoring the cache: - panicked Turbopack at startup (inner_of_upper_lost_follower), 6 runs; - panicked mid-compile of the execute route ("socket connection was closed unexpectedly"), 4 runs; - wedged that compile until the 300s request timeout ("The operation timed out."), 2 runs. No run reached the executor: the cleanup check found no open execution log. - Each app step removes .next/dev before starting, so no app restores another app's cache. - A failing step prints the server log tail, not only a startup failure. - The suite compiles the execute route with a refused request under its own 300s budget and named check; every later request and CLI run is bounded at 60s. A timed-out or dropped request names the route and elapsed time, lands in the report with a null status, and a CLI timeout is reported as one. * fix(ci): start the desktop inbox app from an empty Turbopack dev cache too * fix(e2e): record a status only for a fully read stop-after response * fix(ci): stop each HTTP end-to-end app's whole process session before the next one starts * fix(e2e): compile the desktop executor routes before timing executor calls
…is off (#8669) * feat(mothership): let the worker choose models when the model picker is off With the picker off, chats sent the hidden hosted default selection (GPT-6 Astra), which the worker treats as an explicit pick for both the main agent and subagents. No selection is sent now, so the worker routes every model role itself. A stale client selection is still cleared at admission. Picker-on chats and the Sim Chat block are unchanged. * test(mothership): assert the omitted selection on the serialized worker request * test(mothership): read the serialized request without a JSON round trip
…ts lean (#8625) * perf(cli): load a manual run's draft once and keep embedded CLI results lean - Manual runs reuse the draft state the use case loaded to pick and validate the trigger, instead of reading the draft tables twice more in the execute service and the execution core. The trigger is now chosen from the same snapshot that runs. - Embedded (in-app agent) synchronous runs ask for file references without inline base64 unless --include-file-base64 is passed; --async never sends it. - Embedded logs get leaves the workflow snapshot out unless --include-workflow-state is passed, via a new embeddedRequestDefault contract field. The installed CLI and public API defaults are unchanged. - Add test-cli-run-latency-e2e.ts, which times both commands per segment against a running app and writes a JSON report; CI runs a short pass. * fix(cli): cover followed runs, assert lean results, own CI step for the CLI E2E - Embedded --follow runs also ask for file references only. - The E2E asserts this checkout's embedded results carry no inline file bytes or workflow snapshot; a compared CLI build is measured, not asserted. - Run the CLI E2E against its own self-hosted app: hosted billing admits runs through Redis, which the HTTP E2E job does not provision. * fix(cli): remove the CLI E2E's execution logs and snapshots on cleanup
… cannot trip the update-depth limit (#8702)
This branch was previously deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
workflows run --stop-after(feat(workflows): stop a manual v2 run after a block, andworkflows run --stop-after#8622)