Harden logins and lifecycle, fix Drive reliability, and tighten CI - #12
Merged
Merged
Conversation
…n CI. pip and npm now reuse their caches, each job has a timeout so a hung test cannot hold a runner for six hours, and a newer push to the same ref cancels the older run. start.sh and bin/case go through shellcheck and compose.yaml through `docker compose config`, neither of which ran anywhere before. The SPDX file list is built before the loop, so a git failure fails the job instead of checking nothing.
…environment. A v1.2.0-rc1 tag used to retag latest, since only the semver patterns skip pre-releases. The metadata JSON was pasted into a single-quoted shell string, so a quote in the repository description broke the step.
The desktop image was only ever built by the tag-triggered publish job, so a stale base digest, noVNC checksum or pin surfaced at release time. A path-filtered job now builds the desk, control and ui images without pushing them.
…anges. The advisory bumps so far were found by hand. pip-audit checks the requirements files and npm audit checks Drive's lock file.
The SDK only turns its rebinding check on when bound to loopback, and compose binds 0.0.0.0, so any web page could rebind a name to 127.0.0.1:8788 and call computer_exec. The server now allows the same names cased and Drive do, plus the mcp service name, and compose hands it CASE_PUBLIC_HOST and CASE_ALLOWED_HOSTS. A foreign Host gets 421, a foreign Origin 403.
They were declared `x: int = None`, so the schema said integer and a client that sends "x": null for an unset argument got a validation error. They are `int | None = None` now, and null means unset.
…rest. FastMCP calls sync tools on the event loop, so a 280s computer_login froze every other client and /health until it returned. Tools are now registered through a small decorator that hands the blocking call to anyio's thread pool and leaves the module's functions as they were.
_headers took extra headers no caller passed, and the end of the auth_attempt_wait loop had two branches that both continued. The loop keeps its one early return for a human wait that is out of budget.
It failed on any tool name containing "cred", so a credential_list would have tripped it. It now looks for a write verb next to "cred", which is the invariant SECURITY.md states.
injecting was a bool, so with /login, /login/resume and /auth/submit_challenge overlapping the first to finish reopened the gate while another was still typing. It is now held while any injecting route runs, and a second /login answers 409 instead of resetting the first one's state. The gate was also only checked when /eval, /action and /exec started, so an awaited eval begun before a login could read the password field and return it. Each injection now bumps a generation, and those routes answer 423 on the way out if one started while they ran.
… up. deskd's navigate, field wait, two settles, post-submit wait and TOTP settle add up to about 112s, and cased stopped listening at 95s. A slow login was then failed as a desk error while deskd kept typing and held the gate. cased now waits 125s.
A background job inherits the output pipes, so `sleep 30 & echo started` waited for their EOF until the timeout and reported 124, and a timeout killed only bash and orphaned its children. The command now runs in its own session; /exec stops reading half a second after bash exits and kills the process group on timeout. computer_exec tells agents to start background jobs with nohup and redirected output.
A non-numeric timeout_s or Content-Length, or a /login or /login/resume body without its credential or value, raised inside the handler and came back as a bare 500. Those are 400 bad_request now. navigate() ignored Page.navigate's errorText, so a DNS failure landed on chrome-error:// and login reported a foreign page origin; it now fails with the navigation error.
The blocker watchdog runs every 2s on every awake desk and made two CDP round trips for it. One array expression returns both, with the text cut at the same 3000 chars.
Both listed the CDP targets, picked the first page and opened a websocket to it with the same options. page_ws() does that once.
…h 3.2.
The env file was sourced without export, so nohup cased and the cred
add python never saw CASE_TOKEN and friends, and `up` sourced it a
second time. It is sourced once under set -a now. "${auth[@]}" with
CASE_TOKEN unset died as an unbound variable under set -u on macOS's
bash 3.2; the ${auth[@]+…} form expands to nothing there.
.dockerignore patterns match from the context root, so web/node_modules and the __pycache__ directories under control-plane and mcp were sent to the daemon, and COPY web/ laid the host's node_modules over the one npm ci had just built.
start.sh passed any WxHxD through, but deskd's framebuffer grab only reads 32bpp, so a desk started at 1280x800x16 came up with every screenshot failing. Another depth is now replaced with 24 and logged.
The trap was installed only after everything had started, and PID 1 ignores a signal it has no handler for, so a sleep during startup did nothing until docker's SIGKILL. The quiesce call (6s) and the wait for chromium (8s) also outlasted docker stop's 10s, so the pkill fallback never ran. The trap now comes first, each step is a no-op for what has not started, and the budget is 3s + 4s + 2s.
annotated-doc, annotated-types, pydantic-core, typing-extensions and typing-inspection still floated, although the comment says the transitives are pinned. They are pinned to what pip resolves for the current direct pins, versions only as before.
The pip.conf layer sat after COPY deskd.py, so every deskd edit re-ran it; it now comes before the files that change. The comment above the look assets said configs are seeded only when missing, but start.sh force-seeds once after each LOOK bump.
The scheduler resolves named zones through zoneinfo, which reads the system database first and the tzdata package second. A slim base image may have neither, and then every tz-bearing schedule fell back as if its zone were unknown.
The theme ships only left_ptr and its index.theme named no parent, so text, hand and resize cursors had no themed image to come from. Adwaita is installed with GTK 3, which chromium and xfwm4 pull in.
…owed. MCP HTTP now checks Host against the cased and Drive list, so a proxy that forwards another name gets 421 until that name is in CASE_ALLOWED_HOSTS.
Four test files imported serve.mjs with no CASE_THREADS, so they loaded web/web-ui/threads.json and, when it had threads, rewrote it. Each now points CASE_THREADS at a fresh temp file before the import.
The test script chained the eight files by name, so a new file was skipped until someone remembered to add it, and the web README listed only three. The script now loops over web-ui/test_*.mjs, still one node process per file, and the README points at npm --prefix web test.
The router returned route promises from inside its try without awaiting them, and readBody let the stream's abort reject, so STOP during an attachment upload became an unhandled rejection that ended the process and every running turn with it. schedulesRoute also looked up the computer outside its try, so cased being down did the same. Routes are now awaited, readBody treats a hung-up client as no body, the lookup sits inside the try, and a process-level handler logs anything that still slips through instead of exiting.
These routes used cachedCid, the first awake computer cased listed, because the page never said which computer it was sat at. With two computers, the Files and Credentials tabs showed and changed the other desk's disk and vault. The page now sends computer_id on every such request and the server uses it; a request with no pick gets a 409 instead of being rerouted. Panes opened before the first refresh knew the pick reload once it does.
listen awaited onMessage, and a message that starts a task only settles when the turn ends. Every later message queued behind it, so a steer never reached the running turn and an OTP or approval for a handoff that turn was waiting on sat unread until the wait timed out. Messages are now dispatched the way the Telegram poller does it, with failures still logged.
rateWaitS indexed err.headers like a plain object, but both the OpenAI and Anthropic SDKs hand back a fetch Headers there, so the server's retry-after was never read and every 429 fell back to the exponential guess. It now reads the header through get() when there is one, and still accepts a plain object.
Thirteen test files pointed CASE_HOME at a fixed /tmp/case-*-test that was never removed, so two runs of the same file at once shared one SQLite vault and failed on each other's rows, and a later run started from an earlier run's leftovers. browse and navigate used setdefault and so took an inherited CASE_HOME. Every control-plane test now calls tests/_helpers.isolated_home(), which puts control-plane on sys.path and CASE_HOME in a fresh temporary directory that is removed at exit. The vault permissions test and the runs pruning test keep their scratch files under that directory too.
Twenty files carried their own copy of the loop that calls each test_* function, and test_token, test_dockerd, test_store and test_acceptance_safety called their tests by name instead, so a test added there without also being added to that list never ran. _helpers.run_tests now runs every test_* function in a file, keeps going after a failure, prints each failure and exits non-zero if there was one. The same 508 tests run as before.
test_assist and test_links each defined the same running-computer seed and /v1/desk/check caller, test_auth_attempts and test_lifecycle the same ApiError assertion, and test_auth_mcp, test_mcp_http and test_skills three near-identical ways to re-import case_mcp under a clean environment. They live in tests/_helpers.py now. The case_mcp loader clears the union of the settings the three copies cleared.
wait_attempt re-reads SQLite every 15 seconds in case it missed a publish, but the interval was a literal, so the test for that re-read had to sit out a real five-second wait. It is WAIT_REREAD_S now, with the same value, and the test shortens it.
test_login_success_ungated_reports_success ran the real late-challenge poll, which spent its whole 12-second deadline dialling 127.0.0.1:1; it is mocked now, as the gate beside it already was. The navigate and browse poll loops slept on the wall clock between faked evals, so their tests now give those modules a clock that moves on when slept.
…vent. test_captcha imported store at the top and again inside a helper, and nineteen tests imported cased or login_flow without using them; one side effect had an if whose body was pass, behind a condition padded with an "if False" branch. test_missing_proof_spec_ends_unverified patched events.emit and never looked at it, so it now checks that the unverified outcome is announced as login_completed. The continuation lines of the shared raises helper line up with its call again.
Nothing caught an unused import or a name redefined before use, and test_captcha had collected twenty-one of them. The test job now runs ruff's pyflakes and syntax-error checks over control-plane, mcp, image and tests, with ruff pinned in requirements-dev.txt.
Each test file points CASE_HOME at its own temporary vault when it is imported, but config reads CASE_HOME once, so a single pytest process across the files ends up with one shared vault and tests that expect an empty one fail. The local test commands also run the lint CI runs.
Phone replies such as "yes" used to approve; they are now refused and the handoff keeps waiting, which the Replies and handoffs section did not say.
An /exec, /action or /eval that started before an injection now gets the 423 too, and a second login is refused while one is injecting. SECURITY.md only described requests that arrive during the injection.
pip-audit resolves every -r file into one environment, and the two sets pin different fastapi versions, so the combined run could never resolve.
The runner's own shellcheck is older than the one contributors install and flags `A && B || true` (SC2015), so the same script passed locally and failed in CI. shellcheck now comes from requirements-dev.txt like ruff, and the down branch says what it meant: stop colima on macOS, ignore failure.
Only /login checked for an injection already in flight, so a resume and a challenge submit could type into the same tab at once, and a resume cleared the held login before it knew it could run. Both now answer 409 injection_running first; cased treats that as a soft failure and leaves the handoff pending for a retry.
Typing into a new task or another thread while a turn ran steered the running turn, invisibly, and queued prompts were always sent back to the thread that ran. A message now steers only the turn on screen; typed anywhere else it waits in the queue with its thread, or with a new task, and goes there when the turn ends.
…holds them. The readers blocked until EOF, so every `cmd &` left two threads, two pipe descriptors and up to 2 x CAP of buffer alive for as long as the job ran. They now poll and stop when the call's grace period ends, which closes the pipes as the call did before it stopped waiting for them.
The last run out held the scheduler's single lock while Docker stopped its box, so a slow stop delayed every other computer's schedules. A per-computer gate now orders a starting run's wake after that sleep, and the shared lock only guards the holder counts.
A prompt queued for another thread was sent when openThread resolved, even if a different thread had been picked while it loaded, so it went to the new pick. drainQ now sends only into the queued thread and openThread drains once that thread has loaded; picked away, the prompt stays queued. Prompts queued while their thread was being born take its real id when it arrives.
openThread marks a thread active before its fetch returns, so a queued prompt, or one typed in that moment, was sent into a view the load then replaced, hiding the prompt and its reply. drainQ now waits while the thread is still loading, a prompt typed into a loading thread is queued, and openThread sends it once the thread is up.
A failed load kept the thread active with nothing shown: clicking it again did nothing, and a queued prompt could later be sent into history that never appeared. A failed load now leaves no thread open, so a click or the queue retries it and the prompt waits for a load that succeeds.
A queued prompt for a deleted thread stayed at the head of the queue: each drain retried the thread and stopped there, stranding the prompts behind it. When a thread fails to load because it is gone, or is deleted, its queued prompts start a new task instead and the queue moves on; a server that cannot be reached still leaves them queued for a retry.
Any failed load orphaned the thread's queued prompts, so a 401 after a token change or a 5xx moved them to a new task that was just as likely to fail, and they were lost. Only a 404 does that now; any other failure leaves them queued for their own thread.
After a 401 or a server error the prompt stayed at the head of the queue, nothing retried it, and every prompt behind it waited too. It is now held: drains skip it, the rest of the queue moves on, and it goes out when its thread next loads. Nothing reopens the thread on a timer, since that would switch the view unasked; the chip says it is waiting. Drains also take a prompt for the thread already on screen first, so they never pull the view away from it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What and why
A repo-wide audit, fixed in 114 small commits. Each commit says what was wrong and what happens now; nothing here removes a capability. Highlights:
Host/Origin(under compose it bound0.0.0.0with no rebinding protection, so a web page could callcomputer_exec). Approval handoffs take onlyapproveordeny, and the public/answerdoor returns the public handoff shape. The injection gate also covers/eval,/execand/actioncalls already running when an injection starts, and overlapping injections no longer reopen it early. A refused one-time code is reported as refused; a stale TOTP seed is tried once per 30s window.retry-afteris honoured, Anthropic turns are budgeted by billed tokens, and retried rounds no longer render twice.CASE_LOCAL, which nothing read, is gone. Host-only settings are documented.Upgrading: an MCP behind a reverse proxy needs its hostname in
CASE_ALLOWED_HOSTS, and phone approvals must beapproveordeny. Rebuild the desktop image and recreate existing containers to get the new deskd.Checklist
CLA-SIGNERS) — see CONTRIBUTING.mdSPDX-License-Identifierof their directorytests/test_*.pyexcept acceptance,npm --prefix web test, andtest_deskd.pyon the image's pinsTouches credentials?
Yes:
image/deskd.py(login, resume, submit_challenge),control-plane/store.py, andlinks.py(cookie parsing only). Invariants checked in SECURITY.md:CLEAR_PASSstill runs before the gate reopens. There are tests for overlap and for an eval that raced an injection./answerstill requires the HMAC and no longer returns the screenshot or credential name; link and Assist semantics are unchanged. The acceptance suite was not run; it needs Docker.Generated by Claude Code
The latest changes appear safe to merge.
Summary
The changes since the previous review hold queued prompts when their thread cannot load, let other prompts proceed, and release held prompts when that thread next loads successfully.
Reviews (8) · Last reviewed commit: "Hold a prompt whose thread failed to loa..."