Skip to content

Harden logins and lifecycle, fix Drive reliability, and tighten CI - #12

Merged
ishu86 merged 126 commits into
mainfrom
refactor-cleanup
Sep 24, 2026
Merged

ishu86 merged 126 commits into
mainfrom
refactor-cleanup

Conversation

@ishu86

@ishu86 ishu86 commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

What and why

A repo-wide audit, fixed in 114 small commits. Each commit says what was wrong and what happens now; nothing here removes a capability. Highlights:

  • Security: MCP HTTP now checks Host/Origin (under compose it bound 0.0.0.0 with no rebinding protection, so a web page could call computer_exec). Approval handoffs take only approve or deny, and the public /answer door returns the public handoff shape. The injection gate also covers /eval, /exec and /action calls already running when an injection starts, and overlapping injections no longer reopen it early. A refused one-time code is reported as refused; a stale TOTP seed is tried once per 30s window.
  • Data safety: a wake that lost a race with a sleep or the sweeper deleted the computer's home volume. Wake, sleep and destroy are now serialized per computer, and reconcile skips computers with an operation in flight.
  • Lifecycle and scheduling: failed wakes stop their container, admission is atomic, one schedule no longer sleeps a box another is using, the session keeper leaves boxes someone is using alone, the live view counts as activity, and legacy schedule rows no longer raise every sweep.
  • Drive: an aborted request or a down cased no longer kills the process. Files, Credentials, Schedules and teach act on the picked computer. ntfy replies are handled mid-turn, retry-after is honoured, Anthropic turns are budgeted by billed tokens, and retried rounds no longer render twice.
  • Performance: index-friendly auth-attempt queries, one Docker listing per reconcile, keep-alive desk connections, incremental history trim, one read per screenshot per turn.
  • Config and docs: the usage-stats opt-out, solver proxy, brain timeout and session keeper settings now reach cased under compose. CASE_LOCAL, which nothing read, is gone. Host-only settings are documented.
  • Tests and CI: every test file has an isolated home and a shared runner, and the suite takes 23s instead of 63s with the same 508 tests. CI adds ruff, shellcheck, a compose check, deskd tests against the image's own pins on Python 3.11, image builds on PRs that touch them, and a weekly dependency audit.

Upgrading: an MCP behind a reverse proxy needs its hostname in CASE_ALLOWED_HOSTS, and phone approvals must be approve or deny. Rebuild the desktop image and recreate existing containers to get the new deskd.

Checklist

  • I have signed the CLA (my GitHub username is in CLA-SIGNERS) — see CONTRIBUTING.md
  • New source files carry the SPDX-License-Identifier of their directory
  • Unit tests pass: every tests/test_*.py except acceptance, npm --prefix web test, and test_deskd.py on the image's pins
  • No secrets in the diff: tokens, API keys, ntfy topics, hostnames

Touches credentials?

Yes: image/deskd.py (login, resume, submit_challenge), control-plane/store.py, and links.py (cookie parsing only). Invariants checked in SECURITY.md:

  • The 423 while injecting: a count replaces the boolean, and a generation check refuses routes that raced an injection. CLEAR_PASS still runs before the gate reopens. There are tests for overlap and for an eval that raced an injection.
  • Origin checked before inserting: unchanged. Submitted codes are now checked against page text only, never the URL.
  • Token doors: /answer still requires the HMAC and no longer returns the screenshot or credential name; link and Assist semantics are unchanged. The acceptance suite was not run; it needs Docker.

Generated by Claude Code

RetriggerConfidence Score: 5/5

The latest changes appear safe to merge.

Summary

The changes since the previous review hold queued prompts when their thread cannot load, let other prompts proceed, and release held prompts when that thread next loads successfully.

Reviews (8) · Last reviewed commit: "Hold a prompt whose thread failed to loa..."

…n CI.

pip and npm now reuse their caches, each job has a timeout so a hung
test cannot hold a runner for six hours, and a newer push to the same
ref cancels the older run. start.sh and bin/case go through shellcheck
and compose.yaml through `docker compose config`, neither of which ran
anywhere before. The SPDX file list is built before the loop, so a git
failure fails the job instead of checking nothing.
…environment.

A v1.2.0-rc1 tag used to retag latest, since only the semver patterns
skip pre-releases. The metadata JSON was pasted into a single-quoted
shell string, so a quote in the repository description broke the step.
The desktop image was only ever built by the tag-triggered publish job,
so a stale base digest, noVNC checksum or pin surfaced at release time.
A path-filtered job now builds the desk, control and ui images without
pushing them.
…anges.

The advisory bumps so far were found by hand. pip-audit checks the
requirements files and npm audit checks Drive's lock file.
The SDK only turns its rebinding check on when bound to loopback, and
compose binds 0.0.0.0, so any web page could rebind a name to
127.0.0.1:8788 and call computer_exec. The server now allows the same
names cased and Drive do, plus the mcp service name, and compose hands
it CASE_PUBLIC_HOST and CASE_ALLOWED_HOSTS. A foreign Host gets 421, a
foreign Origin 403.
They were declared `x: int = None`, so the schema said integer and a
client that sends "x": null for an unset argument got a validation
error. They are `int | None = None` now, and null means unset.
…rest.

FastMCP calls sync tools on the event loop, so a 280s computer_login
froze every other client and /health until it returned. Tools are now
registered through a small decorator that hands the blocking call to
anyio's thread pool and leaves the module's functions as they were.
_headers took extra headers no caller passed, and the end of the
auth_attempt_wait loop had two branches that both continued. The loop
keeps its one early return for a human wait that is out of budget.
It failed on any tool name containing "cred", so a credential_list
would have tripped it. It now looks for a write verb next to "cred",
which is the invariant SECURITY.md states.
injecting was a bool, so with /login, /login/resume and
/auth/submit_challenge overlapping the first to finish reopened the
gate while another was still typing. It is now held while any
injecting route runs, and a second /login answers 409 instead of
resetting the first one's state. The gate was also only checked when
/eval, /action and /exec started, so an awaited eval begun before a
login could read the password field and return it. Each injection now
bumps a generation, and those routes answer 423 on the way out if one
started while they ran.
… up.

deskd's navigate, field wait, two settles, post-submit wait and TOTP
settle add up to about 112s, and cased stopped listening at 95s. A slow
login was then failed as a desk error while deskd kept typing and held
the gate. cased now waits 125s.
A background job inherits the output pipes, so `sleep 30 & echo
started` waited for their EOF until the timeout and reported 124, and
a timeout killed only bash and orphaned its children. The command now
runs in its own session; /exec stops reading half a second after bash
exits and kills the process group on timeout. computer_exec tells
agents to start background jobs with nohup and redirected output.
A non-numeric timeout_s or Content-Length, or a /login or /login/resume
body without its credential or value, raised inside the handler and
came back as a bare 500. Those are 400 bad_request now. navigate()
ignored Page.navigate's errorText, so a DNS failure landed on
chrome-error:// and login reported a foreign page origin; it now fails
with the navigation error.
The blocker watchdog runs every 2s on every awake desk and made two CDP
round trips for it. One array expression returns both, with the text
cut at the same 3000 chars.
Both listed the CDP targets, picked the first page and opened a
websocket to it with the same options. page_ws() does that once.
…h 3.2.

The env file was sourced without export, so nohup cased and the cred
add python never saw CASE_TOKEN and friends, and `up` sourced it a
second time. It is sourced once under set -a now. "${auth[@]}" with
CASE_TOKEN unset died as an unbound variable under set -u on macOS's
bash 3.2; the ${auth[@]+…} form expands to nothing there.
.dockerignore patterns match from the context root, so web/node_modules
and the __pycache__ directories under control-plane and mcp were sent
to the daemon, and COPY web/ laid the host's node_modules over the one
npm ci had just built.
start.sh passed any WxHxD through, but deskd's framebuffer grab only
reads 32bpp, so a desk started at 1280x800x16 came up with every
screenshot failing. Another depth is now replaced with 24 and logged.
The trap was installed only after everything had started, and PID 1
ignores a signal it has no handler for, so a sleep during startup did
nothing until docker's SIGKILL. The quiesce call (6s) and the wait for
chromium (8s) also outlasted docker stop's 10s, so the pkill fallback
never ran. The trap now comes first, each step is a no-op for what has
not started, and the budget is 3s + 4s + 2s.
annotated-doc, annotated-types, pydantic-core, typing-extensions and
typing-inspection still floated, although the comment says the
transitives are pinned. They are pinned to what pip resolves for the
current direct pins, versions only as before.
The pip.conf layer sat after COPY deskd.py, so every deskd edit re-ran
it; it now comes before the files that change. The comment above the
look assets said configs are seeded only when missing, but start.sh
force-seeds once after each LOOK bump.
The scheduler resolves named zones through zoneinfo, which reads the
system database first and the tzdata package second. A slim base image
may have neither, and then every tz-bearing schedule fell back as if
its zone were unknown.
The theme ships only left_ptr and its index.theme named no parent, so
text, hand and resize cursors had no themed image to come from.
Adwaita is installed with GTK 3, which chromium and xfwm4 pull in.
…owed.

MCP HTTP now checks Host against the cased and Drive list, so a proxy
that forwards another name gets 421 until that name is in
CASE_ALLOWED_HOSTS.
Four test files imported serve.mjs with no CASE_THREADS, so they loaded
web/web-ui/threads.json and, when it had threads, rewrote it. Each now
points CASE_THREADS at a fresh temp file before the import.
The test script chained the eight files by name, so a new file was
skipped until someone remembered to add it, and the web README listed
only three. The script now loops over web-ui/test_*.mjs, still one node
process per file, and the README points at npm --prefix web test.
The router returned route promises from inside its try without awaiting
them, and readBody let the stream's abort reject, so STOP during an
attachment upload became an unhandled rejection that ended the process
and every running turn with it. schedulesRoute also looked up the
computer outside its try, so cased being down did the same. Routes are
now awaited, readBody treats a hung-up client as no body, the lookup
sits inside the try, and a process-level handler logs anything that
still slips through instead of exiting.
These routes used cachedCid, the first awake computer cased listed,
because the page never said which computer it was sat at. With two
computers, the Files and Credentials tabs showed and changed the other
desk's disk and vault. The page now sends computer_id on every such
request and the server uses it; a request with no pick gets a 409
instead of being rerouted. Panes opened before the first refresh knew
the pick reload once it does.
listen awaited onMessage, and a message that starts a task only settles
when the turn ends. Every later message queued behind it, so a steer
never reached the running turn and an OTP or approval for a handoff that
turn was waiting on sat unread until the wait timed out. Messages are
now dispatched the way the Telegram poller does it, with failures still
logged.
rateWaitS indexed err.headers like a plain object, but both the OpenAI
and Anthropic SDKs hand back a fetch Headers there, so the server's
retry-after was never read and every 429 fell back to the exponential
guess. It now reads the header through get() when there is one, and
still accepts a plain object.
Thirteen test files pointed CASE_HOME at a fixed /tmp/case-*-test that
was never removed, so two runs of the same file at once shared one
SQLite vault and failed on each other's rows, and a later run started
from an earlier run's leftovers. browse and navigate used setdefault
and so took an inherited CASE_HOME. Every control-plane test now calls
tests/_helpers.isolated_home(), which puts control-plane on sys.path
and CASE_HOME in a fresh temporary directory that is removed at exit.
The vault permissions test and the runs pruning test keep their scratch
files under that directory too.
Twenty files carried their own copy of the loop that calls each test_*
function, and test_token, test_dockerd, test_store and
test_acceptance_safety called their tests by name instead, so a test
added there without also being added to that list never ran.
_helpers.run_tests now runs every test_* function in a file, keeps going
after a failure, prints each failure and exits non-zero if there was
one. The same 508 tests run as before.
test_assist and test_links each defined the same running-computer seed
and /v1/desk/check caller, test_auth_attempts and test_lifecycle the
same ApiError assertion, and test_auth_mcp, test_mcp_http and
test_skills three near-identical ways to re-import case_mcp under a
clean environment. They live in tests/_helpers.py now. The case_mcp
loader clears the union of the settings the three copies cleared.
wait_attempt re-reads SQLite every 15 seconds in case it missed a
publish, but the interval was a literal, so the test for that re-read
had to sit out a real five-second wait. It is WAIT_REREAD_S now, with
the same value, and the test shortens it.
test_login_success_ungated_reports_success ran the real late-challenge
poll, which spent its whole 12-second deadline dialling 127.0.0.1:1; it
is mocked now, as the gate beside it already was. The navigate and
browse poll loops slept on the wall clock between faked evals, so
their tests now give those modules a clock that moves on when slept.
…vent.

test_captcha imported store at the top and again inside a helper, and
nineteen tests imported cased or login_flow without using them; one
side effect had an if whose body was pass, behind a condition padded
with an "if False" branch. test_missing_proof_spec_ends_unverified
patched events.emit and never looked at it, so it now checks that the
unverified outcome is announced as login_completed. The continuation
lines of the shared raises helper line up with its call again.
Nothing caught an unused import or a name redefined before use, and
test_captcha had collected twenty-one of them. The test job now runs
ruff's pyflakes and syntax-error checks over control-plane, mcp, image
and tests, with ruff pinned in requirements-dev.txt.
Each test file points CASE_HOME at its own temporary vault when it is
imported, but config reads CASE_HOME once, so a single pytest process
across the files ends up with one shared vault and tests that expect
an empty one fail. The local test commands also run the lint CI runs.
Phone replies such as "yes" used to approve; they are now refused and
the handoff keeps waiting, which the Replies and handoffs section did not
say.
An /exec, /action or /eval that started before an injection now gets
the 423 too, and a second login is refused while one is injecting.
SECURITY.md only described requests that arrive during the injection.
pip-audit resolves every -r file into one environment, and the two sets
pin different fastapi versions, so the combined run could never resolve.
The runner's own shellcheck is older than the one contributors install and
flags `A && B || true` (SC2015), so the same script passed locally and
failed in CI. shellcheck now comes from requirements-dev.txt like ruff, and
the down branch says what it meant: stop colima on macOS, ignore failure.
Comment thread image/deskd.py Outdated
Comment thread web/web-ui/index.html Outdated
Comment thread image/deskd.py
Comment thread control-plane/scheduler.py Outdated
Only /login checked for an injection already in flight, so a resume and
a challenge submit could type into the same tab at once, and a resume
cleared the held login before it knew it could run. Both now answer 409
injection_running first; cased treats that as a soft failure and leaves
the handoff pending for a retry.
Typing into a new task or another thread while a turn ran steered the
running turn, invisibly, and queued prompts were always sent back to the
thread that ran. A message now steers only the turn on screen; typed
anywhere else it waits in the queue with its thread, or with a new task,
and goes there when the turn ends.
…holds them.

The readers blocked until EOF, so every `cmd &` left two threads, two
pipe descriptors and up to 2 x CAP of buffer alive for as long as the job
ran. They now poll and stop when the call's grace period ends, which
closes the pipes as the call did before it stopped waiting for them.
The last run out held the scheduler's single lock while Docker stopped
its box, so a slow stop delayed every other computer's schedules. A
per-computer gate now orders a starting run's wake after that sleep, and
the shared lock only guards the holder counts.
Comment thread web/web-ui/index.html Outdated
A prompt queued for another thread was sent when openThread resolved,
even if a different thread had been picked while it loaded, so it went
to the new pick. drainQ now sends only into the queued thread and
openThread drains once that thread has loaded; picked away, the prompt
stays queued. Prompts queued while their thread was being born take
its real id when it arrives.
Comment thread web/web-ui/index.html Outdated
openThread marks a thread active before its fetch returns, so a queued
prompt, or one typed in that moment, was sent into a view the load then
replaced, hiding the prompt and its reply. drainQ now waits while the
thread is still loading, a prompt typed into a loading thread is queued,
and openThread sends it once the thread is up.
Comment thread web/web-ui/index.html Outdated
A failed load kept the thread active with nothing shown: clicking it
again did nothing, and a queued prompt could later be sent into history
that never appeared. A failed load now leaves no thread open, so a click
or the queue retries it and the prompt waits for a load that succeeds.
Comment thread web/web-ui/index.html Outdated
A queued prompt for a deleted thread stayed at the head of the queue: each
drain retried the thread and stopped there, stranding the prompts behind
it. When a thread fails to load because it is gone, or is deleted, its
queued prompts start a new task instead and the queue moves on; a server
that cannot be reached still leaves them queued for a retry.
Comment thread web/web-ui/index.html Outdated
Any failed load orphaned the thread's queued prompts, so a 401 after a
token change or a 5xx moved them to a new task that was just as likely to
fail, and they were lost. Only a 404 does that now; any other failure
leaves them queued for their own thread.
Comment thread web/web-ui/index.html Outdated
After a 401 or a server error the prompt stayed at the head of the queue,
nothing retried it, and every prompt behind it waited too. It is now held:
drains skip it, the rest of the queue moves on, and it goes out when its
thread next loads. Nothing reopens the thread on a timer, since that would
switch the view unasked; the chip says it is waiting. Drains also take a
prompt for the thread already on screen first, so they never pull the view
away from it.
@ishu86
ishu86 merged commit a81fea1 into main Sep 24, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant