Skip to content

fix(monitoring): log alert sends and record the stall-watchdog live test - #202

Merged
twistedmelonman merged 2 commits into
mainfrom
claude/docs-stall-watchdog-live-test
Sep 23, 2026
Merged

twistedmelonman merged 2 commits into
mainfrom
claude/docs-stall-watchdog-live-test

Conversation

@twistedmelonman

Copy link
Copy Markdown
Member

Follow-up to #201, from deploying stall-watchdog on TILSIT and testing it against a real unanswered privacy prompt on 2026-09-23.

The test worked end to end. A temporary operator LaunchAgent ran an ad-hoc-signed bash copy that listed the NAS mount. The prompt opened at 15:15:03, the watchdog recorded it at 15:16:46, the alert went out at 15:20:56, Don't Allow was clicked at 15:26:47, and the recovery email went out at 15:27:06. An earlier canary answered with Allow after 34 seconds was closed without an email.

It also answered one open question. While the prompt was open, podman ps, the VM's /data, Plex checkFiles=1 and operator's own ls of the NAS all returned in under a second. So an open prompt blocks only the process it was raised for, not other network-volume access. The monitoring README records the timeline, moves both items out of "Not yet proven", and says how to repeat the test.

One gap turned up: the alert email went out, but the watchdog's log had no line for it, only msmtp.log did. The recovery path already logged RESOLVED: <key>. alert_transition now logs ALERT sent: <key>: <subject> when msmtp accepts an alert or reminder. The alert-lib.bats test for that fails against the old library and passes now. The alert-lib, stall-watchdog, plex-watchdog and pia-port-watchdog suites pass.

Not deployed. The deployed alert-lib.sh needs to be replaced with this one, which is a plain copy (sudo install -m 644 -o operator -g staff). No agent restart is needed, because each watchdog run sources it fresh.

Advances #199.

Claude Code Bot added 2 commits September 23, 2026 15:37
During the live stall-watchdog test on 2026-09-23, the alert email went
out at 15:20:56, but the watchdog's own log had no line for it. The only
record was in msmtp.log. The recovery path already logs
`RESOLVED: <key>`; the alert path logged nothing unless the send failed.
alert_transition now logs `ALERT sent: <key>: <subject>` after msmtp
accepts an alert or reminder. A failed send still logs only alert_send's
ERROR line.

In tests/alert-lib.bats, the "entering the bad state" test now checks
for the line. It fails against the old library and passes now. A new
test checks that a failed send logs no "ALERT sent" line.

Advances #199
stall-watchdog was deployed on TILSIT on 2026-09-23 and tested against a
real unanswered prompt: an ad-hoc-signed bash copy, started by a
temporary operator LaunchAgent, that listed the NAS mount. The prompt
opened at 15:15:03, the alert went out at 15:20:56, Don't Allow was
clicked at 15:26:47, and the recovery email went out at 15:27:06. An
earlier canary answered with Allow after 34 s closed without an email.

While the prompt was open, podman ps, the VM's /data, Plex checkFiles=1
and operator's own ls of the NAS all answered in under a second, so an
open prompt blocks only the process it was raised for. The README moves
the live test and that question out of "Not yet proven", records the
timeline, and says how to repeat the test (a new path and identifier
each time, because tccd remembers the answer). It also documents the
new "ALERT sent" log line.

Advances #199
@claude

claude Bot commented Sep 23, 2026

Copy link
Copy Markdown

Logging feature for alert transitions. Adds an "ALERT sent" line to the watchdog's log after each successful email, eliminating the need to cross-check msmtp.log. Tests verify logging appears on success and not on failure. No bugs, regressions, or security issues detected.

VERDICT: PASS

@twistedmelonman
twistedmelonman merged commit 9db45a2 into main Sep 23, 2026
3 checks passed
@twistedmelonman
twistedmelonman deleted the claude/docs-stall-watchdog-live-test branch September 23, 2026 22:47
twistedmelonman added a commit that referenced this pull request Sep 23, 2026
…ks (#203)

Corrects #202. Its README section, and its PR description, say an open
privacy prompt blocks only the process it was raised for. That is wrong.

The claim came from a spot check 28 seconds into the first canary's
prompt, where `podman ps`, the VM's `/data`, Plex `checkFiles` and
operator's `ls` all answered at once. The README then attributed that
check to the second prompt, the one left open for 11 minutes. During
that one, the supervisor log shows `/data` inaccessible inside the
container at 15:16:20, the container stopping at 15:18:42, `podman run
-d` hanging for 180 s from 15:21:20, and the container starting at
15:26:47, the same second the prompt was answered.

The tccd log shows the mechanism. Right after the canary's RESULT
(`17928.161`), tccd evaluated `17928.166`, a request from vfkit with the
stably-signed `gtimeout` as responsible process, and allowed it
immediately. That request had been queued behind the prompt: every one
of these requests comes through `sandboxd` (pid 17928, the first half of
the msgID), which sends them one at a time. So one unanswered prompt
stalls every later network-volume check, including for binaries that
already have a grant. That fits the 09-17 outage, where FileBot,
Transmission and Plex all stalled together.

The README now has those events in the timeline, states the blocking
behaviour and the evidence, explains why the spot check proved nothing,
and warns that a test prompt takes Transmission down while it is open.
Docs only; nothing to deploy.

Advances #199.

Co-authored-by: Claude Code Bot <claude-code@smartwatermelon.github>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant