Repository navigation
fix(monitoring): log alert sends and record the stall-watchdog live test - #202
Merged
Merged
Conversation
added 2 commits
September 23, 2026 15:37
During the live stall-watchdog test on 2026-09-23, the alert email went out at 15:20:56, but the watchdog's own log had no line for it. The only record was in msmtp.log. The recovery path already logs `RESOLVED: <key>`; the alert path logged nothing unless the send failed. alert_transition now logs `ALERT sent: <key>: <subject>` after msmtp accepts an alert or reminder. A failed send still logs only alert_send's ERROR line. In tests/alert-lib.bats, the "entering the bad state" test now checks for the line. It fails against the old library and passes now. A new test checks that a failed send logs no "ALERT sent" line. Advances #199
stall-watchdog was deployed on TILSIT on 2026-09-23 and tested against a real unanswered prompt: an ad-hoc-signed bash copy, started by a temporary operator LaunchAgent, that listed the NAS mount. The prompt opened at 15:15:03, the alert went out at 15:20:56, Don't Allow was clicked at 15:26:47, and the recovery email went out at 15:27:06. An earlier canary answered with Allow after 34 s closed without an email. While the prompt was open, podman ps, the VM's /data, Plex checkFiles=1 and operator's own ls of the NAS all answered in under a second, so an open prompt blocks only the process it was raised for. The README moves the live test and that question out of "Not yet proven", records the timeline, and says how to repeat the test (a new path and identifier each time, because tccd remembers the answer). It also documents the new "ALERT sent" log line. Advances #199
|
Logging feature for alert transitions. Adds an "ALERT sent" line to the watchdog's log after each successful email, eliminating the need to cross-check msmtp.log. Tests verify logging appears on success and not on failure. No bugs, regressions, or security issues detected. VERDICT: PASS |
twistedmelonman
added a commit
that referenced
this pull request
Sep 23, 2026
…ks (#203) Corrects #202. Its README section, and its PR description, say an open privacy prompt blocks only the process it was raised for. That is wrong. The claim came from a spot check 28 seconds into the first canary's prompt, where `podman ps`, the VM's `/data`, Plex `checkFiles` and operator's `ls` all answered at once. The README then attributed that check to the second prompt, the one left open for 11 minutes. During that one, the supervisor log shows `/data` inaccessible inside the container at 15:16:20, the container stopping at 15:18:42, `podman run -d` hanging for 180 s from 15:21:20, and the container starting at 15:26:47, the same second the prompt was answered. The tccd log shows the mechanism. Right after the canary's RESULT (`17928.161`), tccd evaluated `17928.166`, a request from vfkit with the stably-signed `gtimeout` as responsible process, and allowed it immediately. That request had been queued behind the prompt: every one of these requests comes through `sandboxd` (pid 17928, the first half of the msgID), which sends them one at a time. So one unanswered prompt stalls every later network-volume check, including for binaries that already have a grant. That fits the 09-17 outage, where FileBot, Transmission and Plex all stalled together. The README now has those events in the timeline, states the blocking behaviour and the evidence, explains why the spot check proved nothing, and warns that a test prompt takes Transmission down while it is open. Docs only; nothing to deploy. Advances #199. Co-authored-by: Claude Code Bot <claude-code@smartwatermelon.github>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #201, from deploying stall-watchdog on TILSIT and testing it against a real unanswered privacy prompt on 2026-09-23.
The test worked end to end. A temporary operator LaunchAgent ran an ad-hoc-signed bash copy that listed the NAS mount. The prompt opened at 15:15:03, the watchdog recorded it at 15:16:46, the alert went out at 15:20:56, Don't Allow was clicked at 15:26:47, and the recovery email went out at 15:27:06. An earlier canary answered with Allow after 34 seconds was closed without an email.
It also answered one open question. While the prompt was open,
podman ps, the VM's/data, PlexcheckFiles=1and operator's ownlsof the NAS all returned in under a second. So an open prompt blocks only the process it was raised for, not other network-volume access. The monitoring README records the timeline, moves both items out of "Not yet proven", and says how to repeat the test.One gap turned up: the alert email went out, but the watchdog's log had no line for it, only
msmtp.logdid. The recovery path already loggedRESOLVED: <key>.alert_transitionnow logsALERT sent: <key>: <subject>when msmtp accepts an alert or reminder. Thealert-lib.batstest for that fails against the old library and passes now. The alert-lib, stall-watchdog, plex-watchdog and pia-port-watchdog suites pass.Not deployed. The deployed
alert-lib.shneeds to be replaced with this one, which is a plain copy (sudo install -m 644 -o operator -g staff). No agent restart is needed, because each watchdog run sources it fresh.Advances #199.