Skip to content

feat(monitoring): alert when Plex cannot read its media files - #204

Merged
twistedmelonman merged 2 commits into
mainfrom
claude/feat-plex-reachability
Sep 23, 2026
Merged

twistedmelonman merged 2 commits into
mainfrom
claude/feat-plex-reachability

Conversation

@twistedmelonman

Copy link
Copy Markdown
Member

The last part of #199 (PR C in the plan). PR A (#200) fixed alert email under launchd; PR B (#201) added stall-watchdog for privacy prompts, long transmission-done runs and the podman supervisor.

What changes

Each plex-watchdog poll now asks Plex itself whether it can read a media file:

  1. GET /library/recentlyAdded (first 10). Take the first <Video>. TV seasons come back as <Directory>, with no file to check.
  2. GET /library/metadata/<ratingKey>?checkFiles=1, 30 s timeout.

exists="0", accessible="0" or a timeout is a failure. Two in a row send [<host>] Plex cannot read media files through alert_transition, with a 12-hour reminder and a RESOLVED: email on recovery. No movie or episode in the list, any other request error (for example a 404 when the item was removed between the calls), missing attributes, or Plex unreachable skip the check without counting it.

Asking Plex matters: the watchdog runs as /bin/bash, a different TCC identity from Plex, so reading the NAS from the script could pass while Plex is blocked.

Two fixes to the existing script came with it:

  • The check runs right after the prefs fetch, before the hash fast path, which returns early on almost every poll.
  • Step 8 rebuilt state.json from scratch, which would have dropped .transitions and .media_check_failures on every settings change. It now updates its keys in place. The second commit makes the media check start from {} on a corrupt state file; without it the watchdog exited 5 on every run and never recovered.

Testing

  • 11 new tests in tests/plex-watchdog.bats (45 total), run through the launchd-style /bin/bash 3.2 cycle. The curl mock answers each Plex URL from a fixture.
  • Checked against known-bad code: moving the check after the fast path fails 5 tests; rebuilding the state from scratch fails 8; the corrupt-state test fails without its fix.
  • All BATS suites pass; shellcheck and shfmt are clean.
  • On TILSIT: checkFiles=1 on the newest movie returns exists="1" accessible="1" in under a second. The new script, run once as operator in a throwaway HOME against the live Plex, passed the check and exited 0.

Not tested: whether a real privacy prompt makes checkFiles fail. The 2026-09-23 spot check found checkFiles=1 answering while a prompt was open. The monitoring README says this.

Deploy (not done)

Render plex-watchdog.sh with the setup script's values, diff against /Users/operator/.local/bin/plex-watchdog, back it up, install it (mode 700, operator:staff), and kickstart com.tilsit.plex-watchdog.

Closes #199.

Claude Code Bot added 2 commits September 23, 2026 16:39
The last part of #199. On 2026-09-17 a privacy prompt blocked Plex's
access to the NAS for 19 hours. plex-watchdog's only health check asks
whether the Plex API answers, which Plex can do while it cannot open a
single file.

Each plex-watchdog poll now asks Plex itself to check a file. It takes
the first movie or episode from /library/recentlyAdded (TV seasons come
back as <Directory> and have no file), then requests
/library/metadata/<ratingKey>?checkFiles=1 with a 30 s timeout. A <Part>
with exists="0" or accessible="0", or no answer in time, is a failure.
Two in a row send "[<host>] Plex cannot read media files" through
alert_transition, with a reminder every 12 hours and a RESOLVED email
when a check passes. Asking Plex matters: the watchdog's /bin/bash is a
different TCC identity, so reading the NAS from the script could pass
while Plex is blocked.

No movie or episode in the list, any other request error (a 404 when
the item was removed between the two requests), missing attributes, or
Plex being unreachable skip the check without counting it. An open
alert stays open while Plex is unreachable, so there is no false
recovery email.

The check runs right after the prefs fetch, before the hash fast path,
which returns early on almost every poll. Step 8 used to rebuild
state.json from scratch; it now updates its own keys in place, so the
new .media_check_failures and .transitions survive a settings change.

On TILSIT, checkFiles=1 on the newest movie returns exists="1"
accessible="1" in under a second. A run of the new script as operator,
in a throwaway HOME against the live Plex, passed the check and exited
0.

Tests: 10 new tests in tests/plex-watchdog.bats, run through the
launchd-style /bin/bash 3.2 cycle, with a curl mock that answers each
Plex URL from a fixture. Moving the check after the fast path fails 5
of them; going back to rebuilding the state fails 8, including the one
written for it. All BATS suites pass; shellcheck and shfmt are clean.

Docs: plex-watchdog-README.md describes the check; monitoring-README.md
updates the watchdog table and the alert_transition note, and says the
check is expected, but not tested, to catch a prompt that blocks Plex.

Not deployed. Deploy: render plex-watchdog.sh with the setup script's
values, diff against /Users/operator/.local/bin/plex-watchdog, back it
up, install it (mode 700, operator:staff), and kickstart the agent.

Closes #199.
Found after the previous commit. The media check reads state.json and
updates it with jq before anything else. With a corrupt file, jq failed
under set -e and the watchdog exited 5 on every run. Nothing rewrote the
file, so it would never have recovered. The old code rebuilt the file
from scratch in Step 8, so a corrupt file healed on the next full cycle.

The media check now starts from {} when the state is not a JSON object,
the same guard Step 8 already has. The new test writes "not json" to the
state file and expects exit 0 and a rebuilt state; it fails without the
fix. All 45 plex-watchdog tests pass; shellcheck and shfmt are clean.

Advances #199.
@claude

claude Bot commented Sep 23, 2026

Copy link
Copy Markdown

Gate Review — PR #204

State preservation (lines 437–451): Changed from jq -n rebuild to read-modify-write. Verified that piped jq assignments preserve unlisted keys (.media_check_failures, .transitions). Execution order is correct: media check writes first (line 244), Step 8 reads after (line 440).

Media check function (lines 177–273): rc initialized to 0; timeout (exit 28) handled; early returns on skippable errors; xmllint with safe defaults; alert_transition has || true suppressor per CLAUDE.md.

Integration (line 322): Runs after token extraction, no blocking failure paths.

Test coverage: 18 new tests for alert timing, timeout, recovery, skipped cases, corrupted state.

No bugs, regressions, security issues, or data loss risk detected.

VERDICT: PASS

@twistedmelonman
twistedmelonman merged commit f505609 into main Sep 23, 2026
4 checks passed
@twistedmelonman
twistedmelonman deleted the claude/feat-plex-reachability branch September 23, 2026 23:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

alert when the server is blocked on a macOS privacy prompt

1 participant