Repository navigation
docs(monitoring): an open TCC prompt blocks other network-volume checks - #203
Merged
Merged
Conversation
The live-test section added in #202 says an open prompt blocks only the process it was raised for. That is wrong. It rested on a spot check 28 s into the first canary's prompt, and the README then attributed the check to the second prompt. The supervisor log for the second prompt shows the opposite: /data inaccessible inside the container at 15:16:20 (prompt open since 15:15:03), the container stopping at 15:18:42, `podman run -d` hanging for 180 s from 15:21:20, and the container starting at 15:26:47, the same second the prompt was answered. The tccd log shows why: right after the canary's RESULT (17928.161), tccd evaluated 17928.166, a request from vfkit with the stably-signed gtimeout as responsible process, and allowed it at once. It had waited behind the prompt, because sandboxd (pid 17928) sends these requests one at a time. The README now records those events in the timeline, states that an open prompt blocks every later network-volume check, says why the spot check proved nothing, and warns that a test prompt means a Transmission outage while it is open. This matches the 09-17 outage, where FileBot, Transmission and Plex stalled together. Advances #199
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Corrects #202. Its README section, and its PR description, say an open privacy prompt blocks only the process it was raised for. That is wrong.
The claim came from a spot check 28 seconds into the first canary's prompt, where
podman ps, the VM's/data, PlexcheckFilesand operator'slsall answered at once. The README then attributed that check to the second prompt, the one left open for 11 minutes. During that one, the supervisor log shows/datainaccessible inside the container at 15:16:20, the container stopping at 15:18:42,podman run -dhanging for 180 s from 15:21:20, and the container starting at 15:26:47, the same second the prompt was answered.The tccd log shows the mechanism. Right after the canary's RESULT (
17928.161), tccd evaluated17928.166, a request from vfkit with the stably-signedgtimeoutas responsible process, and allowed it immediately. That request had been queued behind the prompt: every one of these requests comes throughsandboxd(pid 17928, the first half of the msgID), which sends them one at a time. So one unanswered prompt stalls every later network-volume check, including for binaries that already have a grant. That fits the 09-17 outage, where FileBot, Transmission and Plex all stalled together.The README now has those events in the timeline, states the blocking behaviour and the evidence, explains why the spot check proved nothing, and warns that a test prompt takes Transmission down while it is open. Docs only; nothing to deploy.
Advances #199.