Skip to content

tailcat: mark magicsock's network up before the first netcheck - #138

Merged
bradfitz merged 1 commit into
tailscale:mainfrom
sophronesis:fix-network-up-before-netcheck
Sep 28, 2026
Merged

bradfitz merged 1 commit into
tailscale:mainfrom
sophronesis:fix-network-up-before-netcheck

Conversation

@sophronesis

Copy link
Copy Markdown
Contributor

In the Nix build sandbox TestServeExec and TestServeExitNode sometimes fail with tailcat Ping: context deadline exceeded. This showed up in the x86_64-linux nixpkgs-review of the v0.7.0 bump (NixOS/nixpkgs#565145).

With no non-loopback interface, magicsock starts with the network down. locoBackend.Start calls SetPrivateKey and SetDERPMap (both start endpoint updates in the background) and only later calls SetNetworkUp(true). If an endpoint update runs in between, updateNetInfo returns early on networkDown() before maybeSetNearestDERP runs. The server then has no home DERP until the periodic re-STUN 20-26s later, and the client's 10s ping gives up before that. The client log shows derp-1 does not know about peer.

This change moves SetNetworkUp(true) in front of SetPrivateKey. On a host with a usable interface the network is already up, so the call does nothing there.

Repro, with the cmd/tailcat tests in a loopback-only netns and a time.Sleep(300 * time.Millisecond) after mc.SetDERPMap(lb.dm) to widen the window:

unshare -rn sh -c 'ip link set lo up && go test -count=1 -run "^(TestServeExec|TestServeExitNode)$" ./cmd/tailcat'

Without the change both tests time out. With it they pass. Without the sleep, the full cmd/tailcat suite passed 8/8 runs in the netns with the change. go vet ./... and the release-tags build are clean. TestSSHSuite in the root package fails on my machine both with and without this change (EnvForwarding/PTYAllocated), which looks like an environment issue rather than something caused here.

nixpkgs context: the failure was reported on the original v0.7.0 bump, NixOS/nixpkgs#565145. NixOS/nixpkgs#567857 carries this change as a local patch, which can be dropped once it's in a tailcat release.

TestServeExec and TestServeExitNode flaked in the Nix build sandbox
(the x86_64-linux review of the v0.7.0 bump in NixOS/nixpkgs#565145)
with "tailcat Ping: context deadline exceeded".

With no non-loopback interface, wgengine starts magicsock with the
network down. locoBackend.Start calls SetPrivateKey and SetDERPMap,
which start endpoint updates in the background, and only later calls
SetNetworkUp(true). If an endpoint update runs in between,
updateNetInfo returns an empty report because the network is down,
before it gets to maybeSetNearestDERP, so the server never picks a
home DERP. SetNetworkUp(true) only connects to an already chosen home,
so nothing fixes that until the periodic re-STUN 20-26s later, and a
client's 10s Ping gives up first; its relay reports the server as not
connected.

Mark the network up first. When the host has a usable interface the
network is already up and the call is a no-op.

Reproduced by running the cmd/tailcat tests in a loopback-only network
namespace (unshare -rn) with a 300ms sleep after SetDERPMap to widen
the window: TestServeExec and TestServeExitNode time out without this
change and pass with it. Without the sleep, the full cmd/tailcat suite
passed 8 of 8 runs in that namespace with this change.

Signed-off-by: sophronesis <oleksandr.buzynnyi@gmail.com>
@bradfitz
bradfitz merged commit 4fb672e into tailscale:main Sep 28, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants