mirror of
https://github.com/tailscale/tailscale.git
synced 2026-10-09 11:52:00 -04:00
If the network comes up during the few milliseconds NewLocalBackend takes, tailscaled starts with its control client paused and never unpauses: the interface state snapshot was taken before the netmon subscription (kept late since #17252), so a change published in between reached nobody. #21261 hit it on fast-booting NixOS microVMs, and it was behind the natlab TestEasyEasy CI flake, where a gokrazy node's DHCP lease landed in that window and "tailscale up" hung at "awaiting unpause". Re-read netmon's state after subscribing rather than subscribing before the snapshot (as #21281 proposed), which would bring back the #17252 data race and let a fresh delta be overwritten by the stale snapshot. netmon.New and the userspace engine had the same shape of gap between their snapshot and their subscription; close those too. Under CPU pressure the hang reproduced in 1 of 22 TestEasyEasy runs before and 0 of 120 after. Thanks to @c-vigo for the detailed diagnosis in #21261 and to @zzz-yu for pointing out this problem and proposing a fix in #21281! Fixes #21261 Updates #14902 Updates #19126 Updates #deflake Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: Ifa21f8055a85321a5afda7800140fc5045c6ce08