mirror of
https://github.com/tailscale/tailscale.git
synced 2026-10-08 11:21:51 -04:00
tstest/natlab/{vmtest,vnet}: bound "tailscale up", dump tailscaled logs on failure
When a natlab node's "tailscale up" hung (TestEasyEasy on CI, twice on main), the only bound was the 10 minute test context, so the failure arrived as go test's timeout panic. The VM console logs weren't dumped (t.Cleanup doesn't run on a panic), and they wouldn't have helped anyway: on gokrazy the console holds only kernel and init output, while tailscaled's stdout/stderr goes to a remote syslog that vnet discards unless the node has VerboseSyslog set. There was no way to see what the stuck node was doing. Bound each node's "tailscale up" in Env.Start to 90 seconds (it takes about a second against the in-process control server), so a stuck node becomes a normal test failure. On failure, dump the tail of each node's tailscaled logs from vnet's fake log.tailscale.com log catcher, which already buffered them per node but exposed them to nothing; add Server.NodeLogs for that. Also add VMTEST_VERBOSE_SYSLOG=1 to stream the guests' syslog into the test output live, the vmtest equivalent of tstest/integration/nat's --log-tailscaled flag. With this, the CI hang reproduced locally under CPU pressure (2 CPUs shared with busy loops) in 1 of 22 runs, and the dumped logs showed tailscaled's control client stuck at "awaiting unpause", which is fixed separately. Updates #deflake Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: Icfd3a2cdc66d924388a722f3a3cd19f86e116921
This commit is contained in:
500 Internal Server Error
Gitea Version: 1.28.0+dev-477-g8b6ad49a5f