Commit Graph
11352 Commits
Author SHA1 Message Date
Luke Kosewski 261cc2f1a8 wgengine/magicsock: fix statticcheck failure
Signed-off-by: Luke Kosewski <lkosewsk@tailscale.com>
2026-09-21 15:48:38 -07:00
Avery PennarunandClaude Opus 4.6 7ce785ef45 ipn/ipnlocal: add TS_FORCE_CACHE_NETMAP envknob to force netmap caching
Allow clients to force netmap caching via the TS_FORCE_CACHE_NETMAP
environment variable, bypassing the requirement for the control server
to grant the cache-network-maps node capability.

This is useful for tsnet users who want faster startup times via
cached netmaps but whose control plane doesn't grant the capability.

Change-Id: I447b57311940d51a5ce9021236add48e0333d0b5
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-21 13:56:23 -07:00
Claus Lensbøl 3f3e56f412 wgengine/netstack: avoid returning from sender goroutines on error (#21332)
Returning from injectToHost and injectToWireGuard leaves the two
goroutines dead and the host without having a way to send traffic.

This is especially relevant for android where the tundev is torn down
and recreated for every call to updateTUN() from a route table change or
roaming between networks.

Fixes #21155

Signed-off-by: Claus Lensbøl <claus@tailscale.com>
2026-09-21 13:56:52 -04:00
Mike Jensen bb94defdd0 fuzz: improve fuzz testing and seed corups (#21340)
This change improves the initial fuzz seeds to get better coverage. It also includes a fix to the geo fuzzing to avoid a harness induced failure when a NaN input is provided.

No actual code logic changes, test only.

Updates tailscale/corp#46608

Change-Id: I0985b6d3a75603927451eea0a25f4cd72c039795

Signed-off-by: Mike Jensen <mikej@tailscale.com>
2026-09-21 09:23:51 -06:00
BeckyPauley 7bf76690f0 cmd/k8s-operator: recover missing egress EndpointSlices (#21174)
* cmd/k8s-operator: move egress EndpointSlice write back into gated provision

PR #20347 moved the EndpointSlice createOrUpdate outside of provision,
causing it to run on every reconcile. This resulted in racing egress-eps on
the EndpointSlice, sometimes causing the TailscaleEgressSvcConfigured to
become stuck as False with the Service not fully updated. Gate it again so
it only runs when a reprovision is required.

Updates #20916

Signed-off-by: Becky Pauley <becky@tailscale.com>

* cmd/k8s-operator: recover missing egress EndpointSlices

Add a watch for EndpointSlices in the egress-services reconciler so a
deleted slice re-triggers a Service reconcile directly. Treat a Service
whose expected per-family EndpointSlice is missing as not up to date so it
re-enters provision and recreates the slice.

Also sort endpoints by Pod UID before writing them in the egress-eps
reconciler, so an unchanged set of ready Pods cannot result in a different
order and trigger an unnecessary Update.

Updates #20916

Signed-off-by: Becky Pauley <becky@tailscale.com>

---------

Signed-off-by: Becky Pauley <becky@tailscale.com>
2026-09-21 11:21:57 +01:00
Brad Fitzpatrick 3014ad828e tstest/natlab/{vmtest,vnet}: bound "tailscale up", dump tailscaled logs on failure
When a natlab node's "tailscale up" hung (TestEasyEasy on CI, twice on
main), the only bound was the 10 minute test context, so the failure
arrived as go test's timeout panic. The VM console logs weren't dumped
(t.Cleanup doesn't run on a panic), and they wouldn't have helped
anyway: on gokrazy the console holds only kernel and init output, while
tailscaled's stdout/stderr goes to a remote syslog that vnet discards
unless the node has VerboseSyslog set. There was no way to see what the
stuck node was doing.

Bound each node's "tailscale up" in Env.Start to 90 seconds (it takes
about a second against the in-process control server), so a stuck node
becomes a normal test failure. On failure, dump the tail of each node's
tailscaled logs from vnet's fake log.tailscale.com log catcher, which
already buffered them per node but exposed them to nothing; add
Server.NodeLogs for that. Also add VMTEST_VERBOSE_SYSLOG=1 to stream
the guests' syslog into the test output live, the vmtest equivalent of
tstest/integration/nat's --log-tailscaled flag.

With this, the CI hang reproduced locally under CPU pressure (2 CPUs
shared with busy loops) in 1 of 22 runs, and the dumped logs showed
tailscaled's control client stuck at "awaiting unpause", which is fixed
separately.

Updates #deflake

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Icfd3a2cdc66d924388a722f3a3cd19f86e116921
2026-09-20 11:02:04 -07:00
Brad Fitzpatrick 7d96cf5a62 tsweb/varz: read memstats via runtime/metrics, export fork metrics
The varz handler got its memstats_* metrics from the expvar package's
"memstats" func, which calls runtime.ReadMemStats and so stopped the
world on every Prometheus scrape. Keep the names but compute them from
runtime/metrics, and never call that func, even from
WritePrometheusExpvar.

While there, export the /tailscale/ metrics from our Go fork (stack
size histogram, stack copy counters, timer zombie counts and lifetime
histogram), which nothing could see before, plus a few upstream ones
with no MemStats equivalent: scheduling latency and GC pause
histograms, live heap, GC and total CPU seconds, thread count, and
mutex wait time. The last replaces derper's hand-rolled version.

Everything read is cheap and a scrape allocates nothing after the
first. Names use a go_runtime_ namespace rather than go_ so they can't
collide with the Prometheus Go client's collector in promvarz binaries.
The runtime's 162-bucket time histograms are reduced to one bucket per
factor of four from 256ns to 1s.

Updates #21300
Updates tailscale/go#189
Updates golang/go#75935

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7e3c9a41f2b85d6e0c4a9b1d3f8e7c2a5b6d4e19
2026-09-19 18:19:20 -07:00
Dep Updater ae69c6b6c9 go.toolchain.rev: bump Go toolchain
* Go toolchain: https://github.com/tailscale/go/compare/32e8826b089fee8cb0c5c4822b9794ca5004f23a...24ee2fd0610e6c505ee4ec061a81215afe119d1f

Triggered by @bradfitz via the bumpdep workflow.

Updates tailscale/corp#29053

Signed-off-by: Dep Updater <noreply+dep-updater@tailscale.com>
2026-09-19 16:52:15 -07:00
Dep Updater 9617640d20 go.mod: bump github.com/bradfitz/go-tool-cache
* github.com/bradfitz/go-tool-cache: v0.0.0-20260909201542-a1c7321be47b to v0.0.0-20260919185303-c660171c910c

Triggered by @bradfitz via the bumpdep workflow.

Updates tailscale/corp#47471

Signed-off-by: Dep Updater <noreply+dep-updater@tailscale.com>
2026-09-19 12:16:27 -07:00
Brad Fitzpatrick 1c778640ea net/netmon, ipn/ipnlocal, wgengine: don't lose a network change during startup
If the network comes up during the few milliseconds NewLocalBackend
takes, tailscaled starts with its control client paused and never
unpauses: the interface state snapshot was taken before the netmon
subscription (kept late since #17252), so a change published in between
reached nobody. #21261 hit it on fast-booting NixOS microVMs, and it was
behind the natlab TestEasyEasy CI flake, where a gokrazy node's DHCP
lease landed in that window and "tailscale up" hung at "awaiting
unpause".

Re-read netmon's state after subscribing rather than subscribing before
the snapshot (as #21281 proposed), which would bring back the #17252
data race and let a fresh delta be overwritten by the stale snapshot.
netmon.New and the userspace engine had the same shape of gap between
their snapshot and their subscription; close those too.

Under CPU pressure the hang reproduced in 1 of 22 TestEasyEasy runs
before and 0 of 120 after.

Thanks to @c-vigo for the detailed diagnosis in #21261 and to @zzz-yu
for pointing out this problem and proposing a fix in #21281!

Fixes #21261
Updates #14902
Updates #19126
Updates #deflake

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ifa21f8055a85321a5afda7800140fc5045c6ce08
2026-09-19 10:46:50 -07:00
Brad Fitzpatrick 574ef3b2a8 gokrazy, .github/workflows: build the natlab image with tool/go
The natlab-basic workflow builds the gokrazy natlab image in its own
step so that the test's rebuild of it (vmtest always rebuilds, so the
baked-in binaries match the source under test) is a build cache hit
rather than a cold build inside go test's -timeout budget. That never
worked: the Makefile ran whatever "go" was on $PATH, the runner's stock
Go, while the go command puts its own $GOROOT/bin first on the test
binary's $PATH, so the rebuild from inside "go test" used tailscale/go.
GOCACHE entries embed the compiler's build ID, so the step warmed
nothing. In practice the in-test rebuild took about 2.5 minutes of the
3 minute -timeout, leaving TestEasyEasy about 20 seconds for booting
two VMs, logging in, and pinging. A passing run on main took 167s. Any
hiccup in the remaining budget, such as the "tailscale up" hang fixed
separately, ended in go test's timeout panic with no useful output.

Make the natlab targets in gokrazy/Makefile use ../tool/go so the step
and the test use the same toolchain and cache. Fix the same mistake in
natlab-test.yml's cache warming step, whose comment documented the
wrong belief about which toolchain the in-test builds use. Raise
natlab-basic's -timeout to match natlab-test.yml so that a hang fails
through vmtest's own bounded waits (which dump the node's logs) instead
of through go test's timeout panic (which dumps nothing about the VMs).

Updates #13038
Updates #deflake

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Iaa6085ec5aa029373204baf75b169ff375c2b355
2026-09-19 10:45:55 -07:00
Brad Fitzpatrick 51682d245b derp/derpserver: drop per-client writer goroutine, start on demand
Each client connection ran two goroutines for its lifetime: the reader
in sclient.run and a sendLoop blocked in a select over its send
queues, pong, peer gone, mesh update, and keepalive channels. Almost
all clients are idle at any moment, so the second goroutine mostly
pinned memory: a 4 KiB stack, a g struct, a sudog per select case,
three channels, and a context and errgroup. At 100k idle connections
that was about 8 KB of a client's 22 KB RSS.

Instead, start up the sendLoop only as needed, letting the goroutine
go away otherwise, like Go 1.28-dev's http2 code
(golang/go@5c51011e82) with similar parking to
https://go.dev/cl/834084 but DERP's producers are all non-blocking, so
a kick bit replaces that http2 code's send count.

Measured with 100k idle TLS connections, server RSS per client went
from 22.3 KB to 14.5 KB (22.9 KB to 15.3 KB after each connection had
carried a packet), goroutines dropped from 2 to 1, and with no change
in BenchmarkSendRecv throughput or allocations and no change in the
time to do 100k serial round trips, each of which parks and wakes the
writer. (it's super cheap to start goroutines)

Updates #21064

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7c3e9a51d4b8f2607a1e5c3d9f8b2a4e6c0d1f3b
2026-09-18 19:19:44 -07:00
Brad Fitzpatrick 630e704763 derp/derpserver: make unique sender cardinality tracking opt-in
Each connected client kept a HyperLogLog sketch of the peers that had
sent it packets, and every relayed packet was inserted into it under a
mutex. That costs memory per client and time on the packet path for a
debug-only estimate that few servers look at.

Keep the accounting but only allocate the sketch when the new
TS_DERP_SENDER_CARDINALITY environment variable is set. When it is
unset, EstimatedUniqueSenders reports 0 and the debug traffic page
omits the field as before.

Updates tailscale/corp#24681

Change-Id: I471bf816e5461069d50b97696e8d91ec3cbe3389
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-18 17:23:22 -07:00
Francois Marier b7a3b1b293 net/portmapper: fix potential race around pcpNonce
All reads from c.mapping need to take place while the lock is held.

Updates #21127

Change-Id: I4782d5ca027584a9e47dc6241af1970348b18c0a
Signed-off-by: Francois Marier <francois@tailscale.com>
2026-09-18 15:01:05 -07:00
Francois Marier 2578d08e8d net/portmapper/pcp: use all-zero address for PCP releases
The PCP spec says that when deleting a mapping, both the external
port and the external address must be zero. While PCP servers are
likely to be lenient in practice, we should do the right thing in
case we encounter strict validation.

Resolves #21363

Change-Id: Ice3467d735def5144a810c39e65642988a830972
Signed-off-by: Francois Marier <francois@tailscale.com>
2026-09-18 10:11:53 -07:00
Fran Bull a9bb6d190b tailcfg: add nodecap conn25-connector-apps
Which will be set to a slice of app names a peer is a connector for.

Updates tailscale/corp#47251

Signed-off-by: Fran Bull <fran@tailscale.com>
2026-09-18 09:13:46 -07:00
Brad Fitzpatrick f87a1b1a82 derp/derpserver: pool received packet payload buffers
Every packet the server relayed allocated a fresh []byte for its
payload in recvPacket or recvForwardPacket and dropped it once the
destination's sendLoop had written it. On one busy server, this was
observed allocating about 160 MB/sec of short-lived garbage, and GC
plus malloc were about 5% of the process CPU profile.

Instead, take payload buffers from a size-classed sync.Pool on the
Server, with power-of-two classes from 1 KiB up to derp.MaxPacketSize,
and return them once the packet has been written, forwarded, or
dropped. sync.Pool holds nothing per connection and is trimmed by the
GC, so idle clients pin no memory; only packets actually in flight
hold a buffer. A compile-time assertion ties the largest size class to
derp.MaxPacketSize, and the get and put helpers panic on sizes outside
the pool's classes rather than indexing past it.

Because the memory is now reused, PacketForwarder implementations must
not retain the payload after ForwardPacket returns. Make that explicit
in the signature: the payload is passed as a new derp.LoanedBytes
value, which exposes only Len, WriteTo, and Clone, so an implementation
has to copy to keep it. derp.Client and derphttp.Client, the real
implementations, already wrote it out synchronously; the test-only
channelFwd now clones.

BenchmarkSendRecv shows one fewer allocation per relayed packet and,
for 1000-byte packets, B/op down from 1278 to 263. ns/op on the
loopback benchmarks is dominated by syscalls and is unchanged within
noise.

Updates #21064

Change-Id: Ie40c82388ddb5d22f75fa828749b53fcaba9adde
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-18 08:55:48 -07:00
Adrian Dewhurst c2dc086468 util/multierr: mark multierr.New as deprecated
We banned use of multierr in various dep tests, so this makes the
situation more obvious if someone stumbles across it.

Change-Id: I17e80880e57e5005fcaea864a365ba264411e802
Signed-off-by: Adrian Dewhurst <adrian@tailscale.com>
2026-09-18 11:46:09 -04:00
Brad Fitzpatrick 47debc5a6f derp/derpserver: don't build debug log arguments on the packet path
sclient.debugLogf and Server.debugLogf check a debug flag before
logging, but Go evaluates and boxes their arguments before the call.
The per-packet call sites in run, handleFrameSendPacket,
handleFrameForwardPacket, sendPkt, recordDrop, and sendPacket's
deferred stats func were therefore calling key.NodePublic.ShortString
and boxing frame headers on every relayed packet, all for messages
that were then discarded.

On one busy server's heap profile, those discarded arguments were
about half of all objects allocated by the process. Guard each hot
call site with the debug flag so nothing is built unless it will be
logged, and document that requirement on both debugLogf methods.

While here, give the sender cardinality sketch its key bytes from a
stack array rather than an AppendTo(nil) allocation per packet.

BenchmarkSendRecv drops from 10 or 11 allocations per relayed packet
to 3, and BenchmarkConcurrentStreams from 11 to 4.

Updates #21064

Change-Id: Ibb7c4fbff546b6f1b40ce0a21d41ae0976705941
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-18 06:52:51 -07:00
Brad Fitzpatrick 1033e714ca cmd/testwrapper: run all package patterns in one go test invocation
testwrapper ran a separate, sequential "go test" invocation for each
package pattern on its command line. That is fine for a single "./..."
argument but not for callers that pass an explicit package list: CI
jobs in the corp repo passing ~200 packages ran ~200 serial go test
processes with no cross-package parallelism and a fixed set of
never-cacheable lookups per process, and spent several times longer
on process startup, package loading, cache lookups, and serial test
binary links than on running tests. See tailscale/corp#48453 for the
details.

Locally, on 203 packages with a fully warm build and test cache, so
measuring only the per-invocation overhead:

  old (203 go test processes):  26.4s
  new (1 go test process):       3.6s  (7.3x faster)

Our own Windows CI job hits the same path: its "sharded:N/M" mode
expands to an explicit list of that shard's packages via listpkgs, so
each shard ran one go test process per package, and Windows process
startup is slower still. Each shard now runs as one invocation.

Fixes tailscale/corp#48453
Updates tailscale/corp#47035

Change-Id: I3f796ff1724af40f93be9f918a7ddfde3bb45a91
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-18 02:56:29 -07:00
Brad Fitzpatrick 3323dc02f4 tsweb: restore AcceptsEncoding
Commit 6608b9a38 removed tsweb.AcceptsEncoding, saying it had no
callers, but it has many callers in the tailscale.io repo. Restore the
function and its test unchanged so those builds work again.

Updates #12170
Updates tailscale/corp#48447 (broken by this)

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7c3e2a9f5b1d4e8a6c0f2b3d9e1a7c5f4b8d2e6a
2026-09-17 21:36:46 -07:00
OSS Updater 28836381da go.mod: update web-client-prebuilt module
Signed-off-by: OSS Updater <noreply+oss-updater@tailscale.com>
2026-09-17 17:55:09 -07:00
yaruk-byte be89526457 tstest/integration: stop skipping Windows integration tests (#21187)
Updates #20750

Signed-off-by: Yaruk Asghar <yaruk@tailscale.com>
2026-09-17 17:11:00 -07:00
James Tucker aa1134d358 go.mod: update golangci-lint to v2
Moving to v2 because v1 references repositories that have been
deleted from GitHub, breaking GOPROXY=direct.

The lint config (.golangci.yml) was already in v2 format; this updates
the tool dependency used by 'make lint' to golangci-lint/v2 v2.13.2,
drops the now-obsolete blank import in internal/tooldeps in favor of a
Go 'tool' directive, and bumps the CI workflow binary to match.

Updates #cleanup

Signed-off-by: James Tucker <jftucker@gmail.com>
2026-09-17 16:55:23 -07:00
James Tucker 6608b9a387 tsweb/compserve, client/web: add zstd, remove brotli for precompressed assets
Adds tsweb/compserve: content negotiation for precompressed static
variants, with a transcode-to-identity fallback for clients that do not
accept an encoding (including when the raw file is absent), and
CompressWriter, which live-compresses dynamic responses with zstd in its
fastest mode, streamed incrementally with no buffering. Negotiation is
q-value and wildcard aware (gzip;q=0 previously matched gzip).

client/web serves its prebuilt embedded assets through compserve,
replacing brotli with zstd; the embedded FS is wrapped in
tsweb/vcstime for conditional-request mod times. tsweb/compserve/gzip.go
keeps transitional serving of gzip variants from pre-zstd file systems
(such as the currently published web-client-prebuilt module):
passthrough to gzip-accepting clients, transcoded to identity otherwise;
it becomes inert once a zstd-only module is published.

util/zstdframe gains pooled GetDecoder and GetStreamingEncoder
(concurrency=1). util/precompress is now a build-time tool, generating
zstd variants only. cmd/tsconnect and cmd/build-webclient consume the
new precompress/compserve split. tsweb.AcceptsEncoding and
tsweb/tswebutil are removed; negotiation lives in compserve and the
deprecated shim had no callers. go.mod bumps web-client-prebuilt.

Also fixes a transcoding bug where http.ServeContent's size probe via
the promoted zstd.Decoder.WriteTo could report a zero length, serving
empty bodies.

Updates tailscale/corp#20099

Signed-off-by: James Tucker <james@tailscale.com>
2026-09-17 15:23:49 -07:00
Dep Updater 178ef3db08 go.toolchain.rev: bump Go toolchain
* Go toolchain: https://github.com/tailscale/go/compare/d030173bb47a6c4a6f885cb56a97dd9eca5fb8b7...32e8826b089fee8cb0c5c4822b9794ca5004f23a

Triggered by @bradfitz via the bumpdep workflow.

Updates tailscale/go#189

Signed-off-by: Dep Updater <noreply+dep-updater@tailscale.com>
2026-09-17 15:03:04 -07:00
Patrick O'Doherty 2e72593fbf feature/clientupdate: require write access for update/install localapi (#21360)
The update/install localapi handler had no PermitWrite check, so any
local user who could reach the localapi socket could make the root
daemon self-update and restart itself. This has been the case since the
endpoint was added in November 2023.

Gate the handler behind PermitWrite so that only root or the operator
user can trigger a self-update, matching the other mutating handlers.

Updates tailscale/corp#48187

Change-Id: Iadfef939f5dab684652cd220e77de63b54bfca2f
Reported-by: Ben Carman <benthecarman@live.com>

Signed-off-by: Patrick O'Doherty <patrick@tailscale.com>
2026-09-17 14:30:40 -07:00
Naman Sood 027e249fcf net/socks5: correctly proxy half-closed TCP connections
Similar to #16462, when we are acting as a TCP proxy, we need to pass
through half-closes correctly since clients and servers will sometimes
close one direction of the connection and still rely on the other
direction working.

Fixes #20883.

Signed-off-by: Naman Sood <mail@nsood.in>
2026-09-17 17:22:32 -04:00
Naman Sood 35848427ce types/nettype: add HalfCloser type
We have multiple situations where we have a `net.Conn` representing a
TCP connection for a proxy and we need access to the underlying
`CloseWrite()` and `CloseRead()` functions to properly pass through
half-closes (see #16462, #20883). Add an interface we can cast to in
order to get access to these functions.

Updates #20883.

Signed-off-by: Naman Sood <mail@nsood.in>
2026-09-17 17:22:32 -04:00
Mike Shaver 16d19a7e5e cmd/tsidp: remove all but the warning from the in-tree tsidp code's README (#21354)
Updates tailscale/tsidp#185

Change-Id: I28cbb5b96a26b25db4311d793fcc7426bc9da9fa

Signed-off-by: Mike Shaver <shaver@tailscale.com>
2026-09-17 14:10:55 -04:00
yaruk-byte 1c7e77f094 tstest/integration: wait longer to remove the staged binaries at teardown (#21207)
The Windows GitHub runners' provisioning daemon, provjobd.exe, opens a
handle to each freshly-written tailscaled.exe. tb.TempDir's own cleanup
retries for a fixed 2s and gives up, failing whichever test happens to tear
down while the handle is held.

Updates #21099

Signed-off-by: Yaruk Asghar <yaruk@tailscale.com>
2026-09-17 09:44:05 -07:00
Mike Jensen 408504af89 net/packet: fix Transport panic from a bad IPv4 IHL (#21331)
`decode4` assigned `q.subofs` before validating it against the declared IP total length. A packet with an IHL past the end of the buffer was rejected but left `subofs` dangling there, so a later Transport call would panic.

This change only store `subofs` once validated. As defense in depth an additional bounds check is added in Transport.

Fuzzing was expanded and improved to get better coverage in `packet.go`.

Credit to @Dev-next-gen for finding and reporting.

Fixes tailscale/corp#48322
Updates tailscale/corp#46608

Change-Id: I198d06b921add9b188daad79d4ca473aed1e3b66

Signed-off-by: Mike Jensen <mikej@tailscale.com>
2026-09-17 10:00:01 -06:00
Alex Chan b16bc957aa ipn,tsnet: replace LocalBackend.NetMap with NetMapNoPeers/NetMapWithPeers
This method has been deprecated for four months; replace all uses with
the Peers or NoPeers equivalents.

Updates #12542

Change-Id: I7b5800d0d92775839dbfb6021751a5dfacc53d2f
Signed-off-by: Alex Chan <alexc@tailscale.com>
2026-09-17 13:43:14 +01:00
Alex Chan ade2dc44b5 tka: add missing test case for "init with an untrusted key"
Updates tailscale/corp#45574

Change-Id: I1ab724b9b4629cd4a5b4bc71bf7417d58e49ee1d
Signed-off-by: Alex Chan <alexc@tailscale.com>
2026-09-17 13:29:16 +01:00
Brad Fitzpatrick 62f3130637 tstest/integration: deflake TestNATPing
TestNATPing called SetMasqueradeAddresses on the test control server
and then immediately read both nodes' status, expecting the new
masqueraded peer addresses to already be there. But the change reaches
the nodes asynchronously via their streaming map responses, so under
load the status check ran before the new map response arrived and the
test failed with "n1 sees n2 as 100.64.0.2; want 100.64.2.1" and the
like. This was the dominant failure mode on the flakes dashboard (11 of
the 18 most recent CI failures) and the only one found in a six hour
Antithesis run (run 30b4de27d89d8d7a651e7b43b3f1ec3f-61-9, 54 failures
in 16,599 runs).

Wait for each node's status to report the expected peer address
instead. Also retry the "tailscale ping" invocations, since a ping can
fail transiently right after a map response changes a peer's addresses
and before the engine is reconfigured; the second most common failure
mode was pings exiting with status 1. Failed pings now include the CLI
output in the error rather than a bare exit status.

Locally, flakestress (32 workers) reproduced the failure in 29 of 210
runs before this change (13.8%) and in 0 of 1,173 runs after.

The remaining Windows-only failure mode on the dashboard, TempDir
cleanup failing because tailscaled.exe is still open, is a
harness-wide issue that affects every integration test and is not
specific to this test.

Fixes #12169
Updates tailscale/corp#47865

Change-Id: I7c3e2b8a5d914f0e6a2b1c9d8e7f6a5b4c3d2e1f
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-16 16:54:02 -07:00
Brad Fitzpatrick b0b1f0f566 go.mod: bump all direct deps to latest
This is the output of the new misc/bumpdeps tool (#21325) run with
--exclude-newer-than-days=7, which asks proxy.golang.org for the newest
version of every direct dependency, ignoring releases younger than a
week in favor of the newest older one, and runs a single go get.
gvisor tracks its "go" branch, wireguard-go its "tailscale" branch,
and golang-x-crypto its "main" branch (the proxy's @latest for it is
a stray v0.91.0 tag from 2024 that predates our acme fork changes).
Indirect deps only moved as far as MVS pulled them.

The week-long cooldown held back gvisor, the gokrazy modules,
chromedp/cdproto, and hashicorp/raft-boltdb/v2, whose only newer
versions are days old; they'll come along next time.

Several upstream changes needed small fixes: nfpm's PrepareForPackager
takes a modification time now (a zero time keeps the old behavior of
using the source file's mtime), esbuild's ServeOptions.Port became an
int while ServeResult.Host became a Hosts slice, client-go's
EventRecorder.Eventf is now recognized by vet as a printf wrapper (so
the k8s-operator calls that passed a preformatted message switch to
Event), google/nftables v0.3.0 reads back the kernel's
NF_NAT_RANGE_PROTO_SPECIFIED flag into a new expr.NAT.Specified field
(so the port map DNAT rule now sets it too or findRule never matches
the rule it just added), and staticcheck v0.8.1 knows encoding/json/v2's
embed tag option, so the two SA5008 suppressions for it are gone.

Two tests assumed old library behavior. client-go's fake clientset now
replays existing objects when a watch starts, as a real apiserver does,
so the k8s-proxy config test must tolerate the loader ignoring that
no-op event before the real reload arrives. fyne.io/systray moved its
dbusmenu object path and answers the first GetLayout with depth 1, so
the systray test now finds the menu via the item's Menu property and
polls until the submenu entries appear.

Then make tidy, make updatedeps, and make kube-generate-all (the
controller-gen bump to v0.22.0 changes doc strings, stops listing
top-level metadata as required, and crd-ref-docs now marks optional
fields).

Updates #8043

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I3f9a2c6e8b1d4705a9e2c7b8d1f4e6a0c2b5d8e3
2026-09-16 16:01:50 -07:00
James Tucker 4f88a7ba46 tsweb/vcstime: add a VCS-time-stamped file system wrapper for embed.FS
embed.FS stores and serves no ModTime: files report a zero ModTime,
which disables HTTP caching semantics when served with net/http: no
Last-Modified header is sent and If-Modified-Since requests are never
answered with a 304.

vcstime.FS wraps an fs.FS (typically an embed.FS) and reports the
vcs.time commit time from the binary's build info as the modification
time of files that have no real timestamp of their own, enabling
http.FS to provide working cache headers for clean builds. Files that
already have a real timestamp are passed through untouched, and
binaries built without VCS stamping degrade to the file system's own
behavior.

Updates tailscale/corp#48172

Signed-off-by: James Tucker <james@tailscale.com>
2026-09-16 14:44:08 -07:00
Brad Fitzpatrick 46332efe57 .github/workflows: use the module proxy for generic bumpdep fetches
The bumpdep workflow ran every go get with GOPROXY=direct, copying the
corp update-oss workflow. For a module like github.com/gokrazy/kernel.amd64
that means cloning a repo full of kernel image blobs, which hung the
first real run of the workflow indefinitely.

Fetch generic modules through the default GOPROXY (proxy.golang.org),
which serves branch names like @main just fine. Keep GOPROXY=direct only
for wireguard-go, which is our own small repo where seeing a just-pushed
commit matters more than the proxy's cache lag. Also give the job a
45 minute timeout so a future hang doesn't hold a runner for six hours.

Updates tailscale/corp#48312

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I3e7d1a9c5b2f48e06d7a1c9e4b3f2a8d6c5e7f10
2026-09-16 13:50:50 -07:00
Brad Fitzpatrick d1c5827ea2 .github/workflows: add bumpdep workflow to bump Go deps and open a PR
This is the OSS analog of the corp repo's update-oss workflow. It's
manually dispatched with a comma-separated list of dependencies and the
URL of the issue motivating the bump, updates each dependency, runs
"make tidy" and "make updatedeps", and opens a pull request from the
tailscale-code-updater app assigned to whoever triggered it.

Three names are special-cased: "go" runs ./pull-toolchain.sh,
"wireguard-go" tracks github.com/tailscale/wireguard-go@tailscale, and
"gvisor" tracks gvisor.dev/gvisor@go. Anything else is treated as a Go
module path and bumped to @latest (or to an explicit path@version). The
issue URL becomes the "Updates" line of the generated commit message.

Updates tailscale/corp#48312

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I6b2f9c1e8d4a7350b2e9f1c4d8a6e2b7f3c5d9a1
2026-09-16 12:58:41 -07:00
Brad Fitzpatrick fa7fb388b5 misc/bumpdeps: add tool to bump all go.mod deps to latest
Bumping deps one "go get foo@latest" at a time is slow with the ~500
modules in go.mod. This tool parses go.mod with x/mod/modfile, asks
proxy.golang.org concurrently for the newest version of each direct
dependency (about a second for the whole file), and then runs a single
"go get" with the modules that actually have something newer. Arguments
narrow it to modules whose path contains one of the given substrings,
so "bumpdeps gvisor" does the obvious thing.

gvisor.dev/gvisor, github.com/tailscale/wireguard-go, and
github.com/tailscale/golang-x-crypto follow branches ("go", "tailscale",
and "main" respectively) rather than tags, so those are resolved via the
proxy's @v/<branch>.info endpoint instead of @latest.

The --exclude-newer-than-days flag is a cooldown in the sense of
https://nesbitt.io/2026/03/04/package-managers-need-to-cool-down.html:
releases younger than that are ignored in favor of the newest one old
enough, so a compromised upstream has to go unnoticed that long before
we'd pick it up. It defaults to off. Branch-tracked modules have nothing
older to fall back to, so they're held until their head has aged.

The tool never downgrades, either by semver (the proxy's @latest can be
older than a pseudo-version we're already on) or by commit time (forks
can carry stray tags that sort above their real development branch, as
golang-x-crypto's v0.91.0 from 2024 does). It skips replaced modules and
modules whose latest release declares a different module path, since go
get rejects those and they need their import paths changed by hand.
Indirect deps are left to MVS by default; -indirect bumps them too, but
that tends to break the build when their importers haven't caught up.

Updates #8043

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7c2e9a41d5f0b83e6a1c4d9f2b7e8a0c3d5f6b1e
2026-09-16 11:30:55 -07:00
Andrew Dunham 158a2476c0 tailcfg: add Node.StableTailnetID and bump capver (#21314)
Add Node.StableTailnetID for control to send the current tailnet's
stable ID on the self node. Expose it as CurrentTailnet.StableID in
LocalAPI status and `tailscale status --json`, and display it in
`tailscale whoami`.

Bump CurrentCapabilityVersion to 148.

Updates #14375

RELNOTE: Show the current tailnet's stable ID in status JSON and whoami.

Change-Id: I0525ff735de8113c8d124045a94e00c19a5a02e2

Signed-off-by: Andrew Dunham <andrew@tailscale.com>
2026-09-16 14:28:46 -04:00
Francois Marier 8b3b8f122e net/portmapper: cleanup old PCP TODOs
Requesting a UDP mapping is the right thing to do since an "all
protocols" (and "all ports") mapping would be akin to requesting
to be a DMZ on that network.

Updates #cleanup

Change-Id: Icc83d4fe14dbec0eb844b093d2d92756d6c05228
Signed-off-by: Francois Marier <francois@tailscale.com>
2026-09-16 11:01:34 -07:00
Brad Fitzpatrick 96492df022 ipn/ipnlocal: log the profile's login name, not the method value
The profile switch log lines passed cp.UserProfile().LoginName without
calling it, so they printed "%!q(func() string=0x7ff7141cd060)" in place
of the login name.

Updates #cleanup

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I4e6a8c0b2d5f7a9c1e3b5d7f9a1c3e5b7d9f1a3c
2026-09-16 09:05:27 -07:00
Brad Fitzpatrick e22eaaebfd control/tsp: don't use bootstrap DNS from the noise client
tsp is a protocol library and should not care about the system's routing
table, interfaces, or how to reach DERP for bootstrap DNS. Those are
tailscaled concerns. Without an explicit DNSCache, ts2021 built a resolver
whose LookupIPFallback did bootstrap DNS over DERP whenever the first dial
to control failed, and that path also crashed on tsp's nil netmon.Monitor.

Provide a plain dnscache.Resolver with no fallback so a failed dial is
just a failed dial.

Updates tailscale/corp#47865

Change-Id: If62ddb90139d5590ca3e0e3f54235bbbc0d8aec2
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-16 07:29:31 -07:00
Brad Fitzpatrick 41a756c1fe net/dnsfallback: don't panic on nil netmon.Monitor in bootstrap DNS path
MakeLookupFunc documents its netMon parameter as optional, but
bootstrapDNSMap passed it straight to netns.NewDialer, which has panicked
on nil since 3672f29a4. Any caller that omitted the monitor and then had a
first dial fail (which is when dnscache consults LookupIPFallback) crashed
with "netns.NewDialer called with nil netMon". control/tsp clients hit this
under Antithesis fault injection.

Use a plain net.Dialer when no monitor is provided and add a regression
test.

Updates tailscale/corp#47865

Change-Id: Ie8255033602dd1f8243c1ced8f453fb7ab9be74d
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-16 07:29:31 -07:00
Michael Ben-Ami 407d32ff42 appc: detect cycles in CNAME chain
Fixes tailscale/corp#48296
Updates tailscale/corp#48187

Signed-off-by: Michael Ben-Ami <mzb@tailscale.com>
2026-09-16 10:08:11 -04:00
maxiscoding28 044abcc9d2 ipn/ipnlocal: make Serve idle connection limit configurable (#21179)
* ipn/ipnlocal: make Serve idle connection limit configurable

Allow Kubernetes operator proxy pods to override the per-host idle connection limit through an environment knob.\n\nUpdates tailscale/tailscale#20875

Signed-off-by: maxiscoding28 <max.winslow@icloud.com>

* ipn/ipnlocal: document Serve idle connection default

Signed-off-by: maxiscoding28 <max.winslow@icloud.com>

---------

Signed-off-by: maxiscoding28 <max.winslow@icloud.com>
2026-09-16 11:01:27 +01:00
Thomas Desrosiers 7add2af9ec prober: cache CRLs across TLS probes (#21286)
The TLS probe fetched and parsed the leaf certificate's CRL on every
run. That is fine when the CRL is small, but some CAs publish CRLs of
several megabytes: the one for the AWS ACM R2M04 intermediate is about
2.5MB, which at the default 15s interval is a continuous 170kB/s per
probed node.

Cache parsed CRLs by distribution point URL and reuse each for up to
an hour, or until its NextUpdate if that comes first. An hour is the
HTTP max-age Let's Encrypt serves on its root CRL. A CRL is cached
only after its signature verifies, every use still re-verifies it
against the probing leaf's issuer (the cache is keyed by URL alone),
and a CRL without a NextUpdate is never cached since it declares no
validity window.

Concurrent probes fetch through singleflight.DoChanContext to avoid
re-fetching a single CRL, with each waiter keeping its own deadline. A
caller that missed the cache re-checks it inside the singleflight
closure, since singleflight dedupes only calls that overlap.

Leaf certificates whose issuer is missing from the presented chain now
fail before any fetch. Previously the probe downloaded the CRL and then
panicked in CheckSignatureFrom, which the prober recovered and recorded
as a probe failure.

Also update the TLS probe's doc comments, which said OCSP where the
code checks a CRL.

Fixes #21310

Signed-off-by: Thomas Desrosiers <git@hive.pw>
2026-09-16 00:08:19 -04:00
Brad Fitzpatrick 678ad167e6 go.toolchain.rev: bump tailscale/go again
For https://github.com/tailscale/go/pull/188

Updates tailscale/corp#29053

Change-Id: I3d42156b3c2ef824b68033031d9be48fe7989176
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-15 20:54:53 -07:00
Brad Fitzpatrick a2263542f2 cmd/testcontrol: add --addr and --ssh-policy flags
The test control server always listened on 127.0.0.1:9911, which is
useless for a node in a VM on the same machine. --addr picks the listen
address; the DERP and STUN servers follow it.

--ssh-policy loads a tailcfg.SSHPolicy from a JSON or HuJSON file and
sends it to every node, which also grants them the SSH node capability so
that "tailscale up --ssh" is accepted. That is what testcontrol.Server
already supported for tests; this exposes it for manual testing of the
Tailscale SSH server.

Updates tailscale/corp#47865

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I2b7c4e9f1a3d5c6e8b0f2a4d6c8e1b3f5a7d9c0e
2026-09-15 20:03:21 -07:00