Commit Graph
11374 Commits
Author SHA1 Message Date
Brad Fitzpatrick e87076a378 ipn/auditlog: store the audit log in the backend's state directory
The store path was hardcoded to %ProgramData%\Tailscale\audit-log.json on
Windows, so a tailscaled run by a regular user with its own --statedir
logged "[unexpected] failed to create audit log store ... Access is denied"
on every profile switch. Use the backend's TailscaleVarRoot when it is
known, falling back to the platform default otherwise. For the Windows
service the state directory is %ProgramData%\Tailscale, so its path is
unchanged; an explicit SetStoreFilePath still wins.

Updates #2791

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I5d7f9b1c3e5a7c9e1b3d5f7a9c1e3b5d7f9a1c3e
2026-09-24 11:52:21 -07:00
Brad Fitzpatrick 6612bed24d ipn/desktop: skip the extension when sessions can't be enumerated
Registering the session state callback enumerates the existing desktop
sessions, which fails for a tailscaled run by an unprivileged user
(WTSEnumerateSessions returns "No more data is available"). That was
reported as "init failed", which reads like a bug. Such a tailscaled has
no other users' sessions to track, so report it as skipping the extension,
the same way an unavailable session manager already is.

Updates #2791

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7f9b1d3e5a7c9e1b3d5f7a9c1e3b5d7f9a1c3e5b
2026-09-24 11:51:51 -07:00
kari-ts 9020fdcd8a logtail: stop retrying uploads after logger is disabled (#21459)
Logger.SetEnabled(false) prevents new log entries from being buffered, but a batch that has already been drained into uploading can continue retrying indefinitely after logging is disabled.  This was seen on a client where log.tailscale.com was blocked and remote client logging had been disabled.
Here we add a check on the disabled state before upload attempts so that existing failed batches are abandoned after logging is disabled.

Updates tailscale/tailscale#21088

Signed-off-by: kari-ts <kari@tailscale.com>
2026-09-24 10:43:53 -07:00
License Updater b1664580dc licenses: update license notices
Signed-off-by: License Updater <noreply+license-updater@tailscale.com>
2026-09-23 18:02:11 -07:00
Patrick O'Doherty 9cdbc3ad3e ssh/tailssh: record stderr on non-PTY sessions (#21369)
On non-PTY exec sessions stdout was wrapped in the recording writer but
stderr was copied to the client raw, so any command output written to
stderr never appeared in the session recording. PTY sessions were
unaffected because they fold stdout and stderr into a single stream that
already flows through the recording writer.

Wrap stderr in rec.writer too, recording it under the "o" (output)
direction, matching how PTY sessions record combined output. The
asciinema cast format has no separate stderr channel, so "o" is the
correct code for players and tsrecorder to render it.

Strengthen TestSSHRecordingNonInteractive to run a command that writes to
both stdout and stderr and assert both appear in the recorded cast lines.

Updates tailscale/corp#48187

Change-Id: I1dc36073722677a85f2b23b331fd55b8df4d342e
Reported-by: Ben Carman <benthecarman@live.com>
Signed-off-by: Patrick O'Doherty <patrick@tailscale.com>
2026-09-23 16:13:52 -07:00
Jordan Whited 8665988efb go.mod,wgengine/wgcfg: add wireguard-go handshake metrics
Updates tailscale/corp#48747

Signed-off-by: Jordan Whited <jordan@tailscale.com>
2026-09-23 15:59:28 -07:00
Mazdak Nasab b6c6c59dde feature/conn25: extend datapath test to verify full packets (#21422)
Extend datapath tests to not only assert headers but
also the full expected packets.

Fixes https://github.com/tailscale/corp/issues/48609

Change-Id: I0e9a1c2f61f2d63eafa430d8e6436908fe927cf4

Signed-off-by: Mazdak Nasab <mazdak.nasab@gmail.com>
2026-09-23 13:52:21 -07:00
Francois Marier 8d43ba6737 tstest/natlab/vmtest: close unused pipe read FD to avoid leak
There are two copies of the read file descriptor (parent & child).
Since the parent copy is unused, we can close it after fork.

Updates #cleanup

Change-Id: Id789e0a04970c9fb8eaeed1f299aa25f29fcf48c
Signed-off-by: Francois Marier <francois@tailscale.com>
2026-09-23 13:04:11 -07:00
François Marier 02de76ba62 tstest/natlab: ignore transport errors during gokrazy update (#21441)
In my testing using TCG, I hit non-EOF network errors that made
this test fail. Since we already poll and check the image version
later, we can treat this response as best-effort.

Resolves #21440

Change-Id: Ia923fc2c199aae7a543d00e06d25a375512daa01
Signed-off-by: Francois Marier <francois@tailscale.com>
2026-09-23 10:50:34 -07:00
Brad Fitzpatrick 0931824b5f feature/androiddns: fall back to getaddrinfo on Android 9 and older
The resnsend dnsproxyd command this package relays raw DNS messages
through was added in Android 10. On Android 9 and older the daemon
answers it with FrameworkListener's text "500 Command not recognized",
which we read as binary: "500 " passed the negative errno check and
"Comm" became the answer length, so every lookup failed with
"androiddns: bogus answer length 1131375981". That's what a tailcat
user hit on a Fire TV stick, which runs Android 9 under Fire OS 7.

Detect that text reply (a binary reply never starts with an ASCII
digit) and switch the process over to the daemon's older getaddrinfo
command, the one bionic's getaddrinfo proxies through. Parse the
single A or AAAA question out of the query, send the command with the
matching address family, and synthesize a DNS answer from the addrinfo
list that comes back, mapping EAI_NONAME to NXDOMAIN and EAI_NODATA to
an empty answer so Go's resolver produces its usual errors. Other
query types return an error saying the Android version can't answer
them.

The reply layout is the field-by-field one netd has used since
Android 6.0 (a 64-bit netd may serve a 32-bit client, so it stopped
sending the raw struct); Android 5.x's raw struct layout is detected
and rejected rather than guessed at. The format was verified against a
32-bit Fire OS 7.7.1.3 device (Android 9, API 28), where a tailcat
build with this change resolves names, fetches its DERP map over
HTTPS, and accepts a connection from another machine.

Updates tailscale/tailcat#126

Change-Id: I7062984c7ec8ce6caf737089e220ef2ca439dcc5
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-23 10:45:09 -07:00
Francois Marier a35aff2895 tstest/natlab: move running instructions to a README
Updates #cleanup

Change-Id: I9580fbfeeef2edf0c678c8ab1ec70131fd20f0b5
Signed-off-by: Francois Marier <francois@tailscale.com>
2026-09-23 10:19:34 -07:00
YewFence 7122baac1b ipnlocal: fix portlist service event type mismatch (#21433)
The portlist extension publishes a PortlistServices event, but the
LocalBackend subscriber accepted the underlying slice type instead. The
event bus matches event types exactly, so endpoint updates were dropped
before reaching Hostinfo.Services.

Add a regression test covering the event-to-Hostinfo path.

Fixes #20192

Signed-off-by: YewFence <hello@yewfence.dev>
2026-09-23 11:55:33 -04:00
Brad Fitzpatrick 63f625397b misc/bumpdeps, .github/workflows: unify the bumpdeps tool and bumpdep workflow
The bumpdep workflow grew 170 lines of inline shell that duplicated
much of what misc/bumpdeps already did in Go (resolving versions,
special branches, running go get), and disagreed with it in small ways
(the workflow fetched wireguard-go with GOPROXY=direct; the tool asked
the proxy for the branch head). It was also untestable except by
extracting the run block and running it by hand.

Move all of that logic into misc/bumpdeps and make the workflow a thin
wrapper that runs it with -github. The tool now:

  - accepts the workflow's argument forms: "go" for the toolchain (via
    ./pull-toolchain.sh), the "wireguard-go" and "gvisor" aliases,
    exact module paths, "path@version", and modules not yet in go.mod,
    alongside its existing substring filters; comma-separated arguments
    are split, so the workflow input passes straight through;
  - finds branch heads with git ls-remote and then asks the proxy for
    the pseudo-version of that commit, so a just-pushed commit is seen
    without ever cloning (GOPROXY=direct on github.com/gokrazy/kernel.*
    hung the workflow for hours);
  - refuses downgrades by diffing go.mod before and after go get,
    rather than grepping go get's output;
  - runs "make tidy" and "make updatedeps" itself (-tidy=false to skip);
  - renders the PR title, body (GitHub compare links derived from the
    module path, including subdirectory-tagged modules), and commit
    message, printing a suggested commit message locally and, with
    -github, writing step outputs and the step summary. -issue accepts
    the issue URL or "#N" form and becomes the "Updates" line.

The branch name changes from a slug of every module path, which
produced names like actions/bumpdep/github.com-gokrazy-kernel.amd64-
main-github.com-gokrazy-kernel.arm64-main-github, to
actions/<workflow>/<actor>/<yyyymmdd-hhmmss> (no actor for scheduled
runs). Each run is a fresh PR; delete-branch cleans up after merge or
close.

The bumpdep workflow gains an optional exclude-newer-than-days input for
the cooldown. All of the formerly-inline logic now has unit tests,
including a fake module proxy.

Also add a gokrazy-bump workflow that runs the same tool every Sunday
night on the direct github.com/gokrazy/* modules (kernels, firmware,
gokrazy itself) and opens a PR labeled run-natlab-tests, so the natlab
VM tests boot the new kernels before merge. Substring filters now skip
indirect dependencies unless -indirect is set, so that "github.com/
gokrazy/" doesn't drag in gokapi; naming an indirect module exactly
still selects it.

Updates tailscale/corp#48312
Updates #8043
Updates #1866

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7c3e91a4d2f58b0e6a1c9d47f3b2e8a5c6d0f1e2
2026-09-23 06:42:52 -07:00
Raj Singh 5f0cf87429 cmd/containerboot: recover from IPN watch closure (#21374)
When containerboot falls behind on the IPN bus, tailscaled closes the
watch. containerboot treated the EOF as fatal and SIGTERMed a healthy
tailscaled, which is easy to hit on large, churny tailnets.

Instead, reconnect and rebuild state from the new watch's initial
status, and only request peer changes in modes that use them. If the
watch can't be reopened for a minute, exit so a dead tailscaled still
restarts the container.

Fixes #21373

Change-Id: Iad7749e4fd0f43eabdb471d6e64bb43f37ff70ff

Signed-off-by: Raj Singh <raj@tailscale.com>
2026-09-23 10:36:01 +01:00
Dep Updater 610b05c58e go.mod: bump github.com/gokrazy/kernel.amd64, github.com/gokrazy/kernel.arm64, github.com/gokrazy/kernel.rpi, github.com/gokrazy/rpi-eeprom, github.com/gokrazy/gokrazy
* github.com/gokrazy/kernel.amd64: v0.0.0-20260705070735-de680abf072b to v0.0.0-20260922084445-21771250d660
* github.com/gokrazy/kernel.arm64: v0.0.0-20260705071517-37841c4d6ff1 to v0.0.0-20260922084859-ac43a676b0b4
* github.com/gokrazy/kernel.rpi: v0.0.0-20251127164438-9778ec0261de to v0.0.0-20260911133309-2cbf751e3f2a
* github.com/gokrazy/rpi-eeprom: v0.0.0-20260518070910-95f7328a8228 to v0.0.0-20260913082024-b80c62cf428d
* github.com/gokrazy/gokrazy: v0.0.0-20260703061218-a4a45a20149d to v0.0.0-20260916140236-39fe3e5557b8

Triggered by @bradfitz via the bumpdep workflow.

Updates #1866

Signed-off-by: Dep Updater <noreply+dep-updater@tailscale.com>
2026-09-22 18:46:04 -07:00
David Bond 8af8f861c0 cmd/k8s-operator/e2e: add tests for kube-apiserver ProxyGroups (#20993)
The e2e suite covered the operator's in-process API server proxy but
not the ProxyGroup-based one. Add tests for both proxy modes that
drive a ConfigMap through its lifecycle via the proxy, verify a
forbidden request is rejected, and check that deleting the ProxyGroup
cleans up its StatefulSet and Tailscale Service.

Fixes tailscale/corp#38009

Change-Id: Ifc0be47ce32dd96f8daa748af6785a8b82ec19a7

Signed-off-by: David Bond <davidsbond93@gmail.com>
2026-09-22 22:39:20 +01:00
Brad Fitzpatrick b3de3b2217 ipn/ipnlocal, net/netns: bind Linux peerapi listener to the tun device
Linux is a weak-host stack, so a LAN-adjacent machine can complete a
TCP handshake with a node's peerapi listener by sending a packet to
the node's Tailscale IP, with no credentials and no tailnet
membership.

macOS and iOS already bind the listener to the tunnel interface, and
Windows is protected by its strong host model, so Linux tun mode was
the only platform that leaked.

Bind the Linux listener to the tunnel interface as well, so the
kernel only answers connections that arrive from the tunnel or from
the local host. A natlab VM test verifies that a same-LAN machine can
no longer complete the handshake, while local and peer peerapi keep
working.

FreeBSD has the same weak-host exposure but no per-socket equivalent,
so handling it there with pf is a TODO (#21419).

Updates tailscale/corp#48248

Reported-By: Samuel Keeley (@keeleysam)
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I5f8501b0938c9f7aa39c4c12ebddd988c72e89bf
2026-09-22 14:10:30 -07:00
John Costa 9ceb71a83b scripts/installer.sh: add Omarchy as an Arch derivative
Fixes #21376

Signed-off-by: John Costa <costajohnt@gmail.com>
2026-09-22 13:53:57 -07:00
Mazdak Nasab 1edf391c9c feature/conn25: add ipv6 test coverage for datapath (#21362)
Add ipv6 test coverage for conn25 datapath by
extending existing tests.

Fixes tailscale/corp#42521

Change-Id: I4c6df35572d96cc80cae3685353522942e2a8546

Signed-off-by: Mazdak Nasab <mazdak.nasab@gmail.com>
2026-09-22 09:36:50 -07:00
Tom Proctor df0bd83055 cmd/k8s-operator/e2e: allow concurrent and partially torn down tests (#21416)
Loosens the connector route assertion to allow concurrent tests against
the same tailnet to more reliably pass, while still asserting the
client itself is advertising those routes and they're recognised in the
API.

Also make createOrUpdate more resilient to a test that didn't fully tear
down on the same cluster previously. By respecting the existing resource
version and finalizers, we can update resources that didn't get deleted
from a previous run.

Updates tailscale/corp#45426

Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com>
2026-09-22 14:23:44 +01:00
Tom Proctor b26751eafc cmd/k8s-operator/e2e: select matching family IP addresses (#21415)
Make sure the tests select matching-family IP addresses for ingress
tests so they're able to pass on IPv6 clusters. We should probably
follow up with another change for the operator that stops clients from
needing to do this sort of selection, but updating the tests is the easy
option to get them passing on IPv6 clusters in the short term.

Updates tailscale/corp#45426

Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com>
2026-09-22 14:23:06 +01:00
Brad Fitzpatrick ee029383f7 net/{pktinfo,stunserver}: reply from the address the request was sent to
If people ran derper on a multi-NIC or multi-address host, STUN replies
could go out from the wrong address. The wildcard UDP socket let the
kernel pick the reply's source address by routing to the client, which
means the default route's address rather than the one the request came
in on. With connmark-based policy routing (e.g. DNAT through a tunnel),
the reply then doesn't match the inbound conntrack entry, goes out the
wrong interface, and the client never sees it. DERP over TCP was fine,
since accepted sockets are pinned to the local address.

Add net/pktinfo, a small Linux-only package that uses IP_PKTINFO and
IPV6_RECVPKTINFO to learn each datagram's destination address and to
reply from it, and use it in the STUN server. Only the source address is
pinned; routing still picks the interface. Other platforms are unchanged.

Fixes #21404

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7b3e9c2d41a8f60e5d9c1b2a3f4e5d6c7b8a9f01
2026-09-21 15:43:30 -07:00
James Tucker 816b661f0d client/web, tsweb/compserve: serve index.html for the root path
GET / in the client/web assets handler normalized to an empty fs
path. The empty name is invalid, and file systems wrapped in fs.Sub
reject it with fs.ErrInvalid instead of fs.ErrNotExist. The
index.html fallback only triggers on fs.ErrNotExist, so it never ran
and the request failed with a 500 "internal error". This broke the
corp smoke test that curls http://100.100.100.100/.

Normalize the empty path to index.html before serving. Also make
compserve.ServeFile return fs.ErrNotExist for invalid paths, per its
documented contract, instead of leaking fs.ErrInvalid.

Both layers have regression tests verified to fail without the fix.

Updates tailscale/corp#20099
Updates #12170

Signed-off-by: James Tucker <james@tailscale.com>
2026-09-21 15:42:33 -07:00
Brad Fitzpatrick 523b626a8e tsweb: implement FlushError on loggingResponseWriter
http.ResponseController.Flush prefers a FlushError method at each level
of the ResponseWriter wrapper chain and falls back to Flusher.Flush,
which can't report an error. loggingResponseWriter only had Flush, so
handlers using ResponseController to flush never saw the underlying
writer's flush errors, such as a stream reset by an expired write
deadline. Add FlushError that delegates via ResponseController so
errors from the wrapped writer propagate.

Updates tailscale/corp#48531

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I259c5de234435f632387261165fe15dcaf34ff45
2026-09-21 13:52:40 -07:00
Claus Lensbøl 3f3e56f412 wgengine/netstack: avoid returning from sender goroutines on error (#21332)
Returning from injectToHost and injectToWireGuard leaves the two
goroutines dead and the host without having a way to send traffic.

This is especially relevant for android where the tundev is torn down
and recreated for every call to updateTUN() from a route table change or
roaming between networks.

Fixes #21155

Signed-off-by: Claus Lensbøl <claus@tailscale.com>
2026-09-21 13:56:52 -04:00
Mike Jensen bb94defdd0 fuzz: improve fuzz testing and seed corups (#21340)
This change improves the initial fuzz seeds to get better coverage. It also includes a fix to the geo fuzzing to avoid a harness induced failure when a NaN input is provided.

No actual code logic changes, test only.

Updates tailscale/corp#46608

Change-Id: I0985b6d3a75603927451eea0a25f4cd72c039795

Signed-off-by: Mike Jensen <mikej@tailscale.com>
2026-09-21 09:23:51 -06:00
BeckyPauley 7bf76690f0 cmd/k8s-operator: recover missing egress EndpointSlices (#21174)
* cmd/k8s-operator: move egress EndpointSlice write back into gated provision

PR #20347 moved the EndpointSlice createOrUpdate outside of provision,
causing it to run on every reconcile. This resulted in racing egress-eps on
the EndpointSlice, sometimes causing the TailscaleEgressSvcConfigured to
become stuck as False with the Service not fully updated. Gate it again so
it only runs when a reprovision is required.

Updates #20916

Signed-off-by: Becky Pauley <becky@tailscale.com>

* cmd/k8s-operator: recover missing egress EndpointSlices

Add a watch for EndpointSlices in the egress-services reconciler so a
deleted slice re-triggers a Service reconcile directly. Treat a Service
whose expected per-family EndpointSlice is missing as not up to date so it
re-enters provision and recreates the slice.

Also sort endpoints by Pod UID before writing them in the egress-eps
reconciler, so an unchanged set of ready Pods cannot result in a different
order and trigger an unnecessary Update.

Updates #20916

Signed-off-by: Becky Pauley <becky@tailscale.com>

---------

Signed-off-by: Becky Pauley <becky@tailscale.com>
2026-09-21 11:21:57 +01:00
Brad Fitzpatrick 3014ad828e tstest/natlab/{vmtest,vnet}: bound "tailscale up", dump tailscaled logs on failure
When a natlab node's "tailscale up" hung (TestEasyEasy on CI, twice on
main), the only bound was the 10 minute test context, so the failure
arrived as go test's timeout panic. The VM console logs weren't dumped
(t.Cleanup doesn't run on a panic), and they wouldn't have helped
anyway: on gokrazy the console holds only kernel and init output, while
tailscaled's stdout/stderr goes to a remote syslog that vnet discards
unless the node has VerboseSyslog set. There was no way to see what the
stuck node was doing.

Bound each node's "tailscale up" in Env.Start to 90 seconds (it takes
about a second against the in-process control server), so a stuck node
becomes a normal test failure. On failure, dump the tail of each node's
tailscaled logs from vnet's fake log.tailscale.com log catcher, which
already buffered them per node but exposed them to nothing; add
Server.NodeLogs for that. Also add VMTEST_VERBOSE_SYSLOG=1 to stream
the guests' syslog into the test output live, the vmtest equivalent of
tstest/integration/nat's --log-tailscaled flag.

With this, the CI hang reproduced locally under CPU pressure (2 CPUs
shared with busy loops) in 1 of 22 runs, and the dumped logs showed
tailscaled's control client stuck at "awaiting unpause", which is fixed
separately.

Updates #deflake

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Icfd3a2cdc66d924388a722f3a3cd19f86e116921
2026-09-20 11:02:04 -07:00
Brad Fitzpatrick 7d96cf5a62 tsweb/varz: read memstats via runtime/metrics, export fork metrics
The varz handler got its memstats_* metrics from the expvar package's
"memstats" func, which calls runtime.ReadMemStats and so stopped the
world on every Prometheus scrape. Keep the names but compute them from
runtime/metrics, and never call that func, even from
WritePrometheusExpvar.

While there, export the /tailscale/ metrics from our Go fork (stack
size histogram, stack copy counters, timer zombie counts and lifetime
histogram), which nothing could see before, plus a few upstream ones
with no MemStats equivalent: scheduling latency and GC pause
histograms, live heap, GC and total CPU seconds, thread count, and
mutex wait time. The last replaces derper's hand-rolled version.

Everything read is cheap and a scrape allocates nothing after the
first. Names use a go_runtime_ namespace rather than go_ so they can't
collide with the Prometheus Go client's collector in promvarz binaries.
The runtime's 162-bucket time histograms are reduced to one bucket per
factor of four from 256ns to 1s.

Updates #21300
Updates tailscale/go#189
Updates golang/go#75935

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7e3c9a41f2b85d6e0c4a9b1d3f8e7c2a5b6d4e19
2026-09-19 18:19:20 -07:00
Dep Updater ae69c6b6c9 go.toolchain.rev: bump Go toolchain
* Go toolchain: https://github.com/tailscale/go/compare/32e8826b089fee8cb0c5c4822b9794ca5004f23a...24ee2fd0610e6c505ee4ec061a81215afe119d1f

Triggered by @bradfitz via the bumpdep workflow.

Updates tailscale/corp#29053

Signed-off-by: Dep Updater <noreply+dep-updater@tailscale.com>
2026-09-19 16:52:15 -07:00
Dep Updater 9617640d20 go.mod: bump github.com/bradfitz/go-tool-cache
* github.com/bradfitz/go-tool-cache: v0.0.0-20260909201542-a1c7321be47b to v0.0.0-20260919185303-c660171c910c

Triggered by @bradfitz via the bumpdep workflow.

Updates tailscale/corp#47471

Signed-off-by: Dep Updater <noreply+dep-updater@tailscale.com>
2026-09-19 12:16:27 -07:00
Brad Fitzpatrick 1c778640ea net/netmon, ipn/ipnlocal, wgengine: don't lose a network change during startup
If the network comes up during the few milliseconds NewLocalBackend
takes, tailscaled starts with its control client paused and never
unpauses: the interface state snapshot was taken before the netmon
subscription (kept late since #17252), so a change published in between
reached nobody. #21261 hit it on fast-booting NixOS microVMs, and it was
behind the natlab TestEasyEasy CI flake, where a gokrazy node's DHCP
lease landed in that window and "tailscale up" hung at "awaiting
unpause".

Re-read netmon's state after subscribing rather than subscribing before
the snapshot (as #21281 proposed), which would bring back the #17252
data race and let a fresh delta be overwritten by the stale snapshot.
netmon.New and the userspace engine had the same shape of gap between
their snapshot and their subscription; close those too.

Under CPU pressure the hang reproduced in 1 of 22 TestEasyEasy runs
before and 0 of 120 after.

Thanks to @c-vigo for the detailed diagnosis in #21261 and to @zzz-yu
for pointing out this problem and proposing a fix in #21281!

Fixes #21261
Updates #14902
Updates #19126
Updates #deflake

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ifa21f8055a85321a5afda7800140fc5045c6ce08
2026-09-19 10:46:50 -07:00
Brad Fitzpatrick 574ef3b2a8 gokrazy, .github/workflows: build the natlab image with tool/go
The natlab-basic workflow builds the gokrazy natlab image in its own
step so that the test's rebuild of it (vmtest always rebuilds, so the
baked-in binaries match the source under test) is a build cache hit
rather than a cold build inside go test's -timeout budget. That never
worked: the Makefile ran whatever "go" was on $PATH, the runner's stock
Go, while the go command puts its own $GOROOT/bin first on the test
binary's $PATH, so the rebuild from inside "go test" used tailscale/go.
GOCACHE entries embed the compiler's build ID, so the step warmed
nothing. In practice the in-test rebuild took about 2.5 minutes of the
3 minute -timeout, leaving TestEasyEasy about 20 seconds for booting
two VMs, logging in, and pinging. A passing run on main took 167s. Any
hiccup in the remaining budget, such as the "tailscale up" hang fixed
separately, ended in go test's timeout panic with no useful output.

Make the natlab targets in gokrazy/Makefile use ../tool/go so the step
and the test use the same toolchain and cache. Fix the same mistake in
natlab-test.yml's cache warming step, whose comment documented the
wrong belief about which toolchain the in-test builds use. Raise
natlab-basic's -timeout to match natlab-test.yml so that a hang fails
through vmtest's own bounded waits (which dump the node's logs) instead
of through go test's timeout panic (which dumps nothing about the VMs).

Updates #13038
Updates #deflake

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Iaa6085ec5aa029373204baf75b169ff375c2b355
2026-09-19 10:45:55 -07:00
Brad Fitzpatrick 51682d245b derp/derpserver: drop per-client writer goroutine, start on demand
Each client connection ran two goroutines for its lifetime: the reader
in sclient.run and a sendLoop blocked in a select over its send
queues, pong, peer gone, mesh update, and keepalive channels. Almost
all clients are idle at any moment, so the second goroutine mostly
pinned memory: a 4 KiB stack, a g struct, a sudog per select case,
three channels, and a context and errgroup. At 100k idle connections
that was about 8 KB of a client's 22 KB RSS.

Instead, start up the sendLoop only as needed, letting the goroutine
go away otherwise, like Go 1.28-dev's http2 code
(golang/go@5c51011e82) with similar parking to
https://go.dev/cl/834084 but DERP's producers are all non-blocking, so
a kick bit replaces that http2 code's send count.

Measured with 100k idle TLS connections, server RSS per client went
from 22.3 KB to 14.5 KB (22.9 KB to 15.3 KB after each connection had
carried a packet), goroutines dropped from 2 to 1, and with no change
in BenchmarkSendRecv throughput or allocations and no change in the
time to do 100k serial round trips, each of which parks and wakes the
writer. (it's super cheap to start goroutines)

Updates #21064

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7c3e9a51d4b8f2607a1e5c3d9f8b2a4e6c0d1f3b
2026-09-18 19:19:44 -07:00
Brad Fitzpatrick 630e704763 derp/derpserver: make unique sender cardinality tracking opt-in
Each connected client kept a HyperLogLog sketch of the peers that had
sent it packets, and every relayed packet was inserted into it under a
mutex. That costs memory per client and time on the packet path for a
debug-only estimate that few servers look at.

Keep the accounting but only allocate the sketch when the new
TS_DERP_SENDER_CARDINALITY environment variable is set. When it is
unset, EstimatedUniqueSenders reports 0 and the debug traffic page
omits the field as before.

Updates tailscale/corp#24681

Change-Id: I471bf816e5461069d50b97696e8d91ec3cbe3389
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-18 17:23:22 -07:00
Francois Marier b7a3b1b293 net/portmapper: fix potential race around pcpNonce
All reads from c.mapping need to take place while the lock is held.

Updates #21127

Change-Id: I4782d5ca027584a9e47dc6241af1970348b18c0a
Signed-off-by: Francois Marier <francois@tailscale.com>
2026-09-18 15:01:05 -07:00
Francois Marier 2578d08e8d net/portmapper/pcp: use all-zero address for PCP releases
The PCP spec says that when deleting a mapping, both the external
port and the external address must be zero. While PCP servers are
likely to be lenient in practice, we should do the right thing in
case we encounter strict validation.

Resolves #21363

Change-Id: Ice3467d735def5144a810c39e65642988a830972
Signed-off-by: Francois Marier <francois@tailscale.com>
2026-09-18 10:11:53 -07:00
Fran Bull a9bb6d190b tailcfg: add nodecap conn25-connector-apps
Which will be set to a slice of app names a peer is a connector for.

Updates tailscale/corp#47251

Signed-off-by: Fran Bull <fran@tailscale.com>
2026-09-18 09:13:46 -07:00
Brad Fitzpatrick f87a1b1a82 derp/derpserver: pool received packet payload buffers
Every packet the server relayed allocated a fresh []byte for its
payload in recvPacket or recvForwardPacket and dropped it once the
destination's sendLoop had written it. On one busy server, this was
observed allocating about 160 MB/sec of short-lived garbage, and GC
plus malloc were about 5% of the process CPU profile.

Instead, take payload buffers from a size-classed sync.Pool on the
Server, with power-of-two classes from 1 KiB up to derp.MaxPacketSize,
and return them once the packet has been written, forwarded, or
dropped. sync.Pool holds nothing per connection and is trimmed by the
GC, so idle clients pin no memory; only packets actually in flight
hold a buffer. A compile-time assertion ties the largest size class to
derp.MaxPacketSize, and the get and put helpers panic on sizes outside
the pool's classes rather than indexing past it.

Because the memory is now reused, PacketForwarder implementations must
not retain the payload after ForwardPacket returns. Make that explicit
in the signature: the payload is passed as a new derp.LoanedBytes
value, which exposes only Len, WriteTo, and Clone, so an implementation
has to copy to keep it. derp.Client and derphttp.Client, the real
implementations, already wrote it out synchronously; the test-only
channelFwd now clones.

BenchmarkSendRecv shows one fewer allocation per relayed packet and,
for 1000-byte packets, B/op down from 1278 to 263. ns/op on the
loopback benchmarks is dominated by syscalls and is unchanged within
noise.

Updates #21064

Change-Id: Ie40c82388ddb5d22f75fa828749b53fcaba9adde
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-18 08:55:48 -07:00
Adrian Dewhurst c2dc086468 util/multierr: mark multierr.New as deprecated
We banned use of multierr in various dep tests, so this makes the
situation more obvious if someone stumbles across it.

Change-Id: I17e80880e57e5005fcaea864a365ba264411e802
Signed-off-by: Adrian Dewhurst <adrian@tailscale.com>
2026-09-18 11:46:09 -04:00
Brad Fitzpatrick 47debc5a6f derp/derpserver: don't build debug log arguments on the packet path
sclient.debugLogf and Server.debugLogf check a debug flag before
logging, but Go evaluates and boxes their arguments before the call.
The per-packet call sites in run, handleFrameSendPacket,
handleFrameForwardPacket, sendPkt, recordDrop, and sendPacket's
deferred stats func were therefore calling key.NodePublic.ShortString
and boxing frame headers on every relayed packet, all for messages
that were then discarded.

On one busy server's heap profile, those discarded arguments were
about half of all objects allocated by the process. Guard each hot
call site with the debug flag so nothing is built unless it will be
logged, and document that requirement on both debugLogf methods.

While here, give the sender cardinality sketch its key bytes from a
stack array rather than an AppendTo(nil) allocation per packet.

BenchmarkSendRecv drops from 10 or 11 allocations per relayed packet
to 3, and BenchmarkConcurrentStreams from 11 to 4.

Updates #21064

Change-Id: Ibb7c4fbff546b6f1b40ce0a21d41ae0976705941
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-18 06:52:51 -07:00
Brad Fitzpatrick 1033e714ca cmd/testwrapper: run all package patterns in one go test invocation
testwrapper ran a separate, sequential "go test" invocation for each
package pattern on its command line. That is fine for a single "./..."
argument but not for callers that pass an explicit package list: CI
jobs in the corp repo passing ~200 packages ran ~200 serial go test
processes with no cross-package parallelism and a fixed set of
never-cacheable lookups per process, and spent several times longer
on process startup, package loading, cache lookups, and serial test
binary links than on running tests. See tailscale/corp#48453 for the
details.

Locally, on 203 packages with a fully warm build and test cache, so
measuring only the per-invocation overhead:

  old (203 go test processes):  26.4s
  new (1 go test process):       3.6s  (7.3x faster)

Our own Windows CI job hits the same path: its "sharded:N/M" mode
expands to an explicit list of that shard's packages via listpkgs, so
each shard ran one go test process per package, and Windows process
startup is slower still. Each shard now runs as one invocation.

Fixes tailscale/corp#48453
Updates tailscale/corp#47035

Change-Id: I3f796ff1724af40f93be9f918a7ddfde3bb45a91
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-18 02:56:29 -07:00
Brad Fitzpatrick 3323dc02f4 tsweb: restore AcceptsEncoding
Commit 6608b9a38 removed tsweb.AcceptsEncoding, saying it had no
callers, but it has many callers in the tailscale.io repo. Restore the
function and its test unchanged so those builds work again.

Updates #12170
Updates tailscale/corp#48447 (broken by this)

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7c3e2a9f5b1d4e8a6c0f2b3d9e1a7c5f4b8d2e6a
2026-09-17 21:36:46 -07:00
OSS Updater 28836381da go.mod: update web-client-prebuilt module
Signed-off-by: OSS Updater <noreply+oss-updater@tailscale.com>
2026-09-17 17:55:09 -07:00
yaruk-byte be89526457 tstest/integration: stop skipping Windows integration tests (#21187)
Updates #20750

Signed-off-by: Yaruk Asghar <yaruk@tailscale.com>
2026-09-17 17:11:00 -07:00
James Tucker aa1134d358 go.mod: update golangci-lint to v2
Moving to v2 because v1 references repositories that have been
deleted from GitHub, breaking GOPROXY=direct.

The lint config (.golangci.yml) was already in v2 format; this updates
the tool dependency used by 'make lint' to golangci-lint/v2 v2.13.2,
drops the now-obsolete blank import in internal/tooldeps in favor of a
Go 'tool' directive, and bumps the CI workflow binary to match.

Updates #cleanup

Signed-off-by: James Tucker <jftucker@gmail.com>
2026-09-17 16:55:23 -07:00
James Tucker 6608b9a387 tsweb/compserve, client/web: add zstd, remove brotli for precompressed assets
Adds tsweb/compserve: content negotiation for precompressed static
variants, with a transcode-to-identity fallback for clients that do not
accept an encoding (including when the raw file is absent), and
CompressWriter, which live-compresses dynamic responses with zstd in its
fastest mode, streamed incrementally with no buffering. Negotiation is
q-value and wildcard aware (gzip;q=0 previously matched gzip).

client/web serves its prebuilt embedded assets through compserve,
replacing brotli with zstd; the embedded FS is wrapped in
tsweb/vcstime for conditional-request mod times. tsweb/compserve/gzip.go
keeps transitional serving of gzip variants from pre-zstd file systems
(such as the currently published web-client-prebuilt module):
passthrough to gzip-accepting clients, transcoded to identity otherwise;
it becomes inert once a zstd-only module is published.

util/zstdframe gains pooled GetDecoder and GetStreamingEncoder
(concurrency=1). util/precompress is now a build-time tool, generating
zstd variants only. cmd/tsconnect and cmd/build-webclient consume the
new precompress/compserve split. tsweb.AcceptsEncoding and
tsweb/tswebutil are removed; negotiation lives in compserve and the
deprecated shim had no callers. go.mod bumps web-client-prebuilt.

Also fixes a transcoding bug where http.ServeContent's size probe via
the promoted zstd.Decoder.WriteTo could report a zero length, serving
empty bodies.

Updates tailscale/corp#20099

Signed-off-by: James Tucker <james@tailscale.com>
2026-09-17 15:23:49 -07:00
Dep Updater 178ef3db08 go.toolchain.rev: bump Go toolchain
* Go toolchain: https://github.com/tailscale/go/compare/d030173bb47a6c4a6f885cb56a97dd9eca5fb8b7...32e8826b089fee8cb0c5c4822b9794ca5004f23a

Triggered by @bradfitz via the bumpdep workflow.

Updates tailscale/go#189

Signed-off-by: Dep Updater <noreply+dep-updater@tailscale.com>
2026-09-17 15:03:04 -07:00
Patrick O'Doherty 2e72593fbf feature/clientupdate: require write access for update/install localapi (#21360)
The update/install localapi handler had no PermitWrite check, so any
local user who could reach the localapi socket could make the root
daemon self-update and restart itself. This has been the case since the
endpoint was added in November 2023.

Gate the handler behind PermitWrite so that only root or the operator
user can trigger a self-update, matching the other mutating handlers.

Updates tailscale/corp#48187

Change-Id: Iadfef939f5dab684652cd220e77de63b54bfca2f
Reported-by: Ben Carman <benthecarman@live.com>

Signed-off-by: Patrick O'Doherty <patrick@tailscale.com>
2026-09-17 14:30:40 -07:00
Naman Sood 027e249fcf net/socks5: correctly proxy half-closed TCP connections
Similar to #16462, when we are acting as a TCP proxy, we need to pass
through half-closes correctly since clients and servers will sometimes
close one direction of the connection and still rely on the other
direction working.

Fixes #20883.

Signed-off-by: Naman Sood <mail@nsood.in>
2026-09-17 17:22:32 -04:00