Commit Graph
2202 Commits
Author SHA1 Message Date
Brad Fitzpatrick e87076a378 ipn/auditlog: store the audit log in the backend's state directory
The store path was hardcoded to %ProgramData%\Tailscale\audit-log.json on
Windows, so a tailscaled run by a regular user with its own --statedir
logged "[unexpected] failed to create audit log store ... Access is denied"
on every profile switch. Use the backend's TailscaleVarRoot when it is
known, falling back to the platform default otherwise. For the Windows
service the state directory is %ProgramData%\Tailscale, so its path is
unchanged; an explicit SetStoreFilePath still wins.

Updates #2791

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I5d7f9b1c3e5a7c9e1b3d5f7a9c1e3b5d7f9a1c3e
2026-09-24 11:52:21 -07:00
Brad Fitzpatrick 6612bed24d ipn/desktop: skip the extension when sessions can't be enumerated
Registering the session state callback enumerates the existing desktop
sessions, which fails for a tailscaled run by an unprivileged user
(WTSEnumerateSessions returns "No more data is available"). That was
reported as "init failed", which reads like a bug. Such a tailscaled has
no other users' sessions to track, so report it as skipping the extension,
the same way an unavailable session manager already is.

Updates #2791

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7f9b1d3e5a7c9e1b3d5f7a9c1e3b5d7f9a1c3e5b
2026-09-24 11:51:51 -07:00
YewFence 7122baac1b ipnlocal: fix portlist service event type mismatch (#21433)
The portlist extension publishes a PortlistServices event, but the
LocalBackend subscriber accepted the underlying slice type instead. The
event bus matches event types exactly, so endpoint updates were dropped
before reaching Hostinfo.Services.

Add a regression test covering the event-to-Hostinfo path.

Fixes #20192

Signed-off-by: YewFence <hello@yewfence.dev>
2026-09-23 11:55:33 -04:00
Brad Fitzpatrick b3de3b2217 ipn/ipnlocal, net/netns: bind Linux peerapi listener to the tun device
Linux is a weak-host stack, so a LAN-adjacent machine can complete a
TCP handshake with a node's peerapi listener by sending a packet to
the node's Tailscale IP, with no credentials and no tailnet
membership.

macOS and iOS already bind the listener to the tunnel interface, and
Windows is protected by its strong host model, so Linux tun mode was
the only platform that leaked.

Bind the Linux listener to the tunnel interface as well, so the
kernel only answers connections that arrive from the tunnel or from
the local host. A natlab VM test verifies that a same-LAN machine can
no longer complete the handshake, while local and peer peerapi keep
working.

FreeBSD has the same weak-host exposure but no per-socket equivalent,
so handling it there with pf is a TODO (#21419).

Updates tailscale/corp#48248

Reported-By: Samuel Keeley (@keeleysam)
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I5f8501b0938c9f7aa39c4c12ebddd988c72e89bf
2026-09-22 14:10:30 -07:00
Brad Fitzpatrick 1c778640ea net/netmon, ipn/ipnlocal, wgengine: don't lose a network change during startup
If the network comes up during the few milliseconds NewLocalBackend
takes, tailscaled starts with its control client paused and never
unpauses: the interface state snapshot was taken before the netmon
subscription (kept late since #17252), so a change published in between
reached nobody. #21261 hit it on fast-booting NixOS microVMs, and it was
behind the natlab TestEasyEasy CI flake, where a gokrazy node's DHCP
lease landed in that window and "tailscale up" hung at "awaiting
unpause".

Re-read netmon's state after subscribing rather than subscribing before
the snapshot (as #21281 proposed), which would bring back the #17252
data race and let a fresh delta be overwritten by the stale snapshot.
netmon.New and the userspace engine had the same shape of gap between
their snapshot and their subscription; close those too.

Under CPU pressure the hang reproduced in 1 of 22 TestEasyEasy runs
before and 0 of 120 after.

Thanks to @c-vigo for the detailed diagnosis in #21261 and to @zzz-yu
for pointing out this problem and proposing a fix in #21281!

Fixes #21261
Updates #14902
Updates #19126
Updates #deflake

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ifa21f8055a85321a5afda7800140fc5045c6ce08
2026-09-19 10:46:50 -07:00
Alex Chan b16bc957aa ipn,tsnet: replace LocalBackend.NetMap with NetMapNoPeers/NetMapWithPeers
This method has been deprecated for four months; replace all uses with
the Peers or NoPeers equivalents.

Updates #12542

Change-Id: I7b5800d0d92775839dbfb6021751a5dfacc53d2f
Signed-off-by: Alex Chan <alexc@tailscale.com>
2026-09-17 13:43:14 +01:00
Andrew Dunham 158a2476c0 tailcfg: add Node.StableTailnetID and bump capver (#21314)
Add Node.StableTailnetID for control to send the current tailnet's
stable ID on the self node. Expose it as CurrentTailnet.StableID in
LocalAPI status and `tailscale status --json`, and display it in
`tailscale whoami`.

Bump CurrentCapabilityVersion to 148.

Updates #14375

RELNOTE: Show the current tailnet's stable ID in status JSON and whoami.

Change-Id: I0525ff735de8113c8d124045a94e00c19a5a02e2

Signed-off-by: Andrew Dunham <andrew@tailscale.com>
2026-09-16 14:28:46 -04:00
Brad Fitzpatrick 96492df022 ipn/ipnlocal: log the profile's login name, not the method value
The profile switch log lines passed cp.UserProfile().LoginName without
calling it, so they printed "%!q(func() string=0x7ff7141cd060)" in place
of the login name.

Updates #cleanup

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I4e6a8c0b2d5f7a9c1e3b5d7f9a1c3e5b7d9f1a3c
2026-09-16 09:05:27 -07:00
maxiscoding28 044abcc9d2 ipn/ipnlocal: make Serve idle connection limit configurable (#21179)
* ipn/ipnlocal: make Serve idle connection limit configurable

Allow Kubernetes operator proxy pods to override the per-host idle connection limit through an environment knob.\n\nUpdates tailscale/tailscale#20875

Signed-off-by: maxiscoding28 <max.winslow@icloud.com>

* ipn/ipnlocal: document Serve idle connection default

Signed-off-by: maxiscoding28 <max.winslow@icloud.com>

---------

Signed-off-by: maxiscoding28 <max.winslow@icloud.com>
2026-09-16 11:01:27 +01:00
leoca ec1e07c737 ipn/ipnlocal: don't panic on an over-long name in the peerAPI DNS debug mode
dnsQueryForName builds the query for GET /dns-query?q=<name>, the peerAPI's
interactive debug mode. The name comes from the peer's query string and the
only thing done to it is appending a trailing dot, but the query is built
with dnsmessage.MustNewName, which panics as soon as the name is longer than
255 bytes.

The panic happens before the query reaches the resolver, so the nameAllowed
filter never gets a say and only sourceAllowed has to be true: any of the
user's own untagged devices, and any peer an extension hook lets through,
such as a client using this node as an exit node with DNS proxying allowed.
http.Server recovers it, so tailscaled survives, but the peer's connection is
dropped instead of answered and the node logs a panic trace per request.

Just under the limit the name was mishandled too: a 255-byte name is accepted
by NewName but rejected by Question when it packs the name, and that error
was dropped, so Finish returned a well-formed query with no question in it
which was then handed to the resolver.

Have dnsQueryForName use NewName, check the error from Question, and return
the error from Finish. handleDNSQuery turns a name it cannot build a query
for into the 400 it already uses for the other malformed-request cases.

Fixes #21307

Change-Id: I7c1a4b7e4f9a1d2c3b5e8f0a6d4c2b9e1f3a7d50
Signed-off-by: leoca <leo.camus23@gmail.com>
2026-09-15 17:53:19 -07:00
Brad Fitzpatrick 2ee809d10d feature: add TS_DISABLE_FEATURE to disable features at runtime
The ts_omit_<name> build tags omit a feature at build time; there has
been no way to do the same at runtime. Some users (either proactively
or in response to a security announcement) might like a way to disable
a feature that's linked-in in their binaries that they're not using.
Then a mitigation announcement can say "set this env var" without
asking users to rebuild or wait for a new release.

This adds env var TS_DISABLE_FEATURE, a comma-separated list of
feature names to disable, and the listed set is reported by the
debug-optional-features LocalAPI endpoint next to the registered set.
The legacy per-feature knobs such as TS_DISABLE_SSH_SERVER and
TS_DISABLE_TAILDROP keep working independently.

A disabled feature behaves as if it had not been linked: it is absent
from feature.IsRegistered, its hooks are unset, and its extensions and
handlers are not registered. Three pieces make that happen:

  * feature.Register now returns bool, false when disabled, and
    feature packages gate their registration init on it. It was added
    to the feature packages that never called it (including taildrop
    and ssh), which also completes the picture reported by
    debug-optional-features. taildrop, routecheck, favorites, and
    serviceclientprefs had registration split across several inits and
    now register from one gated init.
  * ipnext.RegisterExtension ignores a disabled feature's extension.
  * feature.Hook.Set and feature.Hooks.Add walk the call stack and
    silently skip when the calling package under feature/<name> is
    disabled. This covers sub-packages such as
    feature/captiveportal/netcheckhook, which cannot call Register
    themselves without colliding with their parent, and future
    packages whose authors forget the gate.

ssh/tailssh's registrations moved from its inits into tailssh.Register,
called from feature/ssh's gated init. The aws and kube state stores and
syspolicy's Windows store registration are gated too.

feature/register_disable_test.go runs this test binary as a child
process (it links condregister, as tailscaled does) with
TS_DISABLE_FEATURE set to every registered feature at once, and fails
if any of them register anyway, so a feature that ignores the variable
cannot land.

Updates #12614

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I720af6ccab844ae060a9dfd1539fee577fd483e3
2026-09-15 08:37:47 -07:00
Brad Fitzpatrick 8344d97054 control/controlclient,control/ts2021,util/httpbody: bound control response sizes
The map response reader now caps a single message at 256 MiB on the wire
and 1 GiB after zstd decompression. The server-chosen uint32 size prefix
previously let a malicious control server make us allocate up to 4 GiB
before reading any body bytes, and the decoded size was unbounded, so a
small zstd frame could expand into gigabytes of JSON. A 16 MB cap has
been hit by real production traffic before, so both limits sit far above
plausible legitimate sizes. The size-prefixed read moved into a
readMapResponseMessage helper, and the newer control/tsp path already
enforced both kinds of bounds; this brings the long-poll path it
replaces in line, with more generous limits for large tailnets.

ts2021.Client.Do additionally caps every noise response body with
httpbody.LimitSize, so a malicious or buggy control server can't make us
buffer an unbounded response. Client.Do shadows the embedded
http.Client's Do method, so register, set-dns, set-device-attr,
audit-log, all DoNoiseRequest consumers (webclient, tailnet lock, SSH
actions, id-token, feature queries), and the debug CLI get the cap
without per-call-site changes, and future noise endpoints get it for
free.

The cap lives in the new util/httpbody package so other HTTP clients can
adopt the same convention: LimitSize looks the size limit up from
res.Request's context (a Response knows the Request that produced it),
falling back to DefaultMaxSize, 1 MiB, when the context carries no
override. It is like io.LimitReader except that reads past the limit
fail with an error wrapping httpbody.ErrTooLarge instead of silently
truncating, and a body of at most the limit, including one of exactly
the limit, reads back without error: the wrapper probes for EOF once the
limit is exhausted to tell an exactly-at-limit body from an oversize
one. The per-request override, httpbody.WithMaxSize, is a context key,
so transports pick it up with no API changes; LimitSizeTo applies an
explicit limit ignoring any override. Repeated LimitSize or LimitSizeTo
calls replace the previous limit rather than compounding it, so a later
call can raise or remove the limit an earlier one set.

Responses that stream an unbounded number of individually bounded
messages disable the cap with httpbody.WithMaxSize(ctx, 0): the
/machine/map long-poll and control/tsp's map session, whose messages are
already capped per-message (by readMapResponseMessage and decodeMsg, and
by tsp's framedReader and boundedReader). Their non-200 error bodies are
not message streams, so those are capped with LimitSizeTo instead.

The tailnet lock /tka/init/begin, /tka/sync/offer and /tka/affected-sigs
responses can carry per-node key signatures or missing AUMs, which at
100,000 peers reach tens of MB, so they raise the cap to 512 MiB. The
per-response io.LimitedReader decoders that silently truncated those
responses at 1 or 10 MiB are removed: the transport cap is now the
single enforcement point, and it reports oversize bodies instead of
truncating them.

The /key fetch over plain TLS switches from io.LimitReader to
httpbody.LimitSizeTo, so an oversized response reports the problem
instead of producing a confusing truncated-JSON error.

Thanks to Ben Carman for the report!

Updates tailscale/corp#48187

Reported-by: Ben Carman
Change-Id: Ibf95e1ab9e4f0d7ef8866e8c26e62ed2a514455a
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-14 18:58:41 -07:00
M. J. Fromberger 6b0c1de7a0 ipn/ipnlocal: update controlknobs when loading a netmap from cache (#20422)
When a controlclient receives non-keepalive netmap, it updates the controlknobs
based on the capability map of the self node.  We added a new knob based on
tailcfg.NodeAttrCacheNetworkMaps in be2f554dd3, and this ensures we correctly
propagate the attribute to the knob when coming up from a cache as well.

Updates #12639
Updates tailscale/projects#27

Change-Id: I40d34053c9743757f9fddca715379fc63a9ae6a0
Signed-off-by: M. J. Fromberger <fromberger@tailscale.com>
2026-09-14 15:57:00 -07:00
Brad Fitzpatrick ab4262680b ipn, ipn/ipnlocal: remove Notify.NetMap, size initial status to subscription
Nothing reads ipn.Notify.NetMap anymore. The previous commit removed
its runtime (non-initial) emission, and every first-party client is
also off the initial one: the Win32 and WinUI GUIs and the Apple
bridge subscribe with InitialStatus or InitialState plus peer deltas,
Android no longer uses it, and the remaining in-tree subscribers that set
NotifyInitialNetMap (sniproxy and the kube helpers) only did so to get
the initial Notify.SelfChange and discarded the netmap that tailscaled
built, encoded, and shipped for them.

Delete the Notify.NetMap field and the NotifyInitialNetMap bit. The
bit value stays reserved under the name ObsoleteNotifyInitialNetMap and
ValidateNotifyWatchOpt rejects subscriptions that set it, like the
NotifyRateLimit bit removed in the previous commit. NotifyNoNetMap
remains accepted as a no-op because shipping GUIs still set it.

The blessed way to seed a watcher's view is NotifyInitialStatus, but
it unconditionally built O(peers) status entries, which is exactly the
waste this series is deleting for watchers that only care about the
self node. Size the initial status to the subscription instead:
Status.Peer is only populated when the watcher also set
NotifyPeerChanges or NotifyPeerPatches, since only peer-delta
subscribers need a peer baseline to apply deltas to. That matches its
existing first-party users (containerboot and the WinUI GUI both pair
InitialStatus with peer bits).

Migrate sniproxy and the kube helpers to NotifyInitialStatus: they
seed from InitialStatus.Self and react to the (ungated) runtime
Notify.SelfChange messages after that, so their initial message now
carries one PeerStatus instead of a full netmap.

Also drop doc comment references to LocalClient.NetMap, a method that
does not exist; on-demand fetches go through other LocalAPI methods
such as LocalClient.Status.

Updates #12542

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I242992a744c0ffd0be6f27e8c735aa69d5b23b5e
2026-09-14 09:37:30 -07:00
Brad Fitzpatrick 1b0f7b3c6c ipn, ipn/ipnlocal: remove legacy runtime Notify.NetMap emission, rate limiting
The WinUI GUI was the last consumer of the legacy Notify.NetMap field
on runtime (non-initial) IPN bus messages. It was converted to peer
deltas in tailscale/corp#47958, and the GUI and tailscaled ship
together as a unit on Windows, so tailscaled no longer needs to build
and emit full netmaps on the bus on any platform. Remove
goosGetsLegacyNetmapNotify and the code it gated. The initial netmap
(NotifyInitialNetMap) is unaffected; NotifyNoNetMap is now a no-op but
remains accepted for compatibility.

With no runtime netmaps left to rate limit, the NotifyRateLimit
subscription bit is meaningless, so remove the rateLimitingBusSender
machinery too. The bit value stays reserved under the name
ObsoleteNotifyRateLimit and ValidateNotifyWatchOpt now rejects
subscriptions that set it; previously it was only rejected in
combination with new-style delta bits.

Updates #12542

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ic6184ea549726de9d9d56f133aae21e53f62cbc5
2026-09-14 09:37:30 -07:00
Brad Fitzpatrick 5c939a8560 util/dnsname,ipn/ipnlocal: reject hostile DNS names from control
Two fixes for DNS names from a malicious control server, from Ben
Carman's security review:

dnsname.ToFQDN now rejects names containing whitespace or control
characters. ToFQDN previously checked label lengths only, so a search
domain like "evil.com\nnameserver 6.6.6.6" passed validation and was
written verbatim into /etc/resolv.conf by the direct DNS manager, where
the injected line became a real nameserver. The same shape existed on
Windows, where a CRLF in an ExtraRecords name injected lines into the
hosts file. This is deliberately not RFC 1123 hostname validation (see
the existing comment about issue 2024): labels may still contain any
byte that isn't whitespace or a control character.

dnsConfigForNetmap now drops search domains and split-DNS route suffixes
that ToFQDN rejects. It previously logged the error but appended the
zero FQDN anyway, and FQDN.WithoutTrailingDot panics on the empty FQDN
when OS DNS config is written, so one over-long domain from control
crashed tailscaled on every netmap until control sent a valid one.

Thanks to Ben Carman for the report!

Updates tailscale/corp#48187

Reported-by: Ben Carman
Change-Id: I8de40cafacfb66e2b863f097794c42f3ca5c2da9
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-14 07:20:27 -07:00
Mike Jensen c9f5d175b5 ipn/localapi: require local admin for dev-set-state-store (#21222)
State keys written by dev-set-state-store are otherwise gated by their own handlers, for example serve-config requires a local admin before storing a _serve/<profile-id> key.

Writing such a key directly through dev-set-state-store skipped that check. Require`IsLocalAdmin` in the handler, the same check serve-config performs.

Credit to @johnnymiranda for reporting this issue.

Fixes tailscale/corp#47886

Change-Id: Ie82f961b016f78895793d801929a4fa11ebf7fd8

Signed-off-by: Mike Jensen <mikej@tailscale.com>
2026-09-11 14:07:53 -06:00
Adrian Dewhurst f31cfaec32 ipn/ipnlocal: add exceptions to 4via6 filter
Some customers rely on the ability for 4via6 subnet routers to expose
loopback and/or link-local addresses. In situations where the
administrator has deemed these to be safe, accept a list of allowed
addresses in an environment variable named TS_4VIA6_ALLOW_LOCAL.

Fixes #21019

Change-Id: I7b3fe863696ee67aa352c4bd11edcbc8836920e3
Signed-off-by: Adrian Dewhurst <adrian@tailscale.com>
2026-09-11 15:41:26 -04:00
Brad Fitzpatrick 5186c3f41c ipn/ipnlocal: only require peerapi masquerade Host match per address family
isAddressValid rejected all non-masquerade destination addresses
whenever any masquerade address was set for the peer. With a v6-only
masquerade pair, the peer's client still dials the v4 (native, not
masqueraded) peerapi URL, so every fresh peerapi connection got a 403.

The bug was masked by HTTP keep-alive: peerNode is snapshotted per
connection, so connections established before the masquerade was
configured kept validating against the old node view. It surfaced when
a newer tailscale/go toolchain started closing idle netstack
connections, forcing fresh peerapi connections and failing the
TestNATPing v6=true NAT subtests with 403s.

A masquerade address for one family says nothing about the other
family, so require a masquerade address match only for the address
family it applies to, and fall through to the self-addresses check for
the other family.

Fixes #21194

Change-Id: Idd2ed82b81e805b47b6aa6af4b5938984c6fa208
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-10 08:24:34 -07:00
Adrian Dewhurst a4c7902240 feature/conn25, ipn/ipnlocal: restrict conn25 DNS by peercap
Instead of opening all DNS queries with conn25, only permit clients to
access domains of the apps that they have permission to use. Previous
changes ensure that clients report which app they are making a DNS query
for, simplifying the check.

Updates tailscale/corp#40076
Updates tailscale/corp#47585

Change-Id: I40fee220b3f190b2a0db9b3e6cc79f323ef16b73
Signed-off-by: Adrian Dewhurst <adrian@tailscale.com>
2026-09-08 15:57:20 -04:00
Brad Fitzpatrick 0640312e51 ipn/ipnlocal: preserve peer deltas on expiry
Refresh the expiry timer netmap from the live peer state before
reinstalling it, preventing delta updates from being rolled back.

Updates tailscale/corp#47686

Change-Id: Idc738acea82bab5a8ba772084a41e55b38a06bcc
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-04 12:13:40 -07:00
Brad FitzpatrickandBrendan Creane 2ae2808b64 ipn/ipnlocal: don't evict another node's index entries on netmap deltas
When applying netmap deltas, nodeBackend evicted its index entries
(nodeByAddr, nodeByKey, nodeByWGString, nodeByStableID, nodeByName)
derived from a node's last-known value without checking that the entry
still pointed at that node. Control can reassign a churning ephemeral
peer's Tailscale IP (or MagicDNS name) to a newer peer and deliver the
new peer's upsert before the old peer's removal, either in an earlier
MapResponse or reordered within one batch by the NodeID sort in
netmap.MutationsFromMapResponse. The removal then wiped the new
owner's entry.

The peers map itself stayed correct in every ordering, so WireGuard
kept the peer and handshakes succeeded, but WhoIs lookups by IP failed
until the next full netmap rebuilt the indexes. On App Connectors that
surfaced as "peerapi: unknown peer" and refused DNS connections from
affected clients, with a toggle of Tailscale (forcing a full netmap)
as the only recovery.

Make every index eviction conditional on the entry still mapping to
the node being removed or replaced, and add a regression test covering
the cross-batch, intra-batch, and upsert-eviction orderings.

Also add an end-to-end test in tstest/integration showing that a
MapResponse reusing an address is handled incrementally rather than as
a full netmap, and that LocalBackend.WhoIs still resolves the reused
address afterwards, which is the lookup PeerAPI makes before it
accepts a connection.

Updates tailscale/corp#47435

Co-authored-by: Brendan Creane <bcreane@gmail.com>
Signed-off-by: Brendan Creane <bcreane@gmail.com>
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I3f8c2a9d41e07b6a5cd2e94f78b013c6ad2f5e91
2026-09-04 12:13:31 -07:00
Brad Fitzpatrick 31d8badb3b wgengine: support per-peer WireGuard PSKs
Replace the allowed-IP-only peer callback result with wgcfg.PeerConfig.
It carries allowed IPs and an optional pre-shared key through lazy peer
creation and active peer synchronization.

Update wireguard-go for the new peer PSK APIs.

Updates tailscale/tailcat#84

Change-Id: Iacd9d2c74b0b64d690f3cbdf93918686ac6076d7
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-09-03 20:04:09 -07:00
Claus Lensbøl f53c28101a control/controlclient,ipn/ipnlocal,wgengine/{,magicsock}: route disco keys learned over TSMP directly to magicsock (#21016)
Previously, the new disco keys entering the client via TSMP was routed
into controlClient to allow for deduplication and filtering of disco
keys and avoid control overwriting an active TSMP learned key with a
stale key.

A system supporting multiple disco keys was introduced to allow these
keys living side by side, along side a system for selecting an active
egress key with a bias towards keys learned via TSMP. On ingress any
known key is accepted.

This PR makes the switch to key routing, by sending new TSMP learned
disco keys directly into the userspace engine and in turn magicsock,
removing the need for a full layer of deduplication in the
controlClient.

Additionally, the system that previously fully reconfigured clients on
disco key updates coming from control is now using an optimistic
handshake througha newly introduced wireguard-go method,
ScheduleHandshakeOnUserSend implemented in:
https://github.com/tailscale/wireguard-go/pull/81

What this PR does not do is revert the mapSession back to being single
writer. Bringing the mapSession back to this state is desired, however
to ensure the plumbing is done right and easier to reason about, that
will be done in a separate PR.

Updates #20590

Signed-off-by: Claus Lensbøl <claus@tailscale.com>
2026-09-03 09:38:28 -04:00
Brad Fitzpatrick 1293b4f67b wgengine/router/osrouter: add native FreeBSD routing, no-snat support
And add site-to-site natlab vmtests for Linux + FreeBSD.

Fixes #5573
Updates #18897

Change-Id: I9629f46fdd99137f0ae96431487d3a7410e2a4cc
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-08-31 20:31:01 -07:00
James Tucker 57c3357fdb cmd/tailscale: allow IPv6 localhost serve targets
Use net.JoinHostPort when expanding proxy targets so IPv6 loopback addresses retain the required brackets. Accept ::1 as an HTTP and TCP destination in both current and legacy serve implementations.

Fixes #8702

Signed-off-by: James Tucker <jftucker@gmail.com>
2026-08-31 14:25:58 -07:00
Simon Law 0e84b4a3a0 tailcfg: replace int with DERPRegionID for additional type safety (#20646)
Historically, when DERP regions were switched away from strings to
numeric identifiers in PR #14641, tailcfg.Node.HomeDERP was declared
as an int instead of its own type.

This PR declares a new tailcfg.DERPRegionID type, represented by an
int64, and converts the following fields to use this type:

- netcheck.Report.PreferredDERP
- netcheck.Report.RegionLatency
- netcheck.Report.RegionV4Latency
- netcheck.Report.RegionV6Latency
- tailcfg.DERPHomeParams.RegionScore
- tailcfg.DERPMap.Regions
- tailcfg.DERPNode.RegionID
- tailcfg.DERPRegion.RegionID
- tailcfg.NetInfo.PreferredDERP
- tailcfg.Node.HomeDERP
- tailcfg.PeerChange.DERPRegion
- tailcfg.PingResponse.DERPRegionID

Note that the original field was an int, while the new field is backed
by an int64. This change makes DERPRegionID the same size on both
32-bit and 64-bit architectures.

Fixes: #20165

Change-Id: Ic6f795a6d791dd16f756f246d5a02085443e212f

Signed-off-by: Simon Law <sfllaw@tailscale.com>
2026-08-19 16:25:50 -07:00
M. J. Fromberger f3552c29c0 ipn/ipnlocal: fix cache update for peers deleted by netmap deltas (#20851)
After a netmap delta is applied, we scan the mutations for affected peers and
update the cache (if enabled) for those peers. For removals in particular, we
were relying on the node backend to resolve node IDs (provided by the delta
mutation) to stable IDs.

Prior to 65fd320a this happened to work because the node backend would hold on
to all the peers mentioned by the previous full netmap, even after applying
deltas. But that was essentially accidental, and once we fixed it not to do
that, these lookups no longer worked. We need the stable ID, since that is how
the cache is keyed, and now that they're no longer pinned, we were not properly
evicting removed peers from the cache.

To fix this, capture removed peer stable IDs while applying mutations to the
node backend, instead of trying to look them up afterward.

Updates #20796

Change-Id: I14ded78eaf9657645f0869a52460fd3cd86edba6
Signed-off-by: M. J. Fromberger <fromberger@tailscale.com>
2026-08-19 10:31:38 -07:00
Mike Jensen 90ed0bcf4b wgengine/netstack: block 4via6 forwards to host-scoped targets (#20866)
The packet filter only sees a via address's outer ULA, so netstack unconditionally dialed whatever IPv4 was embedded in it.

This change refuses TCP, UDP, and ping relays to host-scoped destinations after UnmapVia.

Credit to the Anthropic infrastructure security team for finding and reporting.

Fixes https://github.com/tailscale/corp/issues/46646

Change-Id: I93d8e27a2eddbce7eb3f8c5fa4677f8de3a8ed9e

Signed-off-by: Mike Jensen <mikej@tailscale.com>
2026-08-19 09:16:08 -06:00
Alex Chan 468241a13f envknob: add SetenvForTest helper for safe environment variable handling
Our tests are inconsistent in how they set environment variables and cleanup.
Missing cleanup logic can leak environment variable state across tests and
change the behaviour of subsequent tests.

This recently caused flakes in feature/acme, where `TestGetCertPEMWithValidity`
leaked `TS_CERT_SHARE_MODE` and `TestAsyncRenewalDedup` to fail inconsistently.

To fix the immediate flake and prevent future leaks, introduce a `SetenvForTest`
helper that handles setting and cleaning up environment variables in tests.

Fixes #20902

Change-Id: I9f8ef45ec875d66034cfd0054311bd69b67e2243
Signed-off-by: Alex Chan <alexc@tailscale.com>
2026-08-19 12:22:26 +01:00
kari-ts d3d14a8831 ipn/{desktop,ipnlocal}: register per-user policy stores on session lifecycle (#20293)
This registers per-user RSOP policy stores when a user logs into a Windows session and ensures that per-user poicy settings are delivered to clients via IPN bus.

The store lifecycle is tied to Windows session lifetime via desktopSessionExt's SessionInitCallback, which handles unattended mode (per-user policies stay enforced after gui disconnection) and multi-user (refcounted across sessions for the same user).

Updates tailscale/corp#42259

Signed-off-by: kari-ts <kari@tailscale.com>
2026-08-17 19:59:42 -07:00
Amal Bansode fd07b9a2b6 net/traffic: fix rendezvous hashing (#20828)
The rendezvous hasher for traffic steering loadbalancing was flawed.
By plainly using the FNV-1a hash value, the result often reflected the
magnitude of the most significant bits in the hash seed, meaning the
hash function was not diffusive (aka missing the Avalanche Effect).

Popular wisdom seems to be that the output of FNV-1a should be mixed
with some large numbers to perturb more output bits. Borrow concepts
from other (Rust, Java) libraries by using the mix13 variant of 64-bit
finalizers by David Stafford.

Modify the fuzz test that asserts this fairness. Adjust a few
constants like client count and candidate count to more closely
reflect real-world scenarios and practical probabilities. Tighten
the bounds for distribution from 50% to +-20%.

Updates tailscale/corp#46471

Signed-off-by: Amal Bansode <amal@tailscale.com>
2026-08-12 14:13:42 -07:00
M. J. Fromberger 93912bbe1d ipn/ipnlocal/netmapcache: invalidate digest caches when removing objects (#20839)
Previously, a delta update that drops a peer would not invalidate the digest
cache for that peer after deleting its cache entry. If a subsequent (later)
delta re-adds that peer with the same content, the digest cache would prevent
us from updating the persistent entry.  Add a test to exercise this, and fix
the bug.

Updates #20795

Change-Id: I4a9ef03e8787a330d6395629251ee61ca2221fdd
Signed-off-by: M. J. Fromberger <fromberger@tailscale.com>
2026-08-12 09:49:18 -07:00
Nick Khyl 1ec348784b ipn/ipnlocal,net/dns/resolver: prevent bare name resolution when MagicDNS is disabled
nodeBackend.nodeByName should always contain both FQDNs and short names,
as it is used in different contexts, including UserDial DNS resolution, which should
be able to resolve unqualified DNS names regardless of the MagicDNS state.

However, net/dns/resolver.Resolver and, by extension, MagicDNSHosts
implementations should only resolve fully qualified domain names,
skipping short names when MagicDNS is disabled for the tailnet.

This fixes it in (*nodeBackend).nodeByFQDNLocked, which is only used
in the MagicDNS paths, and updates the tests.

Fixes #20789

Signed-off-by: Nick Khyl <nickk@tailscale.com>
2026-08-10 17:54:32 -05:00
Simon Law 00699abdfb tailcfg,tailcfg/{nodecap,selfcap}: split capability constants to their own packages (#20639)
Package tailcfg defines the types and constants used by the Tailscale
protocol, but since everything is all in one package, it’s difficult
to sift through the docs: https://pkg.go.dev/tailscale.com/tailcfg

We define and enumerate capabilities as string constants for
tailcfg.NodeCapability and tailcfg.PeerCapability. This PR extracts
them into their own packages:

- tailcfg.CapabilityFileSharing becomes nodecap.FileSharing
- tailcfg.NodeAttrOnlyTCP443 becomes nodecap.OnlyTCP443
- tailcfg.PeerCapabilityTaildrive becomes peercap.Taildrive

We originally intended for CapabilityFoo to grant an entitlement or
permission for Foo, and for NodeAttrBar to configure Bar in the
nodeAttrs section of the policy file. However, there was no technical
enforcement of this convention, so new capabilities have used the
NodeAttr prefix regardless of meaning. Therefore, this PR unifies
tailcfg.CapabilityFoo and tailcfg.NodeAttrBar into a single package as
nodecap.Foo and nodecap.Bar.

Ran `go fix -inline ./...` and committed the changes that replaced
uses of the tailcfg aliases with the authoritative ones.

Updates #20259

Change-Id: Ieb7e7e6c8247c39faf42fdf15c68cdc7c621c730
Signed-off-by: Simon Law <sfllaw@tailscale.com>
2026-08-07 16:30:35 -07:00
Will HannahandJonathan Nobels 15015f19bf net/dns: scope quad-100 on macOS so DoH profiles aren't shadowed (#20775)
* net/dns: scope quad-100 on macOS so DoH profiles aren't shadowed

On sandboxed macOS, an uncovered control ExtraRecord forced quad-100 to
be the primary resolver, proxying all public DNS and shadowing a user's
DoH system profile. Scope quad-100 to its match domains instead, adding
the uncovered host records to MatchDomains so they still resolve while
public names fall through to the OS resolver. quad-100 remains primary
only without a usable base resolver or with non-enumerable MagicDNS
host records.

fixes tailscale/corp#45534

Signed-off-by: Will Hannah <willh@tailscale.com>

* net/dns: move scoped DNS behind an envknob

updates tailscale/corp#45534

Given the sensitivity of this change, let's stuff it behind
a control knob for a release.

Signed-off-by: Jonathan Nobels <jonathan@tailscale.com>

---------

Signed-off-by: Will Hannah <willh@tailscale.com>
Signed-off-by: Jonathan Nobels <jonathan@tailscale.com>
Co-authored-by: Jonathan Nobels <jonathan@tailscale.com>
2026-08-07 15:27:22 -04:00
Brendan Creane 44ec3a1b89 Revert "net/dns: scope quad-100 on macOS so DoH profiles aren't shadowed (#20603)" (#20771)
This reverts commit 7e01825e51.

The change was merged to main accidentally. Reverting so it can go
back through review before landing again.

updates tailscale/corp#45534

Signed-off-by: Brendan Creane <bcreane@gmail.com>
2026-08-06 16:33:23 -07:00
Will HannahandJonathan Nobels 7e01825e51 net/dns: scope quad-100 on macOS so DoH profiles aren't shadowed (#20603)
* net/dns: scope quad-100 on macOS so DoH profiles aren't shadowed

On sandboxed macOS, an uncovered control ExtraRecord forced quad-100 to
be the primary resolver, proxying all public DNS and shadowing a user's
DoH system profile. Scope quad-100 to its match domains instead, adding
the uncovered host records to MatchDomains so they still resolve while
public names fall through to the OS resolver. quad-100 remains primary
only without a usable base resolver or with non-enumerable MagicDNS
host records.

fixes tailscale/corp#45534

Signed-off-by: Will Hannah <willh@tailscale.com>

* net/dns: move scoped DNS behind an envknob

updates tailscale/corp#45534

Given the sensitivity of this change, let's stuff it behind
a control knob for a release.

Signed-off-by: Jonathan Nobels <jonathan@tailscale.com>

---------

Signed-off-by: Will Hannah <willh@tailscale.com>
Signed-off-by: Jonathan Nobels <jonathan@tailscale.com>
Co-authored-by: Jonathan Nobels <jonathan@tailscale.com>
2026-08-06 13:01:16 -07:00
Evan Lowry 2b8a980f1f ipn/ipnlocal: deduplicate RoutableIPs / RequestTags in hostinfo (#20759)
Both RoutableIPs and RequestTags act as a set of values. There have been
cases of misconfigured IAC populating duplicate values, resulting in
more data being sent to control than needs to be.

Updates tailscale/corp#44607

Signed-off-by: Evan Lowry <evan@tailscale.com>
2026-08-05 12:28:47 -03:00
Brad Fitzpatrick 0dfe672b32 ipn/ipnlocal: allow the ingress peer capability for unsigned peers (#20745)
0eb38dc2e (#20561) made peer capability resolution return nothing for
peers with UnsignedPeerAPIOnly set, so that a possibly malicious
control server can't grant capabilities to peers outside the tailnet
lock authority. But Tailscale Funnel ingress nodes are unsigned by
design, and control intentionally grants them
PeerCapabilityIngress, which the peerapi /v0/ingress handler requires.
The result was that every Funnel connection was rejected with a 403
"denied; no ingress cap".

Instead of denying all capabilities to unsigned peers, allowlist
PeerCapabilityIngress specifically. It only permits ingress requests
over the PeerAPI, which unsigned peers can already reach, and the
node only serves them for targets explicitly configured for Funnel.

The tsnet TestFunnel didn't catch the regression because its fake
ingress peer was a normal signed peer. Teach testcontrol to mark a
node as UnsignedPeerAPIOnly (excluding such nodes from traffic-
permitting filter rules, as real control does, so clients don't
discard the packet filter) and make TestFunnel use it so the test
now exercises the same capability checks as production Funnel
traffic.

Updates tailscale/corp#46053
Updates #20739


Change-Id: I3f6b8e2a94d1c07b5a2e9d84f16c30aa79e5d21b

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-08-04 08:43:45 -07:00
Brad Fitzpatrick 4c4d1c35f8 ipn/ipnlocal: avoid deadlocks during shutdown
LocalBackend.Shutdown waits for the ACME refresh loop and active SSH
sessions. Both can be blocked acquiring LocalBackend.mu, so waiting while
holding that mutex deadlocks shutdown.

Detach the SSH server under the mutex, then stop both subsystems after
releasing it. Prevent their work from restarting once shutdown begins, and
serialize repeated Shutdown calls with sync.Once.

Add regression tests that verify subsystem shutdown runs without
LocalBackend.mu held.

Updates tailscale/corp#45964

Change-Id: I37ead4f26fbfb5703a83882668d98a8862ba7d67
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-08-01 13:57:38 -07:00
loowr d645114dd4 ipnauth: set isUnixSock in the ts_omit_unixsocketidentity GetConnIdentity
The ts_omit_unixsocketidentity variant of GetConnIdentity never
performs the *net.UnixConn type assertion, so ConnIdentity.isUnixSock
stays false. ipnserver's Permissions only grants local API access to
unix socket connections, so with this build tag every local API request
is denied read and write access ("status access denied") and the CLI
cannot talk to the daemon at all in --extra-small/--min builds.

Mirror the type assertion from the peercred variant so the omitted
identity build behaves as intended (everyone is an admin when unix
socket identities are compiled out).

Signed-off-by: loowr <loowr@proton.me>
2026-07-31 17:30:11 -07:00
Brad Fitzpatrick 7eeb62415e go.mod: bump staticcheck in prep for Go 1.27, address fallout
Go 1.27 requires this new v0.8.0-rc.1.

But staticcheck 0.8's SA4023 gets stricter and points out that
modifiedExternallyError and handleListenersAccept always return
non-nil errors, and that MonitorHealth's callers don't need a separate
nil check before errors.Is. Simplify all three call sites; no behavior
change.

But then a handful of other places that SA4023 is angry about are
wrong (because it's not considering build tags) and can't be addressed
by ignore directives (again not considering build tags), so we just
disable SA4023 for now.

Updates #20220

Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I2fefe3b986b5798c2e01624a0e9820839d21a569
2026-07-31 08:00:40 -07:00
Naman Sood 78e7e89be6 ipn/ipnlocal: fix comment typo
This function used to have the suffix LockedOnEntry, but no longer has
that since #17804.

Updates #cleanup

Signed-off-by: Naman Sood <mail@nsood.in>
2026-07-30 14:16:53 -04:00
Michael Ben-Ami 0aba4d9460 feature/conn25,ipn/ipnlocal: remove UseWIPCode() guards
Removing this check installs conn25 instance-level hooks whenever the
feature is built in. init-level hooks were always installed, but used
this guard to exit early.

Now all hooks, both instance-level and init-level rely on netmap
configuration (populated tailscale.com/app-connectors-experimental node
attribute, with non-empty apps) to not exit early.

The conn25-shutoff feature flag tells control not to send that node
attribute.

Fixes tailscale/corp#39033

Signed-off-by: Michael Ben-Ami <mzb@tailscale.com>
2026-07-30 12:14:49 -04:00
Claus Lensbøl 3799eaf264 ipn/ipnlocal,wgengine: implement wg-go SetPriorityMessageOnEstablishmentFunc (#20606)
Implements a way to send TSMPDiscoAdverts based on a trigger from
wireguard-go when a rekey happens. This lets us distribute disco keys
consistently, but also sets us up for a minimal message that can be
distributed to other clients.

A benchmark is implemented to make it easier to keep the call cheap and
to avoid locking up anything in wireguard-go.

Updates #20081

Signed-off-by: Claus Lensbøl <claus@tailscale.com>
2026-07-30 09:20:43 -04:00
Simon Law bb498a8393 net/netcheck: extract Compare function for region latencies (#20658)
There are two places in the code where we need to compare the
latencies of a netcheck report. In both cases, the comparison wasn’t
well tested.

This PR extracts that logic into a Compare method of the new
RegionLatency type. This type wraps the map of latency measurements
keyed by region ID.

Updates #cleanup

Change-Id: I7f248988973007c2f452283c1b82f84b03068f77

Signed-off-by: Simon Law <sfllaw@tailscale.com>
2026-07-29 16:51:55 -04:00
Brad Fitzpatrick c0c453334a ipn/ipnlocal: evict stale node indexes when a delta upsert replaces a peer
When a node is renamed in the admin console, control sends peers a
single MapResponse delta: a PeersChanged entry carrying the full
updated node with its new Name, and no new DNSConfig (MagicDNS
records are computed client-side from peer names). That arrives as a
NodeMutationUpsert, but nodeBackend's upsert path only added the new
node's index entries and never removed the replaced node's, so
nodeByName retained the old name, and nodeByAddr, nodeByKey, and
nodeByStableID could likewise go stale if those fields changed.

Since 7e609b258 the quad-100 resolver serves MagicDNS answers on
demand from those live indexes, so a renamed peer's old name kept
resolving until something rebuilt the indexes from a full netmap,
such as toggling Tailscale off and on.

Evict the replaced node's index entries before adding the new ones.
Also consolidate the natlab DNS coverage into a single TestMagicDNS
that boots one VM and exercises extra records, search domains, and
peer add/rename/remove end to end, injecting the same MapResponse
shapes that production control sends.

Updates tailscale/corp#45631

Change-Id: I8a418317d930ec8ce112f7bd19bfd5778117a65e
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
2026-07-27 19:43:11 -07:00
chaosinthecrd 97a75c837d cmd/k8s-operator,ipn/store/kubestore,kube/kubetypes: share ACME account key per tailnet
Introduce a per-tailnet shared ACME account key so that all ingress
ProxyGroup replicas on a tailnet present the same account identity to
Let's Encrypt. This lets renewals claim the ARI "replaces" exemption
from the 50-certs-per-week rate limit, surviving Pod restarts,
ProxyGroup recreation, and cluster migrations.

The operator provisions a "tailscale-acme-accounts" Secret in its
namespace, guarded by a finalizer and a deletion warning event, and
watched so it is recreated promptly if removed. Proxies migrate any
pre-existing per-pod key into the shared Secret on first boot, adopt
the shared key on subsequent boots, and restore it on cert writes if
the Secret was recreated empty. Certs are stamped with the fingerprint
of the issuing account so renewals skip the "replaces" claim when the
account doesn't match.

Opt-in per-ProxyGroup via the tailscale.com/share-acme-account
annotation, or operator-wide via OPERATOR_SHARED_ACME_ACCOUNT_KEY.

Updates #18251
Updates #20288

Signed-off-by: chaosinthecrd <tom@tmlabs.co.uk>
2026-07-27 12:06:41 +01:00
Michael Ben-Ami b062abb1ea appc,ipn/ipnlocal: install conn25 DNS routes when using exit node
We were early-returning when the node was using an exit node, before
Connectors 2025 split DNS routes were calculated and installed.

Now we assemble the routes first, then install them in both exit node
and non-exit-node contexts. The returned resolvers set UseWithExitNode
to true even though as of today, we believe they should be installed in
all cases without regard to that boolean value. With the boolean, we
preserve the flexibility to toggle behavior without touching ipnlocal.

We also add a TODO to turn the extra split DNS route gathering into a
feature hook (tailscale/corp#37125).

This does not affect appc connectors, which receive split DNS routes,
and the UseWithExitNode value directly from control.

Updates #16384

Signed-off-by: Michael Ben-Ami <mzb@tailscale.com>
2026-07-24 09:39:40 -04:00