Contining the trend of making everything be optional that could
possibly be optional, this makes the magicsock UDP underlay transport
be optional. (mostly replacing a bunch of GOOS != "js" checks in the process)
Omitting udptransport (ts_omit_udptransport) makes magicsock DERP-only: no
UDP sockets, no advertised endpoints, and netcheck measures DERP latency
over HTTPS. Omitting only nattraversal keeps disco ping/pong to peers'
advertised endpoints but drops STUN, call-me-maybe, the peer relay client,
the endpoint tracker, and UDP lifetime probing.
Both are plain linker dead-code elimination via buildfeatures
constants. (as oppposed to moving the code all over into feature
packages and indirecting through hooks) The minimal linux/amd64
tailscaled shrinks by 258 KB without nattraversal and 561 KB without
both.
Updates #12614
Change-Id: Ife10804b48a168d9481e4321c18ed8fc85fa71f1
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
After copying from the backend to the client until EOF, forwardTCP shut
down the backend socket's read side. That does nothing on the wire, and
once the backend's FIN has arrived, macOS rejects shutdown(SHUT_RD) with
ENOTCONN (Linux does too once the socket reaches TIME_WAIT), so every
forwarded connection logged "backend -> client close connection: ...
socket is not connected". Keep only the CloseWrite calls, which are what
propagate the half-close.
The regression was introduced by 04d24cdbd4 (#16462), merged 2025-07-07.
Tested on Linux and on macOS 27 (arm64): the wgengine/netstack tests
pass on both. The same ENOTCONN failure was reproduced on macOS via the
identical pattern in net/socks5 (#21522).
Updates #21522
Updates #16462
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I3e9b1d7f4a2c8e6b0f5d9a3c7e1b4f8d2a6c0e5b
FreeBSD handles subnet routes in netstack by default, but the router
still installed its pf NAT rules whenever SNATSubnetRoutes was set,
which is the default. Every FreeBSD node, subnet router or not, loaded
and enabled pf, inserted anchor references into the host's main
ruleset and loaded NAT rules that netstack never needed, since
netstack dials subnet destinations from the host's own addresses.
It also flipped the forwarding sysctls for any advertised route.
Only do either when the kernel path is opted into with
TS_DEBUG_NETSTACK_SUBNETS=false and routes are advertised.
TestSubnetRouterFreeBSDManyFlows ran in netstack mode since the default
changed, so it no longer exercised pf at all. Opt it into the kernel
path and assert the anchor holds NAT rules, and have
TestSubnetRouterFreeBSD assert the default mode leaves pf untouched.
Also drop the stale claim in handleSubnetsInNetstack that the pf NAT
rule never matches; that was the (self) pool bug, since fixed.
Updates #21450
Change-Id: Ibf91a676a026cebb29f64e39a64fe8e75488360b
Signed-off-by: Martin Minkus <martin.minkus@sonic.com>
On Windows, we used to create all routes using the interface IP address as the next hop.
Notably, on Windows 8 and later, the system normalizes that next hop to 0.0.0.0,
resulting in an on-link route.
In #12847, we stopped creating on-link subnet routes because doing so made
the last IP address in the range unreachable. However, we missed a related issue.
As a result, we continued specifying the interface IP address as the next hop
for Tailscale IP routes.
Because Windows normalizes those routes to use an unspecified next hop,
the routes we read back from the system did not match the routes we expected
to see. This caused unnecessary churn, with the same routes being repeatedly
deleted and recreated.
In this PR, we start creating on-link routes with a normalized, unspecified (all-zeroes)
next hop, as required by the API. This ensures that the routes we read back match
the desired routes, preventing unnecessary churn.
https://web.archive.org/web/20260924154040/https://learn.microsoft.com/en-us/windows-hardware/drivers/network/mib-ipforward-row2Fixes#21438
Reported-by: Caleb Crome <caleb.crome@flyzipline.com>
Signed-off-by: Nick Khyl <nickk@tailscale.com>
Returning from injectToHost and injectToWireGuard leaves the two
goroutines dead and the host without having a way to send traffic.
This is especially relevant for android where the tundev is torn down
and recreated for every call to updateTUN() from a route table change or
roaming between networks.
Fixes#21155
Signed-off-by: Claus Lensbøl <claus@tailscale.com>
If the network comes up during the few milliseconds NewLocalBackend
takes, tailscaled starts with its control client paused and never
unpauses: the interface state snapshot was taken before the netmon
subscription (kept late since #17252), so a change published in between
reached nobody. #21261 hit it on fast-booting NixOS microVMs, and it was
behind the natlab TestEasyEasy CI flake, where a gokrazy node's DHCP
lease landed in that window and "tailscale up" hung at "awaiting
unpause".
Re-read netmon's state after subscribing rather than subscribing before
the snapshot (as #21281 proposed), which would bring back the #17252
data race and let a fresh delta be overwritten by the stale snapshot.
netmon.New and the userspace engine had the same shape of gap between
their snapshot and their subscription; close those too.
Under CPU pressure the hang reproduced in 1 of 22 TestEasyEasy runs
before and 0 of 120 after.
Thanks to @c-vigo for the detailed diagnosis in #21261 and to @zzz-yu
for pointing out this problem and proposing a fix in #21281!
Fixes#21261
Updates #14902
Updates #19126
Updates #deflake
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ifa21f8055a85321a5afda7800140fc5045c6ce08
We have multiple situations where we have a `net.Conn` representing a
TCP connection for a proxy and we need access to the underlying
`CloseWrite()` and `CloseRead()` functions to properly pass through
half-closes (see #16462, #20883). Add an interface we can cast to in
order to get access to these functions.
Updates #20883.
Signed-off-by: Naman Sood <mail@nsood.in>
Update gVisor to include its fix for RACK loss detection with coarse
monotonic clocks. Configure netstack with the 500 microsecond clock
resolution used on Windows so RACK accounts for timestamp quantization.
Remove the TCP recovery override that disabled RACK, enabling gVisor's
default RACK behavior on all platforms.
Switch netstack to cubic congestion control. The int overflow in CUBIC
sender cwnd arithmetic that required pinning reno has since been reworked
upstream into float arithmetic with RFC 9438 target clamping.
Align the natlab vnet stack with netstack: enable cubic, and drop the
now-redundant explicit SACK and receive-buffer moderation sets, both of
which are gVisor defaults.
Fixes#9707
Signed-off-by: James Tucker <jftucker@gmail.com>
A tailnet peer can remotely crash tailscaled by sending a DERP-sealed
disco CallMeMaybeVia message with an all-zero ServerDisco key. The
decoder accepts the zero key, and the relay manager later hands it to
DiscoPrivate.Shared, which panics on zero keys. The sender only needs
to be a relay-capable peer in the victim's netmap.
Auditing the other DiscoPrivate.Shared call sites reachable from
decoded messages turned up the same bug on the relay server side.
AllocateUDPRelayEndpointRequest.ClientDisco is attacker-chosen: one
slot must match the sender's disco key, and the other can be zero. It
flows unchecked into udprelay.Server.AllocateEndpoint, which calls
Shared on both client keys and panics in its eventbus subscriber
goroutine. AllocateEndpoint now rejects zero client keys with an error,
which its only caller already handles by logging.
Thanks to Ben Carman for the report!
Updates tailscale/corp#48187
Reported-by: Ben Carman
Change-Id: Ifc0f64d8f63270100b22c06e9f759dce624ab811
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Route packets to dedicated WireGuard, host, and loopback queues when
gVisor writes them. Drain each queue in a separate goroutine, preparing
the outbound path for batching and multiqueue WireGuard delivery.
Updates tailscale/corp#37878
Signed-off-by: Jordan Whited <jordan@tailscale.com>
Adds an opt-in, in-memory aggregator of recent connection-rejection
events (TSMP rejects received from peers, outbound TSMP rejects we emit
on ACL-blocked inbound flows, and pendopen timeouts) keyed by
(direction, proto, peer-address, reason). The aggregated data is exposed
over a new debug-rejects LocalAPI endpoint and a GET /debug/rejects c2n
endpoint, intended for future GUI/CLI consumption when diagnosing why a
connection failed.
Architecture:
- net/connreject holds the data types and a per-LocalBackend
Aggregator (LRU-bounded, default 256 entries on desktop / 32 on
mobile, per direction).
- feature/connreject is a self-registering ipnext.Extension that owns
one Aggregator per LocalBackend, installs note callbacks on the
tundev and engine, subscribes to OnSelfChange to flip the runtime
gate, and serves the LocalAPI/c2n endpoints.
- wgengine.Engine and *tstun.Wrapper each gain a SetConnRejectNote
setter; data-plane sites use a single atomic.Pointer load + nil
check, so the cost when no consumer is installed is one MOV.
Gating:
- Compile-time: ts_omit_connreject build tag (standard
feature/buildfeatures + condregister plumbing). Trims ~41 KB.
- Runtime: nodecap.ConnReject node attribute, off by default
at the control plane. May be removed once the feature is enabled
by default.
Updates CapabilityVersion to 146 (clients understand nodecap.ConnReject
and can serve GET /debug/rejects).
Adds Proto/Src/Dst accessors on flowtrack.Tuple (used by pendopen to
construct events without exposing the tuple's internals to the
aggregator).
Updates #1094
Updates #14802
Change-Id: I83e8f24a7e66fa2d158d128bd25fbe851134941b
Signed-off-by: James Tucker <james@tailscale.com>
This commit bumps the wireguard-go dependency to incorporate changes to
the packet memory model and the tun.Device.Read and conn.ReceiveFunc I/O
interfaces. It updates their implementations accordingly.
These changes improve throughput in all measured benchmarks and reduce
peak RSS in six of eight cases. The two regressions will be addressed in
a follow-up commit that reduces peak RSS below the baseline measured at
1e69418. That work is kept separate to simplify review.
The following throughput and peak RSS benchmarks were performed with
iperf3 between two Intel i5-12400 nodes running Ubuntu 24.04 (Linux 6.8).
The UDP benchmarks did not use UDP GSO on the sender, so they were
roughly equivalent to single packet I/O through wireguard-go.
TCP/1 signifies one TCP stream; TCP/128 signifies 128 parallel TCP
streams.
Throughput (Mb/s)
Test 1e69418 After Change
TCP/1 10,371 11,354 +9.5%
TCP/128 7,886 8,404 +6.6%
UDP/1 2,111 2,853 +35.1%
UDP/128 1,747 2,235 +28.0%
Peak memory (VmHWM, kB)
Test Side 1e69418 After Change
TCP/1 TX 98,240 52,596 -46.5%
RX 287,748 73,384 -74.5%
TCP/128 TX 101,196 52,812 -47.8%
RX 290,420 63,620 -78.1%
UDP/1 TX 58,864 160,840 +173.2%
RX 137,516 49,900 -63.7%
UDP/128 TX 66,148 116,096 +75.5%
RX 154,384 56,556 -63.4%
Updates tailscale/corp#46716
Updates tailscale/corp#22467
Updates tailscale/corp#36989
Updates tailscale/corp#37878
Signed-off-by: Jordan Whited <jordan@tailscale.com>
Upates to controlClient learned keys was logged as coming from TSMP.
Also, nil changes on nil keys were not filtered.
Updates #20494
Signed-off-by: Claus Lensbøl <claus@tailscale.com>
Replace the allowed-IP-only peer callback result with wgcfg.PeerConfig.
It carries allowed IPs and an optional pre-shared key through lazy peer
creation and active peer synchronization.
Update wireguard-go for the new peer PSK APIs.
Updates tailscale/tailcat#84
Change-Id: Iacd9d2c74b0b64d690f3cbdf93918686ac6076d7
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
freebsdRouter.Close ran "ifconfig tailscale0 destroy" while this process
still holds the tun open: the engine closes the router before the tun
device, and FreeBSD's tun(4) makes SIOCIFDESTROY sleep until the last
descriptor closes. tailscaled therefore deadlocks against its own child
ifconfig on every shutdown, and "service tailscaled restart" hangs
forever:
51709 - I /usr/local/bin/tailscaled -port 41641 -tun tailscale0 ...
52455 - I ifconfig tailscale0 destroy
52403 1 I+ /bin/sh /usr/local/etc/rc.d/tailscaled restart
Do only the PF cleanup at Close time. The interface left behind at
process exit is destroyed by the next startup's cleanUp hook
(router.HookCleanUp), which exists for exactly that and runs when
nothing holds the device.
Updates #5573
Change-Id: I7b739d948718bfb460240d2c0dd74658fc40ac82
Signed-off-by: Martin Minkus <martin.minkus@sonic.com>
Previously, the new disco keys entering the client via TSMP was routed
into controlClient to allow for deduplication and filtering of disco
keys and avoid control overwriting an active TSMP learned key with a
stale key.
A system supporting multiple disco keys was introduced to allow these
keys living side by side, along side a system for selecting an active
egress key with a bias towards keys learned via TSMP. On ingress any
known key is accepted.
This PR makes the switch to key routing, by sending new TSMP learned
disco keys directly into the userspace engine and in turn magicsock,
removing the need for a full layer of deduplication in the
controlClient.
Additionally, the system that previously fully reconfigured clients on
disco key updates coming from control is now using an optimistic
handshake througha newly introduced wireguard-go method,
ScheduleHandshakeOnUserSend implemented in:
https://github.com/tailscale/wireguard-go/pull/81
What this PR does not do is revert the mapSession back to being single
writer. Bringing the mapSession back to this state is desired, however
to ensure the plumbing is done right and easier to reason about, that
will be done in a separate PR.
Updates #20590
Signed-off-by: Claus Lensbøl <claus@tailscale.com>
Prior to this change there were two problems with our fuzzing for oss-fuzz:
1. There was an issue if the fuzzing spanned two files (mingled with the testing).
2. The fuzzing needs to be part of the implementation package (no _test packages).
This change fixes that by moving all package fuzzing into a common `fuzz_test.go` within the package.
Updates https://github.com/tailscale/corp/issues/46608
Change-Id: I0b95edcd0df946f723eea32f575c679214c0b202
Signed-off-by: Mike Jensen <mikej@tailscale.com>
When receiving disco traffic from a node, mark that node as having been
seen. Should that disco key not be the active use disco key, switch to
that one as being the active key. Additionally, clear states on the
magicsock connection and prepare for sending a new WG handshake whenever
user data is transmitted.
Sets up for:
- Routing TSMP keys directly into magicsock
- Switching the active connection reset mechanism to the optimistic
handshake
- Cleaning up paths into controlClient
Updates #20494
Signed-off-by: Claus Lensbøl <claus@tailscale.com>
removePFAnchorRef ran at shutdown (and from the startup cleanup hook) and
unconditionally rewrote the main PF ruleset to strip the "tailscale"
anchor references. That is wrong twice over for an operator who
configured those references statically in /etc/pf.conf -- the durable
setup the ensurePFAnchorRef error message recommends:
- it drifts the running ruleset away from /etc/pf.conf on every
tailscaled shutdown, and
- the rewrite reconstructs the ruleset from "pfctl -s" output, which
names PF tables but never their contents, so any table the operator
populates out-of-band is silently emptied (same failure mode
ensurePFAnchorRef now refuses).
Track whether this process inserted the references and only remove what
we added, and even then leave them alone if the ruleset now references
tables. Skipped removal is harmless: callers flush the anchor's contents
first, and a reference to an empty anchor has no effect on traffic. The
flag is process-local by design, so the cross-process startup cleanup
hook never removes references it cannot prove are tailscaled's.
The only-remove-what-we-added semantics were first identified and
implemented by Ross Williams (@overhacked) on a fork of this branch;
this is an independent implementation of the same idea alongside the
table guard.
Updates #5573
Change-Id: Idf579a8b4190c37211782feac4e5bfd057c3ace0
Signed-off-by: Martin Minkus <martin.minkus@sonic.com>
ensurePFAnchorRef reconstructs the main ruleset from "pfctl -sn" and
"pfctl -sr" output and reloads it with "pfctl -f -". That output names any
table a rule references but never prints its contents, so the reconstructed
ruleset re-declares every table as empty. Reloading it silently drops the
addresses of any table the operator populates out-of-band -- a "persist file"
table, "pfctl -T add", pfctl's own automatic tables for interface groups --
and every rule referencing that table then matches nothing.
That is a quiet, security-relevant failure: a ruleset whose "pass ... from
<trusted>" rules still exist but match no addresses looks fine in "pfctl -sr".
There is no way to insert an anchor reference into a running ruleset without
a reload, so detect the case and refuse, pointing the operator at the durable
fix (putting the anchor references in /etc/pf.conf, where they survive reboots
and pf reloads anyway).
Boxes that already have the anchor references configured are unaffected:
ensurePFAnchorRef returns early before reaching this check.
Updates #5573
Change-Id: I6efa9a63e1322af2ec8fd986a774ed76c5d5766e
Signed-off-by: Martin Minkus <martin.minkus@sonic.com>
The FreeBSD subnet-router NAT rule translated with "-> (self)":
nat on ! tailscale0 inet from 100.64.0.0/10 to any -> (self)
In pf, "(self)" is a round-robin pool of every address on the machine,
including tailscale0's own address and loopback, and pf deals each new
state the next address in the pool. Only flows that happen to draw the
egress interface's address work; a flow translated to any other address
gets replies the far end cannot route, and hangs at SYN. With N usable
addresses on the box, roughly (N-1)/N of connections through the subnet
router silently fail.
Observed in a natlab vmtest against a FreeBSD 15.0 subnet router with
four addresses (WAN, LAN, QEMU debug NIC, tailscale0): exactly half of
8 HTTP requests hung, alternating, and the pf state table showed the
failed flows translated to the debug NIC's address and to tailscale0's
own address:
10.0.0.102:51100 (100.64.0.1:35132) -> 10.0.0.103:8080 ESTABLISHED
10.0.2.15:56553 (100.64.0.1:35148) -> 10.0.0.103:8080 SYN_SENT:CLOSED
100.64.0.2:52655 (100.64.0.1:54106) -> 10.0.0.103:8080 SYN_SENT:CLOSED
On a production FreeBSD firewall running this branch, the equivalent
IPv6 rule shows 55 state creations totalling 117 packets (about two
packets per state): SYNs whose replies never came back.
Emit one rule per up, non-loopback, non-Tailscale interface instead,
translating to that interface's own address, per address family only
where the interface holds a usable address of that family:
nat on vtnet0 inet from 100.64.0.0/10 to any -> (vtnet0)
nat on vtnet1 inet from 100.64.0.0/10 to any -> (vtnet1)
which is also the rule form FreeBSD firewall operators write by hand.
With this, the same 8-request test passes 8/8, and the LAN interface's
rule shows 8 states with healthy packet counts (56 packets total).
The interface set is sampled when SNAT is enabled; interfaces added
later are not covered until SNAT is toggled or tailscaled restarts.
Updates #5573
Change-Id: Ife3367124737ce5c8785ca7f920eafca593ec705
Signed-off-by: Martin Minkus <martin.minkus@sonic.com>
The dialer returned by makeHangDialer runs on netstack-owned goroutines
that can outlive the test, so calling tb.Logf from it raced with the
test completing. Use tstest.WhileTestRunningLogger, which stops logging
once the test is done, as makeNetstack already does.
Fixes#21052
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ibbf4288eff48f3ae0bb5812204a1cdfd567322a3
Instead of accessing the key and string directly, put it behind a method
to make it easier to use that as a proxy when we add multiple keys
(control and TSMP origin).
Instead of only updating a single origin for disco keys, teach peermap
how to work with multiple disco key sources so a key can originate from
either control or TSMP.
This sets the work up for switching dynamically between disco keys
later. As of this commit, the keys still arrive for the most part via
the controlClient, but this commit sets us up for:
- Switching dynamically between received keys.
- Route TSMP keys directly into magicsock, not via controlClient.
- Potentially revert controlClient to the state before any TSMP
changes, making it single threaded.
- Use switching keys as trigger for optimistic WG handshakes, instead
of tearing down the full connection and setting it up again.
Updates #20494
Signed-off-by: Claus Lensbøl <claus@tailscale.com>
Invalidate the trust window so we re-run path discovery
instead of coasting on the dead path until trust lapses on its own.
Partially reverts 85bb5f8 removal of resetLocked within updateFromNode.
Fixes#20268
Change-Id: Ib18b31e4329513294b682e698d1a9f426a6a6964
Signed-off-by: Alex Valiushko <alexvaliushko@tailscale.com>
Historically, when DERP regions were switched away from strings to
numeric identifiers in PR #14641, tailcfg.Node.HomeDERP was declared
as an int instead of its own type.
This PR declares a new tailcfg.DERPRegionID type, represented by an
int64, and converts the following fields to use this type:
- netcheck.Report.PreferredDERP
- netcheck.Report.RegionLatency
- netcheck.Report.RegionV4Latency
- netcheck.Report.RegionV6Latency
- tailcfg.DERPHomeParams.RegionScore
- tailcfg.DERPMap.Regions
- tailcfg.DERPNode.RegionID
- tailcfg.DERPRegion.RegionID
- tailcfg.NetInfo.PreferredDERP
- tailcfg.Node.HomeDERP
- tailcfg.PeerChange.DERPRegion
- tailcfg.PingResponse.DERPRegionID
Note that the original field was an int, while the new field is backed
by an int64. This change makes DERPRegionID the same size on both
32-bit and 64-bit architectures.
Fixes: #20165
Change-Id: Ic6f795a6d791dd16f756f246d5a02085443e212f
Signed-off-by: Simon Law <sfllaw@tailscale.com>
The packet filter only sees a via address's outer ULA, so netstack unconditionally dialed whatever IPv4 was embedded in it.
This change refuses TCP, UDP, and ping relays to host-scoped destinations after UnmapVia.
Credit to the Anthropic infrastructure security team for finding and reporting.
Fixes https://github.com/tailscale/corp/issues/46646
Change-Id: I93d8e27a2eddbce7eb3f8c5fa4677f8de3a8ed9e
Signed-off-by: Mike Jensen <mikej@tailscale.com>
Our tests are inconsistent in how they set environment variables and cleanup.
Missing cleanup logic can leak environment variable state across tests and
change the behaviour of subsequent tests.
This recently caused flakes in feature/acme, where `TestGetCertPEMWithValidity`
leaked `TS_CERT_SHARE_MODE` and `TestAsyncRenewalDedup` to fail inconsistently.
To fix the immediate flake and prevent future leaks, introduce a `SetenvForTest`
helper that handles setting and cleaning up environment variables in tests.
Fixes#20902
Change-Id: I9f8ef45ec875d66034cfd0054311bd69b67e2243
Signed-off-by: Alex Chan <alexc@tailscale.com>
Static size checks (iossize) catch binary dirty-page growth but nothing
covered runtime heap cost, which is what actually consumes the iOS
Network Extension's 50 MiB jetsam budget. Bring up a tsnet backend
(with the full condregister feature set, matching shipping clients)
against an in-process testcontrol server and assert live post-GC heap
budgets for (a) backend startup with zero peers and (b) marginal cost
per netmap peer.
The startup test measures 1.3 MiB today and fails loudly on the
conn25 flow-table pre-allocation regression (17 MiB) that jetsam-killed
the iOS extension on large tailnets.
Budgets are deliberately generous (6-12x current measurements) to stay
flake-free while still catching the multi-MiB regressions that matter
for mobile.
A new debugknob enables us to constrain the GSO/GRO batch size to 1 for
these tests so as to avoid the memory allocation associated with those
buffers, which are a known issue with their own work stream.
Updates tailscale/corp#46408
Updates tailscale/corp#18514
Signed-off-by: James Tucker <james@tailscale.com>
Move away from mutating global, exported variables in wireguard-go. Use
device.Option's passed to device.NewDevice, instead. No functional
changes, just API cleanup in preparation of future changes.
Updates tailscale/corp#22467
Updates tailscale/corp#46396
Updates tailscale/corp#37878
Signed-off-by: Jordan Whited <jordan@tailscale.com>
Package tailcfg defines the types and constants used by the Tailscale
protocol, but since everything is all in one package, it’s difficult
to sift through the docs: https://pkg.go.dev/tailscale.com/tailcfg
We define and enumerate capabilities as string constants for
tailcfg.NodeCapability and tailcfg.PeerCapability. This PR extracts
them into their own packages:
- tailcfg.CapabilityFileSharing becomes nodecap.FileSharing
- tailcfg.NodeAttrOnlyTCP443 becomes nodecap.OnlyTCP443
- tailcfg.PeerCapabilityTaildrive becomes peercap.Taildrive
We originally intended for CapabilityFoo to grant an entitlement or
permission for Foo, and for NodeAttrBar to configure Bar in the
nodeAttrs section of the policy file. However, there was no technical
enforcement of this convention, so new capabilities have used the
NodeAttr prefix regardless of meaning. Therefore, this PR unifies
tailcfg.CapabilityFoo and tailcfg.NodeAttrBar into a single package as
nodecap.Foo and nodecap.Bar.
Ran `go fix -inline ./...` and committed the changes that replaced
uses of the tailcfg aliases with the authoritative ones.
Updates #20259
Change-Id: Ieb7e7e6c8247c39faf42fdf15c68cdc7c621c730
Signed-off-by: Simon Law <sfllaw@tailscale.com>
Implements a way to send TSMPDiscoAdverts based on a trigger from
wireguard-go when a rekey happens. This lets us distribute disco keys
consistently, but also sets us up for a minimal message that can be
distributed to other clients.
A benchmark is implemented to make it easier to keep the call cheap and
to avoid locking up anything in wireguard-go.
Updates #20081
Signed-off-by: Claus Lensbøl <claus@tailscale.com>
Unsigned peers aren't covered by tailnet lock, so they must never hold peer capabilities even if the packet filter grants them. This change extends the check for unsigned-peers to ensure full coverage in capabilities.
Fixestailscale/corp#45116
Change-Id: I918af24f0b9855e55921cbdad109cc68e745e125
Signed-off-by: Mike Jensen <mikej@tailscale.com>
Add an AppName field to the DERP ClientInfo so DERP servers can
attribute connections to the application making them, primarily for
best effort stats purposes. The value is plumbed per engine instance
rather than via a process global, so a process hosting multiple stacks
can attribute each one's DERP connections separately:
wgengine.Config.DERPAppName flows through magicsock.Options and
derphttp.Client into the naclbox-sealed ClientInfo JSON. Old servers
ignore the unknown field.
There are no callers in the tree yet setting the name.
Updates tailscale/corp#24454
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ia7d3e9c2b6f8140e5a9d7c3b2e6f1a8d4c0b5e9f
* net/tsaddr: unmap IPv4-mapped IPv6 addrs in IsTailscaleIP
IsTailscaleIP branched on ip.Is4() before checking the CGNAT range, so an
IPv4-mapped IPv6 address (e.g. ::ffff:100.64.0.1) took the IPv6 path and was
tested only against the ULA range, wrongly returning false for a Tailscale
CGNAT address. Unmap at the top so both forms are classified identically;
Unmap is cheap and IsTailscaleIPv4 stays IPv4-only for callers that need it.
Signed-off-by: Brendan Creane <bcreane@gmail.com>
* wgengine/router/osrouter: remove orphaned tailnet addrs on cleanup
The orphan-address sweep added in #20199 ran only inside Router.Set, so the
teardown path (tailscaled --cleanup, and the unconditional cleanup at daemon
start) never removed stale Tailscale addresses a previous instance left on a
persistent tailscale0 -- it only flushed iptables/nftables.
Wire address removal into cleanUp: with no desired config, every Tailscale-range
address on the interface is an orphan, so enumerate and delete them all (IPv4
and IPv6, best-effort) in removeOrphanedAddrsForCleanup.
tailscaleInterfaceAddrs now yields the interface's addresses as an
iter.Seq[netip.Prefix], and the filters compose lazily over it: tailscaleAddrs
(every Tailscale-range address; used by cleanup), deletableAddrs (isDeletableAddr:
Tailscale-range and deletable now, i.e. excluding v6 when v6 is unavailable; used
by the live Set sweep), and orphanedAddrs (drops the desired addresses). The Set
sweep ranges the composed iterator directly, so no throwaway slices are built.
delAddress is made idempotent: it attempts both the loopback-rule teardown and
the address delete and joins their errors, so a missing firewall rule can't leak
the address, and it no longer no-ops on v6 (cleanup relies on that to remove v6
orphans even when this process never brought IPv6 up).
The Set-time sweep is otherwise unchanged; re-running it on network changes
(netmon) for late orphans remains a follow-up (tailscale/corp#43882).
Updates #19974Fixestailscale/corp#44173
Signed-off-by: Brendan Creane <bcreane@gmail.com>
---------
Signed-off-by: Brendan Creane <bcreane@gmail.com>
interfaceV6UsableForTun interpolates the interface name into a /proc path.
A plain filepath.Join + os.Open only cleans the path, so a tunname with
".." (or a symlinked component) could read outside /proc/sys/net/ipv6/conf.
Open under that fixed directory with os.OpenInRoot, which rejects any path
escaping the root (openat-based, so also TOCTOU-resistant), still using
filepath.Join to build the relative name. See https://go.dev/blog/osroot.
Updates #20447
Signed-off-by: Brendan Creane <bcreane@gmail.com>
* wgengine: configure DNS even when router.Set fails
Reconfig configured the router first and returned on any router.Set error,
before the DNS block ran. On a host where router config fails on every
reconfig -- e.g. a tun MTU below 1280 that breaks IPv6, or a kernel missing
netfilter features -- the OS resolver was never told about MagicDNS or the
tailnet search domain, so tailnet names failed to resolve with no DNS error
in the logs.
Record the router error and continue instead of returning on it, still
attempt dns.Set, and join the router, DNS, and VPN-reconfigure errors into
the return value. DNS stays after router config (still needed: some DNS
managers refuse to apply settings before the device has an address); only
the error coupling is broken. Fixes a regression from 84430cdfa (v1.8.0).
Updates #20447
Signed-off-by: Brendan Creane <bcreane@gmail.com>
* wgengine/router/osrouter: gate IPv6 on per-interface support, not just global
getV6Available reported IPv6 usable whenever the netfilter runner reported
global IPv6 support, missing the case where the kernel has IPv6 but has not
enabled it on tailscale0 specifically -- e.g. when the tun MTU is below the
1280-byte IPv6 minimum, so /proc/sys/net/ipv6/conf/tailscale0 never exists
and the v6 address and route adds fail, aborting the whole Set. See #20447.
AND a per-interface check into getV6Available, evaluated per call so a later
Set picks up v6 if the interface gains it. All v6-gated operations funnel
through getV6Available, so Set now skips v6 gracefully instead of erroring.
Also remove the dead r.v6Available field that masked this with its global
name.
Updates #20447
Signed-off-by: Brendan Creane <bcreane@gmail.com>
---------
Signed-off-by: Brendan Creane <bcreane@gmail.com>
Processing a peer add/remove delta still materialized the full netmap
(an O(n) slicesx.MapValues plus sort over all peers, at 10k+
peers in a large tailnet) twice per delta: once in UpdateNetmapDelta
purely to hand the self node to Engine.SetSelfNode, and once in
authReconfigLocked.
Neither needs peers anymore. SetSelfNode gets the self node from the
existing nodeBackend.Self accessor. authReconfigLocked only reads
self-node fields (SelfNode, NodeKey, GetAddresses, HasCap) now that
WireGuard peers ride the incremental route manager and per-peer config
source, so it can use the peers-free NetMap accessor.
That also makes nmcfg.WGCfg vestigial: since wgcfg.Config lost its
Peers field, its peer walk existed only to emit the [v1] skip logs
(expired peers, unselected exit nodes, unaccepted subnet routes),
duplicating filtering the route manager already does. Delete the
package and construct the two-field wgcfg.Config inline. The skip
logs go away; if they're missed, the route manager can log them
incrementally at upsert time instead of rescanning every peer on
every reconfig.
With this, the runtime.DidRange analysis (see the ts_rangehook test)
shows a delta netmap update performing no O(n) range loops except
updateRouteManagerExtras, and the delta phase of that test drops from
1.09s to 0.14s for 400 deltas at n=10000 (from 4.79s at the
start of this effort, before the incremental route manager work).
Updates #12542
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ia0e03ef9db0c988790b2c29de1f0505305e93f58
The wireguard-go device now learns its peer set solely from the live
per-peer config source that LocalBackend installs with
Engine.SetPeerConfigFunc, backed by the route manager. Peers are
created lazily on first packet and converged per peer with
Engine.SyncDevicePeer, so the full-peer-list snapshot in wgcfg.Config
and the diff-and-reconfigure machinery around it (wgcfg.Peer,
ReconfigDevice, and the engine's full device sync in
maybeReconfigWireguardLocked) are dead weight: they duplicated state
that the route manager already owns and forced every netmap change to
rebuild and rehash the entire peer list.
Delete the Peers field and the Peer type from wgcfg, along with
ReconfigDevice and maybeReconfigWireguardLocked. Engine.Reconfig no
longer does any device peer work; it only manages the private key,
addresses, and the non-peer subsystems. Full-netmap application converges the device by
syncing exactly the peers whose routes the route manager reports as
changed or removed.
Updates #12542
Change-Id: Ic776e42cfaa5be6b9329b3d381d5cbde17d7078b
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Previously tstun.Wrapper.SetWGConfig walked wgcfg.Config.Peers on every
netmap to rebuild its own IP-to-peer table for masquerade NAT rewrites
and jailed-peer classification. Now the tun layer instead consumes the
route manager's shared immutable outbound snapshot directly, via a new
Engine.SetPeerRoutes method: LocalBackend pushes the snapshot (plus this
node's native Tailscale addresses) after every route manager commit that
can change it, and per-packet lookups read the interned PeerRoute
attributes from that table.
When no current peer is jailed or masqueraded, LocalBackend installs a
nil table (gated on RouteManager.HasDataPlaneAttrs), preserving the
per-packet nil-check fast path. The exitNodeRequiresMasq machinery is
deleted: its purpose was populating the table with all peers so that
more-specific entries shadow an exit node's /0, and the always-full
route manager table gives that shadowing inherently.
This is the last step before removing the Peers field from wgcfg.Config.
Updates #12542
Change-Id: Ifce09ca929a3f2511303ca1d6efdd583739494ce
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
The BIRD (BGP) integration previously lived half in cmd/tailscaled
(which created a chirp client via a build-tag-gated file on some
platforms) and half in wgengine (which carried the client in its
Config and toggled the "tailscale" protocol as the node gained or
lost primary subnet router duty).
Move it all to a new feature/bird package, installed on the engine
via the new wgengine.HookNewBird hook, like other feature/* packages.
wgengine.Config.BIRDClient (and the wgengine.BIRDClient interface)
are replaced by a BIRDSocket path from which the engine constructs
the feature's Bird handle at startup. The subnet router overlap
detection and protocol toggling move into feature/bird, preserving
the previous ordering: state is recomputed before Reconfig's
ErrNoChanges early return and applied after the router is configured.
tailscaled keeps BIRD support by default on the platforms that
previously had it (linux, darwin, freebsd, openbsd) via
feature/condregister.
Also, add an integration test, as this feature lacked much test
coverage previously.
Updates #12614
Change-Id: I7866a50779e454c87933b358735f7dcd9e2b126f
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Engine.Reconfig previously diffed cfg.Peers disco keys against the
previous config to find restarted peers and flush their WireGuard
sessions, with a TSMP-learned-key map to suppress resets for key
changes that arrived over a working session. That was the last
per-peer state computed from wgcfg.Config.Peers inside the engine,
and it only ran on full reconfigs, so incremental netmap deltas
never got session resets at all.
Move the detection into nodeBackend, which sees every peer change:
full netmaps in SetNetMap and incremental upserts in
UpdateNetmapDelta both now report which peers changed disco keys,
with the same TSMP suppression and mismatch accounting as before.
LocalBackend acts on the result via a new Engine.ResetDevicePeer
method, which just removes the peer from the WireGuard device and
lets the peer lookup func lazily re-create it with fresh state.
LocalBackend.PatchDiscoKey now records TSMP-learned keys in
nodeBackend instead of forwarding to the engine, so the engine's
PatchDiscoKey method and tsmpLearnedDisco map are gone. The
controlclient patchDiscoKeyer interface becomes the exported
DiscoKeyUpdater so LocalBackend can compile-time assert that it
implements it, alongside its NetmapDeltaUpdater friends, replacing
the test that asserted the same of the engine.
This is one of the last steps toward removing Peers from wgcfg.Config.
Updates #12542
Change-Id: I6b42e460f42924816beae89ca43731cb91b66054
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Engine.PeerForIP was pure delegation to the callback that LocalBackend
installs via SetPeerForIPFunc, so external callers going through the
engine were taking a pointless round trip: LocalBackend called
b.e.PeerForIP, which called right back into LocalBackend, and the
netstack UseNetstackForIP hooks in tsnet and tailscaled did the same
dance one layer removed.
Export LocalBackend.PeerForIP and make those callers use it directly.
The engine-internal cold paths (Ping, TSMP disco advertisements,
pendopen diagnostics) still need the lookup and have no netmap of
their own, so SetPeerForIPFunc stays on the interface, but the lookup
method itself is now unexported and gone from the Engine interface.
In tailscaled the UseNetstackForIP hook moves from netstack setup to
just after the LocalBackend is created, since the backend doesn't
exist yet when netstack is wired up.
Updates #12542
Change-Id: Ib1e1a4fa5c84ee0dcb9ce5d1910047f2bab9453c
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>