* derp/derphttp: reject invalid DERP node hostname before proxy CONNECT
When a DERP client reaches a node through an HTTP(S) proxy,
dialNodeUsingProxy writes the CONNECT request by hand and puts
net.JoinHostPort(n.HostName, port) into both the request line and the
Host header. n.HostName comes from the control-supplied DERP map and
net.JoinHostPort does no sanitizing, so a hostname carrying CR/LF was
written verbatim into the plaintext request sent to the proxy. That let
whoever populated the DERP map inject extra headers, or a second
pipelined request, into the connection to the operator's proxy.
Validate n.HostName with httpguts.ValidHostHeader at the top of
dialNodeUsingProxy, before the proxy is dialed, and also reject the
empty hostname, which ValidHostHeader accepts. DNS names and IP
literals continue to work.
Add a table-driven test that runs accepted and rejected hostnames
against a fake proxy and checks the CONNECT target that goes out.
Fixestailscale/corp#48122
Signed-off-by: basavaraj-sm05 <basavaraj@digiscrypt.com>
Co-authored-by: Mike Jensen <mikej@tailscale.com>
Signed-off-by: Mike Jensen <mikej@tailscale.com>
The derper debug pages had no way to see which clients were connected.
The expvar gauges only give counts, /debug/check only says whether the
counts agree, and /debug/traffic only reports connections that moved
bytes since its last tick, and only if ss is installed.
Add /debug/clients/, which by default serves an index page with a form
to pick one of four filters: ?all lists every connection, ?ip=1.2.3.4
and ?cidr=1.2.0.0/16 list connections from an address or prefix, and
?key=nodekey:... lists the connection(s) for one node key. Each row
shows the connection number, key, remote address, connection age,
flags (home, mesh, prober, notideal, dup/active/disabled), protocol
version, app name, per-connection rx/tx packet and byte counts, and
the estimated unique sender count.
Big derpers have far too many connections for one page, so results
are paginated with keyset cursors rather than page numbers: sort=key,
ip, conn, rx, tx, rxpkts, or txpkts (with a leading - for descending)
picks the walk order, limit=N the page size, and after=X resumes after
that value of the sort field. The next-page links add afterconn=N so a
page boundary that falls among connections sharing a value (duplicate
keys, one IP with many ports, equal counters) resumes exactly. Column
headers link to the other sort orders.
The walk under Server.mu does only a filter match, a cursor comparison,
and at most a bounded-heap operation per connection, so connections
before the cursor are discarded without being copied and at most limit
entries are ever kept. Only the summary counts (matching connections
and keys) look at every connection. Snapshots are taken and the page
rendered after the lock is released, so a slow debug client can't
stall the server. A benchmark with 100k connections takes about 10ms
per page.
There were no per-connection traffic counters before, only the
server-wide ones, so sclient gains four atomic.Uint64 counters (rx/tx
packets and bytes, counting data packets like the server-wide ones)
bumped alongside them. That's 32 bytes per connection. For the counter
sorts, the value is loaded once per connection during the walk and
used for both the cursor test and the heap order, so the order stays
consistent while the counters keep changing.
The sclient preferred field becomes an atomic.Bool so the page can
report which connections are the client's home DERP; it was previously
only touched by the run goroutine.
Updates tailscale/corp#48933
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I4e9b7c2d5a83f61b0e7d2c94a5f8b3e16d7c0a29
FreeBSD handles subnet routes in netstack by default, but the router
still installed its pf NAT rules whenever SNATSubnetRoutes was set,
which is the default. Every FreeBSD node, subnet router or not, loaded
and enabled pf, inserted anchor references into the host's main
ruleset and loaded NAT rules that netstack never needed, since
netstack dials subnet destinations from the host's own addresses.
It also flipped the forwarding sysctls for any advertised route.
Only do either when the kernel path is opted into with
TS_DEBUG_NETSTACK_SUBNETS=false and routes are advertised.
TestSubnetRouterFreeBSDManyFlows ran in netstack mode since the default
changed, so it no longer exercised pf at all. Opt it into the kernel
path and assert the anchor holds NAT rules, and have
TestSubnetRouterFreeBSD assert the default mode leaves pf untouched.
Also drop the stale claim in handleSubnetsInNetstack that the pf NAT
rule never matches; that was the (self) pool bug, since fixed.
Updates #21450
Change-Id: Ibf91a676a026cebb29f64e39a64fe8e75488360b
Signed-off-by: Martin Minkus <martin.minkus@sonic.com>
An Ingress annotated with
`tailscale.com/experimental-forward-cluster-traffic-via-ingress` forwards
cluster traffic to the proxy's Pod IP, which is DNATed to the node's Tailscale IP
where serve answers it. On Linux the serve listener is bound to the tunnel
interface and drops that traffic, so set TS_SERVE_ALLOW_ALL_INTERFACES on the
proxy when this annotation is used, which makes serve answer it again.
Document on the annotation how the traffic reaches serve and that it bypasses
tailnet ACLs, and regenerate the CRD and operator manifests.
Updates tailscale/corp#48248
Signed-off-by: chaosinthecrd <tom@tmlabs.co.uk>
When containerboot falls behind on the IPN bus, tailscaled closes the
watch. containerboot treated the EOF as fatal and SIGTERMed a healthy
tailscaled, which is easy to hit on large, churny tailnets.
Instead, reconnect and rebuild state from the new watch's initial
status, and only request peer changes in modes that use them. If the
watch can't be reopened for a minute, exit so a dead tailscaled still
restarts the container.
Fixes#21373
Change-Id: Iad7749e4fd0f43eabdb471d6e64bb43f37ff70ff
Signed-off-by: Raj Singh <raj@tailscale.com>
The e2e suite covered the operator's in-process API server proxy but
not the ProxyGroup-based one. Add tests for both proxy modes that
drive a ConfigMap through its lifecycle via the proxy, verify a
forbidden request is rejected, and check that deleting the ProxyGroup
cleans up its StatefulSet and Tailscale Service.
Fixestailscale/corp#38009
Change-Id: Ifc0be47ce32dd96f8daa748af6785a8b82ec19a7
Signed-off-by: David Bond <davidsbond93@gmail.com>
Linux is a weak-host stack, so a LAN-adjacent machine can complete a
TCP handshake with a node's peerapi listener by sending a packet to
the node's Tailscale IP, with no credentials and no tailnet
membership.
macOS and iOS already bind the listener to the tunnel interface, and
Windows is protected by its strong host model, so Linux tun mode was
the only platform that leaked.
Bind the Linux listener to the tunnel interface as well, so the
kernel only answers connections that arrive from the tunnel or from
the local host. A natlab VM test verifies that a same-LAN machine can
no longer complete the handshake, while local and peer peerapi keep
working.
FreeBSD has the same weak-host exposure but no per-socket equivalent,
so handling it there with pf is a TODO (#21419).
Updates tailscale/corp#48248
Reported-By: Samuel Keeley (@keeleysam)
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I5f8501b0938c9f7aa39c4c12ebddd988c72e89bf
Loosens the connector route assertion to allow concurrent tests against
the same tailnet to more reliably pass, while still asserting the
client itself is advertising those routes and they're recognised in the
API.
Also make createOrUpdate more resilient to a test that didn't fully tear
down on the same cluster previously. By respecting the existing resource
version and finalizers, we can update resources that didn't get deleted
from a previous run.
Updates tailscale/corp#45426
Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com>
Make sure the tests select matching-family IP addresses for ingress
tests so they're able to pass on IPv6 clusters. We should probably
follow up with another change for the operator that stops clients from
needing to do this sort of selection, but updating the tests is the easy
option to get them passing on IPv6 clusters in the short term.
Updates tailscale/corp#45426
Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com>
If people ran derper on a multi-NIC or multi-address host, STUN replies
could go out from the wrong address. The wildcard UDP socket let the
kernel pick the reply's source address by routing to the client, which
means the default route's address rather than the one the request came
in on. With connmark-based policy routing (e.g. DNAT through a tunnel),
the reply then doesn't match the inbound conntrack entry, goes out the
wrong interface, and the client never sees it. DERP over TCP was fine,
since accepted sockets are pinned to the local address.
Add net/pktinfo, a small Linux-only package that uses IP_PKTINFO and
IPV6_RECVPKTINFO to learn each datagram's destination address and to
reply from it, and use it in the STUN server. Only the source address is
pinned; routing still picks the interface. Other platforms are unchanged.
Fixes#21404
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7b3e9c2d41a8f60e5d9c1b2a3f4e5d6c7b8a9f01
* cmd/k8s-operator: move egress EndpointSlice write back into gated provision
PR #20347 moved the EndpointSlice createOrUpdate outside of provision,
causing it to run on every reconcile. This resulted in racing egress-eps on
the EndpointSlice, sometimes causing the TailscaleEgressSvcConfigured to
become stuck as False with the Service not fully updated. Gate it again so
it only runs when a reprovision is required.
Updates #20916
Signed-off-by: Becky Pauley <becky@tailscale.com>
* cmd/k8s-operator: recover missing egress EndpointSlices
Add a watch for EndpointSlices in the egress-services reconciler so a
deleted slice re-triggers a Service reconcile directly. Treat a Service
whose expected per-family EndpointSlice is missing as not up to date so it
re-enters provision and recreates the slice.
Also sort endpoints by Pod UID before writing them in the egress-eps
reconciler, so an unchanged set of ready Pods cannot result in a different
order and trigger an unnecessary Update.
Updates #20916
Signed-off-by: Becky Pauley <becky@tailscale.com>
---------
Signed-off-by: Becky Pauley <becky@tailscale.com>
The varz handler got its memstats_* metrics from the expvar package's
"memstats" func, which calls runtime.ReadMemStats and so stopped the
world on every Prometheus scrape. Keep the names but compute them from
runtime/metrics, and never call that func, even from
WritePrometheusExpvar.
While there, export the /tailscale/ metrics from our Go fork (stack
size histogram, stack copy counters, timer zombie counts and lifetime
histogram), which nothing could see before, plus a few upstream ones
with no MemStats equivalent: scheduling latency and GC pause
histograms, live heap, GC and total CPU seconds, thread count, and
mutex wait time. The last replaces derper's hand-rolled version.
Everything read is cheap and a scrape allocates nothing after the
first. Names use a go_runtime_ namespace rather than go_ so they can't
collide with the Prometheus Go client's collector in promvarz binaries.
The runtime's 162-bucket time histograms are reduced to one bucket per
factor of four from 256ns to 1s.
Updates #21300
Updates tailscale/go#189
Updates golang/go#75935
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7e3c9a41f2b85d6e0c4a9b1d3f8e7c2a5b6d4e19
testwrapper ran a separate, sequential "go test" invocation for each
package pattern on its command line. That is fine for a single "./..."
argument but not for callers that pass an explicit package list: CI
jobs in the corp repo passing ~200 packages ran ~200 serial go test
processes with no cross-package parallelism and a fixed set of
never-cacheable lookups per process, and spent several times longer
on process startup, package loading, cache lookups, and serial test
binary links than on running tests. See tailscale/corp#48453 for the
details.
Locally, on 203 packages with a fully warm build and test cache, so
measuring only the per-invocation overhead:
old (203 go test processes): 26.4s
new (1 go test process): 3.6s (7.3x faster)
Our own Windows CI job hits the same path: its "sharded:N/M" mode
expands to an explicit list of that shard's packages via listpkgs, so
each shard ran one go test process per package, and Windows process
startup is slower still. Each shard now runs as one invocation.
Fixestailscale/corp#48453
Updates tailscale/corp#47035
Change-Id: I3f796ff1724af40f93be9f918a7ddfde3bb45a91
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Adds tsweb/compserve: content negotiation for precompressed static
variants, with a transcode-to-identity fallback for clients that do not
accept an encoding (including when the raw file is absent), and
CompressWriter, which live-compresses dynamic responses with zstd in its
fastest mode, streamed incrementally with no buffering. Negotiation is
q-value and wildcard aware (gzip;q=0 previously matched gzip).
client/web serves its prebuilt embedded assets through compserve,
replacing brotli with zstd; the embedded FS is wrapped in
tsweb/vcstime for conditional-request mod times. tsweb/compserve/gzip.go
keeps transitional serving of gzip variants from pre-zstd file systems
(such as the currently published web-client-prebuilt module):
passthrough to gzip-accepting clients, transcoded to identity otherwise;
it becomes inert once a zstd-only module is published.
util/zstdframe gains pooled GetDecoder and GetStreamingEncoder
(concurrency=1). util/precompress is now a build-time tool, generating
zstd variants only. cmd/tsconnect and cmd/build-webclient consume the
new precompress/compserve split. tsweb.AcceptsEncoding and
tsweb/tswebutil are removed; negotiation lives in compserve and the
deprecated shim had no callers. go.mod bumps web-client-prebuilt.
Also fixes a transcoding bug where http.ServeContent's size probe via
the promoted zstd.Decoder.WriteTo could report a zero length, serving
empty bodies.
Updates tailscale/corp#20099
Signed-off-by: James Tucker <james@tailscale.com>
This is the output of the new misc/bumpdeps tool (#21325) run with
--exclude-newer-than-days=7, which asks proxy.golang.org for the newest
version of every direct dependency, ignoring releases younger than a
week in favor of the newest older one, and runs a single go get.
gvisor tracks its "go" branch, wireguard-go its "tailscale" branch,
and golang-x-crypto its "main" branch (the proxy's @latest for it is
a stray v0.91.0 tag from 2024 that predates our acme fork changes).
Indirect deps only moved as far as MVS pulled them.
The week-long cooldown held back gvisor, the gokrazy modules,
chromedp/cdproto, and hashicorp/raft-boltdb/v2, whose only newer
versions are days old; they'll come along next time.
Several upstream changes needed small fixes: nfpm's PrepareForPackager
takes a modification time now (a zero time keeps the old behavior of
using the source file's mtime), esbuild's ServeOptions.Port became an
int while ServeResult.Host became a Hosts slice, client-go's
EventRecorder.Eventf is now recognized by vet as a printf wrapper (so
the k8s-operator calls that passed a preformatted message switch to
Event), google/nftables v0.3.0 reads back the kernel's
NF_NAT_RANGE_PROTO_SPECIFIED flag into a new expr.NAT.Specified field
(so the port map DNAT rule now sets it too or findRule never matches
the rule it just added), and staticcheck v0.8.1 knows encoding/json/v2's
embed tag option, so the two SA5008 suppressions for it are gone.
Two tests assumed old library behavior. client-go's fake clientset now
replays existing objects when a watch starts, as a real apiserver does,
so the k8s-proxy config test must tolerate the loader ignoring that
no-op event before the real reload arrives. fyne.io/systray moved its
dbusmenu object path and answers the first GetLayout with depth 1, so
the systray test now finds the menu via the item's Menu property and
polls until the submenu entries appear.
Then make tidy, make updatedeps, and make kube-generate-all (the
controller-gen bump to v0.22.0 changes doc strings, stops listing
top-level metadata as required, and crd-ref-docs now marks optional
fields).
Updates #8043
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I3f9a2c6e8b1d4705a9e2c7b8d1f4e6a0c2b5d8e3
Add Node.StableTailnetID for control to send the current tailnet's
stable ID on the self node. Expose it as CurrentTailnet.StableID in
LocalAPI status and `tailscale status --json`, and display it in
`tailscale whoami`.
Bump CurrentCapabilityVersion to 148.
Updates #14375
RELNOTE: Show the current tailnet's stable ID in status JSON and whoami.
Change-Id: I0525ff735de8113c8d124045a94e00c19a5a02e2
Signed-off-by: Andrew Dunham <andrew@tailscale.com>
The test control server always listened on 127.0.0.1:9911, which is
useless for a node in a VM on the same machine. --addr picks the listen
address; the DERP and STUN servers follow it.
--ssh-policy loads a tailcfg.SSHPolicy from a JSON or HuJSON file and
sends it to every node, which also grants them the SSH node capability so
that "tailscale up --ssh" is accepted. That is what testcontrol.Server
already supported for tests; this exposes it for manual testing of the
Tailscale SSH server.
Updates tailscale/corp#47865
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I2b7c4e9f1a3d5c6e8b0f2a4d6c8e1b3f5a7d9c0e
Update gVisor to include its fix for RACK loss detection with coarse
monotonic clocks. Configure netstack with the 500 microsecond clock
resolution used on Windows so RACK accounts for timestamp quantization.
Remove the TCP recovery override that disabled RACK, enabling gVisor's
default RACK behavior on all platforms.
Switch netstack to cubic congestion control. The int overflow in CUBIC
sender cwnd arithmetic that required pinning reno has since been reworked
upstream into float arithmetic with RFC 9438 target clamping.
Align the natlab vnet stack with netstack: enable cubic, and drop the
now-redundant explicit SACK and receive-buffer moderation sets, both of
which are gVisor defaults.
Fixes#9707
Signed-off-by: James Tucker <jftucker@gmail.com>
The tsweb DebugHandler already links to expvar, Prometheus varz, and
pprof, but the only runtime visibility beyond pprof was the handful of
runtime.MemStats fields that varz special-cases. The runtime/metrics
package has far more (GC CPU classes, scheduler latencies, stop-the-world
pause histograms, heap breakdowns, and so on) and is the runtime's
preferred, cheaper interface.
Add a /debug/runtime-metrics handler. By default it serves only an index
of metric names, kinds, and descriptions, which are static, so viewing
the page does not read any values. The index is an HTML table for
browsers (Accept: text/html) and plain text otherwise; format=html or
format=text overrides the sniffing.
Values are read only on request, always as JSON:
- /debug/runtime-metrics/gc/heap/allocs:bytes returns that metric's
bare value, handy for scripts and curl.
- /debug/runtime-metrics?name=NAME (repeatable) returns a JSON object
keyed by metric name. A trailing * matches by prefix, so name=/gc/*
returns the GC metrics and name=* returns everything.
Histograms are objects with counts and buckets arrays; infinite bucket
boundaries, which JSON cannot represent, are the strings "-Inf" and
"+Inf".
In the HTML index, exact metric names link to the bare value form, and
each run of metrics sharing a directory is headed by a row linking every
ancestor directory to its wildcard query (/gc/*, /gc/heap/*, ...).
Responses inherit the debug handler's nosniff, framing, and CSP headers,
and application/json is not a script MIME type, so the values cannot be
pulled in cross-origin as a script by a malicious page.
Like the pprof handlers, this is excluded from js/wasm builds to keep
them small.
Updates #21300
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7c2e9f4a1b8d3e6f0a5c9b2d4e7f1a3c6b8d0e2f
The ts_omit_<name> build tags omit a feature at build time; there has
been no way to do the same at runtime. Some users (either proactively
or in response to a security announcement) might like a way to disable
a feature that's linked-in in their binaries that they're not using.
Then a mitigation announcement can say "set this env var" without
asking users to rebuild or wait for a new release.
This adds env var TS_DISABLE_FEATURE, a comma-separated list of
feature names to disable, and the listed set is reported by the
debug-optional-features LocalAPI endpoint next to the registered set.
The legacy per-feature knobs such as TS_DISABLE_SSH_SERVER and
TS_DISABLE_TAILDROP keep working independently.
A disabled feature behaves as if it had not been linked: it is absent
from feature.IsRegistered, its hooks are unset, and its extensions and
handlers are not registered. Three pieces make that happen:
* feature.Register now returns bool, false when disabled, and
feature packages gate their registration init on it. It was added
to the feature packages that never called it (including taildrop
and ssh), which also completes the picture reported by
debug-optional-features. taildrop, routecheck, favorites, and
serviceclientprefs had registration split across several inits and
now register from one gated init.
* ipnext.RegisterExtension ignores a disabled feature's extension.
* feature.Hook.Set and feature.Hooks.Add walk the call stack and
silently skip when the calling package under feature/<name> is
disabled. This covers sub-packages such as
feature/captiveportal/netcheckhook, which cannot call Register
themselves without colliding with their parent, and future
packages whose authors forget the gate.
ssh/tailssh's registrations moved from its inits into tailssh.Register,
called from feature/ssh's gated init. The aws and kube state stores and
syspolicy's Windows store registration are gated too.
feature/register_disable_test.go runs this test binary as a child
process (it links condregister, as tailscaled does) with
TS_DISABLE_FEATURE set to every registered feature at once, and fails
if any of them register anyway, so a feature that ignores the variable
cannot land.
Updates #12614
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I720af6ccab844ae060a9dfd1539fee577fd483e3
wintun-go loads wintun.dll only from the application directory and
System32. The MSI puts it there, but a plain "go build" tailscaled.exe run
from a terminal has no wintun.dll, and tstunNewWithWindowsRetries then
retried tstun.New for five minutes with nothing in the log but
tstunNew: backoff: 13 msec
tstunNew: backoff: 32 msec
because backoff.BackOff never logs the error it's given, and after the
timeout the function returned context.DeadlineExceeded, discarding the
real error.
Instead, log with helpful text if we detect that.
Updates #21290
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I09dca3317974a977deb9956657a7e36e9ec8c572
The map response reader now caps a single message at 256 MiB on the wire
and 1 GiB after zstd decompression. The server-chosen uint32 size prefix
previously let a malicious control server make us allocate up to 4 GiB
before reading any body bytes, and the decoded size was unbounded, so a
small zstd frame could expand into gigabytes of JSON. A 16 MB cap has
been hit by real production traffic before, so both limits sit far above
plausible legitimate sizes. The size-prefixed read moved into a
readMapResponseMessage helper, and the newer control/tsp path already
enforced both kinds of bounds; this brings the long-poll path it
replaces in line, with more generous limits for large tailnets.
ts2021.Client.Do additionally caps every noise response body with
httpbody.LimitSize, so a malicious or buggy control server can't make us
buffer an unbounded response. Client.Do shadows the embedded
http.Client's Do method, so register, set-dns, set-device-attr,
audit-log, all DoNoiseRequest consumers (webclient, tailnet lock, SSH
actions, id-token, feature queries), and the debug CLI get the cap
without per-call-site changes, and future noise endpoints get it for
free.
The cap lives in the new util/httpbody package so other HTTP clients can
adopt the same convention: LimitSize looks the size limit up from
res.Request's context (a Response knows the Request that produced it),
falling back to DefaultMaxSize, 1 MiB, when the context carries no
override. It is like io.LimitReader except that reads past the limit
fail with an error wrapping httpbody.ErrTooLarge instead of silently
truncating, and a body of at most the limit, including one of exactly
the limit, reads back without error: the wrapper probes for EOF once the
limit is exhausted to tell an exactly-at-limit body from an oversize
one. The per-request override, httpbody.WithMaxSize, is a context key,
so transports pick it up with no API changes; LimitSizeTo applies an
explicit limit ignoring any override. Repeated LimitSize or LimitSizeTo
calls replace the previous limit rather than compounding it, so a later
call can raise or remove the limit an earlier one set.
Responses that stream an unbounded number of individually bounded
messages disable the cap with httpbody.WithMaxSize(ctx, 0): the
/machine/map long-poll and control/tsp's map session, whose messages are
already capped per-message (by readMapResponseMessage and decodeMsg, and
by tsp's framedReader and boundedReader). Their non-200 error bodies are
not message streams, so those are capped with LimitSizeTo instead.
The tailnet lock /tka/init/begin, /tka/sync/offer and /tka/affected-sigs
responses can carry per-node key signatures or missing AUMs, which at
100,000 peers reach tens of MB, so they raise the cap to 512 MiB. The
per-response io.LimitedReader decoders that silently truncated those
responses at 1 or 10 MiB are removed: the transport cap is now the
single enforcement point, and it reports oversize bodies instead of
truncating them.
The /key fetch over plain TLS switches from io.LimitReader to
httpbody.LimitSizeTo, so an oversized response reports the problem
instead of producing a confusing truncated-JSON error.
Thanks to Ben Carman for the report!
Updates tailscale/corp#48187
Reported-by: Ben Carman
Change-Id: Ibf95e1ab9e4f0d7ef8866e8c26e62ed2a514455a
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Nothing reads ipn.Notify.NetMap anymore. The previous commit removed
its runtime (non-initial) emission, and every first-party client is
also off the initial one: the Win32 and WinUI GUIs and the Apple
bridge subscribe with InitialStatus or InitialState plus peer deltas,
Android no longer uses it, and the remaining in-tree subscribers that set
NotifyInitialNetMap (sniproxy and the kube helpers) only did so to get
the initial Notify.SelfChange and discarded the netmap that tailscaled
built, encoded, and shipped for them.
Delete the Notify.NetMap field and the NotifyInitialNetMap bit. The
bit value stays reserved under the name ObsoleteNotifyInitialNetMap and
ValidateNotifyWatchOpt rejects subscriptions that set it, like the
NotifyRateLimit bit removed in the previous commit. NotifyNoNetMap
remains accepted as a no-op because shipping GUIs still set it.
The blessed way to seed a watcher's view is NotifyInitialStatus, but
it unconditionally built O(peers) status entries, which is exactly the
waste this series is deleting for watchers that only care about the
self node. Size the initial status to the subscription instead:
Status.Peer is only populated when the watcher also set
NotifyPeerChanges or NotifyPeerPatches, since only peer-delta
subscribers need a peer baseline to apply deltas to. That matches its
existing first-party users (containerboot and the WinUI GUI both pair
InitialStatus with peer bits).
Migrate sniproxy and the kube helpers to NotifyInitialStatus: they
seed from InitialStatus.Self and react to the (ungated) runtime
Notify.SelfChange messages after that, so their initial message now
carries one PeerStatus instead of a full netmap.
Also drop doc comment references to LocalClient.NetMap, a method that
does not exist; on-demand fetches go through other LocalAPI methods
such as LocalClient.Status.
Updates #12542
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I242992a744c0ffd0be6f27e8c735aa69d5b23b5e
The tsconnect wasm build still has osusergo and netgo since its early
6f5096fa6, copied from our linux static-linking build flags. They are
no-ops on js/wasm: there is no cgo resolver or cgo os/user
implementation to disable, so the pure Go paths are used regardless.
The omitidna and omitpemdecrypt tags (used by the wasm build and by
gocross for darwin and ios) were binary size reduction patches in our
github.com/tailscale/go fork, not upstream Go, and did not survive the
fork's per-release history reboots: omitpemdecrypt only ever existed
on the tailscale.go1.14 and tailscale.go1.15 branches, and omitidna
made it through tailscale.go1.17. Both have been unrecognized no-ops
since we moved to the go1.18 fork branch (927fc3612, March 2022, first
shipped in Tailscale 1.24.0).
Also stop passing tailscale_go explicitly: the tailscale/go fork's
cmd/go sets that tag itself as of 2026-03-31.
Updates #21250
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I10f1244f02a915943ca84e107aff8683b6e771b3
We have visibility into what the gocacheprog is doing via its stats in
the github jobs, and its session on the server side, so this is just a
spammy log line that doesn't print anything very interesting. In
particular, cigocacher is started once for each package when you pass a
list of packages to testwrapper, so it prints many times for jobs where
we filter to a certain set of packages.
Updates #cleanup
Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com>
This is a follow-up to #21029 (aa2681ac5f) to make it a bit stricter
and not cache DNS results until they've passed TLS cert validation,
to weed out DNS servers that are lying (like captive portals).
Because this is done via dnscache.TLSDialer we only catch the control
connection, but that's fine. That's all we need to come back alive
if DNS was down because real system DNS is itself over Tailscale.
The DERP connections should come via IPv4/IPv6 fields in the DERPMap.
And the logging connection isn't important; it'll buffer and catch up
later as needed when DNS is back up.
Updates #21028
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I75f176a0222f04ed52c9de1247deeed9b911f92c
Adds an opt-in, in-memory aggregator of recent connection-rejection
events (TSMP rejects received from peers, outbound TSMP rejects we emit
on ACL-blocked inbound flows, and pendopen timeouts) keyed by
(direction, proto, peer-address, reason). The aggregated data is exposed
over a new debug-rejects LocalAPI endpoint and a GET /debug/rejects c2n
endpoint, intended for future GUI/CLI consumption when diagnosing why a
connection failed.
Architecture:
- net/connreject holds the data types and a per-LocalBackend
Aggregator (LRU-bounded, default 256 entries on desktop / 32 on
mobile, per direction).
- feature/connreject is a self-registering ipnext.Extension that owns
one Aggregator per LocalBackend, installs note callbacks on the
tundev and engine, subscribes to OnSelfChange to flip the runtime
gate, and serves the LocalAPI/c2n endpoints.
- wgengine.Engine and *tstun.Wrapper each gain a SetConnRejectNote
setter; data-plane sites use a single atomic.Pointer load + nil
check, so the cost when no consumer is installed is one MOV.
Gating:
- Compile-time: ts_omit_connreject build tag (standard
feature/buildfeatures + condregister plumbing). Trims ~41 KB.
- Runtime: nodecap.ConnReject node attribute, off by default
at the control plane. May be removed once the feature is enabled
by default.
Updates CapabilityVersion to 146 (clients understand nodecap.ConnReject
and can serve GET /debug/rejects).
Adds Proto/Src/Dst accessors on flowtrack.Tuple (used by pendopen to
construct events without exposing the tuple's internals to the
aggregator).
Updates #1094
Updates #14802
Change-Id: I83e8f24a7e66fa2d158d128bd25fbe851134941b
Signed-off-by: James Tucker <james@tailscale.com>
Add a new modular dnsresolvecache feature that records every successful
DNS resolution from net/dnscache as a JSON file per hostname in
$statedir/dns-cache/, so a later boot with misconfigured DNS can still
find last-known-good IPs for critical hostnames like the control plane.
Files are rewritten only when their contents change, so a file's
modification time records when the answer last changed.
When regular DNS resolution fails, the disk cache is now consulted
before the DERP-based bootstrap DNS in net/dnsfallback. This is the
first step toward removing the DERP-based mechanism: new clientmetrics
(dnscache_disk_fallback_hit, dnscache_disk_fallback_miss,
dnscache_derp_fallback_ok, dnscache_derp_fallback_dial_ok) will tell us
when the DERP path no longer fires in the fleet and can be deleted.
The feature is linked into tailscaled by default (omittable with
ts_omit_dnsresolvecache) and is not included in tsnet.
Updates #21028
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I59228cbf68e3b48dfb1215cdd12bd8166ab14034
Android denies app UIDs both NETLINK_ROUTE (golang/go#40569, #2293)
and /proc/net, so net.Interfaces always fails and netmon.New errors
out before a standalone binary can do anything. The Android app solves
this from Java via netmon.RegisterInterfaceGetter, but raw binaries
run under Termux or a rooted shell have no Java to lean on. This is
the second half of #21129, following the androiddns feature.
I added a netmon fallback hook, consulted only when no interface
getter was registered and net.Interfaces failed, and a new androidbin
feature (ts_omit_androidbin) that implements it: report a single
synthetic interface whose v4 and v6 addresses come from asking the
kernel to route an outbound UDP socket, which sends no packets and is
permitted in the app sandbox. That's enough for magicsock to discover
local endpoints. The fallback also requires runtime evidence of
Android (GOOS=android, or /dev/__properties__ existing for GOOS=linux
binaries running under an Android kernel), so it's inert on regular
Linux.
The androidbin feature also blank imports androiddns, and fixes a
third gap I found while testing: GOOS=linux binaries on Android have
an empty CA root pool, because Go's unix root loader doesn't know
Android's /system/etc/security/cacerts (the GOOS=android loader
does), so all TLS verification fails. On Android it points
SSL_CERT_DIR there unless the user already set it, as Termux's
ca-certificates package does.
The feature is on by default in tailscaled builds on Linux and
Android via condregister, and deliberately not linked into tsnet by
default; tsnet apps and other programs opt in with a blank import of
tailscale.com/feature/androidbin.
I verified on an Android 13 emulator with SELinux enforcing, running
GOOS=linux static binaries under the app UID (run-as), where
net.Interfaces fails with the exact netlinkrib permission denial from
the issue: netmon.New succeeds with the synthetic interface, and a
tailcat binary importing this feature does DNS via dnsproxyd, fetches
its DERP map over TLS using the Android cert store, completes STUN,
selects a DERP region, and prints its address.
Updates #21129
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: If7dbcbb825ecd24bcf6d9d64b1e334dd00700d51
Previously the DERP handler served each connection for its lifetime
on its hijacked connection's net/http handler goroutine. That
goroutine's conn.serve stack frame kept the dead HTTP/1 server state
reachable for the whole DERP connection: the http.conn and its 4KB
bufio.Writer (hijack hands over c.bufw but conn.serve still references
it, so derpserver returning it to its flush pool never made it
collectable), the 4KB bufio.Reader, the upgrade *http.Request with its
parsed headers, and the request context chain. The goroutine also kept
the stack growth from the TLS handshake and HTTP request parsing.
Instead, hand the connection off to a new goroutine and return from
the handler (ala tailscale/corp@dc09e27aef), letting all the HTTP
upgrade state be collected. Give Accept a smaller 1KB frame reader,
draining and releasing the hijacked reader if it contains buffered
bytes from a fast-start client, and a nil bufio.Writer so writes go
through pooled buffers held only for the duration of a write instead
of a per-connection buffer.
Because the handler now returns at handoff time, cmd/derper's
gauge_derper_tls_active_version decrement can no longer be deferred
to handler return: intercept Hijack in the TLS metrics wrapper and,
for hijacked connections (DERP, its WebSocket flavor, and CONNECT),
decrement the gauge once when the hijacked connection closes,
restoring the gauge's connection-lifetime semantics. Teach
derpserver's TCP RTT stats to unwrap the close-hook conn so they
still find the underlying *net.TCPConn.
Also soften the UntypedHexString deprecation notices in types/key to
warnings: the untyped hex string format is the DERP wire protocol's
key encoding, so these call sites are legitimate and permanent, and
a Deprecated marker just makes them light up in editors and linters.
The cautionary text about the format's risks remains.
Measured with 100,000 idle TLS DERP connections on linux/amd64:
standing memory drops from 55.6KB to 32.6KB per connection (-41%).
Updates #21064
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: If05a0c6ea79134807e9e8872861db216
The PeerRelay resource hardcoded UDP port 41641 in the generated
Service, the tailscaled config, status.endpoints and the advertised
static endpoints. That breaks deployments where the load balancer
address or port is not what peers can actually reach, e.g. behind NAT.
Add spec.service.port to choose the UDP port, and spec.staticEndpoints
to advertise extra address:port pairs alongside the discovered load
balancer endpoints. A replica whose only endpoint is a static one
counts as addressed, and a static entry's port wins when it names an
address the load balancer already provides.
The e2e tests now supply static endpoints, which lets them assert the
PeerRelayReady condition goes True in kind, where no cloud controller
ever gives the LoadBalancer Services an address.
Fixes#20821
Change-Id: I463012cd447c81c2849fa653fa18eb82662aaf4f
Signed-off-by: David Bond <davidsbond93@gmail.com>
Move the container image build logic out of the e2e test setup into
cmd/k8s-operator/e2e/internal/build, with a thin CLI wrapper at
cmd/k8s-operator/e2e/build, so CI can publish images once per commit and
fan out into multiple test jobs that consume them. The command skips
images that already exist in the registry, so a retried run doesn't
fail if the registry is immutable.
Also adds support for the test harness leveraging workload identity
federation credentials so we can use GitHub's token issuer for the test
code itself, and AWS' OIDC provider for the operator, and avoid using
any secrets in CI.
Updates tailscale/corp#46577
Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com>
Android doesn't have /etc/resolv.conf. This causes problems for people
running GOOS=linux binaries (or GOOS=android binaries without cgo, so
they don't use Android's bionic libc) in Termux, adb shell, etc.
(Go binaries built with cgo use bionic on Android: golang/go#10714)
16 years ago when I was on the Android team I added a system-wide DNS
cache (dnsproxyd) and made bionic query that, so each Android app
wasn't doing its own DNS resolution. That interface was never meant to
be stable, and I thought that code would be surely dead by now 16
years later, but apparently it lives on, and is more stable now: both
empirically (time, ossification?), and because of how Android's split system
updates work nowadays, the dnsproxyd lives on the other side of bionic,
so they seem to keep it pretty stable. The old bionic<->dnsproxyd APIs
I added 16 years ago are still there, but 8 years ago it got some additional
APIs to query by a DNS packet instead.
So use it! If we find ourselves on Android and without libc access
(and because we don't want to pull in ebitengine/purego with all its
side effects), just query the DNS server like bionic does.
This can be disabled in Linux binaries with ts_omit_androiddns.
Old links:
LineageOS/android_system_netd@007e987feehttps://android.googlesource.com/platform/system/netd/+/007e987fee7e815e0c4bc820f434a632b7a69a9d
("DNS proxy thread in netd.")
aosp-mirror/platform_bionic@a1dbf0b453https://android.googlesource.com/platform/bionic/+/a1dbf0b453801620565e5911f354f82706b0200d
("DNS proxy: the start. proxies getaddrinfo calls.")
Back then I found it cleaner to proxy at the getaddrinfo level rather
than speak in terms of DNS packets. The raw-packet resnsend command I
use here came eight years later, added in November 2018 for Android
10's android_res_nsend NDK API:
LineageOS/android_system_netd@c0c818f448https://android.googlesource.com/platform/system/netd/+/c0c818f448efa90ab1f9b1733fb86c5e22fb894c
("Add resNetworkSend cmd in DnsProxyListener")
Android 10 (codename Q, API level 29, released September 2019) is
therefore the minimum OS version for this to work.
I verified this against the DnsResolver module on an Android 13
emulator with SELinux enforcing, from the shell UID, with both a pure
Go GOOS=android binary and a static GOOS=linux binary: raw queries,
NXDOMAIN handling, the runtime Android detection, and a tailcat binary
reaching DERP with lookups visible in the daemon's logcat output, some
served from my 2010 DNS cache.
Updates #21129
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I0d63763e255a077e4e5745b3e64ba0d78dab6d69
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
This commit bumps the wireguard-go dependency to incorporate changes to
the packet memory model and the tun.Device.Read and conn.ReceiveFunc I/O
interfaces. It updates their implementations accordingly.
These changes improve throughput in all measured benchmarks and reduce
peak RSS in six of eight cases. The two regressions will be addressed in
a follow-up commit that reduces peak RSS below the baseline measured at
1e69418. That work is kept separate to simplify review.
The following throughput and peak RSS benchmarks were performed with
iperf3 between two Intel i5-12400 nodes running Ubuntu 24.04 (Linux 6.8).
The UDP benchmarks did not use UDP GSO on the sender, so they were
roughly equivalent to single packet I/O through wireguard-go.
TCP/1 signifies one TCP stream; TCP/128 signifies 128 parallel TCP
streams.
Throughput (Mb/s)
Test 1e69418 After Change
TCP/1 10,371 11,354 +9.5%
TCP/128 7,886 8,404 +6.6%
UDP/1 2,111 2,853 +35.1%
UDP/128 1,747 2,235 +28.0%
Peak memory (VmHWM, kB)
Test Side 1e69418 After Change
TCP/1 TX 98,240 52,596 -46.5%
RX 287,748 73,384 -74.5%
TCP/128 TX 101,196 52,812 -47.8%
RX 290,420 63,620 -78.1%
UDP/1 TX 58,864 160,840 +173.2%
RX 137,516 49,900 -63.7%
UDP/128 TX 66,148 116,096 +75.5%
RX 154,384 56,556 -63.4%
Updates tailscale/corp#46716
Updates tailscale/corp#22467
Updates tailscale/corp#36989
Updates tailscale/corp#37878
Signed-off-by: Jordan Whited <jordan@tailscale.com>
Once the pkg-types script has generated pkg.d.ts and checked them
against the tsconfig.json and node_modules in cmd/tsconnect, copy
them into the pkgDir so they get bundled into the package.
Updates #19707
Signed-off-by: Gesa Stupperich <gesa@tailscale.com>
Replace the allowed-IP-only peer callback result with wgcfg.PeerConfig.
It carries allowed IPs and an optional pre-shared key through lazy peer
creation and active peer synchronization.
Update wireguard-go for the new peer PSK APIs.
Updates tailscale/tailcat#84
Change-Id: Iacd9d2c74b0b64d690f3cbdf93918686ac6076d7
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Previously, the new disco keys entering the client via TSMP was routed
into controlClient to allow for deduplication and filtering of disco
keys and avoid control overwriting an active TSMP learned key with a
stale key.
A system supporting multiple disco keys was introduced to allow these
keys living side by side, along side a system for selecting an active
egress key with a bias towards keys learned via TSMP. On ingress any
known key is accepted.
This PR makes the switch to key routing, by sending new TSMP learned
disco keys directly into the userspace engine and in turn magicsock,
removing the need for a full layer of deduplication in the
controlClient.
Additionally, the system that previously fully reconfigured clients on
disco key updates coming from control is now using an optimistic
handshake througha newly introduced wireguard-go method,
ScheduleHandshakeOnUserSend implemented in:
https://github.com/tailscale/wireguard-go/pull/81
What this PR does not do is revert the mapSession back to being single
writer. Bringing the mapSession back to this state is desired, however
to ensure the plumbing is done right and easier to reason about, that
will be done in a separate PR.
Updates #20590
Signed-off-by: Claus Lensbøl <claus@tailscale.com>
When the wasmbuild is run from external workflows, it can't derive
the version stamps from its build context. Allow setting the
ProdLDFlags version.longStamp and version.shortStamp from the
VERSION_LONG and VERSION_SHORT environment variables when present.
This way, if the wasmbuild caller already knows them (e.g. from
mkversion), it can pass them in.
This avoids "x.y.z-ERR-BuildInfo" showing up in the admin console's
machine version column.
Updates #19707
Signed-off-by: Gesa Stupperich <gesa@tailscale.com>
When receiving disco traffic from a node, mark that node as having been
seen. Should that disco key not be the active use disco key, switch to
that one as being the active key. Additionally, clear states on the
magicsock connection and prepare for sending a new WG handshake whenever
user data is transmitted.
Sets up for:
- Routing TSMP keys directly into magicsock
- Switching the active connection reset mechanism to the optimistic
handshake
- Cleaning up paths into controlClient
Updates #20494
Signed-off-by: Claus Lensbøl <claus@tailscale.com>
Moving FreeBSD subnet routing from netstack to the kernel changes the
behavior of every existing FreeBSD subnet router, and the kernel path is
not yet ready to be the default:
- The pf NAT rule we install never matches. On a production FreeBSD
15.0 subnet router, "pfctl -vsn -a tailscale" reports 97071
evaluations with 0 packets and 0 translations, and "pfctl -s info"
reports translate: 0. So --snat-subnet-routes=true, the default,
silently performs no source NAT at all.
- Inserting the pf anchor at runtime requires reloading the main
ruleset, which cannot preserve the contents of any table that
ruleset references. See the comment in osrouter.loadPFMainRuleset.
Netstack does its own SNAT in userspace, touches no system state, and is
what FreeBSD has always used, so keep it as the default. The kernel path
can be opted into with TS_DEBUG_NETSTACK_SUBNETS=false; only that path
can serve --snat-subnet-routes=false, which netstack cannot do because
it must rewrite the source address.
With this, TestSubnetRouterFreeBSD passes.
Updates #5573
Signed-off-by: Martin Minkus <martin.minkus@sonic.com>
Use net.JoinHostPort when expanding proxy targets so IPv6 loopback addresses retain the required brackets. Accept ::1 as an HTTP and TCP destination in both current and legacy serve implementations.
Fixes#8702
Signed-off-by: James Tucker <jftucker@gmail.com>
Promote the toolchain from Go 1.26.6 to Go 1.27.0, matching what
go.toolchain.next.rev has been testing. Besides the toolchain files
themselves (updated by pull-toolchain.sh), this bumps the go.mod go
directive, the Dockerfile golang base image, and the README, and
regenerates the depaware.txt files and the gzip assets in
tempfork/spf13/cobra and util/eventbus, whose bytes change with
Go 1.27's rewritten compress/flate.
Also bump golangci-lint to v2.13.1, the first release line built
with Go 1.27; the prebuilt v2.10.1 binary refuses to target a Go
version newer than the one it was built with.
Also bump golang.org/x/net to v0.58.0 (plus the sibling x/ module
upgrades it requires) to pick up upstream commit 8d10596d2624
(http2: avoid deadlocks in wrapped ClientConn state callback),
which we hit during Go 1.27 rc testing.
Also add docs/go-bump-checklist.md for next time.
Updates #20220
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ia3e4c9effafbc91227eed39efb52f1fba1b8d89c
Extend e2e test setup to work with a remote registry and real cluster.
Previously only kind was supported.
Detect cluster node architecture and build images only for the given
architecture to reduce build time. Build for all platforms as the fallback
option. This assumes only one node architecture per cluster, which is
reasonable for our test cases.
By default, --build loads the built images into a kind cluster. Use with
--registry to instead specify and push to a remote registry.
Fixestailscale/corp#46577
Signed-off-by: Becky Pauley <becky@tailscale.com>
The functions used to be used in ipn/ipnlocal but we changed that, so
the functions can be moved into feature/conn25 now.
Updates tailscale/corp#47250
Signed-off-by: Fran Bull <fran@tailscale.com>
Clients can advertise an opaque app name in their ClientInfo but the
server previously did nothing with it.
Constrain app names to at most 32 bytes of printable ASCII, enforced
both in derp.NewClient and by the server when it parses the ClientInfo.
Extend the peerPresent frame, following its existing pattern of
appending optional fields, with a length-prefixed app name after the
flags byte, so trusted mesh watchers (other DERP nodes and stats
tools) can attribute connections by app. Old clients ignore the extra
bytes; old servers send frames without them.
Also add a derper --disallow-app-names flag taking a comma-separated
list of app names whose connections are refused, except for trusted
mesh peers.
Updates tailscale/corp#24454
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I6e721258675145833aafa1355fabf7fc05a5a204
The set-config loop uses uint16 endpoints. When Last is 65535, the increment wraps to zero and file or Unix targets continue indefinitely.
Break after applying Last so every closed range terminates without changing ordinary range behavior.
Fixes#20873
Signed-off-by: Bonobo <github@in9.at>
The service-pg-reconciler reconciles Services annotated for an ingress
ProxyGroup, but its ProxyGroup watch reused ingressProxyGroupFilter,
which is ingressesFromIngressProxyGroup. That handler lists Ingresses
and returns Ingress keys, so when a ProxyGroup became Available the
requests it produced never matched a Service and the reconciler's Get
just came back NotFound.
This fixes this by adding servicesFromIngressProxyGroup, which lists the
Services indexed for the ProxyGroup and returns their keys, matching what
the egress path already does with egressSvcsFromEgressProxyGroup.
Fixes#20944
Signed-off-by: chaosinthecrd <tom@tmlabs.co.uk>