LocalBackend named appc.AppConnector in its appConnector field, in the
field's constructor call, and in the AppConnector accessor, so even
builds with ts_omit_appconnectors linked the appc package and its
dependencies. The buildfeatures.HasAppConnectors checks let the linker
drop the code but could not remove the import.
Add a build-tag-gated type for the field: appcAppConnector is an alias
for appc.AppConnector by default, and in ts_omit_appconnectors builds it
is an empty struct with no-op versions of the methods local.go calls, so
the shared code still type checks. The constructor call moves behind a
newAppConnector helper in the same gated file, and the AppConnector
accessor keeps its *appc.AppConnector signature there, with a nil-
returning version in the omitted build.
This takes tailscale.com/appc out of depaware-min.txt and
depaware-minbox.txt. Add a DepChecker test to keep it out; it also sets
ts_omit_conn25, since conn25 uses appc as well.
Updates #12614
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ia7e2d9c4f1b8e6a3d5c0f2b7e9a1c4d6f8b0e2a5
While a VM reboots, tests poll its agent roughly every 5s. Each
request that times out leaves behind a goroutine blocked in
Server.takeAgentConn, waiting for TTA to connect back. Several can
pile up before a request succeeds and the test ends, and the rest
wait forever.
Making takeAgentConn return at Server shutdown unblocks any goroutine
still waiting when the test ends.
In a full vmtest run, the two dials stranded by the reboot in
TestGokrazyUpdatesItselfToSameImage logged until the package exited
10 minutes later. With this change, "still waiting" log lines in the
same run drop from 267 to 72, and the longest wait goes from 10m10s
to 40s.
Updates #cleanup
Change-Id: Ib4c1d2ce8f2a7b73d3f0c1e6a9e5b2d8f4a1c7e3
Signed-off-by: Francois Marier <francois@tailscale.com>
Only gokrazy builds trust the fake log catcher's TLS certificate, so
Ubuntu, Debian, Fedora and FreeBSD guests failed every upload and the
failure dump of their tailscaled logs was always empty. vnet now also
serves the log catcher over plain HTTP on port 80, and the harness points
those guests at it with TS_LOG_TARGET. Env.NodeLogs returns what a node
has uploaded.
macOS guests still do not upload: tailmac does not pass TailscaledEnv to
tailscaled, so the setting would not reach it.
Updates #13038
Signed-off-by: Brendan Creane <bcreane@gmail.com>
A JSON null for a service or for an endpoint target unmarshals as a nil
pointer, and loadConfigV0 dereferenced it without a check. This made
"tailscale serve set-config" panic. Return an error that names the
service or endpoint instead.
Fixes#21534
Signed-off-by: Brendan Creane <bcreane@gmail.com>
TestSelfSignedDERPHashPinning occasionally fails on a slow vm with:
tailscale up (limit 1m30s): Get "http://unused/up?accept-routes=true": EOF
Because this endpoint has both IPv4 and IPv6 addresses, TTA will fall
back to the second address when running on a slow host and race both
connections. The losing connection will get closed by net.Dialer and
never produce a response.
By keeping the gVisor endpoint with the conn, we can avoid returning
closed connections in takeAgentConnOne().
This fix uncovered a bug in Env.Tailscale(): calling "tailscale
update" via a GET request may be replayed by net.http.Transport and
that will trigger a second update undoing the first. Using POST
disables all such retries.
Update #cleanup
Change-Id: I2f8b7c4d9e1a3b5c7d0e2f4a6b8c1d3e5f7a9b0c
Signed-off-by: Francois Marier <francois@tailscale.com>
dnsConfigForNetmap called appc.AppDNSRoutes directly to synthesize the
conn25 split DNS routes from the self node's capability map, which was
the one place node_backend.go depended on the appc package and was
marked with a TODO to make it a hook.
Add ipnext.Hooks.ExtraDNSRoutes, which returns extension-supplied split
DNS routes. authReconfigLocked calls it and passes the result into
dnsConfigForNetmap, which adds the routes alongside the netmap's own
routes with the same UseWithExitNode handling as before.
The hook takes no arguments: the conn25 extension computes the routes
in its OnSelfChange handler and clears them on a switch to a different
node, and it already tracks the AppConnector.Advertise pref from
ProfileStateChange. Both of those hooks fire before the reconfig that
consults ExtraDNSRoutes, so the state is current. The conditions are
unchanged: no routes when this node itself advertises as an app
connector, and none without the experimental capability.
Updates tailscale/corp#37125
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I6c2e9f4b8a1d3e7f5c0b2a9d8e6f4c3b1a7d5e9f
This field populates RegisterRequest.Expiry, which
controlclient.Direct.doLogout fills with time.Unix(123, 0) to force the
deletion of the device from the server.
Updates tailscale/corp#48031
Signed-off-by: Paul Scott <408401+icio@users.noreply.github.com>
The NL prefix is one of the last references to the "Network Lock" name
which we're cleaning up.
Updates tailscale/corp#37904
Change-Id: I751194e0b49447e97a65ed6975e3da680226ce85
Signed-off-by: Alex Chan <alexc@tailscale.com>
Building cigocacher from the corp repo failed because its go.mod was
missing a reference to go-tool-cache. Migrate cigocacher over to tb so
it's not required.
Updates tailscale/corp#49137
Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com>
logtail's dialer only used dnscache in its fallback path, and built a
new Resolver per dial there, so every new connection to the log server
did a fresh DNS lookup (twice when the first dial failed). On systems
without a caching resolver, that adds up to a lot of queries for
log.tailscale.com.
Share one dnscache.Resolver across all log dialers, and have the
bootstrap path go straight to dnsfallback rather than repeating the
regular lookup. Since the cache TTL is a fixed guess, make
dnscache.Dialer expire a cached entry and look the host up again when
dialing its cached IPs fails, so a moved host isn't stuck on stale IPs.
Updates #15326
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I6a55d72a73ed878faebc15317421a54d9ae293ca
Allow TS_DEBUG_DISABLE_HOSTS_FILE_UPDATES to opt out of Windows hosts entries locally, using the existing tailscaled-env.txt configuration mechanism. Keep the control-plane node attribute authoritative and preserve upstream behavior when the local option is unset or false.
Cover defaults, local opt-out, nil control knobs, and policy precedence with Windows tests. This change is limited to DNS configuration; no route changes or upstream route-fix cherry-pick are included.
Updates tailscale/tailscale#14327
Signed-off-by: Caleb Crome <caleb.crome@flyzipline.com>
The agent and tailscaled are started concurrently by gokrazy and
tta can win that race, in which case it answers 502 Bad Gateway
because tailscaled.sock does not exist yet.
This is most likely to be seen on slower machines using TCG.
Updates #cleanup
Change-Id: I4012bd189adfb2565470c814204fcf18545f1017
Signed-off-by: Francois Marier <francois@tailscale.com>
updates tailscale/corp#33007
When a selected exit node can't carry internet traffic due to a misconfiguration (wrong ID, node
deleted, routes removed, etc), blackhole default routes are installed and all internet traffic is dropped.
That is the correct behavior -- better to drop than leak to the local network -- but it was entirely silent:
the stable-ID lookup in nodeBackend.updateRouteManagerPrefs misses without a log line, and Status
leaves ExitNodeStatus nil. The user sees a healthy Tailscale with no internet. If an admin fat-fingers the
exit node name in an IT policy, for example, it's easy to break every node with zero feedback.
This adds an exit-node-unavailable health warning covering every way the selection can fail to carry
traffic (short of reachability which is a separate concern), reported via exitnodehealth.ArgExitNodeReason.
The warning names the exit node, caching its display name while it is still a peer so the name survives
its departure, and falls back to the stable ID or IP. When the selection is mandated by the ExitNodeID
or ExitNodeIP policy settings, the message tells the user to contact their network administrator instead
of suggesting they pick another exit node.
To test this, set a forced exit node policy with some random ip or node id. The warning has a 5 second
threshold. It should clear as soon as you change to a proper exit node.
Signed-off-by: Jonathan Nobels <jonathan@tailscale.com>
Previously ProxyClass spec.staticEndpoints was only implemented for
ProxyGroups and silently ignored when the ProxyClass was referenced by
a Connector. Extract the NodePort Service provisioning, port
allocation, and node ExternalIP discovery logic shared with ProxyGroup
into a new k8s-operator/reconciler/staticendpoints package, and wire
it into the StatefulSet reconciler so that each Connector replica gets
a per-pod NodePort Service, the discovered endpoints are written to
its tailscaled config as static endpoints, and tailscaled listens on
the Services' target port via the PORT env var. Also reconcile
Connectors on Node changes, clean up NodePort Services on scale down,
Connector deletion, and when static endpoints are removed, and
surface the endpoints in the Connector's device status.
Add e2e tests for static endpoints on both Connectors and ProxyGroups.
They verify the per-replica NodePort Services, the PORT env var, the
endpoints reported in status and advertised to control, and the
cleanup on scale down and on removal of the static endpoints
configuration. The tests skip on clusters whose Nodes have no
ExternalIP addresses, such as kind.
Fixes#18819
Change-Id: Ia134916062471e6c0088692834257aafabd0e713
Signed-off-by: David Bond <davidsbond93@gmail.com>
Compare parsed each numeric field with strconv.ParseUint and panicked
if that failed, so a version with a field of 20 or more digits, such as
a long build stamp, crashed the caller.
Compare digit runs by length after trimming leading zeros, then
lexically. This gives the same results as before for numbers that fit
in a uint64 and works for numbers of any length.
Fixes#21532
Signed-off-by: Raphael Fakhri <153192858+RaphaelFakhri@users.noreply.github.com>
Several features with ts_omit build tags were still linked into
minimal builds because LocalBackend called into them unconditionally,
so the linker couldn't prove them unreachable. Guard those call sites
with their buildfeatures constants:
* app connectors: only subscribe to route update and store events
when app connectors are included; appc is their only publisher.
This also drops AdvertiseRoute and UnadvertiseRoute.
* client update: only subscribe to the tailnet default auto-update
event when auto-updates are supported, as that's all it affects.
* syspolicy: only register the policy change watch when system
policy support is included.
* advertise routes and exit nodes: only validate AdvertiseRoutes
when one of the two features is included.
* netstack: only netstack consults the intercepted TCP ports, so
don't build the port matching func without it.
* ssh: don't check SSH prefs, report the "SSH on but unusable"
health message, or intercept port 22 without SSH support. The
RunSSH pref is still accepted without effect, as before; port 22
was previously intercepted by netstack builds without SSH.
* web client: don't track whether the web client should run when
it's not included.
* tailnet lock: peers marked UnsignedPeerAPIOnly are only special
in that tailnet lock skips their signature checks, and the packet
filter check guards that exemption. Without tailnet lock, all
peers from control are trusted anyway, so skip the check.
This reduces the size of a build_dist.sh --min tailscaled binary on
linux/amd64 by 65,536 bytes (0.5%), about 28 KB of it in ipnlocal and
most of the rest in eventbus and slices generic instantiations for the
events no longer subscribed to.
Updates #12614
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I8f4d4c0779e84f46f3b076d0fb35514c69a52d8f
Add vmtests that boot pairs of gokrazy VMs whose tailscale and tailscaled
are minimal builds and check that IP traffic flows between them over
WireGuard, so we notice if a minimal build stops doing its one job:
build_dist.sh --extra-small (direct path through NATs), that minus NAT
traversal (DERP across NATs, direct on a shared LAN), and DERP-only.
They found that "tailscale up" built with ts_omit_ipnbus returned before
calling Start or EditPrefs, making it a no-op. Fix that, and have it wait
for the Running state by polling tailscaled's status instead of watching
the IPN bus, so scripts can rely on it as with normal builds.
The extra-small feature list moves to featuretags.ExtraSmall so
build_dist.sh (via a new "featuretags --extra-small") and vmtest share
it. The gokrazy builder gains --output and --go-build-tags flags to build
the variant images, and natlabprep --gokrazy prebuilds them all for CI.
Updates #13038
Updates #12614
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I0caa2c7b3adf2399da09b6d07a942f1dc34e8486
Previously, ipconfig /flushdns and ipconfig /registerdns were called
from multiple sites without coordination: the DNS manager's SetDNS
(in a fire-and-forget goroutine), the router's winRouter.Set (synchronously
on every route change), the link change handler, and session unlock.
These could fire concurrently or in rapid succession, spawning multiple
ipconfig processes and putting excessive pressure on the Dnscache service.
In this PR, we introduce a new util/coalescedop.CoalescedOp type that provides
"one in-flight, at most one pending" semantics: Do() is non-blocking,
at most one execution runs at a time, and concurrent/rapid successive
requests coalesce into a single follow-up execution.
We then use CoalescedOp for both ipconfig /flushdns and ipconfig /registerdns,
replacing the fire-and-forget goroutine in windowsManager.SetDNS
and in the session unlock handler in tailscaled.
We also remove the redundant dns.Flush() call from winRouter.Set, since every
meaningful router.Set is followed by dns.Set in Reconfig, which already
flushes.
Updates #21001
Signed-off-by: Nick Khyl <nickk@tailscale.com>
Auto.mapRoutine sleeps in BackOff with c.mu held after PollNetMap fails.
Pausing, restarting the map poll, and shutting down need the same mutex
to cancel that wait, so they can be delayed by the retry timer.
Release c.mu after updating inMapPoll and snapshotting paused. PollNetMap
waits for its map-session callbacks before returning, so accessing the
backoff here cannot race with UpdateFullNetmap resetting it.
Add a synctest regression that checks mutex availability during backoff
and verifies that pausing interrupts the wait without advancing time.
Fixes#21132
Change-Id: I2d5b5a71c5857be60507d4d6ec9e2e0fbd8cf40a
RELNOTE: Reduce pause and shutdown delays during control connection retries.
Signed-off-by: Noah Ingwers <98993329+noah-ing@users.noreply.github.com>
The egress Pods readiness reconciler decided whether a Pod could route
traffic by calling each egress Service's health check through its
ClusterIP up to replicas*3 times and looking for a response from that
Pod. That is O(replicas) calls per Service per reconcile, and a Pod
joining a large ProxyGroup is only sampled 3 times on average, so late
Pods often miss and wait for another 5s requeue. The call only goes
through kube-proxy on the operator's own node, so it does not show that
other nodes route to the Pod either.
egress-eps-reconciler already adds a Pod to an egress Service's
EndpointSlice only once the Pod's state Secret shows the proxy has set up
routing for that Service. This sets the readiness condition once the Pod
is an endpoint in an EndpointSlice of every egress Service for the
ProxyGroup, and removes the health check calls.
Updates tailscale/corp#39464
Signed-off-by: chaosinthecrd <tom@tmlabs.co.uk>
cmd/featuretags --min, cmd/tsconnect/wasmbuild, and a cmd/tailscaled
test each computed "omit every feature except these" on their own.
Move that into featuretags.MinTags so other Go code (such as upcoming
natlab vmtests of minimal builds) can compute the tags without running
cmd/featuretags. No behavior change.
Updates #12614
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ie0a5faf1f1d3e4e035326572b16f20180671d1c4
When quad-100 must be installed as the OS primary resolver (Apple Mode
B), its catch-all forwarders were blended from base-config nameserver
IPs alone: non-standard ports, search domains, and DoH/DoT endpoints
were lost, and a base read that returned no resolvers silently
installed a catch-all with no forwarders at all, breaking all public
DNS while reporting a healthy DNS configuration.
Introduce OSConfig.Resolvers, which OSConfigurators able to recover the
underlying resolver configuration at full fidelity populate for the DNS
manager only: it never reaches the OS, and Equal ignores it so
configuration application is unaffected. Blend it into the catch-all
route verbatim, falling back to plain IP:53 resolvers derived from
Nameservers.
Teach the forwarder to dial arbitrary https:// resolvers instead of
only well-known public providers: at their BootstrapResolution
addresses when present, or at the URL's own host when that is an IP
literal, so enterprise DoH endpoints recovered from the OS can be
forwarded to over DoH rather than plaintext DNS.
Reject an empty base config on Apple Mode B: keep the previous
configuration, use the upstream empty-base health warning and medium
severity, and let an extension-triggered recompile retry instead of
installing a catch-all that cannot forward. Expose the same missing-
resolver error to OS configurators so bridge-reported absence uses this
warning too, while genuine configuration-read failures remain distinct.
Extend the Apple mode tests with empty-base and full-fidelity blend
coverage, and update the iOS primary-mode cases to model the real
NetworkExtension base read (LAN resolvers and search domains) rather
than the silently empty catch-all.
RELNOTE: Fix silent public DNS breakage when Tailscale must be the system DNS resolver on Apple clients.
Updates #20341
Updates tailscale/corp#45534
Updates tailscale/corp#48693
Change-Id: Ie7e277b0fd39d2b91b1e770f8f01d688e02b4d49
Signed-off-by: James Tucker <james@tailscale.com>
Apple tunnels derive global search domains from match domains. Scope
simple iOS and sandboxed macOS configurations to the full MagicDNS route
set; keep custom split resolvers and uncovered forward/PTR records behind
a primary resolver. Respect scoping controls and reject unsafe fallback
when base DNS cannot be read.
Add mode-selection and reverse-coverage tests for both Apple platforms.
RELNOTE: Fix MagicDNS routing and split-DNS search scoping on Apple clients.
Updates tailscale/corp#48693
Change-Id: I50b5641db466f77c19f911255748e12a10d0a5fe
Signed-off-by: James Tucker <james@tailscale.com>
mkversion derives the version from git history, which a shallow clone
lacks (the change count silently comes out as 0) and a tree without
.git can't provide at all. If TS_VERSION_LONG is set, InfoFrom now
takes the version from it plus TS_VERSION_GIT_HASH (and the EXTRA/DATE
variants), runs no git, and derives everything else through the same
mkOutput as the git path. The result's Long must reproduce the input
exactly or InfoFrom errors.
Updates tailscale/corp#44945
Change-Id: I9c1247ea5d1c6c52c823c2205447b5c7ee8e778c
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
The egress Pod readiness reconciler polls each Service's health check
endpoint up to replicas*3 times per Pod. A freshly created Pod is initially
unreachable, so these backoff log lines carry no value and generate
significant noise.
Discard these logs so the per-poll message is no longer logged (the more
useful message "Pod is not yet added as an endpoint for all egress targets,
waiting..." is still logged once per reconcile).
Updates tailscale/tailscale#21079
Signed-off-by: Becky Pauley <becky@tailscale.com>
LowMemory shrank the pending ring buffer, upload batch size, and text
truncation limit for iOS, back when network extensions were limited to
15 MB. All supported iOS versions now allow 50 MB, so the savings are
negligible and the extra code paths are not worth maintaining.
Fixes#13685
Change-Id: Ic0c26f899ce00f105399c7b1e687e2b92194687e
Co-authored-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Signed-off-by: Andrea Gottardo <andrea@gottardo.me>
infoFromDir passed --format=%%ct to git log, which git renders as the
literal string "%ct", so VersionInfo.GitDate was never a timestamp for
builds made directly from a tailscale.com checkout. The corp path in
infoFromCache already used the correct %ct.
Updates tailscale/corp#44945
Change-Id: I28d193019a8dd3d2bdd165da5e4584df1964afd6
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
There were a handful of inconsitencies in natlab with the biggest being.
- Pair and grid tests living under tstest/integration and using the old
test style.
- Inconsistent use of e vs env.
- Methods for actions on node living on the vnet.Node rather than being
a method on vmtest.Env
Clean up these inconsitencies.
Updates #cleanup
Signed-off-by: Claus Lensbøl <claus@tailscale.com>
It's possible for multiple direct connections to be counted and
these tests are successful as long as it's more than one.
Fixes#21578
Change-Id: I979efc92a43f78d0e51d602d96e61828e6fabdff
Signed-off-by: Francois Marier <francois@tailscale.com>
Select the published reco-backed test server with a build tag while
preserving the default implementation and existing caller API. Pin the
public module with its proxy-verified checksums; no workspace is needed.
Accept incremental peer messages in the TSP test. Exercise legacy map
semantics with the default server and version-floor rejection with reco,
including current-version streaming, without manual test skips.
Document GOFLAGS usage so child Go builds inherit the experiment tag.
Updates tailscale/corp#49097
Change-Id: I9b3e086c4c6a9a57a076e76b883422875c843f20
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Peer-change watchers upsert PeersChanged, so peers missing from a full
netmap were never removed from their view.
Updates #15660
Signed-off-by: Kristoffer Dalby <kristoffer@tailscale.com>
A peer removal sent with any MapResponse field that forces a full netmap
rebuild never reaches IPN bus watchers; these tests fail.
Updates #15660
Signed-off-by: Kristoffer Dalby <kristoffer@tailscale.com>
Add three things to the connected clients debug page:
app=NAME narrows any of the existing filters to connections that
advertised that app name, and may be repeated to match any of several
(app=tailcat-server&app=tailcat-client). On its own it applies to all
connections. An empty app= matches connections that sent no app name.
App names in the table link to their filter, and the next-page and
sort links carry the app filter along.
sort=connected walks by connection time, ascending being longest
connected first, with -connected for newest first. The next-page
links use the connection time in Unix nanoseconds as the cursor; by
hand, after= also accepts a duration such as 30m, meaning connections
that have been up that long, which is the natural way to ask for
"everything older than half an hour".
format=json returns the page as a JSON object with the filter
description, the matching connection and key counts, how many
connections remain after the page, the next page's relative URL, and
the client rows, so the page can be walked from curl or a script the
same way a browser follows the next links. Rows gain a connectedAt
timestamp alongside the rounded connected duration.
Updates tailscale/corp#48933
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: Ib5d2e7c40a9f13e8a6d7c2b5f9e0a4d3c8b17e62
Shutting down a socket's read side after its peer's FIN has arrived does
nothing on the wire and fails with ENOTCONN on macOS (and on Linux once
the socket reaches TIME_WAIT), so every cleanly half-closed connection
was logged as a failure and TestTCPHalfClose failed on macOS. Only
half-close the destination after EOF.
When one direction fails, one of the two sockets is dead. If the failure
was writing to it, the other direction is reading from that dead socket
and finishes on its own once it has drained what arrived before the
failure, so leave it alone rather than truncate the data. If the failure
was reading from it, the other direction may be blocked reading the live
peer, so close both connections to unblock it; nothing deliverable is
lost in that case.
Telling the two apart means wrapping the destination, which disables the
kernel splice fast path when both ends are bare TCP connections (about a
third of relay throughput on Linux loopback). That is accepted here: the
common userspace-networking and proxymux configurations wrap the
connections anyway, so they never had the fast path.
The regression was introduced by 027e249fcf (#21359), merged 2026-09-17.
Tested on Linux and on macOS 27 (arm64): the old test failed 49 of 50
runs on macOS; the new tests pass 200 iterations there and pass under
the race detector on Linux.
Fixes#21522
Updates #20883
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I7c2e4a9d1b8f5e3c6a0d2f4b8e1c9a7d3f5b6e2c
conn25.go had grown to hold the extension, the Conn25 type, the config,
the DNS rewriting, and both the client and connector implementations.
Move the client struct and its methods to client.go, and the connector
struct and its methods to connector.go.
appAddr, transitIPExpiryEntry, and connectorTransitIPExpiry are used
only by the connector, so they move to connector.go. addrs moves to
client.go because it is a client-side concept, even though the
extension's send path in conn25.go and addrAssignments.go also refer
to it. Everything else stays in conn25.go.
Test functions covering client and connector behavior move to
client_test.go and connector_test.go correspondingly.
This is a pure code move: no declaration is added, removed, or edited,
and the relative order of everything is unchanged.
Updates #cleanup
Signed-off-by: Michael Ben-Ami <mzb@tailscale.com>
When the tailscale.com/proxy-group annotation is removed from a Service
that was exposed on a ProxyGroup, the HA Service reconciler returned early
on the empty annotation before reaching its cleanup path. The Tailscale
Service was left advertised and the operator's finalizer was never removed,
so the Tailscale Service leaked and the Kubernetes Service wedged forever in
Terminating once deleted.
The annotation is the only place the ProxyGroup name was recorded, and it is
gone by the time cleanup needs it. This commit encodes the ProxyGroup name in
the finalizer, which survives any modification of the resources annotations.
On reconcile, an empty annotation with our finalizer present now recovers the
ProxyGroup name from the finalizer and runs cleanup. Services carrying the bare
legacy finalizer are migrated to the encoded form while the annotation is still
present.
Updates #19922
Signed-off-by: chaosinthecrd <tom@tmlabs.co.uk>
Contining the trend of making everything be optional that could
possibly be optional, this makes the magicsock UDP underlay transport
be optional. (mostly replacing a bunch of GOOS != "js" checks in the process)
Omitting udptransport (ts_omit_udptransport) makes magicsock DERP-only: no
UDP sockets, no advertised endpoints, and netcheck measures DERP latency
over HTTPS. Omitting only nattraversal keeps disco ping/pong to peers'
advertised endpoints but drops STUN, call-me-maybe, the peer relay client,
the endpoint tracker, and UDP lifetime probing.
Both are plain linker dead-code elimination via buildfeatures
constants. (as oppposed to moving the code all over into feature
packages and indirecting through hooks) The minimal linux/amd64
tailscaled shrinks by 258 KB without nattraversal and 561 KB without
both.
Updates #12614
Change-Id: Ife10804b48a168d9481e4321c18ed8fc85fa71f1
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
After copying from the backend to the client until EOF, forwardTCP shut
down the backend socket's read side. That does nothing on the wire, and
once the backend's FIN has arrived, macOS rejects shutdown(SHUT_RD) with
ENOTCONN (Linux does too once the socket reaches TIME_WAIT), so every
forwarded connection logged "backend -> client close connection: ...
socket is not connected". Keep only the CloseWrite calls, which are what
propagate the half-close.
The regression was introduced by 04d24cdbd4 (#16462), merged 2025-07-07.
Tested on Linux and on macOS 27 (arm64): the wgengine/netstack tests
pass on both. The same ENOTCONN failure was reproduced on macOS via the
identical pattern in net/socks5 (#21522).
Updates #21522
Updates #16462
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I3e9b1d7f4a2c8e6b0f5d9a3c7e1b4f8d2a6c0e5b
WriteBatchTo copied the destination IP and port into the pooled
net.UDPAddr handed to sendmmsg(2), but not the zone. For an IPv6
link-local destination the sockaddr then had no scope ID, and Linux
failed the send with "cannot assign requested address". Disco goes
through WriteToUDPAddrPort, which keeps the zone, so magicsock could
pick a link-local path whose WireGuard packets, sent via WriteBatchTo,
never left the host.
Set the zone on every write, so that a destination without one also
clears the zone left over from an earlier write that reused the
pooled address.
Fixes#21411
Co-Authored-By: Jordan Whited <jordan@tailscale.com>
Signed-off-by: chlee <sourcehatchery@gmail.com>
Signed-off-by: Jordan Whited <jordan@tailscale.com>
This adds a tailscaled string flag "--windows-mode" which accepts two
possible values: the empty string (default) to get the normal behavior
(tailscaled running as an admin, usually as a service), and "dev", to
make the safesocket named pipe path be at a location that regular
users (non-admins) can create.
It then modifies the safesocket client side (as used by the CLI) to
try the dev mode path too on failure.
This lets people work on tailscaled.exe+tailscale.exe in a terminal
easily during development, as either an admin or non-admin. (This
used to work prior to the move away from TCP localhost to named pipes
for safesocket on Windows)
Updates #2791
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I9c1e3a5b7d0f2a4c6e8b0d2f4a6c8e0b2d4f6a8c
The pipe's security descriptor named the Administrators group as the
pipe's owner (O:BA) and primary group (G:BA). Only administrators may
assign that SID as an owner, so a tailscaled run by a regular user (in
userspace-networking mode with its own --socket) died at startup with
"This security ID may not be assigned as the owner of this object". The
tests already had to swap in an empty descriptor to run unelevated.
Neither part did anything for us. The primary group is never consulted by
Windows access checks; it exists only for POSIX and NFS compatibility. The
owner grants exactly one thing beyond the DACL: the right to read and
rewrite the pipe's DACL. Whoever creates the pipe already has that in
practice. For the default pipe name under ProtectedPrefix\Administrators,
the creator must be an administrator (only administrators may create pipes
with that prefix), so naming Administrators as owner gave them nothing
new. A regular user creating a pipe of some other name runs the process
that owns it and controls its DACL anyway. Who may connect is decided only
by the DACL, which is unchanged: read and write for Users and LocalSystem.
So drop the owner and group and let the creator own the pipe. While here,
drop the AI, OI, and CI flags too: they concern inheritance to child
objects, which a named pipe doesn't have, and were copied from a file
system descriptor. The tests now use the real descriptor. Also give a
pointer when listening on the default ProtectedPrefix\Administrators pipe
fails with access denied.
Updates #2791
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I8a1c3e5f7b9d2e4a6c8b0d2f4e6a8c0b2d4f6a8c
* derp/derphttp: reject invalid DERP node hostname before proxy CONNECT
When a DERP client reaches a node through an HTTP(S) proxy,
dialNodeUsingProxy writes the CONNECT request by hand and puts
net.JoinHostPort(n.HostName, port) into both the request line and the
Host header. n.HostName comes from the control-supplied DERP map and
net.JoinHostPort does no sanitizing, so a hostname carrying CR/LF was
written verbatim into the plaintext request sent to the proxy. That let
whoever populated the DERP map inject extra headers, or a second
pipelined request, into the connection to the operator's proxy.
Validate n.HostName with httpguts.ValidHostHeader at the top of
dialNodeUsingProxy, before the proxy is dialed, and also reject the
empty hostname, which ValidHostHeader accepts. DNS names and IP
literals continue to work.
Add a table-driven test that runs accepted and rejected hostnames
against a fake proxy and checks the CONNECT target that goes out.
Fixestailscale/corp#48122
Signed-off-by: basavaraj-sm05 <basavaraj@digiscrypt.com>
Co-authored-by: Mike Jensen <mikej@tailscale.com>
Signed-off-by: Mike Jensen <mikej@tailscale.com>
The derper debug pages had no way to see which clients were connected.
The expvar gauges only give counts, /debug/check only says whether the
counts agree, and /debug/traffic only reports connections that moved
bytes since its last tick, and only if ss is installed.
Add /debug/clients/, which by default serves an index page with a form
to pick one of four filters: ?all lists every connection, ?ip=1.2.3.4
and ?cidr=1.2.0.0/16 list connections from an address or prefix, and
?key=nodekey:... lists the connection(s) for one node key. Each row
shows the connection number, key, remote address, connection age,
flags (home, mesh, prober, notideal, dup/active/disabled), protocol
version, app name, per-connection rx/tx packet and byte counts, and
the estimated unique sender count.
Big derpers have far too many connections for one page, so results
are paginated with keyset cursors rather than page numbers: sort=key,
ip, conn, rx, tx, rxpkts, or txpkts (with a leading - for descending)
picks the walk order, limit=N the page size, and after=X resumes after
that value of the sort field. The next-page links add afterconn=N so a
page boundary that falls among connections sharing a value (duplicate
keys, one IP with many ports, equal counters) resumes exactly. Column
headers link to the other sort orders.
The walk under Server.mu does only a filter match, a cursor comparison,
and at most a bounded-heap operation per connection, so connections
before the cursor are discarded without being copied and at most limit
entries are ever kept. Only the summary counts (matching connections
and keys) look at every connection. Snapshots are taken and the page
rendered after the lock is released, so a slow debug client can't
stall the server. A benchmark with 100k connections takes about 10ms
per page.
There were no per-connection traffic counters before, only the
server-wide ones, so sclient gains four atomic.Uint64 counters (rx/tx
packets and bytes, counting data packets like the server-wide ones)
bumped alongside them. That's 32 bytes per connection. For the counter
sorts, the value is loaded once per connection during the walk and
used for both the cursor test and the heap order, so the order stays
consistent while the counters keep changing.
The sclient preferred field becomes an atomic.Bool so the page can
report which connections are the client's home DERP; it was previously
only touched by the run goroutine.
Updates tailscale/corp#48933
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Change-Id: I4e9b7c2d5a83f61b0e7d2c94a5f8b3e16d7c0a29