The derper debug pages had no way to see which clients were connected. The expvar gauges only give counts, /debug/check only says whether the counts agree, and /debug/traffic only reports connections that moved bytes since its last tick, and only if ss is installed. Add /debug/clients/, which by default serves an index page with a form to pick one of four filters: ?all lists every connection, ?ip=1.2.3.4 and ?cidr=1.2.0.0/16 list connections from an address or prefix, and ?key=nodekey:... lists the connection(s) for one node key. Each row shows the connection number, key, remote address, connection age, flags (home, mesh, prober, notideal, dup/active/disabled), protocol version, app name, per-connection rx/tx packet and byte counts, and the estimated unique sender count. Big derpers have far too many connections for one page, so results are paginated with keyset cursors rather than page numbers: sort=key, ip, conn, rx, tx, rxpkts, or txpkts (with a leading - for descending) picks the walk order, limit=N the page size, and after=X resumes after that value of the sort field. The next-page links add afterconn=N so a page boundary that falls among connections sharing a value (duplicate keys, one IP with many ports, equal counters) resumes exactly. Column headers link to the other sort orders. The walk under Server.mu does only a filter match, a cursor comparison, and at most a bounded-heap operation per connection, so connections before the cursor are discarded without being copied and at most limit entries are ever kept. Only the summary counts (matching connections and keys) look at every connection. Snapshots are taken and the page rendered after the lock is released, so a slow debug client can't stall the server. A benchmark with 100k connections takes about 10ms per page. There were no per-connection traffic counters before, only the server-wide ones, so sclient gains four atomic.Uint64 counters (rx/tx packets and bytes, counting data packets like the server-wide ones) bumped alongside them. That's 32 bytes per connection. For the counter sorts, the value is loaded once per connection during the walk and used for both the cursor test and the heap order, so the order stays consistent while the counters keep changing. The sclient preferred field becomes an atomic.Bool so the page can report which connections are the client's home DERP; it was previously only touched by the run goroutine. Updates tailscale/corp#48933 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: I4e9b7c2d5a83f61b0e7d2c94a5f8b3e16d7c0a29
DERP
This directory (and subdirectories) contain the DERP code. The server itself is
in ../cmd/derper.
DERP is a packet relay system (client and servers) where peers are addressed using WireGuard public keys instead of IP addresses.
It relays two types of packets:
-
"Disco" discovery messages (see
../disco) as the a side channel during NAT traversal. -
Encrypted WireGuard packets as the fallback of last resort when UDP is blocked or NAT traversal fails.
DERP Map
Each client receives a "DERP Map" from the coordination server describing the DERP servers the client should try to use.
The client picks its home "DERP home" based on latency. This is done to keep costs low by avoid using cloud load balancers (pricey) or anycast, which would necessarily require server-side routing between DERP regions.
Clients pick their DERP home and report it to the coordination server which shares it to all the peers in the tailnet. When a peer wants to send a packet and it doesn't already have a WireGuard session open, it sends disco messages (some direct, and some over DERP), trying to do the NAT traversal. The client will make connections to multiple DERP regions as needed. Only the DERP home region connection needs to be alive forever.
DERP Regions
Tailscale runs 1 or more DERP nodes (instances of cmd/derper) in various
geographic regions to make sure users have low latency to their DERP home.
Regions generally have multiple nodes per region "meshed" (routing to each
other) together for redundancy: it allows for cloud failures or upgrades without
kicking users out to a higher latency region. Instead, clients will reconnect to
the next node in the region. Each node in the region is required to be meshed
with every other node in the region and forward packets to the other nodes in
the region. Packets are forwarded only one hop within the region. There is no
routing between regions. The assumption is that the mesh TCP connections are
over a VPC that's very fast, low latency, and not charged per byte. The
coordination server assigns the list of nodes in a region as a function of the
tailnet, so all nodes within a tailnet should generally be on the same node and
not require forwarding. Only after a failure do clients of a particular tailnet
get split between nodes in a region and require inter-node forwarding. But over
time it balances back out. There's also an admin-only DERP frame type to force
close the TCP connection of a particular client to force them to reconnect to
their primary if the operator wants to force things to balance out sooner.
(Using the (*derphttp.Client).ClosePeer method, as used by Tailscale's
internal rarely-used cmd/derpprune maintenance tool)
We generally run a minimum of three nodes in a region not for quorum reasons (there's no voting) but just because two is too uncomfortably few for cascading failure reasons: if you're running two nodes at 51% load (CPU, memory, etc) and then one fails, that makes the second one fail. With three or more nodes, you can run each node a bit hotter.