The natlab-basic workflow builds the gokrazy natlab image in its own step so that the test's rebuild of it (vmtest always rebuilds, so the baked-in binaries match the source under test) is a build cache hit rather than a cold build inside go test's -timeout budget. That never worked: the Makefile ran whatever "go" was on $PATH, the runner's stock Go, while the go command puts its own $GOROOT/bin first on the test binary's $PATH, so the rebuild from inside "go test" used tailscale/go. GOCACHE entries embed the compiler's build ID, so the step warmed nothing. In practice the in-test rebuild took about 2.5 minutes of the 3 minute -timeout, leaving TestEasyEasy about 20 seconds for booting two VMs, logging in, and pinging. A passing run on main took 167s. Any hiccup in the remaining budget, such as the "tailscale up" hang fixed separately, ended in go test's timeout panic with no useful output. Make the natlab targets in gokrazy/Makefile use ../tool/go so the step and the test use the same toolchain and cache. Fix the same mistake in natlab-test.yml's cache warming step, whose comment documented the wrong belief about which toolchain the in-test builds use. Raise natlab-basic's -timeout to match natlab-test.yml so that a hang fails through vmtest's own bounded waits (which dump the node's logs) instead of through go test's timeout panic (which dumps nothing about the VMs). Updates #13038 Updates #deflake Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: Iaa6085ec5aa029373204baf75b169ff375c2b355
Tailscale Appliance Gokrazy Image
This is (as of 2024-06-02) a WORK IN PROGRESS (pre-alpha) experiment to package Tailscale as a Gokrazy appliance image for use on both VMs (AWS, GCP, Azure, Proxmox, ...) and Rasperry Pis.
See https://github.com/tailscale/tailscale/issues/1866
Overview
It makes a ~70MB image (about the same size as
tailscale-setup-full-1.66.4.exe and smaller than the combined
Tailscale Android APK) that combines the Linux kernel and Tailscale
and that's it. Nothing written in C. (except optional busybox for
debugging) So no operating system to maintain. Gokrazy has three
partitions: two read-only ones (one active at a time, the other for
updates for the next boot) and one optional stateful, writable
partition that survives upgrades (/perm/)
Initial bootstrap configuration of this appliance will be over either serial or configuration files (auth keys, subnet routes, etc) baked into the image (for Raspberry Pis) or in cloud-init/user-data (for AWS, etc). As of 2024-06-02, AWS user-data config files work.
Quick start
Install dependencies:
$ brew install qemu e2fsprogs
Build + launch:
$ make qemu
That puts serial on stdio. To exit the serial console and escape to
the qemu monitor, type Ctrl-a c. Then type quit in the monitor to
quit.
Building
make image to build just the image (tsapp.img), without uploading it.
UTM
You can also use UTM, but the qemu path above is easier. For UTM, see the UTM instructions.
AWS
Build an AMI
go run build.go --bucket=your-S3-temp-bucket to build an AMI.
Credentials come from the AWS SDK's default chain, so authenticate any way it
recognizes: aws sso login, aws configure, an AWS_PROFILE,
AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY env vars, or
aws-vault exec <profile> -- go run build.go --bucket=... (aws-vault injects
temporary credentials as env vars). If no credentials are found the build stops
with a message telling you how to log in.
Creating an instance
When creating an instance, you need a Nitro machine type to get a
virtual serial console. Notably, that means the t2.* instance types
that AWS pushes as a free option are not new enough. Use t3.* at least.
As of 2024-06-02 this builder tool only supports x86_64 (arm64 should be trivial and will come soon), so don't use a Graviton machine type.
To connect to the serial console, you can either use the web console, or use the CLI like:
$ aws ec2-instance-connect send-serial-console-ssh-public-key --instance-id i-0b4a0eabc43629f13 --serial-port 0 --ssh-public-key file:///your/home/.ssh/id_ed25519.pub --region us-west-2
{
"RequestId": "a93b0ea3-9ff9-45d5-b8ed-b1e70ccc0410",
"Success": true
}
$ ssh i-0b4a0eabc43629f13.port0@serial-console.ec2-instance-connect.us-west-2.aws
Configuring the appliance
The appliance's tailscaled runs with -config=optional:vm:user-data, so a
single AMI supports two ways of joining a tailnet:
-
Declarative (user-data): put a Tailscale config (the
alpha0HuJSON format) in the instance's user-data and the node configures itself on first boot. The minimal config is just an auth key:{ "Version": "alpha0", "AuthKey": "tskey-auth-..." }A config present in user-data locks the CLI (
tailscale set/upare rejected) unless it sets"Locked": false. Add"RemoteConfig": trueto hand full remote management of the node to the tailnet admin (seePrefs.RemoteConfig) — appropriate for admin-owned fleet devices. -
Interactive (serial console): launch the AMI with no user-data. The
optional:prefix means the missing config is not an error, sotailscaledboots unconfigured and you can enroll it over the serial console (connect as above, then runtailscale upand open the printed login URL).
To require config instead (fail to boot if none is present), build an image
whose tailscaled uses -config=vm:user-data without the optional: prefix.