Files
LocalAI/tests/e2e/distributed/cluster_peerlink_test.go
T
Ettore Di Giacinto 764fe46157 test(distributed): prove the worker tunnel end to end, under real inference
Everything this phase built was proven by unit and integration specs. This is
the first run of it against the real binaries: a frontend replica per process,
a worker that binds nothing routable, real inference over the result.

Four scenarios, each with the question "what would make this pass if the tunnel
were doing nothing" answered rather than left open.

A worker with no advertised address is reached through its tunnel. The roster
is asserted to report it advertising nothing, so there is no address a frontend
could have dialled instead, and node_connections is asserted to name the
replica that serves the request.

A request landing on the replica that does NOT own the worker is relayed to the
one that does. With N replicas behind round robin that is (N-1)/N of production
traffic, so it gets the FIRST request for its model: the backend install, the
file staging on the http tag, and the gRPC load and predict all cross the
relay. Which replica owns the tunnel is read from the ownership table through
the production Owner query and mapped to a frontend index through the address
the harness pins per replica; the non-owner is derived from that reading and
asserted to be a non-owner immediately before the request, rather than assumed
from the harness default. Sending the same request to the owner reddens it.

Killing the owning replica re-homes the worker onto the survivor. The worker
dials a balancer rather than a replica, because LOCALAI_REGISTER_TO is resolved
once at boot and is the tunnel endpoint as well as the registration one: aimed
at a single replica, a worker has nowhere to reconnect to when that replica
dies, and the re-home cannot happen at all. Removing the kill reddens it.

And the negative control for the whole suite, which is why the other three mean
anything. Frontend and worker share a host here, so every backend port the
frontend names in a stream target is one it could have dialled directly; if it
did, the first three would pass with the tunnel inert. LOCALAI_WORKER_TUNNEL is
no longer usable for this, because it is a fatal startup error and a worker that
never started says nothing about a worker reachable some other way. The balancer
answers the tunnel connect path itself instead, leaving a worker that registers,
heartbeats, reports healthy and holds no tunnel. It is asserted to have dialled
and been refused, asserted to be held by nobody, and then asserted unreachable
with the refusal naming the missing route. Then the block is lifted, nothing
else changes, and the same request succeeds: that is what attributes the refusal
to the tunnel rather than to any of the ordinary reasons an e2e inference fails.

The fifth spec measures the head-of-line blocking this phase deferred three
times. 128 MiB crosses the session while a warm model is probed back to back,
direct and relayed. Median latency is unchanged, the worst probe is about 3x the
baseline median and about a seventeenth of the transfer window, and the transfer
runs at 415-490 MB/s direct and 222-268 MB/s relayed. A session that
head-of-line blocked would park a probe for the length of the window. Leave the
yamux windows untuned; and note this is loopback, so it says the multiplexing
does not serialise and says nothing about a link with a bandwidth-delay product.

The load spec is measured against a control that the first version did not have.
It passed with the bulk artifact cut to 4 KiB, because the window it read probes
against was mostly cold-load overhead: it would have reported a clean bill on a
session carrying no large message. The same cold load now runs twice, once
empty and once bulk, and the difference between the windows is asserted to be
real before any latency is read from it.

Two defects on the base commit came out of this.

cluster_peerlink_test.go has been red since the relay landed, deterministically,
in isolation and in the suite. It asserted that an accepted peer stream is
refused at once, on the premise that phase 1 installs no relay. The relay
correctly waits fifteen seconds for a frame naming the worker, and the spec's
budget was five. It now writes a relay request for a node no replica holds and
asserts the refusal is ErrNotOwner and specifically not ErrNoConnection, which
is a stronger spec than the one it replaces and the only thing in the e2e suite
that exercises the relay's refusal path.

The harness handed a worker's own HTTP port to a backend process. It took two
ports from freeport and used one as the gRPC base and the other for the file
transfer server; freeport returns adjacent ports often, and the backend
allocator hands out base, base+1, base+2, so the second backend started on a
worker was regularly given the HTTP server's port and died with EADDRINUSE. No
spec had started two backends on one worker before, so it had never fired; the
load spec starts five and it failed about one run in three. Each worker now
reserves a contiguous bind-probed block laid out the way production lays it out,
below the kernel's ephemeral range, with LOCALAI_GRPC_MAX_PORT bounding the
allocator to it. The underlying production defect is not fixed here and is
recorded in the report: allocatePort never checks that a port is free, and its
default range overlaps the ephemeral range on every Linux box.

Constraint 6, whether distributed mode should now refuse to start without an
advertised address, is DEFERRED, and the comment and the docs that described the
cost were understating it. A replica with no advertised address writes no
instances row, and Owner joins a connection against a live instance, so a worker
whose tunnel lands there is unroutable from every OTHER replica while being
registered and healthy. Refusing to start would still be wrong, because the
deployments it would break are single-host ones with no peers to be unreachable
by, and telling those apart at startup is a design with its own specs. Both
places now say what actually happens.

Suite wall clock 592s for 15 specs, up from 502s for 10 of which 2 were red. The
CI budget of 20 minutes does not move.

Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 03:05:12 +00:00

335 lines
14 KiB
Go

package distributed_test
import (
"context"
"fmt"
"io"
"net"
"strings"
"time"
clustersvc "github.com/mudler/LocalAI/core/services/cluster"
"github.com/libp2p/go-yamux/v5"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"gorm.io/driver/postgres"
"gorm.io/gorm"
gormlogger "gorm.io/gorm/logger"
)
const (
// instanceRosterTimeout bounds the wait for a replica's row to appear.
// Registration is synchronous in startup, so this only has to cover the gap
// between /readyz answering and this spec's first query.
instanceRosterTimeout = "30s"
instanceRosterPoll = "500ms"
// deadReplicaTimeout bounds the wait for a survivor to reap a replica that
// was killed: the liveness window plus a sweep interval plus slack. It is
// deliberately derived from the constants rather than a round number, so
// tightening the window shortens the spec instead of leaving it passing for
// the wrong reason.
deadReplicaTimeout = clustersvc.InstanceLiveness + 4*clustersvc.InstanceHeartbeat
// peerDialTimeout bounds one peer dial. Every replica here is a local
// process, so a dial that needs longer has failed, not slowed.
peerDialTimeout = 20 * time.Second
// gracefulDepartureTimeout bounds the wait for a cleanly stopped replica to
// leave the table. It must stay well under InstanceLiveness, which the spec
// asserts: a budget that reached the window would pass on the sweeper doing
// the work and prove nothing about deregistration.
gracefulDepartureTimeout = 15 * time.Second
// peerRefusalTimeout bounds how long a refused stream may take to end. It
// is short on purpose: the refusal is one frame from a replica that has
// already decided, so a stream still open at this point is parked.
peerRefusalTimeout = 5 * time.Second
// unheldNodeID is a worker id no replica holds a tunnel for. It is a
// well-formed id rather than a nonsense string so the refusal it draws is
// the routing answer and not a parse failure.
unheldNodeID = "00000000-0000-0000-0000-00000000dead"
)
// openClusterDB connects to the database the cluster was given, so a spec can
// read the tables the peer link keeps. Nothing serves them over HTTP: they are
// replica-to-replica state, not an admin surface, and inventing an endpoint to
// observe them would be a bigger change than the thing under test.
func openClusterDB(dsn string) *gorm.DB {
GinkgoHelper()
db, err := gorm.Open(postgres.Open(dsn), &gorm.Config{Logger: gormlogger.Discard})
Expect(err).ToNot(HaveOccurred())
DeferCleanup(func() { closeDB(db) })
return db
}
// hostPortOf strips the scheme off a frontend URL, giving the form the
// instances table stores.
func hostPortOf(url string) string {
return strings.TrimPrefix(strings.TrimPrefix(url, "http://"), "https://")
}
// instanceRoster reads the live replica rows, keeping the last error so a
// failing Eventually can name it.
type instanceRoster struct {
registry *clustersvc.Registry
ctx context.Context
lastErr error
lastSaw []clustersvc.Instance
}
func newInstanceRoster(db *gorm.DB) *instanceRoster {
return &instanceRoster{registry: clustersvc.NewRegistry(db), ctx: context.Background()}
}
// addresses returns the advertised address of every live replica, or nil on a
// query error so Eventually keeps trying.
func (r *instanceRoster) addresses() []string {
live, err := r.registry.Live(r.ctx, clustersvc.InstanceLiveness)
if err != nil {
r.lastErr = err
return nil
}
r.lastErr = nil
r.lastSaw = live
addrs := []string{}
for _, instance := range live {
addrs = append(addrs, instance.AdvertisedAddr)
}
return addrs
}
// idAt returns the id of the live replica advertising addr, or "" if no such
// row is present yet.
func (r *instanceRoster) idAt(addr string) string {
for _, instance := range r.lastSaw {
if instance.AdvertisedAddr == addr {
return instance.ID
}
}
return ""
}
func (r *instanceRoster) describe() string {
if r.lastErr != nil {
return fmt.Sprintf("the last read of the instances table failed: %v", r.lastErr)
}
return fmt.Sprintf("the instances table held %d live replica(s): %+v", len(r.lastSaw), r.lastSaw)
}
// awaitReplicas waits for every frontend of c to publish its address and
// returns the roster, positioned on that reading.
func awaitReplicas(roster *instanceRoster, addrs ...string) {
GinkgoHelper()
Eventually(roster.addresses, instanceRosterTimeout, instanceRosterPoll).
Should(ConsistOf(addrs), roster.describe)
}
var _ = Describe("Cluster peer link", Label("Distributed"), Label("Cluster"), func() {
It("publishes an address for every replica that peers can actually dial", func() {
// A wrong implementation registers nothing (the whole of phase 1 had no
// call site until this spec), registers one row for two replicas, or
// records an address nothing can connect to: the bind address of a
// replica behind a service, or the loopback address the route to a
// co-located database would suggest.
c, dsn := startClusterOnFreshDB(2, 0)
roster := newInstanceRoster(openClusterDB(dsn))
awaitReplicas(roster, hostPortOf(c.FrontendURL(0)), hostPortOf(c.FrontendURL(1)))
// "Routable" is not a property of the string. Connect to each address,
// which is the only check that would have caught a replica publishing
// the port it was configured with rather than the one it serves on.
for _, instance := range roster.lastSaw {
conn, err := net.DialTimeout("tcp", instance.AdvertisedAddr, peerDialTimeout)
Expect(err).ToNot(HaveOccurred(),
"replica %s advertises %q, which nothing can connect to", instance.ID, instance.AdvertisedAddr)
Expect(conn.Close()).To(Succeed())
}
})
It("carries a peer stream between two replicas, and refuses one without the cluster token", func() {
// A wrong implementation fails here on WebSocket framing, which is the
// likeliest defect in the peer link: the adapter has to turn
// message-oriented WebSocket frames into the undelimited byte stream
// yamux drives. It also fails if the route was never registered on the
// real server, or if the global session middleware answers it: a peer
// carries no session and no user, only the cluster token.
//
// The stream is opened with the production dialler, resolving the peer
// through the production registry, over a real socket to a real
// process. This spec plays the sibling replica, because phase 1 has
// nothing that makes a frontend dial one on its own.
c, dsn := startClusterOnFreshDB(2, 0)
roster := newInstanceRoster(openClusterDB(dsn))
awaitReplicas(roster, hostPortOf(c.FrontendURL(0)), hostPortOf(c.FrontendURL(1)))
peerID := roster.idAt(hostPortOf(c.FrontendURL(1)))
Expect(peerID).ToNot(BeEmpty())
ctx, cancel := context.WithTimeout(context.Background(), peerDialTimeout)
defer cancel()
pool := clustersvc.NewPeerPool("e2e-peer", c.RegistrationToken(), roster.registry)
DeferCleanup(pool.Close)
// OpenStream is only acknowledged once the far side accepts, so this
// returning at all proves the frontend is accepting streams on the
// session it took, in addition to proving the handshake.
stream, err := pool.Open(ctx, peerID)
Expect(err).ToNot(HaveOccurred())
DeferCleanup(func() { _ = stream.Close() })
// Phase 2 installs the relay on this link, so an accepted stream is one
// the peer is waiting to be told which worker it is for. Name one no
// replica holds and the refusal must come back at once.
//
// This spec used to assert the opposite, that an accepted stream ended
// immediately, because phase 1 had no relay to hand it to. The relay
// made that stale rather than wrong: a stream that says nothing now
// parks for relayHeaderTimeout, which is 15 seconds, and the old
// assertion failed on a five second budget against a replica behaving
// exactly as designed.
Expect(stream.SetWriteDeadline(time.Now().Add(peerRefusalTimeout))).To(Succeed())
Expect(clustersvc.WriteRelayRequest(stream, unheldNodeID, peerRefusalTimeout)).To(Succeed())
Expect(stream.SetReadDeadline(time.Now().Add(peerRefusalTimeout))).To(Succeed())
err = clustersvc.ReadRelayReply(stream)
Expect(err).To(MatchError(clustersvc.ErrNotOwner),
"the peer did not refuse a worker it does not hold: %v", err)
Expect(err).ToNot(MatchError(clustersvc.ErrNoConnection),
"a replica that does not hold a tunnel must not report the worker as absent: that is how a scheduler evicts a healthy worker")
// And the refusal ENDS the stream. A replica that says why and leaves
// the stream open has parked the caller on a request that will never be
// served, which reads as a slow replica rather than a refused request,
// and no deadline on the far side can tell those apart.
Expect(stream.SetReadDeadline(time.Now().Add(peerRefusalTimeout))).To(Succeed())
_, err = stream.Read(make([]byte, 1))
Expect(err).To(SatisfyAny(MatchError(io.EOF), MatchError(yamux.ErrStreamReset)),
"the peer refused the stream and then left it open: %v", err)
// The same dial with the wrong credentials must be refused, otherwise
// the success above says nothing about authentication.
impostor := clustersvc.NewPeerPool("e2e-peer", "not-the-cluster-token", roster.registry)
DeferCleanup(impostor.Close)
_, err = impostor.Open(ctx, peerID)
Expect(err).To(MatchError(clustersvc.ErrPeerUnreachable))
Expect(err).ToNot(MatchError(clustersvc.ErrInstanceNotFound),
"a peer refusing credentials is a live peer; reading it as absence is how a replica evicts healthy workers")
})
It("stops being dialled as soon as a replica shuts down cleanly", func() {
// The crash case below is handled by the sweeper, at the cost of a
// whole liveness window of peers dialling a corpse. A rolling update is
// not a crash: the replica knows it is leaving and says so. Without
// deregistration the two are indistinguishable, and every rolling
// restart spends that window failing peer dials for no reason.
c, dsn := startClusterOnFreshDB(2, 0)
roster := newInstanceRoster(openClusterDB(dsn))
awaitReplicas(roster, hostPortOf(c.FrontendURL(0)), hostPortOf(c.FrontendURL(1)))
departingID := roster.idAt(hostPortOf(c.FrontendURL(1)))
Expect(departingID).ToNot(BeEmpty())
Expect(c.StopFrontendGracefully(1)).To(Succeed())
Eventually(func() bool { return c.FrontendAlive(1) }, "20s", "500ms").Should(BeFalse())
// The budget is deliberately shorter than the liveness window: passing
// it proves the replica announced its departure rather than aged out.
Expect(gracefulDepartureTimeout).To(BeNumerically("<", clustersvc.InstanceLiveness))
Eventually(roster.addresses, gracefulDepartureTimeout, instanceRosterPoll).
Should(ConsistOf(hostPortOf(c.FrontendURL(0))), roster.describe)
// And absence is the RIGHT answer here, unlike the killed case: the
// replica said it was going. A caller may act on this.
ctx, cancel := context.WithTimeout(context.Background(), peerDialTimeout)
defer cancel()
pool := clustersvc.NewPeerPool("e2e-peer", c.RegistrationToken(), roster.registry)
DeferCleanup(pool.Close)
_, err := pool.Open(ctx, departingID)
Expect(err).To(MatchError(clustersvc.ErrInstanceNotFound))
})
It("reports a killed replica as unreachable, reaps what it owned, and evicts no worker", func() {
// This is the absence rule, pinned before phase 2 can depend on it. A
// wrong implementation lets a peer that will not answer surface as node
// absence, and a caller entitled to act on absence then reclaims what
// the peer was running: a network hiccup between two healthy replicas
// evicts healthy workers.
//
// It also pins the reaper: the connection rows a dead replica owned are
// swept by the same sweeper that decides the replica is dead, so the
// two can never disagree about who is alive.
c, dsn := startClusterOnFreshDB(2, 1)
client, err := c.AdminSession(0)
Expect(err).ToNot(HaveOccurred())
// The worker registers with frontend 0, so frontend 1 is the replica
// that can die without taking the worker's registrar with it.
registrar, err := c.WorkerRegistrar(0)
Expect(err).ToNot(HaveOccurred())
Expect(registrar).To(Equal(0), "this spec kills frontend 1 and needs the worker to have registered elsewhere")
probe := newRosterProbe(c, client, 0)
Eventually(probe.healthyNames, nodeRosterTimeout, nodeRosterPoll).
Should(ContainElement(c.WorkerName(0)), probe.describe)
workerID := probe.idOf(c.WorkerName(0))
Expect(workerID).ToNot(BeEmpty())
roster := newInstanceRoster(openClusterDB(dsn))
awaitReplicas(roster, hostPortOf(c.FrontendURL(0)), hostPortOf(c.FrontendURL(1)))
survivorID := roster.idAt(hostPortOf(c.FrontendURL(0)))
doomedID := roster.idAt(hostPortOf(c.FrontendURL(1)))
Expect(survivorID).ToNot(BeEmpty())
Expect(doomedID).ToNot(BeEmpty())
// Give frontend 1 the worker's tunnel. Phase 2 makes the worker do this
// by dialling; here the claim is written directly, because the point
// under test is what happens to the claim when its owner dies.
ctx := context.Background()
epoch, err := roster.registry.Claim(ctx, workerID, doomedID)
Expect(err).ToNot(HaveOccurred())
Expect(epoch).ToNot(BeZero())
Expect(c.KillFrontend(1)).To(Succeed())
Eventually(func() bool { return c.FrontendAlive(1) }, "20s", "500ms").Should(BeFalse())
// The row is still there for the whole liveness window, so this is the
// case that matters: the peer is KNOWN and will not answer.
dialCtx, cancel := context.WithTimeout(ctx, peerDialTimeout)
defer cancel()
pool := clustersvc.NewPeerPool("e2e-peer", c.RegistrationToken(), roster.registry)
DeferCleanup(pool.Close)
_, err = pool.Open(dialCtx, doomedID)
Expect(err).To(MatchError(clustersvc.ErrPeerUnreachable))
Expect(err).ToNot(MatchError(clustersvc.ErrInstanceNotFound),
"a dead replica whose row is still present is unreachable, not absent")
// The survivor sweeps the dead replica and, in the same pass, the claim
// it left behind.
Eventually(roster.addresses, deadReplicaTimeout, instanceRosterPoll).
Should(ConsistOf(hostPortOf(c.FrontendURL(0))), roster.describe)
ownerErr := func() error {
_, _, err := roster.registry.OwnerRow(ctx, workerID)
return err
}
Eventually(ownerErr, deadReplicaTimeout, instanceRosterPoll).
Should(MatchError(clustersvc.ErrNoConnection),
"the claim held by a replica that no longer exists was never reaped")
// And the worker survives the sweep that removed its owner. This is a
// window after the reaping, not a watch over the whole scenario:
// Consistently starts here, so what it rules out is the sweep, or
// anything reacting to it, taking the worker with it.
Consistently(probe.healthyNames, "6s", "1s").
Should(ContainElement(c.WorkerName(0)), probe.describe)
})
})