mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-02 02:54:34 -04:00
chore(distributed): stop telling an operator to run a NATS cluster
Every carrier had already moved and no process opened a bus connection, but the surface an operator reads still described a deployment with a broker in it: a compose service, a 220-line credential-generation script, two CI steps pulling a container nothing started, two flag tables offering --nats-url, an architecture diagram with a NATS box wired to the workers, a join-command generator in the Nodes page that emitted --nats-url for agent workers, and a test suite that stood a NATS server up for specs that no longer used it. That is the one way this programme could still fail invisibly. Every test passes, every binary works, and every production deployment goes on running and paying for infrastructure that carries nothing. Nothing in this repository starts a NATS server any more. The compose file is four services, the docs say to shut the broker down and what to keep, and the e2e suite runs on one PostgreSQL container. The three LOCALAI_NATS_*_TIMEOUT env vars are KEPT, and are now documented twice as being kept. They were never broker settings: each names a control-RPC budget the frontend applies to a worker, still read and still enforced. They carry the prefix only because they arrived with the bus, and renaming them would break every existing deployment for cosmetics. The agent worker's join command was the last surface still emitting the flag, two tasks after the agent worker stopped dialling. The Playwright spec that covered it asserted the opposite of what is now true, so it is inverted rather than deleted, and it reads the rendered command string rather than the component's variables: the variables are what the fix removes, so a spec reading them would have stopped compiling instead of failing, and a compile error is not evidence about what an operator is shown. nats_jwt_test.go and its helpers are deleted. They pinned a real server ENFORCING the minted permissions. The CONTENT of those allow lists is still pinned, untouched, by pkg/natsauth's own suites, including the spec that refuses to let the agent lists go empty, since an empty allow list in NATS means unrestricted. The enforcement half is retired rather than moved: enforcement is a property of a connection, and nothing opens one. The suite's own NATS container goes with them, which the brief left for the next task. Removing the pre-pull while BeforeSuite still ran the image would have defeated the step rather than cleaned it up, and this change removes the last reader of TestInfra.NC. agent_native_executor_test.go and mcp_ci_job_test.go are moved onto infra.Bus() instead of deleted: they were the last two specs building a bridge and a dispatcher on a client nobody uses, which is exactly the drift TestInfra.Bus's own comment warns about. cluster.Options.NatsURL is now fed a deliberately dead address rather than a live container's. Frontends and agent workers still receive LOCALAI_NATS_URL, because that is the coverage for the promise that an existing command line still starts; sourcing it from a running server would have let a regression that actually dialled it pass. The control in cluster_control_test.go keeps its assertion and loses its explanation, which claimed the deployment had a bus and no longer could. One latent spec race surfaced and is fixed: the background-run spec waited for a COUNT of events and then read a snapshot for the terminal status, which is the last event of a run and therefore always arrives after the count is met. Its immediate twin had already been fixed this way. Nothing in production changed. pkg/natsauth keeps its files. It is reachable from production only through the natsauth.Config parameter thread, and that thread is the next task's. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
1 parent
a6b2d7c0ec
commit
d3dfad90b9
23 files changed
+291
-879
No files matched your search
@@ -56,11 +56,12 @@ jobs:
|
||||
- name: Pre-pull test images
|
||||
# Pulling here rather than inside the suite keeps container-start timing
|
||||
# out of the spec timeouts and makes a registry outage read as a
|
||||
# setup failure instead of a test failure. These two are the only images
|
||||
# the suite needs once the testcontainers reaper is disabled below.
|
||||
# setup failure instead of a test failure. This is the only image the
|
||||
# suite needs once the testcontainers reaper is disabled below: the
|
||||
# suite stands up no message broker, because nothing under test dials
|
||||
# one.
|
||||
run: |
|
||||
docker pull postgres:16-alpine
|
||||
docker pull nats:2-alpine
|
||||
- name: Distributed E2E
|
||||
# TESTCONTAINERS_RYUK_DISABLED keeps the pre-pull above meaningful. The
|
||||
# reaper exists to clean up leaked containers on a long-lived host, but
|
||||
@@ -91,9 +92,9 @@ jobs:
|
||||
#
|
||||
# Separate job from tests-e2e-distributed so the fast in-process suite is
|
||||
# not held behind a Go build of local-ai. Serial on purpose: each Ginkgo
|
||||
# process would get its own PostgreSQL and NATS container and each spec
|
||||
# spawns two or three local-ai children, so --procs on an unmeasured runner
|
||||
# is a change to make with numbers, not by default.
|
||||
# process would get its own PostgreSQL container and each spec spawns two or
|
||||
# three local-ai children, so --procs on an unmeasured runner is a change to
|
||||
# make with numbers, not by default.
|
||||
#
|
||||
# The two timeouts bound different things and are not alternatives. Ginkgo's
|
||||
# --timeout=20m bounds the SUITE only; this job timeout must additionally
|
||||
@@ -160,7 +161,6 @@ jobs:
|
||||
# setup failure rather than a test failure.
|
||||
run: |
|
||||
docker pull postgres:16-alpine
|
||||
docker pull nats:2-alpine
|
||||
- name: Cluster E2E
|
||||
env:
|
||||
# No LOCALAI_E2E_BINARY and no separate build step: make test-e2e-cluster
|
||||
|
||||
Reference in new issue
Block a user