mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-13 06:45:26 -04:00
Every carrier had already moved and no process opened a bus connection, but the surface an operator reads still described a deployment with a broker in it: a compose service, a 220-line credential-generation script, two CI steps pulling a container nothing started, two flag tables offering --nats-url, an architecture diagram with a NATS box wired to the workers, a join-command generator in the Nodes page that emitted --nats-url for agent workers, and a test suite that stood a NATS server up for specs that no longer used it. That is the one way this programme could still fail invisibly. Every test passes, every binary works, and every production deployment goes on running and paying for infrastructure that carries nothing. Nothing in this repository starts a NATS server any more. The compose file is four services, the docs say to shut the broker down and what to keep, and the e2e suite runs on one PostgreSQL container. The three LOCALAI_NATS_*_TIMEOUT env vars are KEPT, and are now documented twice as being kept. They were never broker settings: each names a control-RPC budget the frontend applies to a worker, still read and still enforced. They carry the prefix only because they arrived with the bus, and renaming them would break every existing deployment for cosmetics. The agent worker's join command was the last surface still emitting the flag, two tasks after the agent worker stopped dialling. The Playwright spec that covered it asserted the opposite of what is now true, so it is inverted rather than deleted, and it reads the rendered command string rather than the component's variables: the variables are what the fix removes, so a spec reading them would have stopped compiling instead of failing, and a compile error is not evidence about what an operator is shown. nats_jwt_test.go and its helpers are deleted. They pinned a real server ENFORCING the minted permissions. The CONTENT of those allow lists is still pinned, untouched, by pkg/natsauth's own suites, including the spec that refuses to let the agent lists go empty, since an empty allow list in NATS means unrestricted. The enforcement half is retired rather than moved: enforcement is a property of a connection, and nothing opens one. The suite's own NATS container goes with them, which the brief left for the next task. Removing the pre-pull while BeforeSuite still ran the image would have defeated the step rather than cleaned it up, and this change removes the last reader of TestInfra.NC. agent_native_executor_test.go and mcp_ci_job_test.go are moved onto infra.Bus() instead of deleted: they were the last two specs building a bridge and a dispatcher on a client nobody uses, which is exactly the drift TestInfra.Bus's own comment warns about. cluster.Options.NatsURL is now fed a deliberately dead address rather than a live container's. Frontends and agent workers still receive LOCALAI_NATS_URL, because that is the coverage for the promise that an existing command line still starts; sourcing it from a running server would have let a regression that actually dialled it pass. The control in cluster_control_test.go keeps its assertion and loses its explanation, which claimed the deployment had a bus and no longer could. One latent spec race surfaced and is fixed: the background-run spec waited for a COUNT of events and then read a snapshot for the terminal status, which is the last event of a run and therefore always arrives after the count is met. Its immediate twin had already been fixed this way. Nothing in production changed. pkg/natsauth keeps its files. It is reachable from production only through the natsauth.Config parameter thread, and that thread is the next task's. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
228 lines
8.7 KiB
YAML
228 lines
8.7 KiB
YAML
# Docker Compose for LocalAI Distributed Mode
|
|
#
|
|
# Starts a full distributed stack: PostgreSQL, a LocalAI frontend, one
|
|
# llama-cpp backend node and one agent worker.
|
|
#
|
|
# There is no message broker in this file and none is needed. PostgreSQL carries
|
|
# every cross-replica broadcast, and each worker dials one outbound tunnel to the
|
|
# frontend and takes every verb on it.
|
|
#
|
|
# Model files are transferred from the frontend to backend nodes via HTTP
|
|
# — no shared volumes needed between frontend and backends.
|
|
#
|
|
# Usage:
|
|
# docker compose -f docker-compose.distributed.yaml up
|
|
#
|
|
# See docs: https://localai.io/features/distributed-mode/
|
|
|
|
services:
|
|
# --- Infrastructure ---
|
|
|
|
postgres:
|
|
image: quay.io/mudler/localrecall:v0.5.5-postgresql # PostgreSQL with pgvector
|
|
environment:
|
|
POSTGRES_DB: localai
|
|
POSTGRES_USER: localai
|
|
POSTGRES_PASSWORD: localai
|
|
volumes:
|
|
- postgres_data:/var/lib/postgresql
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "pg_isready -U localai"]
|
|
interval: 5s
|
|
timeout: 3s
|
|
retries: 10
|
|
|
|
# --- LocalAI Frontend ---
|
|
# Stateless API server that routes requests to backend nodes.
|
|
# Add more replicas behind a load balancer for HA.
|
|
|
|
localai:
|
|
# image: localai/localai:latest-cpu
|
|
build:
|
|
context: .
|
|
dockerfile: Dockerfile
|
|
args:
|
|
- IMAGE_TYPE=core
|
|
- BASE_IMAGE=ubuntu:24.04
|
|
ports:
|
|
- "8080:8080"
|
|
environment:
|
|
# Distributed mode
|
|
LOCALAI_DISTRIBUTED: "true"
|
|
LOCALAI_AGENT_POOL_EMBEDDING_MODEL: "granite-embedding-107m-multilingual"
|
|
LOCALAI_AGENT_POOL_VECTOR_ENGINE: "postgres"
|
|
LOCALAI_AGENT_POOL_DATABASE_URL: "postgresql://localai:localai@postgres:5432/localai?sslmode=disable"
|
|
LOCALAI_REGISTRATION_TOKEN: "changeme" # Change this in production!
|
|
# Shared-models mode (optional): set when every node mounts the SAME
|
|
# models directory at the SAME path (see "Shared Volume Mode" below).
|
|
# The router then skips gRPC file staging and workers load models
|
|
# directly from the shared volume instead of re-downloading them.
|
|
# LOCALAI_DISTRIBUTED_SHARED_MODELS: "true"
|
|
# Auth (required for distributed mode — must use PostgreSQL)
|
|
LOCALAI_AUTH: "true"
|
|
LOCALAI_AUTH_DATABASE_URL: "postgresql://localai:localai@postgres:5432/localai?sslmode=disable"
|
|
# Force pure-Go DNS resolver. The default cgo resolver follows the
|
|
# container's nsswitch.conf and ends up forwarding to host
|
|
# systemd-resolved (127.0.0.53), which isn't reachable from inside
|
|
# the container, failing every postgres hostname lookup at
|
|
# boot. The pure-Go path reads /etc/resolv.conf directly and uses
|
|
# Docker's embedded DNS at 127.0.0.11.
|
|
GODEBUG: "netdns=go"
|
|
# Paths
|
|
MODELS_PATH: /models
|
|
# Avoid probing remote gallery GGUF metadata during container startup.
|
|
# Remove this line or set a positive limit to opt back into cache warming.
|
|
LOCALAI_VRAM_WARM_LIMIT: "0"
|
|
volumes:
|
|
- frontend_models:/models
|
|
- frontend_data:/data
|
|
depends_on:
|
|
postgres:
|
|
condition: service_healthy
|
|
|
|
# --- Worker Node ---
|
|
# A generic worker that self-registers with the frontend.
|
|
# The same LocalAI image is used — no separate image needed.
|
|
# The SmartRouter tells a worker which backend to install over that worker's
|
|
# own tunnel.
|
|
#
|
|
# Model files are transferred from the frontend via HTTP file staging.
|
|
# The worker has its own independent models volume.
|
|
|
|
worker-1:
|
|
# image: localai/localai:latest-cpu
|
|
build:
|
|
context: .
|
|
dockerfile: Dockerfile
|
|
args:
|
|
- IMAGE_TYPE=core
|
|
- BASE_IMAGE=ubuntu:24.04
|
|
command:
|
|
- worker
|
|
# No published ports and no advertised address: the worker holds one
|
|
# outbound tunnel to the frontend and binds only loopback, so nothing has to
|
|
# reach into this container.
|
|
#
|
|
# No HEALTHCHECK_ENDPOINT override is needed either: the image's healthcheck
|
|
# detects worker mode and derives the port from LOCALAI_SERVE_ADDR below
|
|
# (gRPC base port - 1 = 50050). It runs inside the container, so a loopback
|
|
# bind is enough for it. The worker's /readyz reports 503 while it holds no
|
|
# tunnel session, so `unhealthy` here means the frontend genuinely cannot
|
|
# reach this worker.
|
|
#
|
|
# This worker connects to nothing but the frontend. Everything the frontend
|
|
# asks of it travels the tunnel this container dials out to localai:8080, so
|
|
# the only service it depends on is localai itself.
|
|
environment:
|
|
LOCALAI_SERVE_ADDR: "0.0.0.0:50051"
|
|
DEBUG: "true"
|
|
LOCALAI_REGISTER_TO: "http://localai:8080"
|
|
LOCALAI_NODE_NAME: "worker-1"
|
|
LOCALAI_REGISTRATION_TOKEN: "changeme" # Must match frontend token
|
|
LOCALAI_HEARTBEAT_INTERVAL: "10s"
|
|
GODEBUG: "netdns=go" # See note in localai service
|
|
MODELS_PATH: /models
|
|
volumes:
|
|
- worker_1_models:/models
|
|
depends_on:
|
|
localai:
|
|
condition: service_started
|
|
|
|
# --- GPU Support (NVIDIA) ---
|
|
# Uncomment the following and change the image to a CUDA variant
|
|
# (e.g., localai/localai:latest-gpu-nvidia-cuda-12) to enable GPU.
|
|
#
|
|
# NVIDIA_DRIVER_CAPABILITIES must include `utility` so nvidia-smi / NVML
|
|
# are available inside the container; without it the worker cannot report
|
|
# free VRAM and the Nodes page will show 0 free / total used.
|
|
# `init: true` avoids zombie-reap races that make nvidia-smi flaky.
|
|
#
|
|
# init: true
|
|
# environment:
|
|
# NVIDIA_DRIVER_CAPABILITIES: "compute,utility"
|
|
# deploy:
|
|
# resources:
|
|
# reservations:
|
|
# devices:
|
|
# - driver: nvidia.com/gpu
|
|
# count: all
|
|
# capabilities: [gpu, utility]
|
|
|
|
# --- Shared Volume Mode (optional) ---
|
|
# If all services run on the same Docker host, you can skip gRPC file transfer
|
|
# by sharing a single models volume. Replace the volumes above with:
|
|
#
|
|
# localai:
|
|
# volumes:
|
|
# - shared_models:/models
|
|
# - frontend_data:/data
|
|
#
|
|
# backend-llama-cpp:
|
|
# volumes:
|
|
# - shared_models:/models
|
|
#
|
|
# Then add to the volumes section:
|
|
# shared_models:
|
|
#
|
|
# With shared volumes the model files are already present on every worker at
|
|
# the same path. Set LOCALAI_DISTRIBUTED_SHARED_MODELS=true on the frontend
|
|
# (see its environment above) so the router skips gRPC file staging and the
|
|
# worker loads the model directly from the shared path instead of
|
|
# re-downloading it into a per-model subdirectory.
|
|
|
|
# --- Adding More Workers ---
|
|
# Copy the worker-1 service above and change:
|
|
# - Service name (e.g., worker-2)
|
|
# - LOCALAI_NODE_NAME (must be unique)
|
|
#
|
|
# Nothing else. A worker has no address to make unique: it binds loopback
|
|
# inside its own container and dials out to the frontend. Note that
|
|
# LOCALAI_NODE_NAME really must differ: the registry upserts by name, so two
|
|
# workers sharing one steal each other's row and each other's tunnel credential.
|
|
#
|
|
# Workers are generic: no backend type needed. The SmartRouter installs the
|
|
# required backend over the worker's tunnel when a model request arrives.
|
|
|
|
# --- Agent Worker ---
|
|
# Dedicated process for agent chat execution.
|
|
# The frontend claims a queued run from PostgreSQL and drives it as a
|
|
# streaming control RPC over this container's own outbound tunnel; progress,
|
|
# agent events and the terminal result all come back on that same response.
|
|
# No database access needed: config and skills are sent in the request.
|
|
|
|
agent-worker-1:
|
|
# image: localai/localai:latest-cpu
|
|
build:
|
|
context: .
|
|
dockerfile: Dockerfile
|
|
args:
|
|
- IMAGE_TYPE=core
|
|
- BASE_IMAGE=ubuntu:24.04
|
|
# Install Docker CLI and start agent-worker.
|
|
# The Docker socket is mounted from the host so that MCP stdio servers
|
|
# using "docker run" commands can spawn containers on the host Docker.
|
|
entrypoint: ["/bin/sh", "-c"]
|
|
command:
|
|
- |
|
|
apt-get update -qq && apt-get install -y -qq docker.io >/dev/null 2>&1
|
|
exec /entrypoint.sh agent-worker
|
|
# The agent worker binds its control server on loopback only and publishes
|
|
# no port. The image's healthcheck detects that mode and reports healthy
|
|
# rather than probing a port that will never bind, so no override is needed.
|
|
environment:
|
|
LOCALAI_REGISTER_TO: "http://localai:8080"
|
|
LOCALAI_NODE_NAME: "agent-worker-1"
|
|
LOCALAI_REGISTRATION_TOKEN: "changeme" # Must match frontend token
|
|
GODEBUG: "netdns=go" # See note in localai service
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock
|
|
depends_on:
|
|
localai:
|
|
condition: service_started
|
|
|
|
volumes:
|
|
postgres_data:
|
|
frontend_models:
|
|
frontend_data:
|
|
worker_1_models:
|