mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-30 10:04:32 -04:00
A worker now opens no listener on a routable interface and states no endpoint at registration. Backend processes and the file-transfer server bind loopback, and the frontend reaches both through the tunnel the worker dials. The bind address is built from loopbackHost, the same constant the tunnel's grpc tag dials, so "the worker binds where its tunnel dials" is one fact in one place rather than two literals that can drift. All three advertisement sites are closed, not one: the registration body, RegisterNodeRequest, and the per-backend address in the install reply. That third one was hiding a live bug. stopModelExact refuses a stop whose ExpectedAddress does not match what the worker recorded for the process. The worker recorded 127.0.0.1:port; handleBackendInstall reported advertiseHost:port; the router stored the reported one and sent it straight back. On any worker whose advertise host was not 127.0.0.1, every acknowledged model stop failed with an address mismatch. Nothing caught it because the e2e harness set LOCALAI_ADVERTISE_ADDR=127.0.0.1, which made the rewrite a no-op. Removing the rewrite makes the two strings the same by construction. The brief was wrong about two of the four functions it called dead. effectiveBasePort is the base of the backend port allocator and resolveHTTPAddr is the file server's bind address; deleting them would have deleted the port allocator and the file server. Only the two advertise* helpers were dead, and addr_test.go is rewritten rather than deleted, because the port arithmetic it pinned still needs pinning. NodeModel.Address survives with a narrowed meaning and is renamed WorkerLocalAddress, along with the install reply field that feeds it. The frontend still has to say WHICH backend process on a worker it means, and the port in this string is how it says it: it travels as a stream target and the worker dials its own loopback. The gorm column and the json key stay "address", so neither a migration nor an API break rides along. Every fall-back to the node's address is gone. installBackendOnNode now errors when a worker reports success without naming one, because substituting the now-always-empty node address would name an empty target, and the worker refuses that as an invalid stream, which is classified as the worker answering about its backend. That is the "a present worker reads as something it is not" class this phase forbids. DistributedModelStore.Range had the same shape and was already wrong: it built each remote model's client from the node's base gRPC port, never the port a backend process listens on, so Free and Status went to the wrong place. It uses the replica's address now. BackendNode.Address and HTTPAddress are kept but made provably inert: no writer, no reader that acts on them, and Register force-clears both on re-registration so an upgraded worker's stale advertisement does not outlive its own upgrade in the API and the Nodes page. Dropping the columns is a ~90-site edit across the specs, the e2e suite, the MCP dto and the UI; it is recorded as a follow-up rather than folded in here. A persistent tunnel 401 still does not trigger re-registration, and now for a reason rather than a deferral. Register CLEARS the node's replica rows, so re-registering on a 401 would delete a live worker's rows on every retry, and under the name collision that causes the 401 the two workers would take turns doing it forever: a credential failure causing model reclamation. It also cannot fix the named cause, since a collision is indistinguishable from a restart. The 401 log now names both causes and says nothing can reach this worker, which is true only now that it has no listener. The container healthcheck did not break the way the brief expected, since the listener still exists on loopback and the probe runs inside the container. It did have a real #10987 defect that this change makes the common case: it read LOCALAI_SERVE_ADDR only, while effectiveBasePort reads LOCALAI_ADDR first, so a worker on a non-default base port was probed on 50050 and reported unhealthy while working. It follows the same precedence now. Docs, the compose file and the e2e harness are updated in step: no inbound rule or published port is needed for a worker, the two advertise variables are gone, the remaining address variables are read for their port only, the firewall-the-file-transfer-port warning is narrowed to the LOCALAI_HTTP_ADDR opt-out, and the upgrade-order note no longer claims the worker still listens. The Nodes page showed node.address, which is now always blank, so it shows the node id instead. Eight mutations, all red on a named spec, including reverting the loopback bind, re-adding the address to the registration body, restoring both node-address fall-backs, dropping the force-clear, storing the endpoint's address again, and un-fixing the healthcheck. One of them caught a defect in a spec I had just written: it asserted 200 where the endpoint returns 201, which went unnoticed because core/http/endpoints/localai is not on the task's verify list. It is run here. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
262 lines
9.5 KiB
Go
262 lines
9.5 KiB
Go
package worker
|
|
|
|
import (
|
|
"fmt"
|
|
"math"
|
|
"net"
|
|
"os"
|
|
"strconv"
|
|
"strings"
|
|
|
|
"github.com/mudler/LocalAI/pkg/system"
|
|
"github.com/mudler/LocalAI/pkg/xsysinfo"
|
|
"github.com/mudler/xlog"
|
|
)
|
|
|
|
var (
|
|
totalAvailableVRAM = xsysinfo.TotalAvailableVRAM
|
|
getGPUAggregateInfo = xsysinfo.GetGPUAggregateInfo
|
|
getSystemRAMInfo = xsysinfo.GetSystemRAMInfo
|
|
getCPUInfo = xsysinfo.GetCPUInfo
|
|
)
|
|
|
|
func clampCPUUsage(usage float64) float64 {
|
|
if math.IsNaN(usage) || usage < 0 {
|
|
return 0
|
|
}
|
|
if usage > 100 {
|
|
return 100
|
|
}
|
|
return usage
|
|
}
|
|
|
|
// effectiveBasePort returns the port used as base for gRPC backend processes.
|
|
// Priority: Addr port → ServeAddr port → 50051
|
|
//
|
|
// Only the PORT of those settings is read. Their host halves name an interface
|
|
// this worker no longer binds: every backend listens on loopback and is reached
|
|
// through the tunnel.
|
|
func (cfg *Config) effectiveBasePort() int {
|
|
for _, addr := range []string{cfg.Addr, cfg.ServeAddr} {
|
|
if addr == "" {
|
|
continue
|
|
}
|
|
_, portStr, err := net.SplitHostPort(addr)
|
|
if err != nil {
|
|
xlog.Warn("Invalid worker address; trying the next base-port source", "addr", addr, "error", err)
|
|
continue
|
|
}
|
|
port, err := strconv.Atoi(portStr)
|
|
if err != nil {
|
|
xlog.Warn("Invalid worker port; trying the next base-port source", "addr", addr, "port", portStr, "error", err)
|
|
continue
|
|
}
|
|
if port > 0 && port <= 65535 {
|
|
return port
|
|
}
|
|
xlog.Warn("Worker port is outside the valid range; trying the next base-port source", "addr", addr, "port", port)
|
|
}
|
|
return 50051
|
|
}
|
|
|
|
// effectiveMaxPort returns the last port the gRPC backend allocator may hand
|
|
// out. The range is [basePort, maxPort]; its width is the number of backend
|
|
// processes this worker can run concurrently, minus whatever the port
|
|
// quarantine is holding at the time.
|
|
//
|
|
// An unset, non-positive, or out-of-order value falls back to 65535 — the
|
|
// historical (unbounded) behaviour clamped to something bindable.
|
|
// Misconfiguring this must not shrink the range to nothing and wedge every
|
|
// backend start on the worker.
|
|
func (cfg *Config) effectiveMaxPort(basePort int) int {
|
|
if cfg.GRPCMaxPort <= 0 {
|
|
return defaultMaxPort
|
|
}
|
|
if cfg.GRPCMaxPort > defaultMaxPort {
|
|
xlog.Warn("Configured gRPC max port is above the highest TCP port; clamping",
|
|
"configured", cfg.GRPCMaxPort, "max", defaultMaxPort)
|
|
return defaultMaxPort
|
|
}
|
|
if cfg.GRPCMaxPort < basePort {
|
|
xlog.Warn("Configured gRPC max port is below the base port; ignoring it and using the full range",
|
|
"configured", cfg.GRPCMaxPort, "basePort", basePort, "max", defaultMaxPort)
|
|
return defaultMaxPort
|
|
}
|
|
return cfg.GRPCMaxPort
|
|
}
|
|
|
|
// resolveHTTPAddr returns the address to bind the HTTP file transfer server to.
|
|
// Uses basePort-1 so it doesn't conflict with dynamically allocated gRPC ports
|
|
// which grow upward from basePort.
|
|
//
|
|
// The default is loopback for the same reason backend processes are: the
|
|
// frontend reaches this server over the tunnel, whose http tag dials whatever
|
|
// address this returns. An operator who sets HTTPAddr explicitly still gets
|
|
// exactly that bind (see loopbackAddr, which rewrites only a wildcard), so a
|
|
// deployment that has some other local reason to expose the server can, and
|
|
// nothing in the frontend depends on it.
|
|
func (cfg *Config) resolveHTTPAddr() string {
|
|
if cfg.HTTPAddr != "" {
|
|
return cfg.HTTPAddr
|
|
}
|
|
return net.JoinHostPort(loopbackHost, strconv.Itoa(cfg.effectiveBasePort()-1))
|
|
}
|
|
|
|
// registrationBody builds the JSON body for node registration.
|
|
func (cfg *Config) registrationBody() map[string]any {
|
|
nodeName := cfg.NodeName
|
|
if nodeName == "" {
|
|
hostname, err := os.Hostname()
|
|
if err != nil {
|
|
nodeName = fmt.Sprintf("node-%d", os.Getpid())
|
|
} else {
|
|
nodeName = hostname
|
|
}
|
|
}
|
|
|
|
// Detect GPU info for VRAM-aware scheduling
|
|
totalVRAM, err := totalAvailableVRAM()
|
|
if err != nil {
|
|
xlog.Debug("Failed to detect worker VRAM; registering without GPU capacity", "error", err)
|
|
}
|
|
gpuVendor, err := xsysinfo.DetectGPUVendor()
|
|
if err != nil {
|
|
xlog.Debug("Failed to detect worker GPU vendor; registering without vendor metadata", "error", err)
|
|
}
|
|
// Compute capability (e.g. "12.1" for GB10) lets the router pick per-arch
|
|
// options (e.g. larger physical batch on Blackwell). Detected on the worker
|
|
// because only the worker sees the GPU in distributed mode.
|
|
gpuComputeCap := xsysinfo.NVIDIAComputeCapability()
|
|
// Report our own meta-backend capability so the controller can list the
|
|
// backends this cluster can actually run. The controller cannot infer it
|
|
// from the GPU vendor alone: OS-dependent capabilities (metal, darwin-x86,
|
|
// nvidia-l4t) and the CUDA runtime refinements are only observable here.
|
|
capability := ""
|
|
if systemState, err := system.GetSystemState(); err != nil {
|
|
xlog.Warn("Could not detect system capability for node registration", "error", err)
|
|
} else {
|
|
capability = systemState.DetectedCapability()
|
|
}
|
|
|
|
maxReplicas := cfg.MaxReplicasPerModel
|
|
if maxReplicas < 1 {
|
|
maxReplicas = 1
|
|
}
|
|
// No address and no http_address: this worker has nothing inbound to
|
|
// advertise. It holds one outbound tunnel and the frontend reaches every
|
|
// service on it through that, so an address here would be a value that
|
|
// looks dialable, is stored, is shown, and is never dialled.
|
|
body := map[string]any{
|
|
"name": nodeName,
|
|
"total_vram": totalVRAM,
|
|
"available_vram": totalVRAM, // initially all VRAM is available
|
|
"gpu_vendor": gpuVendor,
|
|
"gpu_compute_capability": gpuComputeCap,
|
|
"capability": capability,
|
|
"max_replicas_per_model": maxReplicas,
|
|
}
|
|
|
|
// Report free space on the filesystem that backs the MODELS directory.
|
|
// That is where staged weights land, so it is the only mount whose free
|
|
// space decides whether this node can accept a model — the scheduler uses
|
|
// it to avoid picking a node that would fail with ENOSPC mid-transfer.
|
|
if diskInfo, err := xsysinfo.GetDiskInfo(cfg.ModelsPath); err != nil {
|
|
// Omitted, not zeroed: total_disk == 0 is how the frontend recognises
|
|
// "this worker does not report disk" and keeps it in rotation.
|
|
xlog.Warn("Failed to detect worker models-path disk capacity; registering without it",
|
|
"path", cfg.ModelsPath, "error", err)
|
|
} else {
|
|
body["total_disk"] = diskInfo.Total
|
|
body["available_disk"] = diskInfo.Available
|
|
}
|
|
|
|
// Report the operator-set budget as a STRING so the server resolves and
|
|
// enforces it against the raw VRAM above. The worker never caps its own
|
|
// total_vram/available_vram, and never touches the xsysinfo process-global
|
|
// budget (that is standalone-only). Omit when unset.
|
|
if cfg.VRAMBudget != "" {
|
|
body["vram_budget"] = cfg.VRAMBudget
|
|
}
|
|
|
|
// Report system RAM independently from VRAM so both discrete-GPU and
|
|
// unified-memory workers expose the capacity visible to the host.
|
|
ramInfo, err := getSystemRAMInfo()
|
|
if err != nil {
|
|
xlog.Debug("Failed to detect worker RAM for registration", "error", err)
|
|
} else {
|
|
body["total_ram"] = ramInfo.Total
|
|
body["available_ram"] = ramInfo.Available
|
|
}
|
|
if cfg.RegistrationToken != "" {
|
|
body["token"] = cfg.RegistrationToken
|
|
}
|
|
|
|
if cpuInfo, err := getCPUInfo(); err != nil {
|
|
xlog.Debug("Failed to sample worker CPU for registration", "error", err)
|
|
} else {
|
|
body["cpu_logical_cores"] = cpuInfo.LogicalCores
|
|
body["cpu_usage_percent"] = clampCPUUsage(cpuInfo.UsagePercent)
|
|
body["cpu_load_1"] = cpuInfo.Load1
|
|
}
|
|
|
|
// Parse and add static node labels. Always include the auto-label
|
|
// `node.replica-slots=N` so AND-selectors in ModelSchedulingConfig can
|
|
// target high-capacity nodes (e.g. {"node.replica-slots":"4"}).
|
|
labels := make(map[string]string)
|
|
if cfg.NodeLabels != "" {
|
|
for _, pair := range strings.Split(cfg.NodeLabels, ",") {
|
|
pair = strings.TrimSpace(pair)
|
|
if k, v, ok := strings.Cut(pair, "="); ok {
|
|
labels[strings.TrimSpace(k)] = strings.TrimSpace(v)
|
|
}
|
|
}
|
|
}
|
|
labels["node.replica-slots"] = strconv.Itoa(maxReplicas)
|
|
body["labels"] = labels
|
|
|
|
return body
|
|
}
|
|
|
|
// heartbeatBody returns the current VRAM/RAM stats for heartbeat payloads.
|
|
//
|
|
// When aggregate VRAM usage is unknown (no GPU, or temporary detection
|
|
// failure), we deliberately OMIT available_vram so the frontend keeps its
|
|
// last good value — overwriting with 0 makes the UI show the node as "fully
|
|
// used", while reporting total-as-available lies to the scheduler about
|
|
// free capacity.
|
|
func (cfg *Config) heartbeatBody() map[string]any {
|
|
body := map[string]any{}
|
|
aggregate := getGPUAggregateInfo()
|
|
if aggregate.TotalVRAM > 0 {
|
|
body["available_vram"] = aggregate.FreeVRAM
|
|
}
|
|
|
|
// RAM availability changes independently from VRAM on discrete-GPU nodes.
|
|
ramInfo, err := getSystemRAMInfo()
|
|
if err != nil {
|
|
xlog.Debug("Failed to detect worker RAM for heartbeat", "error", err)
|
|
} else {
|
|
body["available_ram"] = ramInfo.Available
|
|
}
|
|
|
|
// Free disk changes far faster than VRAM under staging traffic (each
|
|
// accepted model permanently consumes space), so it has to be refreshed
|
|
// every heartbeat rather than only at registration — a node that filled up
|
|
// hours after registering is precisely the case that broke.
|
|
if diskInfo, err := xsysinfo.GetDiskInfo(cfg.ModelsPath); err != nil {
|
|
xlog.Debug("Failed to detect worker models-path disk capacity for heartbeat",
|
|
"path", cfg.ModelsPath, "error", err)
|
|
} else {
|
|
body["total_disk"] = diskInfo.Total
|
|
body["available_disk"] = diskInfo.Available
|
|
}
|
|
|
|
if cpuInfo, err := getCPUInfo(); err != nil {
|
|
xlog.Debug("Failed to sample worker CPU for heartbeat", "error", err)
|
|
} else {
|
|
body["cpu_usage_percent"] = clampCPUUsage(cpuInfo.UsagePercent)
|
|
body["cpu_load_1"] = cpuInfo.Load1
|
|
}
|
|
return body
|
|
}
|