mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-11 08:15:12 -04:00
fix(nodes): never schedule a model onto a node that cannot store it (#11054)
* fix(nodes): never schedule a model onto a node that cannot store it
A worker whose models filesystem was 100% full kept advertising
`status: healthy`, stayed a scheduling candidate, was picked to host a
70 GB video model, accepted the staging request, transferred ~17 GB and
only then failed:
staging .../whisper-large-v3/model.fp32-00001-of-00002.safetensors:
upload to node b7bacbf4-... failed with status 500:
writing file: /models/longcat-video-avatar-1.5/...: no space left on device
The node was at 937G/937G/0-avail. Total elapsed before the truth
surfaced: 16 minutes, for a decision that could never have succeeded.
The worker health signal only ever proved liveness. `/readyz`
(WorkerReadiness/NATSReadiness) checks the NATS link; `status: healthy`
in the registry is driven by heartbeat recency. Node capacity carried
VRAM and RAM but no disk figure at all, and the router compared model
size against VRAM only — nothing anywhere looked at free space on the
filesystem that staging actually writes to.
Report it, then use it:
- Workers now measure the filesystem backing their MODELS directory
(not `/` -- staged weights land in the models path, and that mount is
very often separate) and report `total_disk`/`available_disk` on
registration and on every heartbeat. Free disk moves faster than VRAM
under staging traffic, so the per-heartbeat refresh matters.
- The SmartRouter drops nodes that cannot store the model before it
picks one. The requirement comes from `modelPayloadBytes` -- the same
local paths `stageModelFiles` uploads, already computed for the
size-derived load budget -- plus a 5% / 1 GiB margin, rather than a
fixed percentage of the node's disk. A percentage threshold would take
a small-but-usable node out of rotation for models it could hold, and
on a homogeneous cluster would strand every node at once.
- When no node fits, scheduling fails immediately with an error naming
the requirement and each node's free space, instead of picking one and
discovering it mid-transfer.
Two deliberate non-changes. Low disk does not mark a node `unhealthy`:
the check is per model, so a node too small for one model stays a valid
target for smaller ones. And `total_disk == 0` means "does not report
disk" (pre-upgrade worker, or a failed stat), not "full" -- such nodes
pass through untouched so a rolling upgrade never empties the candidate
pool. A genuinely full node is distinguishable: non-zero total, zero
available. Registry read failures are logged and scheduling continues
unfiltered; a database hiccup must not wedge a cluster.
Free space is surfaced on the node detail page next to VRAM, since the
incident's signature was a node that looked entirely healthy.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]
* feat(nodes): make the disk-headroom check operator-controllable
The admission check added in the previous commit had no off switch. A
scheduler-side veto with no escape hatch is a liability: our size
estimate can be wrong (deduplicating or compressing filesystems, a
backend that fetches its own weights rather than loading the staged
copy), and an operator who hits that has no way out but a downgrade.
Add one knob with two surfaces that share a single source of truth:
- `--distributed-disk-headroom-check` / `LOCALAI_DISTRIBUTED_DISK_HEADROOM_CHECK`
(default true), following the `--distributed-prefix-cache` pattern for
a default-on distributed feature.
- `distributed_disk_headroom_check` in the runtime-settings registry, so
it can be flipped without a restart from `POST /api/settings` and from
Settings -> Distributed in the WebUI.
Both write `DistributedConfig.DiskHeadroomDisabled`, and the SmartRouter
reads that member LIVE on every scheduling decision through a closure
over the application config rather than a value snapshotted at
construction. Env/CLI sets the boot value, the runtime setting overrides
it live, last write wins, and there is exactly one member to read.
Snapshotting would have made the runtime toggle a no-op until restart.
Disabled means WARN, not SKIP. Selection goes back to ignoring free disk
-- byte for byte the pre-check behaviour -- but the check still runs, and
when it would have rejected every node it says so, naming the knob that
suppressed it. Going quiet when switched off would reproduce the exact
condition that made the original incident expensive: a cluster doing
something that could not work and saying nothing. Disabling is also
logged once at startup. Warning only on the total-rejection case keeps
it actionable rather than chatty on a heterogeneous cluster.
Also fixes a false positive in the check itself: shared-models mode
(LOCALAI_DISTRIBUTED_SHARED_MODELS) stages nothing at all -- every node
already mounts this models directory at this path -- so demanding the
full checkpoint size of free space per node would have rejected a
cluster that needs no new bytes. The check is skipped there entirely.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
1 parent
f317da7c0f
commit
6584db992f
28 files changed
+901
-54
No files matched your search
@@ -84,6 +84,12 @@ type RegisterNodeRequest struct {
|
||||
AvailableVRAM uint64 `json:"available_vram,omitempty"`
|
||||
TotalRAM uint64 `json:"total_ram,omitempty"`
|
||||
AvailableRAM uint64 `json:"available_ram,omitempty"`
|
||||
// TotalDisk / AvailableDisk describe the filesystem backing the worker's
|
||||
// MODELS directory (where staged weights land), not the root filesystem.
|
||||
// Omitted by workers that predate the fields; the scheduler treats
|
||||
// total_disk == 0 as "unknown" and leaves such a node in rotation.
|
||||
TotalDisk uint64 `json:"total_disk,omitempty"`
|
||||
AvailableDisk uint64 `json:"available_disk,omitempty"`
|
||||
GPUVendor string `json:"gpu_vendor,omitempty"`
|
||||
// GPUComputeCapability is the worker GPU's compute capability ("major.minor",
|
||||
// e.g. "12.1" for GB10). Used by the router for per-arch option tuning.
|
||||
@@ -176,6 +182,8 @@ func RegisterNodeEndpoint(registry *nodes.NodeRegistry, expectedToken string, au
|
||||
AvailableVRAM: req.AvailableVRAM,
|
||||
TotalRAM: req.TotalRAM,
|
||||
AvailableRAM: req.AvailableRAM,
|
||||
TotalDisk: req.TotalDisk,
|
||||
AvailableDisk: req.AvailableDisk,
|
||||
GPUVendor: req.GPUVendor,
|
||||
GPUComputeCapability: req.GPUComputeCapability,
|
||||
Capability: req.Capability,
|
||||
@@ -372,7 +380,8 @@ func HeartbeatEndpoint(registry *nodes.NodeRegistry) echo.HandlerFunc {
|
||||
_ = c.Bind(&update) // best-effort — empty body is fine
|
||||
|
||||
var updatePtr *nodes.HeartbeatUpdate
|
||||
if update.AvailableVRAM != nil || update.TotalVRAM != nil || update.AvailableRAM != nil || update.GPUVendor != "" {
|
||||
if update.AvailableVRAM != nil || update.TotalVRAM != nil || update.AvailableRAM != nil ||
|
||||
update.AvailableDisk != nil || update.TotalDisk != nil || update.GPUVendor != "" {
|
||||
updatePtr = &update
|
||||
}
|
||||
|
||||
|
||||
@@ -23,6 +23,7 @@
|
||||
"agents": "Agentenaufgaben",
|
||||
"agentpool": "Agenten-Pool",
|
||||
"assistant": "LocalAI Assistant",
|
||||
"distributed": "Verteilt",
|
||||
"responses": "Antworten"
|
||||
}
|
||||
},
|
||||
|
||||
@@ -23,6 +23,7 @@
|
||||
"agents": "Agent Jobs",
|
||||
"agentpool": "Agent Pool",
|
||||
"assistant": "LocalAI Assistant",
|
||||
"distributed": "Distributed",
|
||||
"responses": "Responses"
|
||||
}
|
||||
},
|
||||
|
||||
@@ -23,6 +23,7 @@
|
||||
"agents": "Trabajos de agentes",
|
||||
"agentpool": "Pool de agentes",
|
||||
"assistant": "LocalAI Assistant",
|
||||
"distributed": "Distribuido",
|
||||
"responses": "Respuestas"
|
||||
}
|
||||
},
|
||||
|
||||
@@ -23,6 +23,7 @@
|
||||
"agents": "Agent Job",
|
||||
"agentpool": "Agent Pool",
|
||||
"assistant": "Asisten LocalAI",
|
||||
"distributed": "Terdistribusi",
|
||||
"responses": "Respons"
|
||||
}
|
||||
},
|
||||
|
||||
@@ -23,6 +23,7 @@
|
||||
"agents": "Job degli agenti",
|
||||
"agentpool": "Pool agenti",
|
||||
"assistant": "LocalAI Assistant",
|
||||
"distributed": "Distribuito",
|
||||
"responses": "Risposte"
|
||||
}
|
||||
},
|
||||
|
||||
@@ -23,6 +23,7 @@
|
||||
"agents": "에이전트 작업",
|
||||
"agentpool": "에이전트 풀",
|
||||
"assistant": "LocalAI 어시스턴트",
|
||||
"distributed": "분산",
|
||||
"responses": "응답"
|
||||
}
|
||||
},
|
||||
|
||||
@@ -23,6 +23,7 @@
|
||||
"agents": "智能体任务",
|
||||
"agentpool": "智能体池",
|
||||
"assistant": "LocalAI Assistant",
|
||||
"distributed": "分布式",
|
||||
"responses": "响应"
|
||||
}
|
||||
},
|
||||
|
||||
@@ -96,6 +96,16 @@ export default function NodeDetail() {
|
||||
<span className="cell-mono">{formatVRAM(usedVRAM) || '0'} / {formatVRAM(node.total_vram)}</span>
|
||||
</div>
|
||||
)}
|
||||
{node.total_disk > 0 && (
|
||||
<div>
|
||||
{/* Free space on the worker's MODELS filesystem. A node can look
|
||||
perfectly healthy on VRAM while having nowhere to put the
|
||||
weights, which is why this sits next to VRAM rather than
|
||||
buried in a diagnostics panel. */}
|
||||
<div className="drawer-eyebrow">Models disk free</div>
|
||||
<span className="cell-mono">{formatVRAM(node.available_disk || 0) || '0'} / {formatVRAM(node.total_disk)}</span>
|
||||
</div>
|
||||
)}
|
||||
<div>
|
||||
<div className="drawer-eyebrow">In-flight</div>
|
||||
<span className="cell-mono">{node.in_flight_count || 0}</span>
|
||||
|
||||
@@ -26,6 +26,7 @@ const SECTIONS = [
|
||||
{ id: 'agents', icon: 'fa-tasks', color: 'var(--color-primary)' },
|
||||
{ id: 'agentpool', icon: 'fa-robot', color: 'var(--color-primary)' },
|
||||
{ id: 'assistant', icon: 'fa-user-shield', color: 'var(--color-accent)' },
|
||||
{ id: 'distributed', icon: 'fa-server', color: 'var(--color-accent)' },
|
||||
{ id: 'responses', icon: 'fa-database', color: 'var(--color-accent)' },
|
||||
]
|
||||
|
||||
@@ -640,6 +641,18 @@ export default function Settings() {
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{/* Distributed mode */}
|
||||
<div ref={el => sectionRefs.current.distributed = el} style={{ marginBottom: 'var(--spacing-xl)' }}>
|
||||
<h3 style={{ fontSize: '1rem', fontWeight: 700, display: 'flex', alignItems: 'center', gap: 'var(--spacing-sm)', marginBottom: 'var(--spacing-md)' }}>
|
||||
<i className="fas fa-server" style={{ color: 'var(--color-accent)' }} /> {t('settings.sections.distributed')}
|
||||
</h3>
|
||||
<div className="card">
|
||||
<SettingRow label="Disk headroom check" description="Reject worker nodes that lack free space to store the model, at scheduling time rather than partway through staging. Free space is measured on each worker's models filesystem and compared against the model's own size plus a small margin. Turning this off restores selection that ignores free disk; the check still runs and warns when it would have rejected every node. Takes effect without restart.">
|
||||
<Toggle checked={settings.distributed_disk_headroom_check ?? true} onChange={(v) => update('distributed_disk_headroom_check', v)} />
|
||||
</SettingRow>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{/* Open Responses */}
|
||||
<div ref={el => sectionRefs.current.responses = el} style={{ marginBottom: 'var(--spacing-xl)' }}>
|
||||
<h3 style={{ fontSize: '1rem', fontWeight: 700, display: 'flex', alignItems: 'center', gap: 'var(--spacing-sm)', marginBottom: 'var(--spacing-md)' }}>
|
||||
|
||||
Reference in new issue
Block a user