fix(failover): warn about warm on a remote target, align the spec

The spec promised a load-time warning when a chain marks a remote
target warm, where the flag does nothing; the loader now logs it. The
remote-backend test moves into ModelConfig.IsRemoteProxy so the loader
and the failover manager agree on what is remote.

The spec now says what ships: a load blocked by pinned warm targets
proceeds over the limit after eviction retries, without an error that
names them.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
Ettore Di Giacinto committed 2026-09-27 07:42:20 +00:00
1 parent 8dbe7a0e3f
commit f7f12f8203
6 files changed
+53 -3

No files matched your search

+3
View File
@@ -107,6 +107,9 @@ count toward the active backend limit (`--max-active-backends`) like any
pinned model: LocalAI never evicts them to make room, and if they fill the
limit, a new model still loads rather than being blocked.
`warm` applies only to local targets. On a remote (`cloud-proxy`) target it
has no effect, and LocalAI logs a warning when it loads the chain.
## Realtime pipelines
A pipeline stage can name a chain:
@@ -190,7 +190,9 @@ this.
The manager loads `warm: true` targets at startup and marks them pinned in the
watchdog, so LRU and idle eviction skip them. They still count toward the
active backend limit. When pinned warm targets leave no room for another load,
that load fails with an error that names them. The docs state this.
the loader never evicts them: it retries eviction and then loads the model
anyway, over the limit, with no error that names the warm targets. The docs
state this.
### Events