mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-30 18:14:32 -04:00
fix(failover): warn about warm on a remote target, align the spec
The spec promised a load-time warning when a chain marks a remote target warm, where the flag does nothing; the loader now logs it. The remote-backend test moves into ModelConfig.IsRemoteProxy so the loader and the failover manager agree on what is remote. The spec now says what ships: a load blocked by pinned warm targets proceeds over the limit after eviction retries, without an error that names them. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
1 parent
8dbe7a0e3f
commit
f7f12f8203
6 files changed
+53
-3
No files matched your search
@@ -107,6 +107,9 @@ count toward the active backend limit (`--max-active-backends`) like any
|
||||
pinned model: LocalAI never evicts them to make room, and if they fill the
|
||||
limit, a new model still loads rather than being blocked.
|
||||
|
||||
`warm` applies only to local targets. On a remote (`cloud-proxy`) target it
|
||||
has no effect, and LocalAI logs a warning when it loads the chain.
|
||||
|
||||
## Realtime pipelines
|
||||
|
||||
A pipeline stage can name a chain:
|
||||
|
||||
@@ -190,7 +190,9 @@ this.
|
||||
The manager loads `warm: true` targets at startup and marks them pinned in the
|
||||
watchdog, so LRU and idle eviction skip them. They still count toward the
|
||||
active backend limit. When pinned warm targets leave no room for another load,
|
||||
that load fails with an error that names them. The docs state this.
|
||||
the loader never evicts them: it retries eviction and then loads the model
|
||||
anyway, over the limit, with no error that names the warm targets. The docs
|
||||
state this.
|
||||
|
||||
### Events
|
||||
|
||||
|
||||
Reference in new issue
Block a user