mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-29 09:35:02 -04:00
A chain request reached a cloud-proxy target with the client's model, the chain name, whenever the target set no upstream_model: passthrough forwards the body's model and translate falls back to it. The upstream answered 404, which neither retries nor trips, while the liveness probe, which checks the target's own name, kept passing. PrepareTarget now sets the upstream model of a remote target to proxy.upstream_model or the target name, the same name the probe uses. The request pipeline and realtime chain stages both call it. Assisted-by: Claude:claude-opus-5-5
183 lines
6.8 KiB
Markdown
183 lines
6.8 KiB
Markdown
|
|
+++
|
|
disableToc = false
|
|
title = "Model Failover"
|
|
weight = 15
|
|
url = "/features/model-failover/"
|
|
+++
|
|
|
|
A **failover chain** is a model name that is served by an ordered list of
|
|
other models. LocalAI sends each request to the first healthy target. When a
|
|
target fails, the request moves to the next target, and later requests stay
|
|
there until the first target has recovered.
|
|
|
|
Use it to serve a model from a remote LocalAI or another OpenAI-compatible
|
|
provider, and to fall back to a local model when the remote one is down.
|
|
|
|
## Declaring a chain
|
|
|
|
```yaml
|
|
name: assistant-llm
|
|
failover:
|
|
targets:
|
|
- model: argus-llm # for example a cloud-proxy model
|
|
- model: gemma-local
|
|
warm: true # keep it loaded
|
|
```
|
|
|
|
Clients call `assistant-llm`. Each target is a normal model config. A chain
|
|
has no `backend` and no `parameters.model`.
|
|
|
|
Optional settings, with their defaults:
|
|
|
|
```yaml
|
|
failover:
|
|
probe:
|
|
interval: 15s # how often an idle target is checked
|
|
timeout: 5s
|
|
trip:
|
|
errors: 1 # failures within the window that mark a target down
|
|
window: 30s
|
|
recovery:
|
|
probes: 3 # test requests a target must pass before it is used again
|
|
min_dwell: 60s # minimum time on a lower target before moving back
|
|
```
|
|
|
|
Rules:
|
|
|
|
- A chain needs at least 2 targets. A target can be an alias, but not another
|
|
chain.
|
|
- A chain cannot also set `alias` or `backend`.
|
|
- Responses name the chain as the model. The `X-LocalAI-Served-Model` header
|
|
names the target that served the request.
|
|
- A remote (`cloud-proxy`) target receives its own model name, never the chain
|
|
name: `proxy.upstream_model`, or the target name when `upstream_model` is
|
|
empty. The health check looks for the same name.
|
|
|
|
## How the target is chosen
|
|
|
|
- The active target is the first healthy target in the list.
|
|
- When a target fails, LocalAI marks it down and moves to the next target at
|
|
once.
|
|
- LocalAI moves back to a higher target only when that target has passed
|
|
`recovery.probes` test requests **and** the current target has been active
|
|
for at least `recovery.min_dwell`. This stops an unstable upstream from
|
|
moving traffic back and forth.
|
|
- When all targets are down, the chain is `degraded`. Each request still tries
|
|
every target in order.
|
|
|
|
## Retry inside a request
|
|
|
|
When a target fails before the response starts, LocalAI sends the same
|
|
request to the next target. The client does not see the failure.
|
|
|
|
- LocalAI does not retry after the first byte of a response is sent (for
|
|
example after the first streamed token). The request fails, the target is
|
|
marked down, and the next request uses the next target.
|
|
- A target that is at its concurrency limit (an admission rejection) or that
|
|
is disabled is skipped for that request without being marked down.
|
|
- LocalAI does not retry client errors (4xx), such as a prompt that is too
|
|
long, because the next target would reject it too. A 4xx counts neither as
|
|
a success nor as a failure for the target.
|
|
- Request bodies larger than 32 MiB are not retried.
|
|
|
|
When the primary did not serve the request, the response has the header
|
|
`X-LocalAI-Failover: fallback`, or `X-LocalAI-Failover: degraded` when all
|
|
targets were down.
|
|
|
|
## Health checks
|
|
|
|
| Target | Regular check | Check before moving back |
|
|
|---|---|---|
|
|
| Remote (`cloud-proxy`) | `GET /v1/models` on the upstream lists the model | one small real request, for example a 1-token completion |
|
|
| Local, `warm: true` | the backend answers a health check. A check never loads the model: while it is not loaded, the check passes and real requests judge it | one small real request. While the model is not loaded, the target is used again after `min_dwell` |
|
|
| Local, not warm | none: judged only by real requests; it is never loaded only to check it | none: the target is used again after `min_dwell` |
|
|
|
|
A request that succeeds counts as a check, so a busy target is almost never
|
|
probed.
|
|
|
|
When a target is in more than one chain, its check settings come from the
|
|
first of those chains in name order.
|
|
|
|
## Warm targets
|
|
|
|
`warm: true` loads a local target at startup and protects it from idle and
|
|
LRU eviction, so a switch does not wait for the model to load. Warm targets
|
|
count toward the active backend limit (`--max-active-backends`) like any
|
|
pinned model: LocalAI never evicts them to make room, and if they fill the
|
|
limit, a new model still loads rather than being blocked.
|
|
|
|
## Realtime pipelines
|
|
|
|
A pipeline stage can name a chain:
|
|
|
|
```yaml
|
|
name: assistant
|
|
pipeline:
|
|
vad: silero-vad
|
|
transcription: whisper-chain
|
|
llm: assistant-llm
|
|
tts: voice-chain
|
|
```
|
|
|
|
LocalAI resolves the chain for every call of the stage. When a chain switches,
|
|
the session stays open and keeps its conversation. The next turn uses the new
|
|
target.
|
|
|
|
The session receives a `localai.model.failover` event for each chain stage when
|
|
it starts (`reason: initial`) and each time a chain switches:
|
|
|
|
```json
|
|
{"type":"localai.model.failover","chain":"assistant-llm","stage":"llm",
|
|
"from":"argus-llm","to":"gemma-local","state":"fallback","reason":"trip"}
|
|
```
|
|
|
|
Limits:
|
|
|
|
- Chains are resolved only in full realtime pipelines. A transcription-only or
|
|
sound-detection-only session does not resolve chains yet.
|
|
- After a `session.update` that changes the pipeline, `localai.model.failover`
|
|
events keep describing the chains from session start.
|
|
- A chain used as a router candidate, or as the classifier-mode scoring model,
|
|
is not resolved per call.
|
|
|
|
## Watching failover
|
|
|
|
- `GET /api/failover` lists every chain, its active target and the state of
|
|
each target.
|
|
- `GET /api/failover/{chain}` returns one chain.
|
|
- `GET /api/failover/events` is a server-sent event stream. The first event is
|
|
`snapshot` with the full state. Then `chain.switched` and `target.state`
|
|
events follow.
|
|
- Metrics: `localai_failover_switches_total{chain,from,to,reason}` and
|
|
`localai_failover_target_up{target}`.
|
|
- With tracing on, each skipped target appears in the Traces view with the
|
|
error that made LocalAI skip it.
|
|
|
|
## Pinning a target
|
|
|
|
An admin can force a chain to one target, for example during maintenance:
|
|
|
|
```bash
|
|
curl -X POST http://localhost:8080/api/failover/assistant-llm/pin \
|
|
-H 'Content-Type: application/json' -d '{"target":"gemma-local"}'
|
|
curl -X DELETE http://localhost:8080/api/failover/assistant-llm/pin
|
|
```
|
|
|
|
While a chain is pinned, only the pinned target serves it. Health checks
|
|
continue. A restart removes the pin.
|
|
|
|
## Assistant and MCP
|
|
|
|
The LocalAI Assistant and `local-ai mcp-server` offer `list_failover_chains`,
|
|
`pin_failover_target` and `unpin_failover_target`. Create and edit chains with
|
|
the model config tools, like any other model.
|
|
|
|
## Limits
|
|
|
|
- Failover state is kept in memory by each LocalAI instance. Several frontends
|
|
in distributed mode each keep their own view.
|
|
- Chains do not nest.
|
|
- See also [model aliases]({{%relref "features/model-aliases" %}}) and the
|
|
[realtime API]({{%relref "features/openai-realtime" %}}).
|