Files
LocalAI/docs/content/features/model-failover.md
T
Ettore Di Giacinto 1b6b4b806a docs: document localai-proxy and distributed failover limits
Add the localai-proxy known limits (no grammar or media forwarding,
TTS streams that end cleanly after an upstream failure, the /v1 path in
upstream_url), state that the Unimplemented skip covers the APIs that
answer HTTP 501, and describe a NATS-partitioned leader and pin
re-sync in distributed mode.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00

305 lines
12 KiB
Markdown

+++
disableToc = false
title = "Model Failover"
weight = 15
url = "/features/model-failover/"
+++
A **failover chain** is a model name that is served by an ordered list of
other models. LocalAI sends each request to the first healthy target. When a
target fails, the request moves to the next target, and later requests stay
there until the first target has recovered.
Use it to serve a model from a remote LocalAI or another OpenAI-compatible
provider, and to fall back to a local model when the remote one is down.
## Declaring a chain
```yaml
name: assistant-llm
failover:
targets:
- model: argus-llm # for example a localai-proxy or cloud-proxy model
- model: gemma-local
warm: true # keep it loaded
```
Clients call `assistant-llm`. Each target is a normal model config. A chain
has no `backend` and no `parameters.model`.
Optional settings, with their defaults:
```yaml
failover:
probe:
interval: 15s # how often an idle target is checked
timeout: 5s
trip:
errors: 1 # failures within the window that mark a target down
window: 30s
recovery:
probes: 3 # test requests a target must pass before it is used again
min_dwell: 60s # minimum time on a lower target before moving back
```
Rules:
- A chain needs at least 2 targets. A target can be an alias, but not another
chain.
- A chain cannot also set `alias` or `backend`.
- Responses name the chain as the model. The `X-LocalAI-Served-Model` header
names the target that served the request.
- A remote (`localai-proxy` or `cloud-proxy`) target receives its own model
name, never the chain name: `proxy.upstream_model`, or the target name when
`upstream_model` is empty. The health check looks for the same name.
## How the target is chosen
- The active target is the first healthy target in the list.
- When a target fails, LocalAI marks it down and moves to the next target at
once.
- LocalAI moves back to a higher target only when that target has passed
`recovery.probes` test requests **and** the current target has been active
for at least `recovery.min_dwell`. This stops an unstable upstream from
moving traffic back and forth.
- When all targets are down, the chain is `degraded`. Each request still tries
every target in order.
## Retry inside a request
When a target fails before the response starts, LocalAI sends the same
request to the next target. The client does not see the failure.
- LocalAI does not retry after the first byte of a response is sent (for
example after the first streamed token). The request fails, the target is
marked down, and the next request uses the next target.
- A target that is at its concurrency limit (an admission rejection) or that
is disabled is skipped for that request without being marked down.
- LocalAI does not retry client errors (4xx), such as a prompt that is too
long, because the next target would reject it too. A 4xx counts neither as
a success nor as a failure for the target.
- Request bodies larger than 32 MiB are not retried.
When the primary did not serve the request, the response has the header
`X-LocalAI-Failover: fallback`, or `X-LocalAI-Failover: degraded` when all
targets were down.
## Health checks
| Target | Regular check | Check before moving back |
|---|---|---|
| Remote (`localai-proxy`, `cloud-proxy`) | `GET /v1/models` on the upstream lists the model | one small real request, for example a 1-token completion |
| Local, `warm: true` | the backend answers a health check. A check never loads the model: while it is not loaded, the check passes and real requests judge it | one small real request. While the model is not loaded, the target is used again after `min_dwell` |
| Local, not warm | none: judged only by real requests; it is never loaded only to check it | none: the target is used again after `min_dwell` |
A request that succeeds counts as a check, so a busy target is almost never
probed.
When a target is in more than one chain, its check settings come from the
first of those chains in name order.
## Warm targets
`warm: true` loads a local target at startup and protects it from idle and
LRU eviction, so a switch does not wait for the model to load. Warm targets
count toward the active backend limit (`--max-active-backends`) like any
pinned model: LocalAI never evicts them to make room, and if they fill the
limit, a new model still loads rather than being blocked.
`warm` applies only to local targets. On a remote (`localai-proxy` or
`cloud-proxy`) target it has no effect, and LocalAI logs a warning when it
loads the chain.
## Realtime pipelines
A pipeline stage can name a chain:
```yaml
name: assistant
pipeline:
vad: silero-vad
transcription: whisper-chain
llm: assistant-llm
tts: voice-chain
```
LocalAI resolves the chain for every call of the stage, in full pipelines and
in transcription-only and sound-detection-only sessions. When a chain switches,
the session stays open and keeps its conversation. The next turn uses the new
target.
The session receives a `localai.model.failover` event for each chain stage when
it starts (`reason: initial`) and each time a chain switches:
```json
{"type":"localai.model.failover","chain":"assistant-llm","stage":"llm",
"from":"argus-llm","to":"gemma-local","state":"fallback","reason":"trip"}
```
### Example: stages on a remote LocalAI
This pipeline runs its transcription, LLM and TTS stages on a remote LocalAI
(`argus`) through [`localai-proxy`]({{% relref "operations/cloud-proxy" %}})
models, and uses local models when the remote instance is down. Each stage has
its own chain, so one stage can fail over while the others stay remote.
```yaml
# Remote targets: each one names the model on the upstream LocalAI.
name: argus-stt
backend: localai-proxy
known_usecases: [transcript]
options:
- realtime_pipeline:asr-pipeline # upstream pipeline for live transcription
proxy:
upstream_url: http://argus.lan:8080
upstream_model: parakeet
---
name: argus-llm
backend: localai-proxy
known_usecases: [chat]
proxy:
upstream_url: http://argus.lan:8080
upstream_model: gemma-3-12b
---
name: argus-tts
backend: localai-proxy
known_usecases: [tts]
proxy:
upstream_url: http://argus.lan:8080
upstream_model: kokoro
---
# One chain per stage, remote first, local second.
name: stt-chain
failover:
targets: [{model: argus-stt}, {model: whisper-local, warm: true}]
---
name: llm-chain
failover:
targets: [{model: argus-llm}, {model: gemma-local, warm: true}]
---
name: tts-chain
failover:
targets: [{model: argus-tts}, {model: piper-local}]
---
name: assistant
pipeline:
vad: silero-vad
transcription: stt-chain
llm: llm-chain
tts: tts-chain
```
The example shows the configs as one YAML stream; put each config in its own
file in the models directory. When `argus` stops
answering, the next call of each stage fails over to the local model and the
session receives a `localai.model.failover` event for that stage. A remote
target that does not support a call (it returns `Unimplemented`) is skipped for
that call and is not marked down.
Limits:
- After a `session.update` that changes the pipeline, `localai.model.failover`
events keep describing the chains from session start.
- A chain used as a router candidate, or as the classifier-mode scoring model,
is not resolved per call.
## Watching failover
- `GET /api/failover` lists every chain, its active target and the state of
each target.
- `GET /api/failover/{chain}` returns one chain.
- `GET /api/failover/events` is a server-sent event stream. The first event is
`snapshot` with the full state. Then `chain.switched` and `target.state`
events follow.
- Metrics: `localai_failover_switches_total{chain,from,to,reason}` and
`localai_failover_target_up{target}`.
- With tracing on, each skipped target appears in the Traces view with the
error that made LocalAI skip it.
## Pinning a target
An admin can force a chain to one target, for example during maintenance:
```bash
curl -X POST http://localhost:8080/api/failover/assistant-llm/pin \
-H 'Content-Type: application/json' -d '{"target":"gemma-local"}'
curl -X DELETE http://localhost:8080/api/failover/assistant-llm/pin
```
While a chain is pinned, only the pinned target serves it. Health checks
continue. On a single LocalAI instance, a restart removes the pin. In
[distributed mode](#distributed-mode), pins persist.
## In the UI
- The model editor's template picker has a **Failover Chain** template that
starts a new model with two empty targets. Any model can also add a
`failover.targets` field directly: each row picks a target model and a
**warm** toggle, with move up/down and remove controls. The toggle is
available on every row; `warm` has no effect on a remote (`localai-proxy` or
`cloud-proxy`) target, and LocalAI logs a warning when it loads a chain with
one set. The field flags fewer than two targets, a duplicate target, and the
chain naming itself as one of its own targets.
- Opening an existing chain in the model editor shows a health strip: the
chain's status pill, its active target and how long it has been active,
and a table of every target with its kind, warm flag, status, last probe
time and last error. An admin sees a **Pin** button on each target and an
**Unpin** action for the chain, both behind a confirmation dialog.
- The **Failover** page (`/app/failover`, admin only, linked from the
console navigation) lists every chain with its status pill, active target,
a small pill per target, and time since the last switch. It links each
chain name to its model editor page and shows an empty state linking to
the failover template when no chains exist yet.
- The Installed Models list badges a model that belongs to a chain with
`chain → <active target>`, next to the alias badge.
- All of the above update live from the same event stream as
`GET /api/failover/events` — the strip, the overview page and the badge do
not need a page refresh to reflect a `chain.switched` or `target.state`
event.
## Assistant and MCP
The LocalAI Assistant and `local-ai mcp-server` offer `list_failover_chains`,
`pin_failover_target` and `unpin_failover_target`. Create and edit chains with
the model config tools, like any other model.
## Distributed mode
In [distributed mode]({{%relref "features/distributed-mode" %}}), all frontends
share one failover state:
- Pins apply to the whole cluster. LocalAI stores them in the database, so
they persist across restarts. A pin set on one frontend applies on all.
- Target health and the active target of each chain are shared over NATS, so
all frontends converge on the same target for a chain.
- One frontend, the probe leader, runs the health checks, decides fail-over
and fail-back, and loads warm targets. The leader holds a PostgreSQL
advisory lock and keeps it until it stops or its database connection fails.
Then another frontend takes the lock and becomes the leader: immediately
when the leader shuts down or its process exits, and within about 30 seconds
when the leader's host or network fails.
- Warm targets stay loaded on the workers. The router and the replica
reconciler treat them like pinned models and do not evict them.
- A frontend that starts late gets the current state within 10 seconds,
because the leader sends its full state again every 10 seconds.
- If PostgreSQL is not available, no frontend holds the lock. Health checks
and fail-back stop until the database is back. Requests still fail over to
the next target when a target fails during the request.
- If a frontend cannot start the shared state, it logs an error and manages
failover alone, as a single LocalAI instance does.
- Leadership follows the database lock, not NATS. A leader that loses its NATS
connection but keeps its database connection stays the leader. Until NATS
recovers, the other frontends keep the last decisions they received from it,
and its new decisions do not reach them.
- Each frontend reads the pins from the database again every 30 seconds and
after a NATS reconnect, and applies any change within 10 seconds. A frontend
that missed a pin or an unpin catches up in this way. If a pin cannot be written to the database, LocalAI returns an
error and restores the previous pin.
## Limits
- Chains do not nest.
- See also [model aliases]({{%relref "features/model-aliases" %}}) and the
[realtime API]({{%relref "features/openai-realtime" %}}).