diff --git a/docs/content/features/model-aliases.md b/docs/content/features/model-aliases.md index ed52f0977..292c05374 100644 --- a/docs/content/features/model-aliases.md +++ b/docs/content/features/model-aliases.md @@ -38,6 +38,9 @@ That is the whole config: a `name` (the alias clients call) and an `alias` key - Usage accounting records both sides: requested `gpt-4`, served `my-llama-3`. - Aliases work for every modality (chat, embeddings, audio, images, and so on). +To serve a name from several models with automatic fallback, use a [failover +chain]({{%relref "features/model-failover" %}}). + ## Managing aliases You can create, swap, and remove aliases from any of the management surfaces. diff --git a/docs/content/features/model-failover.md b/docs/content/features/model-failover.md new file mode 100644 index 000000000..b3da3f7d9 --- /dev/null +++ b/docs/content/features/model-failover.md @@ -0,0 +1,179 @@ + ++++ +disableToc = false +title = "Model Failover" +weight = 15 +url = "/features/model-failover/" ++++ + +A **failover chain** is a model name that is served by an ordered list of +other models. LocalAI sends each request to the first healthy target. When a +target fails, the request moves to the next target, and later requests stay +there until the first target has recovered. + +Use it to serve a model from a remote LocalAI or another OpenAI-compatible +provider, and to fall back to a local model when the remote one is down. + +## Declaring a chain + +```yaml +name: assistant-llm +failover: + targets: + - model: argus-llm # for example a cloud-proxy model + - model: gemma-local + warm: true # keep it loaded +``` + +Clients call `assistant-llm`. Each target is a normal model config. A chain +has no `backend` and no `parameters.model`. + +Optional settings, with their defaults: + +```yaml +failover: + probe: + interval: 15s # how often an idle target is checked + timeout: 5s + trip: + errors: 1 # failures within the window that mark a target down + window: 30s + recovery: + probes: 3 # test requests a target must pass before it is used again + min_dwell: 60s # minimum time on a lower target before moving back +``` + +Rules: + +- A chain needs at least 2 targets. A target can be an alias, but not another + chain. +- A chain cannot also set `alias` or `backend`. +- Responses name the chain as the model. The `X-LocalAI-Served-Model` header + names the target that served the request. + +## How the target is chosen + +- The active target is the first healthy target in the list. +- When a target fails, LocalAI marks it down and moves to the next target at + once. +- LocalAI moves back to a higher target only when that target has passed + `recovery.probes` test requests **and** the current target has been active + for at least `recovery.min_dwell`. This stops an unstable upstream from + moving traffic back and forth. +- When all targets are down, the chain is `degraded`. Each request still tries + every target in order. + +## Retry inside a request + +When a target fails before the response starts, LocalAI sends the same +request to the next target. The client does not see the failure. + +- LocalAI does not retry after the first byte of a response is sent (for + example after the first streamed token). The request fails, the target is + marked down, and the next request uses the next target. +- A target that is at its concurrency limit (an admission rejection) or that + is disabled is skipped for that request without being marked down. +- LocalAI does not retry client errors (4xx), such as a prompt that is too + long, because the next target would reject it too. A 4xx counts neither as + a success nor as a failure for the target. +- Request bodies larger than 32 MiB are not retried. + +When the primary did not serve the request, the response has the header +`X-LocalAI-Failover: fallback`, or `X-LocalAI-Failover: degraded` when all +targets were down. + +## Health checks + +| Target | Regular check | Check before moving back | +|---|---|---| +| Remote (`cloud-proxy`) | `GET /v1/models` on the upstream lists the model | one small real request, for example a 1-token completion | +| Local, `warm: true` | the backend answers a health check | one small real request | +| Local, not warm | none: judged only by real requests; it is never loaded only to check it | none: the target is used again after `min_dwell` | + +A request that succeeds counts as a check, so a busy target is almost never +probed. + +When a target is in more than one chain, its check settings come from the +first of those chains in name order. + +## Warm targets + +`warm: true` loads a local target at startup and protects it from idle and +LRU eviction, so a switch does not wait for the model to load. Warm targets +count toward the active backend limit (`--max-active-backends`) like any +pinned model: LocalAI never evicts them to make room, and if they fill the +limit, a new model still loads rather than being blocked. + +## Realtime pipelines + +A pipeline stage can name a chain: + +```yaml +name: assistant +pipeline: + vad: silero-vad + transcription: whisper-chain + llm: assistant-llm + tts: voice-chain +``` + +LocalAI resolves the chain for every call of the stage. When a chain switches, +the session stays open and keeps its conversation. The next turn uses the new +target. + +The session receives a `localai.model.failover` event for each chain stage when +it starts (`reason: initial`) and each time a chain switches: + +```json +{"type":"localai.model.failover","chain":"assistant-llm","stage":"llm", + "from":"argus-llm","to":"gemma-local","state":"fallback","reason":"trip"} +``` + +Limits: + +- Chains are resolved only in full realtime pipelines. A transcription-only or + sound-detection-only session does not resolve chains yet. +- After a `session.update` that changes the pipeline, `localai.model.failover` + events keep describing the chains from session start. +- A chain used as a router candidate, or as the classifier-mode scoring model, + is not resolved per call. + +## Watching failover + +- `GET /api/failover` lists every chain, its active target and the state of + each target. +- `GET /api/failover/{chain}` returns one chain. +- `GET /api/failover/events` is a server-sent event stream. The first event is + `snapshot` with the full state. Then `chain.switched` and `target.state` + events follow. +- Metrics: `localai_failover_switches_total{chain,from,to,reason}` and + `localai_failover_target_up{target}`. +- With tracing on, each skipped target appears in the Traces view with the + error that made LocalAI skip it. + +## Pinning a target + +An admin can force a chain to one target, for example during maintenance: + +```bash +curl -X POST http://localhost:8080/api/failover/assistant-llm/pin \ + -H 'Content-Type: application/json' -d '{"target":"gemma-local"}' +curl -X DELETE http://localhost:8080/api/failover/assistant-llm/pin +``` + +While a chain is pinned, only the pinned target serves it. Health checks +continue. A restart removes the pin. + +## Assistant and MCP + +The LocalAI Assistant and `local-ai mcp-server` offer `list_failover_chains`, +`pin_failover_target` and `unpin_failover_target`. Create and edit chains with +the model config tools, like any other model. + +## Limits + +- Failover state is kept in memory by each LocalAI instance. Several frontends + in distributed mode each keep their own view. +- Chains do not nest. +- See also [model aliases]({{%relref "features/model-aliases" %}}) and the + [realtime API]({{%relref "features/openai-realtime" %}}). diff --git a/docs/content/features/openai-realtime.md b/docs/content/features/openai-realtime.md index 54bec58ab..ac7afb484 100644 --- a/docs/content/features/openai-realtime.md +++ b/docs/content/features/openai-realtime.md @@ -31,6 +31,8 @@ This configuration links the following components: Make sure all referenced models (`silero-vad-ggml`, `whisper-large-turbo`, `qwen3-4b`, `tts-1`) are also installed or defined in your LocalAI instance. +A pipeline stage can name a [failover chain]({{%relref "features/model-failover" %}}); the stage then switches targets without closing the session. + ### Streaming the pipeline By default each stage runs to completion before the next begins: the whole utterance is transcribed, the full LLM reply is generated, then it is synthesized. Each stage can instead be streamed incrementally, which lowers the time-to-first-audio of a turn: diff --git a/docs/content/operations/cloud-proxy.md b/docs/content/operations/cloud-proxy.md index 02af25bd0..bd1d806d4 100644 --- a/docs/content/operations/cloud-proxy.md +++ b/docs/content/operations/cloud-proxy.md @@ -29,6 +29,9 @@ egress remains subject to the same redaction rules a local model would apply. - Use the intelligent router to send small or simple prompts to a local model and complex ones to Claude or GPT-4o. +To fall back to a local model when the upstream is down, list the proxy model +in a [failover chain]({{%relref "features/model-failover" %}}). + ## How it works 1. Request hits LocalAI on `/v1/chat/completions` (OpenAI-shaped) or