docs: design distributed failover, localai-proxy and the chain UI

Failover state is per frontend, remote chain targets only cover chat,
and chains have no UI. Design shared pins, health and chain state for
distributed mode, a localai-proxy backend for every API including live
transcription, a chain editor and health view, and a contributor rule
for distributed-aware state.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
Ettore Di Giacinto committed 2026-09-27 07:42:20 +00:00
1 parent e62854c340
commit b86902e78e
1 file changed
+335
@@ -0,0 +1,335 @@
# Failover chains: distributed mode, localai-proxy backend and WebUI
Date: 2026-09-26
Status: design approved in brainstorming, pending spec review
Builds on: `2026-09-26-model-failover-chains-design.md` (same PR)
## Problem
The failover chains in this PR work on one LocalAI instance. Three gaps
remain:
1. **Distributed mode.** With several frontends, chain definitions converge
(config edits broadcast `cache.invalidate.models`), but runtime state does
not. Each frontend probes, trips and pins on its own. A pin applies only
on the frontend that received it and is lost on restart. `/api/failover`
and its event stream show a different view on each frontend. `warm: true`
pins only a frontend stub, so workers can evict the model, and every
frontend preloads it.
2. **Remote targets cover only chat.** `cloud-proxy` forwards chat and
completions. A remote LocalAI cannot serve transcription, TTS, VAD, sound
detection or the other modalities as a chain target, so a realtime
pipeline cannot fail over per stage between a remote and a local LocalAI.
3. **No UI.** Chains can be edited only as raw JSON in the model editor, and
their health is visible only through the API.
LocalAI also has no rule that makes a feature state how it behaves with
several frontends. The failover feature shipped with per-instance state
because nothing asked the question.
## Goals
- **C. Distributed-aware failover.** Pins, target health and chain state are
the same on every frontend. One frontend probes. Warm targets stay loaded on
workers and are preloaded once.
- **A. `localai-proxy` backend.** A gRPC backend that serves every backend
method with a REST counterpart by calling an upstream LocalAI, including
live transcription through the upstream's realtime API.
- **B. WebUI.** A chain editor field, a chain template, a live health strip
per chain, a "chain" badge in the model list, and a Failover overview page.
- **D. Contributor rule.** `AGENTS.md` and a new `.agents/distributed-state.md`
require every stateful feature to choose and document a distributed mode.
## Non-goals
- Sharing failover state between instances that are not in one distributed
cluster.
- A bridge for backend methods that have no REST or realtime counterpart
upstream (audio encode/decode, metrics, status, fine-tune, quantization).
- Chains as router candidates or as the realtime classifier model.
---
## C. Distributed-aware failover
Standalone mode (no NATS, no PostgreSQL) keeps today's behaviour. Everything
below applies when distributed mode is on.
### Shared state
| State | Writers | Mechanism | Survives restart |
|---|---|---|---|
| Pins | any frontend (REST, MCP) | `syncstate.SyncedMap` named `failover.pins`, key = chain, with a gorm `Store` | yes |
| Target health: state, last error, since, consecutive passes | any frontend on a local transition, and the probe leader | `syncstate.SyncedMap` named `failover.targets`, key = target, NATS only, `Reconcile` on | no |
| Chain state: active target, active since, chain state | the probe leader only | `syncstate.SyncedMap` named `failover.chains`, key = chain, NATS only, `Reconcile` on | no |
- The pins table is created under `advisorylock.KeySchemaMigrate`, the same
way the jobs store creates its tables.
- The manager gets a small `StateSync` dependency. The standalone
implementation is a no-op; the distributed implementation wraps the three
maps. The manager does not import NATS or gorm directly.
- A peer delta is applied through `OnApply`, which changes local state
without publishing again (no echo loops).
### Who does what
- **Every frontend** plans requests from the shared state. `Plan` already
leaves unhealthy targets out of the attempt order, so a target tripped on
another frontend is skipped at once. In-request retry stays local.
- **Any frontend** that sees a real request trip or pass a target publishes
the new target state.
- **The probe leader** runs probes, recovery confirmation, dwell-based
fail-back, the chain recompute and the warm preload. It publishes chain
state. Leadership uses `advisorylock.RunLeaderLoop` with a new key
`failover-prober` and the same 1 s interval as the scheduler.
- **Followers** do not recompute the active target. They adopt the leader's
chain state. When no chain state has arrived yet (start-up), a follower
uses its own recompute until the first delta.
- If the leader stops, another frontend takes the lock on its next tick.
Pins and target health are not affected. Probes and fail-back pause for at
most one tick.
### Events
Each frontend emits `chain.switched` and `target.state` to its own
subscribers (SSE, realtime `localai.model.failover`) when it applies a
change, whether the change is local or from a peer. Every frontend's stream
therefore shows the same events.
### Warm targets
- The SmartRouter's and ReplicaReconciler's pinned-model resolver includes
`WarmTargets()`, so workers never evict a warm target.
- Only the leader preloads warm targets.
- A frontend cannot tell from its local model store whether a worker has a
model loaded. In distributed mode, warm-target liveness asks the node
registry whether a healthy node serves the model. When none does, the probe
is inconclusive: it neither passes nor trips, and real requests decide.
### Tests
- Unit: two managers on the test fakebus (`core/services/testutil`). A pin on
one shows on the other. A trip on one is skipped by the other's plan. Only
the lock holder probes. A new leader resumes fail-back.
- A spec in `tests/e2e/distributed` when its harness supports two frontends
cheaply; otherwise the unit specs are the coverage and the PR says so.
### Docs
`model-failover.md` replaces the "state is per instance" limit with a
"Distributed mode" section that describes the table above.
---
## A. `localai-proxy` backend
### Shape
- `backend/go/localai-proxy` is a separate OCI gallery backend, like
`cloud-proxy`. It is registered in the `Makefile`, `backend/index.yaml` and
`.github/backend-matrix.yml` (Linux amd64/arm64 and Darwin Metal), following
`.agents/adding-backends.md`.
- It reuses cloud-proxy's auth header, HTTP client (no redirects) and
hop-by-hop header helpers. It has no translate mode: the upstream is always
LocalAI.
- `Load` refuses a model without proxy options, so greedy backend probing
never selects it.
### Config
```yaml
name: argus-whisper
backend: localai-proxy
known_usecases: [transcript]
proxy:
upstream_url: https://argus:8080 # base URL; each method appends its path
upstream_model: whisper-large # optional; default: this model's name
api_key_env: ARGUS_KEY
request_timeout_seconds: 60 # applies to non-streaming calls
```
- `core/backend/options.go` passes `ProxyOptions` to `localai-proxy` as well
as `cloud-proxy`.
- `proxy.mode` and `proxy.provider` are ignored, with a load warning.
- A `localai-proxy` model without `known_usecases` loads with a warning, because
usecases decide default-model selection and the failover inference probe.
- The failover prober already treats `localai-proxy` as remote.
`UpstreamBase` accepts a base URL unchanged.
### Method mapping
| Backend method | Upstream endpoint |
|---|---|
| Predict, PredictStream | `/v1/chat/completions`, `/v1/completions` (SSE when streaming) |
| Embedding | `/v1/embeddings` |
| Rerank | `/v1/rerank` |
| TokenizeString, Detokenize | `/v1/tokenize`, `/v1/detokenize` |
| Score | `/api/score` |
| GenerateImage, UpscaleImage | `/v1/images/generations`, `/v1/images/upscale` |
| GenerateVideo | `/video` |
| Generate3D, Animate3D | `/3d/generations`, `/3d/animate` |
| TTS, TTSStream | `/tts` (streaming passes the upstream WAV header and PCM through) |
| SoundGeneration | `/v1/sound-generation` |
| AudioTranscription, AudioTranscriptionStream | `/v1/audio/transcriptions` (`stream=true` for SSE deltas) |
| AudioTranscriptionLive | upstream `/v1/realtime` transcription session (see below) |
| Diarize | `/v1/audio/diarization` |
| VAD | `/v1/vad` |
| SoundDetection | `/v1/audio/classification` |
| Detect, Depth | `/v1/detection`, `/v1/depth` |
| FaceVerify, FaceAnalyze | `/v1/face/verify`, `/v1/face/analyze` |
| VoiceVerify, VoiceAnalyze, VoiceEmbed | `/v1/voice/verify`, `/v1/voice/analyze`, `/v1/voice/embed` |
| Stores* | `/stores/set`, `/stores/get`, `/stores/delete`, `/stores/find` |
| AudioTransform | `/audio/transformations` |
Every request uses the upstream model name (`proxy.upstream_model`, else the
model name), the same derivation as `failover.UpstreamModel`.
Methods with no counterpart (AudioEncode, AudioDecode, AudioToAudioStream,
TokenClassify, GetMetrics, Status, ModelMetadata, fine-tune and quantization)
return gRPC `Unimplemented` with the message
`localai-proxy: <method> has no upstream counterpart`.
### Files
Core passes some inputs and outputs as local paths:
- Inputs (transcription and diarization audio, sound detection `src`, image
`src` and reference images): the proxy reads the file and uploads it as
multipart or base64, as the endpoint expects.
- Outputs (TTS, image, sound generation `dst`): the proxy writes the upstream
result (bytes, or a download of the returned URL, or decoded base64) to
`dst`.
### Live transcription bridge
`AudioTranscriptionLive` opens a WebSocket to the upstream `/v1/realtime`
transcription session:
1. On the first `TranscriptLiveConfig`, send the session update with the
upstream model, language and sample rate, then answer `ready`.
2. Forward each `TranscriptLiveAudio` as `input_audio_buffer.append`
(PCM float to PCM16 base64 at the session rate).
3. Map `conversation.item.input_audio_transcription.delta` to `delta`, and
`...completed` to `delta` (any remaining text) plus `eou: true`.
4. When the gRPC send side closes, commit the buffer, wait for the final
completion, send `final_result` and close.
5. An upstream error or disconnect ends the gRPC stream with `Unavailable`.
Word timings and `eob` are not available from the upstream and stay empty.
A realtime stage whose live session fails reopens on the next chain target at
the next utterance (behaviour from the base spec).
### Core changes
- **Rerank for Go backends.** `pkg/grpc` gets an optional rerank interface and
a server handler, in the same way as `Score`.
- **`Unimplemented` is a capability gap.** The failover retry path (HTTP and
`Manager.Do`) treats gRPC `Unimplemented` like an admission rejection: skip
to the next target for this request, and do not trip the target. Otherwise a
chain of a remote and a local target fails a request the local target can
serve.
### Tests
- Unit: a fake LocalAI `httptest` upstream per method family (request path,
body, model name, auth header, file upload and `dst` write).
- Unit: a fake WebSocket upstream for the live bridge (ready, deltas, eou,
final result, upstream disconnect).
- E2E: `localai-proxy` models that point back at the test server's own mock
models. A realtime pipeline whose stages are chains of a `localai-proxy`
target and a local target completes a turn, and switches stage when the
proxy target fails.
### Docs
A `localai-proxy` section in `docs/content/features/backends.md` (or the page
that documents `cloud-proxy`), and a remote-LocalAI example on
`model-failover.md`.
---
## B. WebUI
### Model editor
- A `failover-targets` field component replaces the JSON editor for
`failover.targets` (`core/config/meta/registry.go` switches the component
name). Each row has a model picker (`SearchableModelSelect`), move up/down,
remove and a **warm** toggle. The toggle is disabled with a tooltip on remote
targets. Inline validation: at least 2 targets, no duplicates, no chain as a
target.
- The probe, trip and recovery fields stay in the Advanced group.
- A **Failover chain** template in `modelTemplates.js`, seeded with two empty
targets. `?template=failover` preselects it.
### Health strip
When the edited model is a chain, a `FailoverChainStatus` component above the
form shows:
- the chain state (primary / fallback / degraded) and the active target, with
the time since it became active;
- for each target: state, kind, warm, last probe and last error;
- **Pin** and **Unpin** for admins, behind a confirm dialog. The pinned target
is marked.
### Model list and overview
- Installed models: a "chain" badge with the active target, in the same way
as the alias badge.
- Operate → Runtime → **Failover**: a dense table with one row per chain
(state, active target, target states, time since the last switch). Each row
links to the chain in the model editor. With no chains, an empty state links
to the Failover chain template.
### Live data
A `useFailoverChains` hook fetches `GET /api/failover`, then opens an
`EventSource` on `/api/failover/events`. `snapshot` replaces the state;
`chain.switched` and `target.state` patch it. The browser reconnects the
stream, and the hook polls every 15 s as a fallback (the Agent Status
pattern). A `failoverApi` group in `src/utils/api.js` holds the calls.
### Conventions
- Design tokens and CSS classes only; no new inline styles (inline-style
ratchet).
- `StatusPill` tones: success for healthy and primary, warning for recovering
and fallback, error for down and degraded, muted for missing.
- Strings in the `models` and `admin` i18n namespaces for all 8 locales.
- Pin controls are hidden when `useAuth().isAdmin` is false.
### Tests
Playwright specs with mocked APIs and a mocked `text/event-stream`: the editor
component, the template, the health strip updating on events, pin controls
for admins and not for other users, and the overview page. UI line coverage
stays at or above `core/http/react-ui/coverage-baseline.txt`.
---
## D. Contributor rule
- New guide `.agents/distributed-state.md`. A feature that keeps runtime
state (in-memory maps, caches, pins, schedulers, background loops, probes)
chooses one mode and documents it:
- **shared**: `syncstate.SyncedMap`, with a `Store` when the state must
survive a restart;
- **single-runner**: an `advisorylock` leader loop;
- **stateless per request**;
- **per-instance**: allowed only with the reason written in the feature's
docs.
The guide gives one real example per mode (finetune jobs, the node health
monitor, open responses, failover chains). Shared and single-runner
features include a fakebus test with two instances.
- `AGENTS.md`: a Quick Reference bullet "Distributed-aware state" and a row in
the Topics table.
- `.agents/api-endpoints-and-auth.md`: a checklist line "Stateful feature:
distributed mode chosen and documented (see distributed-state.md)".
## Order of work
C first (it changes code already in the PR and the event contract the UI
reads), then A (it needs the `Unimplemented` classification and the rerank
handler), then B, then D. All in PR #12285.