chore: drop design specs and plans from the change

Design specs and implementation plans are working notes and are not
kept in the tree.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
This commit is contained in:
Ettore Di Giacinto committed 2026-09-27 19:55:18 +00:00
1 parent 9e95ef0dbd
commit 9bbcde4b1b
11 files changed
-7456

No files matched your search

@@ -1,63 +0,0 @@
# Whisper-Medusa Backend Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Add a dedicated LocalAI speech-to-text backend for aiola Whisper-Medusa checkpoints.
**Architecture:** A Python gRPC backend owns model loading, audio normalization, and Medusa generation. LocalAI's existing `AudioTranscription` RPC remains unchanged; build, gallery, and documentation surfaces follow the existing Python ASR backend pattern.
**Tech Stack:** Python 3.11, PyTorch, torchaudio, transformers, whisper-medusa, gRPC, YAML, Make.
## Global Constraints
- Accept local model paths and Hugging Face model identifiers through `resolve_model_reference`.
- Normalize input audio to mono 16 kHz before generation.
- Default the language to `en` and expose upstream generation regulation options.
- Document upstream's archived status, 30-second clip limit, and checkpoint language limitations.
- Build Linux CPU and NVIDIA CUDA 12 images; do not claim unsupported Darwin or ROCm coverage.
---
### Task 1: Backend behavior
**Files:**
- Create: `backend/python/whisper-medusa/backend.py`
- Create: `backend/python/whisper-medusa/test_unit.py`
**Interfaces:**
- Consumes: LocalAI `LoadModel` and `AudioTranscription` protobuf requests.
- Produces: `BackendServicer`, `_parse_options`, and `_prepare_audio`.
- [ ] Write unit tests for option parsing, mono conversion, resampling, load failure, and transcription.
- [ ] Run `python -m unittest test_unit.py` and confirm it fails because the backend does not exist.
- [ ] Implement the minimal gRPC backend and rerun the unit tests to green.
### Task 2: Packaging and registration
**Files:**
- Create: `backend/python/whisper-medusa/{Makefile,install.sh,protogen.sh,run.sh,test.sh,requirements.txt,requirements-cpu.txt,requirements-cublas12.txt}`
- Modify: `Makefile`
- Modify: `.github/backend-matrix.yml`
- Modify: `backend/index.yaml`
**Interfaces:**
- Consumes: the Python backend Docker build conventions.
- Produces: `whisper-medusa` install/build targets and CPU/CUDA backend images.
- [ ] Add packaging scripts and pinned upstream dependency.
- [ ] Register the backend in Make and backend metadata.
- [ ] Add Linux amd64 CPU and CUDA 12 CI matrix entries.
- [ ] Validate Make and YAML parsing.
### Task 3: Documentation and verification
**Files:**
- Create: `docs/content/features/whisper-medusa.md`
- Modify: `docs/content/features/backends.md`
**Interfaces:**
- Produces: user-facing model YAML and limitation guidance.
- [ ] Document setup, configuration, options, and upstream constraints.
- [ ] Run unit tests, syntax compilation, registration checks, and diff review.
- [ ] Commit with the required `Assisted-by` trailer, push, and open a PR closing issue #3127.
@@ -1,689 +0,0 @@
# Failover: Distributed Mode, localai-proxy and WebUI Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Make failover chains consistent across distributed frontends, add a `localai-proxy` backend that serves every LocalAI API (including live transcription) from a remote LocalAI, add the chain UI, and add the distributed-state contributor rule — all in PR #12285.
**Architecture:** The failover `Manager` gets a `StateSync` dependency (no-op standalone; three `syncstate.SyncedMap`s in distributed mode) and a leader gate (PostgreSQL advisory lock per tick) so one frontend probes and owns chain state. `localai-proxy` is a new Go gRPC backend mapping each backend method to the upstream LocalAI REST endpoint, with a WebSocket bridge to the upstream realtime API for live transcription. The React UI adds an editor field, a template, a live health strip, a badge and an overview page on the existing failover REST/SSE API.
**Tech Stack:** Go (echo, gRPC, gorm, NATS via `syncstate`, gorilla/websocket), Ginkgo/Gomega, React 19 + Vite, Playwright.
**Spec:** `docs/superpowers/specs/2026-09-26-failover-distributed-proxy-ui-design.md` (builds on `docs/superpowers/specs/2026-09-26-model-failover-chains-design.md`)
## Global Constraints
- Worktree `/home/mudler/_git/LocalAI/.wt/failover-chains`, branch `feat/failover-chains`, PR #12285. Push only at the end (the controller does it).
- Commit trailer exactly `Assisted-by: Claude:claude-opus-5-5`. Never `Co-Authored-By` or `Signed-off-by`. `docs/superpowers/` needs `git add -f`.
- Build scope: `go build ./core/... ./pkg/... ./tests/... ./backend/go/localai-proxy/...` — never `go build ./...` (CGo launcher/backends need X11).
- Root `core/http` tests need `LOCALAI_TEST_HTTP_PORT=19391`.
- Logging `github.com/mudler/xlog`; `any` not `interface{}`; comments explain why.
- Standalone mode (no NATS/DB) must behave exactly as before this plan.
- SyncedMap names exactly: `failover.pins`, `failover.targets`, `failover.chains`. Leader republish interval 10 s. Advisory lock key constant `KeyFailoverProber` = 108.
- Backend name exactly `localai-proxy`; backend option `realtime_pipeline:<name>`; `Unimplemented` message format `localai-proxy: <method> has no upstream counterpart`.
- UI: no new inline `style={{}}` (inline-style ratchet); `StatusPill` tones: healthy/primary → success, recovering/fallback → warning, down/degraded → error, missing → muted; strings in i18n for all 8 locales (en, it, es, de, zh-CN, id, ko, pt-BR); pin controls only for `isAdmin`.
- Coverage baselines (`coverage-baseline.txt` 54.2, `core/http/react-ui/coverage-baseline.txt` 40.0) must not go down; never edit them.
## Review Focus
1. A NATS echo of a frontend's own publish must not deadlock or double-emit events (Task 1: "echo of own publish is a no-op").
2. A frontend that joins late (no deltas yet) must still serve requests from local state, then converge on the leader's republish (Task 2: "late joiner converges on heartbeat").
3. Leadership moving between frontends must not reset fail-back timing (shared `activeSince`) (Task 3: "new leader keeps activeSince").
4. A `localai-proxy` target that returns `Unimplemented` must move the request to the next chain target without tripping (Task 4: "Unimplemented skips without trip").
5. An upstream disconnect mid live-transcription must end the gRPC stream with `Unavailable`, not hang (Task 8: "upstream disconnect ends the stream").
---
## Part C — Distributed-aware failover
### Task 1: Manager state-sync hooks and leader gate
**Files:**
- Create: `core/services/failover/statesync.go`, `core/services/failover/statesync_test.go`
- Modify: `core/services/failover/manager.go`, `core/services/failover/schedule.go`
**Interfaces:**
- Consumes: existing `Manager` internals (`setTargetLocked`, `recomputeLocked`, `Pin`, `Unpin`, `emitLocked`, `chainState`, `targetState`, `takeWarmLocked`).
- Produces:
```go
type TargetSnapshot struct {
Target string `json:"target"`
State TargetState `json:"state"`
Reason Reason `json:"reason"`
Error string `json:"error,omitempty"`
ConsecutiveOK int `json:"consecutive_ok"`
Since time.Time `json:"since"`
}
type ChainSnapshot struct {
Chain string `json:"chain"`
Active string `json:"active"`
ActiveSince time.Time `json:"active_since"`
State ChainState `json:"state"`
Reason Reason `json:"reason"`
}
type StateSync interface {
PublishTarget(TargetSnapshot)
PublishChain(ChainSnapshot)
SetPin(chain, target string) error
ClearPin(chain string) error
Pins() map[string]string
}
type LeaderGate func(ctx context.Context, fn func()) bool
func WithLeaderGate(g LeaderGate) Option
func (m *Manager) SetStateSync(s StateSync)
func (m *Manager) ApplyTarget(s TargetSnapshot)
func (m *Manager) ApplyChain(s ChainSnapshot)
func (m *Manager) ApplyPin(chain, target string) // target "" = unpinned
func (m *Manager) IsLeader() bool
func (m *Manager) Republish() // leader: publish every target and chain snapshot
```
Design rules (binding):
- Publishing happens **outside `m.mu`**: a real or fake bus delivers the frontend's own publish back synchronously to `ApplyTarget`/`ApplyChain`/`ApplyPin`, which take `m.mu`. Queue publishes in `m.pending []func()` while locked; every exported mutating method drains and runs them after unlocking (helper `m.unlockAndFlush()`).
- `ApplyTarget` sets state through `setTargetLocked` with `m.applying = true`, so it emits the local `target.state` event but does not publish. An echo whose state equals the local state is a no-op (existing early return).
- Local transitions (`setTargetLocked` with `!m.applying` and `m.sync != nil`) queue `PublishTarget`.
- Chains: when `m.sync != nil && !m.leader`, `recomputeLocked` does not change `ch.active` once the chain has adopted leader state (`ch.adopted`); before adoption it computes locally. `ApplyChain` sets `active`, `activeSince`, `state`, `adopted = true` and emits `chain.switched` (with the snapshot's reason) when `active` or `state` changed. When the leader's recompute changes a chain, it queues `PublishChain`.
- Pins: `Pin`/`Unpin` set local state immediately (read-your-writes), then call `sync.SetPin`/`ClearPin` outside the lock. `ApplyPin` sets/clears `ch.pinned` and recomputes with `ReasonManual`. `SetStateSync` hydrates pins from `s.Pins()`.
- Leader gate: in `Tick`, after `Sync()`, call `gate(ctx, fn)` where `fn` runs probe scheduling + `Reevaluate()`; set `m.leader` to the returned bool. Without a gate (standalone) the manager is always leader and `Tick` behaves exactly as today. Followers still run `Reevaluate()` (local-only when not adopted). On a false→true leader transition set `m.warmPending = true` so `onWarm` fires on the new leader.
- `Republish()` queues `PublishTarget` for every target and `PublishChain` for every chain (leader only; no-op otherwise). `Run` calls it every 10 ticks when leader.
- [ ] **Step 1: Write the failing tests** (`statesync_test.go`, package `failover`, reuse `fakeClock`, `fakeSource`, `local`, `remote`, `chainCfg`, `t`, `errBoom`, `drain` from existing test files)
```go
package failover
import (
"context"
"sync"
"time"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
// loopSync is an in-process StateSync that delivers every publish to all
// managers synchronously, including the publisher (like NATS echo).
type loopSync struct {
mu sync.Mutex
peers []*Manager
pins map[string]string
}
func (l *loopSync) add(m *Manager) { l.mu.Lock(); l.peers = append(l.peers, m); l.mu.Unlock() }
func (l *loopSync) each(f func(*Manager)) {
l.mu.Lock()
ps := append([]*Manager(nil), l.peers...)
l.mu.Unlock()
for _, p := range ps {
f(p)
}
}
func (l *loopSync) PublishTarget(s TargetSnapshot) { l.each(func(m *Manager) { m.ApplyTarget(s) }) }
func (l *loopSync) PublishChain(s ChainSnapshot) { l.each(func(m *Manager) { m.ApplyChain(s) }) }
func (l *loopSync) SetPin(c, t string) error {
l.mu.Lock(); l.pins[c] = t; l.mu.Unlock()
l.each(func(m *Manager) { m.ApplyPin(c, t) })
return nil
}
func (l *loopSync) ClearPin(c string) error {
l.mu.Lock(); delete(l.pins, c); l.mu.Unlock()
l.each(func(m *Manager) { m.ApplyPin(c, "") })
return nil
}
func (l *loopSync) Pins() map[string]string {
l.mu.Lock(); defer l.mu.Unlock()
out := map[string]string{}
for k, v := range l.pins {
out[k] = v
}
return out
}
var _ = Describe("Manager state sync", func() {
var (
clock *fakeClock
src *fakeSource
bus *loopSync
a, b *Manager
leaderIsA bool
gateFor func(isA bool) LeaderGate
ctx = context.Background()
)
BeforeEach(func() {
clock = newFakeClock()
src = newFakeSource(remote("x"), local("y"), chainCfg("chain", nil, t("x"), t("y")))
bus = &loopSync{pins: map[string]string{}}
leaderIsA = true
gateFor = func(isA bool) LeaderGate {
return func(_ context.Context, fn func()) bool {
if isA != leaderIsA {
return false
}
fn()
return true
}
}
a = New(src, WithClock(clock), WithLeaderGate(gateFor(true)))
b = New(src, WithClock(clock), WithLeaderGate(gateFor(false)))
bus.add(a); bus.add(b)
a.SetStateSync(bus); b.SetStateSync(bus)
a.Tick(ctx); b.Tick(ctx)
})
It("echo of own publish is a no-op and emits one event", func() {
events, cancel := a.Subscribe(16)
defer cancel()
a.ReportFailure("x", errBoom)
n := 0
for _, e := range drain(events) {
if e.Type == EventTargetState && e.Target == "x" {
n++
}
}
Expect(n).To(Equal(1))
})
It("a trip on one frontend is skipped by the other's plan", func() {
b.ReportFailure("x", errBoom)
att, err := a.Plan("chain")
Expect(err).ToNot(HaveOccurred())
Expect(att.Target()).To(Equal("y"))
})
It("followers adopt the leader's chain state and emit the switch", func() {
events, cancel := b.Subscribe(16)
defer cancel()
a.ReportFailure("x", errBoom) // leader recomputes and publishes chain state
st, _ := b.ChainStatus("chain")
Expect(st.Active).To(Equal("y"))
var sw []Event
for _, e := range drain(events) {
if e.Type == EventChainSwitched {
sw = append(sw, e)
}
}
Expect(sw).ToNot(BeEmpty())
})
It("a pin on one frontend applies on all", func() {
Expect(b.Pin("chain", "y")).To(Succeed())
st, _ := a.ChainStatus("chain")
Expect(st.Pinned).ToNot(BeNil())
Expect(*st.Pinned).To(Equal("y"))
Expect(a.Unpin("chain")).To(Succeed())
st, _ = b.ChainStatus("chain")
Expect(st.Pinned).To(BeNil())
})
It("hydrates pins when the sync is attached", func() {
bus.pins["chain"] = "y"
c := New(src, WithClock(clock), WithLeaderGate(gateFor(false)))
c.SetStateSync(bus)
st, _ := c.ChainStatus("chain")
Expect(st.Pinned).ToNot(BeNil())
})
It("only the leader probes", func() {
pa, pb := &fakeProber{fail: map[string]error{}}, &fakeProber{fail: map[string]error{}}
a = New(src, WithClock(clock), WithProber(pa), WithLeaderGate(gateFor(true)))
b = New(src, WithClock(clock), WithProber(pb), WithLeaderGate(gateFor(false)))
a.SetStateSync(bus); b.SetStateSync(bus)
a.Tick(ctx); b.Tick(ctx)
Eventually(func() int { return len(pa.take()) }).Should(BeNumerically(">", 0))
Consistently(func() int { return len(pb.take()) }, 200*time.Millisecond).Should(Equal(0))
Expect(a.IsLeader()).To(BeTrue())
Expect(b.IsLeader()).To(BeFalse())
})
It("new leader keeps activeSince across a leadership move", func() {
a.ReportFailure("x", errBoom)
before, _ := b.ChainStatus("chain")
leaderIsA = false
clock.Advance(5 * time.Second)
a.Tick(ctx); b.Tick(ctx)
after, _ := b.ChainStatus("chain")
Expect(after.ActiveSince).To(Equal(before.ActiveSince))
Expect(b.IsLeader()).To(BeTrue())
})
It("republish sends every target and chain", func() {
c := New(src, WithClock(clock), WithLeaderGate(gateFor(false)))
bus.add(c)
c.SetStateSync(bus)
a.ReportFailure("x", errBoom)
a.Republish()
st, _ := c.ChainStatus("chain")
Expect(st.Active).To(Equal("y"))
})
It("standalone manager (no sync, no gate) is always leader", func() {
m := New(src, WithClock(clock))
m.Tick(ctx)
Expect(m.IsLeader()).To(BeTrue())
})
})
```
(`fakeProber` and its `take()` exist in `schedule_test.go`; if the in-flight probe tracking needs waiting, use `Eventually` as shown.)
- [ ] **Step 2: Run to verify failure**
Run: `go test ./core/services/failover/... 2>&1 | tail -5`
Expected: compile failure (`undefined: TargetSnapshot`).
- [ ] **Step 3: Implement** `statesync.go` (types, interface, `WithLeaderGate`, `SetStateSync`, `Apply*`, `IsLeader`, `Republish`, `unlockAndFlush`) and the hooks in `manager.go`/`schedule.go` following the design rules above. Add fields to `Manager`: `sync StateSync`, `gate LeaderGate`, `leader bool` (guarded by `mu`), `applying bool`, `pending []func()`, `ticks int`; to `chainState`: `adopted bool`. Keep every existing test green.
- [ ] **Step 4: Run tests**
Run: `go test -race -count=3 ./core/services/failover/... 2>&1 | tail -10`
Expected: PASS, no races.
- [ ] **Step 5: Commit**
```bash
git add core/services/failover
git commit -m "feat(failover): share state through a sync hook and gate probes on a leader
Assisted-by: Claude:claude-opus-5-5"
```
---
### Task 2: syncstate-backed StateSync with a durable pin store
**Files:**
- Create: `core/services/failover/distsync/distsync.go`, `core/services/failover/distsync/pinstore.go`, `core/services/failover/distsync/distsync_suite_test.go`, `core/services/failover/distsync/distsync_test.go`
**Interfaces:**
- Consumes: Task 1 `StateSync`, `TargetSnapshot`, `ChainSnapshot`, `Manager.ApplyTarget/ApplyChain/ApplyPin`; `syncstate.New/Config/Store`; `messaging.MessagingClient`; `advisorylock.WithLockCtx`, `advisorylock.KeySchemaMigrate`; `testutil.NewFakeBus`.
- Produces:
```go
type PinRecord struct {
Chain string `gorm:"primaryKey" json:"chain"`
Target string `json:"target"`
UpdatedAt time.Time `json:"updated_at"`
}
func (PinRecord) TableName() string { return "failover_pins" }
func NewPinStore(db *gorm.DB) (*PinStore, error) // migrates under KeySchemaMigrate
func New(ctx context.Context, nats messaging.MessagingClient, pins syncstate.Store[string, PinRecord], m *failover.Manager) (*Sync, error)
func (s *Sync) Close() error
// *Sync implements failover.StateSync
```
Rules:
- Three maps: `failover.pins` (key chain, `Store: pins` when non-nil — guard against a typed-nil interface as `finetune/service.go` does), `failover.targets` (key target, NATS only), `failover.chains` (key chain, NATS only). No `Reconcile` on NATS-only maps (it would re-hydrate them empty); the leader's `Republish` covers late joiners.
- `OnApply` for targets/chains calls `m.ApplyTarget`/`m.ApplyChain`; for pins `m.ApplyPin(chain, target)` on "set" and `m.ApplyPin(chain, "")` on "delete".
- `New` starts all three maps, then calls `m.SetStateSync(s)` (which hydrates pins from `s.Pins()`).
- [ ] **Step 1: Write the failing tests** — two managers on one `testutil.NewFakeBus()`, each with its own `distsync.New`, sharing an in-memory `syncstate.Store` fake for pins (a map with a mutex implementing `List/Upsert/Delete`). Specs:
- "a pin on A is visible on B and survives a new instance C built from the same store" (C's `ChainStatus` shows the pin right after `New`).
- "a trip on B makes A's plan skip the target".
- "the leader's chain switch reaches the follower".
- "late joiner converges on heartbeat": build C after A tripped a target; before `A.Republish()` C shows the primary active; after it, the fallback.
- "PinStore round-trips" using a sqlite gorm DB (`gorm.io/driver/sqlite`, `file::memory:`) if the repo already depends on it (check `go.mod`); otherwise skip this spec and test the adapter with the in-memory store only, and say so in the report.
- [ ] **Step 2: Run to verify failure** — `go test ./core/services/failover/distsync/...` → compile failure.
- [ ] **Step 3: Implement** `distsync.go` and `pinstore.go` per the rules (PinStore: `List` = `db.Find`, `Upsert` = `db.Save`, `Delete` = `db.Delete(&PinRecord{Chain: k})`; migration via `advisorylock.WithLockCtx(context.Background(), db, advisorylock.KeySchemaMigrate, func() error { return db.AutoMigrate(&PinRecord{}) })`).
- [ ] **Step 4: Run tests** — `go test -race ./core/services/failover/...` → PASS.
- [ ] **Step 5: Commit**
```bash
git add core/services/failover/distsync
git commit -m "feat(failover): sync pins, target health and chain state over NATS
Assisted-by: Claude:claude-opus-5-5"
```
---
### Task 3: Distributed wiring, warm pins on workers, docs
**Files:**
- Modify: `core/services/advisorylock/keys.go` (add `KeyFailoverProber = 108` with a comment)
- Modify: `core/application/startup.go` (after `application.distributed = distSvc`, ~L294, and before `Run` ~L570), `core/application/distributed.go` (~L380-395, L458: pinned resolver), `core/application/failover.go` (preload only when leader)
- Create: `core/application/failover_distributed.go`, `core/application/failover_distributed_test.go`
- Modify: `docs/content/features/model-failover.md` (replace the per-instance limit with a "Distributed mode" section)
**Interfaces:**
- Consumes: Tasks 1-2; `advisorylock.TryWithLockCtx(ctx, db, key, fn func() error) (bool, error)`; `a.IsDistributed()`, `a.Distributed().Nats`, `a.distributedDB()`; `nodes.PinnedModelResolver` (`GetPinnedModelNames() []string`); `failover.MergePinned`.
- Produces: `func failoverLeaderGate(db *gorm.DB) failover.LeaderGate`; `type failoverPinnedResolver struct{ base nodes.PinnedModelResolver; fm *failover.Manager }` implementing `GetPinnedModelNames()`.
Rules:
- Leader gate: `func(ctx, fn) bool { ok, err := advisorylock.TryWithLockCtx(ctx, db, advisorylock.KeyFailoverProber, func() error { fn(); return nil }); if err != nil { xlog.Warn(...); return false }; return ok }`. The failover manager is constructed before distributed init (startup.go ~L257), so give it the gate through a setter `(*Manager).SetLeaderGate(LeaderGate)` added in this task (mirror of `WithLeaderGate`), then call `distsync.New(ctx, distSvc.Nats, pinStore, application.failoverManager)`; on error log and continue standalone.
- Pinned resolver: wrap `configLoader` so `GetPinnedModelNames()` returns `failover.MergePinned(configLoader.GetPinnedModelNames(), fm.WarmTargets())`, and pass it to both the SmartRouter and the ReplicaReconciler options. Adapt `initDistributed`'s signature minimally (add a `pinned nodes.PinnedModelResolver` parameter or set it after construction — choose the smaller change and explain).
- Warm preload: `applyFailoverWarmTargets` keeps the pin sync, and runs the preload goroutine only when `!a.IsDistributed() || a.failoverManager.IsLeader()`.
- `Run` calls `Republish()` every 10 ticks when leader (implemented in Task 1; verify here).
- [ ] **Step 1: Write the failing tests** (`failover_distributed_test.go`, in the application package's existing test suite):
- `failoverPinnedResolver` merges config pins and warm targets without duplicates.
- `failoverLeaderGate` with a gorm DB that is not PostgreSQL (the advisorylock package falls back to an in-process lock for non-Postgres DBs): two gates on the same DB — while one is inside `fn`, the other returns false; afterwards the other returns true. Use the same DB setup other `core/application` or `advisorylock` tests use (read `core/services/advisorylock/*_test.go` first).
- [ ] **Step 2: Run to verify failure.**
- [ ] **Step 3: Implement** per the rules; add the `keys.go` constant; update docs: in `model-failover.md`, replace the "Failover state is kept in memory by each LocalAI instance" limit with a `## Distributed mode` section: pins are cluster-wide and persisted; target health and chain state are shared; one frontend (the probe leader) probes, fails back and preloads warm targets; warm targets are pinned on workers; a frontend that joins late converges within 10 s.
- [ ] **Step 4: Run tests** — `go test -race ./core/application/... ./core/services/failover/... ./core/services/advisorylock/...` and `go build ./core/... ./pkg/... ./tests/...` → PASS.
- [ ] **Step 5: Commit**
```bash
git add core/services/advisorylock core/application core/services/failover docs/content/features/model-failover.md
git commit -m "feat(failover): run one prober per cluster and pin warm targets on workers
Assisted-by: Claude:claude-opus-5-5"
```
---
## Part A — localai-proxy backend
### Task 4: Core support — Rerank for Go backends, Unimplemented skip, proxy options
**Files:**
- Modify: `pkg/grpc/interface.go` (add `RerankModel`), `pkg/grpc/server.go` (add `Rerank` handler), `pkg/grpc/model_identity_modalities_test.go` (update the note/test that says Go servers have no Rerank)
- Modify: `core/services/failover/classify.go` (add `IsCapabilityGap`), `core/services/failover/manager.go` (`Do` uses it), `core/http/middleware/failover.go` (retry uses it)
- Modify: `core/backend/options.go:545` (proxy options for `localai-proxy`), `core/config/model_config_loader.go` (load-time warnings)
- Test: `pkg/grpc/server_rerank_test.go` (or the package's existing server test file), `core/services/failover/classify_test.go`, `core/services/failover/manager_test.go`, `core/http/middleware/failover_test.go`, `core/config/model_config_loader_test.go`
**Interfaces:**
- Produces:
```go
// pkg/grpc/interface.go
type RerankModel interface {
Rerank(context.Context, *pb.RerankRequest) (*pb.RerankResult, error)
}
// core/services/failover/classify.go
func IsCapabilityGap(err error) bool // gRPC Unimplemented anywhere in the chain
```
- [ ] **Step 1: Write the failing tests**
- `pkg/grpc`: a fake model embedding `base.Base` and implementing `RerankModel` is served by `server.Rerank` (use the package's existing in-process server test pattern for `Score`); a model without it returns `codes.Unimplemented`.
- `classify_test.go`: `IsCapabilityGap(grpcstatus.Error(codes.Unimplemented, "x"))` true; wrapped with `fmt.Errorf("%w")` true; `codes.Unavailable` false; nil false.
- `manager_test.go` ("Unimplemented skips without trip"): `m.Do` where target `a` returns `grpcstatus.Error(codes.Unimplemented, "localai-proxy: X has no upstream counterpart")` and `b` succeeds → `tried == [a b]`, `a` stays `StateHealthy`.
- `failover_test.go` (middleware): handler for `a` returns the same Unimplemented error → served by `b`, `a` healthy.
- `model_config_loader_test.go`: a `backend: localai-proxy` config with `proxy.mode: translate` logs a warning and still loads; one without `known_usecases` loads (warning). Use the loader test's existing log-capture approach if any; otherwise assert only that loading succeeds and note it.
- [ ] **Step 2: Run to verify failure.**
- [ ] **Step 3: Implement**
```go
// pkg/grpc/server.go — copy of the Score handler shape
func (s *server) Rerank(ctx context.Context, in *pb.RerankRequest) (*pb.RerankResult, error) {
if err := s.checkModelIdentity(in); err != nil {
return nil, err
}
rm, ok := s.llm.(RerankModel)
if !ok {
return nil, status.Errorf(codes.Unimplemented, "method Rerank not implemented")
}
if s.llm.Locking() {
s.llm.Lock()
defer s.llm.Unlock()
}
return rm.Rerank(ctx, in)
}
```
(If `checkModelIdentity` does not accept `*pb.RerankRequest`, add the `GetModelIdentity()` method set it needs — `RerankRequest` has field `ModelIdentity`.)
```go
// classify.go
// IsCapabilityGap reports a target that cannot serve this kind of request at
// all. The next target may serve it, and this target is not broken.
func IsCapabilityGap(err error) bool {
if err == nil {
return false
}
st, ok := grpcstatus.FromError(err)
return ok && st.Code() == codes.Unimplemented
}
```
In `Manager.Do`, before the retryable check: `if IsCapabilityGap(err) && !committed.Load() { if !att.Skip() { return err }; continue }`. In `failoverRetry`, treat `failover.IsCapabilityGap(err)` exactly like the admission-rejection branch (spill with `att.Skip()`, no trip). In `options.go` change the condition to `if c.Backend == "cloud-proxy" || c.Backend == "localai-proxy"`. In the loader's load-time pass, `xlog.Warn` for `localai-proxy` configs that set `proxy.mode`/`proxy.provider` (ignored) or have no `known_usecases`.
- [ ] **Step 4: Run tests** — `go test -race ./pkg/grpc/... ./core/services/failover/... ./core/http/middleware/... ./core/config/...` → PASS.
- [ ] **Step 5: Commit**
```bash
git add pkg/grpc core/services/failover core/http/middleware core/backend/options.go core/config
git commit -m "feat(grpc): serve Rerank from Go backends and skip Unimplemented targets
Assisted-by: Claude:claude-opus-5-5"
```
---
### Task 5: localai-proxy backend — skeleton, packaging, text methods
**Files:**
- Create: `backend/go/localai-proxy/{main.go,proxy.go,client.go,text.go,Makefile,package.sh,run.sh}`, `backend/go/localai-proxy/{localai_proxy_suite_test.go,fake_upstream_test.go,text_test.go}`
- Modify: root `Makefile` (6 places, mirroring cloud-proxy: `.NOTPARALLEL`, `TEST_PATHS`, `BACKEND_LOCALAI_PROXY = localai-proxy|golang|.|false|true`, `$(eval $(call generate-docker-build-target,$(BACKEND_LOCALAI_PROXY)))`, `docker-build-localai-proxy` in `docker-build-backends`, `build-localai-proxy-backend`/`clean-localai-proxy-backend` e2e helpers building to `tests/e2e/mock-backend/localai-proxy`), `backend/index.yaml` (meta + cpu/metal latest/development images, mirroring cloud-proxy), `.github/backend-matrix.yml` (amd64, arm64, darwin entries mirroring cloud-proxy), `.gitignore` (the e2e binary)
**Interfaces:**
- Consumes: `pkg/grpc` (`StartServer`, `AIModelRich`, `RerankModel`, `ScoreModel`), `pkg/grpc/base.Base`, `pkg/httpclient.New`.
- Produces: `type LocalAIProxy struct { base.Base; cfg atomic.Pointer[proxyConfig]; client *http.Client }`, `func NewLocalAIProxy() *LocalAIProxy`, `type proxyConfig struct { base, upstreamModel, apiKey, realtimePipeline string; timeout time.Duration }`, helpers `func (p *LocalAIProxy) postJSON(ctx, path string, body, out any) error`, `func (p *LocalAIProxy) postMultipart(ctx, path string, fields map[string]string, fileField, filePath string, out any) error`, `func (p *LocalAIProxy) postStream(ctx, path string, body any) (*http.Response, error)`, `func (p *LocalAIProxy) model(req string) string` (upstream model name), `func unimplemented(method string) error` returning `status.Errorf(codes.Unimplemented, "localai-proxy: %s has no upstream counterpart", method)`.
Rules:
- `Load`: requires `opts.GetProxy()` non-nil and a valid `upstream_url` (base URL; strip a trailing `/`); resolves the API key with the same rules as cloud-proxy's `resolveAPIKey` (copy the function); warns if `mode`/`provider` set; reads `realtime_pipeline:<name>` from `opts.GetOptions()`; `request_timeout_seconds` becomes a per-request context timeout for non-streaming calls; model name = `upstream_model`, else `opts.GetModel()`.
- Auth: `Authorization: Bearer <key>` when a key is set. HTTP client from `httpclient.New()` (no redirects).
- Upstream non-2xx → a gRPC error: 5xx/transport → `codes.Unavailable`; 4xx → `codes.InvalidArgument` (so failover does not trip on client errors); body text in the message (truncated to 500 chars).
- Text methods in `text.go`: `PredictRich`/`PredictStreamRich` via `/v1/chat/completions` when `opts.GetMessages()` is non-empty, else `/v1/completions` with `opts.GetPrompt()` (map tokens, temperature, top_p, top_k, stop, seed; stream parses SSE `data:` lines and sends `pb.Reply{Message}` per delta; do not close the channel); legacy `Predict`/`PredictStream` wrap them; `Embeddings` → `/v1/embeddings` (`input` = `opts.GetEmbeddings()`, returns `data[0].embedding`); `Rerank` → `/v1/rerank`; `TokenizeString` → `/v1/tokenize`; `Score` → `/api/score` (read `core/http/endpoints` for the request shape).
- Methods with no counterpart return `unimplemented("<Method>")`: `AudioEncode`, `AudioDecode`, `AudioToAudioStream`, `TokenClassify`, `ModelMetadata`, fine-tune and quantization methods. `Status` keeps the base implementation.
- [ ] **Step 1: Write the failing tests** — `fake_upstream_test.go`: an `httptest.Server` recording method, path, `Authorization`, and JSON/multipart body, with per-path scripted responses (JSON or SSE). `text_test.go` specs: Load rejects missing proxy options and a bad URL; Load parses `realtime_pipeline`; `PredictRich` hits `/v1/chat/completions` with the upstream model and returns the content; `PredictStreamRich` streams SSE deltas in order; `Embeddings`, `Rerank`, `TokenizeString` hit their paths and map results; a 503 upstream → `codes.Unavailable`; a 400 → `codes.InvalidArgument`; `AudioEncode` → `Unimplemented` with the exact message.
- [ ] **Step 2: Run to verify failure** — `go test ./backend/go/localai-proxy/...`.
- [ ] **Step 3: Implement** per the rules; packaging per the Files list (copy cloud-proxy's `Makefile`/`package.sh`/`run.sh` with the binary renamed).
- [ ] **Step 4: Run tests** — `go test -race ./backend/go/localai-proxy/...`; `make -C backend/go/localai-proxy build`; `make build-localai-proxy-backend` → OK. Validate YAML: `python3 -c "import yaml,sys; yaml.safe_load(open('backend/index.yaml')); yaml.safe_load(open('.github/backend-matrix.yml'))"`.
- [ ] **Step 5: Commit**
```bash
git add backend/go/localai-proxy Makefile backend/index.yaml .github/backend-matrix.yml .gitignore
git commit -m "feat(localai-proxy): add a backend that serves text APIs from a remote LocalAI
Assisted-by: Claude:claude-opus-5-5"
```
---
### Task 6: localai-proxy — audio methods
**Files:**
- Create: `backend/go/localai-proxy/audio.go`, `backend/go/localai-proxy/audio_test.go`
**Interfaces:** consumes Task 5 helpers.
Mapping (request → upstream → result):
| Method | Upstream | Notes |
|---|---|---|
| `TTS(req)` | `POST /tts` JSON `{model, input: req.Text, voice, language}` | write the response bytes to `req.Dst` |
| `TTSStream(req, out)` | `POST /tts` with `stream: true` | copy the chunked `audio/wav` body to `out` as it arrives (header + PCM, unchanged); close `out` per the base contract |
| `SoundGeneration(req)` | `POST /v1/sound-generation` | write bytes to `req.Dst` |
| `AudioTranscription(ctx, req)` | `POST /v1/audio/transcriptions` multipart: `file` from `req.Dst` (input audio path), `model`, `language`, `translate`, `prompt`, `diarize` | map `TranscriptionResult{text, segments, words, language, duration}` to `pb.TranscriptResult` |
| `AudioTranscriptionStream(ctx, req, out)` | same, plus `stream=true` | SSE `transcript.text.delta` → `TranscriptStreamResponse{Delta}`; `transcript.text.done` → `FinalResult`; `error` event → return `codes.Unavailable` |
| `Diarize` | `POST /v1/audio/diarization` multipart | map segments |
| `VAD(req)` | `POST /v1/vad` JSON `{model, audio: req.Audio}` | map `segments[{start,end}]` |
| `SoundDetection(ctx, req)` | `POST /v1/audio/classification` multipart `file` from `req.Src`, `top_k`, `threshold` | map `detections[{index,label,score}]` |
| `AudioTransform` | `POST /audio/transformations` | read the handler for the request shape; write output to the path the request carries |
Read each `pb` request/response message in `backend/backend.proto` and each REST schema in `core/schema` before mapping; keep field names exact.
- [ ] **Step 1: Write the failing tests** — one spec per row using the fake upstream: path, multipart fields (file bytes equal the input file), JSON body, and result mapping; `TTS` writes the upstream bytes to `Dst`; `TTSStream` forwards chunks in order and the first chunk starts with `RIFF`; `AudioTranscriptionStream` emits deltas then the final result; upstream disconnect mid-stream → `codes.Unavailable`.
- [ ] **Step 2: Verify failure.** **Step 3: Implement.** **Step 4:** `go test -race ./backend/go/localai-proxy/...` → PASS.
- [ ] **Step 5: Commit** — `feat(localai-proxy): serve speech, transcription and audio APIs remotely`.
---
### Task 7: localai-proxy — image, video, 3D and vision methods
**Files:**
- Create: `backend/go/localai-proxy/media.go`, `backend/go/localai-proxy/media_test.go`
Mapping:
| Method | Upstream | Notes |
|---|---|---|
| `GenerateImage(req)` | `POST /v1/images/generations` `{model, prompt, negative_prompt, size: "WxH", step, seed, response_format: "b64_json"}`; `req.Src`/`ref_images` sent as base64 in `files`/`ref_images` per `core/schema` | decode `data[0].b64_json` into `req.Dst` |
| `UpscaleImage` | `POST /v1/images/upscale` | same output handling |
| `GenerateVideo` | `POST /video` | write the returned file (b64 or URL download relative to the upstream base) to `Dst` |
| `Generate3D`, `Animate3D` | `POST /3d/generations`, `/3d/animate` | same output handling; implement `AnimationMetadataModel` only if the upstream response carries the metadata |
| `Detect`, `Depth` | `POST /v1/detection`, `/v1/depth` | map results |
| `FaceVerify`, `FaceAnalyze` | `POST /v1/face/verify`, `/v1/face/analyze` | map results |
| `VoiceVerify`, `VoiceAnalyze`, `VoiceEmbed` | `POST /v1/voice/verify`, `/v1/voice/analyze`, `/v1/voice/embed` | map results |
| `StoresSet/Get/Delete/Find` | `POST /stores/set`, `/stores/get`, `/stores/delete`, `/stores/find` | map keys/values |
When the upstream returns a URL instead of b64, download it with the same client (same auth) and write it to `Dst`.
- [ ] **Step 1: Failing tests** — one spec per row with the fake upstream (path, key request fields, `Dst` written from b64 and from a URL). **Step 2** verify failure. **Step 3** implement. **Step 4** `go test -race ./backend/go/localai-proxy/...` → PASS.
- [ ] **Step 5: Commit** — `feat(localai-proxy): serve image, video, 3D and vision APIs remotely`.
---
### Task 8: localai-proxy — live transcription bridge
**Files:**
- Create: `backend/go/localai-proxy/live.go`, `backend/go/localai-proxy/live_test.go`
**Interfaces:** `func (p *LocalAIProxy) AudioTranscriptionLive(in <-chan *pb.TranscriptLiveRequest, out chan<- *pb.TranscriptLiveResponse) error`. Contract (from `pkg/grpc/server.go:477-541`): the backend closes `out`; `in` closes on client EOF; return errors immediately (callers wait for the ready ack).
Protocol (upstream `github.com/gorilla/websocket`):
1. No `realtime_pipeline` → `close(out); return grpcerrors.LiveTranscriptionUnsupported("localai-proxy", "set the realtime_pipeline backend option")` (same helper `base.Base` uses).
2. Read the first `in` message; it must be `config` (else `codes.InvalidArgument`). Rate = `config.sample_rate` or 16000.
3. Dial `ws(s)://<base>/v1/realtime?model=<realtime_pipeline>` with the bearer header. Read `session.created`.
4. Send `{"type":"session.update","session":{"type":"transcription","audio":{"input":{"format":{"type":"audio/pcm","rate":<rate>},"transcription":{"model":"<realtime_pipeline>","language":"<lang>"},"turn_detection":{"type":"server_vad"}}}}}`. On `session.updated` send `TranscriptLiveResponse{Ready: true}`; on `error` return `codes.Unavailable` with its message.
5. Writer goroutine: each `audio.pcm` (float32 in [-1,1]) → PCM16 LE → base64 → `{"type":"input_audio_buffer.append","audio":"..."}`.
6. Reader goroutine: `conversation.item.input_audio_transcription.delta` → `{Delta}`; `...completed` → `{Delta: <transcript minus text already sent for this item>, Eou: true}` and append to the running final text; `...failed` or `error` → end with `codes.Unavailable`.
7. When `in` closes: wait up to 5 s for an in-flight `completed` (track `speech_started` without a matching completion), then send `{FinalResult: {Text: <all completed text>}}`, close the socket, close `out`, return nil.
8. Socket read error before step 7 → close `out`, return `status.Error(codes.Unavailable, ...)`.
- [ ] **Step 1: Failing tests** with an `httptest` WebSocket server (gorilla `Upgrader`) scripting the upstream: unsupported without the option; ready after `session.updated`; audio frames arrive base64-PCM16 of the right length; deltas and a completion map to `Delta`/`Eou`; closing `in` yields `FinalResult` with the concatenated text; "upstream disconnect ends the stream" (server closes mid-session → `Unavailable`, `out` closed, no goroutine left blocked — assert the call returns within 2 s).
- [ ] **Step 2** verify failure. **Step 3** implement. **Step 4** `go test -race ./backend/go/localai-proxy/...` → PASS.
- [ ] **Step 5: Commit** — `feat(localai-proxy): bridge live transcription to the upstream realtime API`.
---
### Task 9: localai-proxy end to end, and docs
**Files:**
- Modify: `tests/e2e/e2e_suite_test.go` (build/locate the `localai-proxy` binary like cloud-proxy), create `tests/e2e/e2e_localai_proxy_test.go`, extend `tests/e2e/realtime_ws_test.go` (label `failover`)
- Modify: docs — the page that documents `cloud-proxy` (find with `grep -rln "cloud-proxy" docs/content`) gets a `localai-proxy` section; `docs/content/features/model-failover.md` gets a remote-LocalAI example with per-stage chains.
Rules:
- Point `localai-proxy` models at the test server itself (`proxy.upstream_url` = the suite's base URL, `upstream_model` = an existing mock model), registered at runtime the way the cloud-proxy/failover e2e specs register models.
- Specs: chat, embeddings, TTS and transcription through `localai-proxy` return 2xx with the upstream model's answer; a chain `[proxy-target, local mock]` where the proxy target's upstream model is `fail-load-…` fails over to the local target; a realtime pipeline whose `llm` stage is a chain `[localai-proxy → mock LLM, mock LLM]` completes a turn and, after the proxy target is made to fail (point it at a `fail-load` upstream model or stop routing), switches stage with a `localai.model.failover` event. Live transcription through the bridge is covered by Task 8's unit specs; add an e2e only if the suite has a pipeline with streaming transcription available.
- Run `make build-localai-proxy-backend build-mock-backend` first.
- [ ] Steps: write specs → run (fail) → implement registration + docs → run `go run github.com/onsi/ginkgo/v2/ginkgo --label-filter=failover -v ./tests/e2e` and the full `!real-models` e2e → commit `test(localai-proxy): proxy APIs and realtime stages end to end`.
---
## Part B — WebUI
### Task 10: Failover data layer and health strip
**Files:**
- Modify: `core/http/react-ui/src/utils/config.js` (endpoints), `src/utils/api.js` (`failoverApi`), `src/components/StatusPill.jsx` (tones)
- Create: `src/hooks/useFailoverChains.js`, `src/components/FailoverChainStatus.jsx`
- Modify: `src/pages/ModelEditor.jsx` (strip after the `me-head` block, before the template selector, when `!isCreateMode` and the chain exists), `src/App.css` (classes), `public/locales/*/models.json` (strings, 8 locales)
- Test: `e2e/failover-health.spec.js`
**Interfaces — Produces:**
```js
// utils/config.js endpoints
failoverChains: '/api/failover',
failoverChain: (name) => `/api/failover/${encodeURIComponent(name)}`,
failoverEvents: '/api/failover/events',
failoverPin: (name) => `/api/failover/${encodeURIComponent(name)}/pin`,
// utils/api.js
export const failoverApi = {
list: () => fetchJSON(API_CONFIG.endpoints.failoverChains),
get: (name) => fetchJSON(API_CONFIG.endpoints.failoverChain(name)),
pin: (name, target) => postJSON(API_CONFIG.endpoints.failoverPin(name), { target }),
unpin: (name) => fetchJSON(API_CONFIG.endpoints.failoverPin(name), { method: 'DELETE' }),
eventsUrl: () => API_CONFIG.endpoints.failoverEvents,
}
// hooks/useFailoverChains.js
export default function useFailoverChains() // → { chains: ChainStatus[], byName: {[name]: ChainStatus}, loading, error, refresh }
```
Hook rules: initial `failoverApi.list()`; `new EventSource(apiUrl(failoverApi.eventsUrl()))`; `snapshot` → replace `chains` from `data.chains`; `chain.switched` → patch that chain's `active`, `state`, `active_since = data.at`; `target.state` → patch that target's `state`, `last_error = data.error` in every chain containing it; `onerror` no-op; poll `list()` every 15 s; close on unmount.
`StatusPill` STATUS map additions: `primary: 'success'`, `fallback: 'warning'`, `recovering: 'warning'`, `degraded: 'error'`, `down: 'error'`, `missing: 'muted'` (`healthy` already maps to success).
`FailoverChainStatus({ chain, onPin, onUnpin, canPin })` renders: chain `StatusPill` + active target + relative "since"; a compact table of targets (model, kind, warm, `StatusPill` state, last probe, last error truncated with title); per-target **Pin** button and a chain-level **Unpin** when `canPin`, each behind `ConfirmDialog`. Classes only.
- [ ] **Step 1: Failing Playwright spec** `e2e/failover-health.spec.js` (import `test` from `./coverage-fixtures.js`; mock `**/api/auth/status`, config metadata, `**/api/failover` with one chain `chain` [a healthy active, b healthy], `**/api/failover/chain`, and `**/api/failover/events` fulfilled with `text/event-stream` body `event: snapshot\ndata: {"chains":[...]}\n\nevent: chain.switched\ndata: {"chain":"chain","from":"a","to":"b","state":"fallback","reason":"trip","at":"2026-09-26T10:00:00Z"}\n\n`). Assertions: `/app/model-editor/chain` shows the strip with the chain state; after the event the active target is `b` and the pill reads fallback; the Pin button is visible with auth disabled (admin) and hidden when auth status reports a non-admin user (mock `/api/auth/me` the way `users-tab-gating.spec.js` does); clicking Pin + confirm POSTs `{target}`.
- [ ] **Step 2** run `cd core/http/react-ui && npx playwright test e2e/failover-health.spec.js` (after `bun run build` if the harness serves the build; follow the repo's UI test instructions in `Makefile` `test-ui*` targets) → FAIL.
- [ ] **Step 3** implement; **Step 4** re-run → PASS; `npm run lint` and `npm run lint:inline-styles` (or the scripts in `package.json`) clean.
- [ ] **Step 5: Commit** — `feat(ui): show live failover chain health in the model editor`.
---
### Task 11: Chain editor field and template
**Files:**
- Create: `core/http/react-ui/src/components/FailoverTargetsEditor.jsx`
- Modify: `src/components/ConfigFieldRenderer.jsx` (branch `component === 'failover-targets'`, same `list-row` wrapper as `router-candidates`), `src/utils/modelTemplates.js` (template), `src/pages/ModelEditor.jsx` (`SECTION_ICONS.failover = 'fa-shuffle'`, `SECTION_COLORS.failover = 'var(--color-accent)'`), `core/config/meta/registry.go` (`failover.targets` `Component: "failover-targets"`), `core/config/meta/registry_test.go` (assert the component), `public/locales/*/modelEditor.json`
- Test: `e2e/failover-editor.spec.js`
`FailoverTargetsEditor({ value, onChange })`: modelled on `RouterCandidatesEditor` — items `{model, warm}`; row = `SearchableModelSelect` (value/onChange), `Toggle` for warm (disabled with a title when the selected model's backend is a proxy — look it up from `useModels()` data if it carries the backend; otherwise leave enabled and rely on the load warning, and say so), move up/down, remove; "Add target" button; inline errors (fewer than 2 targets, duplicate model, the edited model's own name) from `useFormContext()` `formData.name`.
Template entry:
```js
{
id: 'failover',
label: 'Failover Chain',
icon: 'fa-shuffle',
description: 'Serve one model name from an ordered list of models. The first healthy one answers; the next takes over when it fails.',
fields: {
'name': '',
'failover.targets': [{ model: '' }, { model: '' }],
},
},
```
- [ ] Steps: failing spec (template card visible; `?template=failover` shows two target rows; adding/removing/moving rows; duplicate and too-few errors; saving sends `failover.targets` in the PATCH/import body — mock the save endpoint and assert the JSON) → implement → `npx playwright test e2e/failover-editor.spec.js` PASS, `go test ./core/config/meta/...` PASS → commit `feat(ui): edit failover chain targets with a dedicated field`.
---
### Task 12: Chain badge and Failover overview page
**Files:**
- Create: `core/http/react-ui/src/pages/Failover.jsx`
- Modify: `src/pages/InstalledModels.jsx` (chain badge next to the alias badge: `badge badge-info`, icon `fa-shuffle`, text `chain → <active>`; data from `useFailoverChains()`), `src/router.jsx` (`const Failover = page('failover', () => import('./pages/Failover'))`; route `{ path: 'failover', element: <Admin><Failover /></Admin> }`), `src/components/console/consoleConfig.js` (`operate.runtime` item `{ path: '/app/failover', icon: 'fas fa-shuffle', labelKey: 'items.failover', adminOnly: true }`), `public/locales/*/nav.json`, `public/locales/*/models.json`, `src/App.css`
- Test: `e2e/failover-overview.spec.js`
Overview: `useFailoverChains()`; a dense table (chain name link → `/app/model-editor/<name>`, chain `StatusPill`, active target, one small pill per target, time since `active_since`); empty state with a link to `/app/model-editor?template=failover`.
- [ ] Steps: failing spec (nav entry visible for admin; table rows from mocked `/api/failover`; SSE patch updates a row; empty state link; installed-models badge `chain → a`) → implement → specs PASS; run the full UI suite with coverage: `make test-ui-coverage-check` (UI coverage ≥ baseline) → commit `feat(ui): list failover chains and badge chain models`.
---
## Part D — Contributor rule and final verification
### Task 13: distributed-state rule, docs sweep, final verification
**Files:**
- Create: `.agents/distributed-state.md`
- Modify: `AGENTS.md` (Topics table row; Quick Reference bullet), `.agents/api-endpoints-and-auth.md` (checklist line), `docs/content/features/model-failover.md` (UI section: editor, health strip, overview; confirm distributed section from Task 3)
`.agents/distributed-state.md` content (write it in full, following the style of the other `.agents/*.md` guides):
- Title "Distributed-aware state". Why: frontends are stateless replicas; in-memory state diverges silently (the failover chains example).
- The rule: a feature that keeps runtime state (in-memory maps, caches, pins, schedulers, background loops, probes) chooses one mode and documents it on its docs page:
1. **Shared** — `syncstate.SyncedMap` (`core/services/syncstate`); add a `Store` when the state must survive a restart. Example: finetune jobs (`core/services/finetune/service.go`), failover pins (`core/services/failover/distsync`). Gotcha: `Reconcile` without a `Store`/`Loader` re-hydrates the map empty — republish from a leader instead.
2. **Single-runner** — `advisorylock.RunLeaderLoop` / `TryWithLockCtx` (`core/services/advisorylock`); new keys go in `keys.go`. Example: node health monitor (`core/services/nodes/health.go`), failover prober.
3. **Stateless per request** — nothing to share.
4. **Per-instance** — allowed only with the reason written in the feature's docs.
- Tests: shared and single-runner features include a two-instance test on `testutil.NewFakeBus()` (`core/services/testutil/fakebus.go`); note that the fake bus delivers synchronously, including the publisher's own message, so never publish while holding a lock the apply path takes.
- Checklist for PRs.
AGENTS.md Quick Reference bullet:
`- **Distributed-aware state**: any feature that keeps runtime state (maps, caches, pins, schedulers, background loops) must choose shared (syncstate), single-runner (advisorylock), stateless, or documented per-instance behaviour for multi-frontend clusters. See [.agents/distributed-state.md](.agents/distributed-state.md).`
AGENTS.md Topics row:
`| [.agents/distributed-state.md](.agents/distributed-state.md) | Features that keep runtime state — how they must behave with several frontends (syncstate, advisory-lock leaders, fakebus tests) |`
api-endpoints-and-auth.md checklist line (under Quality):
`- [ ] Stateful feature: distributed mode chosen and documented (see [distributed-state.md](distributed-state.md))`
- [ ] Steps: write the files → final verification, in order, reading each output:
```bash
make protogen-go build-mock-backend build-cloud-proxy-backend build-localai-proxy-backend
go vet ./core/services/failover/... ./backend/go/localai-proxy/... ./pkg/grpc/...
go test -race ./core/services/failover/... ./core/application/... ./core/config/... ./core/http/middleware/... ./pkg/grpc/... ./pkg/mcp/localaitools/... ./backend/go/localai-proxy/... ./core/http/endpoints/openai/...
LOCALAI_TEST_HTTP_PORT=19391 go test ./core/http/...
go run github.com/onsi/ginkgo/v2/ginkgo --label-filter='!real-models' -v ./tests/e2e
make test-ui-coverage-check
LOCALAI_TEST_HTTP_PORT=19391 make test-coverage-check
```
Expected: all green; both coverage checks at or above baseline. Record any pre-existing failure (e.g. `make swagger`) with its output tail.
- [ ] Commit — `docs: require distributed-aware state for stateful features`.
File diff suppressed because it is too large. Load diff
@@ -1,73 +0,0 @@
# Configurable copy buffer design
**Date:** 21 August 2026
**Status:** Approved
## Problem
`pkg/xio.Copy` wraps a source reader so a context can stop a copy between
reads. It delegates to `io.Copy`, which uses a 32 KiB buffer for the wrapped
reader and writer types used by model downloads.
Small writes limit model import throughput when the models directory uses an
SMB volume. The development deployment reads large files from the volume at
about 104 MiB/s. A model import writes to the same volume at less than 1 MiB/s.
## Design
Keep `xio.Copy` as the context-aware copy entry point. Add variadic functional
options so existing callers continue to compile without changes.
Add an exported `Option` type and a `WithBufferSize(size int) Option` function.
`Copy` uses a 1 MiB buffer by default. A caller can override the buffer size
with `WithBufferSize`.
If a caller supplies a non-positive buffer size, `Copy` uses the 1 MiB default.
This rule prevents invalid configuration from causing an `io.CopyBuffer`
panic.
`Copy` allocates one buffer for each active call. It passes that buffer to
`io.CopyBuffer`. The context-aware reader continues to check cancellation
before each source read.
The first change does not use `sync.Pool`. A pool adds shared state and retains
large caller-selected buffers. Measurements do not justify that complexity.
## Compatibility
The existing signature gains only a variadic argument:
```go
func Copy(ctx context.Context, dst io.Writer, src io.Reader, options ...Option) (int64, error)
```
All existing calls remain source compatible. Copy results and cancellation
errors do not change.
The default buffer increases temporary memory use by approximately 992 KiB for
each concurrent copy compared with the current 32 KiB buffer.
## Tests and measurement
Add a Ginkgo suite for `pkg/xio`. Tests cover these behaviors:
- `Copy` copies the complete source.
- The default buffer permits reads larger than 32 KiB.
- `WithBufferSize` changes the maximum requested read size.
- A non-positive override uses the default buffer.
- A canceled context stops the copy and returns the context error.
Add a benchmark that runs `Copy` with the default buffer and representative
overrides. The benchmark records throughput and allocations. It does not make
timing assertions.
Run the focused `pkg/xio` suite first. Then run the packages that call
`xio.Copy`: `pkg/downloader` and `pkg/oci`.
## Deployment validation
The code change alone does not alter the running development deployment. After
CI publishes a development image and Flux deploys it, import a large model to
the NAS-backed models directory. Compare the progress rate with the previous
0.7-0.8 MiB/s result.
@@ -1,357 +0,0 @@
# Distributed Model Configuration Revisions
## Problem
Editing a model configuration in a distributed LocalAI deployment can leave the
cluster serving different effective configurations for the same logical model.
The frontend reloads the edited YAML and asks workers to stop the model, but the
existing `backend.stop` message is fire-and-forget. The frontend therefore
removes routing state without knowing whether the worker process stopped.
Separately, the replica reconciler persists `ModelLoadInfo` independently of
live `NodeModel` rows. This is necessary for restoring `min_replicas` after a
worker failure, but the persisted options currently have no relationship to a
specific revision of the model configuration. After an edit, the reconciler can
restore a replica from options captured before the edit.
The observed result was one replica serving a context near 100K while another
served the default 8K context and default parallelism. Requests behaved
differently depending on which replica the router selected. The problem is not
specific to `context_size`: any load-time model option can be stale.
## Goals
- Make all routable replicas of a logical model belong to the current model
configuration revision.
- Prevent the reconciler and late load jobs from restoring options belonging to
an older revision.
- Remove a model from routing before attempting distributed cleanup.
- Confirm that the exact worker process exited before deleting its registry
row.
- Recover safely when a worker or NATS is temporarily unreachable.
- Apply the same lifecycle to raw YAML edits, structured configuration patches,
renames, disabling, and changes received from peer frontends.
- Preserve the existing ability to restore `min_replicas` after ordinary
worker or backend failure when the model configuration has not changed.
- Expose enough state to diagnose why two replicas have different effective
options.
## Non-goals
- Requiring identical hardware-derived options on heterogeneous workers.
- Changing `model.unload`, which remains a memory-release operation.
- Replacing backend administration operations such as backend upgrade, delete,
or stop-all.
- Automatically upgrading workers that do not support the new stop protocol.
- Making arbitrary out-of-band filesystem edits transactional across multiple
machines. Such edits are detected when the model configuration loader next
refreshes the model.
## Configuration identity
Each validated model configuration has a `config_revision`. The revision is a
SHA-256 digest of a canonical semantic representation of the validated model
configuration. Formatting, YAML comments, and map ordering do not affect the
revision. Load-time request overrides and node-specific hardware tuning are not
part of this digest.
Canonicalization must use the typed, validated configuration rather than raw
YAML bytes. The canonical representation includes every field that can affect
model loading or serving. Fields used only to locate the source file or report
runtime status are excluded. The canonical encoder must produce stable field
and map ordering and must distinguish absent values where absence has different
semantics from an explicit zero value.
The revision is carried with the model options from configuration loading into
the distributed router. It is also persisted in:
- `ModelConfigState`, keyed by logical model name, as the currently accepted
revision;
- `ModelLoadInfo`, alongside the serialized `pb.ModelOptions` used for future
reconciliation;
- `NodeModel`, identifying the revision used for that live replica.
Each `NodeModel` also records an `effective_options_hash`, computed from the
fully materialized `pb.ModelOptions` after node-specific hardware defaults and
file-path staging rewrites. This hash is diagnostic only. Two replicas may have
different effective hashes and remain compatible when they share the same
configuration revision.
Rows created by older versions have an empty revision. They remain usable until
the model's first revision-aware configuration mutation. Once a current
revision is recorded, empty-revision rows are stale and cannot be routed.
## Registry invariants
The database is the coordination boundary shared by frontend replicas.
1. At most one current configuration revision exists per logical model name.
2. A `NodeModel` is routable only when it is in the loaded state and its
`config_revision` equals the current `ModelConfigState` revision.
3. A `ModelLoadInfo` row is reconcilable only when its revision equals the
current `ModelConfigState` revision.
4. A load job may publish `NodeModel` or `ModelLoadInfo` state only when its
captured revision still equals the current revision.
5. Advancing the current revision and quarantining prior-revision replica rows
happen in one database transaction.
The load-info upsert becomes compare-and-set rather than unconditional
last-write-wins. If the load's revision is no longer current, the upsert returns
a typed stale-revision error. The load is then abandoned and its worker process
is stopped through the exact stop protocol. A late load can therefore neither
be routed nor overwrite current reconciliation options.
Normal worker death does not change `ModelConfigState` or delete matching
`ModelLoadInfo`; this preserves restart recovery. A configuration mutation
advances `ModelConfigState` and invalidates older load information.
## Configuration mutation lifecycle
All model configuration mutation entry points use one model administration
lifecycle service. The structured PATCH endpoint must no longer bypass local
shutdown behavior.
For an edit that keeps the same logical model name, the service:
1. Validates and persists the new configuration.
2. Reloads it and computes its semantic revision.
3. In one transaction, records the new current revision, marks every replica
from another or empty revision as `unloading`, and removes or supersedes old
`ModelLoadInfo`.
4. Broadcasts the revision-aware invalidation to peer frontends.
5. Starts cleanup for each quarantined replica using exact `model.stop`.
6. Deletes a replica row only after confirmed process termination or confirmed
absence of that exact process.
Marking rows `unloading` precedes network calls. A worker that cannot be reached
therefore cannot continue receiving inference traffic through LocalAI even if
its old backend process is still alive.
The configuration save is durable even if cleanup is incomplete. The endpoint
must not report that saving failed after the new file and revision have
committed. Its response reports that cleanup is pending, and the condition is
also logged and exposed through the existing model/node lifecycle status
surfaces. Subsequent retries finish cleanup.
For rename, the old identity is quarantined and stopped under its old name. The
new identity receives its own current revision. Old load information is not
copied to the new name. Disable performs the same quarantine and cleanup but
does not permit fresh loads while disabled. Delete follows the existing file
deletion lifecycle after exact process cleanup.
Peer invalidation events carry the logical model name, operation, and new
revision. Applying an event is idempotent. A peer that already observes that
revision refreshes its in-memory configuration but does not create a second
cleanup generation.
When the existing configuration watcher detects an out-of-band file change, it
computes the revision after validation and submits the same lifecycle
transition. A parse or validation failure leaves the last accepted revision
current and does not quarantine its replicas. This does not make filesystem
writes atomic, but it ensures a successfully observed external edit cannot
silently bypass revision-aware routing.
## Exact worker process stop
A new request/reply NATS operation, `model.stop`, is separate from the existing
ambiguous `backend.stop` operation.
The request contains:
```text
model_name
process_key
expected_address
force
config_revision
```
`process_key` is the exact supervisor key, including replica index. The
controller derives it from the registry row rather than asking the worker to
resolve a bare backend or model name. `expected_address` prevents a stale row
from stopping an unrelated process after port reuse. `config_revision` is
included for auditability; process key and expected address are the worker-side
identity checks because workers do not own the configuration database.
The reply contains:
```text
matched
freed
terminated
process_key
address
error
```
The worker verifies that both process key and address identify the same
supervised process. A mismatched address is an error and never stops anything.
An absent process is a successful idempotent outcome with `matched=false` and
`terminated=true` because there is no process left to clean up.
For a graceful request, the worker performs bounded gRPC `Free()` and then
terminates the supervised process. A `Free()` failure is recorded but does not
prevent termination. A forced request skips `Free()`. The worker replies only
after the process has exited and its supervisor bookkeeping and port ownership
have been updated.
The existing operations retain their meanings:
- `model.unload` calls gRPC `Free()` without promising process termination;
- `backend.stop` remains an administration and compatibility operation whose
identifier may be a backend name;
- `model.stop` is the only operation used to confirm configuration-generation
cleanup for an exact replica.
Sending both `model.unload` and `model.stop` is unnecessary because graceful
`model.stop` already performs bounded `Free()` before termination.
## Unreachable workers and retry
An `unloading` replica is never routable. Failed `model.stop` attempts retain
the row with its last error, attempt count, and next retry time. A bounded,
backoff-based cleanup loop retries exact stops. Retries are idempotent and are
claimed through the database so multiple frontend replicas do not concurrently
own the same attempt.
The existing recovery paths remain backstops:
- Worker re-registration clears all `NodeModel` rows for that node because a
restarted worker has no surviving supervised backend processes.
- The per-model health monitor removes rows after consecutive unreachable
backend probes.
- Node offline handling prevents scheduling onto a worker with stale
heartbeats.
Cleanup-row removal through any of these paths fires the existing replica
removal hooks. It does not restore stale `ModelLoadInfo` because only the
current revision is eligible for reconciliation.
If a worker keeps heartbeating but does not support `model.stop`, the row stays
quarantined and the error clearly identifies an incompatible worker version.
The system favors temporary unavailability over silently serving an obsolete
configuration. Restarting or upgrading that worker lets re-registration or a
subsequent retry complete cleanup.
## Reconciliation and loading
The reconciler reads the current revision and matching `ModelLoadInfo` in one
consistent operation. If no matching load information exists, it does not use
an older blob. It records a diagnostic explaining that the model must first be
loaded under its current revision.
The next inference request builds options from the current configuration,
captures its revision, and performs the normal install, staging, and load
sequence. On success, it transactionally records the replica and current
`ModelLoadInfo`. The reconciler may then restore additional `min_replicas`
using that revision.
Every scheduling and routing decision rechecks revision eligibility when it
claims a replica. A replica selected immediately before a concurrent edit must
fail the claim after the edit advances the current revision. Existing in-flight
requests may finish; no new request is assigned to the old replica. Graceful
cleanup waits for bounded `Free()` behavior and then terminates it.
## API and observability
Model and node lifecycle responses should expose, where replica details are
already returned:
- current model `config_revision`;
- replica `config_revision`;
- `effective_options_hash`;
- lifecycle state, including `unloading`;
- pending cleanup error and retry time.
Logs for routing, reconciliation, load completion, stale-load rejection, and
cleanup include model name, replica index, node ID, and abbreviated revision.
No serialized model options or request content is added to logs.
The Web UI does not require a new workflow. After saving, it may show that the
configuration is saved while one or more old replicas are still being cleaned
up. User-facing distributed-model documentation explains this state and the
requirement to upgrade workers that lack acknowledged `model.stop` support.
## Rolling upgrades
Database migrations add nullable revision and cleanup columns so old binaries
can continue reading existing rows. New frontends treat missing revisions as
legacy state according to the compatibility rule above.
The new NATS subject avoids changing the semantics of `backend.stop` for old
workers. A new frontend receiving no responder for `model.stop` leaves the
replica quarantined and reports the compatibility problem. It must not fall
back to fire-and-forget `backend.stop`, because doing so would recreate the
original false-success failure.
Deployments should upgrade workers before or together with frontends. Mixed
frontend versions are tolerated at the database level, but old frontends do
not enforce revision-aware routing. Documentation must state that strict
cross-replica consistency is guaranteed only after all frontend replicas run
the revision-aware version.
## Testing
All Go tests use Ginkgo and Gomega.
### Registry tests
- Advancing a revision and quarantining old replicas is atomic.
- Only loaded replicas matching the current revision are returned for routing.
- Empty legacy revisions become stale after a revision-aware mutation.
- Load-info compare-and-set rejects a late old-revision write.
- Matching load information survives ordinary replica removal and worker
failure.
- Re-registration removes quarantined rows without changing current revision or
matching load information.
### Router and reconciler tests
- Given one 8K old-revision replica and one 100K current-revision replica, every
new request routes to the current revision.
- Changing `parallel` produces the same revision transition behavior as changing
`context_size`.
- The reconciler never loads from stale `ModelLoadInfo`.
- A late durable load job cannot publish a stale replica or overwrite current
load information.
- A request racing a configuration edit cannot claim the old generation.
- Heterogeneous effective option hashes remain routable when their
configuration revision matches.
### Worker protocol tests
- Exact process key and address stop the intended process and wait for exit.
- An address mismatch stops nothing.
- An already-absent process returns idempotent success.
- Graceful stop attempts bounded `Free()` and still terminates after a failure.
- Forced stop skips `Free()`.
- Replica port ownership and quarantine are updated before replying.
### Lifecycle tests
- Raw YAML edit, structured PATCH, rename, disable, and peer application all
advance or apply the expected revision and quarantine old replicas.
- A successful stop deletes the matching row.
- A timeout leaves a non-routable `unloading` row with retry state.
- Retry eventually deletes the row after the worker recovers.
- A worker without `model.stop` support produces a visible compatibility error
and never triggers fire-and-forget fallback.
- Partial cleanup does not roll back an already persisted configuration edit.
### Live distributed regression
An integration scenario loads a model on two workers, edits context and
parallel settings, and verifies that no request is routed to an old revision.
After cleanup and reload, every replica reports the current revision. The test
also disconnects one worker during the edit, verifies its replica is
quarantined, reconnects it, and verifies retry or re-registration removes the
stale row.
## Documentation impact
The implementation updates the distributed model lifecycle documentation under
`docs/content/` in the same change. It documents revision consistency,
quarantined cleanup state, rolling-upgrade requirements, and why an edited model
may wait for its first request before `min_replicas` can be restored.
No configuration key or public inference API changes are introduced.
@@ -1,67 +0,0 @@
# Distributed Staging Operations Design
## Problem
`GET /api/operations` reads file-transfer progress from the frontend replica's
in-memory `StagingTracker`. Distributed frontends broadcast tracker updates over
NATS, but those messages are transient. A replica that starts after staging has
begun, temporarily disconnects, or misses an update can return no staging row.
When a browser's one-second polls are balanced across replicas, the operation
therefore appears and disappears.
Distributed cold loads already persist their phase, placement, heartbeat, and
byte progress in PostgreSQL's `model_load_jobs` table. That row is the durable
cluster authority and should provide the baseline operations view.
## Design
Add a `NodeRegistry` query that lists active model-load jobs. The operations
endpoint will use those jobs to build one staging operation per tracking key
when the job is in the `staging` phase. It will then overlay matching local or
NATS-mirrored `StagingTracker` data, because the tracker can contain a fresher
message and filename than the periodically persisted job.
The merge is keyed by the model tracking key. A tracker entry replaces the
database entry's progress and display details rather than creating a duplicate.
Tracker-only entries remain visible for compatibility with staging paths that
do not have a durable load-job row. Database-only entries remain visible on
every replica, which eliminates flicker.
The database row supplies:
- stable operation identity (`staging:<tracking key>`),
- model name and staging phase,
- node name,
- overall progress calculated by `ModelLoadJob.Progress()`, and
- byte counters used by the frontend's ETA calculation.
The tracker overlay supplies its message, filename, node name, progress, and
byte counters when available.
## Failure Handling
If the database query fails, `/api/operations` will log the error and fall back
to the current tracker-only response. An observability failure must not break
the entire operations endpoint or hide unrelated gallery operations.
Only live `staging` rows are included. Pending, backend-installing, loading, and
failed rows are represented by their existing user-facing flows and must not be
mislabelled as file staging.
## Testing
Add focused Ginkgo coverage for:
1. A database-only staging job appears in the operations payload, reproducing
the request landing on a replica that missed all NATS broadcasts.
2. A matching tracker entry overlays the database entry without duplication.
3. Non-staging load jobs do not appear as staging operations.
4. A database read failure retains tracker-only staging operations and the
endpoint still succeeds.
Run the affected Go package tests only; no long build is required.
## Documentation
This corrects consistency of an existing UI operation and introduces no new
API, option, or user workflow. No user documentation change is required.
@@ -1,125 +0,0 @@
# Scheduling Rule Editing and Node Label Reference
## Summary
Improve the React scheduling view so cluster operators can edit existing scheduling rules and inspect node labels without moving back and forth to the Nodes page.
The scheduling page will gain a compact, collapsible node-label reference above the rules table. It will also gain an Edit action that opens the existing scheduling form with the selected rule prefilled. The model name will remain locked while editing because it identifies the rule being updated.
## Goals
- Let operators update an existing scheduling rule in place.
- Make the labels available on each node visible from the scheduling workflow.
- Keep the label reference usable for clusters with many nodes.
- Preserve the existing scheduling API and node API contracts.
- Keep the scheduling rules usable when node-label loading fails.
## Non-goals
- Editing node labels from the scheduling page.
- Renaming the model associated with an existing scheduling rule.
- Adding backend endpoints or changing scheduling semantics.
- Adding a separate scheduling documentation page for this discoverability enhancement.
## User Experience
### Node label reference
A collapsible **Node labels** section appears above the scheduling rules. It loads node data through the existing `nodesApi.list()` client and groups labels by node so operators can tell which selectors match which machines.
The expanded section contains:
- A fuzzy search field that matches node names, label keys, label values, and complete `key=value` text.
- A summary showing the visible result count and total matching node count.
- Node groups containing the node name, operational status, and its `key=value` label chips.
- Five matching nodes initially.
- A **Show 20 more** action when additional matches exist.
Changing the search query resets the visible limit to five. Clearing the query restores the unfiltered result set. The reference can be collapsed to preserve vertical space.
The matching implementation should be lightweight and local to the page. It should normalize searchable node data and support forgiving, case-insensitive token matching without adding a large dependency solely for this feature.
### Editing a scheduling rule
Each scheduling-rule row gains an **Edit** action beside **Delete**. Selecting Edit opens the existing scheduling form above the table and populates every editable field from the selected configuration:
- Scheduling mode
- Node selector
- Minimum and maximum replicas
- Routing policy
- Prefix-cache thresholds
The model selector is replaced by, or presented as, a visibly locked model field while editing. This prevents a rename from creating a second rule while leaving the original in place.
Only one add or edit form may be open at a time. Opening Add clears edit state; opening Edit closes any blank Add form. Cancel closes the form and discards its local changes.
Saving continues to use `nodesApi.setScheduling()`. On success, the page closes the form, shows the existing success toast, and refreshes the scheduling rules. On failure, it shows the error toast and keeps the populated form open so the operator does not lose changes.
## Component Design
### Scheduling form
Refactor `SchedulingForm` to accept an optional existing scheduling configuration. Initial form state will be derived from that configuration, including conversion of a serialized `node_selector` when necessary and derivation of the current mode from `spread_all`, replica values, and selector presence.
The form remains responsible for validation and for producing the existing scheduling request shape. The parent remains responsible for API calls, toast notifications, refreshes, and deciding whether the form is adding or editing.
### Node label reference
Add a focused scheduling-page component for label discovery. It receives node data and owns only presentation state:
- Expanded or collapsed
- Search query
- Visible result limit
Small pure helpers will normalize a node's searchable text and calculate filtered results. Node fetching remains in the scheduling page so loading and retry behavior stay next to the existing scheduling fetch lifecycle.
### Styling
Add scheduling-specific classes to `core/http/react-ui/src/App.css`. Reuse existing design-system tokens and button, input, badge, stack, and text primitives. Do not add static inline styles.
On wide screens, node groups use a responsive compact grid. On narrow screens, they collapse to one column. Search, collapse, pagination, and row actions remain keyboard accessible and expose explicit accessible names.
## Data Flow
1. The page mounts and independently requests scheduling configurations and nodes.
2. Scheduling configurations populate the rules table.
3. Node data populates the label reference; local search and limiting do not trigger network requests.
4. Selecting Edit copies one rule into form state and locks its model identity.
5. Saving posts the existing scheduling payload and refreshes the scheduling list.
6. Node-label retry repeats only the node request and does not disturb scheduling rules or an open scheduling form.
## States and Error Handling
- **Node loading:** Show a compact loading state inside the reference. Do not block the rules table.
- **No nodes:** Explain that no nodes are available yet.
- **Node without labels:** Include it in node-name search results and display **No labels**.
- **No search matches:** Show a clear empty result while preserving the query.
- **Node fetch failure:** Show an inline error with Retry. Scheduling remains fully usable.
- **Malformed selector:** Preserve the current defensive rendering behavior and avoid crashing the edit form; treat an unparseable selector as empty while keeping the rule visible.
- **Save failure:** Preserve all form values and show the existing error toast.
- **Save success:** Close the form and refresh the rules.
## Verification
Add or extend a focused Playwright scheduling spec to cover:
- Labels grouped under the correct nodes.
- Search by node name.
- Search by complete `key=value` text.
- Five-node initial limit and **Show 20 more** expansion.
- Empty-cluster, unlabeled-node, no-match, and failed-loading states.
- Edit opening with the complete rule prefilled.
- Locked model identity during editing.
- Updated values sent through the existing scheduling endpoint.
- Failed saves preserving the open form.
Run the focused Playwright spec, the React inline-style lint, and the production React build. Long repository-wide builds are outside the scope of this frontend-only change.
## Acceptance Criteria
- An operator can edit and save any existing scheduling rule without deleting and recreating it.
- The model identity cannot be changed while editing.
- An operator can inspect labels grouped by node without leaving Scheduling.
- The label reference remains compact with many nodes and supports forgiving search plus progressive expansion.
- A node API failure does not prevent viewing or editing scheduling rules.
- The enhancement works at narrow viewport widths and is keyboard accessible.
@@ -1,158 +0,0 @@
# Request-owned ephemeral staging
## Problem
Distributed requests copy transient inputs below
`<staging>/ephemeral/<category>/<request-id>`. The worker currently removes
these files only when a periodic age sweep considers them stale. A Reachy Mini
sending camera and sound data about once per second created more than 21,000
request directories and filled its Mac worker before the six-hour retention
window elapsed.
Reducing the retention window is insufficient. A time limit bounds residence
time, but the retained bytes still scale with request rate and input size. A
quota sweeper would also have to infer whether an old file is still in use.
Neither rule prevents concurrent uploads from consuming the worker's last free
space.
## Goals
- Give every ephemeral input an explicit owner and release it when that request
finishes, fails, or is cancelled.
- Keep cleanup transport-independent for HTTP and S3/NATS workers.
- Reserve capacity before accepting ephemeral bytes so concurrent requests
cannot consume configured disk headroom.
- Reject a request cleanly when its input does not fit; never evict an input
that a running request may still be reading.
- Recover abandoned files after frontend or worker crashes.
- Never inspect or remove models, data, configuration, or paths outside the
worker's ephemeral staging tree.
## Non-goals
- Retaining request inputs as a cache.
- Evicting persistent model or data files to make an inference request fit.
- Treating modification timestamps as proof that a request is active.
## Request ownership
The `FileStagingClient` already creates one request ID before staging inputs and
waits for synchronous and streaming backend calls to finish. It will track each
ephemeral key before attempting to stage it and defer one request-scoped
release identified by the request ID. Release runs after the backend call
returns, including error and cancellation paths, using a short background
timeout so cancellation of the request does not cancel its cleanup.
Request IDs will use the full UUID rather than the current eight-character
prefix. The worker enumerates only category directories for that validated
request ID and removes each entry with exact, symlink-safe deletion.
`FileStager` will expose an idempotent exact-key `ReleaseRemote` operation and
an optional request-scoped operation. The client uses one fixed-size request
message for the normal path and retains exact-key calls as a rolling-upgrade
fallback:
- HTTP sends one authenticated request containing the fixed-size request ID.
The worker derives and removes that request's exact files, then prunes empty
request and category directories without following symlinks.
- S3/NATS sends one request-reply containing the request ID so the selected
worker evicts the request's local cached files. The frontend then deletes the
matching objects from its tracked exact-key list.
Either deletion may already have happened and still counts as success.
If staging fails partway through a request, the deferred release still includes
the planned key, allowing it to remove a partial file when the transport can
identify one. Cleanup errors are logged and do not replace the inference result.
HTTP and S3 ingress register request operations before any pre-reservation
work. Before enumerating files, the capacity guard marks the request released
and waits for registered operations and admitted writes to finish. Later
operations, reservations, and cache claims for that request are rejected.
Markers expire after one hour and are capped at 16,384 entries, but a marker is
never evicted while its registered operation or cleanup scan is active.
Concurrent operation and cleanup state have the same hard cap. Disk bytes
remain independently bounded by capacity admission.
Cleanup waits within its deadline when all cleanup-pin slots are occupied.
If that deadline expires, existing entries lose active ownership so recovery
can reclaim them; registered ingress for the request remains closed until it
exits.
## Capacity admission
A worker-local ephemeral capacity guard is shared by its HTTP and S3/NATS input
paths. It accounts for both `<staging>/ephemeral`, used by HTTP, and
`<cache>/ephemeral`, used by S3 downloads. Before writing an ephemeral object,
the transport reserves its declared size. HTTP obtains the size from the upload
metadata; S3/NATS obtains it from object metadata. Reservations are serialized
in memory, cover both committed ephemeral bytes and concurrent writes, and are
returned on release or failed transfer.
Admission succeeds only when both conditions remain true after the reservation:
1. Total ephemeral bytes remain below the configured ephemeral staging limit.
2. The filesystem retains the configured minimum free-space headroom.
The guard rejects the transfer before inference when either condition fails.
An input with unknown size is written through a bounded accounting writer that
reserves fixed-size chunks before writing each chunk and stops before crossing
the limit. The existing maximum-upload-size check remains the per-file ceiling.
The limit and headroom are worker settings. By default, ephemeral data may use
the smaller of 10 GiB or 10 percent of filesystem capacity, while the worker
preserves the larger of 1 GiB or 5 percent as free-space headroom. The worker
logs the effective values at startup. A zero or negative operator value selects
the default rather than disabling protection. The guard scans the ephemeral
tree at startup to account for abandoned committed bytes. Filesystem free-space
checks are repeated at reservation time because other processes may share the
volume.
## Crash recovery
The existing periodic cleanup remains as a fallback for ownership messages lost
when a frontend or worker process dies. It uses a one-hour recovery TTL,
performs one startup sweep, and repeats every 15 minutes. It skips every key
held by an active reservation, considers the newest modification time in each
remaining request tree, and does not follow directory symlinks. It removes only
request directories below the registered `<staging>/ephemeral` and
`<cache>/ephemeral` roots.
The recovery window does not control normal storage growth. Request completion
and capacity reservations do. A recovery deletion updates the capacity guard's
accounted bytes.
## Error handling and observability
Admission failures report the requested bytes, current ephemeral usage, limit,
available bytes, and required headroom. Successful release and recovery update
usage counters. Read, stat, and remove failures include the affected path and
allow unrelated cleanup to continue. Missing ephemeral files and directories
are normal for idempotent release.
## Testing
Regression tests will establish the following behavior:
1. Successful, failed, cancelled, and streaming calls issue one request-scoped
worker cleanup only after the backend has returned.
2. Partial staging failures release the planned key without changing the main
error returned to the caller.
3. HTTP and S3/NATS release remove local files; S3/NATS also removes the object.
4. Release rejects persistent keys and path traversal, does not follow
symlinks, and leaves paths outside `ephemeral` untouched.
5. Concurrent reservations cannot exceed the byte limit or free-space
headroom, and failed transfers return their reservations.
6. Unknown-length writes stop at the capacity boundary.
7. Startup accounting includes abandoned ephemeral files, and the recovery
sweep removes only stale, inactive leftovers and updates accounting.
Focused package tests will run with race detection, followed by the relevant
repository lint and vet checks.
## Rollout
The change requires a new LocalAI worker and frontend build because both sides
participate in release. The Mac worker starts by accounting for its existing
backlog and removing recovery-expired files. The deployment check will verify
available space, admission and release logs, stable ephemeral usage under
continuous camera and audio traffic, and successful vision, sound detection,
and transcription requests.
@@ -1,99 +0,0 @@
# EXL3 gallery entries
## Goal
Add four gallery entries that expose the EXL3 configurations validated or
tracked by `vllm.cpp`. Pin each Hugging Face artifact to the revision recorded
by its source or benchmark evidence.
## Entries
### Qwen3.8 target
Add `qwen3.8-27b-exl3-vllm-cpp` for
`Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw`. This entry serves the target without a
draft model.
Use revision `19441ac874c4018295da848e250f23511361cda4`. Configure an 8,192-token
context, 2,048 cache blocks, eight sequences, and 16,384 batched tokens. Disable
prefix caching to match the measured serving configuration.
### Qwen3.8 with DFlash2
Add `qwen3.8-27b-dflash2-exl3-vllm-cpp`. This entry stages the Qwen3.8 target
and `Mia-AiLab/Qwen3.8-27B-DFlash2-EXL3-5.0bpw` at revision
`4f0436269bca761b071f05319e8e04a87cc633f9`.
Configure the `dflash` method with seven speculative tokens. Use the shipped
paged draft route. Apply the same serving limits as the target-only entry.
Tag this entry with `dflash` because it enables speculative decoding. Declare
the target-only entry as its variant. LocalAI can then prefer the faster entry
when the host supports it.
### DeepSeek V4 Flash for Spark
Add `deepseek-v4-flash-spark-exl3-vllm-cpp` for
`0xSero/deepseek-v4-flash-0731-spark`. Use the current repository revision,
`ce5ff0f1efb2e184aafc759d281bfae47d3a359c`. State that the `vllm.cpp`
runtime record used the older revision `22f28d32b9b29b4352eaa380ff8c2c170b2847ab`.
Describe the entry as a Spark and GB10-oriented REAP-K216 checkpoint. State its
large memory requirement and CUDA requirement. Do not claim a completed speed
or correctness gate that the source record does not contain.
### DeepSeek V4 Flash 3.0 bpw
Add `deepseek-v4-flash-exl3-3bpw-vllm-cpp` for
`0xSero/DeepSeek-V4-Flash-0731-EXL3-3.0bpw`. Use the current repository
revision, `e0bf84ac76a5100e8790c22ad10b70b1e2d06d71`.
Tag and describe this entry as experimental. The model card states that the
artifact is structurally complete, but end-to-end generation has not passed.
Keep this entry separate from the Spark entry because the repositories use
different layouts and have different runtime evidence.
## Artifact staging
Use LocalAI's Hugging Face artifact source for each repository. Stage complete
model repositories because these safetensors checkpoints need configuration,
tokenizer, index, and weight files.
Assign the Qwen draft artifact to a companion target. Pass its staged path in
the `vllm-cpp` speculative configuration. Do not download files through backend
startup logic.
## User-visible metadata
Use the `vllm-cpp`, `exl3`, `gpu`, and `cuda` tags on all four entries. Add
architecture, reasoning, tool-calling, and speculative-decoding tags only when
the configured model supports them.
Descriptions must distinguish measured results from unresolved work. The Qwen
DFlash2 description can cite the measured configuration and throughput. The
DeepSeek descriptions must not imply an end-to-end validation that does not
exist.
## Validation
Run the gallery schema and focused gallery tests. Add a focused test if the
artifact or variant structure is not already covered.
Validate these properties:
- Every name is unique.
- Every variant points to an existing entry.
- Each Hugging Face source has a pinned revision.
- The DFlash2 entry stages both repositories and passes the draft path.
- Only the configured DFlash2 entry has the `dflash` tag.
- YAML parsing and gallery loading succeed.
No model download or GPU benchmark is part of this LocalAI change. The
`vllm.cpp` evidence supplies the runtime record.
## Out of scope
- Changes to the `vllm-cpp` backend binaries.
- New EXL3 kernels or model loaders.
- New benchmark claims.
- Gallery entries for unselected EXL3 bit widths.
@@ -1,356 +0,0 @@
# Failover chains: distributed mode, localai-proxy backend and WebUI
Date: 2026-09-26
Status: design approved in brainstorming, pending spec review
Builds on: `2026-09-26-model-failover-chains-design.md` (same PR)
## Problem
The failover chains in this PR work on one LocalAI instance. Three gaps
remain:
1. **Distributed mode.** With several frontends, chain definitions converge
(config edits broadcast `cache.invalidate.models`), but runtime state does
not. Each frontend probes, trips and pins on its own. A pin applies only
on the frontend that received it and is lost on restart. `/api/failover`
and its event stream show a different view on each frontend. `warm: true`
pins only a frontend stub, so workers can evict the model, and every
frontend preloads it.
2. **Remote targets cover only chat.** `cloud-proxy` forwards chat and
completions. A remote LocalAI cannot serve transcription, TTS, VAD, sound
detection or the other modalities as a chain target, so a realtime
pipeline cannot fail over per stage between a remote and a local LocalAI.
3. **No UI.** Chains can be edited only as raw JSON in the model editor, and
their health is visible only through the API.
LocalAI also has no rule that makes a feature state how it behaves with
several frontends. The failover feature shipped with per-instance state
because nothing asked the question.
## Goals
- **C. Distributed-aware failover.** Pins, target health and chain state are
the same on every frontend. One frontend probes. Warm targets stay loaded on
workers and are preloaded once.
- **A. `localai-proxy` backend.** A gRPC backend that serves every backend
method with a REST counterpart by calling an upstream LocalAI, including
live transcription through the upstream's realtime API.
- **B. WebUI.** A chain editor field, a chain template, a live health strip
per chain, a "chain" badge in the model list, and a Failover overview page.
- **D. Contributor rule.** `AGENTS.md` and a new `.agents/distributed-state.md`
require every stateful feature to choose and document a distributed mode.
## Non-goals
- Sharing failover state between instances that are not in one distributed
cluster.
- A bridge for backend methods that have no REST or realtime counterpart
upstream (audio encode/decode, metrics, status, fine-tune, quantization).
- Chains as router candidates or as the realtime classifier model.
---
## C. Distributed-aware failover
Standalone mode (no NATS, no PostgreSQL) keeps today's behaviour. Everything
below applies when distributed mode is on.
### Shared state
| State | Writers | Mechanism | Survives restart |
|---|---|---|---|
| Pins | any frontend (REST, MCP) | `syncstate.SyncedMap` named `failover.pins`, key = chain, with a gorm `Store` | yes |
| Target health: state, last error, since, consecutive passes | any frontend on a local transition, and the probe leader | `syncstate.SyncedMap` named `failover.targets`, key = target, NATS only | no |
| Chain state: active target, active since, chain state | the probe leader only | `syncstate.SyncedMap` named `failover.chains`, key = chain, NATS only | no |
- The pins table is created under `advisorylock.KeySchemaMigrate`, the same
way the jobs store creates its tables.
- The manager gets a small `StateSync` dependency. The standalone
implementation is a no-op; the distributed implementation wraps the three
maps. The manager does not import NATS or gorm directly.
- A peer delta is applied through `OnApply`, which changes local state
without publishing again (no echo loops).
- The NATS-only maps have no `Store`, so a `Reconcile` tick's hydrate would be
a no-op — nothing durable to pull from, so it could never help a late
joiner. Instead the leader republishes every target and chain snapshot every
10 s. A frontend that joins late converges within 10 s and uses its own
state until then.
### Who does what
- **Every frontend** plans requests from the shared state. `Plan` already
leaves unhealthy targets out of the attempt order, so a target tripped on
another frontend is skipped at once. In-request retry stays local.
- **Any frontend** that sees a real request trip or pass a target publishes
the new target state.
- **The probe leader** runs probes, recovery confirmation, dwell-based
fail-back, the chain recompute and the warm preload. It publishes chain
state. Leadership uses `advisorylock.RunLeaderLoop` with a new key
`failover-prober` and the same 1 s interval as the scheduler.
- **Followers** do not recompute the active target. They adopt the leader's
chain state. When no chain state has arrived yet (start-up), a follower
uses its own recompute until the first delta.
- If the leader stops, another frontend takes the lock on its next tick.
Pins and target health are not affected. Probes and fail-back pause for at
most one tick.
### Events
Each frontend emits `chain.switched` and `target.state` to its own
subscribers (SSE, realtime `localai.model.failover`) when it applies a
change, whether the change is local or from a peer. Every frontend's stream
therefore shows the same events.
### Warm targets
- The SmartRouter's and ReplicaReconciler's pinned-model resolver includes
`WarmTargets()`, so workers never evict a warm target.
- Only the leader preloads warm targets.
- A frontend's loaded check sees only its own model stubs. A warm target
without a stub on the leader is treated as not loaded: its liveness passes
and its recovery is inconclusive (it heals after `min_dwell`). Worker health
is left to the node health monitor and to real requests.
### Tests
- Unit: two managers on the test fakebus (`core/services/testutil`). A pin on
one shows on the other. A trip on one is skipped by the other's plan. Only
the lock holder probes. A new leader resumes fail-back.
- A spec in `tests/e2e/distributed` when its harness supports two frontends
cheaply; otherwise the unit specs are the coverage and the PR says so.
### Docs
`model-failover.md` replaces the "state is per instance" limit with a
"Distributed mode" section that describes the table above.
---
## A. `localai-proxy` backend
### Shape
- `backend/go/localai-proxy` is a separate OCI gallery backend, like
`cloud-proxy`. It is registered in the `Makefile`, `backend/index.yaml` and
`.github/backend-matrix.yml` (Linux amd64/arm64 and Darwin Metal), following
`.agents/adding-backends.md`.
- It reuses cloud-proxy's auth header, HTTP client (no redirects) and
hop-by-hop header helpers. It has no translate mode: the upstream is always
LocalAI.
- `Load` refuses a model without proxy options, so greedy backend probing
never selects it.
### Config
```yaml
name: argus-whisper
backend: localai-proxy
known_usecases: [transcript]
proxy:
upstream_url: https://argus:8080 # base URL; each method appends its path
upstream_model: whisper-large # optional; default: this model's name
api_key_env: ARGUS_KEY
request_timeout_seconds: 60 # applies to non-streaming calls
```
- `core/backend/options.go` passes `ProxyOptions` to `localai-proxy` as well
as `cloud-proxy`.
- `proxy.mode` and `proxy.provider` are ignored, with a load warning.
- A `localai-proxy` model without `known_usecases` loads with a warning, because
usecases decide default-model selection and the failover inference probe.
- The failover prober already treats `localai-proxy` as remote.
`UpstreamBase` accepts a base URL unchanged.
### Method mapping
| Backend method | Upstream endpoint |
|---|---|
| Predict, PredictStream | `/v1/chat/completions`, `/v1/completions` (SSE when streaming) |
| Embedding | `/v1/embeddings` |
| Rerank | `/v1/rerank` |
| TokenizeString, Detokenize | `/v1/tokenize`, `/v1/detokenize` |
| Score | `/api/score` |
| GenerateImage, UpscaleImage | `/v1/images/generations`, `/v1/images/upscale` |
| GenerateVideo | `/video` |
| Generate3D, Animate3D | `/3d/generations`, `/3d/animate` |
| TTS, TTSStream | `/tts` (streaming passes the upstream WAV header and PCM through) |
| SoundGeneration | `/v1/sound-generation` |
| AudioTranscription, AudioTranscriptionStream | `/v1/audio/transcriptions` (`stream=true` for SSE deltas) |
| AudioTranscriptionLive | upstream `/v1/realtime` transcription session (see below) |
| Diarize | `/v1/audio/diarization` |
| VAD | `/v1/vad` |
| SoundDetection | `/v1/audio/classification` |
| Detect, Depth | `/v1/detection`, `/v1/depth` |
| FaceVerify, FaceAnalyze | `/v1/face/verify`, `/v1/face/analyze` |
| VoiceVerify, VoiceAnalyze, VoiceEmbed | `/v1/voice/verify`, `/v1/voice/analyze`, `/v1/voice/embed` |
| Stores* | `/stores/set`, `/stores/get`, `/stores/delete`, `/stores/find` |
| AudioTransform | `/audio/transformations` |
Every request uses the upstream model name (`proxy.upstream_model`, else the
model name), the same derivation as `failover.UpstreamModel`.
Methods with no counterpart (AudioEncode, AudioDecode, AudioToAudioStream,
TokenClassify, GetMetrics, Status, ModelMetadata, fine-tune and quantization)
return gRPC `Unimplemented` with the message
`localai-proxy: <method> has no upstream counterpart`.
### Files
Core passes some inputs and outputs as local paths:
- Inputs (transcription and diarization audio, sound detection `src`, image
`src` and reference images): the proxy reads the file and uploads it as
multipart or base64, as the endpoint expects.
- Outputs (TTS, image, sound generation `dst`): the proxy writes the upstream
result (bytes, or a download of the returned URL, or decoded base64) to
`dst`.
### Live transcription bridge
The upstream realtime API needs a pipeline model (VAD and transcription). The
proxy takes it from the model's backend options:
```yaml
options:
- realtime_pipeline:argus-transcribe # an upstream pipeline config
```
Without this option, `AudioTranscriptionLive` returns the standard
"live transcription unsupported" error, and realtime uses its non-live
transcription path for the stage.
With the option, `AudioTranscriptionLive` opens a WebSocket to
`<upstream>/v1/realtime?model=<realtime_pipeline>`:
1. On the first `TranscriptLiveConfig`, send `session.update` with
`type: transcription`, the input rate, the language and server VAD turn
detection. Answer `ready` when `session.updated` arrives.
2. Forward each `TranscriptLiveAudio` as `input_audio_buffer.append`
(PCM float to PCM16 base64 at the session rate).
3. Map `conversation.item.input_audio_transcription.delta` to `delta`, and
`...completed` to `delta` (any remaining text) plus `eou: true`.
4. When the gRPC send side closes, do **not** commit the buffer: the upstream
rejects a manual commit under server VAD. Instead wait up to 5 s for any
turn already in flight (speaking, stopped-but-not-committed, or committed
but not yet completed) to finish on its own, then send `final_result` and
close.
5. An upstream error or disconnect ends the gRPC stream with `Unavailable`.
Word timings and `eob` are not available from the upstream and stay empty.
A realtime stage whose live session fails reopens on the next chain target at
the next utterance (behaviour from the base spec).
### Core changes
- **Rerank for Go backends.** `pkg/grpc` gets an optional rerank interface and
a server handler, in the same way as `Score`.
- **`Unimplemented` is a capability gap.** The failover retry path (HTTP and
`Manager.Do`) treats gRPC `Unimplemented` like an admission rejection: skip
to the next target for this request, and do not trip the target. Otherwise a
chain of a remote and a local target fails a request the local target can
serve.
### Tests
- Unit: a fake LocalAI `httptest` upstream per method family (request path,
body, model name, auth header, file upload and `dst` write).
- Unit: a fake WebSocket upstream for the live bridge (ready, deltas, eou,
final result, upstream disconnect).
- E2E: `localai-proxy` models that point back at the test server's own mock
models. A realtime pipeline whose stages are chains of a `localai-proxy`
target and a local target completes a turn, and switches stage when the
proxy target fails.
### Docs
A `localai-proxy` section in `docs/content/features/backends.md` (or the page
that documents `cloud-proxy`), and a remote-LocalAI example on
`model-failover.md`.
---
## B. WebUI
### Model editor
- A `failover-targets` field component replaces the JSON editor for
`failover.targets` (`core/config/meta/registry.go` switches the component
name). Each row has a model picker (`SearchableModelSelect`), move up/down,
remove and a **warm** toggle. The toggle is disabled with a tooltip on remote
targets. Inline validation: at least 2 targets, no duplicates, no chain as a
target.
- The probe, trip and recovery fields stay in the Advanced group.
- A **Failover chain** template in `modelTemplates.js`, seeded with two empty
targets. `?template=failover` preselects it.
### Health strip
When the edited model is a chain, a `FailoverChainStatus` component above the
form shows:
- the chain state (primary / fallback / degraded) and the active target, with
the time since it became active;
- for each target: state, kind, warm, last probe and last error;
- **Pin** and **Unpin** for admins, behind a confirm dialog. The pinned target
is marked.
### Model list and overview
- Installed models: a "chain" badge with the active target, in the same way
as the alias badge.
- Operate → Runtime → **Failover**: a dense table with one row per chain
(state, active target, target states, time since the last switch). Each row
links to the chain in the model editor. With no chains, an empty state links
to the Failover chain template.
### Live data
A `useFailoverChains` hook fetches `GET /api/failover`, then opens an
`EventSource` on `/api/failover/events`. `snapshot` replaces the state;
`chain.switched` and `target.state` patch it. The browser reconnects the
stream, and the hook polls every 15 s as a fallback (the Agent Status
pattern). A `failoverApi` group in `src/utils/api.js` holds the calls.
### Conventions
- Design tokens and CSS classes only; no new inline styles (inline-style
ratchet).
- `StatusPill` tones: success for healthy and primary, warning for recovering
and fallback, error for down and degraded, muted for missing.
- Strings in the `models` and `admin` i18n namespaces for all 8 locales.
- Pin controls are hidden when `useAuth().isAdmin` is false.
### Tests
Playwright specs with mocked APIs and a mocked `text/event-stream`: the editor
component, the template, the health strip updating on events, pin controls
for admins and not for other users, and the overview page. UI line coverage
stays at or above `core/http/react-ui/coverage-baseline.txt`.
---
## D. Contributor rule
- New guide `.agents/distributed-state.md`. A feature that keeps runtime
state (in-memory maps, caches, pins, schedulers, background loops, probes)
chooses one mode and documents it:
- **shared**: `syncstate.SyncedMap`, with a `Store` when the state must
survive a restart;
- **single-runner**: an `advisorylock` leader loop;
- **stateless per request**;
- **per-instance**: allowed only with the reason written in the feature's
docs.
The guide gives one real example per mode (finetune jobs, the node health
monitor, open responses, failover chains). Shared and single-runner
features include a fakebus test with two instances.
- `AGENTS.md`: a Quick Reference bullet "Distributed-aware state" and a row in
the Topics table.
- `.agents/api-endpoints-and-auth.md`: a checklist line "Stateful feature:
distributed mode chosen and documented (see distributed-state.md)".
## Order of work
C first (it changes code already in the PR and the event contract the UI
reads), then A (it needs the `Unimplemented` classification and the rerank
handler), then B, then D. All in PR #12285.
@@ -1,415 +0,0 @@
# Model failover chains
Date: 2026-09-26
Status: design approved in brainstorming, pending spec review
## Problem
A LocalAI instance that serves a model from a remote upstream (for example a
`cloud-proxy` model that points at a larger LocalAI cluster) has no way to fall
back to a local model when that upstream is unhealthy. Clients that want this
today build it themselves. The wingman voice assistant, for example, keeps an
ordered list of realtime WebSocket endpoints, quarantines an endpoint after
repeated backend errors, and reconnects to the next one. This has three costs:
- Every client reimplements failover, health tracking and recovery.
- The switch happens at the session level. A realtime conversation loses its
context when the client reconnects to another endpoint.
- The client sees at least one failed request before it reacts.
## Goal
A model config can declare an ordered **failover chain** of target models.
LocalAI serves each request for the chain from the highest-priority healthy
target, retries on the next target when a target fails before the response is
committed, probes targets actively, fails back with hysteresis, and tells
clients when the active target changes.
Because a chain is a model name, it works in every place that takes a model
name, including the `llm`, `transcription`, `tts` and `vad` fields of a
realtime pipeline. The realtime session stays on the local instance, so a
switch changes the stage that serves the next turn and the conversation
history survives.
## Non-goals
- A dedicated WebUI chain editor and live health view. That is a follow-up
(sub-project 3). This spec only registers the config fields in the field
metadata registry, so the generic model editor can show them.
- The `localai-proxy` backend (a remote target kind that covers the full
LocalAI API through gRPC). That is a follow-up (sub-project 1). This spec
defines the "remote target" probe path it will plug into.
- Shared failover state across several LocalAI frontends. State is in memory
and per instance.
- Transparent failover after a response has started to stream.
## Config schema
A chain is a model config with a `failover` block. Like an alias, it has no
backend of its own.
```yaml
name: assistant-llm
failover:
targets:
- model: argus-llm # for example a cloud-proxy config
- model: gemma-local
warm: true # load at startup, never evict
probe:
interval: 15s # liveness probe interval
timeout: 5s
trip:
errors: 1 # retryable failures in `window` that mark a target down
window: 30s
recovery:
probes: 3 # consecutive inference probes to confirm recovery
min_dwell: 60s # minimum time on a lower target before fail-back
```
Only `targets` is required. The other values in the example are the defaults.
### Validation
The config loader rejects a chain when:
- `failover` is set together with `alias` or `backend`.
- `targets` has fewer than 2 entries.
- A target does not exist, or the same target is listed twice.
- A target is itself a chain. Chains do not nest.
A target can be an alias. The alias is resolved one hop, as it is today.
The loader logs a warning, and does not reject, when:
- The known usecases of the targets do not overlap (for example an LLM and a
TTS model in one chain). Usecases are often inferred, so this cannot be an
error.
- `warm: true` is set on a remote target. The flag has no effect there.
### Target kinds
The kind decides how a target is probed. It is inferred from the backend of the
target config:
- **remote**: `cloud-proxy`, and `localai-proxy` when it exists.
- **local**: every other backend.
### Naming in responses and accounting
The behaviour matches aliases. Responses echo the chain name. Usage and traces
record `requested=<chain>` and `served=<target>` through the existing
`ContextKeyRequestedModel` and `ContextKeyServedModel` keys.
The upstream of a remote target never sees the chain name. A request served
through a chain reaches a remote target with that target's upstream model:
`proxy.upstream_model`, or the target name when it is empty. This holds in
passthrough and translate mode, and it is the same name the liveness probe
looks for in `/v1/models` (one helper derives both).
## Failover manager
New package: `core/services/failover`. The application creates one `Manager` at
start and keeps it in sync with the model config loader. When a chain is added,
edited or removed (from YAML, the model editor API or the MCP tools), the
manager updates without a restart.
### State
- Health is tracked **per target**. A target that is in two chains is probed
once, and a failure marks it down for both.
- The active target is tracked **per chain**.
Target states:
```
trip (errors within window, from requests or probes)
healthy ───────────────────────────────▶ down
▲ │ liveness probe passes
│ recovery.probes consecutive ▼
└──── inference probes pass ◀──── recovering ──(any failure)──▶ down
```
- At startup, targets are `healthy`. The manager runs one liveness pass
immediately, and in-request retry covers the gap until it completes.
- A target whose config is removed becomes `missing`. It is treated as `down`,
and the chain reports it.
Chain states:
- `primary`: the active target is target 0.
- `fallback`: the active target is a lower target.
- `degraded`: all targets are down.
### Selecting the active target
- The active target is the highest-priority `healthy` target.
- Failover to a lower target is immediate.
- Fail-back to a higher target happens only when that target is `healthy`
(which already needs `recovery.probes` passing inference probes) **and** the
current target has been active for at least `min_dwell`.
- When the chain is `degraded`, requests still try every target in priority
order. A probe can lag behind a recovery, so the manager does not fail fast.
- A manual pin (see API) forces the active target. While a pin is set, probes
continue and report state, but they do not change the active target.
### Probes
| Target | Liveness (steady state) | Recovery confirmation |
|---|---|---|
| remote | `GET <base>/v1/models` returns 2xx and lists the upstream model. `<base>` is the scheme and host of `proxy.upstream_url` plus any path prefix before `/v1`. The upstream model is `proxy.upstream_model`, or the target name when it is empty. `/v1/models` works on any OpenAI-compatible upstream, and `/readyz` exists only on LocalAI. | one minimal real request, chosen by usecase |
| local, `warm: true` | gRPC `HealthCheck` on the loaded backend, with the probe timeout. A probe never loads the model: when the backend is not loaded (the warm preload is still loading it, or a crash removed it), liveness passes and the next real request loads and judges it. | chat and completion: `Predict` with 1 token; embeddings: `Embedding` of `"ping"`; other usecases: `HealthCheck`. A local backend process that answers `HealthCheck` rarely fails only for TTS or transcription. When the backend is not loaded there is nothing to confirm against: the probe neither passes nor trips, and the target returns to `healthy` like a cold one, when `min_dwell` has passed since the trip. |
| local, cold | none: a cold target is judged only by real requests; it is never loaded only to probe it. | none. After a trip, the target returns to `healthy` when `min_dwell` has passed. The next real request is the test. |
Minimal requests by usecase:
- chat and completion: `max_tokens: 1`
- embeddings: the input `"ping"`
- transcription: 200 ms of silence
- TTS: the text `"ok"`
- expensive usecases (image, video, 3D): no inference probe. Liveness is the
confirmation.
Probe load rules:
- One global scheduler ticks every second, without jitter, and starts the
probes that are due. A target shared by chains is probed once. The
scheduler does not wait for a probe: a target whose probe is still running
is skipped, so one slow target does not delay the others.
- A successful real request counts as a liveness pass, so a busy target is
almost never probed.
- Inference probes run only while a target is `recovering`.
- Remote probes use the URL and API key from the target's proxy config.
- Probe results go into the same trip counter as request failures.
When one target is in several chains, its `probe`, `trip` and `recovery`
settings come from the first of those chains in name order. The docs state
this.
### Warm targets
The manager loads `warm: true` targets at startup and marks them pinned in the
watchdog, so LRU and idle eviction skip them. They still count toward the
active backend limit. When pinned warm targets leave no room for another load,
the loader never evicts them: it retries eviction and then loads the model
anyway, over the limit, with no error that names the warm targets. The docs
state this.
### Events
Each change of a target state or of an active target produces an event on an
internal bus:
```
{chain, target, from, to, state, reason, error, at}
```
`reason` is one of `trip`, `recovery`, `manual`, `degraded`, `missing`.
## Request path (HTTP)
### Resolution
In `core/http/middleware/request.go`, next to the alias block, a chain config is
resolved with `mgr.Plan(chain)`. The plan is the ordered list of attempts:
1. the active target,
2. the other `healthy` targets in priority order,
3. the `down` targets, only when the chain is `degraded`.
The middleware stores the plan in the request context and sets
`MODEL_CONFIG` to the config of the first target. Handlers do not change. The
plan is fixed when the request starts, so a config reload does not affect
requests that are in progress.
### In-request retry
A new middleware wraps the handler. It replaces the response writer with one
that records whether the response is committed:
- A response with status 500 or higher is buffered until the handler returns,
as long as no body has been flushed. Error bodies are small, so the wrapper
can discard them.
- A streaming response (SSE, or any flushed body) is committed at the first
flush.
When the handler returns a retryable error and the response is not committed,
the wrapper calls `mgr.ReportFailure(target, err)`, sets the config of the next
target in the plan, and runs the handler again. When the handler succeeds, the
wrapper calls `mgr.ReportSuccess(target)`. When the response is committed and
then fails, the wrapper reports the failure and does not retry.
Request bodies:
- JSON bodies are already parsed into the request context.
- Multipart bodies are cached by `ParseMultipartForm`.
- Other bodies are buffered up to a limit. A larger body gets no in-request
retry. The failure still counts toward the trip.
### Retryable errors
`failover.IsRetryable(err, status)` returns true for:
- connection and dial errors
- timeouts, when the client did not cancel the request
- gRPC `Unavailable`, `Internal`, `DeadlineExceeded` and `Unknown`
- upstream HTTP 5xx
- model load failures
It returns false for client cancellation, 4xx responses and validation errors.
These do not trip a target, because the next target would reject the same
request.
### Handler audit
Each endpoint family must be safe to run again before its response is
committed: chat, completions, embeddings, transcription, TTS, image generation,
rerank, VAD and sound detection. An endpoint that is not safe gets no
in-request retry (its failures still trip the target). The PR lists these
endpoints.
### Response headers
Every response for a chain carries:
- `X-LocalAI-Served-Model: <target>`
- `X-LocalAI-Failover: fallback` or `degraded`, when target 0 did not serve the
request
Plain HTTP clients can see failover without subscribing to events.
## Request path (realtime)
- In `core/http/endpoints/openai/realtime_model.go`, a pipeline stage that names
a chain is resolved **for each call**, not once at session start. This holds
for the full pipeline (`wrappedModel`) and for transcription-only and
sound-detection-only sessions (`transcriptOnlyModel`); both embed the same
stage router. A chain config never reaches the model loader: it has no
backend and would start backend auto-detection. A helper, `mgr.Do(ctx, chain, func(cfg *config.ModelConfig) error)`,
goes through the plan with the same classification as HTTP.
- Streaming stages (`Predict` with a token callback, `TTSStream`,
`TranscribeStream`) wrap the callback. A retry is allowed only until the first
token or audio chunk goes to the client. After that, the turn fails as it does
today, the target is tripped, and the next turn uses the next target.
- `TranscribeLive` is resolved when it opens. A failure in the middle of the
stream ends it in the existing way, and the next utterance opens it again on
the new target.
- The conversation history is in the realtime session on this instance. A
switch of the LLM stage keeps it.
- `Warmup` warms the active target of each chain stage.
## API and events
All endpoints use the global auth middleware. `GET` endpoints and the event
stream need standard auth. The pin endpoints are admin only.
### REST
`GET /api/failover` returns all chains:
```json
{"chains":[{"name":"assistant-llm","state":"fallback","active":"gemma-local",
"active_since":"2026-09-26T10:00:00Z","pinned":null,
"targets":[
{"model":"argus-llm","kind":"remote","warm":false,"state":"recovering",
"consecutive_ok":1,"last_probe":"2026-09-26T10:04:10Z",
"last_error":"503 no healthy nodes"},
{"model":"gemma-local","kind":"local","warm":true,"state":"healthy"}]}]}
```
`GET /api/failover/{chain}` returns one chain.
`POST /api/failover/{chain}/pin` with `{"target": "<model>"}` forces a target.
`DELETE /api/failover/{chain}/pin` removes the pin. The pin is in memory and a
restart clears it.
### Server-sent events
`GET /api/failover/events`:
- The first event is `snapshot`, with the same payload as `GET /api/failover`.
A new client knows the current state without a race against a separate GET.
- Then `chain.switched` with `{chain, from, to, state, reason, at}`, and
`target.state` with `{target, from, to, reason, error, at}`.
- A keepalive comment every 15 s.
### Realtime server event
`localai.model.failover`:
```json
{"type":"localai.model.failover","chain":"assistant-llm","stage":"llm",
"from":"argus-llm","to":"gemma-local","state":"fallback","reason":"trip"}
```
- The server sends it to every session whose pipeline uses the chain when the
chain switches.
- The server also sends it once for each chain stage when the session starts,
with `reason: "initial"` and `from` empty. A client knows at the start whether
it runs on the primary or on a fallback.
- `stage` is one of `llm`, `transcription`, `tts`, `vad`, `sound_detection`.
### Observability
- Metrics: `localai_failover_switches_total{chain,from,to,reason}` and
`localai_failover_target_up{target}`.
- Each failed attempt in a request is recorded in the Traces UI, so a request
served by target 2 shows why target 1 was skipped.
## Capability surfaces
As required by `.agents/api-endpoints-and-auth.md`:
- Handlers in `core/http/endpoints/localai/failover.go` with swagger blocks, tag
`failover`. Routes in `core/http/routes/localai.go`. `make swagger`.
- An `instructionDefs` entry for the new tag in
`core/http/endpoints/localai/api_instructions.go`, and the count in
`api_instructions_test.go`.
- `failover.*` fields in the config field metadata registry
(`core/config/meta/registry.go`), in a new `failover` section next to
`alias`, so the generic model editor can show and edit them.
- MCP tools in `pkg/mcp/localaitools/`: `list_failover_chains`,
`pin_failover_target` and `unpin_failover_target`, in the `inproc` and
`httpapi` clients, the skill prompts, and `toolToHTTPRoute` in
`coverage_test.go`. Chains are created and edited through the existing model
config tools.
- A docs page, `docs/content/features/model-failover.md`, linked from
`model-aliases.md`, `openai-realtime.md` and the cloud-proxy docs.
## Testing
Ginkgo and Gomega, like the rest of LocalAI. The coverage baseline must not go
down.
- **State machine**, with a fake clock: trip; recovery after N inference probes;
`min_dwell` hysteresis; `degraded`; pin and unpin; a target shared by two
chains; a `missing` target.
- **`IsRetryable`**: a table of error and status cases.
- **Config validation**: nested chain, `alias` with `failover`, fewer than 2
targets, missing target, duplicate target, the usecase warning.
- **HTTP integration**: two fake OpenAI-compatible upstreams (`httptest`) behind
`cloud-proxy` target configs.
- Upstream 1 fails. The request is served by upstream 2 with no client error,
`X-LocalAI-Served-Model` is set, and the SSE stream sends `chain.switched`.
- Upstream 1 recovers. Fail-back happens only after N inference probes and
`min_dwell`.
- Upstream 1 fails after the first SSE chunk. There is no retry, the client
gets the error, the target trips, and the next request goes to upstream 2.
- **Handler audit**: one retry test for each endpoint family that proves it is
safe to run again before commit.
- **Realtime**: a pipeline whose LLM stage is a chain of fake backends. The test
checks the `initial` event, the `localai.model.failover` event when the
primary fails, and that the conversation history is kept after the switch.
- **API**: authenticated and unauthenticated access to every endpoint; pin
requires admin.
## Follow-ups
1. `localai-proxy` backend: a fork of `cloud-proxy` that forwards every gRPC
method (Predict, Embedding, AudioTranscription, TTS, GenerateImage, Rerank,
VAD, sound detection) to the REST API of an upstream LocalAI. This lets a
remote model serve any realtime pipeline stage.
2. WebUI: a chain editor and a live health view built on
`GET /api/failover/events`.
3. Wingman: use one local LocalAI endpoint with chains for its pipeline stages,
and react to `localai.model.failover` events instead of its own endpoint
supervisor.