From 9bbcde4b1b41123307288c78ea958ca73e07244b Mon Sep 17 00:00:00 2001 From: Ettore Di Giacinto Date: Sun, 27 Sep 2026 19:55:18 +0000 Subject: [PATCH] chore: drop design specs and plans from the change Design specs and implementation plans are working notes and are not kept in the tree. Signed-off-by: Ettore Di Giacinto Assisted-by: Claude:claude-opus-5-5 [Claude Code] --- .../2026-08-02-whisper-medusa-backend.md | 63 - ...026-09-26-failover-distributed-proxy-ui.md | 689 --- .../plans/2026-09-26-model-failover-chains.md | 5054 ----------------- ...6-08-21-configurable-copy-buffer-design.md | 73 - ...stributed-model-config-revisions-design.md | 357 -- ...1-distributed-staging-operations-design.md | 67 - ...eduling-rule-editing-node-labels-design.md | 125 - ...9-07-ephemeral-staging-retention-design.md | 158 - .../specs/2026-09-07-exl3-gallery-design.md | 99 - ...26-failover-distributed-proxy-ui-design.md | 356 -- ...2026-09-26-model-failover-chains-design.md | 415 -- 11 files changed, 7456 deletions(-) delete mode 100644 docs/superpowers/plans/2026-08-02-whisper-medusa-backend.md delete mode 100644 docs/superpowers/plans/2026-09-26-failover-distributed-proxy-ui.md delete mode 100644 docs/superpowers/plans/2026-09-26-model-failover-chains.md delete mode 100644 docs/superpowers/specs/2026-08-21-configurable-copy-buffer-design.md delete mode 100644 docs/superpowers/specs/2026-08-21-distributed-model-config-revisions-design.md delete mode 100644 docs/superpowers/specs/2026-08-21-distributed-staging-operations-design.md delete mode 100644 docs/superpowers/specs/2026-08-21-scheduling-rule-editing-node-labels-design.md delete mode 100644 docs/superpowers/specs/2026-09-07-ephemeral-staging-retention-design.md delete mode 100644 docs/superpowers/specs/2026-09-07-exl3-gallery-design.md delete mode 100644 docs/superpowers/specs/2026-09-26-failover-distributed-proxy-ui-design.md delete mode 100644 docs/superpowers/specs/2026-09-26-model-failover-chains-design.md diff --git a/docs/superpowers/plans/2026-08-02-whisper-medusa-backend.md b/docs/superpowers/plans/2026-08-02-whisper-medusa-backend.md deleted file mode 100644 index 3e35636e9..000000000 --- a/docs/superpowers/plans/2026-08-02-whisper-medusa-backend.md +++ /dev/null @@ -1,63 +0,0 @@ -# Whisper-Medusa Backend Implementation Plan - -> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. - -**Goal:** Add a dedicated LocalAI speech-to-text backend for aiola Whisper-Medusa checkpoints. - -**Architecture:** A Python gRPC backend owns model loading, audio normalization, and Medusa generation. LocalAI's existing `AudioTranscription` RPC remains unchanged; build, gallery, and documentation surfaces follow the existing Python ASR backend pattern. - -**Tech Stack:** Python 3.11, PyTorch, torchaudio, transformers, whisper-medusa, gRPC, YAML, Make. - -## Global Constraints - -- Accept local model paths and Hugging Face model identifiers through `resolve_model_reference`. -- Normalize input audio to mono 16 kHz before generation. -- Default the language to `en` and expose upstream generation regulation options. -- Document upstream's archived status, 30-second clip limit, and checkpoint language limitations. -- Build Linux CPU and NVIDIA CUDA 12 images; do not claim unsupported Darwin or ROCm coverage. - ---- - -### Task 1: Backend behavior - -**Files:** -- Create: `backend/python/whisper-medusa/backend.py` -- Create: `backend/python/whisper-medusa/test_unit.py` - -**Interfaces:** -- Consumes: LocalAI `LoadModel` and `AudioTranscription` protobuf requests. -- Produces: `BackendServicer`, `_parse_options`, and `_prepare_audio`. - -- [ ] Write unit tests for option parsing, mono conversion, resampling, load failure, and transcription. -- [ ] Run `python -m unittest test_unit.py` and confirm it fails because the backend does not exist. -- [ ] Implement the minimal gRPC backend and rerun the unit tests to green. - -### Task 2: Packaging and registration - -**Files:** -- Create: `backend/python/whisper-medusa/{Makefile,install.sh,protogen.sh,run.sh,test.sh,requirements.txt,requirements-cpu.txt,requirements-cublas12.txt}` -- Modify: `Makefile` -- Modify: `.github/backend-matrix.yml` -- Modify: `backend/index.yaml` - -**Interfaces:** -- Consumes: the Python backend Docker build conventions. -- Produces: `whisper-medusa` install/build targets and CPU/CUDA backend images. - -- [ ] Add packaging scripts and pinned upstream dependency. -- [ ] Register the backend in Make and backend metadata. -- [ ] Add Linux amd64 CPU and CUDA 12 CI matrix entries. -- [ ] Validate Make and YAML parsing. - -### Task 3: Documentation and verification - -**Files:** -- Create: `docs/content/features/whisper-medusa.md` -- Modify: `docs/content/features/backends.md` - -**Interfaces:** -- Produces: user-facing model YAML and limitation guidance. - -- [ ] Document setup, configuration, options, and upstream constraints. -- [ ] Run unit tests, syntax compilation, registration checks, and diff review. -- [ ] Commit with the required `Assisted-by` trailer, push, and open a PR closing issue #3127. diff --git a/docs/superpowers/plans/2026-09-26-failover-distributed-proxy-ui.md b/docs/superpowers/plans/2026-09-26-failover-distributed-proxy-ui.md deleted file mode 100644 index 64d13e241..000000000 --- a/docs/superpowers/plans/2026-09-26-failover-distributed-proxy-ui.md +++ /dev/null @@ -1,689 +0,0 @@ -# Failover: Distributed Mode, localai-proxy and WebUI Implementation Plan - -> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. - -**Goal:** Make failover chains consistent across distributed frontends, add a `localai-proxy` backend that serves every LocalAI API (including live transcription) from a remote LocalAI, add the chain UI, and add the distributed-state contributor rule — all in PR #12285. - -**Architecture:** The failover `Manager` gets a `StateSync` dependency (no-op standalone; three `syncstate.SyncedMap`s in distributed mode) and a leader gate (PostgreSQL advisory lock per tick) so one frontend probes and owns chain state. `localai-proxy` is a new Go gRPC backend mapping each backend method to the upstream LocalAI REST endpoint, with a WebSocket bridge to the upstream realtime API for live transcription. The React UI adds an editor field, a template, a live health strip, a badge and an overview page on the existing failover REST/SSE API. - -**Tech Stack:** Go (echo, gRPC, gorm, NATS via `syncstate`, gorilla/websocket), Ginkgo/Gomega, React 19 + Vite, Playwright. - -**Spec:** `docs/superpowers/specs/2026-09-26-failover-distributed-proxy-ui-design.md` (builds on `docs/superpowers/specs/2026-09-26-model-failover-chains-design.md`) - -## Global Constraints - -- Worktree `/home/mudler/_git/LocalAI/.wt/failover-chains`, branch `feat/failover-chains`, PR #12285. Push only at the end (the controller does it). -- Commit trailer exactly `Assisted-by: Claude:claude-opus-5-5`. Never `Co-Authored-By` or `Signed-off-by`. `docs/superpowers/` needs `git add -f`. -- Build scope: `go build ./core/... ./pkg/... ./tests/... ./backend/go/localai-proxy/...` — never `go build ./...` (CGo launcher/backends need X11). -- Root `core/http` tests need `LOCALAI_TEST_HTTP_PORT=19391`. -- Logging `github.com/mudler/xlog`; `any` not `interface{}`; comments explain why. -- Standalone mode (no NATS/DB) must behave exactly as before this plan. -- SyncedMap names exactly: `failover.pins`, `failover.targets`, `failover.chains`. Leader republish interval 10 s. Advisory lock key constant `KeyFailoverProber` = 108. -- Backend name exactly `localai-proxy`; backend option `realtime_pipeline:`; `Unimplemented` message format `localai-proxy: has no upstream counterpart`. -- UI: no new inline `style={{}}` (inline-style ratchet); `StatusPill` tones: healthy/primary → success, recovering/fallback → warning, down/degraded → error, missing → muted; strings in i18n for all 8 locales (en, it, es, de, zh-CN, id, ko, pt-BR); pin controls only for `isAdmin`. -- Coverage baselines (`coverage-baseline.txt` 54.2, `core/http/react-ui/coverage-baseline.txt` 40.0) must not go down; never edit them. - -## Review Focus - -1. A NATS echo of a frontend's own publish must not deadlock or double-emit events (Task 1: "echo of own publish is a no-op"). -2. A frontend that joins late (no deltas yet) must still serve requests from local state, then converge on the leader's republish (Task 2: "late joiner converges on heartbeat"). -3. Leadership moving between frontends must not reset fail-back timing (shared `activeSince`) (Task 3: "new leader keeps activeSince"). -4. A `localai-proxy` target that returns `Unimplemented` must move the request to the next chain target without tripping (Task 4: "Unimplemented skips without trip"). -5. An upstream disconnect mid live-transcription must end the gRPC stream with `Unavailable`, not hang (Task 8: "upstream disconnect ends the stream"). - ---- - -## Part C — Distributed-aware failover - -### Task 1: Manager state-sync hooks and leader gate - -**Files:** -- Create: `core/services/failover/statesync.go`, `core/services/failover/statesync_test.go` -- Modify: `core/services/failover/manager.go`, `core/services/failover/schedule.go` - -**Interfaces:** -- Consumes: existing `Manager` internals (`setTargetLocked`, `recomputeLocked`, `Pin`, `Unpin`, `emitLocked`, `chainState`, `targetState`, `takeWarmLocked`). -- Produces: - ```go - type TargetSnapshot struct { - Target string `json:"target"` - State TargetState `json:"state"` - Reason Reason `json:"reason"` - Error string `json:"error,omitempty"` - ConsecutiveOK int `json:"consecutive_ok"` - Since time.Time `json:"since"` - } - type ChainSnapshot struct { - Chain string `json:"chain"` - Active string `json:"active"` - ActiveSince time.Time `json:"active_since"` - State ChainState `json:"state"` - Reason Reason `json:"reason"` - } - type StateSync interface { - PublishTarget(TargetSnapshot) - PublishChain(ChainSnapshot) - SetPin(chain, target string) error - ClearPin(chain string) error - Pins() map[string]string - } - type LeaderGate func(ctx context.Context, fn func()) bool - func WithLeaderGate(g LeaderGate) Option - func (m *Manager) SetStateSync(s StateSync) - func (m *Manager) ApplyTarget(s TargetSnapshot) - func (m *Manager) ApplyChain(s ChainSnapshot) - func (m *Manager) ApplyPin(chain, target string) // target "" = unpinned - func (m *Manager) IsLeader() bool - func (m *Manager) Republish() // leader: publish every target and chain snapshot - ``` - -Design rules (binding): -- Publishing happens **outside `m.mu`**: a real or fake bus delivers the frontend's own publish back synchronously to `ApplyTarget`/`ApplyChain`/`ApplyPin`, which take `m.mu`. Queue publishes in `m.pending []func()` while locked; every exported mutating method drains and runs them after unlocking (helper `m.unlockAndFlush()`). -- `ApplyTarget` sets state through `setTargetLocked` with `m.applying = true`, so it emits the local `target.state` event but does not publish. An echo whose state equals the local state is a no-op (existing early return). -- Local transitions (`setTargetLocked` with `!m.applying` and `m.sync != nil`) queue `PublishTarget`. -- Chains: when `m.sync != nil && !m.leader`, `recomputeLocked` does not change `ch.active` once the chain has adopted leader state (`ch.adopted`); before adoption it computes locally. `ApplyChain` sets `active`, `activeSince`, `state`, `adopted = true` and emits `chain.switched` (with the snapshot's reason) when `active` or `state` changed. When the leader's recompute changes a chain, it queues `PublishChain`. -- Pins: `Pin`/`Unpin` set local state immediately (read-your-writes), then call `sync.SetPin`/`ClearPin` outside the lock. `ApplyPin` sets/clears `ch.pinned` and recomputes with `ReasonManual`. `SetStateSync` hydrates pins from `s.Pins()`. -- Leader gate: in `Tick`, after `Sync()`, call `gate(ctx, fn)` where `fn` runs probe scheduling + `Reevaluate()`; set `m.leader` to the returned bool. Without a gate (standalone) the manager is always leader and `Tick` behaves exactly as today. Followers still run `Reevaluate()` (local-only when not adopted). On a false→true leader transition set `m.warmPending = true` so `onWarm` fires on the new leader. -- `Republish()` queues `PublishTarget` for every target and `PublishChain` for every chain (leader only; no-op otherwise). `Run` calls it every 10 ticks when leader. - -- [ ] **Step 1: Write the failing tests** (`statesync_test.go`, package `failover`, reuse `fakeClock`, `fakeSource`, `local`, `remote`, `chainCfg`, `t`, `errBoom`, `drain` from existing test files) - -```go -package failover - -import ( - "context" - "sync" - "time" - - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" -) - -// loopSync is an in-process StateSync that delivers every publish to all -// managers synchronously, including the publisher (like NATS echo). -type loopSync struct { - mu sync.Mutex - peers []*Manager - pins map[string]string -} - -func (l *loopSync) add(m *Manager) { l.mu.Lock(); l.peers = append(l.peers, m); l.mu.Unlock() } -func (l *loopSync) each(f func(*Manager)) { - l.mu.Lock() - ps := append([]*Manager(nil), l.peers...) - l.mu.Unlock() - for _, p := range ps { - f(p) - } -} -func (l *loopSync) PublishTarget(s TargetSnapshot) { l.each(func(m *Manager) { m.ApplyTarget(s) }) } -func (l *loopSync) PublishChain(s ChainSnapshot) { l.each(func(m *Manager) { m.ApplyChain(s) }) } -func (l *loopSync) SetPin(c, t string) error { - l.mu.Lock(); l.pins[c] = t; l.mu.Unlock() - l.each(func(m *Manager) { m.ApplyPin(c, t) }) - return nil -} -func (l *loopSync) ClearPin(c string) error { - l.mu.Lock(); delete(l.pins, c); l.mu.Unlock() - l.each(func(m *Manager) { m.ApplyPin(c, "") }) - return nil -} -func (l *loopSync) Pins() map[string]string { - l.mu.Lock(); defer l.mu.Unlock() - out := map[string]string{} - for k, v := range l.pins { - out[k] = v - } - return out -} - -var _ = Describe("Manager state sync", func() { - var ( - clock *fakeClock - src *fakeSource - bus *loopSync - a, b *Manager - leaderIsA bool - gateFor func(isA bool) LeaderGate - ctx = context.Background() - ) - - BeforeEach(func() { - clock = newFakeClock() - src = newFakeSource(remote("x"), local("y"), chainCfg("chain", nil, t("x"), t("y"))) - bus = &loopSync{pins: map[string]string{}} - leaderIsA = true - gateFor = func(isA bool) LeaderGate { - return func(_ context.Context, fn func()) bool { - if isA != leaderIsA { - return false - } - fn() - return true - } - } - a = New(src, WithClock(clock), WithLeaderGate(gateFor(true))) - b = New(src, WithClock(clock), WithLeaderGate(gateFor(false))) - bus.add(a); bus.add(b) - a.SetStateSync(bus); b.SetStateSync(bus) - a.Tick(ctx); b.Tick(ctx) - }) - - It("echo of own publish is a no-op and emits one event", func() { - events, cancel := a.Subscribe(16) - defer cancel() - a.ReportFailure("x", errBoom) - n := 0 - for _, e := range drain(events) { - if e.Type == EventTargetState && e.Target == "x" { - n++ - } - } - Expect(n).To(Equal(1)) - }) - - It("a trip on one frontend is skipped by the other's plan", func() { - b.ReportFailure("x", errBoom) - att, err := a.Plan("chain") - Expect(err).ToNot(HaveOccurred()) - Expect(att.Target()).To(Equal("y")) - }) - - It("followers adopt the leader's chain state and emit the switch", func() { - events, cancel := b.Subscribe(16) - defer cancel() - a.ReportFailure("x", errBoom) // leader recomputes and publishes chain state - st, _ := b.ChainStatus("chain") - Expect(st.Active).To(Equal("y")) - var sw []Event - for _, e := range drain(events) { - if e.Type == EventChainSwitched { - sw = append(sw, e) - } - } - Expect(sw).ToNot(BeEmpty()) - }) - - It("a pin on one frontend applies on all", func() { - Expect(b.Pin("chain", "y")).To(Succeed()) - st, _ := a.ChainStatus("chain") - Expect(st.Pinned).ToNot(BeNil()) - Expect(*st.Pinned).To(Equal("y")) - Expect(a.Unpin("chain")).To(Succeed()) - st, _ = b.ChainStatus("chain") - Expect(st.Pinned).To(BeNil()) - }) - - It("hydrates pins when the sync is attached", func() { - bus.pins["chain"] = "y" - c := New(src, WithClock(clock), WithLeaderGate(gateFor(false))) - c.SetStateSync(bus) - st, _ := c.ChainStatus("chain") - Expect(st.Pinned).ToNot(BeNil()) - }) - - It("only the leader probes", func() { - pa, pb := &fakeProber{fail: map[string]error{}}, &fakeProber{fail: map[string]error{}} - a = New(src, WithClock(clock), WithProber(pa), WithLeaderGate(gateFor(true))) - b = New(src, WithClock(clock), WithProber(pb), WithLeaderGate(gateFor(false))) - a.SetStateSync(bus); b.SetStateSync(bus) - a.Tick(ctx); b.Tick(ctx) - Eventually(func() int { return len(pa.take()) }).Should(BeNumerically(">", 0)) - Consistently(func() int { return len(pb.take()) }, 200*time.Millisecond).Should(Equal(0)) - Expect(a.IsLeader()).To(BeTrue()) - Expect(b.IsLeader()).To(BeFalse()) - }) - - It("new leader keeps activeSince across a leadership move", func() { - a.ReportFailure("x", errBoom) - before, _ := b.ChainStatus("chain") - leaderIsA = false - clock.Advance(5 * time.Second) - a.Tick(ctx); b.Tick(ctx) - after, _ := b.ChainStatus("chain") - Expect(after.ActiveSince).To(Equal(before.ActiveSince)) - Expect(b.IsLeader()).To(BeTrue()) - }) - - It("republish sends every target and chain", func() { - c := New(src, WithClock(clock), WithLeaderGate(gateFor(false))) - bus.add(c) - c.SetStateSync(bus) - a.ReportFailure("x", errBoom) - a.Republish() - st, _ := c.ChainStatus("chain") - Expect(st.Active).To(Equal("y")) - }) - - It("standalone manager (no sync, no gate) is always leader", func() { - m := New(src, WithClock(clock)) - m.Tick(ctx) - Expect(m.IsLeader()).To(BeTrue()) - }) -}) -``` - -(`fakeProber` and its `take()` exist in `schedule_test.go`; if the in-flight probe tracking needs waiting, use `Eventually` as shown.) - -- [ ] **Step 2: Run to verify failure** - -Run: `go test ./core/services/failover/... 2>&1 | tail -5` -Expected: compile failure (`undefined: TargetSnapshot`). - -- [ ] **Step 3: Implement** `statesync.go` (types, interface, `WithLeaderGate`, `SetStateSync`, `Apply*`, `IsLeader`, `Republish`, `unlockAndFlush`) and the hooks in `manager.go`/`schedule.go` following the design rules above. Add fields to `Manager`: `sync StateSync`, `gate LeaderGate`, `leader bool` (guarded by `mu`), `applying bool`, `pending []func()`, `ticks int`; to `chainState`: `adopted bool`. Keep every existing test green. - -- [ ] **Step 4: Run tests** - -Run: `go test -race -count=3 ./core/services/failover/... 2>&1 | tail -10` -Expected: PASS, no races. - -- [ ] **Step 5: Commit** - -```bash -git add core/services/failover -git commit -m "feat(failover): share state through a sync hook and gate probes on a leader - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 2: syncstate-backed StateSync with a durable pin store - -**Files:** -- Create: `core/services/failover/distsync/distsync.go`, `core/services/failover/distsync/pinstore.go`, `core/services/failover/distsync/distsync_suite_test.go`, `core/services/failover/distsync/distsync_test.go` - -**Interfaces:** -- Consumes: Task 1 `StateSync`, `TargetSnapshot`, `ChainSnapshot`, `Manager.ApplyTarget/ApplyChain/ApplyPin`; `syncstate.New/Config/Store`; `messaging.MessagingClient`; `advisorylock.WithLockCtx`, `advisorylock.KeySchemaMigrate`; `testutil.NewFakeBus`. -- Produces: - ```go - type PinRecord struct { - Chain string `gorm:"primaryKey" json:"chain"` - Target string `json:"target"` - UpdatedAt time.Time `json:"updated_at"` - } - func (PinRecord) TableName() string { return "failover_pins" } - func NewPinStore(db *gorm.DB) (*PinStore, error) // migrates under KeySchemaMigrate - func New(ctx context.Context, nats messaging.MessagingClient, pins syncstate.Store[string, PinRecord], m *failover.Manager) (*Sync, error) - func (s *Sync) Close() error - // *Sync implements failover.StateSync - ``` - -Rules: -- Three maps: `failover.pins` (key chain, `Store: pins` when non-nil — guard against a typed-nil interface as `finetune/service.go` does), `failover.targets` (key target, NATS only), `failover.chains` (key chain, NATS only). No `Reconcile` on NATS-only maps (it would re-hydrate them empty); the leader's `Republish` covers late joiners. -- `OnApply` for targets/chains calls `m.ApplyTarget`/`m.ApplyChain`; for pins `m.ApplyPin(chain, target)` on "set" and `m.ApplyPin(chain, "")` on "delete". -- `New` starts all three maps, then calls `m.SetStateSync(s)` (which hydrates pins from `s.Pins()`). - -- [ ] **Step 1: Write the failing tests** — two managers on one `testutil.NewFakeBus()`, each with its own `distsync.New`, sharing an in-memory `syncstate.Store` fake for pins (a map with a mutex implementing `List/Upsert/Delete`). Specs: - - "a pin on A is visible on B and survives a new instance C built from the same store" (C's `ChainStatus` shows the pin right after `New`). - - "a trip on B makes A's plan skip the target". - - "the leader's chain switch reaches the follower". - - "late joiner converges on heartbeat": build C after A tripped a target; before `A.Republish()` C shows the primary active; after it, the fallback. - - "PinStore round-trips" using a sqlite gorm DB (`gorm.io/driver/sqlite`, `file::memory:`) if the repo already depends on it (check `go.mod`); otherwise skip this spec and test the adapter with the in-memory store only, and say so in the report. - -- [ ] **Step 2: Run to verify failure** — `go test ./core/services/failover/distsync/...` → compile failure. -- [ ] **Step 3: Implement** `distsync.go` and `pinstore.go` per the rules (PinStore: `List` = `db.Find`, `Upsert` = `db.Save`, `Delete` = `db.Delete(&PinRecord{Chain: k})`; migration via `advisorylock.WithLockCtx(context.Background(), db, advisorylock.KeySchemaMigrate, func() error { return db.AutoMigrate(&PinRecord{}) })`). -- [ ] **Step 4: Run tests** — `go test -race ./core/services/failover/...` → PASS. -- [ ] **Step 5: Commit** - -```bash -git add core/services/failover/distsync -git commit -m "feat(failover): sync pins, target health and chain state over NATS - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 3: Distributed wiring, warm pins on workers, docs - -**Files:** -- Modify: `core/services/advisorylock/keys.go` (add `KeyFailoverProber = 108` with a comment) -- Modify: `core/application/startup.go` (after `application.distributed = distSvc`, ~L294, and before `Run` ~L570), `core/application/distributed.go` (~L380-395, L458: pinned resolver), `core/application/failover.go` (preload only when leader) -- Create: `core/application/failover_distributed.go`, `core/application/failover_distributed_test.go` -- Modify: `docs/content/features/model-failover.md` (replace the per-instance limit with a "Distributed mode" section) - -**Interfaces:** -- Consumes: Tasks 1-2; `advisorylock.TryWithLockCtx(ctx, db, key, fn func() error) (bool, error)`; `a.IsDistributed()`, `a.Distributed().Nats`, `a.distributedDB()`; `nodes.PinnedModelResolver` (`GetPinnedModelNames() []string`); `failover.MergePinned`. -- Produces: `func failoverLeaderGate(db *gorm.DB) failover.LeaderGate`; `type failoverPinnedResolver struct{ base nodes.PinnedModelResolver; fm *failover.Manager }` implementing `GetPinnedModelNames()`. - -Rules: -- Leader gate: `func(ctx, fn) bool { ok, err := advisorylock.TryWithLockCtx(ctx, db, advisorylock.KeyFailoverProber, func() error { fn(); return nil }); if err != nil { xlog.Warn(...); return false }; return ok }`. The failover manager is constructed before distributed init (startup.go ~L257), so give it the gate through a setter `(*Manager).SetLeaderGate(LeaderGate)` added in this task (mirror of `WithLeaderGate`), then call `distsync.New(ctx, distSvc.Nats, pinStore, application.failoverManager)`; on error log and continue standalone. -- Pinned resolver: wrap `configLoader` so `GetPinnedModelNames()` returns `failover.MergePinned(configLoader.GetPinnedModelNames(), fm.WarmTargets())`, and pass it to both the SmartRouter and the ReplicaReconciler options. Adapt `initDistributed`'s signature minimally (add a `pinned nodes.PinnedModelResolver` parameter or set it after construction — choose the smaller change and explain). -- Warm preload: `applyFailoverWarmTargets` keeps the pin sync, and runs the preload goroutine only when `!a.IsDistributed() || a.failoverManager.IsLeader()`. -- `Run` calls `Republish()` every 10 ticks when leader (implemented in Task 1; verify here). - -- [ ] **Step 1: Write the failing tests** (`failover_distributed_test.go`, in the application package's existing test suite): - - `failoverPinnedResolver` merges config pins and warm targets without duplicates. - - `failoverLeaderGate` with a gorm DB that is not PostgreSQL (the advisorylock package falls back to an in-process lock for non-Postgres DBs): two gates on the same DB — while one is inside `fn`, the other returns false; afterwards the other returns true. Use the same DB setup other `core/application` or `advisorylock` tests use (read `core/services/advisorylock/*_test.go` first). -- [ ] **Step 2: Run to verify failure.** -- [ ] **Step 3: Implement** per the rules; add the `keys.go` constant; update docs: in `model-failover.md`, replace the "Failover state is kept in memory by each LocalAI instance" limit with a `## Distributed mode` section: pins are cluster-wide and persisted; target health and chain state are shared; one frontend (the probe leader) probes, fails back and preloads warm targets; warm targets are pinned on workers; a frontend that joins late converges within 10 s. -- [ ] **Step 4: Run tests** — `go test -race ./core/application/... ./core/services/failover/... ./core/services/advisorylock/...` and `go build ./core/... ./pkg/... ./tests/...` → PASS. -- [ ] **Step 5: Commit** - -```bash -git add core/services/advisorylock core/application core/services/failover docs/content/features/model-failover.md -git commit -m "feat(failover): run one prober per cluster and pin warm targets on workers - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -## Part A — localai-proxy backend - -### Task 4: Core support — Rerank for Go backends, Unimplemented skip, proxy options - -**Files:** -- Modify: `pkg/grpc/interface.go` (add `RerankModel`), `pkg/grpc/server.go` (add `Rerank` handler), `pkg/grpc/model_identity_modalities_test.go` (update the note/test that says Go servers have no Rerank) -- Modify: `core/services/failover/classify.go` (add `IsCapabilityGap`), `core/services/failover/manager.go` (`Do` uses it), `core/http/middleware/failover.go` (retry uses it) -- Modify: `core/backend/options.go:545` (proxy options for `localai-proxy`), `core/config/model_config_loader.go` (load-time warnings) -- Test: `pkg/grpc/server_rerank_test.go` (or the package's existing server test file), `core/services/failover/classify_test.go`, `core/services/failover/manager_test.go`, `core/http/middleware/failover_test.go`, `core/config/model_config_loader_test.go` - -**Interfaces:** -- Produces: - ```go - // pkg/grpc/interface.go - type RerankModel interface { - Rerank(context.Context, *pb.RerankRequest) (*pb.RerankResult, error) - } - // core/services/failover/classify.go - func IsCapabilityGap(err error) bool // gRPC Unimplemented anywhere in the chain - ``` - -- [ ] **Step 1: Write the failing tests** - - `pkg/grpc`: a fake model embedding `base.Base` and implementing `RerankModel` is served by `server.Rerank` (use the package's existing in-process server test pattern for `Score`); a model without it returns `codes.Unimplemented`. - - `classify_test.go`: `IsCapabilityGap(grpcstatus.Error(codes.Unimplemented, "x"))` true; wrapped with `fmt.Errorf("%w")` true; `codes.Unavailable` false; nil false. - - `manager_test.go` ("Unimplemented skips without trip"): `m.Do` where target `a` returns `grpcstatus.Error(codes.Unimplemented, "localai-proxy: X has no upstream counterpart")` and `b` succeeds → `tried == [a b]`, `a` stays `StateHealthy`. - - `failover_test.go` (middleware): handler for `a` returns the same Unimplemented error → served by `b`, `a` healthy. - - `model_config_loader_test.go`: a `backend: localai-proxy` config with `proxy.mode: translate` logs a warning and still loads; one without `known_usecases` loads (warning). Use the loader test's existing log-capture approach if any; otherwise assert only that loading succeeds and note it. -- [ ] **Step 2: Run to verify failure.** -- [ ] **Step 3: Implement** - ```go - // pkg/grpc/server.go — copy of the Score handler shape - func (s *server) Rerank(ctx context.Context, in *pb.RerankRequest) (*pb.RerankResult, error) { - if err := s.checkModelIdentity(in); err != nil { - return nil, err - } - rm, ok := s.llm.(RerankModel) - if !ok { - return nil, status.Errorf(codes.Unimplemented, "method Rerank not implemented") - } - if s.llm.Locking() { - s.llm.Lock() - defer s.llm.Unlock() - } - return rm.Rerank(ctx, in) - } - ``` - (If `checkModelIdentity` does not accept `*pb.RerankRequest`, add the `GetModelIdentity()` method set it needs — `RerankRequest` has field `ModelIdentity`.) - ```go - // classify.go - // IsCapabilityGap reports a target that cannot serve this kind of request at - // all. The next target may serve it, and this target is not broken. - func IsCapabilityGap(err error) bool { - if err == nil { - return false - } - st, ok := grpcstatus.FromError(err) - return ok && st.Code() == codes.Unimplemented - } - ``` - In `Manager.Do`, before the retryable check: `if IsCapabilityGap(err) && !committed.Load() { if !att.Skip() { return err }; continue }`. In `failoverRetry`, treat `failover.IsCapabilityGap(err)` exactly like the admission-rejection branch (spill with `att.Skip()`, no trip). In `options.go` change the condition to `if c.Backend == "cloud-proxy" || c.Backend == "localai-proxy"`. In the loader's load-time pass, `xlog.Warn` for `localai-proxy` configs that set `proxy.mode`/`proxy.provider` (ignored) or have no `known_usecases`. -- [ ] **Step 4: Run tests** — `go test -race ./pkg/grpc/... ./core/services/failover/... ./core/http/middleware/... ./core/config/...` → PASS. -- [ ] **Step 5: Commit** - -```bash -git add pkg/grpc core/services/failover core/http/middleware core/backend/options.go core/config -git commit -m "feat(grpc): serve Rerank from Go backends and skip Unimplemented targets - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 5: localai-proxy backend — skeleton, packaging, text methods - -**Files:** -- Create: `backend/go/localai-proxy/{main.go,proxy.go,client.go,text.go,Makefile,package.sh,run.sh}`, `backend/go/localai-proxy/{localai_proxy_suite_test.go,fake_upstream_test.go,text_test.go}` -- Modify: root `Makefile` (6 places, mirroring cloud-proxy: `.NOTPARALLEL`, `TEST_PATHS`, `BACKEND_LOCALAI_PROXY = localai-proxy|golang|.|false|true`, `$(eval $(call generate-docker-build-target,$(BACKEND_LOCALAI_PROXY)))`, `docker-build-localai-proxy` in `docker-build-backends`, `build-localai-proxy-backend`/`clean-localai-proxy-backend` e2e helpers building to `tests/e2e/mock-backend/localai-proxy`), `backend/index.yaml` (meta + cpu/metal latest/development images, mirroring cloud-proxy), `.github/backend-matrix.yml` (amd64, arm64, darwin entries mirroring cloud-proxy), `.gitignore` (the e2e binary) - -**Interfaces:** -- Consumes: `pkg/grpc` (`StartServer`, `AIModelRich`, `RerankModel`, `ScoreModel`), `pkg/grpc/base.Base`, `pkg/httpclient.New`. -- Produces: `type LocalAIProxy struct { base.Base; cfg atomic.Pointer[proxyConfig]; client *http.Client }`, `func NewLocalAIProxy() *LocalAIProxy`, `type proxyConfig struct { base, upstreamModel, apiKey, realtimePipeline string; timeout time.Duration }`, helpers `func (p *LocalAIProxy) postJSON(ctx, path string, body, out any) error`, `func (p *LocalAIProxy) postMultipart(ctx, path string, fields map[string]string, fileField, filePath string, out any) error`, `func (p *LocalAIProxy) postStream(ctx, path string, body any) (*http.Response, error)`, `func (p *LocalAIProxy) model(req string) string` (upstream model name), `func unimplemented(method string) error` returning `status.Errorf(codes.Unimplemented, "localai-proxy: %s has no upstream counterpart", method)`. - -Rules: -- `Load`: requires `opts.GetProxy()` non-nil and a valid `upstream_url` (base URL; strip a trailing `/`); resolves the API key with the same rules as cloud-proxy's `resolveAPIKey` (copy the function); warns if `mode`/`provider` set; reads `realtime_pipeline:` from `opts.GetOptions()`; `request_timeout_seconds` becomes a per-request context timeout for non-streaming calls; model name = `upstream_model`, else `opts.GetModel()`. -- Auth: `Authorization: Bearer ` when a key is set. HTTP client from `httpclient.New()` (no redirects). -- Upstream non-2xx → a gRPC error: 5xx/transport → `codes.Unavailable`; 4xx → `codes.InvalidArgument` (so failover does not trip on client errors); body text in the message (truncated to 500 chars). -- Text methods in `text.go`: `PredictRich`/`PredictStreamRich` via `/v1/chat/completions` when `opts.GetMessages()` is non-empty, else `/v1/completions` with `opts.GetPrompt()` (map tokens, temperature, top_p, top_k, stop, seed; stream parses SSE `data:` lines and sends `pb.Reply{Message}` per delta; do not close the channel); legacy `Predict`/`PredictStream` wrap them; `Embeddings` → `/v1/embeddings` (`input` = `opts.GetEmbeddings()`, returns `data[0].embedding`); `Rerank` → `/v1/rerank`; `TokenizeString` → `/v1/tokenize`; `Score` → `/api/score` (read `core/http/endpoints` for the request shape). -- Methods with no counterpart return `unimplemented("")`: `AudioEncode`, `AudioDecode`, `AudioToAudioStream`, `TokenClassify`, `ModelMetadata`, fine-tune and quantization methods. `Status` keeps the base implementation. - -- [ ] **Step 1: Write the failing tests** — `fake_upstream_test.go`: an `httptest.Server` recording method, path, `Authorization`, and JSON/multipart body, with per-path scripted responses (JSON or SSE). `text_test.go` specs: Load rejects missing proxy options and a bad URL; Load parses `realtime_pipeline`; `PredictRich` hits `/v1/chat/completions` with the upstream model and returns the content; `PredictStreamRich` streams SSE deltas in order; `Embeddings`, `Rerank`, `TokenizeString` hit their paths and map results; a 503 upstream → `codes.Unavailable`; a 400 → `codes.InvalidArgument`; `AudioEncode` → `Unimplemented` with the exact message. -- [ ] **Step 2: Run to verify failure** — `go test ./backend/go/localai-proxy/...`. -- [ ] **Step 3: Implement** per the rules; packaging per the Files list (copy cloud-proxy's `Makefile`/`package.sh`/`run.sh` with the binary renamed). -- [ ] **Step 4: Run tests** — `go test -race ./backend/go/localai-proxy/...`; `make -C backend/go/localai-proxy build`; `make build-localai-proxy-backend` → OK. Validate YAML: `python3 -c "import yaml,sys; yaml.safe_load(open('backend/index.yaml')); yaml.safe_load(open('.github/backend-matrix.yml'))"`. -- [ ] **Step 5: Commit** - -```bash -git add backend/go/localai-proxy Makefile backend/index.yaml .github/backend-matrix.yml .gitignore -git commit -m "feat(localai-proxy): add a backend that serves text APIs from a remote LocalAI - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 6: localai-proxy — audio methods - -**Files:** -- Create: `backend/go/localai-proxy/audio.go`, `backend/go/localai-proxy/audio_test.go` - -**Interfaces:** consumes Task 5 helpers. - -Mapping (request → upstream → result): - -| Method | Upstream | Notes | -|---|---|---| -| `TTS(req)` | `POST /tts` JSON `{model, input: req.Text, voice, language}` | write the response bytes to `req.Dst` | -| `TTSStream(req, out)` | `POST /tts` with `stream: true` | copy the chunked `audio/wav` body to `out` as it arrives (header + PCM, unchanged); close `out` per the base contract | -| `SoundGeneration(req)` | `POST /v1/sound-generation` | write bytes to `req.Dst` | -| `AudioTranscription(ctx, req)` | `POST /v1/audio/transcriptions` multipart: `file` from `req.Dst` (input audio path), `model`, `language`, `translate`, `prompt`, `diarize` | map `TranscriptionResult{text, segments, words, language, duration}` to `pb.TranscriptResult` | -| `AudioTranscriptionStream(ctx, req, out)` | same, plus `stream=true` | SSE `transcript.text.delta` → `TranscriptStreamResponse{Delta}`; `transcript.text.done` → `FinalResult`; `error` event → return `codes.Unavailable` | -| `Diarize` | `POST /v1/audio/diarization` multipart | map segments | -| `VAD(req)` | `POST /v1/vad` JSON `{model, audio: req.Audio}` | map `segments[{start,end}]` | -| `SoundDetection(ctx, req)` | `POST /v1/audio/classification` multipart `file` from `req.Src`, `top_k`, `threshold` | map `detections[{index,label,score}]` | -| `AudioTransform` | `POST /audio/transformations` | read the handler for the request shape; write output to the path the request carries | - -Read each `pb` request/response message in `backend/backend.proto` and each REST schema in `core/schema` before mapping; keep field names exact. - -- [ ] **Step 1: Write the failing tests** — one spec per row using the fake upstream: path, multipart fields (file bytes equal the input file), JSON body, and result mapping; `TTS` writes the upstream bytes to `Dst`; `TTSStream` forwards chunks in order and the first chunk starts with `RIFF`; `AudioTranscriptionStream` emits deltas then the final result; upstream disconnect mid-stream → `codes.Unavailable`. -- [ ] **Step 2: Verify failure.** **Step 3: Implement.** **Step 4:** `go test -race ./backend/go/localai-proxy/...` → PASS. -- [ ] **Step 5: Commit** — `feat(localai-proxy): serve speech, transcription and audio APIs remotely`. - ---- - -### Task 7: localai-proxy — image, video, 3D and vision methods - -**Files:** -- Create: `backend/go/localai-proxy/media.go`, `backend/go/localai-proxy/media_test.go` - -Mapping: - -| Method | Upstream | Notes | -|---|---|---| -| `GenerateImage(req)` | `POST /v1/images/generations` `{model, prompt, negative_prompt, size: "WxH", step, seed, response_format: "b64_json"}`; `req.Src`/`ref_images` sent as base64 in `files`/`ref_images` per `core/schema` | decode `data[0].b64_json` into `req.Dst` | -| `UpscaleImage` | `POST /v1/images/upscale` | same output handling | -| `GenerateVideo` | `POST /video` | write the returned file (b64 or URL download relative to the upstream base) to `Dst` | -| `Generate3D`, `Animate3D` | `POST /3d/generations`, `/3d/animate` | same output handling; implement `AnimationMetadataModel` only if the upstream response carries the metadata | -| `Detect`, `Depth` | `POST /v1/detection`, `/v1/depth` | map results | -| `FaceVerify`, `FaceAnalyze` | `POST /v1/face/verify`, `/v1/face/analyze` | map results | -| `VoiceVerify`, `VoiceAnalyze`, `VoiceEmbed` | `POST /v1/voice/verify`, `/v1/voice/analyze`, `/v1/voice/embed` | map results | -| `StoresSet/Get/Delete/Find` | `POST /stores/set`, `/stores/get`, `/stores/delete`, `/stores/find` | map keys/values | - -When the upstream returns a URL instead of b64, download it with the same client (same auth) and write it to `Dst`. - -- [ ] **Step 1: Failing tests** — one spec per row with the fake upstream (path, key request fields, `Dst` written from b64 and from a URL). **Step 2** verify failure. **Step 3** implement. **Step 4** `go test -race ./backend/go/localai-proxy/...` → PASS. -- [ ] **Step 5: Commit** — `feat(localai-proxy): serve image, video, 3D and vision APIs remotely`. - ---- - -### Task 8: localai-proxy — live transcription bridge - -**Files:** -- Create: `backend/go/localai-proxy/live.go`, `backend/go/localai-proxy/live_test.go` - -**Interfaces:** `func (p *LocalAIProxy) AudioTranscriptionLive(in <-chan *pb.TranscriptLiveRequest, out chan<- *pb.TranscriptLiveResponse) error`. Contract (from `pkg/grpc/server.go:477-541`): the backend closes `out`; `in` closes on client EOF; return errors immediately (callers wait for the ready ack). - -Protocol (upstream `github.com/gorilla/websocket`): -1. No `realtime_pipeline` → `close(out); return grpcerrors.LiveTranscriptionUnsupported("localai-proxy", "set the realtime_pipeline backend option")` (same helper `base.Base` uses). -2. Read the first `in` message; it must be `config` (else `codes.InvalidArgument`). Rate = `config.sample_rate` or 16000. -3. Dial `ws(s):///v1/realtime?model=` with the bearer header. Read `session.created`. -4. Send `{"type":"session.update","session":{"type":"transcription","audio":{"input":{"format":{"type":"audio/pcm","rate":},"transcription":{"model":"","language":""},"turn_detection":{"type":"server_vad"}}}}}`. On `session.updated` send `TranscriptLiveResponse{Ready: true}`; on `error` return `codes.Unavailable` with its message. -5. Writer goroutine: each `audio.pcm` (float32 in [-1,1]) → PCM16 LE → base64 → `{"type":"input_audio_buffer.append","audio":"..."}`. -6. Reader goroutine: `conversation.item.input_audio_transcription.delta` → `{Delta}`; `...completed` → `{Delta: , Eou: true}` and append to the running final text; `...failed` or `error` → end with `codes.Unavailable`. -7. When `in` closes: wait up to 5 s for an in-flight `completed` (track `speech_started` without a matching completion), then send `{FinalResult: {Text: }}`, close the socket, close `out`, return nil. -8. Socket read error before step 7 → close `out`, return `status.Error(codes.Unavailable, ...)`. - -- [ ] **Step 1: Failing tests** with an `httptest` WebSocket server (gorilla `Upgrader`) scripting the upstream: unsupported without the option; ready after `session.updated`; audio frames arrive base64-PCM16 of the right length; deltas and a completion map to `Delta`/`Eou`; closing `in` yields `FinalResult` with the concatenated text; "upstream disconnect ends the stream" (server closes mid-session → `Unavailable`, `out` closed, no goroutine left blocked — assert the call returns within 2 s). -- [ ] **Step 2** verify failure. **Step 3** implement. **Step 4** `go test -race ./backend/go/localai-proxy/...` → PASS. -- [ ] **Step 5: Commit** — `feat(localai-proxy): bridge live transcription to the upstream realtime API`. - ---- - -### Task 9: localai-proxy end to end, and docs - -**Files:** -- Modify: `tests/e2e/e2e_suite_test.go` (build/locate the `localai-proxy` binary like cloud-proxy), create `tests/e2e/e2e_localai_proxy_test.go`, extend `tests/e2e/realtime_ws_test.go` (label `failover`) -- Modify: docs — the page that documents `cloud-proxy` (find with `grep -rln "cloud-proxy" docs/content`) gets a `localai-proxy` section; `docs/content/features/model-failover.md` gets a remote-LocalAI example with per-stage chains. - -Rules: -- Point `localai-proxy` models at the test server itself (`proxy.upstream_url` = the suite's base URL, `upstream_model` = an existing mock model), registered at runtime the way the cloud-proxy/failover e2e specs register models. -- Specs: chat, embeddings, TTS and transcription through `localai-proxy` return 2xx with the upstream model's answer; a chain `[proxy-target, local mock]` where the proxy target's upstream model is `fail-load-…` fails over to the local target; a realtime pipeline whose `llm` stage is a chain `[localai-proxy → mock LLM, mock LLM]` completes a turn and, after the proxy target is made to fail (point it at a `fail-load` upstream model or stop routing), switches stage with a `localai.model.failover` event. Live transcription through the bridge is covered by Task 8's unit specs; add an e2e only if the suite has a pipeline with streaming transcription available. -- Run `make build-localai-proxy-backend build-mock-backend` first. - -- [ ] Steps: write specs → run (fail) → implement registration + docs → run `go run github.com/onsi/ginkgo/v2/ginkgo --label-filter=failover -v ./tests/e2e` and the full `!real-models` e2e → commit `test(localai-proxy): proxy APIs and realtime stages end to end`. - ---- - -## Part B — WebUI - -### Task 10: Failover data layer and health strip - -**Files:** -- Modify: `core/http/react-ui/src/utils/config.js` (endpoints), `src/utils/api.js` (`failoverApi`), `src/components/StatusPill.jsx` (tones) -- Create: `src/hooks/useFailoverChains.js`, `src/components/FailoverChainStatus.jsx` -- Modify: `src/pages/ModelEditor.jsx` (strip after the `me-head` block, before the template selector, when `!isCreateMode` and the chain exists), `src/App.css` (classes), `public/locales/*/models.json` (strings, 8 locales) -- Test: `e2e/failover-health.spec.js` - -**Interfaces — Produces:** -```js -// utils/config.js endpoints -failoverChains: '/api/failover', -failoverChain: (name) => `/api/failover/${encodeURIComponent(name)}`, -failoverEvents: '/api/failover/events', -failoverPin: (name) => `/api/failover/${encodeURIComponent(name)}/pin`, -// utils/api.js -export const failoverApi = { - list: () => fetchJSON(API_CONFIG.endpoints.failoverChains), - get: (name) => fetchJSON(API_CONFIG.endpoints.failoverChain(name)), - pin: (name, target) => postJSON(API_CONFIG.endpoints.failoverPin(name), { target }), - unpin: (name) => fetchJSON(API_CONFIG.endpoints.failoverPin(name), { method: 'DELETE' }), - eventsUrl: () => API_CONFIG.endpoints.failoverEvents, -} -// hooks/useFailoverChains.js -export default function useFailoverChains() // → { chains: ChainStatus[], byName: {[name]: ChainStatus}, loading, error, refresh } -``` -Hook rules: initial `failoverApi.list()`; `new EventSource(apiUrl(failoverApi.eventsUrl()))`; `snapshot` → replace `chains` from `data.chains`; `chain.switched` → patch that chain's `active`, `state`, `active_since = data.at`; `target.state` → patch that target's `state`, `last_error = data.error` in every chain containing it; `onerror` no-op; poll `list()` every 15 s; close on unmount. - -`StatusPill` STATUS map additions: `primary: 'success'`, `fallback: 'warning'`, `recovering: 'warning'`, `degraded: 'error'`, `down: 'error'`, `missing: 'muted'` (`healthy` already maps to success). - -`FailoverChainStatus({ chain, onPin, onUnpin, canPin })` renders: chain `StatusPill` + active target + relative "since"; a compact table of targets (model, kind, warm, `StatusPill` state, last probe, last error truncated with title); per-target **Pin** button and a chain-level **Unpin** when `canPin`, each behind `ConfirmDialog`. Classes only. - -- [ ] **Step 1: Failing Playwright spec** `e2e/failover-health.spec.js` (import `test` from `./coverage-fixtures.js`; mock `**/api/auth/status`, config metadata, `**/api/failover` with one chain `chain` [a healthy active, b healthy], `**/api/failover/chain`, and `**/api/failover/events` fulfilled with `text/event-stream` body `event: snapshot\ndata: {"chains":[...]}\n\nevent: chain.switched\ndata: {"chain":"chain","from":"a","to":"b","state":"fallback","reason":"trip","at":"2026-09-26T10:00:00Z"}\n\n`). Assertions: `/app/model-editor/chain` shows the strip with the chain state; after the event the active target is `b` and the pill reads fallback; the Pin button is visible with auth disabled (admin) and hidden when auth status reports a non-admin user (mock `/api/auth/me` the way `users-tab-gating.spec.js` does); clicking Pin + confirm POSTs `{target}`. -- [ ] **Step 2** run `cd core/http/react-ui && npx playwright test e2e/failover-health.spec.js` (after `bun run build` if the harness serves the build; follow the repo's UI test instructions in `Makefile` `test-ui*` targets) → FAIL. -- [ ] **Step 3** implement; **Step 4** re-run → PASS; `npm run lint` and `npm run lint:inline-styles` (or the scripts in `package.json`) clean. -- [ ] **Step 5: Commit** — `feat(ui): show live failover chain health in the model editor`. - ---- - -### Task 11: Chain editor field and template - -**Files:** -- Create: `core/http/react-ui/src/components/FailoverTargetsEditor.jsx` -- Modify: `src/components/ConfigFieldRenderer.jsx` (branch `component === 'failover-targets'`, same `list-row` wrapper as `router-candidates`), `src/utils/modelTemplates.js` (template), `src/pages/ModelEditor.jsx` (`SECTION_ICONS.failover = 'fa-shuffle'`, `SECTION_COLORS.failover = 'var(--color-accent)'`), `core/config/meta/registry.go` (`failover.targets` `Component: "failover-targets"`), `core/config/meta/registry_test.go` (assert the component), `public/locales/*/modelEditor.json` -- Test: `e2e/failover-editor.spec.js` - -`FailoverTargetsEditor({ value, onChange })`: modelled on `RouterCandidatesEditor` — items `{model, warm}`; row = `SearchableModelSelect` (value/onChange), `Toggle` for warm (disabled with a title when the selected model's backend is a proxy — look it up from `useModels()` data if it carries the backend; otherwise leave enabled and rely on the load warning, and say so), move up/down, remove; "Add target" button; inline errors (fewer than 2 targets, duplicate model, the edited model's own name) from `useFormContext()` `formData.name`. - -Template entry: -```js -{ - id: 'failover', - label: 'Failover Chain', - icon: 'fa-shuffle', - description: 'Serve one model name from an ordered list of models. The first healthy one answers; the next takes over when it fails.', - fields: { - 'name': '', - 'failover.targets': [{ model: '' }, { model: '' }], - }, -}, -``` - -- [ ] Steps: failing spec (template card visible; `?template=failover` shows two target rows; adding/removing/moving rows; duplicate and too-few errors; saving sends `failover.targets` in the PATCH/import body — mock the save endpoint and assert the JSON) → implement → `npx playwright test e2e/failover-editor.spec.js` PASS, `go test ./core/config/meta/...` PASS → commit `feat(ui): edit failover chain targets with a dedicated field`. - ---- - -### Task 12: Chain badge and Failover overview page - -**Files:** -- Create: `core/http/react-ui/src/pages/Failover.jsx` -- Modify: `src/pages/InstalledModels.jsx` (chain badge next to the alias badge: `badge badge-info`, icon `fa-shuffle`, text `chain → `; data from `useFailoverChains()`), `src/router.jsx` (`const Failover = page('failover', () => import('./pages/Failover'))`; route `{ path: 'failover', element: }`), `src/components/console/consoleConfig.js` (`operate.runtime` item `{ path: '/app/failover', icon: 'fas fa-shuffle', labelKey: 'items.failover', adminOnly: true }`), `public/locales/*/nav.json`, `public/locales/*/models.json`, `src/App.css` -- Test: `e2e/failover-overview.spec.js` - -Overview: `useFailoverChains()`; a dense table (chain name link → `/app/model-editor/`, chain `StatusPill`, active target, one small pill per target, time since `active_since`); empty state with a link to `/app/model-editor?template=failover`. - -- [ ] Steps: failing spec (nav entry visible for admin; table rows from mocked `/api/failover`; SSE patch updates a row; empty state link; installed-models badge `chain → a`) → implement → specs PASS; run the full UI suite with coverage: `make test-ui-coverage-check` (UI coverage ≥ baseline) → commit `feat(ui): list failover chains and badge chain models`. - ---- - -## Part D — Contributor rule and final verification - -### Task 13: distributed-state rule, docs sweep, final verification - -**Files:** -- Create: `.agents/distributed-state.md` -- Modify: `AGENTS.md` (Topics table row; Quick Reference bullet), `.agents/api-endpoints-and-auth.md` (checklist line), `docs/content/features/model-failover.md` (UI section: editor, health strip, overview; confirm distributed section from Task 3) - -`.agents/distributed-state.md` content (write it in full, following the style of the other `.agents/*.md` guides): -- Title "Distributed-aware state". Why: frontends are stateless replicas; in-memory state diverges silently (the failover chains example). -- The rule: a feature that keeps runtime state (in-memory maps, caches, pins, schedulers, background loops, probes) chooses one mode and documents it on its docs page: - 1. **Shared** — `syncstate.SyncedMap` (`core/services/syncstate`); add a `Store` when the state must survive a restart. Example: finetune jobs (`core/services/finetune/service.go`), failover pins (`core/services/failover/distsync`). Gotcha: `Reconcile` without a `Store`/`Loader` re-hydrates the map empty — republish from a leader instead. - 2. **Single-runner** — `advisorylock.RunLeaderLoop` / `TryWithLockCtx` (`core/services/advisorylock`); new keys go in `keys.go`. Example: node health monitor (`core/services/nodes/health.go`), failover prober. - 3. **Stateless per request** — nothing to share. - 4. **Per-instance** — allowed only with the reason written in the feature's docs. -- Tests: shared and single-runner features include a two-instance test on `testutil.NewFakeBus()` (`core/services/testutil/fakebus.go`); note that the fake bus delivers synchronously, including the publisher's own message, so never publish while holding a lock the apply path takes. -- Checklist for PRs. - -AGENTS.md Quick Reference bullet: -`- **Distributed-aware state**: any feature that keeps runtime state (maps, caches, pins, schedulers, background loops) must choose shared (syncstate), single-runner (advisorylock), stateless, or documented per-instance behaviour for multi-frontend clusters. See [.agents/distributed-state.md](.agents/distributed-state.md).` - -AGENTS.md Topics row: -`| [.agents/distributed-state.md](.agents/distributed-state.md) | Features that keep runtime state — how they must behave with several frontends (syncstate, advisory-lock leaders, fakebus tests) |` - -api-endpoints-and-auth.md checklist line (under Quality): -`- [ ] Stateful feature: distributed mode chosen and documented (see [distributed-state.md](distributed-state.md))` - -- [ ] Steps: write the files → final verification, in order, reading each output: -```bash -make protogen-go build-mock-backend build-cloud-proxy-backend build-localai-proxy-backend -go vet ./core/services/failover/... ./backend/go/localai-proxy/... ./pkg/grpc/... -go test -race ./core/services/failover/... ./core/application/... ./core/config/... ./core/http/middleware/... ./pkg/grpc/... ./pkg/mcp/localaitools/... ./backend/go/localai-proxy/... ./core/http/endpoints/openai/... -LOCALAI_TEST_HTTP_PORT=19391 go test ./core/http/... -go run github.com/onsi/ginkgo/v2/ginkgo --label-filter='!real-models' -v ./tests/e2e -make test-ui-coverage-check -LOCALAI_TEST_HTTP_PORT=19391 make test-coverage-check -``` -Expected: all green; both coverage checks at or above baseline. Record any pre-existing failure (e.g. `make swagger`) with its output tail. -- [ ] Commit — `docs: require distributed-aware state for stateful features`. diff --git a/docs/superpowers/plans/2026-09-26-model-failover-chains.md b/docs/superpowers/plans/2026-09-26-model-failover-chains.md deleted file mode 100644 index a5a498d53..000000000 --- a/docs/superpowers/plans/2026-09-26-model-failover-chains.md +++ /dev/null @@ -1,5054 +0,0 @@ -# Model Failover Chains Implementation Plan - -> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. - -**Goal:** A model config can declare an ordered `failover` chain of target models; LocalAI serves each request from the highest-priority healthy target, retries uncommitted failures on the next target, probes targets, fails back with hysteresis, and publishes switch events. - -**Architecture:** A new `core/services/failover` package holds a `Manager` (per-target health state machine, per-chain active target, probes, event bus). The HTTP request middleware resolves a chain to a target the way it resolves an alias, and a retry wrapper inside `SetModelAndConfig` re-runs the request on the next target while the response is uncommitted. Realtime pipeline stages resolve chains per call through `Manager.Do`. REST, SSE, a realtime server event, metrics and MCP tools expose the state. - -**Tech Stack:** Go, echo v4, Ginkgo v2 + Gomega, OpenTelemetry metrics, gRPC backend interface (`pkg/grpc`). - -**Spec:** `docs/superpowers/specs/2026-09-26-model-failover-chains-design.md` - -## Global Constraints - -- Worktree: `/home/mudler/_git/LocalAI/.wt/failover-chains`, branch `feat/failover-chains`. All paths below are relative to it. -- Commit trailer: `Assisted-by: Claude:claude-opus-5-5`. Never add `Co-Authored-By` or `Signed-off-by` (the human adds the DCO sign-off). Subjects use the repo's conventional style, for example `feat(failover): ...`. -- `docs/superpowers/` is excluded by the local `.git/info/exclude`; add files there with `git add -f`. -- Logging: `github.com/mudler/xlog`. Use `any`, never `interface{}`. Comments explain why, not what. -- Defaults, exactly: probe interval `15s`, probe timeout `5s`, trip errors `1`, trip window `30s`, recovery probes `3`, min dwell `60s`. -- Response headers, exactly: `X-LocalAI-Served-Model`, `X-LocalAI-Failover` (values `fallback`, `degraded`). -- Event names, exactly: SSE `snapshot`, `chain.switched`, `target.state`; realtime `localai.model.failover`. Reasons: `trip`, `recovery`, `manual`, `degraded`, `missing`, `initial`. -- Metrics, exactly: `localai_failover_switches_total{chain,from,to,reason}`, `localai_failover_target_up{target}`. -- Coverage baseline (`coverage-baseline.txt`, 54.2) must not go down. Never edit the baseline. -- Docs change ships in the same PR (`docs/content/features/model-failover.md`). -- Tests need generated protos: run `make protogen-go` once in the worktree before the first `go test`. E2E tests need `make build-mock-backend`. -- Run a package's tests with: `go run github.com/onsi/ginkgo/v2/ginkgo -v ./` (or `go test .//...`). - -## Review Focus - -1. A handler that mutates the parsed request on attempt 1 must not leak that mutation into attempt 2: each attempt re-binds from the replayed body (Task 8, spec "gives each attempt a fresh request"). -2. A client that disconnects mid-request must not trip the target or trigger a retry (Task 8, spec "does not retry when the client cancelled"). -3. A 4xx (bad request, context overflow) must not retry and must not trip the target (Task 8, spec "does not retry or trip on 4xx"). -4. A degraded chain (all targets down) must try every target in priority order and return the last target's error (Task 8, spec "degraded"). -5. A multipart upload (transcription) must be readable again by the second attempt (Task 8, spec "replays a multipart body"). - ---- - -## File Structure - -| File | Responsibility | -|---|---| -| `core/config/model_config_failover.go` (new) | `FailoverConfig` types, defaults, `IsFailover`, `validateFailover`, `WarmFailoverTargets` | -| `core/config/model_config.go` (modify) | `Failover` field, call `validateFailover` from `Validate`, `ProxyConfig.ResolveAPIKey` | -| `core/config/model_config_loader.go` (modify) | `ValidateFailoverTargets`, load-time pruning and usecase warning | -| `core/config/meta/registry.go`, `types.go` (modify) | `failover` section and field metadata | -| `core/services/failover/types.go` (new) | states, reasons, events, status DTOs, `KindOf`, `MergePinned` | -| `core/services/failover/classify.go` (new) | `IsRetryable` | -| `core/services/failover/manager.go` (new) | `Manager`: sync, state machine, plan/attempt, pin, events, status, `Do` | -| `core/services/failover/schedule.go` (new) | `Prober` interface, `Run`, `Tick`, probe scheduling | -| `core/services/failover/prober.go` (new) | `DefaultProber`: remote HTTP and local gRPC probes | -| `core/services/failover/metrics.go` (new) | OTel counter and gauge | -| `core/services/failover/trace.go` (new) | `RecordAttemptTrace` | -| `core/trace/backend_trace.go` (modify) | `BackendTraceFailover` type | -| `core/application/{application.go,startup.go,watchdog.go,failover.go}` | wiring, warm pinning/preload | -| `core/http/middleware/failover.go` (new) | chain resolution, retry wrapper, writer, headers | -| `core/http/middleware/request.go` (modify) | call resolution, wrap `SetModelAndConfig` | -| `core/http/app.go` (modify) | `SetFailoverManager` | -| `core/http/endpoints/localai/failover.go` (new) | REST + SSE handlers | -| `core/http/routes/localai.go` (modify) | routes | -| `core/http/endpoints/localai/api_instructions.go` (modify) | instruction entry | -| `core/http/endpoints/openai/realtime_model.go`, `realtime.go`, `realtime_failover.go` (new), `types/failover.go` (new), `types/server_events.go` | realtime per-call resolution and events | -| `pkg/mcp/localaitools/...` | MCP tools | -| `tests/e2e/...` | e2e specs, mock-backend load-failure trigger, fake upstream `/v1/models` | -| `docs/content/features/model-failover.md` (new) + cross-links | docs | - ---- - -### Task 1: Config schema, per-config validation, field metadata - -**Files:** -- Create: `core/config/model_config_failover.go` -- Modify: `core/config/model_config.go` (field next to `Alias` at line ~76; `Validate()` at line ~1593, before the alias block at ~1650) -- Modify: `core/config/meta/types.go` (`DefaultSections()`), `core/config/meta/registry.go` (next to the `alias` entry at ~414) -- Test: `core/config/model_config_failover_test.go`, `core/config/meta/registry_test.go` - -**Interfaces:** -- Produces: `config.FailoverConfig`, `config.FailoverTarget`, `(ModelConfig).IsFailover() bool`, `(FailoverConfig).ProbeInterval()/ProbeTimeout()/TripWindow()/MinDwell() time.Duration`, `(FailoverConfig).TripErrors()/RecoveryProbes() int`, `(ModelConfig).WarmFailoverTargets() []string`, field `ModelConfig.Failover *FailoverConfig`. - -- [ ] **Step 1: Write the failing tests** - -`core/config/model_config_failover_test.go` (package `config`, like `model_config_test.go`): - -```go -package config - -import ( - "time" - - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" - "gopkg.in/yaml.v3" -) - -var _ = Describe("ModelConfig failover", func() { - chain := func(targets ...string) ModelConfig { - c := ModelConfig{Name: "chain", Failover: &FailoverConfig{}} - for _, t := range targets { - c.Failover.Targets = append(c.Failover.Targets, FailoverTarget{Model: t}) - } - return c - } - - It("parses the YAML block and applies defaults", func() { - var c ModelConfig - Expect(yaml.Unmarshal([]byte(` -name: assistant-llm -failover: - targets: - - model: argus-llm - - model: gemma-local - warm: true - recovery: - probes: 5 -`), &c)).To(Succeed()) - Expect(c.IsFailover()).To(BeTrue()) - Expect(c.Failover.Targets).To(Equal([]FailoverTarget{{Model: "argus-llm"}, {Model: "gemma-local", Warm: true}})) - Expect(c.Failover.ProbeInterval()).To(Equal(15 * time.Second)) - Expect(c.Failover.ProbeTimeout()).To(Equal(5 * time.Second)) - Expect(c.Failover.TripErrors()).To(Equal(1)) - Expect(c.Failover.TripWindow()).To(Equal(30 * time.Second)) - Expect(c.Failover.RecoveryProbes()).To(Equal(5)) - Expect(c.Failover.MinDwell()).To(Equal(60 * time.Second)) - Expect(c.WarmFailoverTargets()).To(Equal([]string{"gemma-local"})) - }) - - It("accepts a valid chain", func() { - c := chain("a", "b") - ok, err := c.Validate() - Expect(err).ToNot(HaveOccurred()) - Expect(ok).To(BeTrue()) - }) - - DescribeTable("rejects invalid chains", - func(mutate func(*ModelConfig), want string) { - c := chain("a", "b") - mutate(&c) - ok, err := c.Validate() - Expect(ok).To(BeFalse()) - Expect(err).To(MatchError(ContainSubstring(want))) - }, - Entry("alias and failover", func(c *ModelConfig) { c.Alias = "x" }, "both alias and failover"), - Entry("backend set", func(c *ModelConfig) { c.Backend = "llama-cpp" }, "must not set backend"), - Entry("one target", func(c *ModelConfig) { c.Failover.Targets = c.Failover.Targets[:1] }, "at least 2 targets"), - Entry("empty target", func(c *ModelConfig) { c.Failover.Targets[1].Model = "" }, "no model"), - Entry("self target", func(c *ModelConfig) { c.Failover.Targets[1].Model = "chain" }, "cannot list itself"), - Entry("duplicate target", func(c *ModelConfig) { c.Failover.Targets[1].Model = "a" }, "twice"), - Entry("bad duration", func(c *ModelConfig) { c.Failover.Probe.Interval = "soon" }, "invalid probe.interval"), - Entry("negative errors", func(c *ModelConfig) { c.Failover.Trip.Errors = -1 }, "trip.errors"), - Entry("no name", func(c *ModelConfig) { c.Name = "" }, "requires a name"), - ) -}) -``` - -In `core/config/meta/registry_test.go`, next to the alias assertions (lines 13-31), add: - -```go - It("registers the failover section", func() { - reg := meta.DefaultRegistry() - Expect(reg).To(HaveKey("failover.targets")) - Expect(reg["failover.targets"].Section).To(Equal("failover")) - var ids []string - for _, s := range meta.DefaultSections() { - ids = append(ids, s.ID) - } - Expect(ids).To(ContainElement("failover")) - }) -``` - -(Match the surrounding style: if the file uses `It` inside an existing `Describe`, put this `It` there.) - -- [ ] **Step 2: Run the tests to verify they fail** - -Run: `go test ./core/config/... 2>&1 | tail -20` -Expected: compile failure, `undefined: FailoverConfig`. - -- [ ] **Step 3: Implement the config types** - -`core/config/model_config_failover.go`: - -```go -package config - -import ( - "fmt" - "time" -) - -// FailoverConfig turns a model config into a failover chain: requests for the -// chain name are served by its highest-priority healthy target. See -// core/services/failover for the runtime side. -type FailoverConfig struct { - Targets []FailoverTarget `yaml:"targets" json:"targets"` - Probe FailoverProbe `yaml:"probe,omitempty" json:"probe,omitempty"` - Trip FailoverTrip `yaml:"trip,omitempty" json:"trip,omitempty"` - Recovery FailoverRecovery `yaml:"recovery,omitempty" json:"recovery,omitempty"` -} - -type FailoverTarget struct { - Model string `yaml:"model" json:"model"` - // Warm keeps a local target loaded and exempt from eviction, so a switch - // does not wait for a cold load. - Warm bool `yaml:"warm,omitempty" json:"warm,omitempty"` -} - -type FailoverProbe struct { - Interval string `yaml:"interval,omitempty" json:"interval,omitempty"` - Timeout string `yaml:"timeout,omitempty" json:"timeout,omitempty"` -} - -type FailoverTrip struct { - Errors int `yaml:"errors,omitempty" json:"errors,omitempty"` - Window string `yaml:"window,omitempty" json:"window,omitempty"` -} - -type FailoverRecovery struct { - Probes int `yaml:"probes,omitempty" json:"probes,omitempty"` - MinDwell string `yaml:"min_dwell,omitempty" json:"min_dwell,omitempty"` -} - -const ( - DefaultFailoverProbeInterval = 15 * time.Second - DefaultFailoverProbeTimeout = 5 * time.Second - DefaultFailoverTripErrors = 1 - DefaultFailoverTripWindow = 30 * time.Second - DefaultFailoverRecoveryProbes = 3 - DefaultFailoverMinDwell = 60 * time.Second -) - -// IsFailover reports whether this config is a failover chain. -func (c ModelConfig) IsFailover() bool { return c.Failover != nil } - -func (f FailoverConfig) ProbeInterval() time.Duration { - return durationOr(f.Probe.Interval, DefaultFailoverProbeInterval) -} -func (f FailoverConfig) ProbeTimeout() time.Duration { - return durationOr(f.Probe.Timeout, DefaultFailoverProbeTimeout) -} -func (f FailoverConfig) TripWindow() time.Duration { - return durationOr(f.Trip.Window, DefaultFailoverTripWindow) -} -func (f FailoverConfig) MinDwell() time.Duration { - return durationOr(f.Recovery.MinDwell, DefaultFailoverMinDwell) -} -func (f FailoverConfig) TripErrors() int { - if f.Trip.Errors <= 0 { - return DefaultFailoverTripErrors - } - return f.Trip.Errors -} -func (f FailoverConfig) RecoveryProbes() int { - if f.Recovery.Probes <= 0 { - return DefaultFailoverRecoveryProbes - } - return f.Recovery.Probes -} - -// WarmFailoverTargets returns the targets marked warm, in chain order. -func (c ModelConfig) WarmFailoverTargets() []string { - if c.Failover == nil { - return nil - } - var out []string - for _, t := range c.Failover.Targets { - if t.Warm { - out = append(out, t.Model) - } - } - return out -} - -func durationOr(s string, def time.Duration) time.Duration { - if s == "" { - return def - } - d, err := time.ParseDuration(s) - if err != nil || d <= 0 { - return def - } - return d -} - -// validateFailover checks what a chain can check without other configs. -// Target existence is checked by ModelConfigLoader.ValidateFailoverTargets. -func (c *ModelConfig) validateFailover() error { - if c.Name == "" { - return fmt.Errorf("failover config requires a name") - } - if c.IsAlias() { - return fmt.Errorf("model %q cannot set both alias and failover", c.Name) - } - if c.Backend != "" || c.Model != "" { - return fmt.Errorf("failover config %q must not set backend or parameters.model: a chain is a pure redirect", c.Name) - } - f := c.Failover - if len(f.Targets) < 2 { - return fmt.Errorf("failover chain %q needs at least 2 targets", c.Name) - } - seen := map[string]bool{} - for _, t := range f.Targets { - switch { - case t.Model == "": - return fmt.Errorf("failover chain %q has a target with no model", c.Name) - case t.Model == c.Name: - return fmt.Errorf("failover chain %q cannot list itself", c.Name) - case seen[t.Model]: - return fmt.Errorf("failover chain %q lists %q twice", c.Name, t.Model) - } - seen[t.Model] = true - } - for key, v := range map[string]string{ - "probe.interval": f.Probe.Interval, - "probe.timeout": f.Probe.Timeout, - "trip.window": f.Trip.Window, - "recovery.min_dwell": f.Recovery.MinDwell, - } { - if v == "" { - continue - } - if d, err := time.ParseDuration(v); err != nil || d <= 0 { - return fmt.Errorf("failover chain %q: invalid %s %q", c.Name, key, v) - } - } - if f.Trip.Errors < 0 { - return fmt.Errorf("failover chain %q: trip.errors must not be negative", c.Name) - } - if f.Recovery.Probes < 0 { - return fmt.Errorf("failover chain %q: recovery.probes must not be negative", c.Name) - } - return nil -} -``` - -In `core/config/model_config.go`, add the field directly after `Alias` (line ~76): - -```go - // Failover makes this config a failover chain over other models. Like an - // alias it has no backend of its own. - Failover *FailoverConfig `yaml:"failover,omitempty" json:"failover,omitempty"` -``` - -In `Validate()`, directly before the `if c.IsAlias() {` block (line ~1650), add: - -```go - if c.IsFailover() { - if err := c.validateFailover(); err != nil { - return false, err - } - return true, nil - } -``` - -If `Validate()` rejects configs without a backend earlier than line 1650 (the artifact check at ~1619 applies to aliases), move this block above that point so a chain returns before any backend-specific check. - -- [ ] **Step 4: Register field metadata** - -In `core/config/meta/types.go` `DefaultSections()`, add after the `alias` line: - -```go - {ID: "failover", Label: "Failover", Icon: "shuffle", Order: 6}, -``` - -(If `shuffle` is not an icon used elsewhere in `DefaultSections()`, reuse `git-merge`.) - -In `core/config/meta/registry.go`, after the `alias` entry, add: - -```go - // --- Failover --- - "failover.targets": { - Section: "failover", - Label: "Failover targets", - Description: "Ordered list of models that serve this chain. The first healthy target serves each request; later targets take over when it fails. Mark a local target warm to keep it loaded.", - Component: "json-editor", - Order: 0, - }, - "failover.probe.interval": { - Section: "failover", Label: "Probe interval", Component: "input", Order: 1, Advanced: true, - Description: "How often an idle target is checked, as a duration (default 15s).", Placeholder: "15s", - }, - "failover.probe.timeout": { - Section: "failover", Label: "Probe timeout", Component: "input", Order: 2, Advanced: true, - Description: "How long one probe may take (default 5s).", Placeholder: "5s", - }, - "failover.trip.errors": { - Section: "failover", Label: "Errors to trip", Component: "number", Order: 3, Advanced: true, - Description: "Failures within the trip window that mark a target down (default 1).", - }, - "failover.trip.window": { - Section: "failover", Label: "Trip window", Component: "input", Order: 4, Advanced: true, - Description: "Window in which failures are counted (default 30s).", Placeholder: "30s", - }, - "failover.recovery.probes": { - Section: "failover", Label: "Recovery probes", Component: "number", Order: 5, Advanced: true, - Description: "Consecutive real test requests a target must pass before it is used again (default 3).", - }, - "failover.recovery.min_dwell": { - Section: "failover", Label: "Minimum time on fallback", Component: "input", Order: 6, Advanced: true, - Description: "Minimum time on a lower target before traffic moves back to a recovered higher one (default 60s).", Placeholder: "60s", - }, -``` - -Run `go test ./core/config/meta/...`. `TestAllFieldsHaveRegistryEntries` lists every reflected path that has no entry. If it names paths different from the seven above (for example `failover.targets[].model`), register those exact paths with matching labels in the same section, and remove entries it reports as unknown. - -- [ ] **Step 5: Run the tests to verify they pass** - -Run: `go test ./core/config/... 2>&1 | tail -20` -Expected: PASS. - -- [ ] **Step 6: Commit** - -```bash -git add core/config -git commit -m "feat(config): add failover chain block to model configs - -A chain is a model config with an ordered list of target models, probe, -trip and recovery settings. Like an alias it has no backend. - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 2: Cross-config validation in the loader and admin paths - -**Files:** -- Modify: `core/config/model_config_loader.go` (new methods next to `ValidateAliasTarget` at ~487; load-time check after the alias check at ~824-840) -- Modify: `core/services/modeladmin/config.go` (next to the `ValidateAliasTarget` calls at ~169 and ~286), `core/http/endpoints/localai/import_model.go` (~186) -- Test: `core/config/model_config_loader_test.go` - -**Interfaces:** -- Consumes: Task 1 types. -- Produces: `(*ModelConfigLoader).ValidateFailoverTargets(cfg *ModelConfig) error`, `(*ModelConfigLoader).FailoverTargetsShareUsecase(cfg *ModelConfig) bool`. - -- [ ] **Step 1: Write the failing tests** - -Append to `core/config/model_config_loader_test.go` (the file seeds `loader.configs` directly, see line ~304): - -```go -var _ = Describe("ModelConfigLoader failover validation", func() { - var loader *ModelConfigLoader - chain := func(targets ...string) *ModelConfig { - c := &ModelConfig{Name: "chain", Failover: &FailoverConfig{}} - for _, t := range targets { - c.Failover.Targets = append(c.Failover.Targets, FailoverTarget{Model: t}) - } - return c - } - - BeforeEach(func() { - loader = NewModelConfigLoader("") - loader.configs["a"] = ModelConfig{Name: "a", Backend: "llama-cpp", KnownUsecaseStrings: []string{"chat"}} - loader.configs["b"] = ModelConfig{Name: "b", Backend: "llama-cpp", KnownUsecaseStrings: []string{"chat"}} - loader.configs["tts"] = ModelConfig{Name: "tts", Backend: "piper", KnownUsecaseStrings: []string{"tts"}} - loader.configs["alias-b"] = ModelConfig{Name: "alias-b", Alias: "b"} - loader.configs["other-chain"] = *chain("a", "b") - loader.configs["alias-chain"] = ModelConfig{Name: "alias-chain", Alias: "other-chain"} - for k, c := range loader.configs { - c.KnownUsecases = GetUsecasesFromYAML(c.KnownUsecaseStrings) - loader.configs[k] = c - } - }) - - It("accepts existing targets and alias targets", func() { - Expect(loader.ValidateFailoverTargets(chain("a", "alias-b"))).To(Succeed()) - }) - It("rejects a missing target", func() { - Expect(loader.ValidateFailoverTargets(chain("a", "nope"))).To(MatchError(ContainSubstring("does not exist"))) - }) - It("rejects a nested chain, directly or through an alias", func() { - Expect(loader.ValidateFailoverTargets(chain("a", "other-chain"))).To(MatchError(ContainSubstring("chains do not nest"))) - Expect(loader.ValidateFailoverTargets(chain("a", "alias-chain"))).To(MatchError(ContainSubstring("chains do not nest"))) - }) - It("reports whether targets share a usecase", func() { - Expect(loader.FailoverTargetsShareUsecase(chain("a", "b"))).To(BeTrue()) - Expect(loader.FailoverTargetsShareUsecase(chain("a", "tts"))).To(BeFalse()) - }) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -Run: `go test ./core/config/... 2>&1 | tail -5` -Expected: `undefined: ... ValidateFailoverTargets`. - -- [ ] **Step 3: Implement** - -In `core/config/model_config_loader.go`: - -```go -// failoverUsecases are the single usecases a chain can share. Checking one -// flag at a time avoids treating "chat+tts" and "tts" as unrelated. -var failoverUsecases = []ModelConfigUsecase{ - FLAG_CHAT, FLAG_COMPLETION, FLAG_EMBEDDINGS, FLAG_RERANK, FLAG_IMAGE, - FLAG_TRANSCRIPT, FLAG_TTS, FLAG_SOUND_GENERATION, FLAG_VAD, FLAG_VIDEO, - FLAG_SOUND_CLASSIFICATION, -} - -// ValidateFailoverTargets checks that every target of a chain exists and is -// not itself a chain. Alias targets are allowed and resolve one hop. -func (bcl *ModelConfigLoader) ValidateFailoverTargets(cfg *ModelConfig) error { - return validateFailoverTargets(cfg, bcl.GetModelConfig) -} - -// FailoverTargetsShareUsecase reports whether all targets of a chain have at -// least one usecase in common. A false result is only a warning: usecases are -// often inferred. -func (bcl *ModelConfigLoader) FailoverTargetsShareUsecase(cfg *ModelConfig) bool { - return failoverTargetsShareUsecase(cfg, bcl.GetModelConfig) -} - -func validateFailoverTargets(cfg *ModelConfig, lookup func(string) (ModelConfig, bool)) error { - if cfg == nil || !cfg.IsFailover() { - return nil - } - for _, t := range cfg.Failover.Targets { - target, ok := lookup(t.Model) - if !ok { - return fmt.Errorf("failover chain %q: target %q does not exist", cfg.Name, t.Model) - } - if target.IsAlias() { - if resolved, ok := lookup(target.Alias); ok { - target = resolved - } - } - if target.IsFailover() { - return fmt.Errorf("failover chain %q: target %q is a chain (chains do not nest)", cfg.Name, t.Model) - } - } - return nil -} - -func failoverTargetsShareUsecase(cfg *ModelConfig, lookup func(string) (ModelConfig, bool)) bool { - if cfg == nil || !cfg.IsFailover() { - return true - } - var targets []ModelConfig - for _, t := range cfg.Failover.Targets { - target, ok := lookup(t.Model) - if !ok { - return true // missing targets are reported by validateFailoverTargets - } - if target.IsAlias() { - if resolved, ok := lookup(target.Alias); ok { - target = resolved - } - } - targets = append(targets, target) - } - for _, u := range failoverUsecases { - all := true - for i := range targets { - if !targets[i].HasUsecases(u) { - all = false - break - } - } - if all { - return true - } - } - return false -} -``` - -In `loadModelConfigsFromPath`, directly after the alias warning loop (~824-840), add a pass over `bcl.configs`. Check whether that code runs with `bcl.Mutex` held: if it does, pass a lock-free lookup (`func(n string) (ModelConfig, bool) { c, ok := bcl.configs[n]; return c, ok }`) instead of `bcl.GetModelConfig`, otherwise it deadlocks. - -```go - for name, cfg := range bcl.configs { - if !cfg.IsFailover() { - continue - } - c := cfg - if err := validateFailoverTargets(&c, lookup); err != nil { - if strict { - return fmt.Errorf("invalid model config %q: %w", name, err) - } - xlog.Error("skipping invalid failover chain", "model", name, "error", err) - delete(bcl.configs, name) - continue - } - if !failoverTargetsShareUsecase(&c, lookup) { - xlog.Warn("failover chain targets share no known usecase", "model", name) - } - } -``` - -Use the strict-mode variable already in scope in that function (it is the one used at line ~813 for `invalid model config`). If the function has no strict flag at that point, drop the `if strict` branch. - -In `core/services/modeladmin/config.go` (both sites) and `core/http/endpoints/localai/import_model.go`, directly after each `ValidateAliasTarget(...)` call, add the equivalent call and return the error the same way the alias error is returned: - -```go - if err := s.Loader.ValidateFailoverTargets(&cfg); err != nil { - return /* same error shape as the ValidateAliasTarget branch above */ err - } -``` - -Copy the exact receiver and variable names from the adjacent `ValidateAliasTarget` call. Do not change what the alias branch returns. - -- [ ] **Step 4: Run tests** - -Run: `go test ./core/config/... ./core/services/modeladmin/... 2>&1 | tail -10` -Expected: PASS. - -- [ ] **Step 5: Commit** - -```bash -git add core/config core/services/modeladmin core/http/endpoints/localai/import_model.go -git commit -m "feat(config): validate failover chain targets across configs - -Reject chains whose targets are missing or are chains, at load and on -create or edit, and warn when the targets share no usecase. - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 3: Failover package types and error classification - -**Files:** -- Create: `core/services/failover/failover_suite_test.go`, `types.go`, `classify.go`, `classify_test.go`, `types_test.go` - -**Interfaces:** -- Produces: types `TargetState`, `ChainState`, `Kind`, `Reason`, `EventType`, `Event`, `TargetStatus`, `ChainStatus`; constants listed below; `KindOf(cfg config.ModelConfig) Kind`; `MergePinned(pinned, warm []string) []string`; `IsRetryable(err error, status int) bool`. - -- [ ] **Step 1: Write the failing tests** - -`core/services/failover/failover_suite_test.go`: - -```go -package failover - -import ( - "testing" - - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" -) - -func TestFailover(t *testing.T) { - RegisterFailHandler(Fail) - RunSpecs(t, "Failover test suite") -} -``` - -`core/services/failover/classify_test.go`: - -```go -package failover - -import ( - "context" - "errors" - "fmt" - "net/http" - - "github.com/labstack/echo/v4" - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" - "google.golang.org/grpc/codes" - grpcstatus "google.golang.org/grpc/status" -) - -var _ = DescribeTable("IsRetryable", - func(err error, status int, want bool) { - Expect(IsRetryable(err, status)).To(Equal(want)) - }, - Entry("nil error, no status", nil, 0, false), - Entry("held 503", nil, http.StatusServiceUnavailable, true), - Entry("held 500", nil, http.StatusInternalServerError, true), - Entry("held 501", nil, http.StatusNotImplemented, false), - Entry("client cancel", context.Canceled, 0, false), - Entry("wrapped client cancel", fmt.Errorf("predict: %w", context.Canceled), 0, false), - Entry("deadline", context.DeadlineExceeded, 0, true), - Entry("echo 502", echo.NewHTTPError(http.StatusBadGateway, "x"), 0, true), - Entry("echo 400", echo.NewHTTPError(http.StatusBadRequest, "x"), 0, false), - Entry("echo 404", echo.NewHTTPError(http.StatusNotFound, "x"), 0, false), - Entry("grpc unavailable", grpcstatus.Error(codes.Unavailable, "x"), 0, true), - Entry("grpc internal", grpcstatus.Error(codes.Internal, "x"), 0, true), - Entry("grpc deadline", grpcstatus.Error(codes.DeadlineExceeded, "x"), 0, true), - Entry("grpc unknown", grpcstatus.Error(codes.Unknown, "x"), 0, true), - Entry("grpc invalid argument", grpcstatus.Error(codes.InvalidArgument, "x"), 0, false), - Entry("cloud-proxy upstream 503", errors.New("cloud-proxy: upstream 503: no healthy nodes"), 0, true), - Entry("cloud-proxy upstream 429 stays 4xx", errors.New("cloud-proxy: upstream 429: slow down"), 0, false), - Entry("context overflow", errors.New("the request exceeds the available context size"), 0, false), - Entry("dial error", errors.New("dial tcp 10.0.0.1:8080: connect: connection refused"), 0, true), -) -``` - -`core/services/failover/types_test.go`: - -```go -package failover - -import ( - "github.com/mudler/LocalAI/core/config" - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" -) - -var _ = Describe("KindOf", func() { - It("treats proxy backends as remote", func() { - Expect(KindOf(config.ModelConfig{Backend: "cloud-proxy"})).To(Equal(KindRemote)) - Expect(KindOf(config.ModelConfig{Backend: "localai-proxy"})).To(Equal(KindRemote)) - Expect(KindOf(config.ModelConfig{Backend: "llama-cpp"})).To(Equal(KindLocal)) - }) -}) - -var _ = Describe("MergePinned", func() { - It("adds warm targets without duplicates", func() { - Expect(MergePinned([]string{"a", "b"}, []string{"b", "c"})).To(Equal([]string{"a", "b", "c"})) - Expect(MergePinned(nil, nil)).To(BeEmpty()) - }) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -Run: `go test ./core/services/failover/... 2>&1 | tail -5` -Expected: compile failure (`undefined: IsRetryable`). - -- [ ] **Step 3: Implement** - -`core/services/failover/types.go`: - -```go -// Package failover serves a model name from an ordered chain of target -// models, moving to the next target when one fails and back when it -// recovers. -package failover - -import ( - "slices" - "time" - - "github.com/mudler/LocalAI/core/config" -) - -type TargetState string - -const ( - StateHealthy TargetState = "healthy" - StateDown TargetState = "down" - StateRecovering TargetState = "recovering" - StateMissing TargetState = "missing" -) - -type ChainState string - -const ( - ChainPrimary ChainState = "primary" - ChainFallback ChainState = "fallback" - ChainDegraded ChainState = "degraded" -) - -type Kind string - -const ( - KindLocal Kind = "local" - KindRemote Kind = "remote" -) - -type Reason string - -const ( - ReasonTrip Reason = "trip" - ReasonRecovery Reason = "recovery" - ReasonManual Reason = "manual" - ReasonDegraded Reason = "degraded" - ReasonMissing Reason = "missing" - ReasonInitial Reason = "initial" -) - -type EventType string - -const ( - EventChainSwitched EventType = "chain.switched" - EventTargetState EventType = "target.state" -) - -// Event is one change of a target state or of a chain's active target. -type Event struct { - Type EventType `json:"type"` - Chain string `json:"chain,omitempty"` - Target string `json:"target,omitempty"` - From string `json:"from"` - To string `json:"to"` - State string `json:"state,omitempty"` - Reason Reason `json:"reason"` - Error string `json:"error,omitempty"` - At time.Time `json:"at"` -} - -type TargetStatus struct { - Model string `json:"model"` - Kind Kind `json:"kind"` - Warm bool `json:"warm"` - State TargetState `json:"state"` - ConsecutiveOK int `json:"consecutive_ok"` - LastProbe *time.Time `json:"last_probe,omitempty"` - LastError string `json:"last_error,omitempty"` -} - -type ChainStatus struct { - Name string `json:"name"` - State ChainState `json:"state"` - Active string `json:"active"` - ActiveSince time.Time `json:"active_since"` - Pinned *string `json:"pinned"` - Targets []TargetStatus `json:"targets"` -} - -// KindOf decides how a target is probed: proxy backends forward to another -// server and are checked over HTTP, everything else runs in this instance. -func KindOf(cfg config.ModelConfig) Kind { - switch cfg.Backend { - case "cloud-proxy", "localai-proxy": - return KindRemote - } - return KindLocal -} - -// MergePinned adds warm failover targets to the config-pinned model list, so -// the watchdog never evicts them. -func MergePinned(pinned, warm []string) []string { - out := slices.Clone(pinned) - for _, w := range warm { - if !slices.Contains(out, w) { - out = append(out, w) - } - } - return out -} -``` - -`core/services/failover/classify.go`: - -```go -package failover - -import ( - "context" - "errors" - "net/http" - "regexp" - "strconv" - "strings" - - "github.com/labstack/echo/v4" - "google.golang.org/grpc/codes" - grpcstatus "google.golang.org/grpc/status" -) - -// cloud-proxy translate mode reports upstream failures as plain text, so the -// status is only visible in the message. -var upstreamStatusRe = regexp.MustCompile(`upstream (\d{3})`) - -// Errors that the next target would reject in the same way. -var requestErrorMarkers = []string{ - "exceeds the available context size", - "is larger than the max context size", - "maximum context length", -} - -// IsRetryable reports whether a failed attempt should move to the next -// target. status is the HTTP status a handler wrote, or 0 when it returned err -// without writing. -func IsRetryable(err error, status int) bool { - if errors.Is(err, context.Canceled) { - return false - } - if status != 0 { - return retryableStatus(status) - } - if err == nil { - return false - } - var he *echo.HTTPError - if errors.As(err, &he) { - return retryableStatus(he.Code) - } - if errors.Is(err, context.DeadlineExceeded) { - return true - } - if st, ok := grpcstatus.FromError(err); ok { - switch st.Code() { - case codes.Unavailable, codes.Internal, codes.DeadlineExceeded, codes.Unknown: - return !isRequestError(st.Message()) - default: - return false - } - } - msg := err.Error() - if m := upstreamStatusRe.FindStringSubmatch(msg); m != nil { - code, _ := strconv.Atoi(m[1]) - return retryableStatus(code) - } - // Anything else is usually a dial or load failure of this target. - return !isRequestError(msg) -} - -func retryableStatus(code int) bool { - return code >= 500 && code != http.StatusNotImplemented -} - -func isRequestError(msg string) bool { - for _, m := range requestErrorMarkers { - if strings.Contains(msg, m) { - return true - } - } - return false -} -``` - -- [ ] **Step 4: Run tests** - -Run: `go test ./core/services/failover/... 2>&1 | tail -5` -Expected: PASS. - -- [ ] **Step 5: Commit** - -```bash -git add core/services/failover -git commit -m "feat(failover): add types and retryable error classification - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 4: Manager state machine, plans, pins, events - -**Files:** -- Create: `core/services/failover/manager.go`, `core/services/failover/manager_test.go`, `core/services/failover/fakes_test.go` - -**Interfaces:** -- Consumes: Task 1 config types, Task 3 types and `IsRetryable`. -- Produces (used by Tasks 5-12): - - `type ConfigSource interface { GetModelConfig(string) (config.ModelConfig, bool); GetAllModelsConfigs() []config.ModelConfig }` - - `type Clock interface { Now() time.Time }` - - `func New(src ConfigSource, opts ...Option) *Manager`; options `WithClock(Clock)`, `WithProber(Prober)`, `WithOnWarmChanged(func([]string))`; the `Prober` interface (declared at the end of `manager.go`, implemented in Task 6) - - `(*Manager).Sync()`, `Reevaluate()`, `Plan(chain string) (*Attempt, error)`, `ReportFailure(target string, err error)`, `ReportSuccess(target string)`, `Pin(chain, target string) error`, `Unpin(chain string) error`, `Status() []ChainStatus`, `ChainStatus(name string) (ChainStatus, bool)`, `Subscribe(buffer int) (<-chan Event, func())`, `WarmTargets() []string`, `Do(ctx, chain string, fn func(ctx context.Context, target string, commit func()) error) error` - - `(*Attempt).Chain() string`, `Target() string`, `Primary() string`, `Degraded() bool`, `Fail(err error) bool`, `Report(err error)`, `Succeed()` - - errors `ErrChainNotFound`, `ErrTargetNotInChain`, `ErrNoTarget` - -- [ ] **Step 1: Write test fakes** - -`core/services/failover/fakes_test.go`: - -```go -package failover - -import ( - "sort" - "sync" - "time" - - "github.com/mudler/LocalAI/core/config" -) - -type fakeClock struct { - mu sync.Mutex - now time.Time -} - -func newFakeClock() *fakeClock { return &fakeClock{now: time.Date(2026, 9, 26, 10, 0, 0, 0, time.UTC)} } -func (c *fakeClock) Now() time.Time { - c.mu.Lock() - defer c.mu.Unlock() - return c.now -} -func (c *fakeClock) Advance(d time.Duration) { - c.mu.Lock() - c.now = c.now.Add(d) - c.mu.Unlock() -} - -type fakeSource struct { - mu sync.Mutex - cfgs map[string]config.ModelConfig -} - -func newFakeSource(cfgs ...config.ModelConfig) *fakeSource { - s := &fakeSource{cfgs: map[string]config.ModelConfig{}} - for _, c := range cfgs { - s.cfgs[c.Name] = c - } - return s -} -func (s *fakeSource) Put(c config.ModelConfig) { s.mu.Lock(); s.cfgs[c.Name] = c; s.mu.Unlock() } -func (s *fakeSource) Delete(name string) { s.mu.Lock(); delete(s.cfgs, name); s.mu.Unlock() } -func (s *fakeSource) GetModelConfig(n string) (config.ModelConfig, bool) { - s.mu.Lock() - defer s.mu.Unlock() - c, ok := s.cfgs[n] - return c, ok -} -func (s *fakeSource) GetAllModelsConfigs() []config.ModelConfig { - s.mu.Lock() - defer s.mu.Unlock() - out := make([]config.ModelConfig, 0, len(s.cfgs)) - for _, c := range s.cfgs { - out = append(out, c) - } - sort.Slice(out, func(i, j int) bool { return out[i].Name < out[j].Name }) - return out -} - -func local(name string) config.ModelConfig { return config.ModelConfig{Name: name, Backend: "llama-cpp"} } -func remote(name string) config.ModelConfig { return config.ModelConfig{Name: name, Backend: "cloud-proxy"} } - -// chainCfg builds a chain; fc may be nil for defaults. -func chainCfg(name string, fc *config.FailoverConfig, targets ...config.FailoverTarget) config.ModelConfig { - f := config.FailoverConfig{} - if fc != nil { - f = *fc - } - f.Targets = targets - return config.ModelConfig{Name: name, Failover: &f} -} - -func t(model string) config.FailoverTarget { return config.FailoverTarget{Model: model} } -func warmT(model string) config.FailoverTarget { return config.FailoverTarget{Model: model, Warm: true} } - -// drain returns the events buffered so far without blocking. -func drain(ch <-chan Event) []Event { - var out []Event - for { - select { - case ev, ok := <-ch: - if !ok { - return out - } - out = append(out, ev) - default: - return out - } - } -} -``` - -- [ ] **Step 2: Write the failing manager tests** - -`core/services/failover/manager_test.go`: - -```go -package failover - -import ( - "context" - "errors" - "time" - - "github.com/mudler/LocalAI/core/config" - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" -) - -var errBoom = errors.New("dial tcp: connection refused") - -var _ = Describe("Manager", func() { - var ( - clock *fakeClock - src *fakeSource - m *Manager - ) - - BeforeEach(func() { - clock = newFakeClock() - src = newFakeSource(remote("a"), local("b"), chainCfg("chain", nil, t("a"), t("b"))) - m = New(src, WithClock(clock)) - }) - - switched := func(evs []Event) []Event { - var out []Event - for _, e := range evs { - if e.Type == EventChainSwitched { - out = append(out, e) - } - } - return out - } - - It("plans the primary first on a fresh chain", func() { - att, err := m.Plan("chain") - Expect(err).ToNot(HaveOccurred()) - Expect(att.Target()).To(Equal("a")) - Expect(att.Primary()).To(Equal("a")) - Expect(att.Degraded()).To(BeFalse()) - st, ok := m.ChainStatus("chain") - Expect(ok).To(BeTrue()) - Expect(st.State).To(Equal(ChainPrimary)) - Expect(st.Targets[0].Kind).To(Equal(KindRemote)) - Expect(st.Targets[1].Kind).To(Equal(KindLocal)) - }) - - It("returns ErrChainNotFound for an unknown chain", func() { - _, err := m.Plan("nope") - Expect(errors.Is(err, ErrChainNotFound)).To(BeTrue()) - }) - - It("trips on the first failure by default and switches with an event", func() { - events, cancel := m.Subscribe(16) - defer cancel() - att, _ := m.Plan("chain") - Expect(att.Fail(errBoom)).To(BeTrue()) - Expect(att.Target()).To(Equal("b")) - st, _ := m.ChainStatus("chain") - Expect(st.Active).To(Equal("b")) - Expect(st.State).To(Equal(ChainFallback)) - Expect(st.Targets[0].State).To(Equal(StateDown)) - Expect(st.Targets[0].LastError).To(ContainSubstring("connection refused")) - sw := switched(drain(events)) - Expect(sw).To(HaveLen(1)) - Expect(sw[0]).To(MatchFields(IgnoreExtras, Fields{ - "Chain": Equal("chain"), "From": Equal("a"), "To": Equal("b"), - "State": Equal("fallback"), "Reason": Equal(ReasonTrip), - })) - }) - - It("counts failures inside the trip window only", func() { - src.Put(chainCfg("chain", &config.FailoverConfig{Trip: config.FailoverTrip{Errors: 2, Window: "30s"}}, t("a"), t("b"))) - m.Sync() - m.ReportFailure("a", errBoom) - clock.Advance(31 * time.Second) - m.ReportFailure("a", errBoom) - st, _ := m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateHealthy)) - clock.Advance(time.Second) - m.ReportFailure("a", errBoom) - st, _ = m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateDown)) - }) - - It("fails back only after recovery probes and min_dwell", func() { - m.Plan("chain") - m.ReportFailure("a", errBoom) - for i := 0; i < 3; i++ { - m.ReportSuccess("a") // a real success counts like a passed inference probe - } - st, _ := m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateHealthy)) - Expect(st.Active).To(Equal("b"), "min_dwell has not passed") - clock.Advance(61 * time.Second) - events, cancel := m.Subscribe(16) - defer cancel() - m.Reevaluate() - st, _ = m.ChainStatus("chain") - Expect(st.Active).To(Equal("a")) - Expect(switched(drain(events))[0].Reason).To(Equal(ReasonRecovery)) - }) - - It("moves up at once when the active target itself goes down", func() { - src.Put(chainCfg("chain", nil, t("a"), t("b"), t("c"))) - src.Put(local("c")) - m.Sync() - m.ReportFailure("a", errBoom) // active: b - for i := 0; i < 3; i++ { - m.ReportSuccess("a") // a healthy again, but dwell not passed - } - m.ReportFailure("b", errBoom) // b down: go to a now, not c - st, _ := m.ChainStatus("chain") - Expect(st.Active).To(Equal("a")) - }) - - It("goes degraded when all targets are down and plans all of them in priority order", func() { - m.ReportFailure("a", errBoom) - m.ReportFailure("b", errBoom) - st, _ := m.ChainStatus("chain") - Expect(st.State).To(Equal(ChainDegraded)) - att, err := m.Plan("chain") - Expect(err).ToNot(HaveOccurred()) - Expect(att.Degraded()).To(BeTrue()) - Expect(att.Target()).To(Equal("a")) - Expect(att.Fail(errBoom)).To(BeTrue()) - Expect(att.Target()).To(Equal("b")) - Expect(att.Fail(errBoom)).To(BeFalse()) - }) - - It("pins a target regardless of health", func() { - Expect(m.Pin("chain", "b")).To(Succeed()) - att, _ := m.Plan("chain") - Expect(att.Target()).To(Equal("b")) - Expect(att.Fail(errBoom)).To(BeFalse(), "a pin allows only the pinned target") - st, _ := m.ChainStatus("chain") - Expect(*st.Pinned).To(Equal("b")) - Expect(st.Active).To(Equal("b")) - Expect(m.Unpin("chain")).To(Succeed()) - st, _ = m.ChainStatus("chain") - Expect(st.Pinned).To(BeNil()) - Expect(errors.Is(m.Pin("chain", "zzz"), ErrTargetNotInChain)).To(BeTrue()) - Expect(errors.Is(m.Pin("nope", "a"), ErrChainNotFound)).To(BeTrue()) - }) - - It("shares target health across chains", func() { - src.Put(chainCfg("chain2", nil, t("a"), t("b"))) - m.Sync() - m.ReportFailure("a", errBoom) - s1, _ := m.ChainStatus("chain") - s2, _ := m.ChainStatus("chain2") - Expect(s1.Active).To(Equal("b")) - Expect(s2.Active).To(Equal("b")) - }) - - It("marks a removed target missing and leaves it out of plans", func() { - m.Plan("chain") - src.Delete("a") - m.Sync() - st, _ := m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateMissing)) - att, _ := m.Plan("chain") - Expect(att.Target()).To(Equal("b")) - Expect(att.Fail(errBoom)).To(BeFalse()) - }) - - It("resets a chain whose target list changed", func() { - m.ReportFailure("a", errBoom) - src.Put(local("c")) - src.Put(chainCfg("chain", nil, t("c"), t("b"))) - m.Sync() - st, _ := m.ChainStatus("chain") - Expect(st.Active).To(Equal("c")) - }) - - It("reports warm local targets and ignores warm on remote ones", func() { - var got []string - m = New(src, WithClock(clock), WithOnWarmChanged(func(w []string) { got = w })) - src.Put(chainCfg("chain", nil, warmT("a"), warmT("b"))) - m.Sync() - Expect(got).To(Equal([]string{"b"})) - Expect(m.WarmTargets()).To(Equal([]string{"b"})) - }) - - It("closes a subscription on cancel", func() { - events, cancel := m.Subscribe(1) - cancel() - _, ok := <-events - Expect(ok).To(BeFalse()) - cancel() // idempotent - }) - - Describe("Do", func() { - It("retries on the next target until commit", func() { - var tried []string - err := m.Do(context.Background(), "chain", func(_ context.Context, target string, commit func()) error { - tried = append(tried, target) - if target == "a" { - return errBoom - } - return nil - }) - Expect(err).ToNot(HaveOccurred()) - Expect(tried).To(Equal([]string{"a", "b"})) - }) - - It("does not retry after commit but still trips the target", func() { - var tried []string - err := m.Do(context.Background(), "chain", func(_ context.Context, target string, commit func()) error { - tried = append(tried, target) - commit() - return errBoom - }) - Expect(err).To(MatchError(errBoom)) - Expect(tried).To(Equal([]string{"a"})) - st, _ := m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateDown)) - }) - - It("does not retry or trip on a non-retryable error", func() { - bad := errors.New("the request exceeds the available context size") - err := m.Do(context.Background(), "chain", func(_ context.Context, _ string, _ func()) error { return bad }) - Expect(err).To(MatchError(bad)) - st, _ := m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateHealthy)) - }) - }) -}) -``` - -The tests use `MatchFields`, so add `. "github.com/onsi/gomega/gstruct"` to the imports. - -- [ ] **Step 3: Run to verify failure** - -Run: `go test ./core/services/failover/... 2>&1 | tail -5` -Expected: compile failure (`undefined: New`). - -- [ ] **Step 4: Implement the manager** - -`core/services/failover/manager.go`: - -```go -package failover - -import ( - "context" - "errors" - "fmt" - "slices" - "sort" - "sync" - "sync/atomic" - "time" - - "github.com/mudler/LocalAI/core/config" - "github.com/mudler/xlog" -) - -var ( - ErrChainNotFound = errors.New("failover chain not found") - ErrTargetNotInChain = errors.New("target is not in this failover chain") - ErrNoTarget = errors.New("failover chain has no usable target") -) - -// ConfigSource is the part of ModelConfigLoader the manager reads. -type ConfigSource interface { - GetModelConfig(name string) (config.ModelConfig, bool) - GetAllModelsConfigs() []config.ModelConfig -} - -type Clock interface{ Now() time.Time } - -type realClock struct{} - -func (realClock) Now() time.Time { return time.Now() } - -type Option func(*Manager) - -func WithClock(c Clock) Option { return func(m *Manager) { m.clock = c } } -func WithProber(p Prober) Option { return func(m *Manager) { m.prober = p } } - -// WithOnWarmChanged is called outside the manager lock when the set of warm -// local targets changes. The application pins and preloads them. -func WithOnWarmChanged(fn func(warm []string)) Option { return func(m *Manager) { m.onWarm = fn } } - -// Manager tracks health per target and the active target per chain. -type Manager struct { - mu sync.Mutex - src ConfigSource - clock Clock - prober Prober - onWarm func([]string) - targets map[string]*targetState - chains map[string]*chainState - subs map[int]chan Event - nextSub int - warm []string - warmPending bool - closed bool -} - -type targetState struct { - name string - kind Kind - warm bool - state TargetState - failures []time.Time - consecutiveOK int - downSince time.Time - lastProbe time.Time - lastActivity time.Time - lastError string - // params come from the first chain, in name order, that lists the target. - params config.FailoverConfig -} - -func (ts *targetState) cold() bool { return ts.kind == KindLocal && !ts.warm } - -type chainState struct { - name string - cfg config.FailoverConfig - targets []string - active int - activeSince time.Time - pinned string - state ChainState -} - -func New(src ConfigSource, opts ...Option) *Manager { - m := &Manager{ - src: src, - clock: realClock{}, - targets: map[string]*targetState{}, - chains: map[string]*chainState{}, - subs: map[int]chan Event{}, - } - for _, o := range opts { - o(m) - } - return m -} - -// Sync reconciles chains with the config source. There is no config-change -// hook in the loader, so this runs on every tick and on a lookup miss. -func (m *Manager) Sync() { - m.mu.Lock() - m.syncLocked() - warm, deliver := m.takeWarmLocked() - m.mu.Unlock() - if deliver && m.onWarm != nil { - m.onWarm(warm) - } -} - -func (m *Manager) syncLocked() { - now := m.clock.Now() - seenChains := map[string]bool{} - claimed := map[string]bool{} - for _, c := range m.src.GetAllModelsConfigs() { - if !c.IsFailover() { - continue - } - seenChains[c.Name] = true - names := make([]string, 0, len(c.Failover.Targets)) - for _, t := range c.Failover.Targets { - names = append(names, t.Model) - } - ch := m.chains[c.Name] - if ch == nil || !slices.Equal(ch.targets, names) { - pinned := "" - if ch != nil && slices.Contains(names, ch.pinned) { - pinned = ch.pinned - } - ch = &chainState{name: c.Name, targets: names, activeSince: now, state: ChainPrimary, pinned: pinned} - m.chains[c.Name] = ch - } - ch.cfg = *c.Failover - for _, t := range c.Failover.Targets { - ts := m.targets[t.Model] - if ts == nil { - ts = &targetState{name: t.Model, state: StateHealthy} - m.targets[t.Model] = ts - } - if !claimed[t.Model] { - claimed[t.Model] = true - ts.params = *c.Failover - ts.warm = false - } - tc, ok := m.lookupTarget(t.Model) - if !ok { - m.setTargetLocked(ts, StateMissing, ReasonMissing, "target config not found") - continue - } - ts.kind = KindOf(tc) - if t.Warm && ts.kind == KindLocal { - ts.warm = true - } - if ts.state == StateMissing { - m.setTargetLocked(ts, StateHealthy, ReasonRecovery, "") - } - } - } - for name := range m.chains { - if !seenChains[name] { - delete(m.chains, name) - } - } - for name := range m.targets { - if !claimed[name] { - delete(m.targets, name) - } - } - for _, ch := range m.chains { - m.recomputeLocked(ch, "") - } - var warm []string - for name, ts := range m.targets { - if ts.warm { - warm = append(warm, name) - } - } - sort.Strings(warm) - if !slices.Equal(warm, m.warm) { - m.warm = warm - m.warmPending = true - } -} - -func (m *Manager) takeWarmLocked() ([]string, bool) { - if !m.warmPending { - return nil, false - } - m.warmPending = false - return slices.Clone(m.warm), true -} - -// lookupTarget returns the config that serves a target, one alias hop deep. -func (m *Manager) lookupTarget(name string) (config.ModelConfig, bool) { - c, ok := m.src.GetModelConfig(name) - if ok && c.IsAlias() { - return m.src.GetModelConfig(c.Alias) - } - return c, ok -} - -func (m *Manager) chainLocked(name string) *chainState { - if ch := m.chains[name]; ch != nil { - return ch - } - m.syncLocked() - return m.chains[name] -} - -// WarmTargets returns the warm local targets, sorted. -func (m *Manager) WarmTargets() []string { - m.mu.Lock() - defer m.mu.Unlock() - return slices.Clone(m.warm) -} - -// Reevaluate recomputes every chain. Dwell-based fail-back needs no event, so -// the scheduler calls this on every tick. -func (m *Manager) Reevaluate() { - m.mu.Lock() - defer m.mu.Unlock() - for _, ch := range m.chains { - m.recomputeLocked(ch, "") - } -} - -func (m *Manager) setTargetLocked(ts *targetState, to TargetState, reason Reason, errMsg string) { - if ts.state == to { - return - } - from := ts.state - now := m.clock.Now() - ts.state = to - switch to { - case StateDown: - ts.downSince = now - ts.consecutiveOK = 0 - ts.failures = nil - case StateRecovering, StateHealthy: - ts.consecutiveOK = 0 - ts.failures = nil - } - m.emitLocked(Event{Type: EventTargetState, Target: ts.name, From: string(from), To: string(to), Reason: reason, Error: errMsg, At: now}) -} - -// recomputeLocked picks the active target. override replaces the reason of a -// resulting switch (pin and unpin are always "manual"). -func (m *Manager) recomputeLocked(ch *chainState, override Reason) { - now := m.clock.Now() - prev := ch.active - next := prev - reason := ReasonTrip - best := -1 - for i, name := range ch.targets { - if ts := m.targets[name]; ts != nil && ts.state == StateHealthy { - best = i - break - } - } - switch { - case ch.pinned != "": - next = slices.Index(ch.targets, ch.pinned) - reason = ReasonManual - case best == -1: - // Nothing is healthy: keep the active target, Plan tries all of them. - case best > prev: - next = best // the active target is not healthy - case best < prev: - cur := m.targets[ch.targets[prev]] - curHealthy := cur != nil && cur.state == StateHealthy - if !curHealthy { - next = best - } else if now.Sub(ch.activeSince) >= ch.cfg.MinDwell() { - next = best - reason = ReasonRecovery - } - } - if override != "" { - reason = override - } - var state ChainState - switch { - case ch.pinned == "" && best == -1: - state = ChainDegraded - case next == 0: - state = ChainPrimary - default: - state = ChainFallback - } - switch { - case next != prev: - ch.active = next - ch.activeSince = now - m.emitLocked(Event{Type: EventChainSwitched, Chain: ch.name, From: ch.targets[prev], To: ch.targets[next], State: string(state), Reason: reason, At: now}) - case state == ChainDegraded && ch.state != ChainDegraded: - m.emitLocked(Event{Type: EventChainSwitched, Chain: ch.name, From: ch.targets[prev], To: ch.targets[next], State: string(state), Reason: ReasonDegraded, At: now}) - } - ch.state = state -} - -func (m *Manager) recomputeForLocked(target string) { - for _, ch := range m.chains { - if slices.Contains(ch.targets, target) { - m.recomputeLocked(ch, "") - } - } -} - -// Attempt walks the targets of one request in order. -type Attempt struct { - m *Manager - chain string - primary string - degraded bool - targets []string - i int -} - -// Plan returns the attempt order for one request: the active target, then the -// other healthy targets. A degraded chain tries every target in priority -// order; a pinned chain only the pinned target. -func (m *Manager) Plan(chain string) (*Attempt, error) { - m.mu.Lock() - defer m.mu.Unlock() - ch := m.chainLocked(chain) - if ch == nil { - return nil, fmt.Errorf("%w: %q", ErrChainNotFound, chain) - } - att := &Attempt{m: m, chain: ch.name, primary: ch.targets[0], degraded: ch.state == ChainDegraded} - usable := func(name string) bool { - ts := m.targets[name] - if ts == nil || ts.state == StateMissing { - return false - } - return att.degraded || ts.state == StateHealthy - } - switch { - case ch.pinned != "": - att.targets = []string{ch.pinned} - case att.degraded: - for _, name := range ch.targets { - if usable(name) { - att.targets = append(att.targets, name) - } - } - default: - active := ch.targets[ch.active] - if usable(active) { - att.targets = append(att.targets, active) - } - for _, name := range ch.targets { - if name != active && usable(name) { - att.targets = append(att.targets, name) - } - } - } - if len(att.targets) == 0 { - return nil, fmt.Errorf("%w: %q", ErrNoTarget, chain) - } - return att, nil -} - -func (a *Attempt) Chain() string { return a.chain } -func (a *Attempt) Target() string { return a.targets[a.i] } -func (a *Attempt) Primary() string { return a.primary } -func (a *Attempt) Degraded() bool { return a.degraded } - -// Fail records err against the current target and moves to the next one. It -// returns false when no target is left. -func (a *Attempt) Fail(err error) bool { - a.m.ReportFailure(a.Target(), err) - if a.i+1 >= len(a.targets) { - return false - } - a.i++ - return true -} - -// Report records err against the current target without moving on: the -// response was already committed, so nothing is left to retry. -func (a *Attempt) Report(err error) { a.m.ReportFailure(a.Target(), err) } - -func (a *Attempt) Succeed() { a.m.ReportSuccess(a.Target()) } - -func (m *Manager) ReportFailure(target string, err error) { - m.mu.Lock() - defer m.mu.Unlock() - ts := m.targets[target] - if ts == nil { - return - } - msg := "" - if err != nil { - msg = err.Error() - } - m.recordFailureLocked(ts, msg) - m.recomputeForLocked(target) -} - -func (m *Manager) recordFailureLocked(ts *targetState, msg string) { - now := m.clock.Now() - ts.lastError = msg - switch ts.state { - case StateRecovering: - m.setTargetLocked(ts, StateDown, ReasonTrip, msg) - case StateHealthy: - cut := now.Add(-ts.params.TripWindow()) - kept := ts.failures[:0] - for _, f := range ts.failures { - if f.After(cut) { - kept = append(kept, f) - } - } - ts.failures = append(kept, now) - if len(ts.failures) >= ts.params.TripErrors() { - m.setTargetLocked(ts, StateDown, ReasonTrip, msg) - } - } -} - -func (m *Manager) ReportSuccess(target string) { - m.mu.Lock() - defer m.mu.Unlock() - ts := m.targets[target] - if ts == nil { - return - } - ts.lastActivity = m.clock.Now() - m.recordPassLocked(ts) - m.recomputeForLocked(target) -} - -// recordPassLocked counts a served request or a passed inference probe. -func (m *Manager) recordPassLocked(ts *targetState) { - switch ts.state { - case StateHealthy: - ts.failures = nil - return - case StateMissing: - return - case StateDown: - if ts.cold() { - // Cold targets are never probed; a served request is proof enough. - m.setTargetLocked(ts, StateHealthy, ReasonRecovery, "") - return - } - m.setTargetLocked(ts, StateRecovering, ReasonRecovery, "") - } - ts.consecutiveOK++ - if ts.consecutiveOK >= ts.params.RecoveryProbes() { - m.setTargetLocked(ts, StateHealthy, ReasonRecovery, "") - } -} - -func (m *Manager) Pin(chain, target string) error { - m.mu.Lock() - defer m.mu.Unlock() - ch := m.chainLocked(chain) - if ch == nil { - return fmt.Errorf("%w: %q", ErrChainNotFound, chain) - } - if !slices.Contains(ch.targets, target) { - return fmt.Errorf("%w: %q", ErrTargetNotInChain, target) - } - ch.pinned = target - m.recomputeLocked(ch, ReasonManual) - return nil -} - -func (m *Manager) Unpin(chain string) error { - m.mu.Lock() - defer m.mu.Unlock() - ch := m.chainLocked(chain) - if ch == nil { - return fmt.Errorf("%w: %q", ErrChainNotFound, chain) - } - ch.pinned = "" - m.recomputeLocked(ch, ReasonManual) - return nil -} - -// Status returns every chain, sorted by name. -func (m *Manager) Status() []ChainStatus { - m.mu.Lock() - defer m.mu.Unlock() - m.syncLocked() - names := make([]string, 0, len(m.chains)) - for name := range m.chains { - names = append(names, name) - } - sort.Strings(names) - out := make([]ChainStatus, 0, len(names)) - for _, name := range names { - out = append(out, m.statusLocked(m.chains[name])) - } - return out -} - -func (m *Manager) ChainStatus(name string) (ChainStatus, bool) { - m.mu.Lock() - defer m.mu.Unlock() - ch := m.chainLocked(name) - if ch == nil { - return ChainStatus{}, false - } - return m.statusLocked(ch), true -} - -func (m *Manager) statusLocked(ch *chainState) ChainStatus { - cs := ChainStatus{Name: ch.name, State: ch.state, Active: ch.targets[ch.active], ActiveSince: ch.activeSince} - if ch.pinned != "" { - p := ch.pinned - cs.Pinned = &p - } - for _, name := range ch.targets { - st := TargetStatus{Model: name} - if ts := m.targets[name]; ts != nil { - st.Kind, st.Warm, st.State = ts.kind, ts.warm, ts.state - st.ConsecutiveOK, st.LastError = ts.consecutiveOK, ts.lastError - if !ts.lastProbe.IsZero() { - lp := ts.lastProbe - st.LastProbe = &lp - } - } - cs.Targets = append(cs.Targets, st) - } - return cs -} - -// Subscribe returns a buffered event channel and a cancel func. A subscriber -// that does not keep up loses events rather than blocking the manager. -func (m *Manager) Subscribe(buffer int) (<-chan Event, func()) { - m.mu.Lock() - defer m.mu.Unlock() - ch := make(chan Event, buffer) - if m.closed { - close(ch) - return ch, func() {} - } - id := m.nextSub - m.nextSub++ - m.subs[id] = ch - var once sync.Once - return ch, func() { - once.Do(func() { - m.mu.Lock() - defer m.mu.Unlock() - if c, ok := m.subs[id]; ok { - delete(m.subs, id) - close(c) - } - }) - } -} - -func (m *Manager) emitLocked(ev Event) { - for _, c := range m.subs { - select { - case c <- ev: - default: - xlog.Warn("failover: dropping event for a slow subscriber", "type", ev.Type, "chain", ev.Chain, "target", ev.Target) - } - } -} - -func (m *Manager) close() { - m.mu.Lock() - defer m.mu.Unlock() - m.closed = true - for id, c := range m.subs { - close(c) - delete(m.subs, id) - } -} - -// Do runs fn against the chain's targets in plan order. fn calls commit once -// output has reached the client; after that a failure is not retried. -func (m *Manager) Do(ctx context.Context, chain string, fn func(ctx context.Context, target string, commit func()) error) error { - att, err := m.Plan(chain) - if err != nil { - return err - } - for { - var committed atomic.Bool - err := fn(ctx, att.Target(), func() { committed.Store(true) }) - switch { - case err == nil: - att.Succeed() - return nil - case ctx.Err() != nil || !IsRetryable(err, 0): - return err - case committed.Load(): - att.Report(err) - return err - case !att.Fail(err): - return err - } - } -} - -// Prober checks targets. Implemented by DefaultProber (prober.go). -type Prober interface { - // Liveness is the cheap steady-state check. - Liveness(ctx context.Context, target config.ModelConfig, kind Kind, warm bool) error - // Inference sends one minimal real request to confirm recovery. - Inference(ctx context.Context, target config.ModelConfig, kind Kind, warm bool) error -} -``` - -- [ ] **Step 5: Run tests** - -Run: `go test -race ./core/services/failover/... 2>&1 | tail -10` -Expected: PASS, no race reports. - -- [ ] **Step 6: Commit** - -```bash -git add core/services/failover -git commit -m "feat(failover): add chain manager with trip, fail-back and pins - -Health is tracked per target and the active target per chain. Fail-back -waits for recovery probes and a minimum time on the fallback. - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 5: Probe scheduling - -**Files:** -- Create: `core/services/failover/schedule.go`, `core/services/failover/schedule_test.go` - -**Interfaces:** -- Consumes: Task 4 `Manager`, `Prober`, internal helpers `recordFailureLocked`, `recordPassLocked`, `setTargetLocked`, `recomputeForLocked`, `lookupTarget`, `close`. -- Produces: `(*Manager).Run(ctx)`, `(*Manager).Tick(ctx)`. - -- [ ] **Step 1: Write the failing tests** - -`core/services/failover/schedule_test.go`: - -```go -package failover - -import ( - "context" - "sync" - "time" - - "github.com/mudler/LocalAI/core/config" - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" -) - -type probeCall struct { - target string - inference bool -} - -type fakeProber struct { - mu sync.Mutex - calls []probeCall - fail map[string]error // target -> error returned by every probe -} - -func (p *fakeProber) record(target string, inference bool) error { - p.mu.Lock() - defer p.mu.Unlock() - p.calls = append(p.calls, probeCall{target, inference}) - return p.fail[target] -} -func (p *fakeProber) Liveness(_ context.Context, c config.ModelConfig, _ Kind, _ bool) error { - return p.record(c.Name, false) -} -func (p *fakeProber) Inference(_ context.Context, c config.ModelConfig, _ Kind, _ bool) error { - return p.record(c.Name, true) -} -func (p *fakeProber) take() []probeCall { - p.mu.Lock() - defer p.mu.Unlock() - out := p.calls - p.calls = nil - return out -} - -var _ = Describe("Manager probes", func() { - var ( - clock *fakeClock - src *fakeSource - prober *fakeProber - m *Manager - ctx = context.Background() - ) - - BeforeEach(func() { - clock = newFakeClock() - prober = &fakeProber{fail: map[string]error{}} - src = newFakeSource(remote("a"), local("b"), local("cold"), - chainCfg("chain", nil, t("a"), warmT("b"))) - m = New(src, WithClock(clock), WithProber(prober)) - }) - - It("probes idle targets on the first tick and not again before the interval", func() { - m.Tick(ctx) - Expect(prober.take()).To(ConsistOf(probeCall{"a", false}, probeCall{"b", false})) - clock.Advance(5 * time.Second) - m.Tick(ctx) - Expect(prober.take()).To(BeEmpty()) - }) - - It("skips the liveness probe for a target with recent traffic", func() { - m.Tick(ctx) - prober.take() - clock.Advance(14 * time.Second) - m.ReportSuccess("a") - clock.Advance(2 * time.Second) - m.Tick(ctx) - Expect(prober.take()).To(ConsistOf(probeCall{"b", false})) - }) - - It("trips a target whose liveness probe fails", func() { - prober.fail["a"] = errBoom - m.Tick(ctx) - st, _ := m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateDown)) - Expect(st.Active).To(Equal("b")) - }) - - It("recovers through liveness, then inference probes, then fails back after dwell", func() { - prober.fail["a"] = errBoom - m.Tick(ctx) - delete(prober.fail, "a") - prober.take() - - clock.Advance(15 * time.Second) - m.Tick(ctx) // liveness passes: recovering - st, _ := m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateRecovering)) - - for i := 0; i < 3; i++ { - clock.Advance(15 * time.Second) - m.Tick(ctx) - } - calls := prober.take() - Expect(calls).To(ContainElement(probeCall{"a", true})) - st, _ = m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateHealthy)) - Expect(st.Active).To(Equal("a"), "60s min_dwell passed during the 4 ticks") - }) - - It("sends a recovering target back down when an inference probe fails", func() { - m.ReportFailure("a", errBoom) - clock.Advance(15 * time.Second) - m.Tick(ctx) // liveness passes: recovering - prober.fail["a"] = errBoom - clock.Advance(15 * time.Second) - m.Tick(ctx) - st, _ := m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateDown)) - }) - - It("never probes a down cold target and restores it after min_dwell", func() { - src.Put(chainCfg("chain", nil, t("cold"), warmT("b"))) - m.Sync() - m.ReportFailure("cold", errBoom) - prober.take() - clock.Advance(30 * time.Second) - m.Tick(ctx) - for _, c := range prober.take() { - Expect(c.target).ToNot(Equal("cold")) - } - st, _ := m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateDown)) - clock.Advance(31 * time.Second) - m.Tick(ctx) - st, _ = m.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(StateHealthy)) - }) - - It("probes a target shared by two chains once per tick", func() { - src.Put(chainCfg("chain2", nil, t("a"), warmT("b"))) - m.Tick(ctx) - calls := prober.take() - n := 0 - for _, c := range calls { - if c.target == "a" { - n++ - } - } - Expect(n).To(Equal(1)) - }) - - It("closes subscriptions when Run stops", func() { - events, _ := m.Subscribe(1) - rctx, cancel := context.WithCancel(ctx) - done := make(chan struct{}) - go func() { m.Run(rctx); close(done) }() - cancel() - Eventually(done).Should(BeClosed()) - Eventually(events).Should(BeClosed()) - }) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -Run: `go test ./core/services/failover/... 2>&1 | tail -5` -Expected: compile failure (`m.Tick undefined`). - -- [ ] **Step 3: Implement** - -`core/services/failover/schedule.go`: - -```go -package failover - -import ( - "context" - "sync" - "time" - - "github.com/mudler/LocalAI/core/config" -) - -// Run drives probes and dwell-based fail-back until ctx ends. -func (m *Manager) Run(ctx context.Context) { - ticker := time.NewTicker(time.Second) - defer ticker.Stop() - m.Tick(ctx) - for { - select { - case <-ctx.Done(): - m.close() - return - case <-ticker.C: - m.Tick(ctx) - } - } -} - -// Tick runs one pass: sync configs, run due probes, recompute chains. It is -// exported so tests can drive the manager without a real ticker. -func (m *Manager) Tick(ctx context.Context) { - m.Sync() - var wg sync.WaitGroup - for _, j := range m.dueProbes() { - wg.Add(1) - go func(j probeJob) { - defer wg.Done() - m.runProbe(ctx, j) - }(j) - } - wg.Wait() - m.Reevaluate() -} - -type probeJob struct { - target string - cfg config.ModelConfig - kind Kind - warm bool - inference bool - timeout time.Duration -} - -func (m *Manager) dueProbes() []probeJob { - m.mu.Lock() - defer m.mu.Unlock() - now := m.clock.Now() - var jobs []probeJob - for _, ts := range m.targets { - interval := ts.params.ProbeInterval() - inference := false - switch ts.state { - case StateMissing: - continue - case StateHealthy: - // A served request is as good as a liveness probe. - if now.Sub(ts.lastActivity) < interval || now.Sub(ts.lastProbe) < interval { - continue - } - case StateDown: - if ts.cold() { - // Loading a cold model only to probe it could evict others. - if now.Sub(ts.downSince) >= ts.params.MinDwell() { - m.setTargetLocked(ts, StateHealthy, ReasonRecovery, "") - m.recomputeForLocked(ts.name) - } - continue - } - if now.Sub(ts.lastProbe) < interval { - continue - } - case StateRecovering: - if ts.cold() || now.Sub(ts.lastProbe) < interval { - continue - } - inference = true - } - if m.prober == nil { - continue - } - cfg, ok := m.lookupTarget(ts.name) - if !ok { - continue - } - ts.lastProbe = now - jobs = append(jobs, probeJob{ - target: ts.name, cfg: cfg, kind: ts.kind, warm: ts.warm, - inference: inference, timeout: ts.params.ProbeTimeout(), - }) - } - return jobs -} - -func (m *Manager) runProbe(ctx context.Context, j probeJob) { - pctx, cancel := context.WithTimeout(ctx, j.timeout) - defer cancel() - var err error - if j.inference { - err = m.prober.Inference(pctx, j.cfg, j.kind, j.warm) - } else { - err = m.prober.Liveness(pctx, j.cfg, j.kind, j.warm) - } - if ctx.Err() != nil { - return // shutting down: a cancelled probe says nothing about the target - } - m.applyProbe(j, err) -} - -func (m *Manager) applyProbe(j probeJob, err error) { - m.mu.Lock() - defer m.mu.Unlock() - ts := m.targets[j.target] - if ts == nil || ts.state == StateMissing { - return - } - if err != nil { - if ts.state == StateDown { - ts.lastError = err.Error() - } else { - m.recordFailureLocked(ts, err.Error()) - } - m.recomputeForLocked(ts.name) - return - } - switch ts.state { - case StateHealthy: - ts.lastActivity = m.clock.Now() - ts.failures = nil - case StateDown: - m.setTargetLocked(ts, StateRecovering, ReasonRecovery, "") - case StateRecovering: - if j.inference { - m.recordPassLocked(ts) - } - } - m.recomputeForLocked(ts.name) -} -``` - -- [ ] **Step 4: Run tests** - -Run: `go test -race ./core/services/failover/... 2>&1 | tail -10` -Expected: PASS. - -- [ ] **Step 5: Commit** - -```bash -git add core/services/failover -git commit -m "feat(failover): schedule liveness and recovery probes - -Idle targets get a liveness probe each interval, recovering targets an -inference probe. Cold local targets are never loaded to be probed. - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 6: Default prober (remote HTTP and local gRPC) - -**Files:** -- Create: `core/services/failover/prober.go`, `core/services/failover/prober_test.go` -- Modify: `core/config/model_config.go` (add `ResolveAPIKey` next to `ProxyConfig`, line ~238-310) -- Test: `core/config/model_config_failover_test.go` (add ResolveAPIKey specs) -- Modify: `docs/superpowers/specs/2026-09-26-model-failover-chains-design.md` (cold-local row: "the model file exists"; see Step 5) - -**Interfaces:** -- Consumes: `Prober` (Task 4), `KindRemote/KindLocal`. -- Produces: `type LoadFunc func(ctx context.Context, cfg config.ModelConfig) (grpc.Backend, error)`, `func NewProber(load LoadFunc, modelPath string) *DefaultProber`, `func UpstreamBase(raw string) (string, error)`, `func UpstreamModel(cfg config.ModelConfig) string`, `(config.ProxyConfig).ResolveAPIKey() (string, error)`. - -- [ ] **Step 1: Write the failing tests** - -Add to `core/config/model_config_failover_test.go`: - -```go -var _ = Describe("ProxyConfig.ResolveAPIKey", func() { - It("reads the env var", func() { - GinkgoT().Setenv("FAILOVER_TEST_KEY", "k1") - Expect(ProxyConfig{APIKeyEnv: "FAILOVER_TEST_KEY"}.ResolveAPIKey()).To(Equal("k1")) - }) - It("fails on an unset env var", func() { - _, err := ProxyConfig{APIKeyEnv: "FAILOVER_TEST_UNSET_KEY"}.ResolveAPIKey() - Expect(err).To(HaveOccurred()) - }) - It("reads and trims the key file", func() { - f := filepath.Join(GinkgoT().TempDir(), "key") - Expect(os.WriteFile(f, []byte(" k2\n"), 0o600)).To(Succeed()) - Expect(ProxyConfig{APIKeyFile: f}.ResolveAPIKey()).To(Equal("k2")) - }) - It("returns empty when nothing is set", func() { - Expect(ProxyConfig{}.ResolveAPIKey()).To(Equal("")) - }) -}) -``` - -(add `os` and `path/filepath` imports) - -`core/services/failover/prober_test.go`: - -```go -package failover - -import ( - "context" - "encoding/json" - "errors" - "io" - "net/http" - "net/http/httptest" - "os" - "path/filepath" - "sync" - - "github.com/mudler/LocalAI/core/config" - "github.com/mudler/LocalAI/pkg/grpc" - pb "github.com/mudler/LocalAI/pkg/grpc/proto" - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" - ggrpc "google.golang.org/grpc" -) - -type fakeUpstream struct { - mu sync.Mutex - srv *httptest.Server - models []string - status int - paths []string - auth string -} - -func newFakeUpstream() *fakeUpstream { - u := &fakeUpstream{status: http.StatusOK} - u.srv = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { - u.mu.Lock() - u.paths = append(u.paths, r.Method+" "+r.URL.Path) - u.auth = r.Header.Get("Authorization") - status, models := u.status, u.models - u.mu.Unlock() - _, _ = io.Copy(io.Discard, r.Body) - if status != http.StatusOK { - w.WriteHeader(status) - return - } - if r.URL.Path == "/v1/models" { - var data []map[string]string - for _, m := range models { - data = append(data, map[string]string{"id": m}) - } - _ = json.NewEncoder(w).Encode(map[string]any{"data": data}) - return - } - _, _ = w.Write([]byte(`{}`)) - })) - return u -} - -type fakeBackend struct { - grpc.Backend - healthy bool - predictErr error - predicted bool -} - -func (b *fakeBackend) HealthCheck(context.Context) (bool, error) { return b.healthy, nil } -func (b *fakeBackend) Predict(context.Context, *pb.PredictOptions, ...ggrpc.CallOption) (*pb.Reply, error) { - b.predicted = true - return &pb.Reply{}, b.predictErr -} - -var _ = Describe("DefaultProber", func() { - var ( - up *fakeUpstream - p *DefaultProber - ctx = context.Background() - ) - - BeforeEach(func() { - up = newFakeUpstream() - DeferCleanup(up.srv.Close) - p = NewProber(nil, "") - }) - - proxied := func(name, upstreamModel string, usecases ...string) config.ModelConfig { - c := config.ModelConfig{Name: name, Backend: "cloud-proxy", KnownUsecaseStrings: usecases} - c.KnownUsecases = config.GetUsecasesFromYAML(usecases) - c.Proxy.UpstreamURL = up.srv.URL + "/v1/chat/completions" - c.Proxy.UpstreamModel = upstreamModel - return c - } - - DescribeTable("UpstreamBase", - func(in, want string) { - got, err := UpstreamBase(in) - Expect(err).ToNot(HaveOccurred()) - Expect(got).To(Equal(want)) - }, - Entry("full endpoint", "https://h:8080/v1/chat/completions", "https://h:8080"), - Entry("path prefix", "https://h/api/v1/chat/completions", "https://h/api"), - Entry("bare host", "https://h", "https://h"), - Entry("bare host slash", "https://h/", "https://h"), - ) - - It("passes liveness when the upstream lists the model", func() { - up.models = []string{"big-llm"} - Expect(p.Liveness(ctx, proxied("argus-llm", "big-llm"), KindRemote, false)).To(Succeed()) - Expect(up.paths).To(ContainElement("GET /v1/models")) - }) - - It("uses the target name when upstream_model is empty", func() { - up.models = []string{"argus-llm"} - Expect(p.Liveness(ctx, proxied("argus-llm", ""), KindRemote, false)).To(Succeed()) - }) - - It("fails liveness when the model is not listed or the upstream errors", func() { - up.models = []string{"other"} - Expect(p.Liveness(ctx, proxied("argus-llm", ""), KindRemote, false)).To(MatchError(ContainSubstring("does not list"))) - up.status = http.StatusServiceUnavailable - Expect(p.Liveness(ctx, proxied("argus-llm", ""), KindRemote, false)).To(MatchError(ContainSubstring("503"))) - }) - - It("sends the API key as a bearer token", func() { - GinkgoT().Setenv("FAILOVER_PROBE_KEY", "sekret") - up.models = []string{"argus-llm"} - c := proxied("argus-llm", "") - c.Proxy.APIKeyEnv = "FAILOVER_PROBE_KEY" - Expect(p.Liveness(ctx, c, KindRemote, false)).To(Succeed()) - Expect(up.auth).To(Equal("Bearer sekret")) - }) - - DescribeTable("remote inference hits the usecase endpoint", - func(usecase, path string) { - Expect(p.Inference(ctx, proxied("m", "", usecase), KindRemote, false)).To(Succeed()) - Expect(up.paths).To(ContainElement("POST " + path)) - }, - Entry("chat", "chat", "/v1/chat/completions"), - Entry("embeddings", "embeddings", "/v1/embeddings"), - Entry("transcription", "transcript", "/v1/audio/transcriptions"), - Entry("tts", "tts", "/v1/audio/speech"), - ) - - It("uses HealthCheck for warm local liveness and Predict for local chat inference", func() { - b := &fakeBackend{healthy: true} - p = NewProber(func(context.Context, config.ModelConfig) (grpc.Backend, error) { return b, nil }, "") - c := config.ModelConfig{Name: "gemma", Backend: "llama-cpp", KnownUsecaseStrings: []string{"chat"}} - c.KnownUsecases = config.GetUsecasesFromYAML(c.KnownUsecaseStrings) - Expect(p.Liveness(ctx, c, KindLocal, true)).To(Succeed()) - b.healthy = false - Expect(p.Liveness(ctx, c, KindLocal, true)).To(HaveOccurred()) - Expect(p.Inference(ctx, c, KindLocal, true)).To(Succeed()) - Expect(b.predicted).To(BeTrue()) - b.predictErr = errors.New("boom") - Expect(p.Inference(ctx, c, KindLocal, true)).To(HaveOccurred()) - }) - - It("checks the model file for cold local liveness without loading", func() { - dir := GinkgoT().TempDir() - p = NewProber(func(context.Context, config.ModelConfig) (grpc.Backend, error) { - Fail("cold liveness must not load the model") - return nil, nil - }, dir) - c := config.ModelConfig{Name: "cold", Backend: "llama-cpp"} - c.Model = "weights.gguf" - Expect(p.Liveness(ctx, c, KindLocal, false)).To(HaveOccurred()) - Expect(os.WriteFile(filepath.Join(dir, "weights.gguf"), []byte("x"), 0o600)).To(Succeed()) - Expect(p.Liveness(ctx, c, KindLocal, false)).To(Succeed()) - c.Model = "org/some-hf-repo" // no extension: downloaded on demand - Expect(p.Liveness(ctx, c, KindLocal, false)).To(Succeed()) - }) -}) -``` - -If `c.Model` is not directly assignable (it lives in an embedded struct), set it through that struct, e.g. `c.PredictionOptions.Model = "weights.gguf"`; use whatever `validateFailover` reads as `c.Model`. - -- [ ] **Step 2: Run to verify failure** - -Run: `go test ./core/config/... ./core/services/failover/... 2>&1 | tail -5` -Expected: compile failure. - -- [ ] **Step 3: Implement `ResolveAPIKey`** - -In `core/config/model_config.go` after the `ProxyConfig` constants: - -```go -// ResolveAPIKey returns the upstream key from api_key_env or api_key_file, or -// "" when neither is set. The cloud-proxy backend applies the same rules. -func (p ProxyConfig) ResolveAPIKey() (string, error) { - switch { - case p.APIKeyEnv != "": - v, ok := os.LookupEnv(p.APIKeyEnv) - if !ok { - return "", fmt.Errorf("proxy api_key_env %q is not set", p.APIKeyEnv) - } - return v, nil - case p.APIKeyFile != "": - b, err := os.ReadFile(p.APIKeyFile) - if err != nil { - return "", fmt.Errorf("proxy api_key_file: %w", err) - } - return strings.TrimSpace(string(b)), nil - } - return "", nil -} -``` - -- [ ] **Step 4: Implement the prober** - -`core/services/failover/prober.go`: - -```go -package failover - -import ( - "bytes" - "context" - "encoding/binary" - "encoding/json" - "errors" - "fmt" - "io" - "io/fs" - "mime/multipart" - "net/http" - "net/url" - "os" - "path/filepath" - "strings" - - "github.com/mudler/LocalAI/core/config" - "github.com/mudler/LocalAI/pkg/grpc" - pb "github.com/mudler/LocalAI/pkg/grpc/proto" -) - -// LoadFunc returns the backend for a local target, loading it if needed. -type LoadFunc func(ctx context.Context, cfg config.ModelConfig) (grpc.Backend, error) - -// DefaultProber probes remote targets over the upstream's OpenAI-compatible -// API and local targets through their gRPC backend. -type DefaultProber struct { - HTTP *http.Client - Load LoadFunc - ModelPath string -} - -func NewProber(load LoadFunc, modelPath string) *DefaultProber { - return &DefaultProber{HTTP: &http.Client{}, Load: load, ModelPath: modelPath} -} - -func (p *DefaultProber) Liveness(ctx context.Context, cfg config.ModelConfig, kind Kind, warm bool) error { - switch { - case kind == KindRemote: - return p.remoteLiveness(ctx, cfg) - case warm: - return p.localHealth(ctx, cfg) - } - return p.coldLiveness(cfg) -} - -func (p *DefaultProber) Inference(ctx context.Context, cfg config.ModelConfig, kind Kind, warm bool) error { - if kind == KindRemote { - return p.remoteInference(ctx, cfg) - } - return p.localInference(ctx, cfg) -} - -// UpstreamBase strips the endpoint path from a cloud-proxy upstream_url: -// everything from "/v1" on, so a path prefix before it survives. -func UpstreamBase(raw string) (string, error) { - u, err := url.Parse(raw) - if err != nil || u.Scheme == "" || u.Host == "" { - return "", fmt.Errorf("invalid upstream_url %q", raw) - } - path := u.Path - if i := strings.Index(path, "/v1"); i >= 0 { - path = path[:i] - } - return u.Scheme + "://" + u.Host + strings.TrimSuffix(path, "/"), nil -} - -// UpstreamModel is the model name the upstream knows the target by. -func UpstreamModel(cfg config.ModelConfig) string { - if cfg.Proxy.UpstreamModel != "" { - return cfg.Proxy.UpstreamModel - } - return cfg.Name -} - -func (p *DefaultProber) authorize(req *http.Request, cfg config.ModelConfig) error { - key, err := cfg.Proxy.ResolveAPIKey() - if err != nil || key == "" { - return err - } - if cfg.Proxy.Provider == config.ProxyProviderAnthropic { - req.Header.Set("x-api-key", key) - req.Header.Set("anthropic-version", "2023-06-01") - return nil - } - req.Header.Set("Authorization", "Bearer "+key) - return nil -} - -func (p *DefaultProber) do(req *http.Request, cfg config.ModelConfig) (*http.Response, error) { - if err := p.authorize(req, cfg); err != nil { - return nil, err - } - resp, err := p.HTTP.Do(req) - if err != nil { - return nil, err - } - if resp.StatusCode/100 != 2 { - resp.Body.Close() - return nil, fmt.Errorf("upstream %s: HTTP %d", req.URL.Path, resp.StatusCode) - } - return resp, nil -} - -func (p *DefaultProber) remoteLiveness(ctx context.Context, cfg config.ModelConfig) error { - base, err := UpstreamBase(cfg.Proxy.UpstreamURL) - if err != nil { - return err - } - req, err := http.NewRequestWithContext(ctx, http.MethodGet, base+"/v1/models", nil) - if err != nil { - return err - } - resp, err := p.do(req, cfg) - if err != nil { - return err - } - defer resp.Body.Close() - var list struct { - Data []struct { - ID string `json:"id"` - } `json:"data"` - } - if err := json.NewDecoder(io.LimitReader(resp.Body, 4<<20)).Decode(&list); err != nil { - return fmt.Errorf("upstream /v1/models: %w", err) - } - want := UpstreamModel(cfg) - for _, d := range list.Data { - if d.ID == want { - return nil - } - } - return fmt.Errorf("upstream does not list model %q", want) -} - -func (p *DefaultProber) remoteInference(ctx context.Context, cfg config.ModelConfig) error { - base, err := UpstreamBase(cfg.Proxy.UpstreamURL) - if err != nil { - return err - } - model := UpstreamModel(cfg) - ping := []map[string]string{{"role": "user", "content": "ping"}} - switch { - case cfg.HasUsecases(config.FLAG_CHAT) || cfg.HasUsecases(config.FLAG_COMPLETION): - if cfg.Proxy.Provider == config.ProxyProviderAnthropic { - return p.postJSON(ctx, cfg, base+"/v1/messages", map[string]any{"model": model, "max_tokens": 1, "messages": ping}) - } - return p.postJSON(ctx, cfg, base+"/v1/chat/completions", map[string]any{"model": model, "max_tokens": 1, "messages": ping}) - case cfg.HasUsecases(config.FLAG_EMBEDDINGS): - return p.postJSON(ctx, cfg, base+"/v1/embeddings", map[string]any{"model": model, "input": "ping"}) - case cfg.HasUsecases(config.FLAG_TRANSCRIPT): - return p.postTranscription(ctx, cfg, base, model) - case cfg.HasUsecases(config.FLAG_TTS): - return p.postJSON(ctx, cfg, base+"/v1/audio/speech", map[string]any{"model": model, "input": "ok"}) - } - // Image, video and other costly usecases: liveness is the confirmation. - return p.remoteLiveness(ctx, cfg) -} - -func (p *DefaultProber) postJSON(ctx context.Context, cfg config.ModelConfig, endpoint string, body any) error { - b, err := json.Marshal(body) - if err != nil { - return err - } - req, err := http.NewRequestWithContext(ctx, http.MethodPost, endpoint, bytes.NewReader(b)) - if err != nil { - return err - } - req.Header.Set("Content-Type", "application/json") - resp, err := p.do(req, cfg) - if err != nil { - return err - } - _, _ = io.Copy(io.Discard, io.LimitReader(resp.Body, 1<<20)) - return resp.Body.Close() -} - -func (p *DefaultProber) postTranscription(ctx context.Context, cfg config.ModelConfig, base, model string) error { - var buf bytes.Buffer - mw := multipart.NewWriter(&buf) - _ = mw.WriteField("model", model) - fw, err := mw.CreateFormFile("file", "probe.wav") - if err != nil { - return err - } - _, _ = fw.Write(silenceWAV()) - if err := mw.Close(); err != nil { - return err - } - req, err := http.NewRequestWithContext(ctx, http.MethodPost, base+"/v1/audio/transcriptions", &buf) - if err != nil { - return err - } - req.Header.Set("Content-Type", mw.FormDataContentType()) - resp, err := p.do(req, cfg) - if err != nil { - return err - } - _, _ = io.Copy(io.Discard, io.LimitReader(resp.Body, 1<<20)) - return resp.Body.Close() -} - -// silenceWAV is 200 ms of 16 kHz mono 16-bit silence. -func silenceWAV() []byte { - const rate, samples = 16000, 3200 - data := samples * 2 - b := make([]byte, 44+data) - copy(b[0:], "RIFF") - binary.LittleEndian.PutUint32(b[4:], uint32(36+data)) - copy(b[8:], "WAVE") - copy(b[12:], "fmt ") - binary.LittleEndian.PutUint32(b[16:], 16) - binary.LittleEndian.PutUint16(b[20:], 1) // PCM - binary.LittleEndian.PutUint16(b[22:], 1) // mono - binary.LittleEndian.PutUint32(b[24:], rate) - binary.LittleEndian.PutUint32(b[28:], rate*2) - binary.LittleEndian.PutUint16(b[32:], 2) - binary.LittleEndian.PutUint16(b[34:], 16) - copy(b[36:], "data") - binary.LittleEndian.PutUint32(b[40:], uint32(data)) - return b -} - -func (p *DefaultProber) localHealth(ctx context.Context, cfg config.ModelConfig) error { - if p.Load == nil { - return errors.New("failover: no backend loader configured") - } - // Load returns the running backend, or starts it again after a crash. - b, err := p.Load(ctx, cfg) - if err != nil { - return err - } - ok, err := b.HealthCheck(ctx) - if err != nil { - return err - } - if !ok { - return errors.New("backend health check failed") - } - return nil -} - -func (p *DefaultProber) localInference(ctx context.Context, cfg config.ModelConfig) error { - if p.Load == nil { - return errors.New("failover: no backend loader configured") - } - b, err := p.Load(ctx, cfg) - if err != nil { - return err - } - switch { - case cfg.HasUsecases(config.FLAG_CHAT) || cfg.HasUsecases(config.FLAG_COMPLETION): - _, err = b.Predict(ctx, &pb.PredictOptions{Prompt: "ping", Tokens: 1}) - return err - case cfg.HasUsecases(config.FLAG_EMBEDDINGS): - _, err = b.Embeddings(ctx, &pb.PredictOptions{Embeddings: "ping"}) - return err - } - // A backend process that answers HealthCheck rarely fails only for TTS or - // transcription, so a real request adds little here. - return p.localHealth(ctx, cfg) -} - -// coldLiveness checks the model file without loading the model. -func (p *DefaultProber) coldLiveness(cfg config.ModelConfig) error { - f := cfg.Model - if f == "" || p.ModelPath == "" || strings.Contains(f, "://") { - return nil - } - path := f - if !filepath.IsAbs(path) { - path = filepath.Join(p.ModelPath, f) - } - if _, err := os.Stat(path); err != nil { - if errors.Is(err, fs.ErrNotExist) && filepath.Ext(f) == "" { - return nil // a repository id, downloaded on demand - } - return fmt.Errorf("model file %s: %w", f, err) - } - return nil -} -``` - -Check the `pb.PredictOptions` field names (`Prompt`, `Tokens`, `Embeddings`) in `pkg/grpc/proto` and adjust only the field names if they differ. - -- [ ] **Step 5: Align the spec with cold liveness** - -In the spec's probe table, change the "local, cold" liveness cell to: "the model file exists (skipped for URLs and repository ids). The model is never loaded only to probe it." (The installed-backend check is left out: the loader installs backends on demand, so a missing backend is not a health signal.) `git add -f` the spec. - -- [ ] **Step 6: Run tests** - -Run: `go test -race ./core/config/... ./core/services/failover/... 2>&1 | tail -10` -Expected: PASS. - -- [ ] **Step 7: Commit** - -```bash -git add core/config core/services/failover -git add -f docs/superpowers/specs -git commit -m "feat(failover): probe remote targets over HTTP and local ones over gRPC - -Remote liveness uses /v1/models, which every OpenAI-compatible upstream -serves. Recovery sends one minimal request for the target's usecase. - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 7: Application wiring, warm targets, metrics, traces - -**Files:** -- Create: `core/services/failover/metrics.go`, `core/services/failover/trace.go`, `core/application/failover.go` -- Modify: `core/services/failover/manager.go` (`emitLocked` records metrics), `core/application/application.go` (field + accessor), `core/application/startup.go` (construct + run), `core/application/watchdog.go` (`SyncPinnedModelsToWatchdog`), `core/trace/backend_trace.go` (new type) -- Test: `core/services/failover/metrics_test.go` - -**Interfaces:** -- Consumes: Tasks 4-6. -- Produces: `(*application.Application).FailoverManager() *failover.Manager`, `failover.RegisterMetrics(m *Manager)`, `failover.RecordAttemptTrace(enabled bool, chain, target string, err error)`, `trace.BackendTraceFailover`. - -- [ ] **Step 1: Write the failing test** - -`core/services/failover/metrics_test.go`: - -```go -package failover - -import ( - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" -) - -var _ = Describe("metrics", func() { - It("registers and records without a meter provider", func() { - src := newFakeSource(remote("a"), local("b"), chainCfg("chain", nil, t("a"), t("b"))) - m := New(src, WithClock(newFakeClock())) - Expect(func() { RegisterMetrics(m) }).ToNot(Panic()) - Expect(func() { m.ReportFailure("a", errBoom) }).ToNot(Panic()) - }) - It("records an attempt trace only when enabled", func() { - Expect(func() { RecordAttemptTrace(false, "chain", "a", errBoom) }).ToNot(Panic()) - Expect(func() { RecordAttemptTrace(true, "chain", "a", errBoom) }).ToNot(Panic()) - }) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -Run: `go test ./core/services/failover/... 2>&1 | tail -5` -Expected: `undefined: RegisterMetrics`. - -- [ ] **Step 3: Implement metrics and traces** - -`core/services/failover/metrics.go`: - -```go -package failover - -import ( - "context" - "sync" - - "go.opentelemetry.io/otel" - "go.opentelemetry.io/otel/attribute" - "go.opentelemetry.io/otel/metric" -) - -var ( - metricsOnce sync.Once - switches metric.Int64Counter -) - -func initMetrics() { - metricsOnce.Do(func() { - meter := otel.Meter("github.com/mudler/LocalAI") - switches, _ = meter.Int64Counter("localai_failover_switches_total", - metric.WithDescription("Failover chain switches between targets")) - }) -} - -func recordSwitch(ev Event) { - initMetrics() - if switches == nil { - return - } - switches.Add(context.Background(), 1, metric.WithAttributes( - attribute.String("chain", ev.Chain), - attribute.String("from", ev.From), - attribute.String("to", ev.To), - attribute.String("reason", string(ev.Reason)), - )) -} - -// RegisterMetrics exports target health as a gauge. The application calls it -// once for its manager; tests create many managers and skip it. -func RegisterMetrics(m *Manager) { - meter := otel.Meter("github.com/mudler/LocalAI") - _, _ = meter.Int64ObservableGauge("localai_failover_target_up", - metric.WithDescription("1 when a failover target is healthy, 0 otherwise"), - metric.WithInt64Callback(func(_ context.Context, o metric.Int64Observer) error { - for name, state := range m.targetStates() { - v := int64(0) - if state == StateHealthy { - v = 1 - } - o.Observe(v, metric.WithAttributes(attribute.String("target", name))) - } - return nil - })) -} -``` - -Add to `manager.go`: - -```go -func (m *Manager) targetStates() map[string]TargetState { - m.mu.Lock() - defer m.mu.Unlock() - out := make(map[string]TargetState, len(m.targets)) - for name, ts := range m.targets { - out[name] = ts.state - } - return out -} -``` - -and at the top of `emitLocked`: - -```go - if ev.Type == EventChainSwitched { - recordSwitch(ev) - } -``` - -In `core/trace/backend_trace.go`, add to the `BackendTraceType` constants: - -```go - BackendTraceFailover BackendTraceType = "failover" -``` - -Check `core/http/react-ui/src` for a label map of backend trace types (`grep -rn "image_generation" core/http/react-ui/src`); if one exists, add `failover: 'Failover'` in the same style. - -`core/services/failover/trace.go`: - -```go -package failover - -import ( - "fmt" - "time" - - "github.com/mudler/LocalAI/core/trace" -) - -// RecordAttemptTrace shows in the Traces UI why a target was skipped. -func RecordAttemptTrace(enabled bool, chain, target string, err error) { - if !enabled || err == nil { - return - } - trace.RecordBackendTrace(trace.BackendTrace{ - Timestamp: time.Now(), - Type: trace.BackendTraceFailover, - ModelName: target, - Summary: fmt.Sprintf("failover chain %s: %s failed, trying the next target", chain, target), - Error: err.Error(), - Data: map[string]any{"chain": chain}, - }) -} -``` - -(If `BackendTrace` field names differ, match `core/trace/backend_trace.go`.) - -- [ ] **Step 4: Wire the application** - -`core/application/application.go`: add field `failoverManager *failover.Manager` to `Application` and: - -```go -// FailoverManager serves failover chains. Never nil after New. -func (a *Application) FailoverManager() *failover.Manager { return a.failoverManager } -``` - -`core/application/failover.go`: - -```go -package application - -import ( - "github.com/mudler/LocalAI/core/backend" - "github.com/mudler/xlog" -) - -// applyFailoverWarmTargets pins warm failover targets in the watchdog and -// loads them, so a switch does not wait for a cold load. -func (a *Application) applyFailoverWarmTargets(warm []string) { - a.SyncPinnedModelsToWatchdog() - for _, name := range warm { - if _, err := backend.PreloadModelByName(a.ApplicationConfig().Context, a.ModelConfigLoader(), a.ModelLoader(), a.ApplicationConfig(), name); err != nil { - xlog.Warn("failover: could not preload warm target", "model", name, "error", err) - } - } -} -``` - -`core/application/watchdog.go` in `SyncPinnedModelsToWatchdog`, before `wd.SetPinnedModels(pinned)`: - -```go - if a.failoverManager != nil { - pinned = failover.MergePinned(pinned, a.failoverManager.WarmTargets()) - } -``` - -`core/application/startup.go`: next to `application.routerRegistry = router.NewRegistry()` (~251), construct the manager: - -```go - application.failoverManager = failover.New(application.ModelConfigLoader(), - failover.WithProber(failover.NewProber(func(ctx context.Context, cfg config.ModelConfig) (grpc.Backend, error) { - return application.ModelLoader().Load(backend.ModelOptions(cfg, options)...) - }, application.ModelLoader().ModelPath)), - failover.WithOnWarmChanged(application.applyFailoverWarmTargets), - ) -``` - -and directly after the `LoadToMemory` preload loop (~539-548), which runs after `initializeWatchdog`, start it: - -```go - failover.RegisterMetrics(application.failoverManager) - go application.failoverManager.Run(options.Context) -``` - -`grpc` here is `github.com/mudler/LocalAI/pkg/grpc`. If `startup.go` already imports a different package as `grpc`, alias this one `lagrpc`. - -- [ ] **Step 5: Build and test** - -Run: `go build ./... && go test ./core/services/failover/... ./core/application/... 2>&1 | tail -10` -Expected: build OK, tests PASS. - -- [ ] **Step 6: Commit** - -```bash -git add core/services/failover core/application core/trace core/http/react-ui/src -git commit -m "feat(failover): run the chain manager and keep warm targets loaded - -Warm local targets are pinned in the watchdog and preloaded. Switches -and target health are exported as metrics, skipped attempts as traces. - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 8: HTTP chain resolution and in-request retry - -**Files:** -- Create: `core/http/middleware/failover.go`, `core/http/middleware/failover_test.go` -- Modify: `core/http/middleware/request.go` (`RequestExtractor` field; `SetModelAndConfig` wraps its body with `failoverRetry`; resolution after the alias block at ~181-195) -- Modify: `core/http/middleware/context_keys.go` (new key) -- Modify: `core/http/app.go:457` (call `SetFailoverManager`) - -**Interfaces:** -- Consumes: `failover.Manager.Plan`, `Attempt` methods, `failover.IsRetryable`, `failover.RecordAttemptTrace`. -- Produces: `(*RequestExtractor).SetFailoverManager(*failover.Manager)`, `ContextKeyFailoverAttempt = "failover.attempt"`, `HeaderServedModel = "X-LocalAI-Served-Model"`, `HeaderFailover = "X-LocalAI-Failover"`, `MaxFailoverReplayBody = 32 << 20`. - -Design note for the implementer: the retry loop wraps the whole `SetModelAndConfig` body plus `next`, so every attempt binds the request again from the replayed body. No route file changes. `failoverWriter` holds back a response with status >= 500 only while the request is resolved to a chain, so other requests are unaffected. - -- [ ] **Step 1: Write the failing tests** - -`core/http/middleware/failover_test.go` (use the same package and suite as `request_config_revision_test.go`): - -```go -package middleware - -import ( - "bytes" - "context" - "errors" - "io" - "mime/multipart" - "net/http" - "net/http/httptest" - "os" - "path/filepath" - "strings" - "sync" - - "github.com/labstack/echo/v4" - "github.com/mudler/LocalAI/core/config" - "github.com/mudler/LocalAI/core/schema" - "github.com/mudler/LocalAI/core/services/failover" - "github.com/mudler/LocalAI/pkg/model" - "github.com/mudler/LocalAI/pkg/system" - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" -) - -var _ = Describe("failover chains in the request pipeline", func() { - var ( - app *echo.Echo - fm *failover.Manager - mu sync.Mutex - calls []string - behavior map[string]func(c echo.Context) error - ) - - served := func(c echo.Context) error { - cfg := c.Get(CONTEXT_LOCALS_KEY_MODEL_CONFIG).(*config.ModelConfig) - return c.JSON(http.StatusOK, map[string]string{"served": cfg.Name}) - } - - handler := func(c echo.Context) error { - cfg := c.Get(CONTEXT_LOCALS_KEY_MODEL_CONFIG).(*config.ModelConfig) - mu.Lock() - calls = append(calls, cfg.Name) - b := behavior[cfg.Name] - mu.Unlock() - if b == nil { - return served(c) - } - return b(c) - } - - post := func(path, body string) *httptest.ResponseRecorder { - req := httptest.NewRequest(http.MethodPost, path, strings.NewReader(body)) - req.Header.Set("Content-Type", "application/json") - rec := httptest.NewRecorder() - app.ServeHTTP(rec, req) - return rec - } - chat := func(model string) *httptest.ResponseRecorder { - return post("/v1/chat/completions", `{"model":"`+model+`","messages":[{"role":"user","content":"hi"}]}`) - } - - BeforeEach(func() { - calls = nil - behavior = map[string]func(c echo.Context) error{} - dir := GinkgoT().TempDir() - write := func(name, body string) { - Expect(os.WriteFile(filepath.Join(dir, name+".yaml"), []byte(body), 0o600)).To(Succeed()) - } - write("a", "name: a\nbackend: fake-a\n") - write("b", "name: b\nbackend: fake-b\n") - write("plain", "name: plain\nbackend: fake-p\n") - write("chain", "name: chain\nfailover:\n targets:\n - model: a\n - model: b\n") - - ss := &system.SystemState{Model: system.Model{ModelsPath: dir}} - appConfig := config.NewApplicationConfig() - appConfig.SystemState = ss - mcl := config.NewModelConfigLoader(dir) - Expect(mcl.LoadModelConfigsFromPath(dir)).To(Succeed()) - re := NewRequestExtractor(mcl, model.NewModelLoader(ss), appConfig) - fm = failover.New(mcl) - re.SetFailoverManager(fm) - - app = echo.New() - // echo's default handler hides internal error messages; the specs - // below check which target's error reached the client. - app.HTTPErrorHandler = func(err error, c echo.Context) { - code := http.StatusInternalServerError - var he *echo.HTTPError - if errors.As(err, &he) { - code = he.Code - } - _ = c.JSON(code, map[string]string{"error": err.Error()}) - } - app.POST("/v1/chat/completions", handler, - re.SetModelAndConfig(func() schema.LocalAIRequest { return new(schema.OpenAIRequest) })) - app.POST("/v1/audio/transcriptions", handler, - re.SetModelAndConfig(func() schema.LocalAIRequest { return new(schema.OpenAIRequest) })) - }) - - It("serves from the next target when the first fails before responding", func() { - behavior["a"] = func(echo.Context) error { return errors.New("dial tcp: connection refused") } - rec := chat("chain") - Expect(rec.Code).To(Equal(http.StatusOK)) - Expect(rec.Body.String()).To(ContainSubstring(`"served":"b"`)) - Expect(rec.Header().Get(HeaderServedModel)).To(Equal("b")) - Expect(rec.Header().Get(HeaderFailover)).To(Equal("fallback")) - Expect(calls).To(Equal([]string{"a", "b"})) - st, _ := fm.ChainStatus("chain") - Expect(st.Active).To(Equal("b")) - }) - - It("serves the primary without the failover header", func() { - rec := chat("chain") - Expect(rec.Header().Get(HeaderServedModel)).To(Equal("a")) - Expect(rec.Header().Get(HeaderFailover)).To(BeEmpty()) - }) - - It("drops a buffered 5xx response and retries", func() { - behavior["a"] = func(c echo.Context) error { - return c.JSON(http.StatusServiceUnavailable, map[string]string{"error": "no healthy nodes"}) - } - rec := chat("chain") - Expect(rec.Code).To(Equal(http.StatusOK)) - Expect(rec.Body.String()).ToNot(ContainSubstring("no healthy nodes")) - }) - - It("does not retry after streaming started, and trips the target", func() { - behavior["a"] = func(c echo.Context) error { - c.Response().Header().Set("Content-Type", "text/event-stream") - _, _ = c.Response().Write([]byte("data: x\n\n")) - c.Response().Flush() - return errors.New("connection reset by peer") - } - rec := chat("chain") - Expect(rec.Body.String()).To(HavePrefix("data: x")) - Expect(calls).To(Equal([]string{"a"})) - st, _ := fm.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(failover.StateDown)) - }) - - It("does not retry or trip on 4xx", func() { - behavior["a"] = func(echo.Context) error { return echo.NewHTTPError(http.StatusBadRequest, "bad") } - rec := chat("chain") - Expect(rec.Code).To(Equal(http.StatusBadRequest)) - Expect(calls).To(Equal([]string{"a"})) - st, _ := fm.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(failover.StateHealthy)) - }) - - It("does not retry when the client cancelled", func() { - ctx, cancel := context.WithCancel(context.Background()) - behavior["a"] = func(echo.Context) error { cancel(); return context.Canceled } - req := httptest.NewRequest(http.MethodPost, "/v1/chat/completions", - strings.NewReader(`{"model":"chain","messages":[{"role":"user","content":"hi"}]}`)).WithContext(ctx) - req.Header.Set("Content-Type", "application/json") - app.ServeHTTP(httptest.NewRecorder(), req) - Expect(calls).To(Equal([]string{"a"})) - st, _ := fm.ChainStatus("chain") - Expect(st.Targets[0].State).To(Equal(failover.StateHealthy)) - }) - - It("gives each attempt a fresh request", func() { - behavior["a"] = func(c echo.Context) error { - in := c.Get(CONTEXT_LOCALS_KEY_LOCALAI_REQUEST).(*schema.OpenAIRequest) - in.Messages = nil - return errors.New("dial tcp: connection refused") - } - behavior["b"] = func(c echo.Context) error { - in := c.Get(CONTEXT_LOCALS_KEY_LOCALAI_REQUEST).(*schema.OpenAIRequest) - Expect(in.Messages).To(HaveLen(1)) - return served(c) - } - Expect(chat("chain").Code).To(Equal(http.StatusOK)) - }) - - It("degraded: tries every target in priority order and returns the last error", func() { - behavior["a"] = func(echo.Context) error { return errors.New("dial tcp: a down") } - behavior["b"] = func(echo.Context) error { return errors.New("dial tcp: b down") } - chat("chain") // trips both - calls = nil - rec := chat("chain") - Expect(calls).To(Equal([]string{"a", "b"})) - Expect(rec.Code).To(Equal(http.StatusInternalServerError)) - Expect(rec.Body.String()).To(ContainSubstring("b down")) - Expect(rec.Header().Get(HeaderFailover)).To(Equal("degraded")) - }) - - It("replays a multipart body for the next target", func() { - var body bytes.Buffer - mw := multipart.NewWriter(&body) - _ = mw.WriteField("model", "chain") - fw, _ := mw.CreateFormFile("file", "a.wav") - _, _ = fw.Write(bytes.Repeat([]byte{7}, 4096)) - Expect(mw.Close()).To(Succeed()) - size := func(c echo.Context) int64 { - fh, err := c.FormFile("file") - Expect(err).ToNot(HaveOccurred()) - f, _ := fh.Open() - n, _ := io.Copy(io.Discard, f) - return n - } - behavior["a"] = func(c echo.Context) error { size(c); return errors.New("dial tcp: refused") } - behavior["b"] = func(c echo.Context) error { - Expect(size(c)).To(Equal(int64(4096))) - return served(c) - } - req := httptest.NewRequest(http.MethodPost, "/v1/audio/transcriptions", &body) - req.Header.Set("Content-Type", mw.FormDataContentType()) - rec := httptest.NewRecorder() - app.ServeHTTP(rec, req) - Expect(rec.Code).To(Equal(http.StatusOK)) - Expect(calls).To(Equal([]string{"a", "b"})) - }) - - It("leaves plain models untouched", func() { - behavior["plain"] = func(c echo.Context) error { - return c.JSON(http.StatusInternalServerError, map[string]string{"error": "boom"}) - } - rec := chat("plain") - Expect(rec.Code).To(Equal(http.StatusInternalServerError)) - Expect(rec.Body.String()).To(ContainSubstring("boom")) - Expect(rec.Header().Get(HeaderServedModel)).To(BeEmpty()) - }) -}) -``` - -If `SetModelAndConfig` rejects `backend: fake-a` configs in this fixture (for example through the existence check), give the fixture configs the backend the revision test uses (`llama-cpp`); the handler never loads a backend. - -- [ ] **Step 2: Run to verify failure** - -Run: `go test ./core/http/middleware/... 2>&1 | tail -5` -Expected: compile failure (`SetFailoverManager undefined`). - -- [ ] **Step 3: Implement** - -In `core/http/middleware/context_keys.go` add: - -```go - // ContextKeyFailoverAttempt holds the *failoverState of a request whose - // model is a failover chain. - ContextKeyFailoverAttempt = "failover.attempt" -``` - -`core/http/middleware/failover.go`: - -```go -package middleware - -import ( - "bufio" - "bytes" - "errors" - "fmt" - "io" - "net" - "net/http" - "slices" - "strings" - - "github.com/labstack/echo/v4" - "github.com/mudler/LocalAI/core/config" - "github.com/mudler/LocalAI/core/services/failover" -) - -const ( - HeaderServedModel = "X-LocalAI-Served-Model" - HeaderFailover = "X-LocalAI-Failover" -) - -// MaxFailoverReplayBody caps the request body kept for a retry. A larger body -// is still served, by one target only. -const MaxFailoverReplayBody = 32 << 20 - -type failoverState struct { - attempt *failover.Attempt -} - -// SetFailoverManager enables failover chains. Without it, a request for a -// chain fails with 503. -func (re *RequestExtractor) SetFailoverManager(m *failover.Manager) { re.failover = m } - -// resolveFailover returns the config of the target that should serve this -// attempt. The first attempt plans the chain; retries reuse the plan. -func (re *RequestExtractor) resolveFailover(c echo.Context, requested string, chain *config.ModelConfig) (*config.ModelConfig, error) { - st, _ := c.Get(ContextKeyFailoverAttempt).(*failoverState) - if st == nil || st.attempt.Chain() != chain.Name { - if re.failover == nil { - return nil, fmt.Errorf("model %q is a failover chain, but failover is not running", chain.Name) - } - att, err := re.failover.Plan(chain.Name) - if err != nil { - return nil, err - } - st = &failoverState{attempt: att} - c.Set(ContextKeyFailoverAttempt, st) - } - for { - cfg, err := re.loadFailoverTarget(st.attempt.Target()) - if err == nil { - c.Set(ContextKeyRequestedModel, requested) - c.Set(ContextKeyServedModel, cfg.Name) - setFailoverHeaders(c.Response().Header(), st.attempt) - return cfg, nil - } - if !st.attempt.Fail(err) { - // Clear the state so the retry wrapper sends this 503 as is. - c.Set(ContextKeyFailoverAttempt, nil) - return nil, err - } - } -} - -func (re *RequestExtractor) loadFailoverTarget(name string) (*config.ModelConfig, error) { - cfg, err := re.modelConfigLoader.LoadModelConfigFileByNameDefaultOptions(name, re.applicationConfig) - if err != nil { - return nil, err - } - resolved, _, err := re.modelConfigLoader.ResolveAlias(cfg) - return resolved, err -} - -func setFailoverHeaders(h http.Header, att *failover.Attempt) { - h.Set(HeaderServedModel, att.Target()) - switch { - case att.Degraded(): - h.Set(HeaderFailover, "degraded") - case att.Target() != att.Primary(): - h.Set(HeaderFailover, "fallback") - default: - h.Del(HeaderFailover) - } -} - -// failoverRetry runs h again on the next target while the response is not -// committed. h is SetModelAndConfig's body plus the rest of the chain, so -// every attempt binds the request again from the replayed body. -func failoverRetry(appConfig *config.ApplicationConfig, h echo.HandlerFunc) echo.HandlerFunc { - return func(c echo.Context) error { - req := c.Request() - src := req.Body - if src == nil { - src = http.NoBody - } - rec := &replayBody{src: src, limit: MaxFailoverReplayBody} - req.Body = rec - resp := c.Response() - orig := resp.Writer - baseHeader := resp.Header().Clone() - defer func() { resp.Writer = orig }() - active := func() bool { - st, _ := c.Get(ContextKeyFailoverAttempt).(*failoverState) - return st != nil - } - tracing := appConfig != nil && appConfig.EnableTracing - for { - w := &failoverWriter{ResponseWriter: orig, active: active} - resp.Writer = w - err := h(c) - st, _ := c.Get(ContextKeyFailoverAttempt).(*failoverState) - if st == nil { - w.release() - return err - } - att := st.attempt - status := w.held - if err == nil && status == 0 { - att.Succeed() - return nil - } - cause := attemptError(err, status, w.body.Bytes()) - retryable := req.Context().Err() == nil && failover.IsRetryable(err, status) - if !retryable || w.committed || !rec.replayable() { - if retryable { - att.Report(cause) - } - w.release() - return err - } - failover.RecordAttemptTrace(tracing, att.Chain(), att.Target(), cause) - if !att.Fail(cause) { - w.release() - return err - } - req.Body = rec.replay() - req.MultipartForm, req.Form, req.PostForm = nil, nil, nil - resetResponse(resp, baseHeader) - } - } -} - -func attemptError(err error, status int, body []byte) error { - if err != nil { - return err - } - msg := strings.TrimSpace(string(body)) - if len(msg) > 200 { - msg = msg[:200] - } - return fmt.Errorf("HTTP %d: %s", status, msg) -} - -func resetResponse(resp *echo.Response, base http.Header) { - h := resp.Header() - for k := range h { - delete(h, k) - } - for k, v := range base { - h[k] = slices.Clone(v) - } - resp.Committed = false - resp.Status = http.StatusOK - resp.Size = 0 -} - -// failoverWriter holds back an error response (status >= 500) of a chain -// request until the handler returns, so the retry can drop it. -type failoverWriter struct { - http.ResponseWriter - active func() bool - held int - body bytes.Buffer - committed bool -} - -func (w *failoverWriter) WriteHeader(code int) { - if w.held != 0 { - return - } - if !w.committed && code >= 500 && w.active() { - w.held = code - return - } - w.committed = true - w.ResponseWriter.WriteHeader(code) -} - -func (w *failoverWriter) Write(b []byte) (int, error) { - if w.held != 0 { - return w.body.Write(b) - } - w.committed = true - return w.ResponseWriter.Write(b) -} - -func (w *failoverWriter) Flush() { - w.release() - w.committed = true - if f, ok := w.ResponseWriter.(http.Flusher); ok { - f.Flush() - } -} - -func (w *failoverWriter) Hijack() (net.Conn, *bufio.ReadWriter, error) { - w.committed = true - h, ok := w.ResponseWriter.(http.Hijacker) - if !ok { - return nil, nil, errors.New("failover: response writer cannot hijack") - } - return h.Hijack() -} - -func (w *failoverWriter) Unwrap() http.ResponseWriter { return w.ResponseWriter } - -// release sends a held error response to the client. -func (w *failoverWriter) release() { - if w.held == 0 { - return - } - code := w.held - w.held = 0 - w.committed = true - w.ResponseWriter.WriteHeader(code) - _, _ = w.ResponseWriter.Write(w.body.Bytes()) - w.body.Reset() -} - -// replayBody records what the handler reads, up to limit, so the body can be -// sent again to the next target. Decoders often stop at the end of the value -// without reading to EOF, so a replay is the recorded bytes followed by -// whatever the previous attempt left unread. -type replayBody struct { - src io.ReadCloser - buf bytes.Buffer - limit int - overflow bool -} - -func (r *replayBody) Read(p []byte) (int, error) { - n, err := r.src.Read(p) - if n > 0 && !r.overflow { - if r.buf.Len()+n > r.limit { - r.overflow = true - r.buf.Reset() - } else { - r.buf.Write(p[:n]) - } - } - return n, err -} - -func (r *replayBody) Close() error { return r.src.Close() } - -// replayable reports whether everything read so far was kept. -func (r *replayBody) replayable() bool { return !r.overflow } - -// replay rewinds to the start of the body and keeps recording, so a third -// attempt can replay too. -func (r *replayBody) replay() io.ReadCloser { - data := bytes.Clone(r.buf.Bytes()) - rest := r.src - r.src = struct { - io.Reader - io.Closer - }{io.MultiReader(bytes.NewReader(data), rest), rest} - r.buf.Reset() - return r -} -``` - -In `core/http/middleware/request.go`: - -1. Add field `failover *failover.Manager` to `RequestExtractor`. -2. In `SetModelAndConfig`, wrap the returned handler: - -```go -func (re *RequestExtractor) SetModelAndConfig(initializer func() schema.LocalAIRequest) echo.MiddlewareFunc { - return func(next echo.HandlerFunc) echo.HandlerFunc { - return failoverRetry(re.applicationConfig, func(c echo.Context) error { - // ... existing body, unchanged ... - }) - } -} -``` - -3. After the alias block (after `cfg = resolved` at ~195) and before the disabled check, add: - -```go - // A failover chain resolves to one of its targets, like an alias. - // failoverRetry re-runs this middleware for the next target. - if cfg != nil && cfg.IsFailover() { - resolved, fErr := re.resolveFailover(c, modelName, cfg) - if fErr != nil { - return c.JSON(http.StatusServiceUnavailable, schema.ErrorResponse{ - Error: &schema.APIError{ - Message: fErr.Error(), - Code: http.StatusServiceUnavailable, - Type: "failover_unavailable", - }, - }) - } - cfg = resolved - } -``` - -In `core/http/app.go` after line 457: - -```go - requestExtractor.SetFailoverManager(application.FailoverManager()) -``` - -- [ ] **Step 4: Run tests** - -Run: `go test -race ./core/http/middleware/... 2>&1 | tail -15` -Expected: PASS, including the existing middleware specs. - -- [ ] **Step 5: Run the wider HTTP suite** - -Run: `make build-mock-backend && go test ./core/http/... 2>&1 | tail -15` -Expected: PASS (the retry wrapper is a no-op for plain models). - -- [ ] **Step 6: Commit** - -```bash -git add core/http -git commit -m "feat(failover): resolve chains per request and retry on the next target - -The retry wraps SetModelAndConfig, so each attempt binds the request -again from a replayed body. A 5xx of a chain request is held back until -the handler returns, and a streamed response is never retried. - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 9: REST and SSE endpoints - -**Files:** -- Create: `core/http/endpoints/localai/failover.go`, `core/http/endpoints/localai/failover_test.go` -- Modify: `core/http/routes/localai.go` (register routes), `core/http/endpoints/localai/api_instructions.go` (+ test count 19 → 20 in `api_instructions_test.go:42`) -- Test: auth coverage next to the existing aliases auth tests (find them with `grep -rn '"/api/aliases"' core/http --include=*_test.go`) - -**Interfaces:** -- Consumes: `failover.Manager` `Status`, `ChainStatus`, `Pin`, `Unpin`, `Subscribe`, errors. -- Produces: routes `GET /api/failover`, `GET /api/failover/events`, `GET /api/failover/:chain`, `POST /api/failover/:chain/pin`, `DELETE /api/failover/:chain/pin`; type `localai.FailoverChainsResponse`. - -- [ ] **Step 1: Write the failing tests** - -`core/http/endpoints/localai/failover_test.go` (match the package name and suite used by the other `_test.go` files in that directory): - -```go -package localai - -import ( - "bufio" - "context" - "encoding/json" - "net/http" - "net/http/httptest" - "strings" - "time" - - "github.com/labstack/echo/v4" - "github.com/mudler/LocalAI/core/config" - "github.com/mudler/LocalAI/core/services/failover" - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" -) - -type mapSource map[string]config.ModelConfig - -func (s mapSource) GetModelConfig(n string) (config.ModelConfig, bool) { c, ok := s[n]; return c, ok } -func (s mapSource) GetAllModelsConfigs() []config.ModelConfig { - var out []config.ModelConfig - for _, c := range s { - out = append(out, c) - } - return out -} - -var _ = Describe("failover endpoints", func() { - var ( - e *echo.Echo - fm *failover.Manager - ) - - BeforeEach(func() { - src := mapSource{ - "a": {Name: "a", Backend: "cloud-proxy"}, - "b": {Name: "b", Backend: "llama-cpp"}, - "chain": {Name: "chain", Failover: &config.FailoverConfig{Targets: []config.FailoverTarget{{Model: "a"}, {Model: "b"}}}}, - } - fm = failover.New(src) - e = echo.New() - e.GET("/api/failover", ListFailoverChainsEndpoint(fm)) - e.GET("/api/failover/events", FailoverEventsEndpoint(fm)) - e.GET("/api/failover/:chain", GetFailoverChainEndpoint(fm)) - e.POST("/api/failover/:chain/pin", PinFailoverTargetEndpoint(fm)) - e.DELETE("/api/failover/:chain/pin", UnpinFailoverTargetEndpoint(fm)) - }) - - do := func(method, path, body string) *httptest.ResponseRecorder { - req := httptest.NewRequest(method, path, strings.NewReader(body)) - req.Header.Set("Content-Type", "application/json") - rec := httptest.NewRecorder() - e.ServeHTTP(rec, req) - return rec - } - - It("lists chains", func() { - rec := do(http.MethodGet, "/api/failover", "") - Expect(rec.Code).To(Equal(http.StatusOK)) - var out FailoverChainsResponse - Expect(json.Unmarshal(rec.Body.Bytes(), &out)).To(Succeed()) - Expect(out.Chains).To(HaveLen(1)) - Expect(out.Chains[0].Active).To(Equal("a")) - }) - - It("gets one chain or 404", func() { - Expect(do(http.MethodGet, "/api/failover/chain", "").Code).To(Equal(http.StatusOK)) - Expect(do(http.MethodGet, "/api/failover/nope", "").Code).To(Equal(http.StatusNotFound)) - }) - - It("pins and unpins", func() { - rec := do(http.MethodPost, "/api/failover/chain/pin", `{"target":"b"}`) - Expect(rec.Code).To(Equal(http.StatusOK)) - st, _ := fm.ChainStatus("chain") - Expect(st.Active).To(Equal("b")) - Expect(do(http.MethodPost, "/api/failover/chain/pin", `{"target":"zzz"}`).Code).To(Equal(http.StatusBadRequest)) - Expect(do(http.MethodPost, "/api/failover/chain/pin", `{}`).Code).To(Equal(http.StatusBadRequest)) - Expect(do(http.MethodPost, "/api/failover/nope/pin", `{"target":"a"}`).Code).To(Equal(http.StatusNotFound)) - Expect(do(http.MethodDelete, "/api/failover/chain/pin", "").Code).To(Equal(http.StatusOK)) - st, _ = fm.ChainStatus("chain") - Expect(st.Pinned).To(BeNil()) - }) - - It("streams a snapshot, then switch events", func() { - srv := httptest.NewServer(e) - defer srv.Close() - ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second) - defer cancel() - req, _ := http.NewRequestWithContext(ctx, http.MethodGet, srv.URL+"/api/failover/events", nil) - resp, err := http.DefaultClient.Do(req) - Expect(err).ToNot(HaveOccurred()) - defer resp.Body.Close() - Expect(resp.Header.Get("Content-Type")).To(HavePrefix("text/event-stream")) - r := bufio.NewReader(resp.Body) - next := func() string { - for { - line, err := r.ReadString('\n') - Expect(err).ToNot(HaveOccurred()) - if strings.HasPrefix(line, "event: ") { - return strings.TrimSpace(strings.TrimPrefix(line, "event: ")) - } - } - } - Expect(next()).To(Equal("snapshot")) - Expect(fm.Pin("chain", "b")).To(Succeed()) - Expect(next()).To(Equal("chain.switched")) - }) -}) -``` - -- [ ] **Step 2: Run to verify failure** - -Run: `go test ./core/http/endpoints/localai/... 2>&1 | tail -5` -Expected: compile failure. - -- [ ] **Step 3: Implement the handlers** - -`core/http/endpoints/localai/failover.go`: - -```go -package localai - -import ( - "encoding/json" - "errors" - "fmt" - "net/http" - "time" - - "github.com/labstack/echo/v4" - "github.com/mudler/LocalAI/core/schema" - "github.com/mudler/LocalAI/core/services/failover" -) - -type FailoverChainsResponse struct { - Chains []failover.ChainStatus `json:"chains"` -} - -type FailoverPinRequest struct { - Target string `json:"target"` -} - -func failoverError(c echo.Context, code int, msg string) error { - return c.JSON(code, schema.ErrorResponse{Error: &schema.APIError{Message: msg, Code: code, Type: "failover_error"}}) -} - -// ListFailoverChainsEndpoint lists failover chains and the health of their targets -// -// @Summary List failover chains and the health of their targets -// @Tags failover -// @Produce json -// @Success 200 {object} FailoverChainsResponse -// @Router /api/failover [get] -func ListFailoverChainsEndpoint(fm *failover.Manager) echo.HandlerFunc { - return func(c echo.Context) error { - return c.JSON(http.StatusOK, FailoverChainsResponse{Chains: fm.Status()}) - } -} - -// GetFailoverChainEndpoint returns one failover chain -// -// @Summary Get one failover chain -// @Tags failover -// @Produce json -// @Param chain path string true "Chain name" -// @Success 200 {object} failover.ChainStatus -// @Failure 404 {object} schema.ErrorResponse -// @Router /api/failover/{chain} [get] -func GetFailoverChainEndpoint(fm *failover.Manager) echo.HandlerFunc { - return func(c echo.Context) error { - st, ok := fm.ChainStatus(c.Param("chain")) - if !ok { - return failoverError(c, http.StatusNotFound, fmt.Sprintf("failover chain %q not found", c.Param("chain"))) - } - return c.JSON(http.StatusOK, st) - } -} - -// PinFailoverTargetEndpoint forces a chain to one target -// -// @Summary Pin a failover chain to one target -// @Tags failover -// @Accept json -// @Produce json -// @Param chain path string true "Chain name" -// @Param request body FailoverPinRequest true "Target to pin" -// @Success 200 {object} failover.ChainStatus -// @Failure 400 {object} schema.ErrorResponse -// @Failure 404 {object} schema.ErrorResponse -// @Router /api/failover/{chain}/pin [post] -func PinFailoverTargetEndpoint(fm *failover.Manager) echo.HandlerFunc { - return func(c echo.Context) error { - var req FailoverPinRequest - if err := c.Bind(&req); err != nil || req.Target == "" { - return failoverError(c, http.StatusBadRequest, "request body must set \"target\"") - } - chain := c.Param("chain") - if err := fm.Pin(chain, req.Target); err != nil { - return pinError(c, err) - } - st, _ := fm.ChainStatus(chain) - return c.JSON(http.StatusOK, st) - } -} - -// UnpinFailoverTargetEndpoint removes a pin -// -// @Summary Remove the pin from a failover chain -// @Tags failover -// @Produce json -// @Param chain path string true "Chain name" -// @Success 200 {object} failover.ChainStatus -// @Failure 404 {object} schema.ErrorResponse -// @Router /api/failover/{chain}/pin [delete] -func UnpinFailoverTargetEndpoint(fm *failover.Manager) echo.HandlerFunc { - return func(c echo.Context) error { - chain := c.Param("chain") - if err := fm.Unpin(chain); err != nil { - return pinError(c, err) - } - st, _ := fm.ChainStatus(chain) - return c.JSON(http.StatusOK, st) - } -} - -func pinError(c echo.Context, err error) error { - switch { - case errors.Is(err, failover.ErrChainNotFound): - return failoverError(c, http.StatusNotFound, err.Error()) - case errors.Is(err, failover.ErrTargetNotInChain): - return failoverError(c, http.StatusBadRequest, err.Error()) - } - return failoverError(c, http.StatusInternalServerError, err.Error()) -} - -// FailoverEventsEndpoint streams failover events -// -// @Summary Stream failover events (server-sent events) -// @Description The first event is "snapshot" with the full state, then "chain.switched" and "target.state" events. -// @Tags failover -// @Produce text/event-stream -// @Success 200 -// @Router /api/failover/events [get] -func FailoverEventsEndpoint(fm *failover.Manager) echo.HandlerFunc { - return func(c echo.Context) error { - // Subscribe before the snapshot so no event falls between the two. - events, cancel := fm.Subscribe(64) - defer cancel() - w := c.Response() - w.Header().Set("Content-Type", "text/event-stream") - w.Header().Set("Cache-Control", "no-cache") - w.Header().Set("Connection", "keep-alive") - w.WriteHeader(http.StatusOK) - send := func(name string, v any) error { - data, err := json.Marshal(v) - if err != nil { - return err - } - if _, err := fmt.Fprintf(w, "event: %s\ndata: %s\n\n", name, data); err != nil { - return err - } - w.Flush() - return nil - } - if err := send("snapshot", FailoverChainsResponse{Chains: fm.Status()}); err != nil { - return nil - } - keepalive := time.NewTicker(15 * time.Second) - defer keepalive.Stop() - for { - select { - case <-c.Request().Context().Done(): - return nil - case <-keepalive.C: - if _, err := fmt.Fprint(w, ": keepalive\n\n"); err != nil { - return nil - } - w.Flush() - case ev, ok := <-events: - if !ok { - return nil - } - if err := send(string(ev.Type), ev); err != nil { - return nil - } - } - } - } -} -``` - -- [ ] **Step 4: Register routes and instructions** - -In `core/http/routes/localai.go`, next to `/api/aliases` (~92), where `app` / `application` is in scope (use the variable holding `*application.Application`): - -```go - fm := application.FailoverManager() - router.GET("/api/failover", localai.ListFailoverChainsEndpoint(fm)) - router.GET("/api/failover/events", localai.FailoverEventsEndpoint(fm)) - router.GET("/api/failover/:chain", localai.GetFailoverChainEndpoint(fm)) - router.POST("/api/failover/:chain/pin", localai.PinFailoverTargetEndpoint(fm), adminMiddleware) - router.DELETE("/api/failover/:chain/pin", localai.UnpinFailoverTargetEndpoint(fm), adminMiddleware) -``` - -In `api_instructions.go` `instructionDefs`, add: - -```go - { - Name: "failover", - Description: "Model failover chains: target health, pinning and switch events", - Tags: []string{"failover"}, - Intro: "A failover chain is a model config with a failover block. Requests for the chain name are served by its highest-priority healthy target; the X-LocalAI-Served-Model response header names it. Subscribe to GET /api/failover/events (SSE) to follow switches.", - }, -``` - -and change `HaveLen(19)` to `HaveLen(20)` in `api_instructions_test.go`. - -- [ ] **Step 5: Auth tests** - -Open the test that covers `/api/aliases` auth (found with the grep above) and add the same cases for: -- `GET /api/failover` without credentials → 401 when auth is enabled; with a user key → 200. -- `POST /api/failover//pin` with a non-admin user → 403; with an admin → 200 (or 404 for an unknown chain, which still proves the admin gate passed). - -Mirror the existing assertions exactly; do not invent a new auth harness. - -- [ ] **Step 6: Swagger and tests** - -Run: `make swagger && go test ./core/http/... 2>&1 | tail -15` -Expected: swagger regenerates without errors; tests PASS. - -- [ ] **Step 7: Commit** - -```bash -git add core/http swagger -git commit -m "feat(failover): expose chain status, pins and events over REST and SSE - -Assisted-by: Claude:claude-opus-5-5" -``` - -(Use the actual swagger output directory if it is not `swagger/`; `git status` shows it.) - ---- - -### Task 10: End-to-end HTTP failover and endpoint audit - -**Files:** -- Modify: `tests/e2e/mock-backend/main.go` (`LoadModel` failure trigger, line ~89) -- Modify: `tests/e2e/cloud_proxy_helpers_test.go` (fake upstream serves `GET /v1/models`) -- Modify: `tests/e2e/e2e_suite_test.go` (model configs) -- Create: `tests/e2e/e2e_failover_test.go` - -**Interfaces:** -- Consumes: the running app from the e2e suite, Task 8 headers, Task 9 endpoints. - -- [ ] **Step 1: Add the mock load-failure trigger** - -In `tests/e2e/mock-backend/main.go` `LoadModel`, before the success return: - -```go - // Lets e2e specs build a failover target whose backend cannot load. - if strings.HasPrefix(in.Model, "fail-load") { - return &pb.Result{Message: "mock: load failure", Success: false}, nil - } -``` - -(add `strings` to the imports if missing) - -- [ ] **Step 2: Make the fake upstream answer `/v1/models`** - -In `newFakeOpenAIUpstream()` (`cloud_proxy_helpers_test.go:40-97`), add a models list and, at the top of the handler: - -```go - if r.Method == http.MethodGet && r.URL.Path == "/v1/models" { - u.mu.Lock() - ids := slices.Clone(u.models) - u.mu.Unlock() - data := make([]map[string]string, 0, len(ids)) - for _, id := range ids { - data = append(data, map[string]string{"id": id}) - } - w.Header().Set("Content-Type", "application/json") - _ = json.NewEncoder(w).Encode(map[string]any{"data": data}) - return - } -``` - -with a `models []string` field and: - -```go -func (u *fakeOpenAIUpstream) SetModels(ids ...string) { - u.mu.Lock() - u.models = ids - u.mu.Unlock() -} -``` - -Use the helper's actual type, mutex and field names. The existing cloud-proxy specs must still pass. - -- [ ] **Step 3: Add configs to the suite** - -In `tests/e2e/e2e_suite_test.go`, where the existing mock-backend configs are written (lines ~105-125), write these additional configs with the same helper and style. For each family `f` in `chat, completion, embeddings, transcription, tts, image, rerank, vad`: - -```yaml -# fail-.yaml -name: fail- -backend: mock-backend -parameters: - model: fail-load- -``` - -```yaml -# chain-.yaml -name: chain- -failover: - targets: - - model: fail- - - model: -``` - -The remote chain needs the fake upstreams' URLs, which exist only at runtime. Write those configs in the `BeforeAll` of the remote spec (Step 4) and load them the way `e2e_cloud_proxy_test.go` registers its cloud-proxy models (read that file first and reuse its mechanism, for example writing YAML to the models dir and calling the loader, or `POST /models/import`). - -- [ ] **Step 4: Write the e2e specs** - -`tests/e2e/e2e_failover_test.go`. Use the suite's base URL variable and HTTP helpers (check `e2e_suite_test.go` for their names; below they are `apiURL` and `http.Post`): - -```go -package e2e_test - -import ( - "bytes" - "encoding/json" - "io" - "mime/multipart" - "net/http" - "time" - - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" -) - -var _ = Describe("Failover chains", Label("failover"), func() { - postJSON := func(path string, body map[string]any) *http.Response { - b, _ := json.Marshal(body) - resp, err := http.Post(apiURL+path, "application/json", bytes.NewReader(b)) - Expect(err).ToNot(HaveOccurred()) - return resp - } - expectServedByMock := func(resp *http.Response) { - defer resp.Body.Close() - body, _ := io.ReadAll(resp.Body) - Expect(resp.StatusCode).To(BeNumerically("<", 300), string(body)) - Expect(resp.Header.Get("X-LocalAI-Served-Model")).ToNot(HavePrefix("fail-")) - Expect(resp.Header.Get("X-LocalAI-Failover")).To(Equal("fallback")) - } - - DescribeTable("retries every endpoint family on the next target", - func(path string, body func(model string) map[string]any) { - expectServedByMock(postJSON(path, body("chain-"+CurrentSpecReport().LeafNodeText))) - }, - Entry("chat", "/v1/chat/completions", func(m string) map[string]any { - return map[string]any{"model": m, "messages": []map[string]string{{"role": "user", "content": "hi"}}} - }), - Entry("completion", "/v1/completions", func(m string) map[string]any { - return map[string]any{"model": m, "prompt": "hi"} - }), - Entry("embeddings", "/v1/embeddings", func(m string) map[string]any { - return map[string]any{"model": m, "input": "hi"} - }), - Entry("tts", "/v1/audio/speech", func(m string) map[string]any { - return map[string]any{"model": m, "input": "hi"} - }), - Entry("image", "/v1/images/generations", func(m string) map[string]any { - return map[string]any{"model": m, "prompt": "a cat", "size": "256x256"} - }), - Entry("rerank", "/v1/rerank", func(m string) map[string]any { - return map[string]any{"model": m, "query": "q", "documents": []string{"a", "b"}} - }), - Entry("vad", "/v1/vad", func(m string) map[string]any { - return map[string]any{"model": m, "audio": []float32{0, 0, 0, 0}} - }), - ) - - It("retries transcription with the multipart body", func() { - var body bytes.Buffer - mw := multipart.NewWriter(&body) - _ = mw.WriteField("model", "chain-transcription") - fw, _ := mw.CreateFormFile("file", "a.wav") - _, _ = fw.Write(testWAV()) - Expect(mw.Close()).To(Succeed()) - resp, err := http.Post(apiURL+"/v1/audio/transcriptions", mw.FormDataContentType(), &body) - Expect(err).ToNot(HaveOccurred()) - expectServedByMock(resp) - }) - - Describe("remote targets", Ordered, func() { - var up1, up2 *fakeOpenAIUpstream - - BeforeAll(func() { - up1, up2 = newFakeOpenAIUpstream(), newFakeOpenAIUpstream() - up1.SetModels("up-1") - up2.SetModels("up-2") - // Register up-1 and up-2 as cloud-proxy passthrough models pointing at - // up1.URL()+"/v1/chat/completions" and up2.URL()+"/v1/chat/completions", - // and chain-remote with probe interval 1s, recovery probes 2 and - // min_dwell 2s, using the same mechanism as e2e_cloud_proxy_test.go. - registerFailoverRemoteModels(up1.URL(), up2.URL()) - }) - - It("fails over when the primary upstream errors and fails back when it recovers", func() { - up1.SetScript(func([]byte) (int, string, string) { return 503, `{"error":"no healthy nodes"}`, "application/json" }) - up2.SetScript(func([]byte) (int, string, string) { - return 200, `{"id":"x","object":"chat.completion","choices":[{"index":0,"message":{"role":"assistant","content":"hi"},"finish_reason":"stop"}]}`, "application/json" - }) - resp := postJSON("/v1/chat/completions", map[string]any{"model": "chain-remote", "messages": []map[string]string{{"role": "user", "content": "hi"}}}) - defer resp.Body.Close() - Expect(resp.StatusCode).To(Equal(200)) - Expect(resp.Header.Get("X-LocalAI-Served-Model")).To(Equal("up-2")) - - up1.SetScript(func([]byte) (int, string, string) { - return 200, `{"id":"x","object":"chat.completion","choices":[{"index":0,"message":{"role":"assistant","content":"hi"},"finish_reason":"stop"}]}`, "application/json" - }) - Eventually(func() string { - r, err := http.Get(apiURL + "/api/failover/chain-remote") - if err != nil { - return "" - } - defer r.Body.Close() - var st struct { - Active string `json:"active"` - } - _ = json.NewDecoder(r.Body).Decode(&st) - return st.Active - }, 30*time.Second, 500*time.Millisecond).Should(Equal("up-1")) - }) - }) -}) -``` - -Add this helper to the same file (200 ms of 16 kHz mono 16-bit silence): - -```go -func testWAV() []byte { - const rate, samples = 16000, 3200 - data := samples * 2 - b := make([]byte, 44+data) - copy(b[0:], "RIFF") - binary.LittleEndian.PutUint32(b[4:], uint32(36+data)) - copy(b[8:], "WAVEfmt ") - binary.LittleEndian.PutUint32(b[16:], 16) - binary.LittleEndian.PutUint16(b[20:], 1) - binary.LittleEndian.PutUint16(b[22:], 1) - binary.LittleEndian.PutUint32(b[24:], rate) - binary.LittleEndian.PutUint32(b[28:], rate*2) - binary.LittleEndian.PutUint16(b[32:], 2) - binary.LittleEndian.PutUint16(b[34:], 16) - copy(b[36:], "data") - binary.LittleEndian.PutUint32(b[40:], uint32(data)) - return b -} -``` - -(add `encoding/binary` to the imports; if the package already defines `testWAV`, use the existing one.) - -Implement `registerFailoverRemoteModels(url1, url2 string)` in the same file using the cloud-proxy registration mechanism found in Step 3. `CurrentSpecReport().LeafNodeText` gives the entry name (`chat`, `completion`, ...), which matches the chain names from Step 3. For VAD and rerank, copy the request shape from existing e2e specs if they differ from the ones above. If the mock backend cannot serve a family at all (the request fails on the mock target too), remove that entry and list the family in the PR description as "covered by unit tests only". - -- [ ] **Step 5: Run the e2e specs** - -Run: `make build-mock-backend && go run github.com/onsi/ginkgo/v2/ginkgo --label-filter=failover -v ./tests/e2e 2>&1 | tail -30` -Expected: PASS. Then run the cloud-proxy specs to check the helper change: `go run github.com/onsi/ginkgo/v2/ginkgo --focus="cloud" -v ./tests/e2e 2>&1 | tail -10`. - -- [ ] **Step 6: Commit** - -```bash -git add tests/e2e -git commit -m "test(failover): cover retry per endpoint family and remote fail-back - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 11: Realtime pipelines - -**Files:** -- Modify: `core/http/endpoints/openai/realtime_model.go` (`RealtimeRoutingContext`, `buildRealtimeRoutingContext` ~860-880, `wrappedModel` ~39-90, `newModel` ~882, stage methods) -- Create: `core/http/endpoints/openai/realtime_failover.go`, `core/http/endpoints/openai/realtime_failover_test.go`, `core/http/endpoints/openai/types/failover.go` -- Modify: `core/http/endpoints/openai/types/server_events.go` (event type constant) -- Modify: `core/http/endpoints/openai/realtime.go` (start events after `newModel` succeeds, ~643-660) -- Test: `tests/e2e/realtime_ws_test.go` (new spec) - -**Interfaces:** -- Consumes: `failover.Manager.Do`, `ChainStatus`, `Subscribe`, `RecordAttemptTrace`. -- Produces: `types.ModelFailoverEvent`, `types.ServerEventTypeModelFailover = "localai.model.failover"`, `(*wrappedModel).stageCall`. - -- [ ] **Step 1: Add the event type** - -In `types/server_events.go`, next to `ServerEventTypeClassifierResult`: - -```go - ServerEventTypeModelFailover ServerEventType = "localai.model.failover" -``` - -`types/failover.go`, following the `ClassifierResultEvent` pattern in `types/classifier.go:426-469` (copy its `ServerEventBase` embedding and `MarshalJSON` shape exactly): - -```go -package types - -import "encoding/json" - -// ModelFailoverEvent tells a client which target serves a pipeline stage -// that names a failover chain. -type ModelFailoverEvent struct { - ServerEventBase - Chain string `json:"chain"` - Stage string `json:"stage"` - From string `json:"from"` - To string `json:"to"` - State string `json:"state"` - Reason string `json:"reason"` -} - -func (ModelFailoverEvent) ServerEventType() ServerEventType { return ServerEventTypeModelFailover } - -func (e ModelFailoverEvent) MarshalJSON() ([]byte, error) { - type alias ModelFailoverEvent - return json.Marshal(struct { - Type ServerEventType `json:"type"` - alias - }{Type: ServerEventTypeModelFailover, alias: alias(e)}) -} -``` - -- [ ] **Step 2: Write the failing unit tests** - -`core/http/endpoints/openai/realtime_failover_test.go` (package and suite of the other tests in that directory; `fakeTransport` is in `realtime_doubles_test.go:29`): - -```go -package openai - -import ( - "context" - "errors" - - "github.com/mudler/LocalAI/core/config" - "github.com/mudler/LocalAI/core/http/endpoints/openai/types" - "github.com/mudler/LocalAI/core/services/failover" - . "github.com/onsi/ginkgo/v2" - . "github.com/onsi/gomega" -) - -type rtSource map[string]config.ModelConfig - -func (s rtSource) GetModelConfig(n string) (config.ModelConfig, bool) { c, ok := s[n]; return c, ok } -func (s rtSource) GetAllModelsConfigs() []config.ModelConfig { - var out []config.ModelConfig - for _, c := range s { - out = append(out, c) - } - return out -} - -var _ = Describe("realtime failover", func() { - var fm *failover.Manager - - BeforeEach(func() { - fm = failover.New(rtSource{ - "a": {Name: "a", Backend: "cloud-proxy"}, - "b": {Name: "b", Backend: "llama-cpp"}, - "chain": {Name: "chain", Failover: &config.FailoverConfig{Targets: []config.FailoverTarget{{Model: "a"}, {Model: "b"}}}}, - }) - }) - - It("routes a chain stage through the plan and retries before commit", func() { - m := &wrappedModel{failover: fm, stageChains: map[string]string{"tts": "chain"}, - stageTargetConfig: func(name string) (*config.ModelConfig, error) { return &config.ModelConfig{Name: name}, nil }} - var tried []string - err := m.stageCall(context.Background(), "tts", nil, func(cfg *config.ModelConfig, _ func()) error { - tried = append(tried, cfg.Name) - if cfg.Name == "a" { - return errors.New("dial tcp: refused") - } - return nil - }) - Expect(err).ToNot(HaveOccurred()) - Expect(tried).To(Equal([]string{"a", "b"})) - }) - - It("calls a plain stage once with its own config", func() { - m := &wrappedModel{} - base := &config.ModelConfig{Name: "plain"} - err := m.stageCall(context.Background(), "tts", base, func(cfg *config.ModelConfig, _ func()) error { - Expect(cfg).To(BeIdenticalTo(base)) - return nil - }) - Expect(err).ToNot(HaveOccurred()) - }) - - It("sends initial events, then switch events, and stops on cancel", func() { - t := &fakeTransport{} - failoverEvents := func() []types.ModelFailoverEvent { - var out []types.ModelFailoverEvent - for _, e := range t.events() { - if fe, ok := e.(types.ModelFailoverEvent); ok { - out = append(out, fe) - } - } - return out - } - stop := startFailoverEvents(t, fm, map[string]string{"llm": "chain"}) - Eventually(failoverEvents).Should(ContainElement(And( - HaveField("Reason", "initial"), HaveField("To", "a"), HaveField("Stage", "llm")))) - fm.ReportFailure("a", errors.New("dial tcp: refused")) - Eventually(failoverEvents).Should(ContainElement(And( - HaveField("Reason", "trip"), HaveField("To", "b")))) - stop() - }) -}) -``` - -`t.events()` must return the events sent so far as `[]types.ServerEvent`. Read `realtime_doubles_test.go`: if `fakeTransport` has no such accessor, add a mutex-guarded `events()` method there. - -- [ ] **Step 3: Run to verify failure** - -Run: `go test ./core/http/endpoints/openai/... 2>&1 | tail -5` -Expected: compile failure (`unknown field failover`). - -- [ ] **Step 4: Implement stage resolution** - -In `realtime_model.go`: - -1. `RealtimeRoutingContext`: add `Failover *failover.Manager`. In `buildRealtimeRoutingContext`, set `Failover: a.FailoverManager()`. -2. `wrappedModel`: add - -```go - // failover and stageChains route pipeline stages that name a failover - // chain; stageChains maps a stage ("llm", "tts", ...) to its chain. - failover *failover.Manager - stageChains map[string]string - stageTargetConfig func(name string) (*config.ModelConfig, error) - appTracing bool -``` - -3. In `newModel`, after each stage config is loaded with `LoadResolvedModelConfig` and before `Validate()`: when the loaded config `IsFailover()`, record the chain and swap in its active target's config, so everything that inspects stage configs at session start (voice, reasoning, templates) sees a real model: - -```go - resolveStage := func(stage string, cfg *config.ModelConfig) (*config.ModelConfig, error) { - if cfg == nil || !cfg.IsFailover() { - return cfg, nil - } - if routing == nil || routing.Failover == nil { - return nil, fmt.Errorf("pipeline %s stage %q is a failover chain, but failover is not running", stage, cfg.Name) - } - st, ok := routing.Failover.ChainStatus(cfg.Name) - if !ok { - return nil, fmt.Errorf("failover chain %q not found", cfg.Name) - } - stageChains[stage] = cfg.Name - return cl.LoadResolvedModelConfig(st.Active, ml.ModelPath, appConfig.ToConfigLoaderOptions()...) - } -``` - -Call it for `vad`, `transcription`, `llm`, `tts` and `sound_detection` right after each load. Declare `stageChains := map[string]string{}` before the loads. When building `&wrappedModel{...}`, set: - -```go - stageChains: stageChains, - stageTargetConfig: func(name string) (*config.ModelConfig, error) { - return cl.LoadResolvedModelConfig(name, ml.ModelPath, appConfig.ToConfigLoaderOptions()...) - }, - appTracing: appConfig.EnableTracing, -``` - -and, when `routing != nil`, `failover: routing.Failover`. - -4. `realtime_failover.go`: - -```go -package openai - -import ( - "context" - "sort" - - "github.com/mudler/LocalAI/core/config" - "github.com/mudler/LocalAI/core/http/endpoints/openai/types" - "github.com/mudler/LocalAI/core/services/failover" -) - -// stageCall runs fn with the config that should serve stage now. A plain -// stage uses base. A chain stage goes through the failover plan and retries -// on the next target until fn calls commit. -func (m *wrappedModel) stageCall(ctx context.Context, stage string, base *config.ModelConfig, fn func(cfg *config.ModelConfig, commit func()) error) error { - chain, ok := m.stageChains[stage] - if !ok || m.failover == nil { - return fn(base, func() {}) - } - return m.failover.Do(ctx, chain, func(ctx context.Context, target string, commit func()) error { - cfg, err := m.stageTargetConfig(target) - if err != nil { - return err - } - err = fn(cfg, commit) - if err != nil { - failover.RecordAttemptTrace(m.appTracing, chain, target, err) - } - return err - }) -} - -// startFailoverEvents tells the client which target serves each chain stage -// now, and again whenever a chain switches. The returned func stops it. -func startFailoverEvents(t Transport, fm *failover.Manager, stageChains map[string]string) func() { - events, cancel := fm.Subscribe(16) - stages := make([]string, 0, len(stageChains)) - for s := range stageChains { - stages = append(stages, s) - } - sort.Strings(stages) - for _, stage := range stages { - chain := stageChains[stage] - if st, ok := fm.ChainStatus(chain); ok { - sendEvent(t, types.ModelFailoverEvent{Chain: chain, Stage: stage, To: st.Active, State: string(st.State), Reason: string(failover.ReasonInitial)}) - } - } - go func() { - for ev := range events { - if ev.Type != failover.EventChainSwitched { - continue - } - for _, stage := range stages { - if stageChains[stage] == ev.Chain { - sendEvent(t, types.ModelFailoverEvent{Chain: ev.Chain, Stage: stage, From: ev.From, To: ev.To, State: ev.State, Reason: string(ev.Reason)}) - } - } - } - }() - return cancel -} -``` - -5. Route each stage method through `stageCall`. Non-streaming example (`TTS`): - -```go -func (m *wrappedModel) TTS(ctx context.Context, text, voice, language string) (string, *proto.Result, error) { - var ( - out string - res *proto.Result - ) - err := m.stageCall(ctx, "tts", m.TTSConfig, func(cfg *config.ModelConfig, _ func()) error { - var err error - out, res, err = backend.ModelTTS(ctx, text, voice, language, "", maps.Clone(m.ttsParams), m.modelLoader, m.appConfig, *cfg) - return err - }) - return out, res, err -} -``` - -Apply the same shape to `VAD` (stage `vad`, `m.VADConfig`), `Transcribe` (`transcription`, `m.TranscriptionConfig`) and `SoundDetection` (`sound_detection`, `m.SoundDetectionConfig`). - -Streaming stages wrap the callback so the first delivered chunk commits: -- `TTSStream`: `onAudio` becomes `func(pcm []byte, sr int) error { commit(); return onAudio(pcm, sr) }`. -- `TranscribeStream`: `onDelta` becomes `func(s string) { commit(); onDelta(s) }`. -- `TranscribeLive`: call `commit()` right after the session opens successfully; the stage call only covers opening. - -`Predict` returns a closure that runs inference later. Keep the early part (message and tool preparation that does not depend on the config) where it is, and move everything that reads `turnCfg` (templating, `routeTurn` result use, `backend.ModelInference`) into the returned closure: - -```go - return func() (backend.LLMResponse, error) { - var resp backend.LLMResponse - err := m.stageCall(ctx, "llm", turnCfg, func(cfg *config.ModelConfig, commit func()) error { - cb := func(s string, u backend.TokenUsage) bool { - commit() - return tokenCallback(s, u) - } - // build predInput from cfg (not turnCfg) here, then: - predict, err := backend.ModelInference(ctx, predInput, messages, images, videos, audios, m.modelLoader, *cfg, m.confLoader, m.appConfig, cb, toolsJSON, toolChoiceJSON, logprobs, topLogprobs, logitBias, nil) - if err != nil { - return err - } - resp, err = predict() - return err - }) - return resp, err - }, nil -``` - -Handle a nil `tokenCallback` (call `commit()` only). Apply `applyPipelineReasoning`/`applyPipelineThinking` to `cfg` inside the closure the same way `newModel` applies them to `cfgLLM`. A chain stage used together with a router (`routeTurn`) is out of scope: when `routeTurn` swapped the config, call the backend directly as today. - -6. In `realtime.go`, directly after the successful `newModel` for the main session (~643-660), start the events for the life of the session handler: - -```go - if wrapped, ok := m.(*wrappedModel); ok && wrapped.failover != nil && len(wrapped.stageChains) > 0 { - stopFailoverEvents := startFailoverEvents(t, wrapped.failover, wrapped.stageChains) - defer stopFailoverEvents() - } -``` - -Check that the enclosing function runs for the whole session (the defer must fire when the session ends, not when setup returns). If it returns earlier, store the stop func on the session and call it where the session is torn down. - -- [ ] **Step 5: Run unit tests** - -Run: `go test -race ./core/http/endpoints/openai/... 2>&1 | tail -15` -Expected: PASS, including the existing realtime specs. - -- [ ] **Step 6: Add the e2e realtime spec** - -In `tests/e2e/realtime_ws_test.go`, add a spec (Label `failover`) next to the existing WebSocket session spec, reusing its connection and turn helpers: -- Config `rt-failover` whose pipeline uses the same VAD/transcription/TTS models as the existing spec and `llm: chain-rt`, where `chain-rt` targets `[fail-rt, ]` and `fail-rt` has `parameters.model: fail-load-rt` (write these in the suite as in Task 10 Step 3). -- Assert: after connect, a `localai.model.failover` event with `stage: llm`, `reason: initial`, `to: fail-rt`. -- Send one user turn: a `localai.model.failover` event with `to` = the mock LLM and `reason: trip` arrives, and the turn completes with `response.done`. -- Send a second turn: it completes, and the conversation still holds the first turn's items (the item ids from the first turn's `conversation.item.created` events are still retrievable, or the second request's messages include them — use whatever the existing spec can observe). - -Run: `go run github.com/onsi/ginkgo/v2/ginkgo --label-filter=failover -v ./tests/e2e 2>&1 | tail -30` -Expected: PASS. - -- [ ] **Step 7: Commit** - -```bash -git add core/http/endpoints/openai tests/e2e -git commit -m "feat(failover): switch realtime pipeline stages per call - -A stage that names a chain is resolved on every call, so a switch keeps -the session and its conversation. Clients get localai.model.failover -events at session start and on every switch. - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 12: MCP admin tools - -**Files:** -- Modify: `pkg/mcp/localaitools/tools.go` (constants, `mutatingToolNames`), `client.go` (interface), `dto.go` (DTOs), `inproc/client.go`, `httpapi/client.go`, `httpapi/routes.go`, `server.go` (register), `prompts/20_tools.md`, `prompts/10_safety.md` -- Create: `pkg/mcp/localaitools/tools_failover.go` -- Modify tests: `coverage_test.go`, `server_test.go`, `fakes_test.go`, `inproc/client_test.go`, `httpapi/client_test.go`, `parity_test.go` -- Modify: where the in-process client is constructed (find with `grep -rn "inproc.Client{\|inproc.New" core pkg`) to pass the failover manager - -**Interfaces:** -- Consumes: `failover.Manager.Status/Pin/Unpin`. -- Produces: tools `list_failover_chains` (read-only), `pin_failover_target`, `unpin_failover_target` (mutating). - -- [ ] **Step 1: Write the failing tests** - -- `coverage_test.go` `toolToHTTPRoute`: - -```go - ToolListFailoverChains: "GET /api/failover", - ToolPinFailoverTarget: "POST /api/failover/:chain/pin", - ToolUnpinFailoverTarget: "DELETE /api/failover/:chain/pin", -``` - -- `server_test.go`: add `ToolListFailoverChains` to `expectedReadOnlyCatalog`; add dispatch rows in the table at ~154: - -```go - {ToolListFailoverChains, map[string]any{}, "ListFailoverChains"}, - {ToolPinFailoverTarget, map[string]any{"chain": "c", "target": "b"}, "PinFailoverTarget"}, - {ToolUnpinFailoverTarget, map[string]any{"chain": "c"}, "UnpinFailoverTarget"}, -``` - -- `fakes_test.go`: fake methods that record the call name like the existing alias fakes (lines ~157-170). -- `inproc/client_test.go` and `httpapi/client_test.go`: one spec each for list and pin, in the style of the alias specs. The httpapi spec asserts the request method and path; the inproc spec builds a `failover.Manager` over a small in-memory source with one chain. - -Run: `go test ./pkg/mcp/localaitools/... 2>&1 | tail -5` -Expected: compile failure. - -- [ ] **Step 2: Implement** - -`tools.go`: add `ToolListFailoverChains = "list_failover_chains"` to the read-only block, `ToolPinFailoverTarget = "pin_failover_target"` and `ToolUnpinFailoverTarget = "unpin_failover_target"` to the mutating block, and both mutating names to `mutatingToolNames`. - -`dto.go`: - -```go -type FailoverTargetInfo struct { - Model string `json:"model"` - Kind string `json:"kind"` - Warm bool `json:"warm"` - State string `json:"state"` - LastError string `json:"last_error,omitempty"` -} - -type FailoverChainInfo struct { - Name string `json:"name"` - State string `json:"state"` - Active string `json:"active"` - Pinned string `json:"pinned,omitempty"` - Targets []FailoverTargetInfo `json:"targets"` -} -``` - -`client.go` interface, next to the alias methods: - -```go - ListFailoverChains(ctx context.Context) ([]FailoverChainInfo, error) - PinFailoverTarget(ctx context.Context, chain, target string) error - UnpinFailoverTarget(ctx context.Context, chain string) error -``` - -`inproc/client.go`: add field `Failover *failover.Manager` to the client struct, set it where the client is constructed (from `application.FailoverManager()`), and: - -```go -func (c *Client) ListFailoverChains(_ context.Context) ([]localaitools.FailoverChainInfo, error) { - out := []localaitools.FailoverChainInfo{} - if c.Failover == nil { - return out, nil - } - for _, ch := range c.Failover.Status() { - info := localaitools.FailoverChainInfo{Name: ch.Name, State: string(ch.State), Active: ch.Active} - if ch.Pinned != nil { - info.Pinned = *ch.Pinned - } - for _, t := range ch.Targets { - info.Targets = append(info.Targets, localaitools.FailoverTargetInfo{ - Model: t.Model, Kind: string(t.Kind), Warm: t.Warm, State: string(t.State), LastError: t.LastError, - }) - } - out = append(out, info) - } - return out, nil -} - -func (c *Client) PinFailoverTarget(_ context.Context, chain, target string) error { - if c.Failover == nil { - return errors.New("failover is not running") - } - return c.Failover.Pin(chain, target) -} - -func (c *Client) UnpinFailoverTarget(_ context.Context, chain string) error { - if c.Failover == nil { - return errors.New("failover is not running") - } - return c.Failover.Unpin(chain) -} -``` - -`httpapi/routes.go`: `routeFailover = "/api/failover"`. `httpapi/client.go`: - -```go -func (c *Client) ListFailoverChains(ctx context.Context) ([]localaitools.FailoverChainInfo, error) { - var out struct { - Chains []localaitools.FailoverChainInfo `json:"chains"` - } - if err := c.do(ctx, http.MethodGet, routeFailover, nil, &out); err != nil { - return nil, err - } - return out.Chains, nil -} - -func (c *Client) PinFailoverTarget(ctx context.Context, chain, target string) error { - return c.do(ctx, http.MethodPost, routeFailover+"/"+url.PathEscape(chain)+"/pin", map[string]string{"target": target}, nil) -} - -func (c *Client) UnpinFailoverTarget(ctx context.Context, chain string) error { - return c.do(ctx, http.MethodDelete, routeFailover+"/"+url.PathEscape(chain)+"/pin", nil, nil) -} -``` - -The REST list returns `pinned` as a JSON string or null and `failover.TargetStatus` fields; `FailoverChainInfo` decodes the fields it shares. Check that `c.do` accepts a body map and a nil out; match its real signature. - -`tools_failover.go`, following `tools_aliases.go`: - -```go -package localaitools - -import ( - "context" - - "github.com/modelcontextprotocol/go-sdk/mcp" -) - -func registerFailoverTools(s *mcp.Server, client LocalAIClient, opts Options) { - mcp.AddTool(s, &mcp.Tool{ - Name: ToolListFailoverChains, - Description: "List model failover chains, the target serving each one now, and the health of every target.", - }, func(ctx context.Context, _ *mcp.CallToolRequest, _ struct{}) (*mcp.CallToolResult, any, error) { - chains, err := client.ListFailoverChains(ctx) - if err != nil { - return errorResult(err), nil, nil - } - return jsonResult(chains) - }) - - if opts.DisableMutating { - return - } - - mcp.AddTool(s, &mcp.Tool{ - Name: ToolPinFailoverTarget, - Description: "Force a failover chain to serve every request from one target, regardless of health, until it is unpinned.", - }, func(ctx context.Context, _ *mcp.CallToolRequest, args struct { - Chain string `json:"chain" jsonschema:"failover chain name"` - Target string `json:"target" jsonschema:"target model to pin"` - }) (*mcp.CallToolResult, any, error) { - if err := client.PinFailoverTarget(ctx, args.Chain, args.Target); err != nil { - return errorResult(err), nil, nil - } - return jsonResult(map[string]string{"chain": args.Chain, "pinned": args.Target}) - }) - - mcp.AddTool(s, &mcp.Tool{ - Name: ToolUnpinFailoverTarget, - Description: "Remove the pin from a failover chain so health decides the target again.", - }, func(ctx context.Context, _ *mcp.CallToolRequest, args struct { - Chain string `json:"chain" jsonschema:"failover chain name"` - }) (*mcp.CallToolResult, any, error) { - if err := client.UnpinFailoverTarget(ctx, args.Chain); err != nil { - return errorResult(err), nil, nil - } - return jsonResult(map[string]string{"chain": args.Chain, "pinned": ""}) - }) -} -``` - -Match the import path of the MCP SDK and the exact signatures of `errorResult`/`jsonResult` used in `tools_aliases.go`. - -`server.go`: add `registerFailoverTools(s, client, opts)` to the register list (~45-56). - -`prompts/20_tools.md`: under `## Read-only` add ``- `list_failover_chains` — List failover chains, their active target and target health.``; under `## Mutating` add ``- `pin_failover_target` — Force a failover chain to one target.`` and ``- `unpin_failover_target` — Remove a failover pin.``. `prompts/10_safety.md` line 5: add both mutating names to the backticked list. - -- [ ] **Step 3: Run tests** - -Run: `go test ./pkg/mcp/localaitools/... 2>&1 | tail -10` -Expected: PASS (including `TestToolHTTPRouteMappingComplete`, the prompts test and parity). - -- [ ] **Step 4: Commit** - -```bash -git add pkg/mcp core -git commit -m "feat(failover): add MCP tools to list chains and pin targets - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -### Task 13: Documentation and final verification - -**Files:** -- Create: `docs/content/features/model-failover.md` -- Modify: `docs/content/features/model-aliases.md`, `docs/content/features/openai-realtime.md`, and the cloud-proxy section (find it with `grep -rln "cloud-proxy" docs/content`; add the link on the page that documents the `proxy:` block) - -- [ ] **Step 1: Write the docs page** - -`docs/content/features/model-failover.md`: - -````markdown -+++ -disableToc = false -title = "Model Failover" -weight = 15 -url = "/features/model-failover/" -+++ - -A **failover chain** is a model name that is served by an ordered list of -other models. LocalAI sends each request to the first healthy target. When a -target fails, the request moves to the next target, and later requests stay -there until the first target has recovered. - -Use it to serve a model from a remote LocalAI or another OpenAI-compatible -provider, and to fall back to a local model when the remote one is down. - -## Declaring a chain - -```yaml -name: assistant-llm -failover: - targets: - - model: argus-llm # for example a cloud-proxy model - - model: gemma-local - warm: true # keep it loaded -``` - -Clients call `assistant-llm`. Each target is a normal model config. A chain -has no `backend` and no `parameters.model`. - -Optional settings, with their defaults: - -```yaml -failover: - probe: - interval: 15s # how often an idle target is checked - timeout: 5s - trip: - errors: 1 # failures within the window that mark a target down - window: 30s - recovery: - probes: 3 # test requests a target must pass before it is used again - min_dwell: 60s # minimum time on a lower target before moving back -``` - -Rules: - -- A chain needs at least 2 targets. A target can be an alias, but not another - chain. -- A chain cannot also set `alias` or `backend`. -- Responses name the chain as the model. The `X-LocalAI-Served-Model` header - names the target that served the request. - -## How the target is chosen - -- The active target is the first healthy target in the list. -- When a target fails, LocalAI marks it down and moves to the next target at - once. -- LocalAI moves back to a higher target only when that target has passed - `recovery.probes` test requests **and** the current target has been active - for at least `recovery.min_dwell`. This stops an unstable upstream from - moving traffic back and forth. -- When all targets are down, the chain is `degraded`. Each request still tries - every target in order. - -## Retry inside a request - -When a target fails before the response starts, LocalAI sends the same -request to the next target. The client does not see the failure. - -- LocalAI does not retry after the first byte of a response is sent (for - example after the first streamed token). The request fails, the target is - marked down, and the next request uses the next target. -- LocalAI does not retry client errors (4xx), such as a prompt that is too - long, because the next target would reject it too. -- Request bodies larger than 32 MiB are not retried. - -When the primary did not serve the request, the response has the header -`X-LocalAI-Failover: fallback`, or `X-LocalAI-Failover: degraded` when all -targets were down. - -## Health checks - -| Target | Regular check | Check before moving back | -|---|---|---| -| Remote (`cloud-proxy`) | `GET /v1/models` on the upstream lists the model | one small real request, for example a 1-token completion | -| Local, `warm: true` | the backend answers a health check | one small real request | -| Local, not warm | the model file exists | none: the target is used again after `min_dwell` | - -A request that succeeds counts as a check, so a busy target is almost never -probed. A model that is not warm is never loaded only to check it. - -When a target is in more than one chain, its check settings come from the -first of those chains in name order. - -## Warm targets - -`warm: true` loads a local target at startup and protects it from idle and -LRU eviction, so a switch does not wait for the model to load. Warm targets -count toward the active backend limit. If warm targets fill that limit, other -models cannot load, and the error names the warm targets. - -## Realtime pipelines - -A pipeline stage can name a chain: - -```yaml -name: assistant -pipeline: - vad: silero-vad - transcription: whisper-chain - llm: assistant-llm - tts: voice-chain -``` - -LocalAI resolves the chain for every call of the stage. When a chain switches, -the session stays open and keeps its conversation. The next turn uses the new -target. - -The session receives a `localai.model.failover` event for each chain stage when -it starts (`reason: initial`) and each time a chain switches: - -```json -{"type":"localai.model.failover","chain":"assistant-llm","stage":"llm", - "from":"argus-llm","to":"gemma-local","state":"fallback","reason":"trip"} -``` - -A chain used as a candidate of a router in a realtime pipeline is not -resolved per call. - -## Watching failover - -- `GET /api/failover` lists every chain, its active target and the state of - each target. -- `GET /api/failover/{chain}` returns one chain. -- `GET /api/failover/events` is a server-sent event stream. The first event is - `snapshot` with the full state. Then `chain.switched` and `target.state` - events follow. -- Metrics: `localai_failover_switches_total{chain,from,to,reason}` and - `localai_failover_target_up{target}`. -- With tracing on, each skipped target appears in the Traces view with the - error that made LocalAI skip it. - -## Pinning a target - -An admin can force a chain to one target, for example during maintenance: - -```bash -curl -X POST http://localhost:8080/api/failover/assistant-llm/pin \ - -H 'Content-Type: application/json' -d '{"target":"gemma-local"}' -curl -X DELETE http://localhost:8080/api/failover/assistant-llm/pin -``` - -While a chain is pinned, only the pinned target serves it. Health checks -continue. A restart removes the pin. - -## Assistant and MCP - -The LocalAI Assistant and `local-ai mcp-server` offer `list_failover_chains`, -`pin_failover_target` and `unpin_failover_target`. Create and edit chains with -the model config tools, like any other model. - -## Limits - -- Failover state is kept in memory by each LocalAI instance. Several frontends - in distributed mode each keep their own view. -- Chains do not nest. -- See also [model aliases]({{%relref "features/model-aliases" %}}) and the - [realtime API]({{%relref "features/openai-realtime" %}}). -```` - -- [ ] **Step 2: Cross-links** - -- `model-aliases.md`: add at the end of `## Rules and behavior`: `To serve a name from several models with automatic fallback, use a [failover chain]({{%relref "features/model-failover" %}}).` -- `openai-realtime.md`: in the pipeline section add: `A pipeline stage can name a [failover chain]({{%relref "features/model-failover" %}}); the stage then switches targets without closing the session.` -- The page that documents `proxy:` / `cloud-proxy`: add `To fall back to a local model when the upstream is down, list the proxy model in a [failover chain]({{%relref "features/model-failover" %}}).` - -- [ ] **Step 3: Final verification** - -Run, in order, and read each output: - -```bash -make protogen-go build-mock-backend -go vet ./core/... ./pkg/mcp/... -go test -race ./core/services/failover/... ./core/config/... ./core/http/... ./pkg/mcp/localaitools/... -go run github.com/onsi/ginkgo/v2/ginkgo --label-filter=failover -v ./tests/e2e -make swagger && git diff --stat -- swagger -make test-coverage-check -``` - -Expected: vet clean, all tests PASS, swagger has no uncommitted changes, coverage at or above the baseline. If coverage dropped, add tests; never edit `coverage-baseline.txt`. - -- [ ] **Step 4: Commit** - -```bash -git add docs/content -git commit -m "docs: document model failover chains - -Assisted-by: Claude:claude-opus-5-5" -``` - ---- - -## PR notes (for the finishing step) - -The PR description must include: -- The MCP decision: tools added for list, pin and unpin; chain create/edit reuses the model config tools. -- Endpoint families without in-request retry, if Task 10 found any. -- The limits from the docs page (per-instance state, router candidates in realtime). -- A note that the human submitter adds `Signed-off-by` (DCO). diff --git a/docs/superpowers/specs/2026-08-21-configurable-copy-buffer-design.md b/docs/superpowers/specs/2026-08-21-configurable-copy-buffer-design.md deleted file mode 100644 index 1085af7fe..000000000 --- a/docs/superpowers/specs/2026-08-21-configurable-copy-buffer-design.md +++ /dev/null @@ -1,73 +0,0 @@ -# Configurable copy buffer design - -**Date:** 21 August 2026 -**Status:** Approved - -## Problem - -`pkg/xio.Copy` wraps a source reader so a context can stop a copy between -reads. It delegates to `io.Copy`, which uses a 32 KiB buffer for the wrapped -reader and writer types used by model downloads. - -Small writes limit model import throughput when the models directory uses an -SMB volume. The development deployment reads large files from the volume at -about 104 MiB/s. A model import writes to the same volume at less than 1 MiB/s. - -## Design - -Keep `xio.Copy` as the context-aware copy entry point. Add variadic functional -options so existing callers continue to compile without changes. - -Add an exported `Option` type and a `WithBufferSize(size int) Option` function. -`Copy` uses a 1 MiB buffer by default. A caller can override the buffer size -with `WithBufferSize`. - -If a caller supplies a non-positive buffer size, `Copy` uses the 1 MiB default. -This rule prevents invalid configuration from causing an `io.CopyBuffer` -panic. - -`Copy` allocates one buffer for each active call. It passes that buffer to -`io.CopyBuffer`. The context-aware reader continues to check cancellation -before each source read. - -The first change does not use `sync.Pool`. A pool adds shared state and retains -large caller-selected buffers. Measurements do not justify that complexity. - -## Compatibility - -The existing signature gains only a variadic argument: - -```go -func Copy(ctx context.Context, dst io.Writer, src io.Reader, options ...Option) (int64, error) -``` - -All existing calls remain source compatible. Copy results and cancellation -errors do not change. - -The default buffer increases temporary memory use by approximately 992 KiB for -each concurrent copy compared with the current 32 KiB buffer. - -## Tests and measurement - -Add a Ginkgo suite for `pkg/xio`. Tests cover these behaviors: - -- `Copy` copies the complete source. -- The default buffer permits reads larger than 32 KiB. -- `WithBufferSize` changes the maximum requested read size. -- A non-positive override uses the default buffer. -- A canceled context stops the copy and returns the context error. - -Add a benchmark that runs `Copy` with the default buffer and representative -overrides. The benchmark records throughput and allocations. It does not make -timing assertions. - -Run the focused `pkg/xio` suite first. Then run the packages that call -`xio.Copy`: `pkg/downloader` and `pkg/oci`. - -## Deployment validation - -The code change alone does not alter the running development deployment. After -CI publishes a development image and Flux deploys it, import a large model to -the NAS-backed models directory. Compare the progress rate with the previous -0.7-0.8 MiB/s result. - diff --git a/docs/superpowers/specs/2026-08-21-distributed-model-config-revisions-design.md b/docs/superpowers/specs/2026-08-21-distributed-model-config-revisions-design.md deleted file mode 100644 index 6d1de2e63..000000000 --- a/docs/superpowers/specs/2026-08-21-distributed-model-config-revisions-design.md +++ /dev/null @@ -1,357 +0,0 @@ -# Distributed Model Configuration Revisions - -## Problem - -Editing a model configuration in a distributed LocalAI deployment can leave the -cluster serving different effective configurations for the same logical model. -The frontend reloads the edited YAML and asks workers to stop the model, but the -existing `backend.stop` message is fire-and-forget. The frontend therefore -removes routing state without knowing whether the worker process stopped. - -Separately, the replica reconciler persists `ModelLoadInfo` independently of -live `NodeModel` rows. This is necessary for restoring `min_replicas` after a -worker failure, but the persisted options currently have no relationship to a -specific revision of the model configuration. After an edit, the reconciler can -restore a replica from options captured before the edit. - -The observed result was one replica serving a context near 100K while another -served the default 8K context and default parallelism. Requests behaved -differently depending on which replica the router selected. The problem is not -specific to `context_size`: any load-time model option can be stale. - -## Goals - -- Make all routable replicas of a logical model belong to the current model - configuration revision. -- Prevent the reconciler and late load jobs from restoring options belonging to - an older revision. -- Remove a model from routing before attempting distributed cleanup. -- Confirm that the exact worker process exited before deleting its registry - row. -- Recover safely when a worker or NATS is temporarily unreachable. -- Apply the same lifecycle to raw YAML edits, structured configuration patches, - renames, disabling, and changes received from peer frontends. -- Preserve the existing ability to restore `min_replicas` after ordinary - worker or backend failure when the model configuration has not changed. -- Expose enough state to diagnose why two replicas have different effective - options. - -## Non-goals - -- Requiring identical hardware-derived options on heterogeneous workers. -- Changing `model.unload`, which remains a memory-release operation. -- Replacing backend administration operations such as backend upgrade, delete, - or stop-all. -- Automatically upgrading workers that do not support the new stop protocol. -- Making arbitrary out-of-band filesystem edits transactional across multiple - machines. Such edits are detected when the model configuration loader next - refreshes the model. - -## Configuration identity - -Each validated model configuration has a `config_revision`. The revision is a -SHA-256 digest of a canonical semantic representation of the validated model -configuration. Formatting, YAML comments, and map ordering do not affect the -revision. Load-time request overrides and node-specific hardware tuning are not -part of this digest. - -Canonicalization must use the typed, validated configuration rather than raw -YAML bytes. The canonical representation includes every field that can affect -model loading or serving. Fields used only to locate the source file or report -runtime status are excluded. The canonical encoder must produce stable field -and map ordering and must distinguish absent values where absence has different -semantics from an explicit zero value. - -The revision is carried with the model options from configuration loading into -the distributed router. It is also persisted in: - -- `ModelConfigState`, keyed by logical model name, as the currently accepted - revision; -- `ModelLoadInfo`, alongside the serialized `pb.ModelOptions` used for future - reconciliation; -- `NodeModel`, identifying the revision used for that live replica. - -Each `NodeModel` also records an `effective_options_hash`, computed from the -fully materialized `pb.ModelOptions` after node-specific hardware defaults and -file-path staging rewrites. This hash is diagnostic only. Two replicas may have -different effective hashes and remain compatible when they share the same -configuration revision. - -Rows created by older versions have an empty revision. They remain usable until -the model's first revision-aware configuration mutation. Once a current -revision is recorded, empty-revision rows are stale and cannot be routed. - -## Registry invariants - -The database is the coordination boundary shared by frontend replicas. - -1. At most one current configuration revision exists per logical model name. -2. A `NodeModel` is routable only when it is in the loaded state and its - `config_revision` equals the current `ModelConfigState` revision. -3. A `ModelLoadInfo` row is reconcilable only when its revision equals the - current `ModelConfigState` revision. -4. A load job may publish `NodeModel` or `ModelLoadInfo` state only when its - captured revision still equals the current revision. -5. Advancing the current revision and quarantining prior-revision replica rows - happen in one database transaction. - -The load-info upsert becomes compare-and-set rather than unconditional -last-write-wins. If the load's revision is no longer current, the upsert returns -a typed stale-revision error. The load is then abandoned and its worker process -is stopped through the exact stop protocol. A late load can therefore neither -be routed nor overwrite current reconciliation options. - -Normal worker death does not change `ModelConfigState` or delete matching -`ModelLoadInfo`; this preserves restart recovery. A configuration mutation -advances `ModelConfigState` and invalidates older load information. - -## Configuration mutation lifecycle - -All model configuration mutation entry points use one model administration -lifecycle service. The structured PATCH endpoint must no longer bypass local -shutdown behavior. - -For an edit that keeps the same logical model name, the service: - -1. Validates and persists the new configuration. -2. Reloads it and computes its semantic revision. -3. In one transaction, records the new current revision, marks every replica - from another or empty revision as `unloading`, and removes or supersedes old - `ModelLoadInfo`. -4. Broadcasts the revision-aware invalidation to peer frontends. -5. Starts cleanup for each quarantined replica using exact `model.stop`. -6. Deletes a replica row only after confirmed process termination or confirmed - absence of that exact process. - -Marking rows `unloading` precedes network calls. A worker that cannot be reached -therefore cannot continue receiving inference traffic through LocalAI even if -its old backend process is still alive. - -The configuration save is durable even if cleanup is incomplete. The endpoint -must not report that saving failed after the new file and revision have -committed. Its response reports that cleanup is pending, and the condition is -also logged and exposed through the existing model/node lifecycle status -surfaces. Subsequent retries finish cleanup. - -For rename, the old identity is quarantined and stopped under its old name. The -new identity receives its own current revision. Old load information is not -copied to the new name. Disable performs the same quarantine and cleanup but -does not permit fresh loads while disabled. Delete follows the existing file -deletion lifecycle after exact process cleanup. - -Peer invalidation events carry the logical model name, operation, and new -revision. Applying an event is idempotent. A peer that already observes that -revision refreshes its in-memory configuration but does not create a second -cleanup generation. - -When the existing configuration watcher detects an out-of-band file change, it -computes the revision after validation and submits the same lifecycle -transition. A parse or validation failure leaves the last accepted revision -current and does not quarantine its replicas. This does not make filesystem -writes atomic, but it ensures a successfully observed external edit cannot -silently bypass revision-aware routing. - -## Exact worker process stop - -A new request/reply NATS operation, `model.stop`, is separate from the existing -ambiguous `backend.stop` operation. - -The request contains: - -```text -model_name -process_key -expected_address -force -config_revision -``` - -`process_key` is the exact supervisor key, including replica index. The -controller derives it from the registry row rather than asking the worker to -resolve a bare backend or model name. `expected_address` prevents a stale row -from stopping an unrelated process after port reuse. `config_revision` is -included for auditability; process key and expected address are the worker-side -identity checks because workers do not own the configuration database. - -The reply contains: - -```text -matched -freed -terminated -process_key -address -error -``` - -The worker verifies that both process key and address identify the same -supervised process. A mismatched address is an error and never stops anything. -An absent process is a successful idempotent outcome with `matched=false` and -`terminated=true` because there is no process left to clean up. - -For a graceful request, the worker performs bounded gRPC `Free()` and then -terminates the supervised process. A `Free()` failure is recorded but does not -prevent termination. A forced request skips `Free()`. The worker replies only -after the process has exited and its supervisor bookkeeping and port ownership -have been updated. - -The existing operations retain their meanings: - -- `model.unload` calls gRPC `Free()` without promising process termination; -- `backend.stop` remains an administration and compatibility operation whose - identifier may be a backend name; -- `model.stop` is the only operation used to confirm configuration-generation - cleanup for an exact replica. - -Sending both `model.unload` and `model.stop` is unnecessary because graceful -`model.stop` already performs bounded `Free()` before termination. - -## Unreachable workers and retry - -An `unloading` replica is never routable. Failed `model.stop` attempts retain -the row with its last error, attempt count, and next retry time. A bounded, -backoff-based cleanup loop retries exact stops. Retries are idempotent and are -claimed through the database so multiple frontend replicas do not concurrently -own the same attempt. - -The existing recovery paths remain backstops: - -- Worker re-registration clears all `NodeModel` rows for that node because a - restarted worker has no surviving supervised backend processes. -- The per-model health monitor removes rows after consecutive unreachable - backend probes. -- Node offline handling prevents scheduling onto a worker with stale - heartbeats. - -Cleanup-row removal through any of these paths fires the existing replica -removal hooks. It does not restore stale `ModelLoadInfo` because only the -current revision is eligible for reconciliation. - -If a worker keeps heartbeating but does not support `model.stop`, the row stays -quarantined and the error clearly identifies an incompatible worker version. -The system favors temporary unavailability over silently serving an obsolete -configuration. Restarting or upgrading that worker lets re-registration or a -subsequent retry complete cleanup. - -## Reconciliation and loading - -The reconciler reads the current revision and matching `ModelLoadInfo` in one -consistent operation. If no matching load information exists, it does not use -an older blob. It records a diagnostic explaining that the model must first be -loaded under its current revision. - -The next inference request builds options from the current configuration, -captures its revision, and performs the normal install, staging, and load -sequence. On success, it transactionally records the replica and current -`ModelLoadInfo`. The reconciler may then restore additional `min_replicas` -using that revision. - -Every scheduling and routing decision rechecks revision eligibility when it -claims a replica. A replica selected immediately before a concurrent edit must -fail the claim after the edit advances the current revision. Existing in-flight -requests may finish; no new request is assigned to the old replica. Graceful -cleanup waits for bounded `Free()` behavior and then terminates it. - -## API and observability - -Model and node lifecycle responses should expose, where replica details are -already returned: - -- current model `config_revision`; -- replica `config_revision`; -- `effective_options_hash`; -- lifecycle state, including `unloading`; -- pending cleanup error and retry time. - -Logs for routing, reconciliation, load completion, stale-load rejection, and -cleanup include model name, replica index, node ID, and abbreviated revision. -No serialized model options or request content is added to logs. - -The Web UI does not require a new workflow. After saving, it may show that the -configuration is saved while one or more old replicas are still being cleaned -up. User-facing distributed-model documentation explains this state and the -requirement to upgrade workers that lack acknowledged `model.stop` support. - -## Rolling upgrades - -Database migrations add nullable revision and cleanup columns so old binaries -can continue reading existing rows. New frontends treat missing revisions as -legacy state according to the compatibility rule above. - -The new NATS subject avoids changing the semantics of `backend.stop` for old -workers. A new frontend receiving no responder for `model.stop` leaves the -replica quarantined and reports the compatibility problem. It must not fall -back to fire-and-forget `backend.stop`, because doing so would recreate the -original false-success failure. - -Deployments should upgrade workers before or together with frontends. Mixed -frontend versions are tolerated at the database level, but old frontends do -not enforce revision-aware routing. Documentation must state that strict -cross-replica consistency is guaranteed only after all frontend replicas run -the revision-aware version. - -## Testing - -All Go tests use Ginkgo and Gomega. - -### Registry tests - -- Advancing a revision and quarantining old replicas is atomic. -- Only loaded replicas matching the current revision are returned for routing. -- Empty legacy revisions become stale after a revision-aware mutation. -- Load-info compare-and-set rejects a late old-revision write. -- Matching load information survives ordinary replica removal and worker - failure. -- Re-registration removes quarantined rows without changing current revision or - matching load information. - -### Router and reconciler tests - -- Given one 8K old-revision replica and one 100K current-revision replica, every - new request routes to the current revision. -- Changing `parallel` produces the same revision transition behavior as changing - `context_size`. -- The reconciler never loads from stale `ModelLoadInfo`. -- A late durable load job cannot publish a stale replica or overwrite current - load information. -- A request racing a configuration edit cannot claim the old generation. -- Heterogeneous effective option hashes remain routable when their - configuration revision matches. - -### Worker protocol tests - -- Exact process key and address stop the intended process and wait for exit. -- An address mismatch stops nothing. -- An already-absent process returns idempotent success. -- Graceful stop attempts bounded `Free()` and still terminates after a failure. -- Forced stop skips `Free()`. -- Replica port ownership and quarantine are updated before replying. - -### Lifecycle tests - -- Raw YAML edit, structured PATCH, rename, disable, and peer application all - advance or apply the expected revision and quarantine old replicas. -- A successful stop deletes the matching row. -- A timeout leaves a non-routable `unloading` row with retry state. -- Retry eventually deletes the row after the worker recovers. -- A worker without `model.stop` support produces a visible compatibility error - and never triggers fire-and-forget fallback. -- Partial cleanup does not roll back an already persisted configuration edit. - -### Live distributed regression - -An integration scenario loads a model on two workers, edits context and -parallel settings, and verifies that no request is routed to an old revision. -After cleanup and reload, every replica reports the current revision. The test -also disconnects one worker during the edit, verifies its replica is -quarantined, reconnects it, and verifies retry or re-registration removes the -stale row. - -## Documentation impact - -The implementation updates the distributed model lifecycle documentation under -`docs/content/` in the same change. It documents revision consistency, -quarantined cleanup state, rolling-upgrade requirements, and why an edited model -may wait for its first request before `min_replicas` can be restored. - -No configuration key or public inference API changes are introduced. diff --git a/docs/superpowers/specs/2026-08-21-distributed-staging-operations-design.md b/docs/superpowers/specs/2026-08-21-distributed-staging-operations-design.md deleted file mode 100644 index 856340e69..000000000 --- a/docs/superpowers/specs/2026-08-21-distributed-staging-operations-design.md +++ /dev/null @@ -1,67 +0,0 @@ -# Distributed Staging Operations Design - -## Problem - -`GET /api/operations` reads file-transfer progress from the frontend replica's -in-memory `StagingTracker`. Distributed frontends broadcast tracker updates over -NATS, but those messages are transient. A replica that starts after staging has -begun, temporarily disconnects, or misses an update can return no staging row. -When a browser's one-second polls are balanced across replicas, the operation -therefore appears and disappears. - -Distributed cold loads already persist their phase, placement, heartbeat, and -byte progress in PostgreSQL's `model_load_jobs` table. That row is the durable -cluster authority and should provide the baseline operations view. - -## Design - -Add a `NodeRegistry` query that lists active model-load jobs. The operations -endpoint will use those jobs to build one staging operation per tracking key -when the job is in the `staging` phase. It will then overlay matching local or -NATS-mirrored `StagingTracker` data, because the tracker can contain a fresher -message and filename than the periodically persisted job. - -The merge is keyed by the model tracking key. A tracker entry replaces the -database entry's progress and display details rather than creating a duplicate. -Tracker-only entries remain visible for compatibility with staging paths that -do not have a durable load-job row. Database-only entries remain visible on -every replica, which eliminates flicker. - -The database row supplies: - -- stable operation identity (`staging:`), -- model name and staging phase, -- node name, -- overall progress calculated by `ModelLoadJob.Progress()`, and -- byte counters used by the frontend's ETA calculation. - -The tracker overlay supplies its message, filename, node name, progress, and -byte counters when available. - -## Failure Handling - -If the database query fails, `/api/operations` will log the error and fall back -to the current tracker-only response. An observability failure must not break -the entire operations endpoint or hide unrelated gallery operations. - -Only live `staging` rows are included. Pending, backend-installing, loading, and -failed rows are represented by their existing user-facing flows and must not be -mislabelled as file staging. - -## Testing - -Add focused Ginkgo coverage for: - -1. A database-only staging job appears in the operations payload, reproducing - the request landing on a replica that missed all NATS broadcasts. -2. A matching tracker entry overlays the database entry without duplication. -3. Non-staging load jobs do not appear as staging operations. -4. A database read failure retains tracker-only staging operations and the - endpoint still succeeds. - -Run the affected Go package tests only; no long build is required. - -## Documentation - -This corrects consistency of an existing UI operation and introduces no new -API, option, or user workflow. No user documentation change is required. diff --git a/docs/superpowers/specs/2026-08-21-scheduling-rule-editing-node-labels-design.md b/docs/superpowers/specs/2026-08-21-scheduling-rule-editing-node-labels-design.md deleted file mode 100644 index 950aa7645..000000000 --- a/docs/superpowers/specs/2026-08-21-scheduling-rule-editing-node-labels-design.md +++ /dev/null @@ -1,125 +0,0 @@ -# Scheduling Rule Editing and Node Label Reference - -## Summary - -Improve the React scheduling view so cluster operators can edit existing scheduling rules and inspect node labels without moving back and forth to the Nodes page. - -The scheduling page will gain a compact, collapsible node-label reference above the rules table. It will also gain an Edit action that opens the existing scheduling form with the selected rule prefilled. The model name will remain locked while editing because it identifies the rule being updated. - -## Goals - -- Let operators update an existing scheduling rule in place. -- Make the labels available on each node visible from the scheduling workflow. -- Keep the label reference usable for clusters with many nodes. -- Preserve the existing scheduling API and node API contracts. -- Keep the scheduling rules usable when node-label loading fails. - -## Non-goals - -- Editing node labels from the scheduling page. -- Renaming the model associated with an existing scheduling rule. -- Adding backend endpoints or changing scheduling semantics. -- Adding a separate scheduling documentation page for this discoverability enhancement. - -## User Experience - -### Node label reference - -A collapsible **Node labels** section appears above the scheduling rules. It loads node data through the existing `nodesApi.list()` client and groups labels by node so operators can tell which selectors match which machines. - -The expanded section contains: - -- A fuzzy search field that matches node names, label keys, label values, and complete `key=value` text. -- A summary showing the visible result count and total matching node count. -- Node groups containing the node name, operational status, and its `key=value` label chips. -- Five matching nodes initially. -- A **Show 20 more** action when additional matches exist. - -Changing the search query resets the visible limit to five. Clearing the query restores the unfiltered result set. The reference can be collapsed to preserve vertical space. - -The matching implementation should be lightweight and local to the page. It should normalize searchable node data and support forgiving, case-insensitive token matching without adding a large dependency solely for this feature. - -### Editing a scheduling rule - -Each scheduling-rule row gains an **Edit** action beside **Delete**. Selecting Edit opens the existing scheduling form above the table and populates every editable field from the selected configuration: - -- Scheduling mode -- Node selector -- Minimum and maximum replicas -- Routing policy -- Prefix-cache thresholds - -The model selector is replaced by, or presented as, a visibly locked model field while editing. This prevents a rename from creating a second rule while leaving the original in place. - -Only one add or edit form may be open at a time. Opening Add clears edit state; opening Edit closes any blank Add form. Cancel closes the form and discards its local changes. - -Saving continues to use `nodesApi.setScheduling()`. On success, the page closes the form, shows the existing success toast, and refreshes the scheduling rules. On failure, it shows the error toast and keeps the populated form open so the operator does not lose changes. - -## Component Design - -### Scheduling form - -Refactor `SchedulingForm` to accept an optional existing scheduling configuration. Initial form state will be derived from that configuration, including conversion of a serialized `node_selector` when necessary and derivation of the current mode from `spread_all`, replica values, and selector presence. - -The form remains responsible for validation and for producing the existing scheduling request shape. The parent remains responsible for API calls, toast notifications, refreshes, and deciding whether the form is adding or editing. - -### Node label reference - -Add a focused scheduling-page component for label discovery. It receives node data and owns only presentation state: - -- Expanded or collapsed -- Search query -- Visible result limit - -Small pure helpers will normalize a node's searchable text and calculate filtered results. Node fetching remains in the scheduling page so loading and retry behavior stay next to the existing scheduling fetch lifecycle. - -### Styling - -Add scheduling-specific classes to `core/http/react-ui/src/App.css`. Reuse existing design-system tokens and button, input, badge, stack, and text primitives. Do not add static inline styles. - -On wide screens, node groups use a responsive compact grid. On narrow screens, they collapse to one column. Search, collapse, pagination, and row actions remain keyboard accessible and expose explicit accessible names. - -## Data Flow - -1. The page mounts and independently requests scheduling configurations and nodes. -2. Scheduling configurations populate the rules table. -3. Node data populates the label reference; local search and limiting do not trigger network requests. -4. Selecting Edit copies one rule into form state and locks its model identity. -5. Saving posts the existing scheduling payload and refreshes the scheduling list. -6. Node-label retry repeats only the node request and does not disturb scheduling rules or an open scheduling form. - -## States and Error Handling - -- **Node loading:** Show a compact loading state inside the reference. Do not block the rules table. -- **No nodes:** Explain that no nodes are available yet. -- **Node without labels:** Include it in node-name search results and display **No labels**. -- **No search matches:** Show a clear empty result while preserving the query. -- **Node fetch failure:** Show an inline error with Retry. Scheduling remains fully usable. -- **Malformed selector:** Preserve the current defensive rendering behavior and avoid crashing the edit form; treat an unparseable selector as empty while keeping the rule visible. -- **Save failure:** Preserve all form values and show the existing error toast. -- **Save success:** Close the form and refresh the rules. - -## Verification - -Add or extend a focused Playwright scheduling spec to cover: - -- Labels grouped under the correct nodes. -- Search by node name. -- Search by complete `key=value` text. -- Five-node initial limit and **Show 20 more** expansion. -- Empty-cluster, unlabeled-node, no-match, and failed-loading states. -- Edit opening with the complete rule prefilled. -- Locked model identity during editing. -- Updated values sent through the existing scheduling endpoint. -- Failed saves preserving the open form. - -Run the focused Playwright spec, the React inline-style lint, and the production React build. Long repository-wide builds are outside the scope of this frontend-only change. - -## Acceptance Criteria - -- An operator can edit and save any existing scheduling rule without deleting and recreating it. -- The model identity cannot be changed while editing. -- An operator can inspect labels grouped by node without leaving Scheduling. -- The label reference remains compact with many nodes and supports forgiving search plus progressive expansion. -- A node API failure does not prevent viewing or editing scheduling rules. -- The enhancement works at narrow viewport widths and is keyboard accessible. diff --git a/docs/superpowers/specs/2026-09-07-ephemeral-staging-retention-design.md b/docs/superpowers/specs/2026-09-07-ephemeral-staging-retention-design.md deleted file mode 100644 index e84d36193..000000000 --- a/docs/superpowers/specs/2026-09-07-ephemeral-staging-retention-design.md +++ /dev/null @@ -1,158 +0,0 @@ -# Request-owned ephemeral staging - -## Problem - -Distributed requests copy transient inputs below -`/ephemeral//`. The worker currently removes -these files only when a periodic age sweep considers them stale. A Reachy Mini -sending camera and sound data about once per second created more than 21,000 -request directories and filled its Mac worker before the six-hour retention -window elapsed. - -Reducing the retention window is insufficient. A time limit bounds residence -time, but the retained bytes still scale with request rate and input size. A -quota sweeper would also have to infer whether an old file is still in use. -Neither rule prevents concurrent uploads from consuming the worker's last free -space. - -## Goals - -- Give every ephemeral input an explicit owner and release it when that request - finishes, fails, or is cancelled. -- Keep cleanup transport-independent for HTTP and S3/NATS workers. -- Reserve capacity before accepting ephemeral bytes so concurrent requests - cannot consume configured disk headroom. -- Reject a request cleanly when its input does not fit; never evict an input - that a running request may still be reading. -- Recover abandoned files after frontend or worker crashes. -- Never inspect or remove models, data, configuration, or paths outside the - worker's ephemeral staging tree. - -## Non-goals - -- Retaining request inputs as a cache. -- Evicting persistent model or data files to make an inference request fit. -- Treating modification timestamps as proof that a request is active. - -## Request ownership - -The `FileStagingClient` already creates one request ID before staging inputs and -waits for synchronous and streaming backend calls to finish. It will track each -ephemeral key before attempting to stage it and defer one request-scoped -release identified by the request ID. Release runs after the backend call -returns, including error and cancellation paths, using a short background -timeout so cancellation of the request does not cancel its cleanup. - -Request IDs will use the full UUID rather than the current eight-character -prefix. The worker enumerates only category directories for that validated -request ID and removes each entry with exact, symlink-safe deletion. - -`FileStager` will expose an idempotent exact-key `ReleaseRemote` operation and -an optional request-scoped operation. The client uses one fixed-size request -message for the normal path and retains exact-key calls as a rolling-upgrade -fallback: - -- HTTP sends one authenticated request containing the fixed-size request ID. - The worker derives and removes that request's exact files, then prunes empty - request and category directories without following symlinks. -- S3/NATS sends one request-reply containing the request ID so the selected - worker evicts the request's local cached files. The frontend then deletes the - matching objects from its tracked exact-key list. - Either deletion may already have happened and still counts as success. - -If staging fails partway through a request, the deferred release still includes -the planned key, allowing it to remove a partial file when the transport can -identify one. Cleanup errors are logged and do not replace the inference result. - -HTTP and S3 ingress register request operations before any pre-reservation -work. Before enumerating files, the capacity guard marks the request released -and waits for registered operations and admitted writes to finish. Later -operations, reservations, and cache claims for that request are rejected. -Markers expire after one hour and are capped at 16,384 entries, but a marker is -never evicted while its registered operation or cleanup scan is active. -Concurrent operation and cleanup state have the same hard cap. Disk bytes -remain independently bounded by capacity admission. -Cleanup waits within its deadline when all cleanup-pin slots are occupied. -If that deadline expires, existing entries lose active ownership so recovery -can reclaim them; registered ingress for the request remains closed until it -exits. - -## Capacity admission - -A worker-local ephemeral capacity guard is shared by its HTTP and S3/NATS input -paths. It accounts for both `/ephemeral`, used by HTTP, and -`/ephemeral`, used by S3 downloads. Before writing an ephemeral object, -the transport reserves its declared size. HTTP obtains the size from the upload -metadata; S3/NATS obtains it from object metadata. Reservations are serialized -in memory, cover both committed ephemeral bytes and concurrent writes, and are -returned on release or failed transfer. - -Admission succeeds only when both conditions remain true after the reservation: - -1. Total ephemeral bytes remain below the configured ephemeral staging limit. -2. The filesystem retains the configured minimum free-space headroom. - -The guard rejects the transfer before inference when either condition fails. -An input with unknown size is written through a bounded accounting writer that -reserves fixed-size chunks before writing each chunk and stops before crossing -the limit. The existing maximum-upload-size check remains the per-file ceiling. - -The limit and headroom are worker settings. By default, ephemeral data may use -the smaller of 10 GiB or 10 percent of filesystem capacity, while the worker -preserves the larger of 1 GiB or 5 percent as free-space headroom. The worker -logs the effective values at startup. A zero or negative operator value selects -the default rather than disabling protection. The guard scans the ephemeral -tree at startup to account for abandoned committed bytes. Filesystem free-space -checks are repeated at reservation time because other processes may share the -volume. - -## Crash recovery - -The existing periodic cleanup remains as a fallback for ownership messages lost -when a frontend or worker process dies. It uses a one-hour recovery TTL, -performs one startup sweep, and repeats every 15 minutes. It skips every key -held by an active reservation, considers the newest modification time in each -remaining request tree, and does not follow directory symlinks. It removes only -request directories below the registered `/ephemeral` and -`/ephemeral` roots. - -The recovery window does not control normal storage growth. Request completion -and capacity reservations do. A recovery deletion updates the capacity guard's -accounted bytes. - -## Error handling and observability - -Admission failures report the requested bytes, current ephemeral usage, limit, -available bytes, and required headroom. Successful release and recovery update -usage counters. Read, stat, and remove failures include the affected path and -allow unrelated cleanup to continue. Missing ephemeral files and directories -are normal for idempotent release. - -## Testing - -Regression tests will establish the following behavior: - -1. Successful, failed, cancelled, and streaming calls issue one request-scoped - worker cleanup only after the backend has returned. -2. Partial staging failures release the planned key without changing the main - error returned to the caller. -3. HTTP and S3/NATS release remove local files; S3/NATS also removes the object. -4. Release rejects persistent keys and path traversal, does not follow - symlinks, and leaves paths outside `ephemeral` untouched. -5. Concurrent reservations cannot exceed the byte limit or free-space - headroom, and failed transfers return their reservations. -6. Unknown-length writes stop at the capacity boundary. -7. Startup accounting includes abandoned ephemeral files, and the recovery - sweep removes only stale, inactive leftovers and updates accounting. - -Focused package tests will run with race detection, followed by the relevant -repository lint and vet checks. - -## Rollout - -The change requires a new LocalAI worker and frontend build because both sides -participate in release. The Mac worker starts by accounting for its existing -backlog and removing recovery-expired files. The deployment check will verify -available space, admission and release logs, stable ephemeral usage under -continuous camera and audio traffic, and successful vision, sound detection, -and transcription requests. diff --git a/docs/superpowers/specs/2026-09-07-exl3-gallery-design.md b/docs/superpowers/specs/2026-09-07-exl3-gallery-design.md deleted file mode 100644 index b4175655f..000000000 --- a/docs/superpowers/specs/2026-09-07-exl3-gallery-design.md +++ /dev/null @@ -1,99 +0,0 @@ -# EXL3 gallery entries - -## Goal - -Add four gallery entries that expose the EXL3 configurations validated or -tracked by `vllm.cpp`. Pin each Hugging Face artifact to the revision recorded -by its source or benchmark evidence. - -## Entries - -### Qwen3.8 target - -Add `qwen3.8-27b-exl3-vllm-cpp` for -`Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw`. This entry serves the target without a -draft model. - -Use revision `19441ac874c4018295da848e250f23511361cda4`. Configure an 8,192-token -context, 2,048 cache blocks, eight sequences, and 16,384 batched tokens. Disable -prefix caching to match the measured serving configuration. - -### Qwen3.8 with DFlash2 - -Add `qwen3.8-27b-dflash2-exl3-vllm-cpp`. This entry stages the Qwen3.8 target -and `Mia-AiLab/Qwen3.8-27B-DFlash2-EXL3-5.0bpw` at revision -`4f0436269bca761b071f05319e8e04a87cc633f9`. - -Configure the `dflash` method with seven speculative tokens. Use the shipped -paged draft route. Apply the same serving limits as the target-only entry. - -Tag this entry with `dflash` because it enables speculative decoding. Declare -the target-only entry as its variant. LocalAI can then prefer the faster entry -when the host supports it. - -### DeepSeek V4 Flash for Spark - -Add `deepseek-v4-flash-spark-exl3-vllm-cpp` for -`0xSero/deepseek-v4-flash-0731-spark`. Use the current repository revision, -`ce5ff0f1efb2e184aafc759d281bfae47d3a359c`. State that the `vllm.cpp` -runtime record used the older revision `22f28d32b9b29b4352eaa380ff8c2c170b2847ab`. - -Describe the entry as a Spark and GB10-oriented REAP-K216 checkpoint. State its -large memory requirement and CUDA requirement. Do not claim a completed speed -or correctness gate that the source record does not contain. - -### DeepSeek V4 Flash 3.0 bpw - -Add `deepseek-v4-flash-exl3-3bpw-vllm-cpp` for -`0xSero/DeepSeek-V4-Flash-0731-EXL3-3.0bpw`. Use the current repository -revision, `e0bf84ac76a5100e8790c22ad10b70b1e2d06d71`. - -Tag and describe this entry as experimental. The model card states that the -artifact is structurally complete, but end-to-end generation has not passed. -Keep this entry separate from the Spark entry because the repositories use -different layouts and have different runtime evidence. - -## Artifact staging - -Use LocalAI's Hugging Face artifact source for each repository. Stage complete -model repositories because these safetensors checkpoints need configuration, -tokenizer, index, and weight files. - -Assign the Qwen draft artifact to a companion target. Pass its staged path in -the `vllm-cpp` speculative configuration. Do not download files through backend -startup logic. - -## User-visible metadata - -Use the `vllm-cpp`, `exl3`, `gpu`, and `cuda` tags on all four entries. Add -architecture, reasoning, tool-calling, and speculative-decoding tags only when -the configured model supports them. - -Descriptions must distinguish measured results from unresolved work. The Qwen -DFlash2 description can cite the measured configuration and throughput. The -DeepSeek descriptions must not imply an end-to-end validation that does not -exist. - -## Validation - -Run the gallery schema and focused gallery tests. Add a focused test if the -artifact or variant structure is not already covered. - -Validate these properties: - -- Every name is unique. -- Every variant points to an existing entry. -- Each Hugging Face source has a pinned revision. -- The DFlash2 entry stages both repositories and passes the draft path. -- Only the configured DFlash2 entry has the `dflash` tag. -- YAML parsing and gallery loading succeed. - -No model download or GPU benchmark is part of this LocalAI change. The -`vllm.cpp` evidence supplies the runtime record. - -## Out of scope - -- Changes to the `vllm-cpp` backend binaries. -- New EXL3 kernels or model loaders. -- New benchmark claims. -- Gallery entries for unselected EXL3 bit widths. diff --git a/docs/superpowers/specs/2026-09-26-failover-distributed-proxy-ui-design.md b/docs/superpowers/specs/2026-09-26-failover-distributed-proxy-ui-design.md deleted file mode 100644 index 7ae52561b..000000000 --- a/docs/superpowers/specs/2026-09-26-failover-distributed-proxy-ui-design.md +++ /dev/null @@ -1,356 +0,0 @@ -# Failover chains: distributed mode, localai-proxy backend and WebUI - -Date: 2026-09-26 -Status: design approved in brainstorming, pending spec review -Builds on: `2026-09-26-model-failover-chains-design.md` (same PR) - -## Problem - -The failover chains in this PR work on one LocalAI instance. Three gaps -remain: - -1. **Distributed mode.** With several frontends, chain definitions converge - (config edits broadcast `cache.invalidate.models`), but runtime state does - not. Each frontend probes, trips and pins on its own. A pin applies only - on the frontend that received it and is lost on restart. `/api/failover` - and its event stream show a different view on each frontend. `warm: true` - pins only a frontend stub, so workers can evict the model, and every - frontend preloads it. -2. **Remote targets cover only chat.** `cloud-proxy` forwards chat and - completions. A remote LocalAI cannot serve transcription, TTS, VAD, sound - detection or the other modalities as a chain target, so a realtime - pipeline cannot fail over per stage between a remote and a local LocalAI. -3. **No UI.** Chains can be edited only as raw JSON in the model editor, and - their health is visible only through the API. - -LocalAI also has no rule that makes a feature state how it behaves with -several frontends. The failover feature shipped with per-instance state -because nothing asked the question. - -## Goals - -- **C. Distributed-aware failover.** Pins, target health and chain state are - the same on every frontend. One frontend probes. Warm targets stay loaded on - workers and are preloaded once. -- **A. `localai-proxy` backend.** A gRPC backend that serves every backend - method with a REST counterpart by calling an upstream LocalAI, including - live transcription through the upstream's realtime API. -- **B. WebUI.** A chain editor field, a chain template, a live health strip - per chain, a "chain" badge in the model list, and a Failover overview page. -- **D. Contributor rule.** `AGENTS.md` and a new `.agents/distributed-state.md` - require every stateful feature to choose and document a distributed mode. - -## Non-goals - -- Sharing failover state between instances that are not in one distributed - cluster. -- A bridge for backend methods that have no REST or realtime counterpart - upstream (audio encode/decode, metrics, status, fine-tune, quantization). -- Chains as router candidates or as the realtime classifier model. - ---- - -## C. Distributed-aware failover - -Standalone mode (no NATS, no PostgreSQL) keeps today's behaviour. Everything -below applies when distributed mode is on. - -### Shared state - -| State | Writers | Mechanism | Survives restart | -|---|---|---|---| -| Pins | any frontend (REST, MCP) | `syncstate.SyncedMap` named `failover.pins`, key = chain, with a gorm `Store` | yes | -| Target health: state, last error, since, consecutive passes | any frontend on a local transition, and the probe leader | `syncstate.SyncedMap` named `failover.targets`, key = target, NATS only | no | -| Chain state: active target, active since, chain state | the probe leader only | `syncstate.SyncedMap` named `failover.chains`, key = chain, NATS only | no | - -- The pins table is created under `advisorylock.KeySchemaMigrate`, the same - way the jobs store creates its tables. -- The manager gets a small `StateSync` dependency. The standalone - implementation is a no-op; the distributed implementation wraps the three - maps. The manager does not import NATS or gorm directly. -- A peer delta is applied through `OnApply`, which changes local state - without publishing again (no echo loops). -- The NATS-only maps have no `Store`, so a `Reconcile` tick's hydrate would be - a no-op — nothing durable to pull from, so it could never help a late - joiner. Instead the leader republishes every target and chain snapshot every - 10 s. A frontend that joins late converges within 10 s and uses its own - state until then. - -### Who does what - -- **Every frontend** plans requests from the shared state. `Plan` already - leaves unhealthy targets out of the attempt order, so a target tripped on - another frontend is skipped at once. In-request retry stays local. -- **Any frontend** that sees a real request trip or pass a target publishes - the new target state. -- **The probe leader** runs probes, recovery confirmation, dwell-based - fail-back, the chain recompute and the warm preload. It publishes chain - state. Leadership uses `advisorylock.RunLeaderLoop` with a new key - `failover-prober` and the same 1 s interval as the scheduler. -- **Followers** do not recompute the active target. They adopt the leader's - chain state. When no chain state has arrived yet (start-up), a follower - uses its own recompute until the first delta. -- If the leader stops, another frontend takes the lock on its next tick. - Pins and target health are not affected. Probes and fail-back pause for at - most one tick. - -### Events - -Each frontend emits `chain.switched` and `target.state` to its own -subscribers (SSE, realtime `localai.model.failover`) when it applies a -change, whether the change is local or from a peer. Every frontend's stream -therefore shows the same events. - -### Warm targets - -- The SmartRouter's and ReplicaReconciler's pinned-model resolver includes - `WarmTargets()`, so workers never evict a warm target. -- Only the leader preloads warm targets. -- A frontend's loaded check sees only its own model stubs. A warm target - without a stub on the leader is treated as not loaded: its liveness passes - and its recovery is inconclusive (it heals after `min_dwell`). Worker health - is left to the node health monitor and to real requests. - -### Tests - -- Unit: two managers on the test fakebus (`core/services/testutil`). A pin on - one shows on the other. A trip on one is skipped by the other's plan. Only - the lock holder probes. A new leader resumes fail-back. -- A spec in `tests/e2e/distributed` when its harness supports two frontends - cheaply; otherwise the unit specs are the coverage and the PR says so. - -### Docs - -`model-failover.md` replaces the "state is per instance" limit with a -"Distributed mode" section that describes the table above. - ---- - -## A. `localai-proxy` backend - -### Shape - -- `backend/go/localai-proxy` is a separate OCI gallery backend, like - `cloud-proxy`. It is registered in the `Makefile`, `backend/index.yaml` and - `.github/backend-matrix.yml` (Linux amd64/arm64 and Darwin Metal), following - `.agents/adding-backends.md`. -- It reuses cloud-proxy's auth header, HTTP client (no redirects) and - hop-by-hop header helpers. It has no translate mode: the upstream is always - LocalAI. -- `Load` refuses a model without proxy options, so greedy backend probing - never selects it. - -### Config - -```yaml -name: argus-whisper -backend: localai-proxy -known_usecases: [transcript] -proxy: - upstream_url: https://argus:8080 # base URL; each method appends its path - upstream_model: whisper-large # optional; default: this model's name - api_key_env: ARGUS_KEY - request_timeout_seconds: 60 # applies to non-streaming calls -``` - -- `core/backend/options.go` passes `ProxyOptions` to `localai-proxy` as well - as `cloud-proxy`. -- `proxy.mode` and `proxy.provider` are ignored, with a load warning. -- A `localai-proxy` model without `known_usecases` loads with a warning, because - usecases decide default-model selection and the failover inference probe. -- The failover prober already treats `localai-proxy` as remote. - `UpstreamBase` accepts a base URL unchanged. - -### Method mapping - -| Backend method | Upstream endpoint | -|---|---| -| Predict, PredictStream | `/v1/chat/completions`, `/v1/completions` (SSE when streaming) | -| Embedding | `/v1/embeddings` | -| Rerank | `/v1/rerank` | -| TokenizeString, Detokenize | `/v1/tokenize`, `/v1/detokenize` | -| Score | `/api/score` | -| GenerateImage, UpscaleImage | `/v1/images/generations`, `/v1/images/upscale` | -| GenerateVideo | `/video` | -| Generate3D, Animate3D | `/3d/generations`, `/3d/animate` | -| TTS, TTSStream | `/tts` (streaming passes the upstream WAV header and PCM through) | -| SoundGeneration | `/v1/sound-generation` | -| AudioTranscription, AudioTranscriptionStream | `/v1/audio/transcriptions` (`stream=true` for SSE deltas) | -| AudioTranscriptionLive | upstream `/v1/realtime` transcription session (see below) | -| Diarize | `/v1/audio/diarization` | -| VAD | `/v1/vad` | -| SoundDetection | `/v1/audio/classification` | -| Detect, Depth | `/v1/detection`, `/v1/depth` | -| FaceVerify, FaceAnalyze | `/v1/face/verify`, `/v1/face/analyze` | -| VoiceVerify, VoiceAnalyze, VoiceEmbed | `/v1/voice/verify`, `/v1/voice/analyze`, `/v1/voice/embed` | -| Stores* | `/stores/set`, `/stores/get`, `/stores/delete`, `/stores/find` | -| AudioTransform | `/audio/transformations` | - -Every request uses the upstream model name (`proxy.upstream_model`, else the -model name), the same derivation as `failover.UpstreamModel`. - -Methods with no counterpart (AudioEncode, AudioDecode, AudioToAudioStream, -TokenClassify, GetMetrics, Status, ModelMetadata, fine-tune and quantization) -return gRPC `Unimplemented` with the message -`localai-proxy: has no upstream counterpart`. - -### Files - -Core passes some inputs and outputs as local paths: - -- Inputs (transcription and diarization audio, sound detection `src`, image - `src` and reference images): the proxy reads the file and uploads it as - multipart or base64, as the endpoint expects. -- Outputs (TTS, image, sound generation `dst`): the proxy writes the upstream - result (bytes, or a download of the returned URL, or decoded base64) to - `dst`. - -### Live transcription bridge - -The upstream realtime API needs a pipeline model (VAD and transcription). The -proxy takes it from the model's backend options: - -```yaml -options: - - realtime_pipeline:argus-transcribe # an upstream pipeline config -``` - -Without this option, `AudioTranscriptionLive` returns the standard -"live transcription unsupported" error, and realtime uses its non-live -transcription path for the stage. - -With the option, `AudioTranscriptionLive` opens a WebSocket to -`/v1/realtime?model=`: - -1. On the first `TranscriptLiveConfig`, send `session.update` with - `type: transcription`, the input rate, the language and server VAD turn - detection. Answer `ready` when `session.updated` arrives. -2. Forward each `TranscriptLiveAudio` as `input_audio_buffer.append` - (PCM float to PCM16 base64 at the session rate). -3. Map `conversation.item.input_audio_transcription.delta` to `delta`, and - `...completed` to `delta` (any remaining text) plus `eou: true`. -4. When the gRPC send side closes, do **not** commit the buffer: the upstream - rejects a manual commit under server VAD. Instead wait up to 5 s for any - turn already in flight (speaking, stopped-but-not-committed, or committed - but not yet completed) to finish on its own, then send `final_result` and - close. -5. An upstream error or disconnect ends the gRPC stream with `Unavailable`. - Word timings and `eob` are not available from the upstream and stay empty. - -A realtime stage whose live session fails reopens on the next chain target at -the next utterance (behaviour from the base spec). - -### Core changes - -- **Rerank for Go backends.** `pkg/grpc` gets an optional rerank interface and - a server handler, in the same way as `Score`. -- **`Unimplemented` is a capability gap.** The failover retry path (HTTP and - `Manager.Do`) treats gRPC `Unimplemented` like an admission rejection: skip - to the next target for this request, and do not trip the target. Otherwise a - chain of a remote and a local target fails a request the local target can - serve. - -### Tests - -- Unit: a fake LocalAI `httptest` upstream per method family (request path, - body, model name, auth header, file upload and `dst` write). -- Unit: a fake WebSocket upstream for the live bridge (ready, deltas, eou, - final result, upstream disconnect). -- E2E: `localai-proxy` models that point back at the test server's own mock - models. A realtime pipeline whose stages are chains of a `localai-proxy` - target and a local target completes a turn, and switches stage when the - proxy target fails. - -### Docs - -A `localai-proxy` section in `docs/content/features/backends.md` (or the page -that documents `cloud-proxy`), and a remote-LocalAI example on -`model-failover.md`. - ---- - -## B. WebUI - -### Model editor - -- A `failover-targets` field component replaces the JSON editor for - `failover.targets` (`core/config/meta/registry.go` switches the component - name). Each row has a model picker (`SearchableModelSelect`), move up/down, - remove and a **warm** toggle. The toggle is disabled with a tooltip on remote - targets. Inline validation: at least 2 targets, no duplicates, no chain as a - target. -- The probe, trip and recovery fields stay in the Advanced group. -- A **Failover chain** template in `modelTemplates.js`, seeded with two empty - targets. `?template=failover` preselects it. - -### Health strip - -When the edited model is a chain, a `FailoverChainStatus` component above the -form shows: - -- the chain state (primary / fallback / degraded) and the active target, with - the time since it became active; -- for each target: state, kind, warm, last probe and last error; -- **Pin** and **Unpin** for admins, behind a confirm dialog. The pinned target - is marked. - -### Model list and overview - -- Installed models: a "chain" badge with the active target, in the same way - as the alias badge. -- Operate → Runtime → **Failover**: a dense table with one row per chain - (state, active target, target states, time since the last switch). Each row - links to the chain in the model editor. With no chains, an empty state links - to the Failover chain template. - -### Live data - -A `useFailoverChains` hook fetches `GET /api/failover`, then opens an -`EventSource` on `/api/failover/events`. `snapshot` replaces the state; -`chain.switched` and `target.state` patch it. The browser reconnects the -stream, and the hook polls every 15 s as a fallback (the Agent Status -pattern). A `failoverApi` group in `src/utils/api.js` holds the calls. - -### Conventions - -- Design tokens and CSS classes only; no new inline styles (inline-style - ratchet). -- `StatusPill` tones: success for healthy and primary, warning for recovering - and fallback, error for down and degraded, muted for missing. -- Strings in the `models` and `admin` i18n namespaces for all 8 locales. -- Pin controls are hidden when `useAuth().isAdmin` is false. - -### Tests - -Playwright specs with mocked APIs and a mocked `text/event-stream`: the editor -component, the template, the health strip updating on events, pin controls -for admins and not for other users, and the overview page. UI line coverage -stays at or above `core/http/react-ui/coverage-baseline.txt`. - ---- - -## D. Contributor rule - -- New guide `.agents/distributed-state.md`. A feature that keeps runtime - state (in-memory maps, caches, pins, schedulers, background loops, probes) - chooses one mode and documents it: - - **shared**: `syncstate.SyncedMap`, with a `Store` when the state must - survive a restart; - - **single-runner**: an `advisorylock` leader loop; - - **stateless per request**; - - **per-instance**: allowed only with the reason written in the feature's - docs. - - The guide gives one real example per mode (finetune jobs, the node health - monitor, open responses, failover chains). Shared and single-runner - features include a fakebus test with two instances. -- `AGENTS.md`: a Quick Reference bullet "Distributed-aware state" and a row in - the Topics table. -- `.agents/api-endpoints-and-auth.md`: a checklist line "Stateful feature: - distributed mode chosen and documented (see distributed-state.md)". - -## Order of work - -C first (it changes code already in the PR and the event contract the UI -reads), then A (it needs the `Unimplemented` classification and the rerank -handler), then B, then D. All in PR #12285. diff --git a/docs/superpowers/specs/2026-09-26-model-failover-chains-design.md b/docs/superpowers/specs/2026-09-26-model-failover-chains-design.md deleted file mode 100644 index 0765ddb97..000000000 --- a/docs/superpowers/specs/2026-09-26-model-failover-chains-design.md +++ /dev/null @@ -1,415 +0,0 @@ -# Model failover chains - -Date: 2026-09-26 -Status: design approved in brainstorming, pending spec review - -## Problem - -A LocalAI instance that serves a model from a remote upstream (for example a -`cloud-proxy` model that points at a larger LocalAI cluster) has no way to fall -back to a local model when that upstream is unhealthy. Clients that want this -today build it themselves. The wingman voice assistant, for example, keeps an -ordered list of realtime WebSocket endpoints, quarantines an endpoint after -repeated backend errors, and reconnects to the next one. This has three costs: - -- Every client reimplements failover, health tracking and recovery. -- The switch happens at the session level. A realtime conversation loses its - context when the client reconnects to another endpoint. -- The client sees at least one failed request before it reacts. - -## Goal - -A model config can declare an ordered **failover chain** of target models. -LocalAI serves each request for the chain from the highest-priority healthy -target, retries on the next target when a target fails before the response is -committed, probes targets actively, fails back with hysteresis, and tells -clients when the active target changes. - -Because a chain is a model name, it works in every place that takes a model -name, including the `llm`, `transcription`, `tts` and `vad` fields of a -realtime pipeline. The realtime session stays on the local instance, so a -switch changes the stage that serves the next turn and the conversation -history survives. - -## Non-goals - -- A dedicated WebUI chain editor and live health view. That is a follow-up - (sub-project 3). This spec only registers the config fields in the field - metadata registry, so the generic model editor can show them. -- The `localai-proxy` backend (a remote target kind that covers the full - LocalAI API through gRPC). That is a follow-up (sub-project 1). This spec - defines the "remote target" probe path it will plug into. -- Shared failover state across several LocalAI frontends. State is in memory - and per instance. -- Transparent failover after a response has started to stream. - -## Config schema - -A chain is a model config with a `failover` block. Like an alias, it has no -backend of its own. - -```yaml -name: assistant-llm -failover: - targets: - - model: argus-llm # for example a cloud-proxy config - - model: gemma-local - warm: true # load at startup, never evict - probe: - interval: 15s # liveness probe interval - timeout: 5s - trip: - errors: 1 # retryable failures in `window` that mark a target down - window: 30s - recovery: - probes: 3 # consecutive inference probes to confirm recovery - min_dwell: 60s # minimum time on a lower target before fail-back -``` - -Only `targets` is required. The other values in the example are the defaults. - -### Validation - -The config loader rejects a chain when: - -- `failover` is set together with `alias` or `backend`. -- `targets` has fewer than 2 entries. -- A target does not exist, or the same target is listed twice. -- A target is itself a chain. Chains do not nest. - -A target can be an alias. The alias is resolved one hop, as it is today. - -The loader logs a warning, and does not reject, when: - -- The known usecases of the targets do not overlap (for example an LLM and a - TTS model in one chain). Usecases are often inferred, so this cannot be an - error. -- `warm: true` is set on a remote target. The flag has no effect there. - -### Target kinds - -The kind decides how a target is probed. It is inferred from the backend of the -target config: - -- **remote**: `cloud-proxy`, and `localai-proxy` when it exists. -- **local**: every other backend. - -### Naming in responses and accounting - -The behaviour matches aliases. Responses echo the chain name. Usage and traces -record `requested=` and `served=` through the existing -`ContextKeyRequestedModel` and `ContextKeyServedModel` keys. - -The upstream of a remote target never sees the chain name. A request served -through a chain reaches a remote target with that target's upstream model: -`proxy.upstream_model`, or the target name when it is empty. This holds in -passthrough and translate mode, and it is the same name the liveness probe -looks for in `/v1/models` (one helper derives both). - -## Failover manager - -New package: `core/services/failover`. The application creates one `Manager` at -start and keeps it in sync with the model config loader. When a chain is added, -edited or removed (from YAML, the model editor API or the MCP tools), the -manager updates without a restart. - -### State - -- Health is tracked **per target**. A target that is in two chains is probed - once, and a failure marks it down for both. -- The active target is tracked **per chain**. - -Target states: - -``` - trip (errors within window, from requests or probes) - healthy ───────────────────────────────▶ down - ▲ │ liveness probe passes - │ recovery.probes consecutive ▼ - └──── inference probes pass ◀──── recovering ──(any failure)──▶ down -``` - -- At startup, targets are `healthy`. The manager runs one liveness pass - immediately, and in-request retry covers the gap until it completes. -- A target whose config is removed becomes `missing`. It is treated as `down`, - and the chain reports it. - -Chain states: - -- `primary`: the active target is target 0. -- `fallback`: the active target is a lower target. -- `degraded`: all targets are down. - -### Selecting the active target - -- The active target is the highest-priority `healthy` target. -- Failover to a lower target is immediate. -- Fail-back to a higher target happens only when that target is `healthy` - (which already needs `recovery.probes` passing inference probes) **and** the - current target has been active for at least `min_dwell`. -- When the chain is `degraded`, requests still try every target in priority - order. A probe can lag behind a recovery, so the manager does not fail fast. -- A manual pin (see API) forces the active target. While a pin is set, probes - continue and report state, but they do not change the active target. - -### Probes - -| Target | Liveness (steady state) | Recovery confirmation | -|---|---|---| -| remote | `GET /v1/models` returns 2xx and lists the upstream model. `` is the scheme and host of `proxy.upstream_url` plus any path prefix before `/v1`. The upstream model is `proxy.upstream_model`, or the target name when it is empty. `/v1/models` works on any OpenAI-compatible upstream, and `/readyz` exists only on LocalAI. | one minimal real request, chosen by usecase | -| local, `warm: true` | gRPC `HealthCheck` on the loaded backend, with the probe timeout. A probe never loads the model: when the backend is not loaded (the warm preload is still loading it, or a crash removed it), liveness passes and the next real request loads and judges it. | chat and completion: `Predict` with 1 token; embeddings: `Embedding` of `"ping"`; other usecases: `HealthCheck`. A local backend process that answers `HealthCheck` rarely fails only for TTS or transcription. When the backend is not loaded there is nothing to confirm against: the probe neither passes nor trips, and the target returns to `healthy` like a cold one, when `min_dwell` has passed since the trip. | -| local, cold | none: a cold target is judged only by real requests; it is never loaded only to probe it. | none. After a trip, the target returns to `healthy` when `min_dwell` has passed. The next real request is the test. | - -Minimal requests by usecase: - -- chat and completion: `max_tokens: 1` -- embeddings: the input `"ping"` -- transcription: 200 ms of silence -- TTS: the text `"ok"` -- expensive usecases (image, video, 3D): no inference probe. Liveness is the - confirmation. - -Probe load rules: - -- One global scheduler ticks every second, without jitter, and starts the - probes that are due. A target shared by chains is probed once. The - scheduler does not wait for a probe: a target whose probe is still running - is skipped, so one slow target does not delay the others. -- A successful real request counts as a liveness pass, so a busy target is - almost never probed. -- Inference probes run only while a target is `recovering`. -- Remote probes use the URL and API key from the target's proxy config. -- Probe results go into the same trip counter as request failures. - -When one target is in several chains, its `probe`, `trip` and `recovery` -settings come from the first of those chains in name order. The docs state -this. - -### Warm targets - -The manager loads `warm: true` targets at startup and marks them pinned in the -watchdog, so LRU and idle eviction skip them. They still count toward the -active backend limit. When pinned warm targets leave no room for another load, -the loader never evicts them: it retries eviction and then loads the model -anyway, over the limit, with no error that names the warm targets. The docs -state this. - -### Events - -Each change of a target state or of an active target produces an event on an -internal bus: - -``` -{chain, target, from, to, state, reason, error, at} -``` - -`reason` is one of `trip`, `recovery`, `manual`, `degraded`, `missing`. - -## Request path (HTTP) - -### Resolution - -In `core/http/middleware/request.go`, next to the alias block, a chain config is -resolved with `mgr.Plan(chain)`. The plan is the ordered list of attempts: - -1. the active target, -2. the other `healthy` targets in priority order, -3. the `down` targets, only when the chain is `degraded`. - -The middleware stores the plan in the request context and sets -`MODEL_CONFIG` to the config of the first target. Handlers do not change. The -plan is fixed when the request starts, so a config reload does not affect -requests that are in progress. - -### In-request retry - -A new middleware wraps the handler. It replaces the response writer with one -that records whether the response is committed: - -- A response with status 500 or higher is buffered until the handler returns, - as long as no body has been flushed. Error bodies are small, so the wrapper - can discard them. -- A streaming response (SSE, or any flushed body) is committed at the first - flush. - -When the handler returns a retryable error and the response is not committed, -the wrapper calls `mgr.ReportFailure(target, err)`, sets the config of the next -target in the plan, and runs the handler again. When the handler succeeds, the -wrapper calls `mgr.ReportSuccess(target)`. When the response is committed and -then fails, the wrapper reports the failure and does not retry. - -Request bodies: - -- JSON bodies are already parsed into the request context. -- Multipart bodies are cached by `ParseMultipartForm`. -- Other bodies are buffered up to a limit. A larger body gets no in-request - retry. The failure still counts toward the trip. - -### Retryable errors - -`failover.IsRetryable(err, status)` returns true for: - -- connection and dial errors -- timeouts, when the client did not cancel the request -- gRPC `Unavailable`, `Internal`, `DeadlineExceeded` and `Unknown` -- upstream HTTP 5xx -- model load failures - -It returns false for client cancellation, 4xx responses and validation errors. -These do not trip a target, because the next target would reject the same -request. - -### Handler audit - -Each endpoint family must be safe to run again before its response is -committed: chat, completions, embeddings, transcription, TTS, image generation, -rerank, VAD and sound detection. An endpoint that is not safe gets no -in-request retry (its failures still trip the target). The PR lists these -endpoints. - -### Response headers - -Every response for a chain carries: - -- `X-LocalAI-Served-Model: ` -- `X-LocalAI-Failover: fallback` or `degraded`, when target 0 did not serve the - request - -Plain HTTP clients can see failover without subscribing to events. - -## Request path (realtime) - -- In `core/http/endpoints/openai/realtime_model.go`, a pipeline stage that names - a chain is resolved **for each call**, not once at session start. This holds - for the full pipeline (`wrappedModel`) and for transcription-only and - sound-detection-only sessions (`transcriptOnlyModel`); both embed the same - stage router. A chain config never reaches the model loader: it has no - backend and would start backend auto-detection. A helper, `mgr.Do(ctx, chain, func(cfg *config.ModelConfig) error)`, - goes through the plan with the same classification as HTTP. -- Streaming stages (`Predict` with a token callback, `TTSStream`, - `TranscribeStream`) wrap the callback. A retry is allowed only until the first - token or audio chunk goes to the client. After that, the turn fails as it does - today, the target is tripped, and the next turn uses the next target. -- `TranscribeLive` is resolved when it opens. A failure in the middle of the - stream ends it in the existing way, and the next utterance opens it again on - the new target. -- The conversation history is in the realtime session on this instance. A - switch of the LLM stage keeps it. -- `Warmup` warms the active target of each chain stage. - -## API and events - -All endpoints use the global auth middleware. `GET` endpoints and the event -stream need standard auth. The pin endpoints are admin only. - -### REST - -`GET /api/failover` returns all chains: - -```json -{"chains":[{"name":"assistant-llm","state":"fallback","active":"gemma-local", - "active_since":"2026-09-26T10:00:00Z","pinned":null, - "targets":[ - {"model":"argus-llm","kind":"remote","warm":false,"state":"recovering", - "consecutive_ok":1,"last_probe":"2026-09-26T10:04:10Z", - "last_error":"503 no healthy nodes"}, - {"model":"gemma-local","kind":"local","warm":true,"state":"healthy"}]}]} -``` - -`GET /api/failover/{chain}` returns one chain. - -`POST /api/failover/{chain}/pin` with `{"target": ""}` forces a target. -`DELETE /api/failover/{chain}/pin` removes the pin. The pin is in memory and a -restart clears it. - -### Server-sent events - -`GET /api/failover/events`: - -- The first event is `snapshot`, with the same payload as `GET /api/failover`. - A new client knows the current state without a race against a separate GET. -- Then `chain.switched` with `{chain, from, to, state, reason, at}`, and - `target.state` with `{target, from, to, reason, error, at}`. -- A keepalive comment every 15 s. - -### Realtime server event - -`localai.model.failover`: - -```json -{"type":"localai.model.failover","chain":"assistant-llm","stage":"llm", - "from":"argus-llm","to":"gemma-local","state":"fallback","reason":"trip"} -``` - -- The server sends it to every session whose pipeline uses the chain when the - chain switches. -- The server also sends it once for each chain stage when the session starts, - with `reason: "initial"` and `from` empty. A client knows at the start whether - it runs on the primary or on a fallback. -- `stage` is one of `llm`, `transcription`, `tts`, `vad`, `sound_detection`. - -### Observability - -- Metrics: `localai_failover_switches_total{chain,from,to,reason}` and - `localai_failover_target_up{target}`. -- Each failed attempt in a request is recorded in the Traces UI, so a request - served by target 2 shows why target 1 was skipped. - -## Capability surfaces - -As required by `.agents/api-endpoints-and-auth.md`: - -- Handlers in `core/http/endpoints/localai/failover.go` with swagger blocks, tag - `failover`. Routes in `core/http/routes/localai.go`. `make swagger`. -- An `instructionDefs` entry for the new tag in - `core/http/endpoints/localai/api_instructions.go`, and the count in - `api_instructions_test.go`. -- `failover.*` fields in the config field metadata registry - (`core/config/meta/registry.go`), in a new `failover` section next to - `alias`, so the generic model editor can show and edit them. -- MCP tools in `pkg/mcp/localaitools/`: `list_failover_chains`, - `pin_failover_target` and `unpin_failover_target`, in the `inproc` and - `httpapi` clients, the skill prompts, and `toolToHTTPRoute` in - `coverage_test.go`. Chains are created and edited through the existing model - config tools. -- A docs page, `docs/content/features/model-failover.md`, linked from - `model-aliases.md`, `openai-realtime.md` and the cloud-proxy docs. - -## Testing - -Ginkgo and Gomega, like the rest of LocalAI. The coverage baseline must not go -down. - -- **State machine**, with a fake clock: trip; recovery after N inference probes; - `min_dwell` hysteresis; `degraded`; pin and unpin; a target shared by two - chains; a `missing` target. -- **`IsRetryable`**: a table of error and status cases. -- **Config validation**: nested chain, `alias` with `failover`, fewer than 2 - targets, missing target, duplicate target, the usecase warning. -- **HTTP integration**: two fake OpenAI-compatible upstreams (`httptest`) behind - `cloud-proxy` target configs. - - Upstream 1 fails. The request is served by upstream 2 with no client error, - `X-LocalAI-Served-Model` is set, and the SSE stream sends `chain.switched`. - - Upstream 1 recovers. Fail-back happens only after N inference probes and - `min_dwell`. - - Upstream 1 fails after the first SSE chunk. There is no retry, the client - gets the error, the target trips, and the next request goes to upstream 2. -- **Handler audit**: one retry test for each endpoint family that proves it is - safe to run again before commit. -- **Realtime**: a pipeline whose LLM stage is a chain of fake backends. The test - checks the `initial` event, the `localai.model.failover` event when the - primary fails, and that the conversation history is kept after the switch. -- **API**: authenticated and unauthenticated access to every endpoint; pin - requires admin. - -## Follow-ups - -1. `localai-proxy` backend: a fork of `cloud-proxy` that forwards every gRPC - method (Predict, Embedding, AudioTranscription, TTS, GenerateImage, Rerank, - VAD, sound detection) to the REST API of an upstream LocalAI. This lets a - remote model serve any realtime pipeline stage. -2. WebUI: a chain editor and a live health view built on - `GET /api/failover/events`. -3. Wingman: use one local LocalAI endpoint with chains for its pipeline stages, - and react to `localai.model.failover` events instead of its own endpoint - supervisor.