A leader whose host died without closing its connection kept the
advisory lock for about two hours of OS keepalive defaults, and no other
frontend could probe. The lock session now sets short TCP keepalives and
tcp_user_timeout, so the server drops it within about 30 seconds.
Shutdown now closes the lock for good, so a tick that runs after it
cannot take the lock back.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A lock taken per tick passed between frontends on almost every tick, so
several frontends probed at once and each change of leader re-sent the
warm set and all state. The leader now holds a dedicated PostgreSQL
session with the advisory lock and keeps it until it shuts down or the
session dies.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A pin arrives from the sync layer once. Dropping it when the chain or
target is unknown here left this frontend routing differently from the
cluster whenever its config lagged or a chain was re-created.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Failover state is per frontend, remote chain targets only cover chat,
and chains have no UI. Design shared pins, health and chain state for
distributed mode, a localai-proxy backend for every API including live
transcription, a chain editor and health view, and a contributor rule
for distributed-aware state.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Pass the API key environment lookup through ApplicationConfig to satisfy
core configuration lint. Keep credential resolution dynamic and exclude
the callback from serialization.
Handle the five close results reported by errcheck.
Assisted-by: Codex:gpt-6 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The spec promised a load-time warning when a chain marks a remote
target warm, where the flag does nothing; the loader now logs it. The
remote-backend test moves into ModelConfig.IsRemoteProxy so the loader
and the failover manager agree on what is remote.
The spec now says what ships: a load blocked by pinned warm targets
proceeds over the limit after eviction retries, without an error that
names them.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
make build-cloud-proxy-backend writes it next to the mock backend,
which is ignored; the cloud-proxy binary showed up as untracked.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Go resends custom headers such as x-api-key when it follows a redirect,
also to another host, so a redirecting upstream could receive the
target's API key elsewhere. The probe client now treats a redirect as
the response, which fails the probe as a non-2xx status.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Transcription-only and sound-detection-only realtime sessions passed a
chain config straight to the model loader. It has no backend, so the
loader fell back to greedy backend auto-detection: slow, and ending in
an unhelpful error. Sound-only sessions are a main use of chains.
The stage routing of the full pipeline moves into a stageRouter that
both realtime model kinds embed. Every stage resolves to the chain's
active target at build time and goes through the failover plan per
call. The session sends failover events for any model with chain
stages, and restarts them when a transcription session.update swaps
the model.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A chain request reached a cloud-proxy target with the client's model,
the chain name, whenever the target set no upstream_model: passthrough
forwards the body's model and translate falls back to it. The upstream
answered 404, which neither retries nor trips, while the liveness
probe, which checks the target's own name, kept passing.
PrepareTarget now sets the upstream model of a remote target to
proxy.upstream_model or the target name, the same name the probe uses.
The request pipeline and realtime chain stages both call it.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A warm target's liveness probe called ModelLoader.Load, which blocked
until the model finished loading (while the warm preload loaded it
too). Tick waited for every probe, so all probing froze, and the probe
then ran HealthCheck on an expired context and tripped the target at
every startup.
The prober now takes a function that returns the running backend
without loading it. A target that is not loaded passes liveness; its
recovery is neither confirmed nor failed and it returns to healthy
after min_dwell, like a cold target. Tick no longer waits for probes:
each probe applies its own result and a target whose probe is running
is skipped.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The request path called HasChains on every request, and with no chains
it scanned the config source each time: the loader's lock plus a copy
and sort of every config, forever, on every installation without
chains. Sync now keeps an atomic flag and HasChains reads only that.
A chain added since the last sync is still served because Plan syncs
on a miss; only in-request retry waits for the next tick (at most 1s).
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A stage that names a chain is resolved on every call, so a switch keeps
the session and its conversation. Clients get localai.model.failover
events at session start and on every switch.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A missing file is not a reliable signal: models download on first use,
some backends need no file, and dotted names like Phi-3.5-mini look like
paths. Marking such a fallback down removed the retry a chain exists
for. Cold targets are now judged only by real requests.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
An admission rejection or a disabled target moves the request to the
next target through Attempt.Skip, which records no failure. A 4xx
response no longer counts as a success. Requests skip body recording
when no chain is configured, and stop it once the model is known not
to be a chain.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The retry wraps SetModelAndConfig, so each attempt binds the request
again from a replayed body. A 5xx of a chain request is held back until
the handler returns, and a streamed response is never retried.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
applyFailoverWarmTargets ran on the manager's single scheduler goroutine
(Sync -> Tick), so a slow or hung PreloadModelByName call froze probing
and fail-back for every chain. Keep the watchdog pin synchronous but run
the preload loop in its own goroutine. Adds a seam (preloadModelByName)
so a unit test can substitute a blocking loader and assert the callback
still returns promptly.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Warm local targets are pinned in the watchdog and preloaded. Switches
and target health are exported as metrics, skipped attempts as traces.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Matches cloud-proxy's resolveAPIKey exactly: os.Getenv + empty check
rather than os.LookupEnv, so a variable that's set but empty errors
instead of probing unauthenticated.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Remote liveness uses /v1/models, which every OpenAI-compatible upstream
serves. Recovery sends one minimal request for the target's usecase.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Idle targets get a liveness probe each interval, recovering targets an
inference probe. Cold local targets are never loaded to be probed.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
recomputeLocked only fired the event on an active-target change or on
entering degraded. When the active target itself recovered while every
target was down, the chain silently left degraded with no event, so
SSE/realtime consumers tracking chain.switched.state got stuck on
"degraded".
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
errcheck flagged two side-effect-only m.Plan calls in tests, and unused
flagged close(), which Task 5's probe scheduler wires in.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Health is tracked per target and the active target per chain. Fail-back
waits for recovery probes and a minimum time on the fallback.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Reject chains whose targets are missing or are chains, at load and on
create or edit, and warn when the targets share no usecase.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A chain is a model config with an ordered list of target models, probe,
trip and recovery settings. Like an alias it has no backend.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
/readyz exists only on LocalAI upstreams, and a cloud-proxy target can
point at any OpenAI-compatible provider.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A model served from a remote upstream has no local fallback when that
upstream is unhealthy, so clients reimplement failover and lose realtime
context when they switch endpoints.
Define a failover block on model configs: an ordered list of targets with
active probes, in-request retry before the response is committed,
fail-back with hysteresis, and switch events over REST, SSE and the
realtime socket.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(gallery): remove invalid Qwen-Image chat entry
The entry sends diffusion weights to llama.cpp as a chat model.
Remove it and document the existing image-generation alternatives.
Assisted-by: Codex:gpt-6
* feat(gallery): add Hemmingway-1 GGUF variants
Add Q4_K_M and Q8_0 builds for llama.cpp with embedded chat templates.
Record the upstream CC BY-NC 4.0 license and installation instructions.
Verify both SHA256 values against Hugging Face LFS metadata and headers.
Assisted-by: Codex:gpt-6
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(gallery): read metadata for system-path backends, enabling variant aliases
Problem:
- ListSystemBackends only read metadata.json for user-managed backends;
the system-path scan (LOCALAI_BACKENDS_SYSTEM_PATH) was a bare
directory walk with Metadata hardcoded nil
- system-packaged backends (distro packages installing several
accelerator builds of one backend) could not declare aliases or meta
indirection at all, while gallery-installed backends could
- surfaced while packaging LocalAI for Gentoo: the packages install
cpu-/rocm-/vulkan-audio-cpp as system backends aliased to audio-cpp,
which the server ignored
Change:
- scan each root separately, clean the system collection against the
user-managed one, merge, then build and resolve — precedence lives in
one explicit step
- alias candidates carry their own metadata: the resolved alias entry
can never pair one installation's executable with another's metadata,
and it reports the chosen candidate's origin (IsSystem)
- deterministic resolution: entries build in sorted name order and
candidates sort by name at the resolution site, independent of scan
order
Precedence (user-managed always wins):
- a user-managed backend hides a same-named system backend entirely
- a user-managed variant takes over its whole alias family: the alias
resolves among user-managed variants only and the system family's
concrete names disappear — family versions move together, and a stale
system variant may not work with newer models, so it must not stay
reachable
- a system variant's alias never hijacks a name that exists as a
user-managed backend
Tests: Ginkgo regressions for system-path aliasing, same-name hiding,
family takeover, and the full metadata permutation matrix of
cross-root name collisions (both directions, with and without
metadata on each side).
Docs: new "Backend Directory Format" section (run.sh, metadata.json,
alias resolution — previously undocumented for user-managed backends
too) and "System-Provided Backends" with the precedence rules.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
* fix(gallery): preserve managed meta backends
A system alias can replace a user-managed meta backend during discovery.
Protect meta entries with the same precedence guard as concrete backends.
Add a regression test and clarify the documented precedence.
Assisted-by: Codex:GPT-6
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
---------
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29
Replace the model-specific SystemOne gRPC approach with a generic Score
RPC extension. The pre-existing Score RPC (previously unused by any
backend) now carries question_type and response_json fields:
- question_type="systemone" routes kev/laya decision-pipeline requests
through the unified vllm_decide C ABI (v29), returning the full
response JSON in response_json.
- question_type empty routes cua-s1-forms candidate scoring through the
same vllm_decide ABI, returning CandidateScore probabilities.
The vllm-cpp backend's Score() method calls vllm_decide and dispatches
by architecture internally. The /v1/systemone HTTP endpoint checks
whether the model's backend supports Score; if so, it forwards the raw
request JSON and returns the backend response as-is. Other backends
fall through to the existing NER-based path.
This mirrors the vllm.cpp C ABI refactor (PR #3301) that replaced
vllm_systemone + vllm_score with a single vllm_decide function. The
purego bindings bump abiVersion from 27 to 29 and resolve vllm_decide
and vllm_decide_free symbols.
Also fixes validModelPath to accept cua-s1-forms.json and
rl_agent_config.json alongside config.json, matching the engine's
model_loader.cpp config-filename ordering.
AI-Assisted: true
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore: ⬆️ Update mudler/vllm.cpp to e28ec46c6 (fix macOS -Werror build)
Bumps vllm.cpp to e28ec46c6 which fixes a -Wnull-conversion error in
qwen3_5.cpp:12483 that broke the macOS Metal CI build under -Werror.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Add five nemo-speech-cpp gallery entries for the diarization
capability introduced by the NeMo-Speech.cpp bump in #12257:
- nemo-speech-cpp-sortformer-diarization-v2: standalone streaming
Sortformer 4-speaker diarization (nvidia/diar_streaming_sortformer_4spk-v2).
Serves /v1/audio/diarization with known_usecases: [diarization].
- nemo-speech-cpp-nemotron-3.5-asr-streaming: standalone multilingual
streaming ASR (nvidia/nemotron-3.5-asr-streaming-0.6b).
- nemo-speech-cpp-nemotron-3.5-asr-streaming-diarized: Nemotron ASR
with the sortformer attached via the diar_model option, giving
per-word speaker tags on /v1/audio/transcriptions.
- nemo-speech-cpp-parakeet-tdt-0.6b-v3: standalone multilingual ASR,
25 languages (nvidia/parakeet-tdt-0.6b-v3).
- nemo-speech-cpp-parakeet-tdt-0.6b-v3-diarized: Parakeet v3 ASR
with the sortformer attached via the diar_model option, giving
per-word speaker tags on /v1/audio/transcriptions.
No backend code changes: the sortformer to familyDiarization mapping,
the diar_model option, and the MethodDiarize gRPC method already exist.
All five entries verified end-to-end against a running LocalAI instance.
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>