Commit Graph
1248 Commits
Author SHA1 Message Date
Ettore Di Giacinto d8571a8ee9 fix(failover): spill 429 admission rejections to the next target
#12113 changed admission control to reject with 429 instead of 503.
failoverWriter only held back responses with status >= 500, so a 429
rejection reached the client and the chain never spilled to its next
target.

Hold 429 as well. An admission rejection is still flagged and spills
without tripping the target. Any other 429 is not retryable, so it is
released to the client unchanged.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-28 15:34:29 +00:00
Ettore Di Giacinto 9c156656bd Merge PR #12285: feat(failover): serve a model name from a chain of local and remote targets
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 15:32:30 +00:00
mudler-agentandEttore Di Giacinto 590512d9eb feat(distributed): report worker version and show models in node inspector (#12328)
* ci: bump Hugo from 0.146.3 to 0.166.0

The hugo-theme-relearn submodule was bumped to 9.1.x in #12096,
which requires Hugo >= 0.165.0. The pinned 0.146.3 broke the docs
site build with a template error in alias.html that could not
evaluate the Locale field on langs.Language.

Bump HUGO_VERSION to 0.166.0 (latest stable) to satisfy the
theme minimum and resolve the alias.html template error.

Assisted-by: nib:claude-sonnet-4.5 [bash] [read] [edit]

* feat(distributed): report worker version and show models in node inspector

Workers now send their LocalAI build version and git commit at
registration. The controller stores them on BackendNode and exposes
them through the existing node list/detail API responses.

The node inspector side pane now fetches and renders the list of
loaded models (name, state, in-flight) instead of showing only a
count, matching what the node detail page already displays.

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 13:13:31 +02:00
mudler-agentandEttore Di Giacinto 50c284fdcc fix: make the remaining VerifyPath checks effective (#12326)
utils.VerifyPath joins its argument onto the base path, so a path that
the caller already joined always passes. Several callers gave it joined
paths, and their checks could not fail:

- modeladmin (config view, patch, edit, pin and state): the config file
  path from the loader. A config loaded from outside the models
  directory (--models-config-file) could be pinned, and the pin wrote
  the outside file. The patch and state paths stopped later, in the
  mutation snapshot, with a different error.
- core/backend/tts.go: the model path joined onto the models path.
- The trellis2cpp and stablediffusion-ggml backends: option paths
  (*_path) joined onto the model path. A "../" value outside the model
  directory was accepted.

Add utils.VerifyResolvedPath for a full path. modeladmin and tts use
it. The backends now check the relative option value before they join
it. A rename in modeladmin checks the new relative name.

For models from a config file outside the models directory, the admin
API and web UI now return ErrPathNotTrusted for view, edit, pin, and
enable or disable. The docs describe this.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 10:04:22 +02:00
mudler-agentandEttore Di Giacinto bc01ef2350 fix(gallery): keep model deletion inside the models directory (#12324)
listModelFiles gave utils.VerifyPath paths that it had already joined
onto the models directory. VerifyPath joins its argument onto the base
again, so an absolute path always passes and none of the four checks
could fail. Model deletion then removed files outside the models
directory:

- A model name such as "../outside/victim" removed
  outside/victim.yaml. The in-process MCP delete_model tool passes the
  name from the tool call without a check.
- A gallery file that lists a files: entry with "../" removed that
  file.

listModelFiles now gives VerifyPath the relative names.

InTrustedRoot also looped forever when a relative path was outside a
relative root. filepath.Dir stops at "." for a relative path, and the
loop waited for "/". The loop now stops when Dir returns its input.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 09:28:04 +02:00
mudler-agentandEttore Di Giacinto 97ad8f1d70 fix(distributed): stage the files a model install declares (#12309)
* fix(distributed): stage every shard of a split GGUF

A split GGUF is configured by its first shard only. llama.cpp opens the
other "-0000N-of-0000M.gguf" files from the same directory by name. The
router staged only the configured path, so the worker received shard 1
and the load failed with "failed to load GGUF split".

The router now stages the remaining shards next to the first one. A
missing shard fails the load and names the file. The file count for
progress and the payload size also include all shards. The payload size
feeds the load deadline and the disk-headroom check. For a 111 GB model
whose first shard is 10 MB, both were sized for less than 1 GB.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): stage the files a model install declares

Replace the split GGUF file name matching with the model's own file
list. A gallery install or an import records every file of the model in
._gallery_<name>.yaml (files:), and a config can list more under
download_files:. The router now stages all of these files, not only the
files that the config's path fields name. This includes the other
shards of a split GGUF, which llama.cpp opens by name.

The application gives the router a resolver that reads the two lists.
The resolver looks up the files by model name when it stages them, so a
replica that the reconciler loads from saved load options gets the same
files. backend.proto does not change.

A declared file that is missing on the frontend is skipped with a
warning. The load deadline and the disk headroom check include the
declared files.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 08:34:56 +02:00
leilei3167 0e4347037f fix: point docker-compose default at gallery phi-2-chat (#11987)
The quickstart compose file still requested phi-2, which is no longer in
the gallery. Use phi-2-chat instead and fix model preload error wrapping
so discover/install failures report the real error instead of %!w(<nil>).
Keep earlier model failures when discovery fails for another model.
Document the Compose gallery default.

Fixes #11974

Signed-off-by: lei_lei <imleilei123@gmail.com>
2026-09-28 04:49:32 +02:00
Ettore Di Giacinto 0565fc06af Merge PR #12302: chore(deps): bump LocalAGI to 7e0947d (no-RAG-DB crash fix, tool filters, per-collection models)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 01:48:00 +00:00
f82efdb43b fix(models): fallback to application config default context size in /v1/models/capabilities (#12202) (#12216)
* fix(models): fallback to application config default context size (#12202)

Honor appConfig.ContextSize in /v1/models/capabilities when model context_size is unset.

* docs(models): explain context size fallback

Describe the application default used by capability discovery and
preserve the distinction between total context and per-request limits.

Assisted-by: Codex:GPT-6

* fix(models): apply the default context size only when context_size is unset

The request path applies the application default context size only
when a model leaves context_size unset. An explicit 0 or -1 falls
through to the backend fallback. The capabilities endpoint now does
the same, so it reports the value the backend uses.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 00:52:56 +02:00
Tai Anandlocalai-org-maint-bot eec532704c fix(functions): honor function_arguments_key when building the tool grammar (#11677)
* fix(functions): honor function_arguments_key when building the tool grammar

All four call sites of `Functions.ToJSONStructure(name, args string)` pass
`FunctionsConfig.FunctionNameKey` as *both* arguments, so
`FunctionArgumentsKey` never reaches the grammar generator.

`ToJSONStructure` writes the two properties into the same map:

    property[nameKey] = FunctionName{Const: function.Name}
    property[argsKey] = Argument{...}

When `nameKey == argsKey` the second assignment overwrites the first, so a
model configured with `function_name_key` gets a grammar carrying only the
arguments object -- the `{"const": "<function name>"}` constraint is gone and
the grammar can no longer express which function was called.

With `function_name_key: function`, the generated property set collapses from

    {"function": {"const": "get_weather"}, "arguments": {...}}

to

    {"function": {"type": "object", "properties": {...}}}

Setting only `function_arguments_key` is equally broken in the other
direction: the grammar keeps emitting `arguments` while `ParseFunctionCall`
(pkg/functions/parse.go) looks up the configured key, so the parsed call comes
back with its arguments empty.

The default configuration is unaffected -- with both keys empty
`ToJSONStructure` falls back to `name`/`arguments` for both parameters, which
is why this went unnoticed.

The existing `ToJSONStructure()` unit test already calls the helper with two
distinct keys, so only the call sites were wrong. Extend that test with a case
that keeps both custom keys distinct and asserts the two properties survive.

Signed-off-by: Anai-Guo <antai12232931@outlook.com>

* test(functions): cover configured grammar keys

Route grammar construction through FunctionsConfig so the regression test
covers the key wiring used by every endpoint.

Assisted-by: Codex:gpt-5

* chore: empty commit to trigger workflow approval

Signed-off-by: Tai An <antai12232931@outlook.com>

---------

Signed-off-by: Anai-Guo <antai12232931@outlook.com>
Signed-off-by: Tai An <antai12232931@outlook.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-28 00:52:51 +02:00
Stefan Walcz 950c271710 [router] fix: re-seed the knn corpus index when the vector store comes back empty (#12267)
fix(router): re-seed the knn corpus index when the vector store comes back empty

The corpus manager records a store as synced by file fingerprint and
embedding fingerprint. The local-store backend behind it is an in-memory
gRPC process the model loader may evict (active-backend cap, memory
pressure) or the idle watchdog may kill, and relaunch on the next
request — empty. The file is unchanged, so EnsureLoaded returned early
and the router went blind: every probe fell back with similarity 0 while
corpus/stats kept reporting the full count.

Measured on a production router (LOCALAI_MAX_ACTIVE_BACKENDS=6, four
resident models + two router stores): loading any further backend
evicted a store, and the idle watchdog killed both after 15 minutes;
/stores/find returned 0 hits against a 100-line corpus file whose stored
vectors matched fresh embeddings with cosine 1.000.

Two parts, because the knn classifier is built once and cached
(GetOrBuildClassifier), so the sync at build time is otherwise the only
one for the process lifetime:

- corpus.Manager remembers one vector it inserted (probe) and, on the
  synced path, asks the live index for it. A miss means the index was
  relaunched — fall through and re-seed from the file (no re-embedding).
- The router middleware wraps the knn classifier's store so every
  lookup runs EnsureLoaded first; the loader gets the raw store, so its
  probe never re-enters the wrapper. A sync error fails the lookup
  closed, like the build-time load.

Specs: corpus package (relaunched empty store is re-seeded under an
unchanged file), middleware (relaunched index behind the cached
classifier is re-seeded instead of falling back; the spec is red without
the wrapper). The test fake now answers Search for inserted vectors.
Folds in the maintainer's follow-up (router-corpus-reseed-after-store-relaunch): reviewed and accepted.


Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-09-28 00:52:47 +02:00
52a6d62bbc fix: return correct HTTP status codes for saturation and no-nodes-available (#12113)
* feat: return 429 when backends are saturated

When backends are at capacity (per-model max_concurrent or the
process-wide --max-concurrent-backend-requests ceiling), the response
was 503. The OpenAI SDK, litellm, and most agent harnesses key on 429
for rate-limit backoff and treat 503 as a hard error.

Both saturation paths now return 429 with the existing Retry-After
header and type: "rate_limit_error" in the JSON body. The per-model
admission middleware keeps admission_rejected as the code field so
existing alerts that match on it still fire.

Non-saturation 503s are unchanged: model cold-loading (with progress
body), model-load failure cooldown, PII detector fail-closed, and
classifier unavailable. These mean "not ready" rather than "busy".

Assisted-by: AGENT:regolo/glm5.2 [TOOL]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix: return 503 when scheduler has no available nodes

When the scheduler cannot find any healthy node to serve a model —
all nodes are full and eviction cannot free a slot, or a node_selector
excludes every candidate — the error fell through to 500. A 500 tells
clients something is broken when the condition is transient and
retryable.

The router now wraps these errors with a new ErrNoAvailableNodes
sentinel. The HTTP error handler maps it to 503 via applyNoAvailableNodes,
following the same pattern as applyBackendAdmission (429). Unrelated
scheduler errors (DB timeouts, registry lookups) still return 500.

Three return sites are wrapped:
- resolveSelectorCandidates: selector matches zero healthy nodes
- scheduleNewModel eviction-busy: all models have in-flight requests
- scheduleNewModel eviction-failed: eviction itself errored

The existing scheduleAndLoad wrapper ("no available nodes: %w") preserves
the sentinel through the chain via errors.Is, as does ModelRouterAdapter.

Assisted-by: AGENT:regolo/glm5.2 [TOOL]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(http): use Ginkgo for admission tests

Replace forbidden testing.T calls with Ginkgo and Gomega so the lint
check accepts the admission handler tests.

Assisted-by: Codex:GPT-6 forbidigo

* fix(middleware): show the recorded status for admission rejections

The admission audit row now records 429, but the Middleware page still
printed a hard-coded 503. Read the status from the event, and update
the two package comments that still said 503.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-28 00:11:51 +02:00
Tai AnandEttore Di Giacinto 5b10d36b5f fix(quantization): pin the producing backend on imported quantized models (#11875) (#11879)
* fix(quantization): pin the producing backend on imported quantized models (#11875)

ImportModel hands the copied GGUF to importers.ImportLocalPath, which detects
the file format and defaults every GGUF to `backend: llama-cpp`. For a model
this service just produced with a backend stock llama.cpp cannot read, the
generated config names an engine that cannot load the file, and the import
silently registers an unloadable model. Correcting `backend:` by hand makes the
same file work.

The job record already carries the backend that served StartQuantization, so
carry it into the config instead of keeping the detected default. The gallery
publishes a quantizer as a release channel of the engine that runs its output
("llama-cpp-quantization" is llama.cpp's quantizer, whose GGUF is served by
"llama-cpp"), so the channel suffix is stripped to get the serving backend.
A backend that both quantizes and serves ("rocmfp4") carries no suffix and
passes through unchanged, as do pinned hardware variants ("rocm-rocmfp4"),
which are valid values for a config's backend field. An empty job backend
leaves the detected default in place.

Also replace the importer's generic "Fine-tuned model (GGUF)" description for
this path: the model was quantized, not fine-tuned, and the job knows the type.

Signed-off-by: Tai An <antai12232931@outlook.com>

* style: restore trailing newline in service.go for gofmt

Signed-off-by: Anai Guo <antai12232931@outlook.com>

* style(quantization): restore trailing newline in service.go

gofmt requires the file to end with a newline; the previous style commit
did not actually add it.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: Tai An <antai12232931@outlook.com>
Signed-off-by: Anai Guo <antai12232931@outlook.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 00:11:46 +02:00
leilei3167 afebecc63d fix(ollama): report on-disk size for /api/tags and /api/ps (#11989)
Hardcoding size/size_vram as 0 made Ollama clients treat loaded models as
free. Prefer ModelFileName+ModelPath Stat when available, and omit size_vram
(and size) when the value is unknown instead of emitting literal zeros.
Resolve each listed model by its stored ID so tagged variants use their
own weights.

Fixes #11969

Signed-off-by: lei_lei <imleilei123@gmail.com>
2026-09-28 00:11:41 +02:00
Ettore Di Giacinto 84e2fc5eac feat(agents): support tool lists and required tool in distributed mode
LocalAGI 7e0947d added allowed_tools/excluded_tools and the
required_tool_before_finish gate. Single-node agents get them through
LocalAGI's runtime, but the distributed executor drives cogito directly
and its static config meta did not list the fields, so the agent form
hid them and the worker ignored them.

The distributed config now parses the tool lists from a JSON array or a
comma/newline separated string, and the meta entries match LocalAGI's.
The executor filters the knowledge base, skill and MCP tools (MCP via
cogito.WithMCPToolFilter) before the model sees them, and re-prompts the
model when it answers before the required tool returned "ok": true, up
to the configured number of reminders.

LocalAGI keeps its filter and gate helpers unexported, so a minimal copy
lives in core/services/agents/toolpolicy.go. A spec compares the meta
entries with LocalAGI's to catch drift.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:49:12 +00:00
Ettore Di Giacinto 00bd5ee121 fix(failover): skip a target that has been edited into a chain
A chain is checked for nested chains when it is saved, but not when
one of its targets is later edited into a chain. Requests then served
the inner chain's config as the target, which has no backend and
triggers backend auto-detection.

Mark such a target missing so no plan picks it, and skip it without a
trip in the HTTP path when it is pinned.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:19:09 +00:00
Ettore Di Giacinto 8474296cb1 fix(failover): keep the aliased model of a warm target loaded
The warm list held target names as the chain lists them. For an alias
target that is the alias, but the preloader loads the alias stub (no
backend, no model) and the eviction guard compares against loaded
model names, which never include an alias. A warm alias target was
neither preloaded nor protected from eviction.

Report the model that serves each warm target instead.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:18:30 +00:00
Ettore Di Giacinto 1b03600437 fix(failover): drop the echo of an own target publish
A manager snapshots a target publish when it queues it and applies
every echo the sync layer sends back. With one recovery probe, a
success moves a target from down to recovering to healthy under one
lock and queues two publishes. The echo of "recovering" then arrived
after the target was healthy, rolled it back, switched the chain away
with reason trip and restarted the dwell timer.

Tag each target snapshot with the publishing manager and ignore own
echoes. The manager already holds that state or a newer one.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:17:52 +00:00
Ettore Di Giacinto f20cc16033 fix(openai): reject an unknown audio response_format as a 400 before the backend runs
Transcription and diarization checked response_format only after the
backend had run, and returned a plain error for an unknown value.
Failover counts a plain error as a target failure, so one request with
a bad response_format ran the backend on every target of a chain and
tripped all of them.

Check the format before the backend runs and answer 400.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:17:06 +00:00
Ettore Di Giacinto dae9a431e8 Merge remote-tracking branch 'origin/master' into feat/failover-chains
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 19:57:52 +00:00
Ettore Di Giacinto 7804e4591c fix(syncstate): do not keep a value in memory when its Store write fails
SyncedMap.Set and Delete changed memory before they wrote the Store.
When the write failed they returned the error, but the unpersisted
value stayed in memory. Callers that re-read the map then applied it
again. The failover manager rolled back a failed pin, but its periodic
pin re-sync read the pin back from the map and re-applied it on that
frontend until the database came back.

Write the Store first and change memory only after it succeeds. A
failed Set or Delete now leaves memory and peers as they were, so the
map always matches what the Store holds. This also stops a failed
create of a fine-tune, quantization or agent task from leaving an
orphan entry that the API had reported as failed.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 19:57:13 +00:00
localai-org-maint-botandlocalai-org-maint-bot 490b952d06 feat(gallery): publish signed OCI fallbacks (#12182)
* feat(gallery): publish signed OCI fallbacks

Publish both official gallery indexes with their local base configs so
an outage of the HTTP and GitHub sources can fall back to Quay.

Keep artifact signing policies separate from backend image policies,
and expose each moving gallery tag only after its digest is signed.

Assisted-by: Codex:gpt-6

* fix(gallery): confine packaged files to selected roots

Use directory-scoped file access to reject symlink escapes during gallery packaging. Create private bundle files for the publishing runner.

Assisted-by: Codex:GPT-6

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:38 +02:00
localai-org-maint-botandlocalai-org-maint-bot f154bd990a feat(system): report per-model DRM VRAM (#12026)
* feat(system): report per-model DRM VRAM

Expose optional resident device memory for local backend process trees.
Deduplicate DRM clients and omit unsupported or incomplete readings.
Document accounting limits and preserve a measured zero in JSON.

Closes #11970.

Assisted-by: Codex:gpt-6

* fix(system): document trusted procfs reads

Scope G304 annotations to paths built from the fixed procfs root,
integer process IDs, and kernel directory entries. These reads accept
no user-controlled path components.

Assisted-by: Codex:GPT-6 gosec

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:34 +02:00
localai-org-maint-botandlocalai-org-maint-bot 5794495a37 fix(responses): preserve streamed output items (#12048)
Keep each message and reasoning item at its announced output index.
Include the answer in completed responses with reasoning or fallback
function calls, and retain reasoning supplied through backend deltas.

Add regression coverage for stream indices, final output, plain text,
and automatic tool parsing.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:29 +02:00
localai-org-maint-botandlocalai-org-maint-bot 0e52bb657e fix(responses): wait for complete JSON tool calls (#12001)
Partial JSON parsing heals a name-only chunk into a tool call. The
stream emits that call with empty arguments and skips later chunks.

Require complete JSON before emitting terminal tool-call events.
Preserve complete calls before an unfinished trailing call, and count
only actual tool calls. Add split-chunk regression tests and docs.

Refs #11635. The non-streaming report remains unconfirmed.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:24 +02:00
localai-org-maint-botandlocalai-org-maint-bot 4e94c914c9 fix(swagger): describe backend metadata as an object (#12178)
Swag cannot resolve json.RawMessage in OpenAIResponse and aborts the daily
schema generation. Set its Swagger type without changing JSON encoding,
and regenerate the checked-in specifications.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:06:54 +02:00
Ettore Di Giacinto 27ceda46b0 refactor: name the realtime pipeline stages with constants
Review asked for constants instead of the stage literals ("vad",
"transcription", "llm", "tts", "sound_detection") passed to resolveStage,
stageCall and isChainStage, so the uses can be cross-checked. Add
PipelineStage* constants next to the Pipeline type in core/config: the
names match its yaml keys, and core/backend (preload roles) and the openai
realtime endpoint both need them.

Use them in realtime_model.go (stage routing and preload roles),
realtime.go (the voice_recognition preload role) and core/backend
preload.go. model_failover events take their stage from the stageChains
keys, so they now carry the constants too. The failover tests use the
constants for inputs and keep literal wire values in their event
assertions.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 8db6c15fe0 refactor: name the proxy backends with constants
Review asked for constants instead of the "cloud-proxy" and
"localai-proxy" literals. Add CloudProxyBackend and LocalAIProxyBackend next
to the other backend-name constants in pkg/model (WhisperBackend,
TransformersBackend, ...), which core/config already imports, and use them
in every production check: the proxy options builder, IsRemoteProxy,
IsCloudProxyBackendPassthrough, the PII defaults, the localai-proxy backend
hook and loader warning, and the PII middleware metadata in the routes and
the in-process MCP client.

The proxy options builder now calls IsRemoteProxy() instead of repeating
the two-backend check, so the set of proxy backends is defined once.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 14c692098e fix(openai): answer 400 for a missing or malformed audio upload
TranscriptEndpoint, DiarizationEndpoint and SoundClassificationEndpoint
returned the raw c.FormFile("file") error. Echo turns a non-HTTPError into
a 500, so a request with no multipart boundary or no file field looked like
a server fault. This became visible in tests/e2e once the suite registered
a transcription-capable model (lp-transcription) at runtime: the
default-model middleware then resolves a model, the request reaches the
handler, and "should return mocked transcription" got a 500 in some spec
orders (seed 1790493709).

Read the upload through a small uploadedFile helper that maps any
FormFile failure to 400 with the field name and parser reason. Server-side
failures after that (temp dir, file create, copy) stay 500. The image and
LocalAI upload endpoints already return 400 here.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 55f7d5bfdf fix(e2e): wire proxy api_key_env lookup into the e2e in-process app
19f6a18b8 moved credential-env resolution behind an explicit lookup
(ApplicationConfig.ProxyAPIKeyEnvLookup / config.WithProxyAPIKeyEnvLookup),
wired only at the CLI boundary (core/cli/run.go). The e2e suite builds its
Application in-process without that option, so the failover prober could
never resolve a remote target's api_key_env, remote liveness never passed,
and "fails over ... and fails back" hung waiting for chain-remote to
recover. Pass config.WithProxyAPIKeyEnvLookup(os.Getenv) there too, same as
the CLI. worker/federated commands don't serve proxy/failover configs and
tests/e2e-ui never sets api_key_env, so neither needs the lookup.

Also make the misconfiguration itself easier to diagnose: the prober now
logs a one-time xlog.Warn per api_key_env when a remote target sets it but
no lookup is configured, instead of only surfacing it as a per-probe
"is unset" error indistinguishable from a genuinely empty env var.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 4a7c96f533 fix(advisorylock): bound HeldLock.TryAcquire on a hung database
The failover prober gate calls TryAcquire with the application context,
which never ends. A database that stopped answering blocked the
scheduler goroutine for good, with the lock's mutex held. Bound the
session open and the lock query with the same timeout as Verify.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 5ecaf8f792 fix(failover): count a chain switch once per cluster
Every frontend recorded localai_failover_switches_total for the switches
it adopted from the leader, so a cluster of N frontends counted each
switch N times. Only the leader, which decides the switch, records it.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 4f801dfc99 fix(failover): roll back a pin that could not be shared
Pin and Unpin apply the change locally first for read-your-writes. When
the shared write then failed, the local pin stayed, so this frontend
served a target the others did not. Restore the previous pin on error,
unless a newer change arrived meanwhile.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto ed9951cf7c fix(failover): re-sync pins on every frontend after a missed delta
A NATS reconnect re-hydrates the pins map from the DB without OnApply, so
a frontend that missed an unpin kept serving the old pin and flip-flopped
with the leader's republish. Every frontend now reconciles the manager's
pins with the shared set every ten ticks and after a reconnect, and the
pins map re-reads the DB every 30 s to repair a delta dropped without a
reconnect.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto e61bd2d78e fix(failover): do not trip targets on a gRPC message size limit
ResourceExhausted is retried and trips the target, which is right for a
rate limit or an out-of-memory backend. A payload over the gRPC message
cap is ResourceExhausted too, but every target rejects it the same way,
so it tripped the whole chain. Classify it as a request error.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 5ed67be0bf fix(failover): skip a target that answers HTTP 501 instead of failing the chain
The non-OpenAI endpoints (depth, detection, face_*, voice_*, images,
video, 3d) map a backend's gRPC Unimplemented to an echo 501 without the
gRPC status. The retry loop did not see a capability gap, and IsRetryable
is false for 501, so the client got 501 and the next target was never
tried. Treat a returned or written 501 as a capability gap.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 452a3a0cbe fix(docs): correct warm-toggle and Reconcile-without-Store claims
The failover chain editor leaves the warm toggle enabled on every row
(the model list has no backend field to gate on) and relies on the
server warning instead, so the docs describing it as disabled for
remote targets were wrong. Separately, syncstate's hydrate() returns
early with no Store or Loader, so a Reconcile tick is a no-op rather
than one that empties the map — correct that claim everywhere it was
repeated (contributor guide, distsync comment, design spec).

No behavior change.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 7b11af7f7d feat(ui): list failover chains and badge chain models
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto bdd4a73d05 feat(ui): edit failover chain targets with a dedicated field
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto d65b4d3e6a feat(ui): show live failover chain health in the model editor
Editing a failover chain now shows its state, the target serving it,
and a per-target health table under the editor header. The strip reads
GET /api/failover and follows /api/failover/events, with a 15 s re-list
to cover SSE reconnect gaps. Admins can pin a target or unpin the chain
after a confirmation.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 882f51fd62 fix(failover): trip a rate-limited or exhausted target instead of skipping it
Treating ResourceExhausted as a capability gap skipped the target
without counting a failure, so a target that stays rate limited or out
of memory kept its traffic. It is now an ordinary retryable failure:
the request moves to the next target and the exhausted one trips.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 6df1767133 fix(localai-proxy): keep rerank, stream errors and chat intact through the proxy
Rerank no longer sends top_n 0, which the upstream rejects. A
mid-stream upstream error frame now fails the call instead of ending
it as a short success. Temperature 0 is forwarded. An upstream 429
becomes ResourceExhausted, which failover skips like Unimplemented.

A localai-proxy config sends its own name upstream when upstream_model
is unset, and a chat proxy defaults to the tokenizer template so chat
reaches the upstream as messages.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 509cca35e0 feat(grpc): serve Rerank from Go backends and skip Unimplemented targets
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto e59fb854ee fix(failover): free the leader lock soon after the leader's host dies
A leader whose host died without closing its connection kept the
advisory lock for about two hours of OS keepalive defaults, and no other
frontend could probe. The lock session now sets short TCP keepalives and
tcp_user_timeout, so the server drops it within about 30 seconds.

Shutdown now closes the lock for good, so a tick that runs after it
cannot take the lock back.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto b5792e4d17 fix(failover): keep the probe leader until its session ends
A lock taken per tick passed between frontends on almost every tick, so
several frontends probed at once and each change of leader re-sent the
warm set and all state. The leader now holds a dedicated PostgreSQL
session with the advisory lock and keeps it until it shuts down or the
session dies.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto d30c33a074 feat(failover): run one prober per cluster and pin warm targets on workers
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto 7b88674fc5 feat(failover): sync pins, target health and chain state over NATS
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto 5f98dfcf95 fix(failover): keep shared pins for chains this frontend has not loaded yet
A pin arrives from the sync layer once. Dropping it when the chain or
target is unknown here left this frontend routing differently from the
cluster whenever its config lagged or a chain was re-created.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto 65c22c53bc feat(failover): share state through a sync hook and gate probes on a leader
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
localai-org-maint-bot e62854c340 fix(failover): pass credential lookup from CLI
Pass the API key environment lookup through ApplicationConfig to satisfy
core configuration lint. Keep credential resolution dynamic and exclude
the callback from serialization.

Handle the five close results reported by errcheck.

Assisted-by: Codex:gpt-6 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00