Commit Graph
317 Commits
Author SHA1 Message Date
Ettore Di Giacinto 9c156656bd Merge PR #12285: feat(failover): serve a model name from a chain of local and remote targets
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 15:32:30 +00:00
Stefan Walcz bc1d9924de fix(llama-cpp): keep llama.cpp's default cache_ram instead of no limit (#12297)
grpc-server.cpp forced params.cache_ram_mib = -1 (no limit) since #7009.
Since v4.3 kv_unified and cache_idle_slots are on by default, so every
distinct prompt now leaves its slot KV state in the host-side prompt
cache, and without a limit the backend grows until the host runs out of
memory.

Measured on gfx1151 (Strix Halo, 128 GB), llama-cpp backend, one request
at a time, 100 distinct prompts of ~2000 characters plus a fixed system
prompt, max_tokens 200:

  model                        cache_ram     RSS loaded -> after 100
  gemma-4-26B-A4B (q8_0 KV)    -1 (default)  1.4 GB -> 25.3 GB
  Qwen3.6-35B-A3B (q8_0 KV)    -1 (default)  1.1 GB -> 19.5 GB
  gemma-4-26B-A4B              -1, same prompt 100x  1.4 GB -> 1.6 GB
  gemma-4-26B-A4B              4096          1.4 GB -> 5.4 GB (flat from
                                             request 20 on, same latency)
  Qwen3.6-35B-A3B              4096          1.1 GB -> 5.1 GB (flat)

The memory is not released when idle. In production a document
classification pass pushed the daily chat model to 34 GB RSS overnight.

Drop the override so llama.cpp's own default (8192 MiB) applies; the
cache_ram option still accepts -1 for users who want no limit. Update
both docs tables (the option reference and the prompt-cache table) and
note what -1 does.


Assisted-by: Claude:claude-opus-5-5
Assisted-by: Codex:GPT-6

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-09-28 11:37:24 +02:00
mudler-agentandEttore Di Giacinto 97ad8f1d70 fix(distributed): stage the files a model install declares (#12309)
* fix(distributed): stage every shard of a split GGUF

A split GGUF is configured by its first shard only. llama.cpp opens the
other "-0000N-of-0000M.gguf" files from the same directory by name. The
router staged only the configured path, so the worker received shard 1
and the load failed with "failed to load GGUF split".

The router now stages the remaining shards next to the first one. A
missing shard fails the load and names the file. The file count for
progress and the payload size also include all shards. The payload size
feeds the load deadline and the disk-headroom check. For a 111 GB model
whose first shard is 10 MB, both were sized for less than 1 GB.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): stage the files a model install declares

Replace the split GGUF file name matching with the model's own file
list. A gallery install or an import records every file of the model in
._gallery_<name>.yaml (files:), and a config can list more under
download_files:. The router now stages all of these files, not only the
files that the config's path fields name. This includes the other
shards of a split GGUF, which llama.cpp opens by name.

The application gives the router a resolver that reads the two lists.
The resolver looks up the files by model name when it stages them, so a
replica that the reconciler loads from saved load options gets the same
files. backend.proto does not change.

A declared file that is missing on the frontend is skipped with a
warning. The load deadline and the disk headroom check include the
declared files.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 08:34:56 +02:00
Pratik Gandhi ea10f26f57 docs: fix 11 dead links in the documentation (#12320)
- compatibility table: voxtral.c, OuteTTS and VoxCPM repos live under
  antirez, edwko and OpenBMB
- distributed inferencing: llama.cpp RPC README moved to tools/rpc under ggml-org
- customize-model: the phi-2 example config moved to LocalAI-examples, and
  embedded/models was replaced by the gallery
- model-gallery: malformed URL; link the gallery index
- GPU acceleration: ROCm install guide moved
- integrations: Wave Terminal docs page moved to ai-presets

Assisted-by: Claude:claude-fable-5-1

Signed-off-by: Pratik Gandhi <travpreneur@gmail.com>
2026-09-28 08:33:15 +02:00
61f4f67b75 sglang backend: pass through thinking_budget + require_reasoning (#12193)
* sglang backend: pass through thinking_budget + require_reasoning

sglang's raw Engine.async_generate() API (which this backend calls
directly, bypassing sglang's own OpenAI server) supports a precise,
tokenizer-derived reasoning-length budget via
sampling_params["custom_params"]["thinking_budget"] plus
require_reasoning=True, gated behind --enable-strict-thinking. Neither
was reachable through LocalAI: this backend built sampling_params only
from a fixed field mapping (temperature, top_p, ...) with no custom_params
key, and never passed require_reasoning to async_generate at all.

- LoadModel now reads a model-level "thinking_budget" option (same
  mechanism as the existing tool_parser/reasoning_parser options), and
  _build_sampling_params adds it as custom_params.thinking_budget on
  every request when configured.
- _new_reasoning_parser already derives, from the rendered prompt, whether
  the model's chat template pre-opened a reasoning block (Qwen3-style
  templates append <think> to the prompt instead of letting the model
  emit it) -- the same signal sglang's own OpenAI server computes from
  per-template config to decide require_reasoning. This backend has no
  template manager, so it now returns that signal too and _predict
  forwards it to async_generate(require_reasoning=...).

Verified against production (NVFP4, sm_121, Qwen3.6-35B-A3B) via a raw
Engine.async_generate() call bypassing this backend: 301 reasoning
tokens against a 300-token budget, clean completion, ~27s. Not yet
verified through this backend's own gRPC path end-to-end (no local
CUDA/sglang environment available here) -- existing + new unit tests in
test.py cover the pure-Python merge/passthrough logic only.

Scope note: require_reasoning is derived only from the existing
prompt-suffix heuristic, not sglang's full per-template
_get_reasoning_from_request decision tree (minimax-m3/hunyuan special
cases etc.) -- this backend has no template manager to evaluate that
tree against, and the prompt-suffix check is the one heuristic already
validated in this file (test_reasoning_parser_forced_when_template_prefills_think_tag).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: honour a model-level reasoning_default

A model YAML can already carry "parameters: reasoning_effort:", but that
value only reaches this backend when a *caller* sets it per request (the Go
side turns it into Metadata["enable_thinking"]). As a model-level default it
is silently dropped: a config reading "reasoning_effort: none" still produces
full reasoning on every request, so the config says one thing and the model
does another.

That gap is expensive in practice. On a self-hosted Qwen3.6-35B-A3B the
reasoning phase consumed the entire max_tokens budget before any content was
produced - 90% of code completions came back empty at max_tokens=768, and the
server log filled with "backend produced only reasoning, retrying". The
config looked like reasoning was off the whole time.

This adds "reasoning_default:off" (or ":on") on the same model-level
options: mechanism as thinking_budget. A per-request value always wins; the
default only fills in when the request is silent.

Measured on the stack above (sglang 0.5.20, NVFP4, GB10/sm_121) after
applying it:
  default (nothing set)          -> 0 chars reasoning, 27 tokens
  "reasoning_effort": "none"     -> 0 chars reasoning, 27 tokens
  metadata enable_thinking=true  -> capped at the 512-token thinking_budget,
                                    541 tokens total, finish_reason stop

Tests: three cases added to backend/python/sglang/test.py covering the
default, per-request override in both directions, and the unconfigured case
(which must leave the template untouched).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: validate thinking_budget instead of crashing LoadModel

Addresses the review on this PR:

- `int(thinking_budget)` raised on values like "5000.0" or "abc" and took
  LoadModel down. The option is now parsed by _parse_thinking_budget():
  integral numbers in any spelling are accepted, anything else is ignored
  with a warning on stderr.
- Zero and negative budgets are ignored with a warning instead of being
  passed to sglang, where they have no defined meaning. Turning reasoning
  off is what reasoning_default:off is for.
- A load-time warning when thinking_budget is set but enable_strict_thinking
  is not in engine_args, since sglang then ignores the budget silently.
- Tests for integral spellings, unset, zero, negative, non-integer and the
  strict-thinking warning.

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): explain reasoning options

Document the reasoning budget, strict-thinking requirement, and
precedence of request metadata over the model-level default.

Also note that the budget has to stay well below max_tokens (otherwise
it never triggers and the reply can end up empty), and that
POST /models/reload or a backend-only restart does not pick up changed
options; LocalAI itself has to be restarted.

Assisted-by: Codex:GPT-6
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): clarify configuration reloads

Distinguish rereading model configuration from updating a running backend. Keep the full LocalAI restart recommendation for changed reasoning options.

Assisted-by: Codex:GPT-6

* sglang backend: only pass require_reasoning when sglang supports it

Engine.async_generate() gained the require_reasoning keyword in sglang
0.5.13 and takes no **kwargs. The CPU profile builds v0.5.11 from source
and the other profiles only set a >=0.5.11 floor, so passing the keyword
unconditionally made every request fail with TypeError. Detect support
once at import time, as the file already does for sampling_seed.

enable_strict_thinking first appears in sglang 0.5.12; fix the comment.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 04:49:37 +02:00
Ettore Di Giacinto 0565fc06af Merge PR #12302: chore(deps): bump LocalAGI to 7e0947d (no-RAG-DB crash fix, tool filters, per-collection models)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 01:48:00 +00:00
f82efdb43b fix(models): fallback to application config default context size in /v1/models/capabilities (#12202) (#12216)
* fix(models): fallback to application config default context size (#12202)

Honor appConfig.ContextSize in /v1/models/capabilities when model context_size is unset.

* docs(models): explain context size fallback

Describe the application default used by capability discovery and
preserve the distinction between total context and per-request limits.

Assisted-by: Codex:GPT-6

* fix(models): apply the default context size only when context_size is unset

The request path applies the application default context size only
when a model leaves context_size unset. An explicit 0 or -1 falls
through to the backend fallback. The capabilities endpoint now does
the same, so it reports the value the backend uses.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 00:52:56 +02:00
Ettore Di Giacinto 84e2fc5eac feat(agents): support tool lists and required tool in distributed mode
LocalAGI 7e0947d added allowed_tools/excluded_tools and the
required_tool_before_finish gate. Single-node agents get them through
LocalAGI's runtime, but the distributed executor drives cogito directly
and its static config meta did not list the fields, so the agent form
hid them and the worker ignored them.

The distributed config now parses the tool lists from a JSON array or a
comma/newline separated string, and the meta entries match LocalAGI's.
The executor filters the knowledge base, skill and MCP tools (MCP via
cogito.WithMCPToolFilter) before the model sees them, and re-prompts the
model when it answers before the required tool returned "ok": true, up
to the configured number of reminders.

LocalAGI keeps its filter and gate helpers unexported, so a minimal copy
lives in core/services/agents/toolpolicy.go. A spec compares the meta
entries with LocalAGI's to catch drift.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:49:12 +00:00
Ettore Di Giacinto b7155a9d97 chore(deps): bump LocalAGI to 7e0947d
Pick up the LocalAGI PRs merged after f2a2af4:
- per-collection embedding and reranker models, locked per collection
  so one agent's upload or rerank no longer stalls the others (#499)
- required_tool_before_finish: a tool the agent must call successfully
  before it may answer (#495)
- allowed_tools / excluded_tools per agent, applied to MCP tools too
  (#480)

Document the new agent settings. They show up in the single-node agent
form, which reads LocalAGI's config metadata; distributed mode keeps its
own field list and does not offer them yet.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:37:03 +00:00
Ettore Di Giacinto dae9a431e8 Merge remote-tracking branch 'origin/master' into feat/failover-chains
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 19:57:52 +00:00
Ettore Di Giacinto 7f821ab7ef Merge PR #12287: chore(gallery): add Sharp-Spark 4B variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

# Conflicts:
#	docs/content/features/model-gallery.md
2026-09-27 19:49:34 +00:00
Ettore Di Giacinto 4693ccf737 Merge PR #12293: chore(gallery): add Swift 1.5 GSQ-RCO variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

# Conflicts:
#	docs/content/features/model-gallery.md
2026-09-27 19:49:33 +00:00
Ettore Di Giacinto aed7b7823a Merge PR #12295: chore(gallery): add ThinkingCap Qwen3.8 variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

# Conflicts:
#	docs/content/features/model-gallery.md
2026-09-27 19:49:32 +00:00
Ettore Di Giacinto a45dd81e22 Merge PR #12296: chore(gallery): add Agention Qwen3.8 variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

# Conflicts:
#	docs/content/features/model-gallery.md
2026-09-27 19:49:32 +00:00
Ettore Di Giacinto ae6ccb5f52 Merge PR #12298: chore(gallery): add Qwopus Flash V2 variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 19:48:59 +00:00
Ettore Di Giacinto 1819c33f5f Merge PR #12300: chore(gallery): add Cyber-Tiel-Coder variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 19:48:58 +00:00
localai-org-maint-botandlocalai-org-maint-bot 490b952d06 feat(gallery): publish signed OCI fallbacks (#12182)
* feat(gallery): publish signed OCI fallbacks

Publish both official gallery indexes with their local base configs so
an outage of the HTTP and GitHub sources can fall back to Quay.

Keep artifact signing policies separate from backend image policies,
and expose each moving gallery tag only after its digest is signed.

Assisted-by: Codex:gpt-6

* fix(gallery): confine packaged files to selected roots

Use directory-scoped file access to reject symlink escapes during gallery packaging. Create private bundle files for the publishing runner.

Assisted-by: Codex:GPT-6

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:38 +02:00
localai-org-maint-botandlocalai-org-maint-bot 5794495a37 fix(responses): preserve streamed output items (#12048)
Keep each message and reasoning item at its announced output index.
Include the answer in completed responses with reasoning or fallback
function calls, and retain reasoning supplied through backend deltas.

Add regression coverage for stream indices, final output, plain text,
and automatic tool parsing.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:29 +02:00
localai-org-maint-botandlocalai-org-maint-bot 0e52bb657e fix(responses): wait for complete JSON tool calls (#12001)
Partial JSON parsing heals a name-only chunk into a tool call. The
stream emits that call with empty arguments and skips later chunks.

Require complete JSON before emitting terminal tool-call events.
Preserve complete calls before an unfinished trailing call, and count
only actual tool calls. Add split-chunk regression tests and docs.

Refs #11635. The non-streaming report remains unconfirmed.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:24 +02:00
localai-org-maint-botandlocalai-org-maint-bot 1b1bd0f069 fix(compose): request NVIDIA compute capability (#11990)
The legacy NVIDIA device reservation requests utility without compute.
Docker derives driver capabilities from that list, leaving CUDA libraries
unavailable even when monitoring works.

Include compute in the legacy example and clarify the matching docs.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:07:08 +02:00
localai-org-maint-bot 065f9691fa chore(gallery): add Cyber-Tiel-Coder variants
Add Q4 and Q8 MTP builds with a shared vision projector and installation docs.

Assisted-by: Codex:gpt-6
2026-09-27 16:05:55 +00:00
localai-org-maint-bot 6043e5e0cb chore(gallery): add Qwopus Flash V2 variants
Add Q4_K_M and Q8_0 builds with vision and MTP decoding. Pin the
weights and projector to a verified Hugging Face revision.

Assisted-by: Codex:gpt-6
2026-09-27 12:04:40 +00:00
localai-org-maint-bot dcddb641f0 chore(gallery): add Agention Qwen3.8 variants
Add IQ4_XS and Q4_K_M GGUF builds with a BF16 vision projector.
Pin verified artifacts and document installation and variant selection.

Assisted-by: Codex:gpt-6
2026-09-27 08:05:46 +00:00
Ettore Di Giacinto 1b6b4b806a docs: document localai-proxy and distributed failover limits
Add the localai-proxy known limits (no grammar or media forwarding,
TTS streams that end cleanly after an upstream failure, the /v1 path in
upstream_url), state that the Unimplemented skip covers the APIs that
answer HTTP 501, and describe a NATS-partitioned leader and pin
re-sync in distributed mode.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 452a3a0cbe fix(docs): correct warm-toggle and Reconcile-without-Store claims
The failover chain editor leaves the warm toggle enabled on every row
(the model list has no backend field to gate on) and relies on the
server warning instead, so the docs describing it as disabled for
remote targets were wrong. Separately, syncstate's hydrate() returns
early with no Store or Loader, so a Reconcile tick is a no-op rather
than one that empties the map — correct that claim everywhere it was
repeated (contributor guide, distsync comment, design spec).

No behavior change.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 00d9d80897 docs: require distributed-aware state for stateful features
Add the distributed-aware state contributor rule: any feature that
keeps runtime state must choose shared (syncstate), single-runner
(advisorylock), stateless, or documented per-instance behaviour, so it
behaves correctly across multiple frontends instead of diverging
silently. Also sweeps the failover/localai-proxy docs for gaps found
along the way: the UI (chain editor field, health strip, overview
page, chain badge), the 429->ResourceExhausted trip and 501->skip
mappings, and a spec correction for the live-transcription bridge's
actual close behavior.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 4a2a6180a2 test(localai-proxy): proxy APIs and realtime stages end to end
The e2e suite now registers the localai-proxy binary and points proxy
models back at the test server itself, so a request leaves LocalAI
through the backend, returns over REST and is answered by a mock model.
Chat, embeddings, TTS and transcription through the proxy return the
upstream model's answer; a chain whose proxy target's upstream model
fails to load serves from the local target; and a realtime pipeline
whose LLM stage is a chain on a remote target completes a turn, then
switches to the local target with a localai.model.failover trip event
when a gate in front of the upstream starts answering 503.

The docs describe the localai-proxy backend next to cloud-proxy and add
a per-stage remote LocalAI example to the failover page.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto e59fb854ee fix(failover): free the leader lock soon after the leader's host dies
A leader whose host died without closing its connection kept the
advisory lock for about two hours of OS keepalive defaults, and no other
frontend could probe. The lock session now sets short TCP keepalives and
tcp_user_timeout, so the server drops it within about 30 seconds.

Shutdown now closes the lock for good, so a tick that runs after it
cannot take the lock back.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto b5792e4d17 fix(failover): keep the probe leader until its session ends
A lock taken per tick passed between frontends on almost every tick, so
several frontends probed at once and each change of leader re-sent the
warm set and all state. The leader now holds a dedicated PostgreSQL
session with the advisory lock and keeps it until it shuts down or the
session dies.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto d30c33a074 feat(failover): run one prober per cluster and pin warm targets on workers
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto f7f12f8203 fix(failover): warn about warm on a remote target, align the spec
The spec promised a load-time warning when a chain marks a remote
target warm, where the flag does nothing; the loader now logs it. The
remote-backend test moves into ModelConfig.IsRemoteProxy so the loader
and the failover manager agree on what is remote.

The spec now says what ships: a load blocked by pinned warm targets
proceeds over the limit after eviction retries, without an error that
names them.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto 132bb216a1 fix(failover): resolve chains in transcription and sound-only sessions
Transcription-only and sound-detection-only realtime sessions passed a
chain config straight to the model loader. It has no backend, so the
loader fell back to greedy backend auto-detection: slow, and ending in
an unhelpful error. Sound-only sessions are a main use of chains.

The stage routing of the full pipeline moves into a stageRouter that
both realtime model kinds embed. Every stage resolves to the chain's
active target at build time and goes through the failover plan per
call. The session sends failover events for any model with chain
stages, and restarts them when a transcription session.update swaps
the model.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto 2ecbaab010 fix(failover): send remote targets their own upstream model
A chain request reached a cloud-proxy target with the client's model,
the chain name, whenever the target set no upstream_model: passthrough
forwards the body's model and translate falls back to it. The upstream
answered 404, which neither retries nor trips, while the liveness
probe, which checks the target's own name, kept passing.

PrepareTarget now sets the upstream model of a remote target to
proxy.upstream_model or the target name, the same name the probe uses.
The request pipeline and realtime chain stages both call it.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto 5d2b90b848 fix(failover): never load a warm target inside a probe
A warm target's liveness probe called ModelLoader.Load, which blocked
until the model finished loading (while the warm preload loaded it
too). Tick waited for every probe, so all probing froze, and the probe
then ran HealthCheck on an expired context and tripped the target at
every startup.

The prober now takes a function that returns the running backend
without loading it. A target that is not loaded passes liveness; its
recovery is neither confirmed nor failed and it returns to healthy
after min_dwell, like a cold target. Tick no longer waits for probes:
each probe applies its own result and a target whose probe is running
is skipped.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto dcc7bd8548 docs: document model failover chains
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
localai-org-maint-bot 7460312d23 chore(gallery): add ThinkingCap Qwen3.8 variants
Add Q4_K_M and Q8_0 builds with the F16 vision projector and install docs.
Pin artifact revisions and verify SHA256 against HF LFS metadata and HTTP
headers.

Assisted-by: Codex:gpt-6
2026-09-27 04:04:34 +00:00
localai-org-maint-bot dc4db4b119 chore(gallery): add Swift 1.5 GSQ-RCO variants
Add four text-only llama.cpp builds with pinned download URLs and
verified checksums. Document variant selection and the model license.

Assisted-by: Codex:gpt-6
2026-09-27 00:07:50 +00:00
localai-org-maint-bot e35a967970 chore(gallery): add Sharp-Spark 4B variants
Add Q4, Q5, and Q6 builds with the embedded chat template.
Pin artifact revisions and document installation.

Assisted-by: Codex:gpt-6
2026-09-26 20:05:14 +00:00
localai-org-maint-botandlocalai-org-maint-bot 92b8f1d8ed chore(gallery): add MiMo distill Qwen 9B variants (#12282)
Add Q4_K_M and Q8_0 builds with the F16 vision projector and pinned
artifact URLs. Document installation and explicit variant selection.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-26 18:54:08 +02:00
localai-org-maint-botandlocalai-org-maint-bot dbdd2a4101 chore(gallery): add Hemmingway and remove invalid chat entry (#12278)
* fix(gallery): remove invalid Qwen-Image chat entry

The entry sends diffusion weights to llama.cpp as a chat model.
Remove it and document the existing image-generation alternatives.

Assisted-by: Codex:gpt-6

* feat(gallery): add Hemmingway-1 GGUF variants

Add Q4_K_M and Q8_0 builds for llama.cpp with embedded chat templates.
Record the upstream CC BY-NC 4.0 license and installation instructions.
Verify both SHA256 values against Hugging Face LFS metadata and headers.

Assisted-by: Codex:gpt-6

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-26 15:28:41 +02:00
Plamen K. Kosseffandlocalai-org-maint-bot 6cfc99196d feat(gallery): read metadata for system-path backends, enabling variant aliases (#12141)
* feat(gallery): read metadata for system-path backends, enabling variant aliases

Problem:
- ListSystemBackends only read metadata.json for user-managed backends;
  the system-path scan (LOCALAI_BACKENDS_SYSTEM_PATH) was a bare
  directory walk with Metadata hardcoded nil
- system-packaged backends (distro packages installing several
  accelerator builds of one backend) could not declare aliases or meta
  indirection at all, while gallery-installed backends could
- surfaced while packaging LocalAI for Gentoo: the packages install
  cpu-/rocm-/vulkan-audio-cpp as system backends aliased to audio-cpp,
  which the server ignored

Change:
- scan each root separately, clean the system collection against the
  user-managed one, merge, then build and resolve — precedence lives in
  one explicit step
- alias candidates carry their own metadata: the resolved alias entry
  can never pair one installation's executable with another's metadata,
  and it reports the chosen candidate's origin (IsSystem)
- deterministic resolution: entries build in sorted name order and
  candidates sort by name at the resolution site, independent of scan
  order

Precedence (user-managed always wins):
- a user-managed backend hides a same-named system backend entirely
- a user-managed variant takes over its whole alias family: the alias
  resolves among user-managed variants only and the system family's
  concrete names disappear — family versions move together, and a stale
  system variant may not work with newer models, so it must not stay
  reachable
- a system variant's alias never hijacks a name that exists as a
  user-managed backend

Tests: Ginkgo regressions for system-path aliasing, same-name hiding,
family takeover, and the full metadata permutation matrix of
cross-root name collisions (both directions, with and without
metadata on each side).

Docs: new "Backend Directory Format" section (run.sh, metadata.json,
alias resolution — previously undocumented for user-managed backends
too) and "System-Provided Backends" with the precedence rules.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>

* fix(gallery): preserve managed meta backends

A system alias can replace a user-managed meta backend during discovery.
Protect meta entries with the same precedence guard as concrete backends.
Add a regression test and clarify the documented precedence.

Assisted-by: Codex:GPT-6
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>

---------

Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-26 11:33:51 +00:00
mudler-agentandEttore Di Giacinto 543fb4bd24 fix(gallery): verification follow-ups for oci:// galleries (#12243)
* fix(gallery): verification follow-ups for oci:// galleries

Follow-ups from the post-merge review of #12238 and #12239.

Only a policy decision is a refusal now. cosignverify wraps
ErrPolicyRejected around a failed signature check, an identity or
source-repository mismatch, a not_before cutoff and a missing or
unparseable bundle. A TUF, registry or network failure during
verification, or a timeout, is an outage: the gallery falls back to the
copy verified under the current policy, as it does when the registry is
down.

An oci:// gallery with a verification block, or any oci:// gallery under
strict integrity, is no longer answered by an https://, github: or
file:// mirror. Such a mirror is ignored with a warning, because nothing
can check its signature. The index of an HTTP gallery, whose policy only
covers its backend images, is cached under the URL-only name again, so no
unchecked body is stored under a policy-keyed name.

The in-memory index cache key now includes the policy. After a runtime
policy change the index is fetched again, and entries with a relative url
install again.

The registry digest lookups after install and upgrade, and in the
upgrade check, run only for real registry references (new
URI.LooksLikeRegistryOCI), not for ollama:// or ocifile://.

The refusal message names strict integrity when that is the cause, and
the gallery name is no longer repeated.

Specs pin the URL-only cache name for galleries without a policy, a fixed
key for a fixed policy, and that every GalleryVerification field changes
the key. The docs describe refusal, outage, mirrors and strict integrity.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): reset listings on gallery changes, classify referrer outages

Review follow-ups for this PR.

The React UI lists from AvailableGalleryModelsCached, which is keyed by
nothing. A gallery change through the settings API or a
runtime_settings.json edit now drops that listing when the model or
backend gallery configuration differs. Before, the UI kept the old list,
with local paths into the old policy's tree, until the next background
refresh, or for good when the new policy refused the gallery.

In cosignverify, a referrer the registry fails to serve now makes the
lookup an outage whatever other referrers failed and in any order, since
the unread one may be the valid signature. An invalid policy (Validate in
NewVerifier, an unparseable not_before) is ErrPolicyRejected, because no
fetch can make it usable.

The docs say that only an oci:// gallery with a verification block skips
non-OCI mirrors, and list an unusable policy as a refusal.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-24 21:20:21 +02:00
mudler-agentandEttore Di Giacinto f5c4083d7e fix(gallery): tie oci:// gallery caches to the verification policy (#12239)
The unpacked oci:// gallery cache and the last known good copy of the
index were keyed on the gallery URL only. After an operator tightened a
gallery's verification policy (added source_repository, moved not_before
forward), content verified under the older policy, or under none, was
still served for up to an hour from the unpacked cache, and indefinitely
from the last known good copy while fetches failed. A fetch refused by
signature verification also fell back to that last known good copy, so
a refusal became a silent downgrade. Turning strict integrity on did not
stop an unverified cached copy from being served either.

Name both caches by the URL plus a stable hash of the policy. A gallery
without a policy keeps its old URL-only name, so existing caches stay
usable. A fetch refused by the policy, or by strict integrity, is now
reported and never answered with a cached copy; a network failure still
falls back, but only to a copy verified under the current policy. The
strict integrity check runs before the cache is read.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-24 15:43:51 +02:00
mudler-agentandEttore Di Giacinto be0671c635 feat(gallery): optionally pin the signing certificate's source repository (#12235)
* feat(cosignverify): optionally pin the certificate's source repository

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): source_repository in the verification policy

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(gallery): when source_repository is checked; test the issuer

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-24 09:04:55 +02:00
mudler-agentandEttore Di Giacinto 16b60a766d docs(gallery): remove per-model documentation sections (#12222)
The gallery agent started adding documentation sections for individual
models (NeoHorse, Maple-Preview, Hy-MT2, Instella-MoE, Spark-X2.5,
Occamy, etc.) to the model-gallery page. This clutters the general
gallery documentation with model-specific install instructions and
descriptions that belong in the gallery index or model cards, not in
the feature docs.

Remove all per-model sections. The page now covers only the gallery
infrastructure: how galleries work, how to add them, the API, variants,
and the stable-diffusion/whisper examples that were already there.

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-23 12:39:11 +02:00
21c5495a99 feat(vllm-cpp): add GLiNER2.5 NER via TokenClassify (#12140)
* feat(vllm-cpp): add GLiNER2.5 NER via TokenClassify

Wire the vllm-cpp backend to the C ABI NER surface (vllm_gliner_ner,
ABI v27) so LocalAI can serve zero-shot named entity recognition through
the existing TokenClassify gRPC method.

backend.go: TokenClassify method on *VllmCpp calls vllm_gliner_ner with
the text and labels, copies the C-owned entity array into protobuf
TokenClassifyEntity messages, and frees the result.

govllmcpp.go: cNerEntity and cNerResult Go POD mirrors matching the C
structs; vllmGlinerNer and vllmNerResultFree purego bindings; abiVersion
bumped 26 -> 27.

options.go: ner_labels, ner_threshold, ner_max_width parsed from
engine_args.

pkg/grpc: ClassifyModel interface and TokenClassify server handler
(follows the Embedding locking pattern).

core/config: vllm-cpp backend declares MethodTokenClassify and
UsecaseTokenClassify.

docs/content/features/vllm-cpp.md: NER section documenting the
engine_args keys and the host-forward contract.

Assisted-by: MAKI:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): correct NER pointer lint directive

Use the govet directive for the C-owned NER array, matching the other
purego pointer conversions. The array remains valid until its deferred
free; the misspelled directive caused CI to flag this conversion.

Assisted-by: Codex:gpt-6 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vllm-cpp): add kev-compatible SystemOne API endpoints

Add POST /v1/systemone, /v1/systemone/permute, and
/v1/systemone/separate to LocalAI, mirroring the kev project's
structured-extraction API. Each endpoint runs zero-shot NER over the
rendered state text and builds kev-compatible answers for three question
types: noul (binary entity presence), choice (pick one option), and
score (pick one level).

The TokenClassifyRequest proto gains a `repeated string labels` field so
each question can supply its own labels at inference time, and
TokenClassifier gains TokenClassifyWithLabels for per-call label
selection. The vllm-cpp backend uses request labels when non-empty,
falling back to configured ner_labels then the built-in defaults.

Helpers (renderState, softmax, choiceConfidence, scoreConfidence, r2)
are ported from kev/api.py and mirrored in vllm.cpp's api_server.cpp so
both servers produce the same answer shape.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): suppress gosec G404 on seeded permutation RNG

The SystemOne permute endpoint uses math/rand with a caller-supplied
seed for reproducible option permutations, matching kev's random.seed.
gosec flags this as G404 (weak RNG). Add #nosec with a comment naming
the intent: this is reproducibility, not cryptography.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(vllm-cpp): bump vllm.cpp pin to GLiNER2.5 merge commit

Advance VLLM_CPP_VERSION from f3cd97e to 5058268d, the commit that
landed GLiNER2.5 zero-shot NER support (PR #3224) in vllm.cpp. This
brings the DeBERTa v2 encoder, GLiNER2 boundary head, C ABI NER
functions, and server endpoints into the LocalAI vllm-cpp backend.
The ABI version (27) and Go struct mirrors already match.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): use instruction text as NER label in SystemOne handler

The SystemOne handler was passing question IDs as NER labels for noul
questions and bare key names for choice questions, so the model never
matched any entities. Port the label mapping from vllm.cpp's
ParseSystemOneBody:

- noul: use the rendered instructions field (with instr alias) as the
  NER label, not the question ID
- choice: use optionText(name, desc) — "name: description" or "name"
  when the description is null/empty — not the bare key
- score: already correct (rendered criteria text)
- permute: shuffle indices and build parallel key/label arrays so the
  NER call uses the optionText labels while the response is keyed by
  the original option names

Also add the instructions field to the SystemOneQuestion schema struct
(accepted alongside the instr backward-compat alias).

Verified end-to-end against the real GLiNER2.5 model: noul questions
now find "Apple Inc. is" (organization, 0.999) and "Tim Cook is"
(person, 0.852) where they previously returned zero entities.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-23 12:31:28 +02:00
f9965e5e7d batch(gallery): merge 14 gallery model-addition PRs (#12221)
* chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* feat(gallery): add NeoHorse-1-9B GGUF variants

Add the official Q4_K_M, Q5_K_M, and Q8_0 builds with revision-pinned
weights and verified SHA256 values.

Assisted-by: Codex:gpt-6

* chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* feat(gallery): add Qwen3.8 35B Distill variants

Add Q4_K_M, Q5_K_M, and Q8_0 builds with the vision projector.
Pin publisher revisions and document installation and variant selection.

Assisted-by: Codex:gpt-6

* feat(gallery): add ByteShape Qwen3.8 variants

Offer five ShapeLearn GGUF builds with a vision projector and MTP.
Pin downloads to a verified HF revision and document variant selection.

Assisted-by: Codex:gpt-6

* feat(gallery): add Flash Next GSQ-RCO variants

Offer Q2_0, IQ2_XS, and IQ3_XXS builds with both model shards and the
vision projector. Pin and verify download hashes and document how to
select each variant.

Assisted-by: Codex:GPT-6

* feat(gallery): add Occamy-1.0 GGUF variants

Add the publisher's Q4_K_M and Q8_0 builds with the F16 vision projector.
Link the builds as variants and pin downloads to a verified revision.
Document installation and the source tokenizer's NFC requirement.

Assisted-by: Codex:gpt-6

* fix(gallery): set MiniCPM5 context at the top level

The Q4 and Q8 overrides place context_size inside parameters, where
PredictionOptions ignores it. Move it beside parameters so both
builds use the intended 8,192-token context, matching F16.

Assisted-by: Codex:GPT-6

* feat(gallery): add Hy-MT2 7B GGUF variants

Offer the official Q4_K_M, Q6_K, and Q8_0 builds for translation.
Pin the downloads and document installation and translation prompts.

Assisted-by: Codex:gpt-6

* feat(gallery): add Maple-Preview GGUF variants

Offer four ternary builds through the existing llama.cpp backend.
Use the publisher's CPU settings and embedded chat template.
Pin downloads and verify SHA256 values against two HF metadata sources.
Document installation and explicit variant selection.

Assisted-by: Codex:gpt-6

* feat(gallery): add Qwen3.8 Cyber GGUF variants

Offer IQ4_XS and Q8_0 builds with the matching BF16 vision projector.
Pin download revisions and document automatic and explicit selection.

Assisted-by: Codex:GPT-6

* feat(gallery): add official NeoHorse 4B variants

Offer the official Q5_K_M and BF16 GGUF builds alongside the existing
community quantizations. Pin both downloads and document variant selection.

Assisted-by: Codex:gpt-6

* chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-23 12:24:35 +02:00
leilei3167andmudler-agent 7e0c5d957c docs(gpu): drop false auto-set claim for gfx1151 env vars (#12109)
The ROCm/hipblas image does not set HSA_OVERRIDE_GFX_VERSION,
ROCBLAS_USE_HIPBLASLT, HSA_XNACK, or HSA_ENABLE_SDMA. Remove the
misleading parenthetical so readers know to pass them explicitly.

Fixes #12071

Assisted-by: Cursor:Composer

Signed-off-by: lei_lei <imleilei123@gmail.com>
Co-authored-by: mudler-agent <mudler-bot@c3os.io>
2026-09-23 12:24:05 +02:00
Plamen K. Kosseff e7306a087a feat(audio-cpp): AUDIOCPP_DEFAULT_BACKEND fallback for models without a backend option (#12133)
Models whose options carry no explicit backend: open their session on the
CPU backend even in accelerator images. The gallery entries carry
backend:best since #11892; this covers hand-written model configurations
the same way, per deployment: the environment variable supplies the
fallback, an explicit backend: option always wins (merged beside the
existing threads and maingpu fallbacks), and validation reuses the
option parser.

Assisted-by: Claude:claude-fable-5

Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
2026-09-23 12:16:31 +02:00
Richard Palethorpe 7034d353b1 feat(kimodo): Add observability hooks (#12184)
feat(kimodocpp): add API and backend request observability

Capture animation requests, phase timings, output metadata, and failures in traces. Record correct API error statuses and cover completed, running, failed, and disabled tracing.

Assisted-by: Codex:GPT-6

Signed-off-by: Richard Palethorpe <io@richiejp.com>
2026-09-22 07:06:32 +00:00