Commit Graph
368 Commits
Author SHA1 Message Date
localai-org-maint-bot c4ab987e3e chore: merge master into distributed transport PR
Preserve failover support alongside the PostgreSQL broadcast carrier.

Assisted-by: Codex:gpt-6 golangci-lint
2026-09-28 16:06:47 +00:00
Ettore Di Giacinto 9c156656bd Merge PR #12285: feat(failover): serve a model name from a chain of local and remote targets
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 15:32:30 +00:00
localai-org-maint-bot 4151ceda85 chore: merge master into distributed transport PR
Keep the version-reporting import from master and omit the unused
sanitize import after the transport changes.

Assisted-by: Codex:GPT-6
2026-09-28 12:02:15 +00:00
Stefan Walcz bc1d9924de fix(llama-cpp): keep llama.cpp's default cache_ram instead of no limit (#12297)
grpc-server.cpp forced params.cache_ram_mib = -1 (no limit) since #7009.
Since v4.3 kv_unified and cache_idle_slots are on by default, so every
distinct prompt now leaves its slot KV state in the host-side prompt
cache, and without a limit the backend grows until the host runs out of
memory.

Measured on gfx1151 (Strix Halo, 128 GB), llama-cpp backend, one request
at a time, 100 distinct prompts of ~2000 characters plus a fixed system
prompt, max_tokens 200:

  model                        cache_ram     RSS loaded -> after 100
  gemma-4-26B-A4B (q8_0 KV)    -1 (default)  1.4 GB -> 25.3 GB
  Qwen3.6-35B-A3B (q8_0 KV)    -1 (default)  1.1 GB -> 19.5 GB
  gemma-4-26B-A4B              -1, same prompt 100x  1.4 GB -> 1.6 GB
  gemma-4-26B-A4B              4096          1.4 GB -> 5.4 GB (flat from
                                             request 20 on, same latency)
  Qwen3.6-35B-A3B              4096          1.1 GB -> 5.1 GB (flat)

The memory is not released when idle. In production a document
classification pass pushed the daily chat model to 34 GB RSS overnight.

Drop the override so llama.cpp's own default (8192 MiB) applies; the
cache_ram option still accepts -1 for users who want no limit. Update
both docs tables (the option reference and the prompt-cache table) and
note what -1 does.


Assisted-by: Claude:claude-opus-5-5
Assisted-by: Codex:GPT-6

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-09-28 11:37:24 +02:00
mudler-agentandEttore Di Giacinto 97ad8f1d70 fix(distributed): stage the files a model install declares (#12309)
* fix(distributed): stage every shard of a split GGUF

A split GGUF is configured by its first shard only. llama.cpp opens the
other "-0000N-of-0000M.gguf" files from the same directory by name. The
router staged only the configured path, so the worker received shard 1
and the load failed with "failed to load GGUF split".

The router now stages the remaining shards next to the first one. A
missing shard fails the load and names the file. The file count for
progress and the payload size also include all shards. The payload size
feeds the load deadline and the disk-headroom check. For a 111 GB model
whose first shard is 10 MB, both were sized for less than 1 GB.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): stage the files a model install declares

Replace the split GGUF file name matching with the model's own file
list. A gallery install or an import records every file of the model in
._gallery_<name>.yaml (files:), and a config can list more under
download_files:. The router now stages all of these files, not only the
files that the config's path fields name. This includes the other
shards of a split GGUF, which llama.cpp opens by name.

The application gives the router a resolver that reads the two lists.
The resolver looks up the files by model name when it stages them, so a
replica that the reconciler loads from saved load options gets the same
files. backend.proto does not change.

A declared file that is missing on the frontend is skipped with a
warning. The load deadline and the disk headroom check include the
declared files.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 08:34:56 +02:00
Pratik Gandhi ea10f26f57 docs: fix 11 dead links in the documentation (#12320)
- compatibility table: voxtral.c, OuteTTS and VoxCPM repos live under
  antirez, edwko and OpenBMB
- distributed inferencing: llama.cpp RPC README moved to tools/rpc under ggml-org
- customize-model: the phi-2 example config moved to LocalAI-examples, and
  embedded/models was replaced by the gallery
- model-gallery: malformed URL; link the gallery index
- GPU acceleration: ROCm install guide moved
- integrations: Wave Terminal docs page moved to ai-presets

Assisted-by: Claude:claude-fable-5-1

Signed-off-by: Pratik Gandhi <travpreneur@gmail.com>
2026-09-28 08:33:15 +02:00
61f4f67b75 sglang backend: pass through thinking_budget + require_reasoning (#12193)
* sglang backend: pass through thinking_budget + require_reasoning

sglang's raw Engine.async_generate() API (which this backend calls
directly, bypassing sglang's own OpenAI server) supports a precise,
tokenizer-derived reasoning-length budget via
sampling_params["custom_params"]["thinking_budget"] plus
require_reasoning=True, gated behind --enable-strict-thinking. Neither
was reachable through LocalAI: this backend built sampling_params only
from a fixed field mapping (temperature, top_p, ...) with no custom_params
key, and never passed require_reasoning to async_generate at all.

- LoadModel now reads a model-level "thinking_budget" option (same
  mechanism as the existing tool_parser/reasoning_parser options), and
  _build_sampling_params adds it as custom_params.thinking_budget on
  every request when configured.
- _new_reasoning_parser already derives, from the rendered prompt, whether
  the model's chat template pre-opened a reasoning block (Qwen3-style
  templates append <think> to the prompt instead of letting the model
  emit it) -- the same signal sglang's own OpenAI server computes from
  per-template config to decide require_reasoning. This backend has no
  template manager, so it now returns that signal too and _predict
  forwards it to async_generate(require_reasoning=...).

Verified against production (NVFP4, sm_121, Qwen3.6-35B-A3B) via a raw
Engine.async_generate() call bypassing this backend: 301 reasoning
tokens against a 300-token budget, clean completion, ~27s. Not yet
verified through this backend's own gRPC path end-to-end (no local
CUDA/sglang environment available here) -- existing + new unit tests in
test.py cover the pure-Python merge/passthrough logic only.

Scope note: require_reasoning is derived only from the existing
prompt-suffix heuristic, not sglang's full per-template
_get_reasoning_from_request decision tree (minimax-m3/hunyuan special
cases etc.) -- this backend has no template manager to evaluate that
tree against, and the prompt-suffix check is the one heuristic already
validated in this file (test_reasoning_parser_forced_when_template_prefills_think_tag).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: honour a model-level reasoning_default

A model YAML can already carry "parameters: reasoning_effort:", but that
value only reaches this backend when a *caller* sets it per request (the Go
side turns it into Metadata["enable_thinking"]). As a model-level default it
is silently dropped: a config reading "reasoning_effort: none" still produces
full reasoning on every request, so the config says one thing and the model
does another.

That gap is expensive in practice. On a self-hosted Qwen3.6-35B-A3B the
reasoning phase consumed the entire max_tokens budget before any content was
produced - 90% of code completions came back empty at max_tokens=768, and the
server log filled with "backend produced only reasoning, retrying". The
config looked like reasoning was off the whole time.

This adds "reasoning_default:off" (or ":on") on the same model-level
options: mechanism as thinking_budget. A per-request value always wins; the
default only fills in when the request is silent.

Measured on the stack above (sglang 0.5.20, NVFP4, GB10/sm_121) after
applying it:
  default (nothing set)          -> 0 chars reasoning, 27 tokens
  "reasoning_effort": "none"     -> 0 chars reasoning, 27 tokens
  metadata enable_thinking=true  -> capped at the 512-token thinking_budget,
                                    541 tokens total, finish_reason stop

Tests: three cases added to backend/python/sglang/test.py covering the
default, per-request override in both directions, and the unconfigured case
(which must leave the template untouched).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: validate thinking_budget instead of crashing LoadModel

Addresses the review on this PR:

- `int(thinking_budget)` raised on values like "5000.0" or "abc" and took
  LoadModel down. The option is now parsed by _parse_thinking_budget():
  integral numbers in any spelling are accepted, anything else is ignored
  with a warning on stderr.
- Zero and negative budgets are ignored with a warning instead of being
  passed to sglang, where they have no defined meaning. Turning reasoning
  off is what reasoning_default:off is for.
- A load-time warning when thinking_budget is set but enable_strict_thinking
  is not in engine_args, since sglang then ignores the budget silently.
- Tests for integral spellings, unset, zero, negative, non-integer and the
  strict-thinking warning.

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): explain reasoning options

Document the reasoning budget, strict-thinking requirement, and
precedence of request metadata over the model-level default.

Also note that the budget has to stay well below max_tokens (otherwise
it never triggers and the reply can end up empty), and that
POST /models/reload or a backend-only restart does not pick up changed
options; LocalAI itself has to be restarted.

Assisted-by: Codex:GPT-6
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): clarify configuration reloads

Distinguish rereading model configuration from updating a running backend. Keep the full LocalAI restart recommendation for changed reasoning options.

Assisted-by: Codex:GPT-6

* sglang backend: only pass require_reasoning when sglang supports it

Engine.async_generate() gained the require_reasoning keyword in sglang
0.5.13 and takes no **kwargs. The CPU profile builds v0.5.11 from source
and the other profiles only set a >=0.5.11 floor, so passing the keyword
unconditionally made every request fail with TypeError. Detect support
once at import time, as the file already does for sampling_seed.

enable_strict_thinking first appears in sglang 0.5.12; fix the comment.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 04:49:37 +02:00
localai-org-maint-bot 733eeda123 chore: merge master into distributed transport PR
Keep the newer SQLite dependency from master to resolve the conflict.

Assisted-by: Codex:gpt-6
2026-09-28 02:02:45 +00:00
Ettore Di Giacinto 0565fc06af Merge PR #12302: chore(deps): bump LocalAGI to 7e0947d (no-RAG-DB crash fix, tool filters, per-collection models)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 01:48:00 +00:00
f82efdb43b fix(models): fallback to application config default context size in /v1/models/capabilities (#12202) (#12216)
* fix(models): fallback to application config default context size (#12202)

Honor appConfig.ContextSize in /v1/models/capabilities when model context_size is unset.

* docs(models): explain context size fallback

Describe the application default used by capability discovery and
preserve the distinction between total context and per-request limits.

Assisted-by: Codex:GPT-6

* fix(models): apply the default context size only when context_size is unset

The request path applies the application default context size only
when a model leaves context_size unset. An explicit 0 or -1 falls
through to the backend fallback. The capabilities endpoint now does
the same, so it reports the value the backend uses.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 00:52:56 +02:00
Ettore Di Giacinto 84e2fc5eac feat(agents): support tool lists and required tool in distributed mode
LocalAGI 7e0947d added allowed_tools/excluded_tools and the
required_tool_before_finish gate. Single-node agents get them through
LocalAGI's runtime, but the distributed executor drives cogito directly
and its static config meta did not list the fields, so the agent form
hid them and the worker ignored them.

The distributed config now parses the tool lists from a JSON array or a
comma/newline separated string, and the meta entries match LocalAGI's.
The executor filters the knowledge base, skill and MCP tools (MCP via
cogito.WithMCPToolFilter) before the model sees them, and re-prompts the
model when it answers before the required tool returned "ok": true, up
to the configured number of reminders.

LocalAGI keeps its filter and gate helpers unexported, so a minimal copy
lives in core/services/agents/toolpolicy.go. A spec compares the meta
entries with LocalAGI's to catch drift.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:49:12 +00:00
Ettore Di Giacinto b7155a9d97 chore(deps): bump LocalAGI to 7e0947d
Pick up the LocalAGI PRs merged after f2a2af4:
- per-collection embedding and reranker models, locked per collection
  so one agent's upload or rerank no longer stalls the others (#499)
- required_tool_before_finish: a tool the agent must call successfully
  before it may answer (#495)
- allowed_tools / excluded_tools per agent, applied to MCP tools too
  (#480)

Document the new agent settings. They show up in the single-node agent
form, which reads LocalAGI's config metadata; distributed mode keeps its
own field list and does not offer them yet.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:37:03 +00:00
Ettore Di Giacinto dae9a431e8 Merge remote-tracking branch 'origin/master' into feat/failover-chains
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 19:57:52 +00:00
Ettore Di Giacinto 7f821ab7ef Merge PR #12287: chore(gallery): add Sharp-Spark 4B variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

# Conflicts:
#	docs/content/features/model-gallery.md
2026-09-27 19:49:34 +00:00
Ettore Di Giacinto 4693ccf737 Merge PR #12293: chore(gallery): add Swift 1.5 GSQ-RCO variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

# Conflicts:
#	docs/content/features/model-gallery.md
2026-09-27 19:49:33 +00:00
Ettore Di Giacinto aed7b7823a Merge PR #12295: chore(gallery): add ThinkingCap Qwen3.8 variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

# Conflicts:
#	docs/content/features/model-gallery.md
2026-09-27 19:49:32 +00:00
Ettore Di Giacinto a45dd81e22 Merge PR #12296: chore(gallery): add Agention Qwen3.8 variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

# Conflicts:
#	docs/content/features/model-gallery.md
2026-09-27 19:49:32 +00:00
Ettore Di Giacinto ae6ccb5f52 Merge PR #12298: chore(gallery): add Qwopus Flash V2 variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 19:48:59 +00:00
Ettore Di Giacinto 1819c33f5f Merge PR #12300: chore(gallery): add Cyber-Tiel-Coder variants
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 19:48:58 +00:00
localai-org-maint-botandlocalai-org-maint-bot 490b952d06 feat(gallery): publish signed OCI fallbacks (#12182)
* feat(gallery): publish signed OCI fallbacks

Publish both official gallery indexes with their local base configs so
an outage of the HTTP and GitHub sources can fall back to Quay.

Keep artifact signing policies separate from backend image policies,
and expose each moving gallery tag only after its digest is signed.

Assisted-by: Codex:gpt-6

* fix(gallery): confine packaged files to selected roots

Use directory-scoped file access to reject symlink escapes during gallery packaging. Create private bundle files for the publishing runner.

Assisted-by: Codex:GPT-6

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:38 +02:00
localai-org-maint-botandlocalai-org-maint-bot 5794495a37 fix(responses): preserve streamed output items (#12048)
Keep each message and reasoning item at its announced output index.
Include the answer in completed responses with reasoning or fallback
function calls, and retain reasoning supplied through backend deltas.

Add regression coverage for stream indices, final output, plain text,
and automatic tool parsing.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:29 +02:00
localai-org-maint-botandlocalai-org-maint-bot 0e52bb657e fix(responses): wait for complete JSON tool calls (#12001)
Partial JSON parsing heals a name-only chunk into a tool call. The
stream emits that call with empty arguments and skips later chunks.

Require complete JSON before emitting terminal tool-call events.
Preserve complete calls before an unfinished trailing call, and count
only actual tool calls. Add split-chunk regression tests and docs.

Refs #11635. The non-streaming report remains unconfirmed.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:24 +02:00
localai-org-maint-botandlocalai-org-maint-bot 1b1bd0f069 fix(compose): request NVIDIA compute capability (#11990)
The legacy NVIDIA device reservation requests utility without compute.
Docker derives driver capabilities from that list, leaving CUDA libraries
unavailable even when monitoring works.

Include compute in the legacy example and clarify the matching docs.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:07:08 +02:00
localai-org-maint-bot 065f9691fa chore(gallery): add Cyber-Tiel-Coder variants
Add Q4 and Q8 MTP builds with a shared vision projector and installation docs.

Assisted-by: Codex:gpt-6
2026-09-27 16:05:55 +00:00
localai-org-maint-bot 6043e5e0cb chore(gallery): add Qwopus Flash V2 variants
Add Q4_K_M and Q8_0 builds with vision and MTP decoding. Pin the
weights and projector to a verified Hugging Face revision.

Assisted-by: Codex:gpt-6
2026-09-27 12:04:40 +00:00
localai-org-maint-bot dcddb641f0 chore(gallery): add Agention Qwen3.8 variants
Add IQ4_XS and Q4_K_M GGUF builds with a BF16 vision projector.
Pin verified artifacts and document installation and variant selection.

Assisted-by: Codex:gpt-6
2026-09-27 08:05:46 +00:00
Ettore Di Giacinto 1b6b4b806a docs: document localai-proxy and distributed failover limits
Add the localai-proxy known limits (no grammar or media forwarding,
TTS streams that end cleanly after an upstream failure, the /v1 path in
upstream_url), state that the Unimplemented skip covers the APIs that
answer HTTP 501, and describe a NATS-partitioned leader and pin
re-sync in distributed mode.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 452a3a0cbe fix(docs): correct warm-toggle and Reconcile-without-Store claims
The failover chain editor leaves the warm toggle enabled on every row
(the model list has no backend field to gate on) and relies on the
server warning instead, so the docs describing it as disabled for
remote targets were wrong. Separately, syncstate's hydrate() returns
early with no Store or Loader, so a Reconcile tick is a no-op rather
than one that empties the map — correct that claim everywhere it was
repeated (contributor guide, distsync comment, design spec).

No behavior change.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 00d9d80897 docs: require distributed-aware state for stateful features
Add the distributed-aware state contributor rule: any feature that
keeps runtime state must choose shared (syncstate), single-runner
(advisorylock), stateless, or documented per-instance behaviour, so it
behaves correctly across multiple frontends instead of diverging
silently. Also sweeps the failover/localai-proxy docs for gaps found
along the way: the UI (chain editor field, health strip, overview
page, chain badge), the 429->ResourceExhausted trip and 501->skip
mappings, and a spec correction for the live-transcription bridge's
actual close behavior.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 4a2a6180a2 test(localai-proxy): proxy APIs and realtime stages end to end
The e2e suite now registers the localai-proxy binary and points proxy
models back at the test server itself, so a request leaves LocalAI
through the backend, returns over REST and is answered by a mock model.
Chat, embeddings, TTS and transcription through the proxy return the
upstream model's answer; a chain whose proxy target's upstream model
fails to load serves from the local target; and a realtime pipeline
whose LLM stage is a chain on a remote target completes a turn, then
switches to the local target with a localai.model.failover trip event
when a gate in front of the upstream starts answering 503.

The docs describe the localai-proxy backend next to cloud-proxy and add
a per-stage remote LocalAI example to the failover page.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto e59fb854ee fix(failover): free the leader lock soon after the leader's host dies
A leader whose host died without closing its connection kept the
advisory lock for about two hours of OS keepalive defaults, and no other
frontend could probe. The lock session now sets short TCP keepalives and
tcp_user_timeout, so the server drops it within about 30 seconds.

Shutdown now closes the lock for good, so a tick that runs after it
cannot take the lock back.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto b5792e4d17 fix(failover): keep the probe leader until its session ends
A lock taken per tick passed between frontends on almost every tick, so
several frontends probed at once and each change of leader re-sent the
warm set and all state. The leader now holds a dedicated PostgreSQL
session with the advisory lock and keeps it until it shuts down or the
session dies.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto d30c33a074 feat(failover): run one prober per cluster and pin warm targets on workers
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto f7f12f8203 fix(failover): warn about warm on a remote target, align the spec
The spec promised a load-time warning when a chain marks a remote
target warm, where the flag does nothing; the loader now logs it. The
remote-backend test moves into ModelConfig.IsRemoteProxy so the loader
and the failover manager agree on what is remote.

The spec now says what ships: a load blocked by pinned warm targets
proceeds over the limit after eviction retries, without an error that
names them.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto 132bb216a1 fix(failover): resolve chains in transcription and sound-only sessions
Transcription-only and sound-detection-only realtime sessions passed a
chain config straight to the model loader. It has no backend, so the
loader fell back to greedy backend auto-detection: slow, and ending in
an unhelpful error. Sound-only sessions are a main use of chains.

The stage routing of the full pipeline moves into a stageRouter that
both realtime model kinds embed. Every stage resolves to the chain's
active target at build time and goes through the failover plan per
call. The session sends failover events for any model with chain
stages, and restarts them when a transcription session.update swaps
the model.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto 2ecbaab010 fix(failover): send remote targets their own upstream model
A chain request reached a cloud-proxy target with the client's model,
the chain name, whenever the target set no upstream_model: passthrough
forwards the body's model and translate falls back to it. The upstream
answered 404, which neither retries nor trips, while the liveness
probe, which checks the target's own name, kept passing.

PrepareTarget now sets the upstream model of a remote target to
proxy.upstream_model or the target name, the same name the probe uses.
The request pipeline and realtime chain stages both call it.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto 5d2b90b848 fix(failover): never load a warm target inside a probe
A warm target's liveness probe called ModelLoader.Load, which blocked
until the model finished loading (while the warm preload loaded it
too). Tick waited for every probe, so all probing froze, and the probe
then ran HealthCheck on an expired context and tripped the target at
every startup.

The prober now takes a function that returns the running backend
without loading it. A target that is not loaded passes liveness; its
recovery is neither confirmed nor failed and it returns to healthy
after min_dwell, like a cold target. Tick no longer waits for probes:
each probe applies its own result and a target whose probe is running
is skipped.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
Ettore Di Giacinto dcc7bd8548 docs: document model failover chains
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:20 +00:00
localai-org-maint-bot 7460312d23 chore(gallery): add ThinkingCap Qwen3.8 variants
Add Q4_K_M and Q8_0 builds with the F16 vision projector and install docs.
Pin artifact revisions and verify SHA256 against HF LFS metadata and HTTP
headers.

Assisted-by: Codex:gpt-6
2026-09-27 04:04:34 +00:00
Ettore Di Giacinto 0cb7b7e70d fix(distributed): prove agent worker credentials
Keep pending agent workers alive with authenticated heartbeats while preserving approval as the tunnel and execution boundary. Exercise the real binary credential handoff through an authenticated inference and verify that credential remains non-admin.

Assisted-by: Codex:GPT-5 [apply_patch] [exec_command]
2026-09-27 03:05:13 +00:00
localai-org-maint-botandEttore Di Giacinto f041bee1cb fix(ui): complete node operation journeys (#12070)
Worker setup now uses a focused drawer without interrupting fleet
monitoring. Both generated worker commands retain their NATS settings.

Log views preserve their launch context and show one-based replica
labels. Bulk controls disclose selections outside the current view.

Assisted-by: Codex:gpt-5 Playwright ESLint

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 03:05:13 +00:00
localai-org-maint-bot b2be52e6f7 fix(worker): resolve staging directory symlinks
Resolve allowed directories before comparing them with resolved files.
Otherwise staging rejects valid files under macOS temporary paths.
Cover aliased roots, sibling paths, and symlinks escaping the root.

Assisted-by: Codex:gpt-6
2026-09-27 03:05:13 +00:00
Ettore Di Giacinto df30b1a0c4 fix(distributed): back a failed claim off instead of respinning it
The claim queue's attempts counter grew without bound and nothing read it. At
the default two-second poll a permanently undispatchable row cost about 43000
UPDATEs a day, and it cost more than writes: rows are claimed oldest first, so
the oldest stuck row was re-claimed ahead of every newer one on every tick and
held a dispatch slot while it failed. One poison row starved the queue behind
it.

No dead letter, and that is the decision rather than the omission. Read
settleClaim: the only outcome that releases a claim is one where NOTHING was
learned about the work. No agent worker was connected, the tunnel broke, a peer
could not be reached, the stream was refused before the request body left this
replica. Not one of those is a worker saying it ran the job and it failed, and
an attempt ceiling would turn "the fleet was away long enough" into a job
failure nobody reported, which is the collapse this whole design exists to
prevent pointed at work instead of at nodes. The one verdict available here,
that no build of any worker serves this kind, is already settled as an answer.

So the retry stays unbounded and the RATE does not. Each release stamps the row
with the earliest it may be claimed again, doubling from two seconds to a cap
of sixty, computed in the release statement from the row's own attempts count
and stamped on the DATABASE clock, because that is the clock competing replicas
order the queue on. Queued work becomes claimable again within one cap of the
fleet returning, and a stuck row no longer holds the head of the queue. A claim
released by the reap carries no delay at all: that work was never handed to
anyone, so there is nothing to back off from.

Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 03:05:13 +00:00
Ettore Di Giacinto 4aa0288c6e fix(worker): check a gRPC port is free before handing it out
The backend port allocator allocated from its own bookkeeping alone. That
bookkeeping records what this worker did with a port, and the collision it
cannot see is with something this worker never did: the default base port is
50051, inside Linux's default ephemeral range of 32768 to 60999, so the kernel
hands ports in this range to outbound connections and to anything that binds
port 0. A backend handed one of those dies on bind, and the frontend sees a
backend that will not start.

Every candidate is now probed by binding the exact address the backend will
listen on, in all four allocation branches: the key's own port, the free pool,
a grown port and a stolen one. Probing the free pool matters as much as
probing a grown port, because a port this worker released is exactly as
available to the kernel as one it never used.

A candidate that fails the probe is quarantined rather than blacklisted, since
whatever holds it is usually an ephemeral connection that gives it back, and
its affinity claim is dropped so an unbindable port does not stay reserved for
the key that last held it. Exhaustion now says how many candidates were
skipped, which is what tells an operator "something else is in my range" from
"my range is too narrow".

This does not remove the race and cannot: between the probe and the child's
bind the kernel can still give the port away. It removes the far larger window
in which the allocator hands out a port the kernel gave away minutes ago,
which was the whole of the observed one-in-three harness flake. The e2e
harness comment that recorded the missing check is corrected, and the docs say
how to move the range out of the ephemeral one entirely.

Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 03:05:13 +00:00
Ettore Di Giacinto 07208b9180 fix(distributed): say that skills and collections are replica-local
Skills and RAG collections had no cross-replica invalidation, and the two
builders that would have published one were deleted earlier in this branch
because nothing called them. Wiring one now would be wrong, not merely late.

Both features are derived entirely from the frontend's own state directory.
A skills.Service indexes <state dir>/skills, a collections backend enumerates
<state dir>/collections and holds one handle per collection it found there,
and no replica reads or writes another replica's copy of either. In
distributed mode PostgreSQL carries a skill's NAME and description in
skills_metadata, and nothing else: Get, Search, Export and the resource verbs
all read local files. So a peer told to drop a cache entry would rebuild it
from a directory that does not hold the change. For a postgres-engine
collection it would be worse than a no-op, since re-deriving one on a replica
with no local index file yields a collection that answers with an empty file
list against a populated vector store. What is missing is shared storage, not
a broadcast.

Recorded rather than left silent: the two cache fields say why nothing
invalidates them, a distributed frontend logs the limitation once at startup,
and the docs name the two deployments that avoid it. The new spec pins the
premise, so a change that moved either directory onto storage every replica
mounts reddens and the decision gets taken again.

Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 03:05:13 +00:00
Ettore Di Giacinto d72b591a78 feat(cluster): make a peer prove which replica it is
GET /api/cluster/peer authenticated with the deployment's shared
registration token and took the dialling replica's id from ?id= on trust.
Every worker holds that token, so anything holding it could open a peer
link as any replica: relay through it to every worker tunnel that replica
owns, displace a real replica's inbound link by declaring its id, and
point the roughly 31 GiB per-session receive window at one replica.

Validating the id against the instances table does not fix this, because
the attack declares a real replica's id. So the route now checks two
credentials and needs both. The shared token still says the dialler
belongs to this deployment; a new per-replica credential says which
replica it is.

The credential follows the per-node worker credential rather than
inventing a second mechanism: crypto/rand.Text, stored only as a hex
SHA-256, compared in constant time, with no fallback to the shared token.
It differs in the stronger direction. A worker's credential is minted by
the frontend and handed over once; a replica writes its own instances
row, so it mints its own secret, publishes only the hash in the same
statement that publishes its address, and never sends the plaintext
anywhere but the peer dial.

A peer that presents no credential is refused, not waved through. An old
replica and an attacker holding the shared token send the same request,
so accepting the first accepts the second; there is no safe downgrade
here, only a quiet one. The refusal is made loud instead, on both sides,
naming the upgrade rather than the network. On the documented
frontend-first order a new replica still dials an old one; an old replica
cannot dial a new one, which costs relayed requests that land on a
not-yet-restarted replica and surfaces as no route, never as absence.

A rejected peer gets its own sentinel, ErrPeerRejected, whose unwrap
chain carries ErrPeerUnreachable as well and no absence sentinel at all.
Keeping the older sentinel means no existing consumer changes behaviour;
the cause stays out of the chain, so absence cannot escape through it and
nothing can read an authorization failure as a worker that went away.

One consequence beyond the fix: a replica with no advertised address has
no instances row, so it now cannot dial out either. It was already
unreachable inward. The startup error and the docs say so.

Registry.Register, NewMembership, NewPeerPool, PeerHandler and
RegisterClusterRoutes all gained required arguments, so the identity
cannot be dropped without a compile failure.

Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 03:05:13 +00:00
Ettore Di Giacinto eb1656ed9c chore(distributed): take the nats-io modules out of the build
Distributed mode has not dialled a message broker since the control plane
moved onto the workers' own outward tunnels and every fan-out family moved
onto PostgreSQL LISTEN/NOTIFY. What was left was the dependency itself, and
the code that existed only to feed it.

Dropped from go.mod: nats-io/jwt/v2, nats-io/nats.go, nats-io/nkeys,
nats-io/nuid and testcontainers-go/modules/nats, along with the fourteen
indirect requires that only the NATS testcontainer pulled in. go.sum carries
no nats line either, so the removal is not the partial kind where the require
goes and the checksum stays.

Deleted with them: pkg/natsauth in full, the broker client's remaining
options and TLS files, the per-node JWT minting on both the register and the
approve path, and the natsauth.Config parameter threaded through the node
routes. The credential manager is renamed and stripped rather than deleted,
because it still holds the tunnel token that every re-registration rotates.

The bus flags stay accepted and ignored, and are now hidden, on every command
that had them, so an existing unit file, compose file or Helm values file
still starts on the day of the upgrade. What is not kept is the validation
that REQUIRED one: a distributed frontend started with no bus URL is no
longer fatal. The TLS paths lose type:"existingfile" deliberately, so a
certificate deleted along with the broker cannot fail a startup.

One operator-visible behaviour change: --nats-require-auth no longer makes an
agent worker wait through admin approval. Ask for that wait with
--distributed-require-auth, which already implied it. It is documented in the
migration section and pinned from both sides.

A deployment now needs PostgreSQL and the frontends' own HTTP listener, and
nothing else.

coverage-baseline.txt moves from 54.2 to 62.0.

Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 03:05:13 +00:00
Ettore Di Giacinto e4e5dbd227 chore(distributed): stop telling an operator to run a NATS cluster
Every carrier had already moved and no process opened a bus connection, but
the surface an operator reads still described a deployment with a broker in
it: a compose service, a 220-line credential-generation script, two CI steps
pulling a container nothing started, two flag tables offering --nats-url, an
architecture diagram with a NATS box wired to the workers, a join-command
generator in the Nodes page that emitted --nats-url for agent workers, and a
test suite that stood a NATS server up for specs that no longer used it.

That is the one way this programme could still fail invisibly. Every test
passes, every binary works, and every production deployment goes on running
and paying for infrastructure that carries nothing.

Nothing in this repository starts a NATS server any more. The compose file is
four services, the docs say to shut the broker down and what to keep, and the
e2e suite runs on one PostgreSQL container.

The three LOCALAI_NATS_*_TIMEOUT env vars are KEPT, and are now documented
twice as being kept. They were never broker settings: each names a control-RPC
budget the frontend applies to a worker, still read and still enforced. They
carry the prefix only because they arrived with the bus, and renaming them
would break every existing deployment for cosmetics.

The agent worker's join command was the last surface still emitting the flag,
two tasks after the agent worker stopped dialling. The Playwright spec that
covered it asserted the opposite of what is now true, so it is inverted rather
than deleted, and it reads the rendered command string rather than the
component's variables: the variables are what the fix removes, so a spec
reading them would have stopped compiling instead of failing, and a compile
error is not evidence about what an operator is shown.

nats_jwt_test.go and its helpers are deleted. They pinned a real server
ENFORCING the minted permissions. The CONTENT of those allow lists is still
pinned, untouched, by pkg/natsauth's own suites, including the spec that
refuses to let the agent lists go empty, since an empty allow list in NATS
means unrestricted. The enforcement half is retired rather than moved:
enforcement is a property of a connection, and nothing opens one.

The suite's own NATS container goes with them, which the brief left for the
next task. Removing the pre-pull while BeforeSuite still ran the image would
have defeated the step rather than cleaned it up, and this change removes the
last reader of TestInfra.NC. agent_native_executor_test.go and
mcp_ci_job_test.go are moved onto infra.Bus() instead of deleted: they were
the last two specs building a bridge and a dispatcher on a client nobody uses,
which is exactly the drift TestInfra.Bus's own comment warns about.

cluster.Options.NatsURL is now fed a deliberately dead address rather than a
live container's. Frontends and agent workers still receive LOCALAI_NATS_URL,
because that is the coverage for the promise that an existing command line
still starts; sourcing it from a running server would have let a regression
that actually dialled it pass. The control in cluster_control_test.go keeps
its assertion and loses its explanation, which claimed the deployment had a
bus and no longer could.

One latent spec race surfaced and is fixed: the background-run spec waited for
a COUNT of events and then read a snapshot for the terminal status, which is
the last event of a run and therefore always arrives after the count is met.
Its immediate twin had already been fixed this way. Nothing in production
changed.

pkg/natsauth keeps its files. It is reachable from production only through the
natsauth.Config parameter thread, and that thread is the next task's.

Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 03:05:13 +00:00
Ettore Di Giacinto 793c4e211f fix(distributed): read a departed agent tunnel as the routing fact it is
Task 4 gave agent workers tunnels and deliberately left the NodeType skip in
HealthMonitor.tunnelDeparted, with a spec asserting that an agent node whose
presence reader answers PresenceGone is NOT marked unhealthy. That spec was
scaffolding. It was true while an agent worker took its jobs and its verbs over
the message bus: a departure row for one said nothing about whether it could
work, and an early bug in the new tunnel client could otherwise have demoted a
fleet of healthy agent workers.

There is no bus. An agent worker is reachable through its tunnel and through
nothing else, so a departed agent tunnel means exactly what a departed backend
tunnel means: no live replica holds it, the departure has outlived the reconnect
grace, and that is a routing fact the scheduler and a reaper may act on. The
skip would now hide the only symptom an unreachable agent worker has. This is
the deliberate removal Task 4's M6 predicted, and task-4-report.md is where that
mutation already stands recorded red against the spec this commit deletes.

The skip existed at ONE site. router_liveness.go has none: its candidates come
from queries that already filter node_type = 'backend'. The two skips in
managers_distributed.go stay, because an agent worker still runs no backend
processes, so it has no backend to list and no backend op to apply.

Two node types can depart now, which is why the second half exists. Before this,
one type could depart and every per-node cache a departure left stale was
dropped from wherever its owner happened to notice, so a reader could not tell
which caches a demotion invalidated by reading the demotion path. Departure gets
ONE notification point. DepartureNotifier is edge triggered, because the monitor
runs on a ticker and a departed node stays departed; its subscribers are NAMED,
because what has to be caught is a forgotten cache and a count can say only that
one of four is missing; and NewHealthMonitor takes it as a required positional
argument, so a caller that does not pass one fails to compile.

Four caches subscribe: prefix-cache affinity in every model, probe freshness at
every address, in-flight staging operations, and the per-node breakdown of every
open gallery operation. The prefix-cache one is registered only when
prefix-cache routing is enabled, so --distributed-prefix-cache=false stays a
true no-op. The notification carries the node's name as well as its id, because
the staging tracker keys on the name and the other two key on the id, and a
subscriber should not have to read the registry from inside an eviction hook.

A departure notification is an act on absence, so it fires only on the routing
fact. A tunnel lost inside the grace, a worker that never dialled, a presence
query that failed and a stale heartbeat all announce nothing, asserted per node
type. The stale-heartbeat branch is excluded on purpose: it already marks the
node offline, which deletes its rows and runs the registry's replica-removed
hooks, so firing there too would double-evict and make the notification mean two
different things at its subscribers.

Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 03:05:13 +00:00
Ettore Di Giacinto da8cf9c9fd feat(distributed): carry an agent cancel on the worker's own tunnel
agent.<name>.cancel was the last family on a message bus, and the only
reason an agent worker dialled one. Its subscriber is the worker running
the execution, and a worker has no database, so the family could not move
to the PostgreSQL fan-out carrier: a cancel published there would reach no
worker while reporting that it had been sent.

It is a control verb now. An agent worker mounts workerctl.PathAgentCancel
on the loopback control plane behind its tunnel and applies the cancel to
the same registry the executor registers a run on. The frontend issues it
through nodes.AgentControlClient.CancelAgentRun.

That call is a FAN-OUT and not a pick, because nothing records which worker
holds a given execution: the claim row names the claiming replica, and it
is deleted when the run ends. Every agent worker a live replica can reach
is asked over its own tunnel, relayed by the peer mesh when a peer holds
it, and each worker answers only for itself.

The answers stay apart, which is why this family was held back. A cancel a
worker made is nil. A cancel some worker could not be asked is
ErrAgentCancelUndelivered, which is neither a refusal nor a missing run. A
cancel every reachable worker declined to own is ErrAgentRunNotOnAnyWorker.
A deployment with no agent worker is ErrNoAgentWorker. Neither new sentinel
wraps ErrWorkerUnroutable and neither is a worker answer, so nothing is
reaped, demoted or evicted because of a cancel.

A worker in the ABSENT CONNECTION condition, one whose tunnel was lost
inside the reconnect grace, counts as undelivered. It is not retried in the
call and not queued: a retry would spend a budget the caller did not
choose, and a queue would need durable state whose only consumer is a run
whose control stream went with the tunnel. A worker whose departure has
outlived the grace is the one routing fact a caller may act on and is
excluded, or a single retired agent node would make every cancel
undelivered for ever.

The fan-out reads a different node set from the pick. A draining worker
takes no new work but is still finishing what it holds, so it is offered
the cancel; a pending one is refused by the tunnel route on every dial and
is not.

With that, nothing in LocalAI connects to NATS. The agent worker's dial,
its credential ladder and its refresh loop are gone, and so is the
frontend's cancel carrier. LOCALAI_NATS_URL is accepted and ignored
everywhere, and distributed mode no longer requires it.

Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 03:05:13 +00:00