Design specs and implementation plans are working notes and are not
kept in the tree.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* feat(gallery): publish signed OCI fallbacks
Publish both official gallery indexes with their local base configs so
an outage of the HTTP and GitHub sources can fall back to Quay.
Keep artifact signing policies separate from backend image policies,
and expose each moving gallery tag only after its digest is signed.
Assisted-by: Codex:gpt-6
* fix(gallery): confine packaged files to selected roots
Use directory-scoped file access to reject symlink escapes during gallery packaging. Create private bundle files for the publishing runner.
Assisted-by: Codex:GPT-6
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(system): report per-model DRM VRAM
Expose optional resident device memory for local backend process trees.
Deduplicate DRM clients and omit unsupported or incomplete readings.
Document accounting limits and preserve a measured zero in JSON.
Closes#11970.
Assisted-by: Codex:gpt-6
* fix(system): document trusted procfs reads
Scope G304 annotations to paths built from the fixed procfs root,
integer process IDs, and kernel directory entries. These reads accept
no user-controlled path components.
Assisted-by: Codex:GPT-6 gosec
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Keep each message and reasoning item at its announced output index.
Include the answer in completed responses with reasoning or fallback
function calls, and retain reasoning supplied through backend deltas.
Add regression coverage for stream indices, final output, plain text,
and automatic tool parsing.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Partial JSON parsing heals a name-only chunk into a tool call. The
stream emits that call with empty arguments and skips later chunks.
Require complete JSON before emitting terminal tool-call events.
Preserve complete calls before an unfinished trailing call, and count
only actual tool calls. Add split-chunk regression tests and docs.
Refs #11635. The non-streaming report remains unconfirmed.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The legacy NVIDIA device reservation requests utility without compute.
Docker derives driver capabilities from that list, leaving CUDA libraries
unavailable even when monitoring works.
Include compute in the legacy example and clarify the matching docs.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Older Go linkers stamp pure-Go hosts with SDK metadata that disables
modern Metal APIs. Select Go 1.27 for Darwin builds and document the
backend rebuild requirement.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The link text still said chatbot-ui, but it now pointed at the examples
repository root. Link the configurations directory, which holds the
example model config files, and describe it as such.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Add IQ4_XS and Q4_K_M GGUF builds with a BF16 vision projector.
Pin verified artifacts and document installation and variant selection.
Assisted-by: Codex:gpt-6
Add the localai-proxy known limits (no grammar or media forwarding,
TTS streams that end cleanly after an upstream failure, the /v1 path in
upstream_url), state that the Unimplemented skip covers the APIs that
answer HTTP 501, and describe a NATS-partitioned leader and pin
re-sync in distributed mode.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The failover chain editor leaves the warm toggle enabled on every row
(the model list has no backend field to gate on) and relies on the
server warning instead, so the docs describing it as disabled for
remote targets were wrong. Separately, syncstate's hydrate() returns
early with no Store or Loader, so a Reconcile tick is a no-op rather
than one that empties the map — correct that claim everywhere it was
repeated (contributor guide, distsync comment, design spec).
No behavior change.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Add the distributed-aware state contributor rule: any feature that
keeps runtime state must choose shared (syncstate), single-runner
(advisorylock), stateless, or documented per-instance behaviour, so it
behaves correctly across multiple frontends instead of diverging
silently. Also sweeps the failover/localai-proxy docs for gaps found
along the way: the UI (chain editor field, health strip, overview
page, chain badge), the 429->ResourceExhausted trip and 501->skip
mappings, and a spec correction for the live-transcription bridge's
actual close behavior.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The e2e suite now registers the localai-proxy binary and points proxy
models back at the test server itself, so a request leaves LocalAI
through the backend, returns over REST and is answered by a mock model.
Chat, embeddings, TTS and transcription through the proxy return the
upstream model's answer; a chain whose proxy target's upstream model
fails to load serves from the local target; and a realtime pipeline
whose LLM stage is a chain on a remote target completes a turn, then
switches to the local target with a localai.model.failover trip event
when a gate in front of the upstream starts answering 503.
The docs describe the localai-proxy backend next to cloud-proxy and add
a per-stage remote LocalAI example to the failover page.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A leader whose host died without closing its connection kept the
advisory lock for about two hours of OS keepalive defaults, and no other
frontend could probe. The lock session now sets short TCP keepalives and
tcp_user_timeout, so the server drops it within about 30 seconds.
Shutdown now closes the lock for good, so a tick that runs after it
cannot take the lock back.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A lock taken per tick passed between frontends on almost every tick, so
several frontends probed at once and each change of leader re-sent the
warm set and all state. The leader now holds a dedicated PostgreSQL
session with the advisory lock and keeps it until it shuts down or the
session dies.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Failover state is per frontend, remote chain targets only cover chat,
and chains have no UI. Design shared pins, health and chain state for
distributed mode, a localai-proxy backend for every API including live
transcription, a chain editor and health view, and a contributor rule
for distributed-aware state.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The spec promised a load-time warning when a chain marks a remote
target warm, where the flag does nothing; the loader now logs it. The
remote-backend test moves into ModelConfig.IsRemoteProxy so the loader
and the failover manager agree on what is remote.
The spec now says what ships: a load blocked by pinned warm targets
proceeds over the limit after eviction retries, without an error that
names them.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Transcription-only and sound-detection-only realtime sessions passed a
chain config straight to the model loader. It has no backend, so the
loader fell back to greedy backend auto-detection: slow, and ending in
an unhelpful error. Sound-only sessions are a main use of chains.
The stage routing of the full pipeline moves into a stageRouter that
both realtime model kinds embed. Every stage resolves to the chain's
active target at build time and goes through the failover plan per
call. The session sends failover events for any model with chain
stages, and restarts them when a transcription session.update swaps
the model.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A chain request reached a cloud-proxy target with the client's model,
the chain name, whenever the target set no upstream_model: passthrough
forwards the body's model and translate falls back to it. The upstream
answered 404, which neither retries nor trips, while the liveness
probe, which checks the target's own name, kept passing.
PrepareTarget now sets the upstream model of a remote target to
proxy.upstream_model or the target name, the same name the probe uses.
The request pipeline and realtime chain stages both call it.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A warm target's liveness probe called ModelLoader.Load, which blocked
until the model finished loading (while the warm preload loaded it
too). Tick waited for every probe, so all probing froze, and the probe
then ran HealthCheck on an expired context and tripped the target at
every startup.
The prober now takes a function that returns the running backend
without loading it. A target that is not loaded passes liveness; its
recovery is neither confirmed nor failed and it returns to healthy
after min_dwell, like a cold target. Tick no longer waits for probes:
each probe applies its own result and a target whose probe is running
is skipped.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A missing file is not a reliable signal: models download on first use,
some backends need no file, and dotted names like Phi-3.5-mini look like
paths. Marking such a fallback down removed the retry a chain exists
for. Cold targets are now judged only by real requests.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Remote liveness uses /v1/models, which every OpenAI-compatible upstream
serves. Recovery sends one minimal request for the target's usecase.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
/readyz exists only on LocalAI upstreams, and a cloud-proxy target can
point at any OpenAI-compatible provider.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A model served from a remote upstream has no local fallback when that
upstream is unhealthy, so clients reimplement failover and lose realtime
context when they switch endpoints.
Define a failover block on model configs: an ordered list of targets with
active probes, in-request retry before the response is committed,
fail-back with hysteresis, and switch events over REST, SSE and the
realtime socket.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Add Q4_K_M and Q8_0 builds with the F16 vision projector and install docs.
Pin artifact revisions and verify SHA256 against HF LFS metadata and HTTP
headers.
Assisted-by: Codex:gpt-6
Add four text-only llama.cpp builds with pinned download URLs and
verified checksums. Document variant selection and the model license.
Assisted-by: Codex:gpt-6
Add Q4_K_M and Q8_0 builds with the F16 vision projector and pinned
artifact URLs. Document installation and explicit variant selection.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* fix(gallery): remove invalid Qwen-Image chat entry
The entry sends diffusion weights to llama.cpp as a chat model.
Remove it and document the existing image-generation alternatives.
Assisted-by: Codex:gpt-6
* feat(gallery): add Hemmingway-1 GGUF variants
Add Q4_K_M and Q8_0 builds for llama.cpp with embedded chat templates.
Record the upstream CC BY-NC 4.0 license and installation instructions.
Verify both SHA256 values against Hugging Face LFS metadata and headers.
Assisted-by: Codex:gpt-6
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(gallery): read metadata for system-path backends, enabling variant aliases
Problem:
- ListSystemBackends only read metadata.json for user-managed backends;
the system-path scan (LOCALAI_BACKENDS_SYSTEM_PATH) was a bare
directory walk with Metadata hardcoded nil
- system-packaged backends (distro packages installing several
accelerator builds of one backend) could not declare aliases or meta
indirection at all, while gallery-installed backends could
- surfaced while packaging LocalAI for Gentoo: the packages install
cpu-/rocm-/vulkan-audio-cpp as system backends aliased to audio-cpp,
which the server ignored
Change:
- scan each root separately, clean the system collection against the
user-managed one, merge, then build and resolve — precedence lives in
one explicit step
- alias candidates carry their own metadata: the resolved alias entry
can never pair one installation's executable with another's metadata,
and it reports the chosen candidate's origin (IsSystem)
- deterministic resolution: entries build in sorted name order and
candidates sort by name at the resolution site, independent of scan
order
Precedence (user-managed always wins):
- a user-managed backend hides a same-named system backend entirely
- a user-managed variant takes over its whole alias family: the alias
resolves among user-managed variants only and the system family's
concrete names disappear — family versions move together, and a stale
system variant may not work with newer models, so it must not stay
reachable
- a system variant's alias never hijacks a name that exists as a
user-managed backend
Tests: Ginkgo regressions for system-path aliasing, same-name hiding,
family takeover, and the full metadata permutation matrix of
cross-root name collisions (both directions, with and without
metadata on each side).
Docs: new "Backend Directory Format" section (run.sh, metadata.json,
alias resolution — previously undocumented for user-managed backends
too) and "System-Provided Backends" with the precedence rules.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
* fix(gallery): preserve managed meta backends
A system alias can replace a user-managed meta backend during discovery.
Protect meta entries with the same precedence guard as concrete backends.
Add a regression test and clarify the documented precedence.
Assisted-by: Codex:GPT-6
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
---------
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* fix(gallery): verification follow-ups for oci:// galleries
Follow-ups from the post-merge review of #12238 and #12239.
Only a policy decision is a refusal now. cosignverify wraps
ErrPolicyRejected around a failed signature check, an identity or
source-repository mismatch, a not_before cutoff and a missing or
unparseable bundle. A TUF, registry or network failure during
verification, or a timeout, is an outage: the gallery falls back to the
copy verified under the current policy, as it does when the registry is
down.
An oci:// gallery with a verification block, or any oci:// gallery under
strict integrity, is no longer answered by an https://, github: or
file:// mirror. Such a mirror is ignored with a warning, because nothing
can check its signature. The index of an HTTP gallery, whose policy only
covers its backend images, is cached under the URL-only name again, so no
unchecked body is stored under a policy-keyed name.
The in-memory index cache key now includes the policy. After a runtime
policy change the index is fetched again, and entries with a relative url
install again.
The registry digest lookups after install and upgrade, and in the
upgrade check, run only for real registry references (new
URI.LooksLikeRegistryOCI), not for ollama:// or ocifile://.
The refusal message names strict integrity when that is the cause, and
the gallery name is no longer repeated.
Specs pin the URL-only cache name for galleries without a policy, a fixed
key for a fixed policy, and that every GalleryVerification field changes
the key. The docs describe refusal, outage, mirrors and strict integrity.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(gallery): reset listings on gallery changes, classify referrer outages
Review follow-ups for this PR.
The React UI lists from AvailableGalleryModelsCached, which is keyed by
nothing. A gallery change through the settings API or a
runtime_settings.json edit now drops that listing when the model or
backend gallery configuration differs. Before, the UI kept the old list,
with local paths into the old policy's tree, until the next background
refresh, or for good when the new policy refused the gallery.
In cosignverify, a referrer the registry fails to serve now makes the
lookup an outage whatever other referrers failed and in any order, since
the unread one may be the valid signature. An invalid policy (Validate in
NewVerifier, an unparseable not_before) is ErrPolicyRejected, because no
fetch can make it usable.
The docs say that only an oci:// gallery with a verification block skips
non-OCI mirrors, and list an unusable policy as a refusal.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
The unpacked oci:// gallery cache and the last known good copy of the
index were keyed on the gallery URL only. After an operator tightened a
gallery's verification policy (added source_repository, moved not_before
forward), content verified under the older policy, or under none, was
still served for up to an hour from the unpacked cache, and indefinitely
from the last known good copy while fetches failed. A fetch refused by
signature verification also fell back to that last known good copy, so
a refusal became a silent downgrade. Turning strict integrity on did not
stop an unverified cached copy from being served either.
Name both caches by the URL plus a stable hash of the policy. A gallery
without a policy keeps its old URL-only name, so existing caches stay
usable. A fetch refused by the policy, or by strict integrity, is now
reported and never answered with a cached copy; a network failure still
falls back, but only to a copy verified under the current policy. The
strict integrity check runs before the cache is read.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>