Older Go linkers stamp pure-Go hosts with SDK metadata that disables
modern Metal APIs. Select Go 1.27 for Darwin builds and document the
backend rebuild requirement.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* ⬆️ Update TheTom/llama-cpp-turboquant
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(turboquant): drop the upstreamed D512 patch
Upstream c0e227c guards D512 declarations, dispatch, and instances with
GGML_USE_HIP. This prevents the CUDA shared-memory overflow that our
patch addressed. The old patch now rejects the guarded source.
Remove the obsolete patch for the pinned a3d5603d revision. The remaining
patch series applies successfully, and the build-target test passes.
Assisted-by: Codex:gpt-6
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Picks up mudler/LocalAGI 3ce0a08 "fix(mcp): re-dial an MCP session the
server has dropped". Agents open their MCP sessions once, when they are
created; when the MCP server restarts it forgets them and the go-sdk
client does not reconnect by itself, so the agent kept a dead session -
or, behind a server that revives unknown session IDs, a stale tool
list - until LocalAI restarted.
Only go.mod/go.sum change; core/services/agentpool builds and vets
against the new version.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
Add Q4_K_M and Q8_0 builds with the F16 vision projector and pinned
artifact URLs. Document installation and explicit variant selection.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The cached gallery listing refreshed each entry's installed flag with
one os.Stat per entry, under the cache's global write lock. A gallery
holds about 1,900 entries. On a models directory on SMB, one refresh
took about 14s. The listing and every row's VRAM estimate run this
refresh, and the lock serialized them, so the models page took
minutes to load.
The installed check now lists the models directory once and looks up
each entry in that listing. On the same SMB share the listing takes
about 75ms. The answers match os.Stat: a symlink counts only when its
target exists, and names with a path separator still use os.Stat. The
listing runs before the lock is taken, so the lock covers only the
flag updates.
Concurrent callers on a cold cache now share one upstream load. Before
this, each caller fetched the gallery index and the configs itself.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
The hugo-theme-relearn submodule was bumped to 9.1.x in #12096,
which requires Hugo >= 0.165.0. The pinned 0.146.3 broke the docs
site build with a template error in alias.html that could not
evaluate the Locale field on langs.Language.
Bump HUGO_VERSION to 0.166.0 (latest stable) to satisfy the
theme minimum and resolve the alias.html template error.
Assisted-by: nib:claude-sonnet-4.5 [bash] [read] [edit]
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* ⬆️ Update leejet/stable-diffusion.cpp
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(stablediffusion-ggml): adapt to upstream tiling struct rename
Upstream commit 2f88688 renamed the sd_tiling_params_t fields from
tile_size_x/y to tile_size_w/h and rel_size_x/y to rel_size_w/h.
Update the gosd.cpp wrappers to match so the C++ backend compiles.
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* fix(gallery): remove invalid Qwen-Image chat entry
The entry sends diffusion weights to llama.cpp as a chat model.
Remove it and document the existing image-generation alternatives.
Assisted-by: Codex:gpt-6
* feat(gallery): add Hemmingway-1 GGUF variants
Add Q4_K_M and Q8_0 builds for llama.cpp with embedded chat templates.
Record the upstream CC BY-NC 4.0 license and installation instructions.
Verify both SHA256 values against Hugging Face LFS metadata and headers.
Assisted-by: Codex:gpt-6
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(gallery): read metadata for system-path backends, enabling variant aliases
Problem:
- ListSystemBackends only read metadata.json for user-managed backends;
the system-path scan (LOCALAI_BACKENDS_SYSTEM_PATH) was a bare
directory walk with Metadata hardcoded nil
- system-packaged backends (distro packages installing several
accelerator builds of one backend) could not declare aliases or meta
indirection at all, while gallery-installed backends could
- surfaced while packaging LocalAI for Gentoo: the packages install
cpu-/rocm-/vulkan-audio-cpp as system backends aliased to audio-cpp,
which the server ignored
Change:
- scan each root separately, clean the system collection against the
user-managed one, merge, then build and resolve — precedence lives in
one explicit step
- alias candidates carry their own metadata: the resolved alias entry
can never pair one installation's executable with another's metadata,
and it reports the chosen candidate's origin (IsSystem)
- deterministic resolution: entries build in sorted name order and
candidates sort by name at the resolution site, independent of scan
order
Precedence (user-managed always wins):
- a user-managed backend hides a same-named system backend entirely
- a user-managed variant takes over its whole alias family: the alias
resolves among user-managed variants only and the system family's
concrete names disappear — family versions move together, and a stale
system variant may not work with newer models, so it must not stay
reachable
- a system variant's alias never hijacks a name that exists as a
user-managed backend
Tests: Ginkgo regressions for system-path aliasing, same-name hiding,
family takeover, and the full metadata permutation matrix of
cross-root name collisions (both directions, with and without
metadata on each side).
Docs: new "Backend Directory Format" section (run.sh, metadata.json,
alias resolution — previously undocumented for user-managed backends
too) and "System-Provided Backends" with the precedence rules.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
* fix(gallery): preserve managed meta backends
A system alias can replace a user-managed meta backend during discovery.
Protect meta entries with the same precedence guard as concrete backends.
Add a regression test and clarify the documented precedence.
Assisted-by: Codex:GPT-6
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
---------
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29
Replace the model-specific SystemOne gRPC approach with a generic Score
RPC extension. The pre-existing Score RPC (previously unused by any
backend) now carries question_type and response_json fields:
- question_type="systemone" routes kev/laya decision-pipeline requests
through the unified vllm_decide C ABI (v29), returning the full
response JSON in response_json.
- question_type empty routes cua-s1-forms candidate scoring through the
same vllm_decide ABI, returning CandidateScore probabilities.
The vllm-cpp backend's Score() method calls vllm_decide and dispatches
by architecture internally. The /v1/systemone HTTP endpoint checks
whether the model's backend supports Score; if so, it forwards the raw
request JSON and returns the backend response as-is. Other backends
fall through to the existing NER-based path.
This mirrors the vllm.cpp C ABI refactor (PR #3301) that replaced
vllm_systemone + vllm_score with a single vllm_decide function. The
purego bindings bump abiVersion from 27 to 29 and resolve vllm_decide
and vllm_decide_free symbols.
Also fixes validModelPath to accept cua-s1-forms.json and
rl_agent_config.json alongside config.json, matching the engine's
model_loader.cpp config-filename ordering.
AI-Assisted: true
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore: ⬆️ Update mudler/vllm.cpp to e28ec46c6 (fix macOS -Werror build)
Bumps vllm.cpp to e28ec46c6 which fixes a -Wnull-conversion error in
qwen3_5.cpp:12483 that broke the macOS Metal CI build under -Werror.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Add five nemo-speech-cpp gallery entries for the diarization
capability introduced by the NeMo-Speech.cpp bump in #12257:
- nemo-speech-cpp-sortformer-diarization-v2: standalone streaming
Sortformer 4-speaker diarization (nvidia/diar_streaming_sortformer_4spk-v2).
Serves /v1/audio/diarization with known_usecases: [diarization].
- nemo-speech-cpp-nemotron-3.5-asr-streaming: standalone multilingual
streaming ASR (nvidia/nemotron-3.5-asr-streaming-0.6b).
- nemo-speech-cpp-nemotron-3.5-asr-streaming-diarized: Nemotron ASR
with the sortformer attached via the diar_model option, giving
per-word speaker tags on /v1/audio/transcriptions.
- nemo-speech-cpp-parakeet-tdt-0.6b-v3: standalone multilingual ASR,
25 languages (nvidia/parakeet-tdt-0.6b-v3).
- nemo-speech-cpp-parakeet-tdt-0.6b-v3-diarized: Parakeet v3 ASR
with the sortformer attached via the diar_model option, giving
per-word speaker tags on /v1/audio/transcriptions.
No backend code changes: the sortformer to familyDiarization mapping,
the diar_model option, and the MethodDiarize gRPC method already exist.
All five entries verified end-to-end against a running LocalAI instance.
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Add gallery entries for three vllm.cpp-backed models:
- laya-vllm-cpp: multilingual non-autoregressive System 1 decision
model (ModernBERT-large, 421M params). Served via POST /v1/systemone.
Weights from convaiinnovations/laya (Apache-2.0).
- cua-s1-forms-vllm-cpp: one-pass option scorer for GUI form filling.
Served via POST /api/score. Weights from cua-ai/cua-s1-forms (MIT).
- gliner2.5-vllm-cpp: zero-shot named entity recognition and structured
extraction (GLiNER2.5, mDeBERTa-v3-base, 287M params). Served via the
token classification endpoint. Weights from fastino/gliner2.5-multi-v1
(Apache-2.0).
All entries set backend: vllm-cpp and declare CPU+GPU tags.
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
llama.cpp picks a new random media marker per backend process. LocalAI
cached the first probe on the model config and skipped later probes when
MediaMarker was non-empty, so after SINGLE_ACTIVE eviction/reload the
prompt still used the stale marker and mtmd_tokenize failed (0 markers
vs 1 bitmap).
Re-probe whenever the model was not already resident before Load, while
still skipping the RPC on warm cache hits.
Fixes#12246
Assisted-by: Cursor:composer-2.5
Signed-off-by: leilei3167 <imleilei123@gmail.com>
* fix(gallery): verification follow-ups for oci:// galleries
Follow-ups from the post-merge review of #12238 and #12239.
Only a policy decision is a refusal now. cosignverify wraps
ErrPolicyRejected around a failed signature check, an identity or
source-repository mismatch, a not_before cutoff and a missing or
unparseable bundle. A TUF, registry or network failure during
verification, or a timeout, is an outage: the gallery falls back to the
copy verified under the current policy, as it does when the registry is
down.
An oci:// gallery with a verification block, or any oci:// gallery under
strict integrity, is no longer answered by an https://, github: or
file:// mirror. Such a mirror is ignored with a warning, because nothing
can check its signature. The index of an HTTP gallery, whose policy only
covers its backend images, is cached under the URL-only name again, so no
unchecked body is stored under a policy-keyed name.
The in-memory index cache key now includes the policy. After a runtime
policy change the index is fetched again, and entries with a relative url
install again.
The registry digest lookups after install and upgrade, and in the
upgrade check, run only for real registry references (new
URI.LooksLikeRegistryOCI), not for ollama:// or ocifile://.
The refusal message names strict integrity when that is the cause, and
the gallery name is no longer repeated.
Specs pin the URL-only cache name for galleries without a policy, a fixed
key for a fixed policy, and that every GalleryVerification field changes
the key. The docs describe refusal, outage, mirrors and strict integrity.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(gallery): reset listings on gallery changes, classify referrer outages
Review follow-ups for this PR.
The React UI lists from AvailableGalleryModelsCached, which is keyed by
nothing. A gallery change through the settings API or a
runtime_settings.json edit now drops that listing when the model or
backend gallery configuration differs. Before, the UI kept the old list,
with local paths into the old policy's tree, until the next background
refresh, or for good when the new policy refused the gallery.
In cosignverify, a referrer the registry fails to serve now makes the
lookup an outage whatever other referrers failed and in any order, since
the unread one may be the valid signature. An invalid policy (Validate in
NewVerifier, an unparseable not_before) is ErrPolicyRejected, because no
fetch can make it usable.
The docs say that only an oci:// gallery with a verification block skips
non-OCI mirrors, and list an unusable policy as a refusal.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
The unpacked oci:// gallery cache and the last known good copy of the
index were keyed on the gallery URL only. After an operator tightened a
gallery's verification policy (added source_repository, moved not_before
forward), content verified under the older policy, or under none, was
still served for up to an hour from the unpacked cache, and indefinitely
from the last known good copy while fetches failed. A fetch refused by
signature verification also fell back to that last known good copy, so
a refusal became a silent downgrade. Turning strict integrity on did not
stop an unverified cached copy from being served either.
Name both caches by the URL plus a stable hash of the policy. A gallery
without a policy keeps its old URL-only name, so existing caches stay
usable. A fetch refused by the policy, or by strict integrity, is now
reported and never answered with a cached copy; a network failure still
falls back, but only to a copy verified under the current policy. The
strict integrity check runs before the cache is read.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
A backend installed from an oci://host/repo:tag or oci://host/repo@sha256
URI was downloaded correctly, but the digest lookup that follows the
install (and the one after an upgrade) passed the raw URI to the registry
client. The client does not know the oci:// scheme: with a port in the
host it failed to parse the reference, without one it read "oci" as the
registry host and queried https://oci/v2/. The install still succeeded,
so the only trace was a warning and an empty digest in metadata.json,
which made the next upgrade check report an upgrade for no reason.
Add downloader.URI.OCIReference and route every consumer that hands an
OCI URI to a registry client through it: the install and upgrade digest
lookups, the upgrade check, the OCI download path, the oci:// gallery
fetch and the llama.cpp importer.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>