Commit Graph
8160 Commits
Author SHA1 Message Date
Ettore Di Giacinto 024f069787 Merge PR #12180: chore(model-gallery): propose variant groupings for review
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 19:48:58 +00:00
mudler-agentandEttore Di Giacinto 6bda88cc88 fix(kokoros): add missing animate3_d stub to Backend trait impl (#12301)
* fix(kokoros): add missing animate3_d stub to Backend trait impl

#12095 added the Animate3D RPC to backend.proto, but the kokoros
service never got a matching method. The tonic-generated Backend trait
now requires it, so kokoros fails to build with E0046 whenever the
full backend matrix runs.

Return Unimplemented, as the other unsupported RPCs do.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

* fix(kokoros): fill new Result fields with defaults

backend.proto added a metadata field to Result, so the struct literals
in the kokoros service no longer name every field and fail to compile.
Spread Default::default() into them, so later additive proto fields do
not break the build again.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 21:38:15 +02:00
localai-org-maint-botandlocalai-org-maint-bot 490b952d06 feat(gallery): publish signed OCI fallbacks (#12182)
* feat(gallery): publish signed OCI fallbacks

Publish both official gallery indexes with their local base configs so
an outage of the HTTP and GitHub sources can fall back to Quay.

Keep artifact signing policies separate from backend image policies,
and expose each moving gallery tag only after its digest is signed.

Assisted-by: Codex:gpt-6

* fix(gallery): confine packaged files to selected roots

Use directory-scoped file access to reject symlink escapes during gallery packaging. Create private bundle files for the publishing runner.

Assisted-by: Codex:GPT-6

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:38 +02:00
localai-org-maint-botandlocalai-org-maint-bot f154bd990a feat(system): report per-model DRM VRAM (#12026)
* feat(system): report per-model DRM VRAM

Expose optional resident device memory for local backend process trees.
Deduplicate DRM clients and omit unsupported or incomplete readings.
Document accounting limits and preserve a measured zero in JSON.

Closes #11970.

Assisted-by: Codex:gpt-6

* fix(system): document trusted procfs reads

Scope G304 annotations to paths built from the fixed procfs root,
integer process IDs, and kernel directory entries. These reads accept
no user-controlled path components.

Assisted-by: Codex:GPT-6 gosec

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:34 +02:00
localai-org-maint-botandlocalai-org-maint-bot 5794495a37 fix(responses): preserve streamed output items (#12048)
Keep each message and reasoning item at its announced output index.
Include the answer in completed responses with reasoning or fallback
function calls, and retain reasoning supplied through backend deltas.

Add regression coverage for stream indices, final output, plain text,
and automatic tool parsing.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:29 +02:00
localai-org-maint-botandlocalai-org-maint-bot 0e52bb657e fix(responses): wait for complete JSON tool calls (#12001)
Partial JSON parsing heals a name-only chunk into a tool call. The
stream emits that call with empty arguments and skips later chunks.

Require complete JSON before emitting terminal tool-call events.
Preserve complete calls before an unfinished trailing call, and count
only actual tool calls. Add split-chunk regression tests and docs.

Refs #11635. The non-streaming report remains unconfirmed.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:18:24 +02:00
Leoy b9634e0339 fix(modelartifacts): reuse committed sibling files for narrowed allow_patterns (#11484)
A request with narrower allow_patterns hashes to a different CacheKey
than an already-committed broader sibling, so committedResult misses and
materializeLocked re-fetches files the sibling already holds. After the
own-tree reuseMaterializedFile miss, consult committed sibling trees
for the same Source (type+endpoint+repo+revision), re-verify the file
through verifyDownloadedFile (full SHA-256, never size-only), and
hard-link it into the writer's staging snapshot (copy fallback only on
EXDEV). Each file is matched individually against the sibling's
manifest, so a broader request can never inherit a narrower sibling's
gaps as if complete.

The sibling manifest set is loaded and source-matched once per
materialization (files indexed by path) instead of once per staged
file, so a models volume with 20 committed artifacts and a 300-file
snapshot does one manifest pass rather than ~6000 reads and JSON
parses. The sibling-reuse behavior cases live in the package's
registered Ginkgo suite so repository test conventions apply.

Refs #11047

Signed-off-by: supermario_leo <leo.stack@outlook.com>
2026-09-27 21:10:35 +02:00
localai-org-maint-botandlocalai-org-maint-bot 1b1bd0f069 fix(compose): request NVIDIA compute capability (#11990)
The legacy NVIDIA device reservation requests utility without compute.
Docker derives driver capabilities from that list, leaving CUDA libraries
unavailable even when monitoring works.

Include compute in the legacy example and clarify the matching docs.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:07:08 +02:00
localai-org-maint-botandmudler c6f1e96a7d chore(website): refresh the counters (#12039)
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 21:07:04 +02:00
localai-org-maint-botandlocalai-org-maint-bot 657cf9ca83 fix(ci): retain backend digests for release retries (#12160)
The v4.10.0 ace-step and VibeVoice merge jobs started just after their
digest artifacts expired. Keep the small digest artifacts for seven
days so a multi-day release matrix can finish publishing its images.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:06:59 +02:00
localai-org-maint-botandlocalai-org-maint-bot 4e94c914c9 fix(swagger): describe backend metadata as an object (#12178)
Swag cannot resolve json.RawMessage in OpenAIResponse and aborts the daily
schema generation. Set its Swagger type without changing JSON encoding,
and regenerate the checked-in specifications.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:06:54 +02:00
localai-org-maint-botandlocalai-org-maint-bot a7a6bc2963 fix(ci): use Go 1.27 for Darwin backends (#12284)
Older Go linkers stamp pure-Go hosts with SDK metadata that disables
modern Metal APIs. Select Go 1.27 for Darwin builds and document the
backend rebuild requirement.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:06:50 +02:00
dependabot[bot] e6b2309b3f chore(deps): bump grpcio from 1.83.0 to 1.84.0 in /backend/python/transformers (#12102)
chore(deps): bump grpcio in /backend/python/transformers

Bumps [grpcio](https://github.com/grpc/grpc) from 1.83.0 to 1.84.0.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.83.0...v1.84.0)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.84.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:06:07 +02:00
dependabot[bot] de203c2fb5 chore(deps): update transformers requirement from >=5.15.1 to >=5.17.0 in /backend/python/transformers (#12105)
chore(deps): update transformers requirement

Updates the requirements on [transformers](https://github.com/huggingface/transformers) to permit the latest version.
- [Release notes](https://github.com/huggingface/transformers/releases)
- [Commits](https://github.com/huggingface/transformers/compare/v5.15.1...v5.17.0)

---
updated-dependencies:
- dependency-name: transformers
  dependency-version: 5.17.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:06:03 +02:00
dependabot[bot] 1cacecc460 chore(deps): update numpy requirement from >=2.5.2 to >=2.5.3 in /backend/python/transformers (#12244)
chore(deps): update numpy requirement in /backend/python/transformers

Updates the requirements on [numpy](https://github.com/numpy/numpy) to permit the latest version.
- [Release notes](https://github.com/numpy/numpy/releases)
- [Changelog](https://github.com/numpy/numpy/blob/main/doc/RELEASE_WALKTHROUGH.rst)
- [Commits](https://github.com/numpy/numpy/compare/v2.5.2...v2.5.3)

---
updated-dependencies:
- dependency-name: numpy
  dependency-version: 2.5.3
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:05:59 +02:00
dependabot[bot] 13a01e657a chore(deps): bump sentence-transformers from 5.7.0 to 6.1.0 in /backend/python/transformers (#12245)
chore(deps): bump sentence-transformers in /backend/python/transformers

Bumps [sentence-transformers](https://github.com/huggingface/sentence-transformers) from 5.7.0 to 6.1.0.
- [Release notes](https://github.com/huggingface/sentence-transformers/releases)
- [Commits](https://github.com/huggingface/sentence-transformers/compare/v5.7.0...v6.1.0)

---
updated-dependencies:
- dependency-name: sentence-transformers
  dependency-version: 6.1.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:05:55 +02:00
4bc5f292fe chore: ⬆️ Update TheTom/llama-cpp-turboquant to a3d5603d110bda29222d2011596cdc84d7fa532d (#12232)
* ⬆️ Update TheTom/llama-cpp-turboquant

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(turboquant): drop the upstreamed D512 patch

Upstream c0e227c guards D512 declarations, dispatch, and instances with
GGML_USE_HIP. This prevents the CUDA shared-memory overflow that our
patch addressed. The old patch now rejects the guarded source.

Remove the obsolete patch for the pinned a3d5603d revision. The remaining
patch series applies successfully, and the build-target test passes.

Assisted-by: Codex:gpt-6

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:05:49 +02:00
localai-org-maint-botandmudler 08827cfd5e chore: ⬆️ Update mudler/vllm.cpp to c3bebc357385990f721af66a3a6c69328dd4fc6c (#12252)
⬆️ Update mudler/vllm.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 21:05:45 +02:00
localai-org-maint-botandmudler a592e23778 chore: ⬆️ Update ggml-org/llama.cpp to 95887577ab5fead779581a7030a83c7752ff3234 (#12272)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 21:05:40 +02:00
Stefan Walcz 7690b06789 chore(deps): bump LocalAGI to 8253de9 (re-dial dropped MCP sessions) (#12299)
Picks up mudler/LocalAGI 3ce0a08 "fix(mcp): re-dial an MCP session the
server has dropped". Agents open their MCP sessions once, when they are
created; when the MCP server restarts it forgets them and the go-sdk
client does not reconnect by itself, so the agent kept a dead session -
or, behind a server that revives unknown session IDs, a stale tool
list - until LocalAI restarted.

Only go.mod/go.sum change; core/services/agentpool builds and vets
against the new version.


Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-09-27 20:58:01 +02:00
localai-org-maint-botandmudler c9e822215a chore(model-gallery): ⬆️ update checksum (#12290)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:25:27 +02:00
localai-org-maint-botandmudler 4524765b9f chore: ⬆️ Update 0xShug0/audio.cpp to 94bd4656399180befc141b17bd6696bf84df0a9f (#12289)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:25:09 +02:00
localai-org-maint-botandmudler fc6df9efc3 chore: ⬆️ Update mudler/parakeet.cpp to 2bf88954dc628b32835734e2e9159550a75a1dc6 (#12291)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:24:52 +02:00
localai-org-maint-botandmudler 01017dcdd6 chore: ⬆️ Update CrispStrobe/CrispASR to 013ae1624dc40ecf059065d577180722439f804e (#12292)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:24:25 +02:00
localai-org-maint-botandmudler 9ea9277ee6 chore: ⬆️ Update ikawrakow/ik_llama.cpp to cdf232cc17e410e60c1bc3b85516c4a41199b662 (#12288)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:24:03 +02:00
localai-org-maint-botandlocalai-org-maint-bot 92b8f1d8ed chore(gallery): add MiMo distill Qwen 9B variants (#12282)
Add Q4_K_M and Q8_0 builds with the F16 vision projector and pinned
artifact URLs. Document installation and explicit variant selection.

Assisted-by: Codex:GPT-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-26 18:54:08 +02:00
localai-org-maint-botandEttore Di Giacinto a95da0a46c fix(gallery): read the models dir once per gallery listing (#12283)
The cached gallery listing refreshed each entry's installed flag with
one os.Stat per entry, under the cache's global write lock. A gallery
holds about 1,900 entries. On a models directory on SMB, one refresh
took about 14s. The listing and every row's VRAM estimate run this
refresh, and the lock serialized them, so the models page took
minutes to load.

The installed check now lists the models directory once and looks up
each entry in that listing. On the same SMB share the listing takes
about 75ms. The answers match os.Stat: a symlink counts only when its
target exists, and names with a path separator still use os.Stat. The
listing runs before the lock is taken, so the lock covers only the
flag updates.

Concurrent callers on a cold cache now share one upstream load. Before
this, each caller fetched the gallery index and the configs itself.


Assisted-by: Claude:claude-opus-5-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-26 18:53:52 +02:00
localai-org-maint-botandEttore Di Giacinto 64c5670382 ci: bump Hugo from 0.146.3 to 0.166.0 (#12281)
The hugo-theme-relearn submodule was bumped to 9.1.x in #12096,
which requires Hugo >= 0.165.0. The pinned 0.146.3 broke the docs
site build with a template error in alias.html that could not
evaluate the Locale field on langs.Language.

Bump HUGO_VERSION to 0.166.0 (latest stable) to satisfy the
theme minimum and resolve the alias.html template error.

Assisted-by: nib:claude-sonnet-4.5 [bash] [read] [edit]

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-26 17:19:38 +02:00
e2b617104d chore: ⬆️ Update leejet/stable-diffusion.cpp to 2f886889e6e8b78738d6b87f7191f6018557c551 (#12274)
* ⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(stablediffusion-ggml): adapt to upstream tiling struct rename

Upstream commit 2f88688 renamed the sd_tiling_params_t fields from
tile_size_x/y to tile_size_w/h and rel_size_x/y to rel_size_w/h.
Update the gosd.cpp wrappers to match so the C++ backend compiles.

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-26 17:19:13 +02:00
localai-org-maint-botandlocalai-org-maint-bot dbdd2a4101 chore(gallery): add Hemmingway and remove invalid chat entry (#12278)
* fix(gallery): remove invalid Qwen-Image chat entry

The entry sends diffusion weights to llama.cpp as a chat model.
Remove it and document the existing image-generation alternatives.

Assisted-by: Codex:gpt-6

* feat(gallery): add Hemmingway-1 GGUF variants

Add Q4_K_M and Q8_0 builds for llama.cpp with embedded chat templates.
Record the upstream CC BY-NC 4.0 license and installation instructions.
Verify both SHA256 values against Hugging Face LFS metadata and headers.

Assisted-by: Codex:gpt-6

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-26 15:28:41 +02:00
Plamen K. Kosseffandlocalai-org-maint-bot 6cfc99196d feat(gallery): read metadata for system-path backends, enabling variant aliases (#12141)
* feat(gallery): read metadata for system-path backends, enabling variant aliases

Problem:
- ListSystemBackends only read metadata.json for user-managed backends;
  the system-path scan (LOCALAI_BACKENDS_SYSTEM_PATH) was a bare
  directory walk with Metadata hardcoded nil
- system-packaged backends (distro packages installing several
  accelerator builds of one backend) could not declare aliases or meta
  indirection at all, while gallery-installed backends could
- surfaced while packaging LocalAI for Gentoo: the packages install
  cpu-/rocm-/vulkan-audio-cpp as system backends aliased to audio-cpp,
  which the server ignored

Change:
- scan each root separately, clean the system collection against the
  user-managed one, merge, then build and resolve — precedence lives in
  one explicit step
- alias candidates carry their own metadata: the resolved alias entry
  can never pair one installation's executable with another's metadata,
  and it reports the chosen candidate's origin (IsSystem)
- deterministic resolution: entries build in sorted name order and
  candidates sort by name at the resolution site, independent of scan
  order

Precedence (user-managed always wins):
- a user-managed backend hides a same-named system backend entirely
- a user-managed variant takes over its whole alias family: the alias
  resolves among user-managed variants only and the system family's
  concrete names disappear — family versions move together, and a stale
  system variant may not work with newer models, so it must not stay
  reachable
- a system variant's alias never hijacks a name that exists as a
  user-managed backend

Tests: Ginkgo regressions for system-path aliasing, same-name hiding,
family takeover, and the full metadata permutation matrix of
cross-root name collisions (both directions, with and without
metadata on each side).

Docs: new "Backend Directory Format" section (run.sh, metadata.json,
alias resolution — previously undocumented for user-managed backends
too) and "System-Provided Backends" with the precedence rules.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>

* fix(gallery): preserve managed meta backends

A system alias can replace a user-managed meta backend during discovery.
Protect meta entries with the same precedence guard as concrete backends.
Add a regression test and clarify the documented precedence.

Assisted-by: Codex:GPT-6
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>

---------

Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-26 11:33:51 +00:00
localai-org-maint-botandmudler 9fa672faee chore: ⬆️ Update CrispStrobe/CrispASR to 6b78932d09765406ba0e0154d95bc6289246ceee (#12273)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:56 +02:00
localai-org-maint-botandmudler d270c2823c chore: ⬆️ Update PrismML-Eng/llama.cpp to adfffbe41b2cabcd51fff326ab045662265062bb (#12271)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:43 +02:00
localai-org-maint-botandmudler 8f29d5d271 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 1aaf7105be6e55a97fa4a9fd6f5bd362b08436dc (#12270)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:32 +02:00
localai-org-maint-botandmudler ded329854c chore: ⬆️ Update 0xShug0/audio.cpp to e79205f3e0083d04e812e1a4a376f71be97e9a22 (#12269)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:15 +02:00
mudler-agentandEttore Di Giacinto 42c58a5838 feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29 (#12247)
* feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29

Replace the model-specific SystemOne gRPC approach with a generic Score
RPC extension. The pre-existing Score RPC (previously unused by any
backend) now carries question_type and response_json fields:

- question_type="systemone" routes kev/laya decision-pipeline requests
  through the unified vllm_decide C ABI (v29), returning the full
  response JSON in response_json.
- question_type empty routes cua-s1-forms candidate scoring through the
  same vllm_decide ABI, returning CandidateScore probabilities.

The vllm-cpp backend's Score() method calls vllm_decide and dispatches
by architecture internally. The /v1/systemone HTTP endpoint checks
whether the model's backend supports Score; if so, it forwards the raw
request JSON and returns the backend response as-is. Other backends
fall through to the existing NER-based path.

This mirrors the vllm.cpp C ABI refactor (PR #3301) that replaced
vllm_systemone + vllm_score with a single vllm_decide function. The
purego bindings bump abiVersion from 27 to 29 and resolve vllm_decide
and vllm_decide_free symbols.

Also fixes validModelPath to accept cua-s1-forms.json and
rl_agent_config.json alongside config.json, matching the engine's
model_loader.cpp config-filename ordering.

AI-Assisted: true
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore: ⬆️ Update mudler/vllm.cpp to e28ec46c6 (fix macOS -Werror build)

Bumps vllm.cpp to e28ec46c6 which fixes a -Wnull-conversion error in
qwen3_5.cpp:12483 that broke the macOS Metal CI build under -Werror.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-25 23:27:27 +02:00
localai-org-maint-botandmudler 1b203fca9f chore(model-gallery): ⬆️ update checksum (#12275)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 23:27:02 +02:00
mudler-agentandEttore Di Giacinto 401c091053 feat(gallery): add nemo-speech-cpp diarization and ASR models (#12265)
Add five nemo-speech-cpp gallery entries for the diarization
capability introduced by the NeMo-Speech.cpp bump in #12257:

- nemo-speech-cpp-sortformer-diarization-v2: standalone streaming
  Sortformer 4-speaker diarization (nvidia/diar_streaming_sortformer_4spk-v2).
  Serves /v1/audio/diarization with known_usecases: [diarization].

- nemo-speech-cpp-nemotron-3.5-asr-streaming: standalone multilingual
  streaming ASR (nvidia/nemotron-3.5-asr-streaming-0.6b).

- nemo-speech-cpp-nemotron-3.5-asr-streaming-diarized: Nemotron ASR
  with the sortformer attached via the diar_model option, giving
  per-word speaker tags on /v1/audio/transcriptions.

- nemo-speech-cpp-parakeet-tdt-0.6b-v3: standalone multilingual ASR,
  25 languages (nvidia/parakeet-tdt-0.6b-v3).

- nemo-speech-cpp-parakeet-tdt-0.6b-v3-diarized: Parakeet v3 ASR
  with the sortformer attached via the diar_model option, giving
  per-word speaker tags on /v1/audio/transcriptions.

No backend code changes: the sortformer to familyDiarization mapping,
the diar_model option, and the MethodDiarize gRPC method already exist.

All five entries verified end-to-end against a running LocalAI instance.

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-25 18:10:26 +02:00
localai-org-maint-botandmudler c511b6dadf chore: ⬆️ Update ggml-org/llama.cpp to 84e76d8a23162eca70490da131945ebec1f09bf4 (#12258)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 14:59:38 +02:00
mudler-agentandEttore Di Giacinto 1bfd842b84 feat(gallery): add vllm-cpp entries for laya, cua-s1-forms, and gliner2.5 (#12240)
Add gallery entries for three vllm.cpp-backed models:

- laya-vllm-cpp: multilingual non-autoregressive System 1 decision
  model (ModernBERT-large, 421M params). Served via POST /v1/systemone.
  Weights from convaiinnovations/laya (Apache-2.0).

- cua-s1-forms-vllm-cpp: one-pass option scorer for GUI form filling.
  Served via POST /api/score. Weights from cua-ai/cua-s1-forms (MIT).

- gliner2.5-vllm-cpp: zero-shot named entity recognition and structured
  extraction (GLiNER2.5, mDeBERTa-v3-base, 287M params). Served via the
  token classification endpoint. Weights from fastino/gliner2.5-multi-v1
  (Apache-2.0).

All entries set backend: vllm-cpp and declare CPU+GPU tags.

Assisted-by: Maki:regolo/glm5.2 [maki]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-25 12:20:52 +02:00
localai-org-maint-botandmudler 1768dac662 chore: ⬆️ Update PrismML-Eng/llama.cpp to 842b1880415d6f508f03b789e5ce70194def7bfd (#12250)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:29:04 +02:00
leilei3167 a95ef3e7f3 fix(backend): re-probe MediaMarker after cold vision model load (#12254)
llama.cpp picks a new random media marker per backend process. LocalAI
cached the first probe on the model config and skipped later probes when
MediaMarker was non-empty, so after SINGLE_ACTIVE eviction/reload the
prompt still used the stale marker and mtmd_tokenize failed (0 markers
vs 1 bitmap).

Re-probe whenever the model was not already resident before Load, while
still skipping the RPC on warm cache hits.

Fixes #12246

Assisted-by: Cursor:composer-2.5

Signed-off-by: leilei3167 <imleilei123@gmail.com>
2026-09-25 09:16:33 +02:00
localai-org-maint-botandmudler 9a0afba315 chore(model-gallery): ⬆️ update checksum (#12256)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:12:52 +02:00
localai-org-maint-botandmudler 4668e9d141 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 8ab42195a05a9d48a3942b17568c1f3a876e133a (#12249)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:08:07 +02:00
localai-org-maint-botandmudler c678654d3e chore: ⬆️ Update 0xShug0/audio.cpp to 857de2366ed74bdb2c37f85259089e3a0a6b8cb0 (#12248)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:07:58 +02:00
localai-org-maint-botandmudler 1b961c0aca chore: ⬆️ Update ikawrakow/ik_llama.cpp to 20f7a72edd7049fe5a87eef2b5e9a50ae109ca4b (#12251)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:07:50 +02:00
localai-org-maint-botandmudler b010fd9b46 chore: ⬆️ Update ggml-org/whisper.cpp to d09f61a708f3487afa956ff578e60eae5e7a233c (#12253)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:04:37 +02:00
localai-org-maint-botandmudler b600f1b34d chore: ⬆️ Update leejet/stable-diffusion.cpp to b167b942f77ecb17e7f78e163a8c32ff7ac95c10 (#12255)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:57:57 +02:00
localai-org-maint-botandmudler a51bce57d6 chore: ⬆️ Update NVIDIA/NeMo-Speech.cpp to 97a15afa5caa9bce5baaa86c1184103877af4101 (#12257)
⬆️ Update NVIDIA/NeMo-Speech.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:18:18 +02:00
localai-org-maint-botandmudler 690a95afea chore: ⬆️ Update CrispStrobe/CrispASR to acc08e3bd3e5c17a3852115f3efa0e1ab30bc47a (#12259)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:18:04 +02:00