Commit Graph
2188 Commits
Author SHA1 Message Date
Ettore Di Giacinto 7c2df9b3dd fix(cachyllama): link SSD cache implementation
CachyLLaMA splits persistent prompt-cache support into separate source files. Include those implementations in the monolithic LocalAI gRPC adapter when the fork provides them so the ARM64 fallback build resolves the page-manager symbols.

Assisted-by: Codex:gpt-5 [Codex]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 03:02:23 +00:00
Ettore Di Giacinto b4e8d9b15b feat(backends): add CachyLLaMA support
Add CachyLLaMA as a GGUF-compatible llama.cpp fork backend with CPU and Vulkan builds on Linux and Metal on Apple silicon. Expose it through model import, document persistent SSD cache options, and wire CI path filtering and dependency updates.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 03:02:23 +00:00
61f4f67b75 sglang backend: pass through thinking_budget + require_reasoning (#12193)
* sglang backend: pass through thinking_budget + require_reasoning

sglang's raw Engine.async_generate() API (which this backend calls
directly, bypassing sglang's own OpenAI server) supports a precise,
tokenizer-derived reasoning-length budget via
sampling_params["custom_params"]["thinking_budget"] plus
require_reasoning=True, gated behind --enable-strict-thinking. Neither
was reachable through LocalAI: this backend built sampling_params only
from a fixed field mapping (temperature, top_p, ...) with no custom_params
key, and never passed require_reasoning to async_generate at all.

- LoadModel now reads a model-level "thinking_budget" option (same
  mechanism as the existing tool_parser/reasoning_parser options), and
  _build_sampling_params adds it as custom_params.thinking_budget on
  every request when configured.
- _new_reasoning_parser already derives, from the rendered prompt, whether
  the model's chat template pre-opened a reasoning block (Qwen3-style
  templates append <think> to the prompt instead of letting the model
  emit it) -- the same signal sglang's own OpenAI server computes from
  per-template config to decide require_reasoning. This backend has no
  template manager, so it now returns that signal too and _predict
  forwards it to async_generate(require_reasoning=...).

Verified against production (NVFP4, sm_121, Qwen3.6-35B-A3B) via a raw
Engine.async_generate() call bypassing this backend: 301 reasoning
tokens against a 300-token budget, clean completion, ~27s. Not yet
verified through this backend's own gRPC path end-to-end (no local
CUDA/sglang environment available here) -- existing + new unit tests in
test.py cover the pure-Python merge/passthrough logic only.

Scope note: require_reasoning is derived only from the existing
prompt-suffix heuristic, not sglang's full per-template
_get_reasoning_from_request decision tree (minimax-m3/hunyuan special
cases etc.) -- this backend has no template manager to evaluate that
tree against, and the prompt-suffix check is the one heuristic already
validated in this file (test_reasoning_parser_forced_when_template_prefills_think_tag).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: honour a model-level reasoning_default

A model YAML can already carry "parameters: reasoning_effort:", but that
value only reaches this backend when a *caller* sets it per request (the Go
side turns it into Metadata["enable_thinking"]). As a model-level default it
is silently dropped: a config reading "reasoning_effort: none" still produces
full reasoning on every request, so the config says one thing and the model
does another.

That gap is expensive in practice. On a self-hosted Qwen3.6-35B-A3B the
reasoning phase consumed the entire max_tokens budget before any content was
produced - 90% of code completions came back empty at max_tokens=768, and the
server log filled with "backend produced only reasoning, retrying". The
config looked like reasoning was off the whole time.

This adds "reasoning_default:off" (or ":on") on the same model-level
options: mechanism as thinking_budget. A per-request value always wins; the
default only fills in when the request is silent.

Measured on the stack above (sglang 0.5.20, NVFP4, GB10/sm_121) after
applying it:
  default (nothing set)          -> 0 chars reasoning, 27 tokens
  "reasoning_effort": "none"     -> 0 chars reasoning, 27 tokens
  metadata enable_thinking=true  -> capped at the 512-token thinking_budget,
                                    541 tokens total, finish_reason stop

Tests: three cases added to backend/python/sglang/test.py covering the
default, per-request override in both directions, and the unconfigured case
(which must leave the template untouched).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: validate thinking_budget instead of crashing LoadModel

Addresses the review on this PR:

- `int(thinking_budget)` raised on values like "5000.0" or "abc" and took
  LoadModel down. The option is now parsed by _parse_thinking_budget():
  integral numbers in any spelling are accepted, anything else is ignored
  with a warning on stderr.
- Zero and negative budgets are ignored with a warning instead of being
  passed to sglang, where they have no defined meaning. Turning reasoning
  off is what reasoning_default:off is for.
- A load-time warning when thinking_budget is set but enable_strict_thinking
  is not in engine_args, since sglang then ignores the budget silently.
- Tests for integral spellings, unset, zero, negative, non-integer and the
  strict-thinking warning.

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): explain reasoning options

Document the reasoning budget, strict-thinking requirement, and
precedence of request metadata over the model-level default.

Also note that the budget has to stay well below max_tokens (otherwise
it never triggers and the reply can end up empty), and that
POST /models/reload or a backend-only restart does not pick up changed
options; LocalAI itself has to be restarted.

Assisted-by: Codex:GPT-6
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): clarify configuration reloads

Distinguish rereading model configuration from updating a running backend. Keep the full LocalAI restart recommendation for changed reasoning options.

Assisted-by: Codex:GPT-6

* sglang backend: only pass require_reasoning when sglang supports it

Engine.async_generate() gained the require_reasoning keyword in sglang
0.5.13 and takes no **kwargs. The CPU profile builds v0.5.11 from source
and the other profiles only set a >=0.5.11 floor, so passing the keyword
unconditionally made every request fail with TypeError. Detect support
once at import time, as the file already does for sampling_seed.

enable_strict_thinking first appears in sglang 0.5.12; fix the comment.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 04:49:37 +02:00
dependabot[bot] eb97d5f549 chore(deps): bump grpcio from 1.83.1 to 1.84.0 in /backend/python/coqui (#12106)
Bumps [grpcio](https://github.com/grpc/grpc) from 1.83.1 to 1.84.0.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.83.1...v1.84.0)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.84.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 23:29:48 +02:00
mudler-agentandEttore Di Giacinto 6bda88cc88 fix(kokoros): add missing animate3_d stub to Backend trait impl (#12301)
* fix(kokoros): add missing animate3_d stub to Backend trait impl

#12095 added the Animate3D RPC to backend.proto, but the kokoros
service never got a matching method. The tonic-generated Backend trait
now requires it, so kokoros fails to build with E0046 whenever the
full backend matrix runs.

Return Unimplemented, as the other unsupported RPCs do.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

* fix(kokoros): fill new Result fields with defaults

backend.proto added a metadata field to Result, so the struct literals
in the kokoros service no longer name every field and fail to compile.
Spread Default::default() into them, so later additive proto fields do
not break the build again.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 21:38:15 +02:00
dependabot[bot] e6b2309b3f chore(deps): bump grpcio from 1.83.0 to 1.84.0 in /backend/python/transformers (#12102)
chore(deps): bump grpcio in /backend/python/transformers

Bumps [grpcio](https://github.com/grpc/grpc) from 1.83.0 to 1.84.0.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.83.0...v1.84.0)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.84.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:06:07 +02:00
dependabot[bot] de203c2fb5 chore(deps): update transformers requirement from >=5.15.1 to >=5.17.0 in /backend/python/transformers (#12105)
chore(deps): update transformers requirement

Updates the requirements on [transformers](https://github.com/huggingface/transformers) to permit the latest version.
- [Release notes](https://github.com/huggingface/transformers/releases)
- [Commits](https://github.com/huggingface/transformers/compare/v5.15.1...v5.17.0)

---
updated-dependencies:
- dependency-name: transformers
  dependency-version: 5.17.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:06:03 +02:00
dependabot[bot] 1cacecc460 chore(deps): update numpy requirement from >=2.5.2 to >=2.5.3 in /backend/python/transformers (#12244)
chore(deps): update numpy requirement in /backend/python/transformers

Updates the requirements on [numpy](https://github.com/numpy/numpy) to permit the latest version.
- [Release notes](https://github.com/numpy/numpy/releases)
- [Changelog](https://github.com/numpy/numpy/blob/main/doc/RELEASE_WALKTHROUGH.rst)
- [Commits](https://github.com/numpy/numpy/compare/v2.5.2...v2.5.3)

---
updated-dependencies:
- dependency-name: numpy
  dependency-version: 2.5.3
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:05:59 +02:00
dependabot[bot] 13a01e657a chore(deps): bump sentence-transformers from 5.7.0 to 6.1.0 in /backend/python/transformers (#12245)
chore(deps): bump sentence-transformers in /backend/python/transformers

Bumps [sentence-transformers](https://github.com/huggingface/sentence-transformers) from 5.7.0 to 6.1.0.
- [Release notes](https://github.com/huggingface/sentence-transformers/releases)
- [Commits](https://github.com/huggingface/sentence-transformers/compare/v5.7.0...v6.1.0)

---
updated-dependencies:
- dependency-name: sentence-transformers
  dependency-version: 6.1.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:05:55 +02:00
4bc5f292fe chore: ⬆️ Update TheTom/llama-cpp-turboquant to a3d5603d110bda29222d2011596cdc84d7fa532d (#12232)
* ⬆️ Update TheTom/llama-cpp-turboquant

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(turboquant): drop the upstreamed D512 patch

Upstream c0e227c guards D512 declarations, dispatch, and instances with
GGML_USE_HIP. This prevents the CUDA shared-memory overflow that our
patch addressed. The old patch now rejects the guarded source.

Remove the obsolete patch for the pinned a3d5603d revision. The remaining
patch series applies successfully, and the build-target test passes.

Assisted-by: Codex:gpt-6

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:05:49 +02:00
localai-org-maint-botandmudler 08827cfd5e chore: ⬆️ Update mudler/vllm.cpp to c3bebc357385990f721af66a3a6c69328dd4fc6c (#12252)
⬆️ Update mudler/vllm.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 21:05:45 +02:00
localai-org-maint-botandmudler a592e23778 chore: ⬆️ Update ggml-org/llama.cpp to 95887577ab5fead779581a7030a83c7752ff3234 (#12272)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 21:05:40 +02:00
localai-org-maint-botandmudler 4524765b9f chore: ⬆️ Update 0xShug0/audio.cpp to 94bd4656399180befc141b17bd6696bf84df0a9f (#12289)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:25:09 +02:00
localai-org-maint-botandmudler fc6df9efc3 chore: ⬆️ Update mudler/parakeet.cpp to 2bf88954dc628b32835734e2e9159550a75a1dc6 (#12291)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:24:52 +02:00
localai-org-maint-botandmudler 01017dcdd6 chore: ⬆️ Update CrispStrobe/CrispASR to 013ae1624dc40ecf059065d577180722439f804e (#12292)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:24:25 +02:00
localai-org-maint-botandmudler 9ea9277ee6 chore: ⬆️ Update ikawrakow/ik_llama.cpp to cdf232cc17e410e60c1bc3b85516c4a41199b662 (#12288)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:24:03 +02:00
e2b617104d chore: ⬆️ Update leejet/stable-diffusion.cpp to 2f886889e6e8b78738d6b87f7191f6018557c551 (#12274)
* ⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(stablediffusion-ggml): adapt to upstream tiling struct rename

Upstream commit 2f88688 renamed the sd_tiling_params_t fields from
tile_size_x/y to tile_size_w/h and rel_size_x/y to rel_size_w/h.
Update the gosd.cpp wrappers to match so the C++ backend compiles.

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-26 17:19:13 +02:00
localai-org-maint-botandmudler 9fa672faee chore: ⬆️ Update CrispStrobe/CrispASR to 6b78932d09765406ba0e0154d95bc6289246ceee (#12273)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:56 +02:00
localai-org-maint-botandmudler d270c2823c chore: ⬆️ Update PrismML-Eng/llama.cpp to adfffbe41b2cabcd51fff326ab045662265062bb (#12271)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:43 +02:00
localai-org-maint-botandmudler 8f29d5d271 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 1aaf7105be6e55a97fa4a9fd6f5bd362b08436dc (#12270)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:32 +02:00
localai-org-maint-botandmudler ded329854c chore: ⬆️ Update 0xShug0/audio.cpp to e79205f3e0083d04e812e1a4a376f71be97e9a22 (#12269)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:15 +02:00
mudler-agentandEttore Di Giacinto 42c58a5838 feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29 (#12247)
* feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29

Replace the model-specific SystemOne gRPC approach with a generic Score
RPC extension. The pre-existing Score RPC (previously unused by any
backend) now carries question_type and response_json fields:

- question_type="systemone" routes kev/laya decision-pipeline requests
  through the unified vllm_decide C ABI (v29), returning the full
  response JSON in response_json.
- question_type empty routes cua-s1-forms candidate scoring through the
  same vllm_decide ABI, returning CandidateScore probabilities.

The vllm-cpp backend's Score() method calls vllm_decide and dispatches
by architecture internally. The /v1/systemone HTTP endpoint checks
whether the model's backend supports Score; if so, it forwards the raw
request JSON and returns the backend response as-is. Other backends
fall through to the existing NER-based path.

This mirrors the vllm.cpp C ABI refactor (PR #3301) that replaced
vllm_systemone + vllm_score with a single vllm_decide function. The
purego bindings bump abiVersion from 27 to 29 and resolve vllm_decide
and vllm_decide_free symbols.

Also fixes validModelPath to accept cua-s1-forms.json and
rl_agent_config.json alongside config.json, matching the engine's
model_loader.cpp config-filename ordering.

AI-Assisted: true
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore: ⬆️ Update mudler/vllm.cpp to e28ec46c6 (fix macOS -Werror build)

Bumps vllm.cpp to e28ec46c6 which fixes a -Wnull-conversion error in
qwen3_5.cpp:12483 that broke the macOS Metal CI build under -Werror.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-25 23:27:27 +02:00
localai-org-maint-botandmudler c511b6dadf chore: ⬆️ Update ggml-org/llama.cpp to 84e76d8a23162eca70490da131945ebec1f09bf4 (#12258)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 14:59:38 +02:00
localai-org-maint-botandmudler 1768dac662 chore: ⬆️ Update PrismML-Eng/llama.cpp to 842b1880415d6f508f03b789e5ce70194def7bfd (#12250)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:29:04 +02:00
localai-org-maint-botandmudler 4668e9d141 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 8ab42195a05a9d48a3942b17568c1f3a876e133a (#12249)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:08:07 +02:00
localai-org-maint-botandmudler c678654d3e chore: ⬆️ Update 0xShug0/audio.cpp to 857de2366ed74bdb2c37f85259089e3a0a6b8cb0 (#12248)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:07:58 +02:00
localai-org-maint-botandmudler 1b961c0aca chore: ⬆️ Update ikawrakow/ik_llama.cpp to 20f7a72edd7049fe5a87eef2b5e9a50ae109ca4b (#12251)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:07:50 +02:00
localai-org-maint-botandmudler b010fd9b46 chore: ⬆️ Update ggml-org/whisper.cpp to d09f61a708f3487afa956ff578e60eae5e7a233c (#12253)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:04:37 +02:00
localai-org-maint-botandmudler b600f1b34d chore: ⬆️ Update leejet/stable-diffusion.cpp to b167b942f77ecb17e7f78e163a8c32ff7ac95c10 (#12255)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:57:57 +02:00
localai-org-maint-botandmudler a51bce57d6 chore: ⬆️ Update NVIDIA/NeMo-Speech.cpp to 97a15afa5caa9bce5baaa86c1184103877af4101 (#12257)
⬆️ Update NVIDIA/NeMo-Speech.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:18:18 +02:00
localai-org-maint-botandmudler 690a95afea chore: ⬆️ Update CrispStrobe/CrispASR to acc08e3bd3e5c17a3852115f3efa0e1ab30bc47a (#12259)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:18:04 +02:00
localai-org-maint-botandmudler 3a63699f5c chore: ⬆️ Update vllm-project/vllm cu130 wheel to 0.30.0 (#12214)
⬆️ Update vllm-project/vllm cu130 wheel

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 21:20:07 +02:00
localai-org-maint-botandmudler c6a574ef1b chore: ⬆️ Update vllm-metal (darwin) to v0.30.0 (#12225)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 21:19:30 +02:00
localai-org-maint-botandmudler 11b3b184ea chore: ⬆️ Update PrismML-Eng/llama.cpp to 0324c66521960d67aa7da8687fb1453a79a6565c (#12226)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 21:19:15 +02:00
localai-org-maint-botandmudler e143f14551 chore: ⬆️ Update ggml-org/whisper.cpp to a664346ea5c6dddff3e61a2b7b32dd4514613f50 (#12227)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 21:19:02 +02:00
localai-org-maint-botandmudler fdd19c7f76 chore: ⬆️ Update ggml-org/llama.cpp to d2e54583c7452353eb35d40431281f6ee984332f (#12228)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 21:18:47 +02:00
localai-org-maint-botandmudler df45e6cab8 chore: ⬆️ Update 0xShug0/audio.cpp to 9bdd1d908bbd128e9eb405f5a8e38d0defb84c72 (#12224)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 17:05:19 +02:00
localai-org-maint-botandmudler 8588020d77 chore: ⬆️ Update mudler/vllm.cpp to b24f8094cba9b4f02df71bcff8d41ddc7e88b4ef (#12229)
⬆️ Update mudler/vllm.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 09:32:52 +02:00
localai-org-maint-botandmudler 694e1ea3cc chore: ⬆️ Update CrispStrobe/CrispASR to 97a35a6e519fda1835f3c8353516384aa8710b8c (#12230)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 09:32:33 +02:00
localai-org-maint-botandmudler 91b462db24 chore: ⬆️ Update ikawrakow/ik_llama.cpp to f3d6e6e3020ddfebad60113845bf521620766da5 (#12233)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 09:05:08 +02:00
mudler-agentandEttore Di Giacinto 81eaca8768 fix(turboquant): extend D512 flash-attn patch to all turbo V types (#12234)
The previous patch only removed DECL_FATTN_VEC_CASE_D512 for turbo2_0
and turbo3_0 V cache types. turbo4_0 also overflows shared memory
(0x10100 bytes > 0xc000 max), causing ptxas errors on CUDA 12/13.

Additionally, the previous patch was incomplete: it only removed the
template instantiations but not the dispatch calls in fattn.cu or the
extern declarations in fattn-vec.cuh. This caused linker errors
(undefined reference to ggml_cuda_flash_attn_ext_vec_case_d512).

This patch removes all three layers for all turbo V types:
- Template instance .cu files (DECL_FATTN_VEC_CASE_D512)
- Dispatch calls in fattn.cu (FATTN_VEC_CASE_D512)
- Extern declarations in fattn-vec.cuh (extern DECL_FATTN_VEC_CASE_D512)

Signed-off-by: mudler <mudler@localai.io>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-24 08:01:28 +02:00
2ccda5ba92 chore: ⬆️ Update TheTom/llama-cpp-turboquant to 4deec5587b2963af00bdf80884f3337e02eb7d64 (#12154)
* ⬆️ Update TheTom/llama-cpp-turboquant

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(turboquant): patch D512 flash-attn shared memory overflow

turboquant 4deec55 added DECL_FATTN_VEC_CASE_D512 for TURBO2_0 and
TURBO3_0 V cache types. The D=512 kernel template with these types
allocates 65 KB of shared memory, exceeding the 48 KB GPU limit:

  ptxas error: Entry function uses too much shared data
  (0x10100 bytes, 0xc000 max)

Carry the fix as a patch under backend/cpp/turboquant/patches/ until
TheTom/llama-cpp-turboquant#386 is merged upstream.

TURBO4_0 (4-bit) does not overflow and is left unchanged.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-23 21:12:38 +00:00
9ad18c8c67 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 3ac485d0688fc684f5dcf2c95283b745220f9dcc (#12191)
* ⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(omnivoice-cpp): workaround GGML_SOURCE_DIR for nested builds

omnivoice.cpp 3ac485d changed from a relative add_subdirectory(ggml)
to CMAKE_SOURCE_DIR-based path resolution. CMAKE_SOURCE_DIR points to
the top-level project, not the current subdirectory, so when
omnivoice is consumed via add_subdirectory() the build fails:

  add_subdirectory given source ".../omnivoice-cpp/ggml"
  which is not an existing directory.

Set GGML_SOURCE_DIR to the correct path before add_subdirectory so
the upstream code picks it up from the cache. This is a workaround
until ServeurpersoCom/omnivoice.cpp#20 is merged upstream.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-23 22:13:58 +02:00
localai-org-maint-botandEttore Di Giacinto d71a22f1ae chore: ⬆️ Update mudler/vllm.cpp to d4738d241271b4d10134a6499f97337f20fcf8ce (#12175)
vllm.cpp d4738d2 bumped VLLM_ABI_VERSION from 26 to 27. The
abi-check target caught the mismatch: the Go struct mirrors in
govllmcpp.go still declared v26.

Update abiVersion, the header comment, and the test expectation
to v27. The struct layout did not change between the two versions,
so no offset adjustments are needed.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-23 21:21:53 +02:00
21c5495a99 feat(vllm-cpp): add GLiNER2.5 NER via TokenClassify (#12140)
* feat(vllm-cpp): add GLiNER2.5 NER via TokenClassify

Wire the vllm-cpp backend to the C ABI NER surface (vllm_gliner_ner,
ABI v27) so LocalAI can serve zero-shot named entity recognition through
the existing TokenClassify gRPC method.

backend.go: TokenClassify method on *VllmCpp calls vllm_gliner_ner with
the text and labels, copies the C-owned entity array into protobuf
TokenClassifyEntity messages, and frees the result.

govllmcpp.go: cNerEntity and cNerResult Go POD mirrors matching the C
structs; vllmGlinerNer and vllmNerResultFree purego bindings; abiVersion
bumped 26 -> 27.

options.go: ner_labels, ner_threshold, ner_max_width parsed from
engine_args.

pkg/grpc: ClassifyModel interface and TokenClassify server handler
(follows the Embedding locking pattern).

core/config: vllm-cpp backend declares MethodTokenClassify and
UsecaseTokenClassify.

docs/content/features/vllm-cpp.md: NER section documenting the
engine_args keys and the host-forward contract.

Assisted-by: MAKI:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): correct NER pointer lint directive

Use the govet directive for the C-owned NER array, matching the other
purego pointer conversions. The array remains valid until its deferred
free; the misspelled directive caused CI to flag this conversion.

Assisted-by: Codex:gpt-6 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vllm-cpp): add kev-compatible SystemOne API endpoints

Add POST /v1/systemone, /v1/systemone/permute, and
/v1/systemone/separate to LocalAI, mirroring the kev project's
structured-extraction API. Each endpoint runs zero-shot NER over the
rendered state text and builds kev-compatible answers for three question
types: noul (binary entity presence), choice (pick one option), and
score (pick one level).

The TokenClassifyRequest proto gains a `repeated string labels` field so
each question can supply its own labels at inference time, and
TokenClassifier gains TokenClassifyWithLabels for per-call label
selection. The vllm-cpp backend uses request labels when non-empty,
falling back to configured ner_labels then the built-in defaults.

Helpers (renderState, softmax, choiceConfidence, scoreConfidence, r2)
are ported from kev/api.py and mirrored in vllm.cpp's api_server.cpp so
both servers produce the same answer shape.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): suppress gosec G404 on seeded permutation RNG

The SystemOne permute endpoint uses math/rand with a caller-supplied
seed for reproducible option permutations, matching kev's random.seed.
gosec flags this as G404 (weak RNG). Add #nosec with a comment naming
the intent: this is reproducibility, not cryptography.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(vllm-cpp): bump vllm.cpp pin to GLiNER2.5 merge commit

Advance VLLM_CPP_VERSION from f3cd97e to 5058268d, the commit that
landed GLiNER2.5 zero-shot NER support (PR #3224) in vllm.cpp. This
brings the DeBERTa v2 encoder, GLiNER2 boundary head, C ABI NER
functions, and server endpoints into the LocalAI vllm-cpp backend.
The ABI version (27) and Go struct mirrors already match.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): use instruction text as NER label in SystemOne handler

The SystemOne handler was passing question IDs as NER labels for noul
questions and bare key names for choice questions, so the model never
matched any entities. Port the label mapping from vllm.cpp's
ParseSystemOneBody:

- noul: use the rendered instructions field (with instr alias) as the
  NER label, not the question ID
- choice: use optionText(name, desc) — "name: description" or "name"
  when the description is null/empty — not the bare key
- score: already correct (rendered criteria text)
- permute: shuffle indices and build parallel key/label arrays so the
  NER call uses the optionText labels while the response is keyed by
  the original option names

Also add the instructions field to the SystemOneQuestion schema struct
(accepted alongside the instr backward-compat alias).

Verified end-to-end against the real GLiNER2.5 model: noul questions
now find "Apple Inc. is" (organization, 0.999) and "Tim Cook is"
(person, 0.852) where they previously returned zero entities.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-23 12:31:28 +02:00
X c5216b36af fix(vllm-omni): remove invalid syntax in test.py (#12135)
Signed-off-by: hendrixx-cnc <tjhendrx@icloud.com>
2026-09-23 12:24:21 +02:00
pos-ei-don 0b09cff673 fix(sglang): support msgspec-based ServerArgs (sglang >= 0.5.20) (#12155)
sglang 0.5.20 moved its config tier from dataclasses to msgspec.Struct
(sgl-project/sglang#38753). _apply_engine_args validates engine_args keys
via dataclasses.fields(ServerArgs), which raises TypeError there. That
call runs on every LoadModel, so no model loads at all on the sglang
backend once sglang >= 0.5.20 is installed, and the error surfaces as a
generic "Unexpected <class 'TypeError'>" that does not name the cause.

Introspect both shapes: msgspec structs carry their field names in
__struct_fields__, so key validation and the close-match suggestion keep
working, and older dataclass-based sglang stays supported.

Adds a test that pins the msgspec path with a stand-in, so it is covered
regardless of which sglang version is installed.

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
2026-09-23 12:23:35 +02:00
Plamen K. Kosseff e7306a087a feat(audio-cpp): AUDIOCPP_DEFAULT_BACKEND fallback for models without a backend option (#12133)
Models whose options carry no explicit backend: open their session on the
CPU backend even in accelerator images. The gallery entries carry
backend:best since #11892; this covers hand-written model configurations
the same way, per deployment: the environment variable supplies the
fallback, an explicit backend: option always wins (merged beside the
existing threads and maingpu fallbacks), and validation reuses the
option parser.

Assisted-by: Claude:claude-fable-5

Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
2026-09-23 12:16:31 +02:00
localai-org-maint-botandmudler 2e279920b0 chore: ⬆️ Update CrispStrobe/CrispASR to 18d74132d22fa7c967181720310d4ba1df9b7bf0 (#12213)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-23 11:04:36 +02:00
localai-org-maint-botandmudler c93bd4da85 chore: ⬆️ Update leejet/stable-diffusion.cpp to c92d73c408515c94beef32161bb5960764fde7a0 (#12212)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-23 11:04:23 +02:00