Commit Graph
2208 Commits
Author SHA1 Message Date
localai-org-maint-botandmudler 84a2b209a0 chore: ⬆️ Update ggml-org/llama.cpp to 4da6337767f973e2b4d0797e5b323d77d8565e4a (#12318)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 21:41:46 +02:00
Ettore Di Giacinto 9c156656bd Merge PR #12285: feat(failover): serve a model name from a chain of local and remote targets
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 15:32:30 +00:00
localai-org-maint-botandmudler 102fbe0c78 chore: ⬆️ Update leejet/stable-diffusion.cpp to 3f8527a46c54ecf4cb4ed6003da8e8982283c73c (#12317)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 13:10:28 +02:00
localai-org-maint-botandmudler 197311f033 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to ead199a2bc4c53a57cac90095ae049a111d9e98d (#12316)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 13:10:11 +02:00
localai-org-maint-botandmudler b449ad3828 chore: ⬆️ Update 0xShug0/audio.cpp to 77491a33c589c53ff18add050095cf35647c8213 (#12315)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 13:09:57 +02:00
localai-org-maint-botandmudler 7f139c8add chore: ⬆️ Update CrispStrobe/CrispASR to ec98831d0776ec8a16ccaf93955693eb7ecfbec3 (#12314)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 13:09:32 +02:00
localai-org-maint-botandmudler adbff0a44c chore: ⬆️ Update ikawrakow/ik_llama.cpp to ed27bf7ed25e637692e89cd341d802522a2cee8a (#12313)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 13:09:15 +02:00
Stefan Walcz bc1d9924de fix(llama-cpp): keep llama.cpp's default cache_ram instead of no limit (#12297)
grpc-server.cpp forced params.cache_ram_mib = -1 (no limit) since #7009.
Since v4.3 kv_unified and cache_idle_slots are on by default, so every
distinct prompt now leaves its slot KV state in the host-side prompt
cache, and without a limit the backend grows until the host runs out of
memory.

Measured on gfx1151 (Strix Halo, 128 GB), llama-cpp backend, one request
at a time, 100 distinct prompts of ~2000 characters plus a fixed system
prompt, max_tokens 200:

  model                        cache_ram     RSS loaded -> after 100
  gemma-4-26B-A4B (q8_0 KV)    -1 (default)  1.4 GB -> 25.3 GB
  Qwen3.6-35B-A3B (q8_0 KV)    -1 (default)  1.1 GB -> 19.5 GB
  gemma-4-26B-A4B              -1, same prompt 100x  1.4 GB -> 1.6 GB
  gemma-4-26B-A4B              4096          1.4 GB -> 5.4 GB (flat from
                                             request 20 on, same latency)
  Qwen3.6-35B-A3B              4096          1.1 GB -> 5.1 GB (flat)

The memory is not released when idle. In production a document
classification pass pushed the daily chat model to 34 GB RSS overnight.

Drop the override so llama.cpp's own default (8192 MiB) applies; the
cache_ram option still accepts -1 for users who want no limit. Update
both docs tables (the option reference and the prompt-cache table) and
note what -1 does.


Assisted-by: Claude:claude-opus-5-5
Assisted-by: Codex:GPT-6

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-09-28 11:37:24 +02:00
mudler-agentandEttore Di Giacinto 50c284fdcc fix: make the remaining VerifyPath checks effective (#12326)
utils.VerifyPath joins its argument onto the base path, so a path that
the caller already joined always passes. Several callers gave it joined
paths, and their checks could not fail:

- modeladmin (config view, patch, edit, pin and state): the config file
  path from the loader. A config loaded from outside the models
  directory (--models-config-file) could be pinned, and the pin wrote
  the outside file. The patch and state paths stopped later, in the
  mutation snapshot, with a different error.
- core/backend/tts.go: the model path joined onto the models path.
- The trellis2cpp and stablediffusion-ggml backends: option paths
  (*_path) joined onto the model path. A "../" value outside the model
  directory was accepted.

Add utils.VerifyResolvedPath for a full path. modeladmin and tts use
it. The backends now check the relative option value before they join
it. A rename in modeladmin checks the new relative name.

For models from a config file outside the models directory, the admin
API and web UI now return ErrPathNotTrusted for view, edit, pin, and
enable or disable. The docs describe this.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 10:04:22 +02:00
61f4f67b75 sglang backend: pass through thinking_budget + require_reasoning (#12193)
* sglang backend: pass through thinking_budget + require_reasoning

sglang's raw Engine.async_generate() API (which this backend calls
directly, bypassing sglang's own OpenAI server) supports a precise,
tokenizer-derived reasoning-length budget via
sampling_params["custom_params"]["thinking_budget"] plus
require_reasoning=True, gated behind --enable-strict-thinking. Neither
was reachable through LocalAI: this backend built sampling_params only
from a fixed field mapping (temperature, top_p, ...) with no custom_params
key, and never passed require_reasoning to async_generate at all.

- LoadModel now reads a model-level "thinking_budget" option (same
  mechanism as the existing tool_parser/reasoning_parser options), and
  _build_sampling_params adds it as custom_params.thinking_budget on
  every request when configured.
- _new_reasoning_parser already derives, from the rendered prompt, whether
  the model's chat template pre-opened a reasoning block (Qwen3-style
  templates append <think> to the prompt instead of letting the model
  emit it) -- the same signal sglang's own OpenAI server computes from
  per-template config to decide require_reasoning. This backend has no
  template manager, so it now returns that signal too and _predict
  forwards it to async_generate(require_reasoning=...).

Verified against production (NVFP4, sm_121, Qwen3.6-35B-A3B) via a raw
Engine.async_generate() call bypassing this backend: 301 reasoning
tokens against a 300-token budget, clean completion, ~27s. Not yet
verified through this backend's own gRPC path end-to-end (no local
CUDA/sglang environment available here) -- existing + new unit tests in
test.py cover the pure-Python merge/passthrough logic only.

Scope note: require_reasoning is derived only from the existing
prompt-suffix heuristic, not sglang's full per-template
_get_reasoning_from_request decision tree (minimax-m3/hunyuan special
cases etc.) -- this backend has no template manager to evaluate that
tree against, and the prompt-suffix check is the one heuristic already
validated in this file (test_reasoning_parser_forced_when_template_prefills_think_tag).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: honour a model-level reasoning_default

A model YAML can already carry "parameters: reasoning_effort:", but that
value only reaches this backend when a *caller* sets it per request (the Go
side turns it into Metadata["enable_thinking"]). As a model-level default it
is silently dropped: a config reading "reasoning_effort: none" still produces
full reasoning on every request, so the config says one thing and the model
does another.

That gap is expensive in practice. On a self-hosted Qwen3.6-35B-A3B the
reasoning phase consumed the entire max_tokens budget before any content was
produced - 90% of code completions came back empty at max_tokens=768, and the
server log filled with "backend produced only reasoning, retrying". The
config looked like reasoning was off the whole time.

This adds "reasoning_default:off" (or ":on") on the same model-level
options: mechanism as thinking_budget. A per-request value always wins; the
default only fills in when the request is silent.

Measured on the stack above (sglang 0.5.20, NVFP4, GB10/sm_121) after
applying it:
  default (nothing set)          -> 0 chars reasoning, 27 tokens
  "reasoning_effort": "none"     -> 0 chars reasoning, 27 tokens
  metadata enable_thinking=true  -> capped at the 512-token thinking_budget,
                                    541 tokens total, finish_reason stop

Tests: three cases added to backend/python/sglang/test.py covering the
default, per-request override in both directions, and the unconfigured case
(which must leave the template untouched).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: validate thinking_budget instead of crashing LoadModel

Addresses the review on this PR:

- `int(thinking_budget)` raised on values like "5000.0" or "abc" and took
  LoadModel down. The option is now parsed by _parse_thinking_budget():
  integral numbers in any spelling are accepted, anything else is ignored
  with a warning on stderr.
- Zero and negative budgets are ignored with a warning instead of being
  passed to sglang, where they have no defined meaning. Turning reasoning
  off is what reasoning_default:off is for.
- A load-time warning when thinking_budget is set but enable_strict_thinking
  is not in engine_args, since sglang then ignores the budget silently.
- Tests for integral spellings, unset, zero, negative, non-integer and the
  strict-thinking warning.

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): explain reasoning options

Document the reasoning budget, strict-thinking requirement, and
precedence of request metadata over the model-level default.

Also note that the budget has to stay well below max_tokens (otherwise
it never triggers and the reply can end up empty), and that
POST /models/reload or a backend-only restart does not pick up changed
options; LocalAI itself has to be restarted.

Assisted-by: Codex:GPT-6
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): clarify configuration reloads

Distinguish rereading model configuration from updating a running backend. Keep the full LocalAI restart recommendation for changed reasoning options.

Assisted-by: Codex:GPT-6

* sglang backend: only pass require_reasoning when sglang supports it

Engine.async_generate() gained the require_reasoning keyword in sglang
0.5.13 and takes no **kwargs. The CPU profile builds v0.5.11 from source
and the other profiles only set a >=0.5.11 floor, so passing the keyword
unconditionally made every request fail with TypeError. Detect support
once at import time, as the file already does for sampling_seed.

enable_strict_thinking first appears in sglang 0.5.12; fix the comment.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 04:49:37 +02:00
dependabot[bot] eb97d5f549 chore(deps): bump grpcio from 1.83.1 to 1.84.0 in /backend/python/coqui (#12106)
Bumps [grpcio](https://github.com/grpc/grpc) from 1.83.1 to 1.84.0.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.83.1...v1.84.0)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.84.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 23:29:48 +02:00
Ettore Di Giacinto 10f7b1ca60 fix(localai-proxy): cancel the upstream request when the caller goes away
The proxy sent every chat and completion request with
context.Background, and the gRPC server gave rich backends no context
at all. When a client disconnected, or failover gave up on the target,
the upstream kept generating to the end, which costs tokens on a paid
or shared upstream. A silent upstream held the backend forever.

Add the optional AIModelRichContext interface. The gRPC server prefers
it and passes the call's context, like the Score and Rerank
extensions. The proxy implements it, so the upstream request ends with
the gRPC call.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:16:02 +00:00
Ettore Di Giacinto 19c1717bdb fix(localai-proxy): report streamed token usage and forward token embeddings
A streamed chat or completion reply never carried token counts: the
upstream LocalAI sends the usage trailer only when the request sets
stream_options.include_usage, and the proxy did not set it. Set it on
every streamed request.

Embeddings of tokenized input arrive in EmbeddingTokens with an empty
Embeddings string, so the proxy embedded an empty string. Send the
tokens as a token list instead.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 20:04:03 +00:00
Ettore Di Giacinto dae9a431e8 Merge remote-tracking branch 'origin/master' into feat/failover-chains
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-27 19:57:52 +00:00
mudler-agentandEttore Di Giacinto 6bda88cc88 fix(kokoros): add missing animate3_d stub to Backend trait impl (#12301)
* fix(kokoros): add missing animate3_d stub to Backend trait impl

#12095 added the Animate3D RPC to backend.proto, but the kokoros
service never got a matching method. The tonic-generated Backend trait
now requires it, so kokoros fails to build with E0046 whenever the
full backend matrix runs.

Return Unimplemented, as the other unsupported RPCs do.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

* fix(kokoros): fill new Result fields with defaults

backend.proto added a metadata field to Result, so the struct literals
in the kokoros service no longer name every field and fail to compile.
Spread Default::default() into them, so later additive proto fields do
not break the build again.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 21:38:15 +02:00
dependabot[bot] e6b2309b3f chore(deps): bump grpcio from 1.83.0 to 1.84.0 in /backend/python/transformers (#12102)
chore(deps): bump grpcio in /backend/python/transformers

Bumps [grpcio](https://github.com/grpc/grpc) from 1.83.0 to 1.84.0.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.83.0...v1.84.0)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.84.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:06:07 +02:00
dependabot[bot] de203c2fb5 chore(deps): update transformers requirement from >=5.15.1 to >=5.17.0 in /backend/python/transformers (#12105)
chore(deps): update transformers requirement

Updates the requirements on [transformers](https://github.com/huggingface/transformers) to permit the latest version.
- [Release notes](https://github.com/huggingface/transformers/releases)
- [Commits](https://github.com/huggingface/transformers/compare/v5.15.1...v5.17.0)

---
updated-dependencies:
- dependency-name: transformers
  dependency-version: 5.17.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:06:03 +02:00
dependabot[bot] 1cacecc460 chore(deps): update numpy requirement from >=2.5.2 to >=2.5.3 in /backend/python/transformers (#12244)
chore(deps): update numpy requirement in /backend/python/transformers

Updates the requirements on [numpy](https://github.com/numpy/numpy) to permit the latest version.
- [Release notes](https://github.com/numpy/numpy/releases)
- [Changelog](https://github.com/numpy/numpy/blob/main/doc/RELEASE_WALKTHROUGH.rst)
- [Commits](https://github.com/numpy/numpy/compare/v2.5.2...v2.5.3)

---
updated-dependencies:
- dependency-name: numpy
  dependency-version: 2.5.3
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:05:59 +02:00
dependabot[bot] 13a01e657a chore(deps): bump sentence-transformers from 5.7.0 to 6.1.0 in /backend/python/transformers (#12245)
chore(deps): bump sentence-transformers in /backend/python/transformers

Bumps [sentence-transformers](https://github.com/huggingface/sentence-transformers) from 5.7.0 to 6.1.0.
- [Release notes](https://github.com/huggingface/sentence-transformers/releases)
- [Commits](https://github.com/huggingface/sentence-transformers/compare/v5.7.0...v6.1.0)

---
updated-dependencies:
- dependency-name: sentence-transformers
  dependency-version: 6.1.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-27 21:05:55 +02:00
4bc5f292fe chore: ⬆️ Update TheTom/llama-cpp-turboquant to a3d5603d110bda29222d2011596cdc84d7fa532d (#12232)
* ⬆️ Update TheTom/llama-cpp-turboquant

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(turboquant): drop the upstreamed D512 patch

Upstream c0e227c guards D512 declarations, dispatch, and instances with
GGML_USE_HIP. This prevents the CUDA shared-memory overflow that our
patch addressed. The old patch now rejects the guarded source.

Remove the obsolete patch for the pinned a3d5603d revision. The remaining
patch series applies successfully, and the build-target test passes.

Assisted-by: Codex:gpt-6

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-27 21:05:49 +02:00
localai-org-maint-botandmudler 08827cfd5e chore: ⬆️ Update mudler/vllm.cpp to c3bebc357385990f721af66a3a6c69328dd4fc6c (#12252)
⬆️ Update mudler/vllm.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 21:05:45 +02:00
localai-org-maint-botandmudler a592e23778 chore: ⬆️ Update ggml-org/llama.cpp to 95887577ab5fead779581a7030a83c7752ff3234 (#12272)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 21:05:40 +02:00
localai-org-maint-botandmudler 4524765b9f chore: ⬆️ Update 0xShug0/audio.cpp to 94bd4656399180befc141b17bd6696bf84df0a9f (#12289)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:25:09 +02:00
localai-org-maint-botandmudler fc6df9efc3 chore: ⬆️ Update mudler/parakeet.cpp to 2bf88954dc628b32835734e2e9159550a75a1dc6 (#12291)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:24:52 +02:00
localai-org-maint-botandmudler 01017dcdd6 chore: ⬆️ Update CrispStrobe/CrispASR to 013ae1624dc40ecf059065d577180722439f804e (#12292)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:24:25 +02:00
localai-org-maint-botandmudler 9ea9277ee6 chore: ⬆️ Update ikawrakow/ik_llama.cpp to cdf232cc17e410e60c1bc3b85516c4a41199b662 (#12288)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-27 13:24:03 +02:00
Ettore Di Giacinto 53ce024efc fix(localai-proxy): clear the gosec findings
Code scanning flagged seven issues in the new backend:

- G115 text.go: tool-call indexes and tokenize lengths come from the
  upstream server as int and were cast straight to int32. Add clampInt32 so
  an absurd upstream value saturates instead of wrapping.
- G115 live.go: the int16 -> uint16 cast in PCM16 encoding is a deliberate
  two's-complement reinterpretation of an already clamped sample; mark it
  with #nosec and say so.
- G304 proxy.go, media.go, client.go: api_key_file comes from the model
  config, and the media input and output paths are files core staged or
  chose for the call. None are caller-supplied. Clean the paths and add
  #nosec with that reason, as core/gallery and the sound classification
  endpoint already do.
- G306 media.go: write generated media 0o600. Core runs as the same user
  and serves the file itself.

gosec reports 0 issues for backend/go/localai-proxy and
core/services/failover. The G104 once reported for failover/prober.go is
no longer present.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto fe8519d7b5 fix(localai-proxy): treat upstream_url as the server root, ignoring /v1
The failover prober cuts upstream_url at /v1, but Load kept the path, so
an upstream_url ending in /v1 probed healthy while every request went to
/v1/v1/... and got a 404, which never trips the target. Cut the path at
/v1 in Load too, with a warning.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto e66d87b887 fix(localai-proxy): bound live transcription setup and track in-flight turns precisely
A caller that gives up before the ready ack now ends the call with Canceled
and closes the upstream socket, and setup has a 3 minute default bound when
request_timeout_seconds is unset, so a hung upstream cannot hold the call
or block failover. A closing session now waits only for a turn the upstream
committed (or is still speaking), not for turns it discarded. Audio held
before the ready ack is capped at 5 s, and NaN samples become silence.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto c2375a7db8 feat(localai-proxy): bridge live transcription to the upstream realtime API
AudioTranscriptionLive opens <upstream>/v1/realtime?model=<realtime_pipeline>
as a transcription session with server VAD, forwards PCM as base64 PCM16
appends, and maps transcription deltas and completions to Delta/Eou. Closing
the send side waits briefly for an in-flight utterance, then sends the final
text. An upstream error, failed transcription or disconnect ends the stream
with Unavailable so failover reopens on the next target.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 37023939ad fix(localai-proxy): wrap bare base64 image inputs, block unreachable Depth exports, and validate download URLs
Detect/Depth/FaceVerify/FaceAnalyze forwarded the bare base64 core hands the
backend, but the upstream's own REST handlers only accept a URL or a
data:...;base64, string, so every real call 400'd. Wrap the payload as a
data URI (sniffing its MIME type) before sending it.

Depth requests for exports/dst now return Unimplemented: those files are
written to the upstream's own local disk and are unreachable from here, so
failover should move to a local target instead.

Generation replies that hand back a URL are now re-fetched by path only,
checked against the upstream's known generated-content prefixes, instead of
stripping the configured base as a literal string prefix — the old approach
broke (or silently trusted an arbitrary host) the moment the upstream
advertised a different base via LOCALAI_BASE_URL, a reverse proxy, or
X-Forwarded-Host.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto be9c039e65 feat(localai-proxy): serve image, video, 3D and vision APIs remotely
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 87fd7da531 feat(localai-proxy): serve speech, transcription and audio APIs remotely
The proxy now forwards TTS, streaming TTS, sound generation,
transcription (plain and streaming), diarization, VAD, sound
detection and audio transforms to the upstream LocalAI.

Streaming TTS passes the upstream WAV bytes through unchanged. A
streaming transcription that stops before its final frame, or sends
an error frame, fails with Unavailable instead of ending as a short
success. Transcription always sends diarize, because the upstream
treats a missing field as true. Audio transforms also download the
separation stems the upstream names and write them beside Dst.
Sound generation from a source clip returns Unimplemented, because
the REST endpoint has no field for the clip.

The multipart helper now takes repeated fields and several files.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 882f51fd62 fix(failover): trip a rate-limited or exhausted target instead of skipping it
Treating ResourceExhausted as a capability gap skipped the target
without counting a failure, so a target that stays rate limited or out
of memory kept its traffic. It is now an ordinary retryable failure:
the request moves to the next target and the exhausted one trips.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 6df1767133 fix(localai-proxy): keep rerank, stream errors and chat intact through the proxy
Rerank no longer sends top_n 0, which the upstream rejects. A
mid-stream upstream error frame now fails the call instead of ending
it as a short success. Temperature 0 is forwarded. An upstream 429
becomes ResourceExhausted, which failover skips like Unimplemented.

A localai-proxy config sends its own name upstream when upstream_model
is unset, and a chat proxy defaults to the tokenizer template so chat
reaches the upstream as messages.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 6d5c9600d7 feat(localai-proxy): add a backend that serves text APIs from a remote LocalAI
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
e2b617104d chore: ⬆️ Update leejet/stable-diffusion.cpp to 2f886889e6e8b78738d6b87f7191f6018557c551 (#12274)
* ⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(stablediffusion-ggml): adapt to upstream tiling struct rename

Upstream commit 2f88688 renamed the sd_tiling_params_t fields from
tile_size_x/y to tile_size_w/h and rel_size_x/y to rel_size_w/h.
Update the gosd.cpp wrappers to match so the C++ backend compiles.

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-26 17:19:13 +02:00
localai-org-maint-botandmudler 9fa672faee chore: ⬆️ Update CrispStrobe/CrispASR to 6b78932d09765406ba0e0154d95bc6289246ceee (#12273)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:56 +02:00
localai-org-maint-botandmudler d270c2823c chore: ⬆️ Update PrismML-Eng/llama.cpp to adfffbe41b2cabcd51fff326ab045662265062bb (#12271)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:43 +02:00
localai-org-maint-botandmudler 8f29d5d271 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 1aaf7105be6e55a97fa4a9fd6f5bd362b08436dc (#12270)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:32 +02:00
localai-org-maint-botandmudler ded329854c chore: ⬆️ Update 0xShug0/audio.cpp to e79205f3e0083d04e812e1a4a376f71be97e9a22 (#12269)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:15 +02:00
mudler-agentandEttore Di Giacinto 42c58a5838 feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29 (#12247)
* feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29

Replace the model-specific SystemOne gRPC approach with a generic Score
RPC extension. The pre-existing Score RPC (previously unused by any
backend) now carries question_type and response_json fields:

- question_type="systemone" routes kev/laya decision-pipeline requests
  through the unified vllm_decide C ABI (v29), returning the full
  response JSON in response_json.
- question_type empty routes cua-s1-forms candidate scoring through the
  same vllm_decide ABI, returning CandidateScore probabilities.

The vllm-cpp backend's Score() method calls vllm_decide and dispatches
by architecture internally. The /v1/systemone HTTP endpoint checks
whether the model's backend supports Score; if so, it forwards the raw
request JSON and returns the backend response as-is. Other backends
fall through to the existing NER-based path.

This mirrors the vllm.cpp C ABI refactor (PR #3301) that replaced
vllm_systemone + vllm_score with a single vllm_decide function. The
purego bindings bump abiVersion from 27 to 29 and resolve vllm_decide
and vllm_decide_free symbols.

Also fixes validModelPath to accept cua-s1-forms.json and
rl_agent_config.json alongside config.json, matching the engine's
model_loader.cpp config-filename ordering.

AI-Assisted: true
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore: ⬆️ Update mudler/vllm.cpp to e28ec46c6 (fix macOS -Werror build)

Bumps vllm.cpp to e28ec46c6 which fixes a -Wnull-conversion error in
qwen3_5.cpp:12483 that broke the macOS Metal CI build under -Werror.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-25 23:27:27 +02:00
localai-org-maint-botandmudler c511b6dadf chore: ⬆️ Update ggml-org/llama.cpp to 84e76d8a23162eca70490da131945ebec1f09bf4 (#12258)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 14:59:38 +02:00
localai-org-maint-botandmudler 1768dac662 chore: ⬆️ Update PrismML-Eng/llama.cpp to 842b1880415d6f508f03b789e5ce70194def7bfd (#12250)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:29:04 +02:00
localai-org-maint-botandmudler 4668e9d141 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 8ab42195a05a9d48a3942b17568c1f3a876e133a (#12249)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:08:07 +02:00
localai-org-maint-botandmudler c678654d3e chore: ⬆️ Update 0xShug0/audio.cpp to 857de2366ed74bdb2c37f85259089e3a0a6b8cb0 (#12248)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:07:58 +02:00
localai-org-maint-botandmudler 1b961c0aca chore: ⬆️ Update ikawrakow/ik_llama.cpp to 20f7a72edd7049fe5a87eef2b5e9a50ae109ca4b (#12251)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:07:50 +02:00
localai-org-maint-botandmudler b010fd9b46 chore: ⬆️ Update ggml-org/whisper.cpp to d09f61a708f3487afa956ff578e60eae5e7a233c (#12253)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:04:37 +02:00
localai-org-maint-botandmudler b600f1b34d chore: ⬆️ Update leejet/stable-diffusion.cpp to b167b942f77ecb17e7f78e163a8c32ff7ac95c10 (#12255)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:57:57 +02:00
localai-org-maint-botandmudler a51bce57d6 chore: ⬆️ Update NVIDIA/NeMo-Speech.cpp to 97a15afa5caa9bce5baaa86c1184103877af4101 (#12257)
⬆️ Update NVIDIA/NeMo-Speech.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:18:18 +02:00