grpc-server.cpp forced params.cache_ram_mib = -1 (no limit) since #7009.
Since v4.3 kv_unified and cache_idle_slots are on by default, so every
distinct prompt now leaves its slot KV state in the host-side prompt
cache, and without a limit the backend grows until the host runs out of
memory.
Measured on gfx1151 (Strix Halo, 128 GB), llama-cpp backend, one request
at a time, 100 distinct prompts of ~2000 characters plus a fixed system
prompt, max_tokens 200:
model cache_ram RSS loaded -> after 100
gemma-4-26B-A4B (q8_0 KV) -1 (default) 1.4 GB -> 25.3 GB
Qwen3.6-35B-A3B (q8_0 KV) -1 (default) 1.1 GB -> 19.5 GB
gemma-4-26B-A4B -1, same prompt 100x 1.4 GB -> 1.6 GB
gemma-4-26B-A4B 4096 1.4 GB -> 5.4 GB (flat from
request 20 on, same latency)
Qwen3.6-35B-A3B 4096 1.1 GB -> 5.1 GB (flat)
The memory is not released when idle. In production a document
classification pass pushed the daily chat model to 34 GB RSS overnight.
Drop the override so llama.cpp's own default (8192 MiB) applies; the
cache_ram option still accepts -1 for users who want no limit. Update
both docs tables (the option reference and the prompt-cache table) and
note what -1 does.
Assisted-by: Claude:claude-opus-5-5
Assisted-by: Codex:GPT-6
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
utils.VerifyPath joins its argument onto the base path, so a path that
the caller already joined always passes. Several callers gave it joined
paths, and their checks could not fail:
- modeladmin (config view, patch, edit, pin and state): the config file
path from the loader. A config loaded from outside the models
directory (--models-config-file) could be pinned, and the pin wrote
the outside file. The patch and state paths stopped later, in the
mutation snapshot, with a different error.
- core/backend/tts.go: the model path joined onto the models path.
- The trellis2cpp and stablediffusion-ggml backends: option paths
(*_path) joined onto the model path. A "../" value outside the model
directory was accepted.
Add utils.VerifyResolvedPath for a full path. modeladmin and tts use
it. The backends now check the relative option value before they join
it. A rename in modeladmin checks the new relative name.
For models from a config file outside the models directory, the admin
API and web UI now return ErrPathNotTrusted for view, edit, pin, and
enable or disable. The docs describe this.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* sglang backend: pass through thinking_budget + require_reasoning
sglang's raw Engine.async_generate() API (which this backend calls
directly, bypassing sglang's own OpenAI server) supports a precise,
tokenizer-derived reasoning-length budget via
sampling_params["custom_params"]["thinking_budget"] plus
require_reasoning=True, gated behind --enable-strict-thinking. Neither
was reachable through LocalAI: this backend built sampling_params only
from a fixed field mapping (temperature, top_p, ...) with no custom_params
key, and never passed require_reasoning to async_generate at all.
- LoadModel now reads a model-level "thinking_budget" option (same
mechanism as the existing tool_parser/reasoning_parser options), and
_build_sampling_params adds it as custom_params.thinking_budget on
every request when configured.
- _new_reasoning_parser already derives, from the rendered prompt, whether
the model's chat template pre-opened a reasoning block (Qwen3-style
templates append <think> to the prompt instead of letting the model
emit it) -- the same signal sglang's own OpenAI server computes from
per-template config to decide require_reasoning. This backend has no
template manager, so it now returns that signal too and _predict
forwards it to async_generate(require_reasoning=...).
Verified against production (NVFP4, sm_121, Qwen3.6-35B-A3B) via a raw
Engine.async_generate() call bypassing this backend: 301 reasoning
tokens against a 300-token budget, clean completion, ~27s. Not yet
verified through this backend's own gRPC path end-to-end (no local
CUDA/sglang environment available here) -- existing + new unit tests in
test.py cover the pure-Python merge/passthrough logic only.
Scope note: require_reasoning is derived only from the existing
prompt-suffix heuristic, not sglang's full per-template
_get_reasoning_from_request decision tree (minimax-m3/hunyuan special
cases etc.) -- this backend has no template manager to evaluate that
tree against, and the prompt-suffix check is the one heuristic already
validated in this file (test_reasoning_parser_forced_when_template_prefills_think_tag).
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
* sglang backend: honour a model-level reasoning_default
A model YAML can already carry "parameters: reasoning_effort:", but that
value only reaches this backend when a *caller* sets it per request (the Go
side turns it into Metadata["enable_thinking"]). As a model-level default it
is silently dropped: a config reading "reasoning_effort: none" still produces
full reasoning on every request, so the config says one thing and the model
does another.
That gap is expensive in practice. On a self-hosted Qwen3.6-35B-A3B the
reasoning phase consumed the entire max_tokens budget before any content was
produced - 90% of code completions came back empty at max_tokens=768, and the
server log filled with "backend produced only reasoning, retrying". The
config looked like reasoning was off the whole time.
This adds "reasoning_default:off" (or ":on") on the same model-level
options: mechanism as thinking_budget. A per-request value always wins; the
default only fills in when the request is silent.
Measured on the stack above (sglang 0.5.20, NVFP4, GB10/sm_121) after
applying it:
default (nothing set) -> 0 chars reasoning, 27 tokens
"reasoning_effort": "none" -> 0 chars reasoning, 27 tokens
metadata enable_thinking=true -> capped at the 512-token thinking_budget,
541 tokens total, finish_reason stop
Tests: three cases added to backend/python/sglang/test.py covering the
default, per-request override in both directions, and the unconfigured case
(which must leave the template untouched).
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
* sglang backend: validate thinking_budget instead of crashing LoadModel
Addresses the review on this PR:
- `int(thinking_budget)` raised on values like "5000.0" or "abc" and took
LoadModel down. The option is now parsed by _parse_thinking_budget():
integral numbers in any spelling are accepted, anything else is ignored
with a warning on stderr.
- Zero and negative budgets are ignored with a warning instead of being
passed to sglang, where they have no defined meaning. Turning reasoning
off is what reasoning_default:off is for.
- A load-time warning when thinking_budget is set but enable_strict_thinking
is not in engine_args, since sglang then ignores the budget silently.
- Tests for integral spellings, unset, zero, negative, non-integer and the
strict-thinking warning.
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
* docs(sglang): explain reasoning options
Document the reasoning budget, strict-thinking requirement, and
precedence of request metadata over the model-level default.
Also note that the budget has to stay well below max_tokens (otherwise
it never triggers and the reply can end up empty), and that
POST /models/reload or a backend-only restart does not pick up changed
options; LocalAI itself has to be restarted.
Assisted-by: Codex:GPT-6
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
* docs(sglang): clarify configuration reloads
Distinguish rereading model configuration from updating a running backend. Keep the full LocalAI restart recommendation for changed reasoning options.
Assisted-by: Codex:GPT-6
* sglang backend: only pass require_reasoning when sglang supports it
Engine.async_generate() gained the require_reasoning keyword in sglang
0.5.13 and takes no **kwargs. The CPU profile builds v0.5.11 from source
and the other profiles only set a >=0.5.11 floor, so passing the keyword
unconditionally made every request fail with TypeError. Detect support
once at import time, as the file already does for sampling_seed.
enable_strict_thinking first appears in sglang 0.5.12; fix the comment.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
The proxy sent every chat and completion request with
context.Background, and the gRPC server gave rich backends no context
at all. When a client disconnected, or failover gave up on the target,
the upstream kept generating to the end, which costs tokens on a paid
or shared upstream. A silent upstream held the backend forever.
Add the optional AIModelRichContext interface. The gRPC server prefers
it and passes the call's context, like the Score and Rerank
extensions. The proxy implements it, so the upstream request ends with
the gRPC call.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
A streamed chat or completion reply never carried token counts: the
upstream LocalAI sends the usage trailer only when the request sets
stream_options.include_usage, and the proxy did not set it. Set it on
every streamed request.
Embeddings of tokenized input arrive in EmbeddingTokens with an empty
Embeddings string, so the proxy embedded an empty string. Send the
tokens as a token list instead.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* fix(kokoros): add missing animate3_d stub to Backend trait impl
#12095 added the Animate3D RPC to backend.proto, but the kokoros
service never got a matching method. The tonic-generated Backend trait
now requires it, so kokoros fails to build with E0046 whenever the
full backend matrix runs.
Return Unimplemented, as the other unsupported RPCs do.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* fix(kokoros): fill new Result fields with defaults
backend.proto added a metadata field to Result, so the struct literals
in the kokoros service no longer name every field and fail to compile.
Spread Default::default() into them, so later additive proto fields do
not break the build again.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* ⬆️ Update TheTom/llama-cpp-turboquant
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(turboquant): drop the upstreamed D512 patch
Upstream c0e227c guards D512 declarations, dispatch, and instances with
GGML_USE_HIP. This prevents the CUDA shared-memory overflow that our
patch addressed. The old patch now rejects the guarded source.
Remove the obsolete patch for the pinned a3d5603d revision. The remaining
patch series applies successfully, and the build-target test passes.
Assisted-by: Codex:gpt-6
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Code scanning flagged seven issues in the new backend:
- G115 text.go: tool-call indexes and tokenize lengths come from the
upstream server as int and were cast straight to int32. Add clampInt32 so
an absurd upstream value saturates instead of wrapping.
- G115 live.go: the int16 -> uint16 cast in PCM16 encoding is a deliberate
two's-complement reinterpretation of an already clamped sample; mark it
with #nosec and say so.
- G304 proxy.go, media.go, client.go: api_key_file comes from the model
config, and the media input and output paths are files core staged or
chose for the call. None are caller-supplied. Clean the paths and add
#nosec with that reason, as core/gallery and the sound classification
endpoint already do.
- G306 media.go: write generated media 0o600. Core runs as the same user
and serves the file itself.
gosec reports 0 issues for backend/go/localai-proxy and
core/services/failover. The G104 once reported for failover/prober.go is
no longer present.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The failover prober cuts upstream_url at /v1, but Load kept the path, so
an upstream_url ending in /v1 probed healthy while every request went to
/v1/v1/... and got a 404, which never trips the target. Cut the path at
/v1 in Load too, with a warning.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A caller that gives up before the ready ack now ends the call with Canceled
and closes the upstream socket, and setup has a 3 minute default bound when
request_timeout_seconds is unset, so a hung upstream cannot hold the call
or block failover. A closing session now waits only for a turn the upstream
committed (or is still speaking), not for turns it discarded. Audio held
before the ready ack is capped at 5 s, and NaN samples become silence.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
AudioTranscriptionLive opens <upstream>/v1/realtime?model=<realtime_pipeline>
as a transcription session with server VAD, forwards PCM as base64 PCM16
appends, and maps transcription deltas and completions to Delta/Eou. Closing
the send side waits briefly for an in-flight utterance, then sends the final
text. An upstream error, failed transcription or disconnect ends the stream
with Unavailable so failover reopens on the next target.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Detect/Depth/FaceVerify/FaceAnalyze forwarded the bare base64 core hands the
backend, but the upstream's own REST handlers only accept a URL or a
data:...;base64, string, so every real call 400'd. Wrap the payload as a
data URI (sniffing its MIME type) before sending it.
Depth requests for exports/dst now return Unimplemented: those files are
written to the upstream's own local disk and are unreachable from here, so
failover should move to a local target instead.
Generation replies that hand back a URL are now re-fetched by path only,
checked against the upstream's known generated-content prefixes, instead of
stripping the configured base as a literal string prefix — the old approach
broke (or silently trusted an arbitrary host) the moment the upstream
advertised a different base via LOCALAI_BASE_URL, a reverse proxy, or
X-Forwarded-Host.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The proxy now forwards TTS, streaming TTS, sound generation,
transcription (plain and streaming), diarization, VAD, sound
detection and audio transforms to the upstream LocalAI.
Streaming TTS passes the upstream WAV bytes through unchanged. A
streaming transcription that stops before its final frame, or sends
an error frame, fails with Unavailable instead of ending as a short
success. Transcription always sends diarize, because the upstream
treats a missing field as true. Audio transforms also download the
separation stems the upstream names and write them beside Dst.
Sound generation from a source clip returns Unimplemented, because
the REST endpoint has no field for the clip.
The multipart helper now takes repeated fields and several files.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Treating ResourceExhausted as a capability gap skipped the target
without counting a failure, so a target that stays rate limited or out
of memory kept its traffic. It is now an ordinary retryable failure:
the request moves to the next target and the exhausted one trips.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Rerank no longer sends top_n 0, which the upstream rejects. A
mid-stream upstream error frame now fails the call instead of ending
it as a short success. Temperature 0 is forwarded. An upstream 429
becomes ResourceExhausted, which failover skips like Unimplemented.
A localai-proxy config sends its own name upstream when upstream_model
is unset, and a chat proxy defaults to the tokenizer template so chat
reaches the upstream as messages.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* ⬆️ Update leejet/stable-diffusion.cpp
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(stablediffusion-ggml): adapt to upstream tiling struct rename
Upstream commit 2f88688 renamed the sd_tiling_params_t fields from
tile_size_x/y to tile_size_w/h and rel_size_x/y to rel_size_w/h.
Update the gosd.cpp wrappers to match so the C++ backend compiles.
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29
Replace the model-specific SystemOne gRPC approach with a generic Score
RPC extension. The pre-existing Score RPC (previously unused by any
backend) now carries question_type and response_json fields:
- question_type="systemone" routes kev/laya decision-pipeline requests
through the unified vllm_decide C ABI (v29), returning the full
response JSON in response_json.
- question_type empty routes cua-s1-forms candidate scoring through the
same vllm_decide ABI, returning CandidateScore probabilities.
The vllm-cpp backend's Score() method calls vllm_decide and dispatches
by architecture internally. The /v1/systemone HTTP endpoint checks
whether the model's backend supports Score; if so, it forwards the raw
request JSON and returns the backend response as-is. Other backends
fall through to the existing NER-based path.
This mirrors the vllm.cpp C ABI refactor (PR #3301) that replaced
vllm_systemone + vllm_score with a single vllm_decide function. The
purego bindings bump abiVersion from 27 to 29 and resolve vllm_decide
and vllm_decide_free symbols.
Also fixes validModelPath to accept cua-s1-forms.json and
rl_agent_config.json alongside config.json, matching the engine's
model_loader.cpp config-filename ordering.
AI-Assisted: true
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore: ⬆️ Update mudler/vllm.cpp to e28ec46c6 (fix macOS -Werror build)
Bumps vllm.cpp to e28ec46c6 which fixes a -Wnull-conversion error in
qwen3_5.cpp:12483 that broke the macOS Metal CI build under -Werror.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>