The proxy sent every chat and completion request with
context.Background, and the gRPC server gave rich backends no context
at all. When a client disconnected, or failover gave up on the target,
the upstream kept generating to the end, which costs tokens on a paid
or shared upstream. A silent upstream held the backend forever.
Add the optional AIModelRichContext interface. The gRPC server prefers
it and passes the call's context, like the Score and Rerank
extensions. The proxy implements it, so the upstream request ends with
the gRPC call.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
A streamed chat or completion reply never carried token counts: the
upstream LocalAI sends the usage trailer only when the request sets
stream_options.include_usage, and the proxy did not set it. Set it on
every streamed request.
Embeddings of tokenized input arrive in EmbeddingTokens with an empty
Embeddings string, so the proxy embedded an empty string. Send the
tokens as a token list instead.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* fix(kokoros): add missing animate3_d stub to Backend trait impl
#12095 added the Animate3D RPC to backend.proto, but the kokoros
service never got a matching method. The tonic-generated Backend trait
now requires it, so kokoros fails to build with E0046 whenever the
full backend matrix runs.
Return Unimplemented, as the other unsupported RPCs do.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* fix(kokoros): fill new Result fields with defaults
backend.proto added a metadata field to Result, so the struct literals
in the kokoros service no longer name every field and fail to compile.
Spread Default::default() into them, so later additive proto fields do
not break the build again.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* ⬆️ Update TheTom/llama-cpp-turboquant
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(turboquant): drop the upstreamed D512 patch
Upstream c0e227c guards D512 declarations, dispatch, and instances with
GGML_USE_HIP. This prevents the CUDA shared-memory overflow that our
patch addressed. The old patch now rejects the guarded source.
Remove the obsolete patch for the pinned a3d5603d revision. The remaining
patch series applies successfully, and the build-target test passes.
Assisted-by: Codex:gpt-6
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Code scanning flagged seven issues in the new backend:
- G115 text.go: tool-call indexes and tokenize lengths come from the
upstream server as int and were cast straight to int32. Add clampInt32 so
an absurd upstream value saturates instead of wrapping.
- G115 live.go: the int16 -> uint16 cast in PCM16 encoding is a deliberate
two's-complement reinterpretation of an already clamped sample; mark it
with #nosec and say so.
- G304 proxy.go, media.go, client.go: api_key_file comes from the model
config, and the media input and output paths are files core staged or
chose for the call. None are caller-supplied. Clean the paths and add
#nosec with that reason, as core/gallery and the sound classification
endpoint already do.
- G306 media.go: write generated media 0o600. Core runs as the same user
and serves the file itself.
gosec reports 0 issues for backend/go/localai-proxy and
core/services/failover. The G104 once reported for failover/prober.go is
no longer present.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The failover prober cuts upstream_url at /v1, but Load kept the path, so
an upstream_url ending in /v1 probed healthy while every request went to
/v1/v1/... and got a 404, which never trips the target. Cut the path at
/v1 in Load too, with a warning.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A caller that gives up before the ready ack now ends the call with Canceled
and closes the upstream socket, and setup has a 3 minute default bound when
request_timeout_seconds is unset, so a hung upstream cannot hold the call
or block failover. A closing session now waits only for a turn the upstream
committed (or is still speaking), not for turns it discarded. Audio held
before the ready ack is capped at 5 s, and NaN samples become silence.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
AudioTranscriptionLive opens <upstream>/v1/realtime?model=<realtime_pipeline>
as a transcription session with server VAD, forwards PCM as base64 PCM16
appends, and maps transcription deltas and completions to Delta/Eou. Closing
the send side waits briefly for an in-flight utterance, then sends the final
text. An upstream error, failed transcription or disconnect ends the stream
with Unavailable so failover reopens on the next target.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Detect/Depth/FaceVerify/FaceAnalyze forwarded the bare base64 core hands the
backend, but the upstream's own REST handlers only accept a URL or a
data:...;base64, string, so every real call 400'd. Wrap the payload as a
data URI (sniffing its MIME type) before sending it.
Depth requests for exports/dst now return Unimplemented: those files are
written to the upstream's own local disk and are unreachable from here, so
failover should move to a local target instead.
Generation replies that hand back a URL are now re-fetched by path only,
checked against the upstream's known generated-content prefixes, instead of
stripping the configured base as a literal string prefix — the old approach
broke (or silently trusted an arbitrary host) the moment the upstream
advertised a different base via LOCALAI_BASE_URL, a reverse proxy, or
X-Forwarded-Host.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The proxy now forwards TTS, streaming TTS, sound generation,
transcription (plain and streaming), diarization, VAD, sound
detection and audio transforms to the upstream LocalAI.
Streaming TTS passes the upstream WAV bytes through unchanged. A
streaming transcription that stops before its final frame, or sends
an error frame, fails with Unavailable instead of ending as a short
success. Transcription always sends diarize, because the upstream
treats a missing field as true. Audio transforms also download the
separation stems the upstream names and write them beside Dst.
Sound generation from a source clip returns Unimplemented, because
the REST endpoint has no field for the clip.
The multipart helper now takes repeated fields and several files.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Treating ResourceExhausted as a capability gap skipped the target
without counting a failure, so a target that stays rate limited or out
of memory kept its traffic. It is now an ordinary retryable failure:
the request moves to the next target and the exhausted one trips.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Rerank no longer sends top_n 0, which the upstream rejects. A
mid-stream upstream error frame now fails the call instead of ending
it as a short success. Temperature 0 is forwarded. An upstream 429
becomes ResourceExhausted, which failover skips like Unimplemented.
A localai-proxy config sends its own name upstream when upstream_model
is unset, and a chat proxy defaults to the tokenizer template so chat
reaches the upstream as messages.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* ⬆️ Update leejet/stable-diffusion.cpp
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(stablediffusion-ggml): adapt to upstream tiling struct rename
Upstream commit 2f88688 renamed the sd_tiling_params_t fields from
tile_size_x/y to tile_size_w/h and rel_size_x/y to rel_size_w/h.
Update the gosd.cpp wrappers to match so the C++ backend compiles.
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29
Replace the model-specific SystemOne gRPC approach with a generic Score
RPC extension. The pre-existing Score RPC (previously unused by any
backend) now carries question_type and response_json fields:
- question_type="systemone" routes kev/laya decision-pipeline requests
through the unified vllm_decide C ABI (v29), returning the full
response JSON in response_json.
- question_type empty routes cua-s1-forms candidate scoring through the
same vllm_decide ABI, returning CandidateScore probabilities.
The vllm-cpp backend's Score() method calls vllm_decide and dispatches
by architecture internally. The /v1/systemone HTTP endpoint checks
whether the model's backend supports Score; if so, it forwards the raw
request JSON and returns the backend response as-is. Other backends
fall through to the existing NER-based path.
This mirrors the vllm.cpp C ABI refactor (PR #3301) that replaced
vllm_systemone + vllm_score with a single vllm_decide function. The
purego bindings bump abiVersion from 27 to 29 and resolve vllm_decide
and vllm_decide_free symbols.
Also fixes validModelPath to accept cua-s1-forms.json and
rl_agent_config.json alongside config.json, matching the engine's
model_loader.cpp config-filename ordering.
AI-Assisted: true
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore: ⬆️ Update mudler/vllm.cpp to e28ec46c6 (fix macOS -Werror build)
Bumps vllm.cpp to e28ec46c6 which fixes a -Wnull-conversion error in
qwen3_5.cpp:12483 that broke the macOS Metal CI build under -Werror.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
The previous patch only removed DECL_FATTN_VEC_CASE_D512 for turbo2_0
and turbo3_0 V cache types. turbo4_0 also overflows shared memory
(0x10100 bytes > 0xc000 max), causing ptxas errors on CUDA 12/13.
Additionally, the previous patch was incomplete: it only removed the
template instantiations but not the dispatch calls in fattn.cu or the
extern declarations in fattn-vec.cuh. This caused linker errors
(undefined reference to ggml_cuda_flash_attn_ext_vec_case_d512).
This patch removes all three layers for all turbo V types:
- Template instance .cu files (DECL_FATTN_VEC_CASE_D512)
- Dispatch calls in fattn.cu (FATTN_VEC_CASE_D512)
- Extern declarations in fattn-vec.cuh (extern DECL_FATTN_VEC_CASE_D512)
Signed-off-by: mudler <mudler@localai.io>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>