A chat-only model now gets a 400 naming known_usecases: [systemone] instead
of a backend error. Configs declaring no usecases and token_classify models
stay allowed so existing laya and GLiNER setups keep working.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): load diarization and CED models and companions
Repin PARAKEET_VERSION to parakeet.cpp PR #75's head, which adds
parakeet_capi_model_kind (ABI v8). Bind the new diarization, sound
event and combined scene stream C symbols through the same
purego.Dlsym probe pattern already used for the batched JSON entry
point, so the backend still loads against an older libparakeet.so.
Load now classifies the loaded GGUF by role (ASR, diarization or
sound) via parakeet_capi_model_kind and can load up to two companion
models from Options[] (asr_model:, diarization_model:, sound_model:,
paths resolved against opts.ModelPath), verifying each companion's
kind and freeing every context opened so far on any failure. Free
releases the primary and every companion. AudioTranscription now
names the loaded role when it is not ASR instead of a generic model
not loaded error. The dynamic batcher starts only when an ASR context
ends up loaded, primary or companion.
This is groundwork only: the Diarize and SoundDetection RPCs and the
live scene stream that actually use these new roles land in later
commits.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): reset role fields on a failed companion load
loadRoles' freeLoaded only released the C contexts it had opened; it
left ctxPtr/diarCtx/tagCtx and companions pointing at those now-freed
contexts, so a later Free() on the same instance would double-free.
Zero all four alongside the CppFree calls.
Also route AudioTranscriptionStream and AudioTranscriptionLive through
notASRError when ctxPtr is unset but a diarization or sound model is
loaded, matching AudioTranscription: both used to return the generic
model-not-loaded error instead of naming the loaded role.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): add speaker diarization
Implement the Diarize RPC for the parakeet-cpp Go backend, wired to
Nemotron-3-Diarization through libparakeet.so's diarization C-API.
Plain diarization uses parakeet_capi_diarize_pcm; when include_text is
set and an ASR companion is loaded, parakeet_capi_transcribe_and_
diarize_json fills each segment's text instead. Speaker labels are the
decimal index, or "unknown" for -1 (no diarized speaker overlaps).
min_duration_off merges same-speaker segments across a short gap
before min_duration_on drops the segments still too short, then ids
are renumbered. num_speakers/min_speakers/max_speakers/clustering_
threshold have no Sortformer equivalent and are logged at debug
instead of rejected.
Verified against the real Nemotron-3-Diarization + parakeet-tdt_ctc-
110m checkpoints on the two_speakers.wav fixture: correct A-B-A-B
speaker segmentation and matching speaker-attributed transcripts.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): add sound event detection
Wire the SoundDetection RPC to the CED tagger context (p.tagCtx)
loaded by Task 1's role classification. It runs the whole clip
through a one-shot parakeet_capi_sound_stream_* session (window
10s, hop 10s, top_k set to the tagger's class count so every
drained window carries a full score list), averages each class's
score across the drained windows, sorts descending, then applies
the request's threshold and top_k (0 keeps every class).
No tagCtx returns FailedPrecondition; a libparakeet.so missing the
sound_stream symbols returns Unimplemented. Every C call runs under
engineMu, and the stream is always freed, even when a feed or drain
call fails partway through.
Verified against a real ced-tiny-q8_0.gguf on the rooster.wav demo
clip: "Chicken, rooster" tops the list at score 0.91.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): cancel sound detection mid-feed, shrink the lock
SoundDetection now checks ctx before each 10 s feed slice (mirroring
driver.go's feedSlices) and returns Canceled if the caller gave up,
so a long clip can be interrupted instead of feeding to completion
regardless. The stream is still freed on every path, cancellation
included.
Also narrow engineMu to the C calls: the drained JSON document is
now decoded after the lock is released, splitting soundStreamScores
into a locked soundStreamDrain (opts, begin, feed, drain, free) and
an unlocked json.Unmarshal.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): stream speaker and sound events during live transcription
Add two additive proto fields, LiveSpeakerSegment and LiveSoundEvent,
repeated on TranscriptLiveResponse. When a diarization or sound
companion model is loaded, AudioTranscriptionLive now runs a no-ASR
scene stream (parakeet_capi_scene_stream_begin) beside the ASR
streaming session, feeding it the same PCM slices and forwarding any
closed speaker or sound events alongside the matching ASR delta, or
on their own when a slice has no ASR output.
The scene stream is freed and reopened on a mid-stream Config reset,
flushed with is_last before the closing FinalResult, and degrades
gracefully (a warning, not an error) when begin or a later feed call
fails, so live transcription keeps working ASR-only. Existing live
behavior is unchanged when no companion is configured, and no scene
C call is made in that case.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): keep scene events off the ASR critical path in live
Emit each slice's ASR result right after the ASR feed, before the
scene feed for that slice runs, so a companion diarization/sound
model never adds scene compute latency in front of the delta or
<EOU> that drives realtime turn detection. Closed speakers/sounds go
out afterward as their own response, so a slice with both now
produces two responses, ASR first. The live feed log line now
reports ASR and scene wall time separately.
Re-check the diarization/sound contexts a scene stream was begun
with against the live contexts before every feed, under the same
lock: Free() can race between an ASR feed and the matching scene
feed and free the model the stream borrows. A mismatch now returns
without touching the C side. Freeing the stream itself stays
unconditional; the scene stream's destructor only releases its own
buffers and never touches the borrowed contexts.
Also recover a panicking stub inside the live test goroutine instead
of crashing the test binary, and reset the live decode-lag tracker on
a mid-stream config reset, matching what its own comment already
promised.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(realtime): surface live speaker and sound events
Carry the backend's closed speaker segments and sound events
(TranscriptLiveResponse fields 7/8) through LiveTranscriptionEvent
as LiveSpeakerSegment/LiveSoundEvent (nanoseconds mapped to
seconds), and forward them from the semantic_vad live path.
Each speaker segment emits
conversation.item.input_audio_transcription.segment with speaker,
start, end and empty text under the turn's item id. Each sound
event emits conversation.item.sound_detection with one tag
(label, score = peak, index) and the event's new optional
start/end seconds fields, omitted when unset so the existing
unary/windowed sound-detection path is unaffected.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(realtime): keep start/end on a zero-second transcription segment
ConversationItemInputAudioTranscriptionSegmentEvent.Start/End used
omitempty, so a speaker segment starting at 0.0s dropped its
"start" key. Nothing emitted this event before the live scene-event
path, so drop omitempty: the segment always carries real times.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(gallery): add parakeet-cpp diarization, CED and realtime scene models
Add gallery entries for the new parakeet-cpp capabilities: standalone
Nemotron-3-Diarization, the same paired with the Parakeet TDT+CTC
110M ASR model for speaker-attributed text, CED-Tiny and CED-Base
sound classifiers, and a realtime scene bundle combining the
streaming EOU ASR model with diarization and sound companions.
SHA256 taken from the Hub API; licenses from each model card
(openmdw-1.1 for Nemotron-3-Diarization, apache-2.0 for CED,
cc-by-4.0 for the Parakeet ASR models).
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs: document parakeet-cpp diarization, sound detection and live scene events
Cover the new parakeet-cpp capabilities across the feature pages:
Nemotron-3-Diarization as a diarization backend (with and without
speaker text, the ignored speaker-count hints, the Sortformer
voice-like-sound quirk), CED as a sound classification backend, the
asr_model/diarization_model/sound_model/diarization_latency companion
options, and the realtime live speaker/sound events (event shapes,
the speech-turn-only limitation, and using this or
pipeline.sound_detection but not both).
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(gallery): correct the realtime-scene license and wording nits
parakeet-cpp-realtime-scene mistakenly copied cc-by-4.0 from the
existing realtime_eou_120m-v1 entry; the model card lists the NVIDIA
open model license instead. Switch to the gallery's usual spelling
for that license and keep the diarization/CED licenses called out in
the description.
Also: audio-diarization.md now says getting per-segment text needs
both an asr_model companion and include_text=true on the request, and
audio-to-text.md's option table reads "Use on" (a pairing the loader
does not enforce) instead of "Allowed on".
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): reject a companion role that duplicates the primary's
loadRoles let a companion option (asr_model:/diarization_model:/
sound_model:) assign into a role field the primary already occupied,
for example asr_model: on an already-ASR primary. The companion's
context silently overwrote ctxPtr/diarCtx/tagCtx, and Free() only
walks those three fields, so the original primary context was never
freed again.
Reject a companion whose role the primary already holds before its
GGUF is even loaded, freeing everything loadRoles opened so far, the
same way a wrong-kind companion is already rejected.
Also warn, rather than silently fall through, when
parakeet_capi_model_kind reports PARAKEET_MODEL_KIND_NONE for a
successfully loaded primary; the primary is still treated as ASR,
matching today's behavior.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): cap live scene sound score retention
sceneBegin started the live diarization/sound companion stream with
the C API's default sound options, whose top_k keeps 5 scores per
window forever until drained. The live scene path never drains sound
scores (only the offline SoundDetection RPC does, with its own fresh
stream), so this window queue on the C side grew for the whole
session's lifetime.
Set opts.Sound.TopK = 0 before starting the scene stream: this
disables score retention while leaving sound event detection (onset/
offset), which the live path actually consumes, unaffected.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): merge diarization segments per speaker, harden Diarize
mergeCloseSegments only compared neighbors in the single start-sorted
segment list, so two same-speaker segments never merged once another
speaker's turn fell between them (A, B, A): the short B segment broke
the adjacency the merge relied on. Group segments by speaker first,
merge within each speaker's own start-ordered run, then re-sort the
result by start so interleaved speakers come back out in timeline
order.
Also harden Diarize's entry points the same way streamFeedDoc/
sceneFeed already are: diarizeCall re-checks p.diarCtx (and, on the
include_text path, p.ctxPtr) under engineMu right before the C call,
so a Free() racing between Diarize's own checks and the lock can no
longer reach the C side with a freed context. When the include_text
call returns NULL, last_error is now read from both contexts and
whichever came back non-empty is reported, since either side of the
pairing can be the one that failed. A WAV decode failure is reported
as InvalidArgument instead of an unwrapped/untyped error.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): harden SoundDetection's engine checks
soundStreamDrain ran every C call under engineMu but never re-checked
p.tagCtx there, so a Free() racing between SoundDetection's own
tagCtx==0 check and this lock could still reach the C side with a
freed context. Re-check p.tagCtx under the lock and return
ModelNotLoaded when it was cleared, mirroring diarizeCall's own
re-check. A WAV decode failure is now reported as InvalidArgument
instead of an unwrapped/untyped error.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* test(parakeet-cpp): cover a mid-session scene feed failure
feedSlicesScene already degrades gracefully when a scene feed call
fails mid-session: it frees the broken stream and carries the ASR-only
session forward. Add a spec covering that path end to end: the scene
stream is freed exactly once, later audio slices still produce ASR
responses, and no speaker/sound events appear before or after the
failure.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs: fix the parakeet-cpp companion role table and realtime scene docs
audio-to-text.md's companion option table read "Use on" with a note
that the loader did not enforce the pairing; it now rejects a
companion whose role duplicates the primary's, so restore the
"Allowed on" wording and describe the real enforcement.
openai-realtime.md's live speaker/sound section claimed a mid-stream
session.update resets the companion stream and that it flushes on
session close; neither happens, since the realtime core opens one
live stream (and so one scene stream) per speech turn and closes it
at that turn's commit, with no mid-stream Config in between. Document
that lifecycle instead, state precisely that start/end are seconds
from the start of the turn's own audio, and note that the diarization
model starts a fresh session every turn, so a speaker index is only
meaningful within one turn. The example sound tag ("Rooster", index
17) did not match any real CED label; index 17 in ced-tiny-q8_0.gguf
is "Baby laughter". Replaced with "Chicken, rooster" at its real
index, 99.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): use CED's real index for Chicken, rooster
The scene feed comment and the live test's canned document gave
"Chicken, rooster" index 365. In CED's AudioSet label list it is 99,
which is also what the realtime docs show.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(realtime): call the test event accessor
The scene-event tests range over a method instead of its returned slice.
Call the synchronized accessor so the OpenAI test package compiles.
Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(parakeet-cpp): pin parakeet.cpp master with sound events
mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and
#74 (the missing <algorithm> include that broke the image builds) are
on master now. Pin 6dea76a instead of the #75 PR head, and update the
header comment the bump bot reads.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(parakeet-cpp): pin parakeet.cpp with ced.cpp on main
parakeet.cpp #76 moved its ced.cpp submodule from the head of
localai-org/ced.cpp#3 (a branch-only commit) to ced.cpp main, where
#3 landed with an identical tree. Pin 623a968 so the image builds no
longer depend on that branch.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(transcription): carry speaker labels on words and streamed segments
A diarizing backend could label transcript segments, but two paths
dropped the label: TranscriptWord had no speaker field, so live
transcription words and word-level timestamps could not carry one, and
the stream=true transcript.text.done event left the speaker out of
its segments.
TranscriptWord gains an optional speaker (proto field 4, additive).
It flows through the live event and result mapping, the JSON word
output of the endpoint and the CLI, and transcript.text.done now
includes a segment's speaker when there is one. Empty labels are
omitted, so responses without diarization are unchanged.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
(cherry picked from commit 2f0049f979)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(importers): detect the parakeet.cpp diarization GGUF
The Nemotron-3-Diarization GGUFs are published in
mudler/parakeet-cpp-gguf as nemotron-3-diarization-<quant>.gguf. The
parakeet-cpp importer did not recognise that name, so a direct
`local-ai models import` of the file fell through to another importer.
A direct URL to the file now imports with the diarization usecase. A
repo import still picks ASR weights when the repo also ships the
diarization model, and falls back to the diarization weights only
when there are no others.
Ported from #12323.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(config): advertise diarization and sound detection for parakeet-cpp
The capability table listed parakeet-cpp as transcription only, though
the backend now answers Diarize (Nemotron-3-Diarization) and
SoundDetection (CED) depending on the model kind it loads.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): label transcript segments with the diarization companion
A diarization_model companion only fed live speaker events and
Diarize; /v1/audio/transcriptions ignored it.
With the companion attached and diarize=true (the OpenAI endpoint's
default), unary transcription now labels each segment with its
speaker and splits segments at speaker turns; with word timestamps
each word carries its speaker. The stream=true final result labels
each utterance with the speaker who said most of it. Both use the
checkpoint's own diarization over the whole clip, as NeMo's diarize()
does. Words take the speaker whose segments overlap them most, or the
nearest segment within 0.5 s, the same rule as parakeet.cpp's
speaker-attributed ASR.
Docs: the diarization_model row and a paragraph on transcript
speakers; Nemotron-3-Diarization handles up to 8 speakers.
Ported from #12323.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(realtime): speaker segments from committed-turn transcription
Speaker events reached a realtime session only from the live
semantic_vad path, which needs a cache-aware streaming transcription
model. Committed-turn transcription (server_vad, or any offline
model) always asked the backend for diarize=false and dropped the
segments' speakers.
pipeline.diarization (off by default) asks the transcription model for
speaker labels on each committed turn and emits every labelled segment
as a conversation.item.input_audio_transcription.segment event, with
its text, before the turn's completed event. It is opt-in because some
backends fail a diarization request they cannot serve.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(gallery): add parakeet-cpp-realtime-scene-tdt
parakeet-cpp-realtime-scene pairs the streaming EOU model with the
diarization and CED companions; its speaker and sound events need a
cache-aware streaming model. This entry does the same with Parakeet
TDT 0.6B v3 (multilingual, offline) for realtime under server_vad:
set it as both transcription and sound_detection and turn on
pipeline.diarization, and each committed turn gets speaker segments
and sound tags from one parakeet-cpp backend.
Files and sha256 match the Hub and are shared with the existing TDT v3,
diarization and CED-Tiny entries. A real-model spec checks the
combination on a clip with two speakers and a rooster: A-B-A-B speaker
turns, and "Chicken, rooster" among the sound tags. The test loader
now binds the sound entry points like main.go.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(gallery): add CED-Base variants of the parakeet-cpp scene models
parakeet-cpp-realtime-scene and parakeet-cpp-realtime-scene-tdt ship
with CED-Tiny. The -base variants use CED-Base (86M), which tags sounds
more confidently (on the rooster clip "Crowing" 0.65 against 0.49 for
Tiny).
Measured on CPU over a 37 s clip: the live diarization + sound stream
runs at 0.125 of real time with CED-Base against 0.103 with CED-Tiny,
because diarization dominates; sound detection per committed turn costs
0.031 against 0.005. The realtime docs list both and note that any CED
size works as sound_model.
Files and sha256 match the Hub and are shared with the existing
parakeet-cpp-ced-base entry. The TDT variant passes the real-model
scene spec with CED-Base (A-B-A-B speakers, "Chicken, rooster" found).
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(config): register pipeline.diarization in the config metadata
TestAllFieldsHaveRegistryEntries fails on the branch because the new
pipeline.diarization field has no registry entry. Add one so the model
editor shows it as a toggle next to the sound detection options.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
#12113 changed admission control to reject with 429 instead of 503.
failoverWriter only held back responses with status >= 500, so a 429
rejection reached the client and the chain never spilled to its next
target.
Hold 429 as well. An admission rejection is still flagged and spills
without tripping the target. Any other 429 is not retryable, so it is
released to the client unchanged.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* ci: bump Hugo from 0.146.3 to 0.166.0
The hugo-theme-relearn submodule was bumped to 9.1.x in #12096,
which requires Hugo >= 0.165.0. The pinned 0.146.3 broke the docs
site build with a template error in alias.html that could not
evaluate the Locale field on langs.Language.
Bump HUGO_VERSION to 0.166.0 (latest stable) to satisfy the
theme minimum and resolve the alias.html template error.
Assisted-by: nib:claude-sonnet-4.5 [bash] [read] [edit]
* feat(distributed): report worker version and show models in node inspector
Workers now send their LocalAI build version and git commit at
registration. The controller stores them on BackendNode and exposes
them through the existing node list/detail API responses.
The node inspector side pane now fetches and renders the list of
loaded models (name, state, in-flight) instead of showing only a
count, matching what the node detail page already displays.
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* fix(models): fallback to application config default context size (#12202)
Honor appConfig.ContextSize in /v1/models/capabilities when model context_size is unset.
* docs(models): explain context size fallback
Describe the application default used by capability discovery and
preserve the distinction between total context and per-request limits.
Assisted-by: Codex:GPT-6
* fix(models): apply the default context size only when context_size is unset
The request path applies the application default context size only
when a model leaves context_size unset. An explicit 0 or -1 falls
through to the backend fallback. The capabilities endpoint now does
the same, so it reports the value the backend uses.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* fix(functions): honor function_arguments_key when building the tool grammar
All four call sites of `Functions.ToJSONStructure(name, args string)` pass
`FunctionsConfig.FunctionNameKey` as *both* arguments, so
`FunctionArgumentsKey` never reaches the grammar generator.
`ToJSONStructure` writes the two properties into the same map:
property[nameKey] = FunctionName{Const: function.Name}
property[argsKey] = Argument{...}
When `nameKey == argsKey` the second assignment overwrites the first, so a
model configured with `function_name_key` gets a grammar carrying only the
arguments object -- the `{"const": "<function name>"}` constraint is gone and
the grammar can no longer express which function was called.
With `function_name_key: function`, the generated property set collapses from
{"function": {"const": "get_weather"}, "arguments": {...}}
to
{"function": {"type": "object", "properties": {...}}}
Setting only `function_arguments_key` is equally broken in the other
direction: the grammar keeps emitting `arguments` while `ParseFunctionCall`
(pkg/functions/parse.go) looks up the configured key, so the parsed call comes
back with its arguments empty.
The default configuration is unaffected -- with both keys empty
`ToJSONStructure` falls back to `name`/`arguments` for both parameters, which
is why this went unnoticed.
The existing `ToJSONStructure()` unit test already calls the helper with two
distinct keys, so only the call sites were wrong. Extend that test with a case
that keeps both custom keys distinct and asserts the two properties survive.
Signed-off-by: Anai-Guo <antai12232931@outlook.com>
* test(functions): cover configured grammar keys
Route grammar construction through FunctionsConfig so the regression test
covers the key wiring used by every endpoint.
Assisted-by: Codex:gpt-5
* chore: empty commit to trigger workflow approval
Signed-off-by: Tai An <antai12232931@outlook.com>
---------
Signed-off-by: Anai-Guo <antai12232931@outlook.com>
Signed-off-by: Tai An <antai12232931@outlook.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
fix(router): re-seed the knn corpus index when the vector store comes back empty
The corpus manager records a store as synced by file fingerprint and
embedding fingerprint. The local-store backend behind it is an in-memory
gRPC process the model loader may evict (active-backend cap, memory
pressure) or the idle watchdog may kill, and relaunch on the next
request — empty. The file is unchanged, so EnsureLoaded returned early
and the router went blind: every probe fell back with similarity 0 while
corpus/stats kept reporting the full count.
Measured on a production router (LOCALAI_MAX_ACTIVE_BACKENDS=6, four
resident models + two router stores): loading any further backend
evicted a store, and the idle watchdog killed both after 15 minutes;
/stores/find returned 0 hits against a 100-line corpus file whose stored
vectors matched fresh embeddings with cosine 1.000.
Two parts, because the knn classifier is built once and cached
(GetOrBuildClassifier), so the sync at build time is otherwise the only
one for the process lifetime:
- corpus.Manager remembers one vector it inserted (probe) and, on the
synced path, asks the live index for it. A miss means the index was
relaunched — fall through and re-seed from the file (no re-embedding).
- The router middleware wraps the knn classifier's store so every
lookup runs EnsureLoaded first; the loader gets the raw store, so its
probe never re-enters the wrapper. A sync error fails the lookup
closed, like the build-time load.
Specs: corpus package (relaunched empty store is re-seeded under an
unchanged file), middleware (relaunched index behind the cached
classifier is re-seeded instead of falling back; the spec is red without
the wrapper). The test fake now answers Search for inserted vectors.
Folds in the maintainer's follow-up (router-corpus-reseed-after-store-relaunch): reviewed and accepted.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
* feat: return 429 when backends are saturated
When backends are at capacity (per-model max_concurrent or the
process-wide --max-concurrent-backend-requests ceiling), the response
was 503. The OpenAI SDK, litellm, and most agent harnesses key on 429
for rate-limit backoff and treat 503 as a hard error.
Both saturation paths now return 429 with the existing Retry-After
header and type: "rate_limit_error" in the JSON body. The per-model
admission middleware keeps admission_rejected as the code field so
existing alerts that match on it still fire.
Non-saturation 503s are unchanged: model cold-loading (with progress
body), model-load failure cooldown, PII detector fail-closed, and
classifier unavailable. These mean "not ready" rather than "busy".
Assisted-by: AGENT:regolo/glm5.2 [TOOL]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix: return 503 when scheduler has no available nodes
When the scheduler cannot find any healthy node to serve a model —
all nodes are full and eviction cannot free a slot, or a node_selector
excludes every candidate — the error fell through to 500. A 500 tells
clients something is broken when the condition is transient and
retryable.
The router now wraps these errors with a new ErrNoAvailableNodes
sentinel. The HTTP error handler maps it to 503 via applyNoAvailableNodes,
following the same pattern as applyBackendAdmission (429). Unrelated
scheduler errors (DB timeouts, registry lookups) still return 500.
Three return sites are wrapped:
- resolveSelectorCandidates: selector matches zero healthy nodes
- scheduleNewModel eviction-busy: all models have in-flight requests
- scheduleNewModel eviction-failed: eviction itself errored
The existing scheduleAndLoad wrapper ("no available nodes: %w") preserves
the sentinel through the chain via errors.Is, as does ModelRouterAdapter.
Assisted-by: AGENT:regolo/glm5.2 [TOOL]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* test(http): use Ginkgo for admission tests
Replace forbidden testing.T calls with Ginkgo and Gomega so the lint
check accepts the admission handler tests.
Assisted-by: Codex:GPT-6 forbidigo
* fix(middleware): show the recorded status for admission rejections
The admission audit row now records 429, but the Middleware page still
printed a hard-coded 503. Read the status from the event, and update
the two package comments that still said 503.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Hardcoding size/size_vram as 0 made Ollama clients treat loaded models as
free. Prefer ModelFileName+ModelPath Stat when available, and omit size_vram
(and size) when the value is unknown instead of emitting literal zeros.
Resolve each listed model by its stored ID so tagged variants use their
own weights.
Fixes#11969
Signed-off-by: lei_lei <imleilei123@gmail.com>
A chain is checked for nested chains when it is saved, but not when
one of its targets is later edited into a chain. Requests then served
the inner chain's config as the target, which has no backend and
triggers backend auto-detection.
Mark such a target missing so no plan picks it, and skip it without a
trip in the HTTP path when it is pinned.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Transcription and diarization checked response_format only after the
backend had run, and returned a plain error for an unknown value.
Failover counts a plain error as a target failure, so one request with
a bad response_format ran the backend on every target of a chain and
tripped all of them.
Check the format before the backend runs and answer 400.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* feat(system): report per-model DRM VRAM
Expose optional resident device memory for local backend process trees.
Deduplicate DRM clients and omit unsupported or incomplete readings.
Document accounting limits and preserve a measured zero in JSON.
Closes#11970.
Assisted-by: Codex:gpt-6
* fix(system): document trusted procfs reads
Scope G304 annotations to paths built from the fixed procfs root,
integer process IDs, and kernel directory entries. These reads accept
no user-controlled path components.
Assisted-by: Codex:GPT-6 gosec
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Keep each message and reasoning item at its announced output index.
Include the answer in completed responses with reasoning or fallback
function calls, and retain reasoning supplied through backend deltas.
Add regression coverage for stream indices, final output, plain text,
and automatic tool parsing.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Partial JSON parsing heals a name-only chunk into a tool call. The
stream emits that call with empty arguments and skips later chunks.
Require complete JSON before emitting terminal tool-call events.
Preserve complete calls before an unfinished trailing call, and count
only actual tool calls. Add split-chunk regression tests and docs.
Refs #11635. The non-streaming report remains unconfirmed.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Review asked for constants instead of the stage literals ("vad",
"transcription", "llm", "tts", "sound_detection") passed to resolveStage,
stageCall and isChainStage, so the uses can be cross-checked. Add
PipelineStage* constants next to the Pipeline type in core/config: the
names match its yaml keys, and core/backend (preload roles) and the openai
realtime endpoint both need them.
Use them in realtime_model.go (stage routing and preload roles),
realtime.go (the voice_recognition preload role) and core/backend
preload.go. model_failover events take their stage from the stageChains
keys, so they now carry the constants too. The failover tests use the
constants for inputs and keep literal wire values in their event
assertions.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Review asked for constants instead of the "cloud-proxy" and
"localai-proxy" literals. Add CloudProxyBackend and LocalAIProxyBackend next
to the other backend-name constants in pkg/model (WhisperBackend,
TransformersBackend, ...), which core/config already imports, and use them
in every production check: the proxy options builder, IsRemoteProxy,
IsCloudProxyBackendPassthrough, the PII defaults, the localai-proxy backend
hook and loader warning, and the PII middleware metadata in the routes and
the in-process MCP client.
The proxy options builder now calls IsRemoteProxy() instead of repeating
the two-backend check, so the set of proxy backends is defined once.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
TranscriptEndpoint, DiarizationEndpoint and SoundClassificationEndpoint
returned the raw c.FormFile("file") error. Echo turns a non-HTTPError into
a 500, so a request with no multipart boundary or no file field looked like
a server fault. This became visible in tests/e2e once the suite registered
a transcription-capable model (lp-transcription) at runtime: the
default-model middleware then resolves a model, the request reaches the
handler, and "should return mocked transcription" got a 500 in some spec
orders (seed 1790493709).
Read the upload through a small uploadedFile helper that maps any
FormFile failure to 400 with the field name and parser reason. Server-side
failures after that (temp dir, file create, copy) stay 500. The image and
LocalAI upload endpoints already return 400 here.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The non-OpenAI endpoints (depth, detection, face_*, voice_*, images,
video, 3d) map a backend's gRPC Unimplemented to an echo 501 without the
gRPC status. The retry loop did not see a capability gap, and IsRetryable
is false for 501, so the client got 501 and the next target was never
tried. Treat a returned or written 501 as a capability gap.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Editing a failover chain now shows its state, the target serving it,
and a per-target health table under the editor header. The strip reads
GET /api/failover and follows /api/failover/events, with a 15 s re-list
to cover SSE reconnect gaps. Admins can pin a target or unpin the chain
after a confirmation.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Treating ResourceExhausted as a capability gap skipped the target
without counting a failure, so a target that stays rate limited or out
of memory kept its traffic. It is now an ordinary retryable failure:
the request moves to the next target and the exhausted one trips.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Rerank no longer sends top_n 0, which the upstream rejects. A
mid-stream upstream error frame now fails the call instead of ending
it as a short success. Temperature 0 is forwarded. An upstream 429
becomes ResourceExhausted, which failover skips like Unimplemented.
A localai-proxy config sends its own name upstream when upstream_model
is unset, and a chat proxy defaults to the tokenizer template so chat
reaches the upstream as messages.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Pass the API key environment lookup through ApplicationConfig to satisfy
core configuration lint. Keep credential resolution dynamic and exclude
the callback from serialization.
Handle the five close results reported by errcheck.
Assisted-by: Codex:gpt-6 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Transcription-only and sound-detection-only realtime sessions passed a
chain config straight to the model loader. It has no backend, so the
loader fell back to greedy backend auto-detection: slow, and ending in
an unhelpful error. Sound-only sessions are a main use of chains.
The stage routing of the full pipeline moves into a stageRouter that
both realtime model kinds embed. Every stage resolves to the chain's
active target at build time and goes through the failover plan per
call. The session sends failover events for any model with chain
stages, and restarts them when a transcription session.update swaps
the model.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A chain request reached a cloud-proxy target with the client's model,
the chain name, whenever the target set no upstream_model: passthrough
forwards the body's model and translate falls back to it. The upstream
answered 404, which neither retries nor trips, while the liveness
probe, which checks the target's own name, kept passing.
PrepareTarget now sets the upstream model of a remote target to
proxy.upstream_model or the target name, the same name the probe uses.
The request pipeline and realtime chain stages both call it.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The request path called HasChains on every request, and with no chains
it scanned the config source each time: the loader's lock plus a copy
and sort of every config, forever, on every installation without
chains. Sync now keeps an atomic flag and HasChains reads only that.
A chain added since the last sync is still served because Plan syncs
on a miss; only in-request retry waits for the next tick (at most 1s).
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A stage that names a chain is resolved on every call, so a switch keeps
the session and its conversation. Clients get localai.model.failover
events at session start and on every switch.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
An admission rejection or a disabled target moves the request to the
next target through Attempt.Skip, which records no failure. A 4xx
response no longer counts as a success. Requests skip body recording
when no chain is configured, and stop it once the model is known not
to be a chain.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The retry wraps SetModelAndConfig, so each attempt binds the request
again from a replayed body. A 5xx of a chain request is held back until
the handler returns, and a streamed response is never retried.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Warm local targets are pinned in the watchdog and preloaded. Switches
and target health are exported as metrics, skipped attempts as traces.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Reject chains whose targets are missing or are chains, at load and on
create or edit, and warn when the targets share no usecase.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29
Replace the model-specific SystemOne gRPC approach with a generic Score
RPC extension. The pre-existing Score RPC (previously unused by any
backend) now carries question_type and response_json fields:
- question_type="systemone" routes kev/laya decision-pipeline requests
through the unified vllm_decide C ABI (v29), returning the full
response JSON in response_json.
- question_type empty routes cua-s1-forms candidate scoring through the
same vllm_decide ABI, returning CandidateScore probabilities.
The vllm-cpp backend's Score() method calls vllm_decide and dispatches
by architecture internally. The /v1/systemone HTTP endpoint checks
whether the model's backend supports Score; if so, it forwards the raw
request JSON and returns the backend response as-is. Other backends
fall through to the existing NER-based path.
This mirrors the vllm.cpp C ABI refactor (PR #3301) that replaced
vllm_systemone + vllm_score with a single vllm_decide function. The
purego bindings bump abiVersion from 27 to 29 and resolve vllm_decide
and vllm_decide_free symbols.
Also fixes validModelPath to accept cua-s1-forms.json and
rl_agent_config.json alongside config.json, matching the engine's
model_loader.cpp config-filename ordering.
AI-Assisted: true
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore: ⬆️ Update mudler/vllm.cpp to e28ec46c6 (fix macOS -Werror build)
Bumps vllm.cpp to e28ec46c6 which fixes a -Wnull-conversion error in
qwen3_5.cpp:12483 that broke the macOS Metal CI build under -Werror.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* fix(gallery): verification follow-ups for oci:// galleries
Follow-ups from the post-merge review of #12238 and #12239.
Only a policy decision is a refusal now. cosignverify wraps
ErrPolicyRejected around a failed signature check, an identity or
source-repository mismatch, a not_before cutoff and a missing or
unparseable bundle. A TUF, registry or network failure during
verification, or a timeout, is an outage: the gallery falls back to the
copy verified under the current policy, as it does when the registry is
down.
An oci:// gallery with a verification block, or any oci:// gallery under
strict integrity, is no longer answered by an https://, github: or
file:// mirror. Such a mirror is ignored with a warning, because nothing
can check its signature. The index of an HTTP gallery, whose policy only
covers its backend images, is cached under the URL-only name again, so no
unchecked body is stored under a policy-keyed name.
The in-memory index cache key now includes the policy. After a runtime
policy change the index is fetched again, and entries with a relative url
install again.
The registry digest lookups after install and upgrade, and in the
upgrade check, run only for real registry references (new
URI.LooksLikeRegistryOCI), not for ollama:// or ocifile://.
The refusal message names strict integrity when that is the cause, and
the gallery name is no longer repeated.
Specs pin the URL-only cache name for galleries without a policy, a fixed
key for a fixed policy, and that every GalleryVerification field changes
the key. The docs describe refusal, outage, mirrors and strict integrity.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(gallery): reset listings on gallery changes, classify referrer outages
Review follow-ups for this PR.
The React UI lists from AvailableGalleryModelsCached, which is keyed by
nothing. A gallery change through the settings API or a
runtime_settings.json edit now drops that listing when the model or
backend gallery configuration differs. Before, the UI kept the old list,
with local paths into the old policy's tree, until the next background
refresh, or for good when the new policy refused the gallery.
In cosignverify, a referrer the registry fails to serve now makes the
lookup an outage whatever other referrers failed and in any order, since
the unread one may be the valid signature. An invalid policy (Validate in
NewVerifier, an unparseable not_before) is ErrPolicyRejected, because no
fetch can make it usable.
The docs say that only an oci:// gallery with a verification block skips
non-OCI mirrors, and list an unusable policy as a refusal.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* feat(vllm-cpp): add GLiNER2.5 NER via TokenClassify
Wire the vllm-cpp backend to the C ABI NER surface (vllm_gliner_ner,
ABI v27) so LocalAI can serve zero-shot named entity recognition through
the existing TokenClassify gRPC method.
backend.go: TokenClassify method on *VllmCpp calls vllm_gliner_ner with
the text and labels, copies the C-owned entity array into protobuf
TokenClassifyEntity messages, and frees the result.
govllmcpp.go: cNerEntity and cNerResult Go POD mirrors matching the C
structs; vllmGlinerNer and vllmNerResultFree purego bindings; abiVersion
bumped 26 -> 27.
options.go: ner_labels, ner_threshold, ner_max_width parsed from
engine_args.
pkg/grpc: ClassifyModel interface and TokenClassify server handler
(follows the Embedding locking pattern).
core/config: vllm-cpp backend declares MethodTokenClassify and
UsecaseTokenClassify.
docs/content/features/vllm-cpp.md: NER section documenting the
engine_args keys and the host-forward contract.
Assisted-by: MAKI:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(vllm-cpp): correct NER pointer lint directive
Use the govet directive for the C-owned NER array, matching the other
purego pointer conversions. The array remains valid until its deferred
free; the misspelled directive caused CI to flag this conversion.
Assisted-by: Codex:gpt-6 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(vllm-cpp): add kev-compatible SystemOne API endpoints
Add POST /v1/systemone, /v1/systemone/permute, and
/v1/systemone/separate to LocalAI, mirroring the kev project's
structured-extraction API. Each endpoint runs zero-shot NER over the
rendered state text and builds kev-compatible answers for three question
types: noul (binary entity presence), choice (pick one option), and
score (pick one level).
The TokenClassifyRequest proto gains a `repeated string labels` field so
each question can supply its own labels at inference time, and
TokenClassifier gains TokenClassifyWithLabels for per-call label
selection. The vllm-cpp backend uses request labels when non-empty,
falling back to configured ner_labels then the built-in defaults.
Helpers (renderState, softmax, choiceConfidence, scoreConfidence, r2)
are ported from kev/api.py and mirrored in vllm.cpp's api_server.cpp so
both servers produce the same answer shape.
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(vllm-cpp): suppress gosec G404 on seeded permutation RNG
The SystemOne permute endpoint uses math/rand with a caller-supplied
seed for reproducible option permutations, matching kev's random.seed.
gosec flags this as G404 (weak RNG). Add #nosec with a comment naming
the intent: this is reproducibility, not cryptography.
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(vllm-cpp): bump vllm.cpp pin to GLiNER2.5 merge commit
Advance VLLM_CPP_VERSION from f3cd97e to 5058268d, the commit that
landed GLiNER2.5 zero-shot NER support (PR #3224) in vllm.cpp. This
brings the DeBERTa v2 encoder, GLiNER2 boundary head, C ABI NER
functions, and server endpoints into the LocalAI vllm-cpp backend.
The ABI version (27) and Go struct mirrors already match.
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(vllm-cpp): use instruction text as NER label in SystemOne handler
The SystemOne handler was passing question IDs as NER labels for noul
questions and bare key names for choice questions, so the model never
matched any entities. Port the label mapping from vllm.cpp's
ParseSystemOneBody:
- noul: use the rendered instructions field (with instr alias) as the
NER label, not the question ID
- choice: use optionText(name, desc) — "name: description" or "name"
when the description is null/empty — not the bare key
- score: already correct (rendered criteria text)
- permute: shuffle indices and build parallel key/label arrays so the
NER call uses the optionText labels while the response is keyed by
the original option names
Also add the instructions field to the SystemOneQuestion schema struct
(accepted alongside the instr backward-compat alias).
Verified end-to-end against the real GLiNER2.5 model: noul questions
now find "Apple Inc. is" (organization, 0.999) and "Tim Cook is"
(person, 0.852) where they previously returned zero entities.
Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* fix(capabilities): report the per-request context with split KV slots
With parallel slots and kv_unified:false, llama.cpp gives each slot
n_ctx/n_parallel, padded up to a multiple of 256. /models/capabilities
still reported the full n_ctx. A client that budgets a request against
context_size then overflows at a fraction of it.
EffectiveRequestContextSize returns the per-slot size in that case and
the full context otherwise. With the unified KV cache, the grpc-server
default, one request may use all of n_ctx. The capabilities endpoint
and the router's prompt trimmer now use it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(openai): return an HTTP error when a stream fails before any chunk
A streamed chat request set the SSE headers, then waited for the
backend. When the backend failed before the first token, LocalAI sent
a 200 with a `data: {"error":...}` chunk and [DONE]. Clients that do
not parse error chunks saw an empty reply. cogito's LocalAI client was
one of them: nib users got "streaming decision produced no content"
instead of the context overflow that caused it.
Nothing has been written at that point, so the handler now returns the
error as a normal HTTP response. A failure after the first chunk keeps
the in-stream error chunk.
A prompt that exceeds the context is now a 400 on both paths, as in
the OpenAI API and llama-server, and no longer a 500. The message is
kept whole, because clients read the token counts from it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* test(e2e): check the error from closing the response body
golangci-lint's errcheck flags the unchecked resp.Body.Close in the
new pre-stream error helper.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
feat(kimodocpp): add API and backend request observability
Capture animation requests, phase timings, output metadata, and failures in traces. Record correct API error statuses and cover completed, running, failed, and disabled tracing.
Assisted-by: Codex:GPT-6
Signed-off-by: Richard Palethorpe <io@richiejp.com>
Track successful header authentication before allowing cross-site requests
to bypass CSRF checks. Arbitrary headers on unauthenticated servers and
cookie-authenticated requests no longer grant an exemption.
Share the production CSRF middleware with multipart tests, add regression
coverage for credential sources, and document the exemption behavior.
Assisted-by: Codex:gpt-6 golangci-lint
Signed-off-by: Richard Palethorpe <io@richiejp.com>
On a single-node install nothing in Operate listed the models loaded on
this machine or let an admin stop one. The System page that did was
retired in #11548, and its replacements (the Nodes workbench) only work
in distributed mode. The Nodes page also mis-detected single-node mode:
the cluster routes are not registered there, so /api/nodes answers 404,
but only 503 was treated as "distributed off", which sent every
single-node install to the empty worker-registration card. The rail hid
the entry anyway.
Nodes route on a single node becomes "This machine":
- the Nodes page's VRAM / RAM / CPU / models-disk gauges, fed from this
host by mapping /api/resources onto the worker heartbeat fields
- a memory bar splitting host RAM by running model
- a running-models table (backend, RSS, CPU share, uptime, PID) with
search, sorting, logs and a confirmed Stop
- the distributed setup behind an "Add machines" button
The Operate overview gains a "Running now" preview (heaviest five, with
Stop) on single node and a pointer to Nodes > Running models on a
cluster. The rail shows "This machine" in Runtime with a running count.
Backend, additive only:
- /system: each loaded model carries a `process` block (pid, rss_bytes,
memory_percent, cpu_percent, started_at). A sampler keeps one gopsutil
handle per PID so CPU is the share since the previous poll rather than
the lifetime average; it is omitted on the first reading.
- /api/resources: host `cpu` and models-path `disk`, the same readings
workers send in their heartbeat.
Also fixes the fleet tables widening the page on phones: the headers'
absolutely positioned sr-only labels escaped the scroll wrapper.
Assisted-by: Claude:claude-opus-5 [Playwright]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
An alias config is a pure redirect with no backend of its own, so the
capabilities listing described it from its stub: no capabilities, no
modalities, and the default 4096 context_size. Clients that size their
context budget from this endpoint (nib, for one) then compacted every
turn against a model that really serves 100k.
Resolve the alias and report the target's capabilities, modalities and
context_size under the alias's id. A dangling or chained alias now
reports no enrichment instead of defaults no model runs with.
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Report input tokens and frame-step output units in response metadata and
record them through the existing usage accounting pipeline. Preserve the
accounting rule and model-specific dimensions as JSON without extending
the gRPC schema for each modality.
Expose animation usage only under metadata.usage, validate counts before
recording, and document the response contract and loaded-model location.
Add coverage for transport, defaults, failures, persistence, and recording
requests once with statistics enabled or disabled.
Assisted-by: Codex:GPT-6
Signed-off-by: Richard Palethorpe <io@richiejp.com>
The global ::selection used --color-primary-light (14% primary), which
composites to about 1.1:1 against the dark page ground — selected text
was nearly indistinguishable from unselected. Give selection its own
token in both palettes and align the CodeMirror themes with the same
strengths.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
Expose a negative_prompt string parameter on the image endpoints,
matching Stable Diffusion WebUI / vLLM-Omni conventions. When both the
negative_prompt parameter and a '|'-suffixed negative prompt in the
main prompt are present, they are joined with a comma so callers can
keep a global negative prompt in negative_prompt and add per-image
negative tags after '|'.
Assisted-by: Pi: DeepSeek V4 Pro
Signed-off-by: Fedor Zuev <Fedor.Zuev@gmail.com>
* fix(vulkan): preserve host ICD discovery for packaged backends
Add bundled Mesa manifests through VK_ADD_DRIVER_FILES instead of replacing the system driver list. Merge inherited and model-specific additive paths while preserving explicit operator overrides, with regression coverage.
Assisted-by: Codex:gpt-5 golangci-lint
Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* feat(3d): add Kimodo CPU and Vulkan animation backend
Introduce a distinct animation capability and model-described 3D operations, with a typed /3d/animate API, RPC transport, distributed media staging, permissions, and tracing.
Add a persistent kimodo.cpp adapter, skeleton GLB export, CPU/Vulkan packages, model and backend galleries, importer support, CI builds, and documentation. Adapt Studio inputs to each model and provide real-time skeleton playback, seeking, and history.
Cover backend validation, packaging, API behavior, importer inventories, distributed staging, and Studio workflows. Validate real-model CPU/Vulkan generation and deploy the integration to the local QA instance.
Assisted-by: Codex:gpt-5 golangci-lint
Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* feat(kimodocpp): adopt monolithic encoders and resident inference
Update upstream for resident weights, packed execution paths, and cached motion graphs. Default to all 32 text layers while retaining configurable streaming and legacy bundle support.
Use monolithic Q8_0 encoders by default and offer all six published quantizations through the gallery and importer. Refresh pinned hashes, tests, and documentation; remove the obsolete thread patch and ensure cached source checkouts follow the upstream pin.
Validated CPU and Vulkan generation, lower-bit streaming, gallery/importer suites, packaging, lint, and cold/warm Studio generation on localai-dev.
Assisted-by: Codex:gpt-5 golangci-lint
Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Richard Palethorpe <io@richiejp.com>
---------
Signed-off-by: Richard Palethorpe <io@richiejp.com>