Commit Graph
8326 Commits
Author SHA1 Message Date
Ettore Di Giacinto 246ca859d3 fix(vllm-cpp): annotate the hf_overrides config.json read for gosec
G304 flags reading a path built from a variable. The directory is the model
directory from the operator's own model config, not a request input, so it is
annotated the way the other backends do it.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 18:31:38 +00:00
mudler-agentandEttore Di Giacinto b540c3e1fd chore(vllm-cpp): bump to 967883486 (ABI v30), add hf_overrides and Tev1 entries, fix vllm-cpp gallery installs (#12379)
* chore(vllm-cpp): bump vllm.cpp to 967883486 (ABI v30)

Moves the pin from c3bebc357 to 967883486. On top of the Nimble decision
adapter and the Qwen3.5 vision-loader fix, this brings Tev1 on
/v1/systemone and vllm_decide (opt-in through a "Tev1Model" architecture
in config.json), a tokenizer/ subdirectory fallback so the Laya HF
snapshot loads as downloaded, a stop-token fix, a logprobs fix under async
scheduling and a pinned parakeet.cpp fetch for the diarization build.

ABI v30 only adds the diarization and speaker-attributed ASR entry
points; no existing struct or signature changed, so the purego mirrors
keep their layout and only abiVersion moves to 30. Between 4479dc99f and
967883486 vllm.h changed only in a comment.

v30 turns VLLM_CPP_WITH_DIARIZATION on by default. The fetch is pinned
now, but ON still downloads parakeet.cpp at configure time and links a
second ggml into libvllm for calls this backend never makes, so build
with the option off: the symbols stay present as refusing stubs.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

* feat(vllm-cpp): add the hf_overrides engine arg

vLLM parity: engine_args.hf_overrides is a JSON object of top-level
config.json keys merged over the model directory's own config.json. The
main use is opting a published checkpoint into an engine adapter its
config does not name, such as {"architectures": ["Tev1Model"]} on the
Tev1 snapshots, which declare Qwen3_5ForConditionalGeneration.

The C ABI has no override input and the engine reads config.json from
the directory it is given, so Load builds a private overlay directory:
the merged config.json plus a symlink to every other entry of the model
directory, and passes that to the engine. The download is never written.
Free, a failed load and the next Load remove the overlay.
validModelPath and the DFlash draft resolution still see the real
directory.

A value that is not a JSON object, a .gguf model or a directory without
config.json fails the load instead of being skipped like an unknown
engine_args key, because loading the unmodified config would serve a
different architecture than the one configured.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

* fix(gallery): nest vllm-cpp artifacts under overrides

artifacts: is a model-config key, and the installer reads model-config
keys only from overrides:. Five vllm-cpp entries (laya, gliner25-decide,
qwen3-vl-4b, cua-s1-forms and gliner2.5) declared it at the entry top
level, where it is silently dropped: the install reports success, writes
a config whose model is the bare HF repo id and downloads nothing, and
vllm-cpp (which does not infer artifacts) then fails the first load with
"model path not found".

Move each block under overrides:, and add a guard test that refuses a
top-level artifacts: key in the index.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

* feat(gallery): add Tev1 4B and 0.8B on vllm-cpp

Two decisions entries for Together AI's Tev1 checkpoints, pinned to the
current HF revisions. Tev1 is autoregressive: vllm.cpp answers
/v1/systemone by scoring the option letters, and the same engine still
serves chat completions. The published config.json names
Qwen3_5ForConditionalGeneration, so each entry sets
hf_overrides: {architectures: [Tev1Model]} to enable the decision route
without editing the download. known_usecases is [decisions] only, since
a declared decisions list is authoritative for reservation.

The descriptions state what was checked: agreement with transformers on
CPU over seven questions (4B 7/7, max probability difference 0.0004;
0.8B 6/7 with one near tie), CPU-only for the decision route, and a
fine-tune license the model card says is still being finalized, so no
license key is set.

The Decisions API page lists both entries, drops the note that Tev1
does not serve /v1/systemone and documents the 24-option limit (Ollama
allows 26).

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 20:22:49 +02:00
mudler-agent 1d7023be1a test(agentpool): pin the standalone agent contract (#12378)
Adds a Ginkgo contract suite for the standalone (LocalAGI-backed) agent service in core/services/agentpool, with a fake OpenAI-compatible LLM and a harness that boots a real AgentPoolService. It pins CRUD, per-user isolation, export and import, pause and resume, the chat and SSE event contract, status and observables, persistence and the raw-key agent lookup. No production code changes.

Specs that pin a known gap are named known gap or known defect, so later phases flip them on purpose.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-30 19:23:24 +02:00
Ettore Di Giacinto f378fe89d0 Merge pull request #12373 from mudler/feat/systemone-capability
feat: decisions usecase for decision models, with gallery tagging
2026-09-30 16:29:58 +00:00
localai-org-maint-botandmudler 7e0d836d21 chore: ⬆️ Update TheTom/llama-cpp-turboquant to bcb85fc3ae85efa0f5f392c6c880dfc524923860 (#12344)
⬆️ Update TheTom/llama-cpp-turboquant

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-30 16:53:22 +02:00
localai-org-maint-botandmudler 355c39968c chore: ⬆️ Update mudler/parakeet.cpp to 623a968bccbd2214588df398fcce687cd4218dea (#12347)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-30 16:53:09 +02:00
Ettore Di Giacinto 5613572f38 feat(systemone): validate requests and align the docs with Ollama's contract
All three routes now validate the request before it reaches a model: body
size (413 over 64 KiB), state, question count, blank ids, option and level
counts, and noul criteria keys. Forwarded decision requests skipped this
before, so a malformed question surfaced as a backend error.

The docs claimed the wire shape matches Ollama's. Field names and question
types do; confidence, error shape, keep_alive and state rendering differ, and
the docs now say so.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 14:31:50 +00:00
localai-org-maint-botandmudler 64291c6cd0 chore: ⬆️ Update 0xShug0/audio.cpp to ed96b7307c8daba2ebcf7912af928825f6b14cb9 (#12362)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-30 16:17:43 +02:00
localai-org-maint-botandmudler 82aaea3a4d chore: ⬆️ Update CrispStrobe/CrispASR to be202c472503a5c7f1d3e568c420865cad02f1c3 (#12363)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-30 16:17:29 +02:00
localai-org-maint-botandmudler c61e316d1f chore: ⬆️ Update ikawrakow/ik_llama.cpp to 0821d62a8b356bd1db3c6765551a30bfcc44a6de (#12364)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-30 16:17:15 +02:00
localai-org-maint-botandmudler da37a4d929 chore: ⬆️ Update localai-org/ced.cpp to 61dec2ab0106f2047ee40062a7075dbf08c523d0 (#12365)
⬆️ Update localai-org/ced.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-30 16:17:00 +02:00
Ettore Di Giacinto 70ce62901f refactor: name the capability decisions instead of systemone
The usecase describes what a model can do, and the category is the Decisions
API. SystemOne stays as the wire contract: the /v1/systemone routes, the
Score RPC question_type and the swagger tag are unchanged. The usecase,
flag, auth feature, UI label, gallery tags and docs page are now decisions.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 14:14:09 +00:00
mudler-agentandEttore Di Giacinto bd6863af81 fix(react-ui): extract text from PDF attachments in chat and home (#12374)
The React UI read every non-media attachment with file.text(). For a PDF
that decodes the binary bytes as UTF-8, so the model received raw
"%PDF ... stream ... endobj" noise instead of the document. The legacy
Alpine UI ran pdf.js; that step was not ported when the React UI
replaced it, but both file pickers still advertise .pdf.

Add a shared readAttachmentText helper that routes PDFs through
pdfjs-dist and reads other files as before. pdf.js and its worker load
on first use, so the main bundle does not grow. A PDF that cannot be
parsed or has no text layer (scanned, encrypted, damaged) is rejected
with a toast instead of being attached as an empty or garbage file.

Cover the chat and home paths with Playwright specs that build a real
PDF in the test.

Assisted-by: Claude Code:claude-sonnet-5-5 [playwright] [eslint]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 16:13:23 +02:00
mudler-agentandEttore Di Giacinto 4d0317db8d chore(deps): bump nib to v0.12.1 (#12372)
nib v0.12.0 called xlog.SetLogger with a *slog.Logger, which does not
build against the xlog v0.0.6 LocalAI uses. v0.12.1 fixes that
(mudler/nib#137).

nib's ApprovalMode is now a named string type, so the chat tests
compare against nibtypes.ApprovalAuto and ApprovalPrompt instead of
untyped strings, and run.go sets the exported constant.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 16:12:35 +02:00
localai-org-maint-botandmudler b8fdb50291 chore(model-gallery): ⬆️ update checksum (#12366)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-30 15:56:27 +02:00
Ettore Di Giacinto b3d65fd538 fix(systemone): route NER models to the NER path and refuse decision models on permute and separate
vllm_decide refuses NER architectures and the NER entry point refuses decision
architectures, so each model kind 500ed on half of the routes. A token_classify
model now goes to the NER path on /v1/systemone, and /permute and /separate
return 400 for decision models. Docs and instructions state which kind serves
which route.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 11:55:26 +00:00
Ettore Di Giacinto c7f278dd0d docs: document the systemone usecase and decisions API
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 11:37:48 +00:00
Ettore Di Giacinto c2af8d56ea feat(gallery): tag vllm-cpp entries by capability and add decision and vision models
laya declares the systemone usecase instead of chat. The gated Qwen3.6 27B
NVFP4 entries gain vision; the 35B-A3B entries gain it as experimental
because image input is not token-gated. Adds GLiNER2.5-Decide and
Qwen3-VL-4B. A guard test keeps capability tags and known_usecases in
agreement for every vllm-cpp entry.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 11:37:13 +00:00
Ettore Di Giacinto b4852d62d2 feat(systemone): register the decisions API on auth and instructions
Adds a default-on systemone route feature for the three /v1/systemone
routes and an /api/instructions area for them. No MCP tool is added: the
endpoints run inference and are not admin install/edit actions, and the
route-map test still passes.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 11:32:03 +00:00
Ettore Di Giacinto 36846466e4 feat(ui): show the systemone usecase on installed models
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 11:29:23 +00:00
Ettore Di Giacinto e03cf8dac8 feat(systemone): refuse models that do not declare the usecase
A chat-only model now gets a 400 naming known_usecases: [systemone] instead
of a backend error. Configs declaring no usecases and token_classify models
stay allowed so existing laya and GLiNER setups keep working.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 11:27:40 +00:00
Ettore Di Giacinto 84c83a70dd feat(config): add systemone usecase for decision models
Explicit-only, reserving usecase like score and token_classify: a declared
list is authoritative and the heuristic never guesses it. vllm-cpp now
lists systemone and vision as possible usecases.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-30 11:26:22 +00:00
2fa36e7147 feat(parakeet-cpp): speaker diarization, sound detection and live scene events (#12335)
* feat(parakeet-cpp): load diarization and CED models and companions

Repin PARAKEET_VERSION to parakeet.cpp PR #75's head, which adds
parakeet_capi_model_kind (ABI v8). Bind the new diarization, sound
event and combined scene stream C symbols through the same
purego.Dlsym probe pattern already used for the batched JSON entry
point, so the backend still loads against an older libparakeet.so.

Load now classifies the loaded GGUF by role (ASR, diarization or
sound) via parakeet_capi_model_kind and can load up to two companion
models from Options[] (asr_model:, diarization_model:, sound_model:,
paths resolved against opts.ModelPath), verifying each companion's
kind and freeing every context opened so far on any failure. Free
releases the primary and every companion. AudioTranscription now
names the loaded role when it is not ASR instead of a generic model
not loaded error. The dynamic batcher starts only when an ASR context
ends up loaded, primary or companion.

This is groundwork only: the Diarize and SoundDetection RPCs and the
live scene stream that actually use these new roles land in later
commits.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): reset role fields on a failed companion load

loadRoles' freeLoaded only released the C contexts it had opened; it
left ctxPtr/diarCtx/tagCtx and companions pointing at those now-freed
contexts, so a later Free() on the same instance would double-free.
Zero all four alongside the CppFree calls.

Also route AudioTranscriptionStream and AudioTranscriptionLive through
notASRError when ctxPtr is unset but a diarization or sound model is
loaded, matching AudioTranscription: both used to return the generic
model-not-loaded error instead of naming the loaded role.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): add speaker diarization

Implement the Diarize RPC for the parakeet-cpp Go backend, wired to
Nemotron-3-Diarization through libparakeet.so's diarization C-API.

Plain diarization uses parakeet_capi_diarize_pcm; when include_text is
set and an ASR companion is loaded, parakeet_capi_transcribe_and_
diarize_json fills each segment's text instead. Speaker labels are the
decimal index, or "unknown" for -1 (no diarized speaker overlaps).
min_duration_off merges same-speaker segments across a short gap
before min_duration_on drops the segments still too short, then ids
are renumbered. num_speakers/min_speakers/max_speakers/clustering_
threshold have no Sortformer equivalent and are logged at debug
instead of rejected.

Verified against the real Nemotron-3-Diarization + parakeet-tdt_ctc-
110m checkpoints on the two_speakers.wav fixture: correct A-B-A-B
speaker segmentation and matching speaker-attributed transcripts.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): add sound event detection

Wire the SoundDetection RPC to the CED tagger context (p.tagCtx)
loaded by Task 1's role classification. It runs the whole clip
through a one-shot parakeet_capi_sound_stream_* session (window
10s, hop 10s, top_k set to the tagger's class count so every
drained window carries a full score list), averages each class's
score across the drained windows, sorts descending, then applies
the request's threshold and top_k (0 keeps every class).

No tagCtx returns FailedPrecondition; a libparakeet.so missing the
sound_stream symbols returns Unimplemented. Every C call runs under
engineMu, and the stream is always freed, even when a feed or drain
call fails partway through.

Verified against a real ced-tiny-q8_0.gguf on the rooster.wav demo
clip: "Chicken, rooster" tops the list at score 0.91.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): cancel sound detection mid-feed, shrink the lock

SoundDetection now checks ctx before each 10 s feed slice (mirroring
driver.go's feedSlices) and returns Canceled if the caller gave up,
so a long clip can be interrupted instead of feeding to completion
regardless. The stream is still freed on every path, cancellation
included.

Also narrow engineMu to the C calls: the drained JSON document is
now decoded after the lock is released, splitting soundStreamScores
into a locked soundStreamDrain (opts, begin, feed, drain, free) and
an unlocked json.Unmarshal.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): stream speaker and sound events during live transcription

Add two additive proto fields, LiveSpeakerSegment and LiveSoundEvent,
repeated on TranscriptLiveResponse. When a diarization or sound
companion model is loaded, AudioTranscriptionLive now runs a no-ASR
scene stream (parakeet_capi_scene_stream_begin) beside the ASR
streaming session, feeding it the same PCM slices and forwarding any
closed speaker or sound events alongside the matching ASR delta, or
on their own when a slice has no ASR output.

The scene stream is freed and reopened on a mid-stream Config reset,
flushed with is_last before the closing FinalResult, and degrades
gracefully (a warning, not an error) when begin or a later feed call
fails, so live transcription keeps working ASR-only. Existing live
behavior is unchanged when no companion is configured, and no scene
C call is made in that case.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): keep scene events off the ASR critical path in live

Emit each slice's ASR result right after the ASR feed, before the
scene feed for that slice runs, so a companion diarization/sound
model never adds scene compute latency in front of the delta or
<EOU> that drives realtime turn detection. Closed speakers/sounds go
out afterward as their own response, so a slice with both now
produces two responses, ASR first. The live feed log line now
reports ASR and scene wall time separately.

Re-check the diarization/sound contexts a scene stream was begun
with against the live contexts before every feed, under the same
lock: Free() can race between an ASR feed and the matching scene
feed and free the model the stream borrows. A mismatch now returns
without touching the C side. Freeing the stream itself stays
unconditional; the scene stream's destructor only releases its own
buffers and never touches the borrowed contexts.

Also recover a panicking stub inside the live test goroutine instead
of crashing the test binary, and reset the live decode-lag tracker on
a mid-stream config reset, matching what its own comment already
promised.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(realtime): surface live speaker and sound events

Carry the backend's closed speaker segments and sound events
(TranscriptLiveResponse fields 7/8) through LiveTranscriptionEvent
as LiveSpeakerSegment/LiveSoundEvent (nanoseconds mapped to
seconds), and forward them from the semantic_vad live path.

Each speaker segment emits
conversation.item.input_audio_transcription.segment with speaker,
start, end and empty text under the turn's item id. Each sound
event emits conversation.item.sound_detection with one tag
(label, score = peak, index) and the event's new optional
start/end seconds fields, omitted when unset so the existing
unary/windowed sound-detection path is unaffected.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(realtime): keep start/end on a zero-second transcription segment

ConversationItemInputAudioTranscriptionSegmentEvent.Start/End used
omitempty, so a speaker segment starting at 0.0s dropped its
"start" key. Nothing emitted this event before the live scene-event
path, so drop omitempty: the segment always carries real times.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(gallery): add parakeet-cpp diarization, CED and realtime scene models

Add gallery entries for the new parakeet-cpp capabilities: standalone
Nemotron-3-Diarization, the same paired with the Parakeet TDT+CTC
110M ASR model for speaker-attributed text, CED-Tiny and CED-Base
sound classifiers, and a realtime scene bundle combining the
streaming EOU ASR model with diarization and sound companions.

SHA256 taken from the Hub API; licenses from each model card
(openmdw-1.1 for Nemotron-3-Diarization, apache-2.0 for CED,
cc-by-4.0 for the Parakeet ASR models).

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: document parakeet-cpp diarization, sound detection and live scene events

Cover the new parakeet-cpp capabilities across the feature pages:
Nemotron-3-Diarization as a diarization backend (with and without
speaker text, the ignored speaker-count hints, the Sortformer
voice-like-sound quirk), CED as a sound classification backend, the
asr_model/diarization_model/sound_model/diarization_latency companion
options, and the realtime live speaker/sound events (event shapes,
the speech-turn-only limitation, and using this or
pipeline.sound_detection but not both).

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): correct the realtime-scene license and wording nits

parakeet-cpp-realtime-scene mistakenly copied cc-by-4.0 from the
existing realtime_eou_120m-v1 entry; the model card lists the NVIDIA
open model license instead. Switch to the gallery's usual spelling
for that license and keep the diarization/CED licenses called out in
the description.

Also: audio-diarization.md now says getting per-segment text needs
both an asr_model companion and include_text=true on the request, and
audio-to-text.md's option table reads "Use on" (a pairing the loader
does not enforce) instead of "Allowed on".

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): reject a companion role that duplicates the primary's

loadRoles let a companion option (asr_model:/diarization_model:/
sound_model:) assign into a role field the primary already occupied,
for example asr_model: on an already-ASR primary. The companion's
context silently overwrote ctxPtr/diarCtx/tagCtx, and Free() only
walks those three fields, so the original primary context was never
freed again.

Reject a companion whose role the primary already holds before its
GGUF is even loaded, freeing everything loadRoles opened so far, the
same way a wrong-kind companion is already rejected.

Also warn, rather than silently fall through, when
parakeet_capi_model_kind reports PARAKEET_MODEL_KIND_NONE for a
successfully loaded primary; the primary is still treated as ASR,
matching today's behavior.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): cap live scene sound score retention

sceneBegin started the live diarization/sound companion stream with
the C API's default sound options, whose top_k keeps 5 scores per
window forever until drained. The live scene path never drains sound
scores (only the offline SoundDetection RPC does, with its own fresh
stream), so this window queue on the C side grew for the whole
session's lifetime.

Set opts.Sound.TopK = 0 before starting the scene stream: this
disables score retention while leaving sound event detection (onset/
offset), which the live path actually consumes, unaffected.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): merge diarization segments per speaker, harden Diarize

mergeCloseSegments only compared neighbors in the single start-sorted
segment list, so two same-speaker segments never merged once another
speaker's turn fell between them (A, B, A): the short B segment broke
the adjacency the merge relied on. Group segments by speaker first,
merge within each speaker's own start-ordered run, then re-sort the
result by start so interleaved speakers come back out in timeline
order.

Also harden Diarize's entry points the same way streamFeedDoc/
sceneFeed already are: diarizeCall re-checks p.diarCtx (and, on the
include_text path, p.ctxPtr) under engineMu right before the C call,
so a Free() racing between Diarize's own checks and the lock can no
longer reach the C side with a freed context. When the include_text
call returns NULL, last_error is now read from both contexts and
whichever came back non-empty is reported, since either side of the
pairing can be the one that failed. A WAV decode failure is reported
as InvalidArgument instead of an unwrapped/untyped error.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): harden SoundDetection's engine checks

soundStreamDrain ran every C call under engineMu but never re-checked
p.tagCtx there, so a Free() racing between SoundDetection's own
tagCtx==0 check and this lock could still reach the C side with a
freed context. Re-check p.tagCtx under the lock and return
ModelNotLoaded when it was cleared, mirroring diarizeCall's own
re-check. A WAV decode failure is now reported as InvalidArgument
instead of an unwrapped/untyped error.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(parakeet-cpp): cover a mid-session scene feed failure

feedSlicesScene already degrades gracefully when a scene feed call
fails mid-session: it frees the broken stream and carries the ASR-only
session forward. Add a spec covering that path end to end: the scene
stream is freed exactly once, later audio slices still produce ASR
responses, and no speaker/sound events appear before or after the
failure.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: fix the parakeet-cpp companion role table and realtime scene docs

audio-to-text.md's companion option table read "Use on" with a note
that the loader did not enforce the pairing; it now rejects a
companion whose role duplicates the primary's, so restore the
"Allowed on" wording and describe the real enforcement.

openai-realtime.md's live speaker/sound section claimed a mid-stream
session.update resets the companion stream and that it flushes on
session close; neither happens, since the realtime core opens one
live stream (and so one scene stream) per speech turn and closes it
at that turn's commit, with no mid-stream Config in between. Document
that lifecycle instead, state precisely that start/end are seconds
from the start of the turn's own audio, and note that the diarization
model starts a fresh session every turn, so a speaker index is only
meaningful within one turn. The example sound tag ("Rooster", index
17) did not match any real CED label; index 17 in ced-tiny-q8_0.gguf
is "Baby laughter". Replaced with "Chicken, rooster" at its real
index, 99.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): use CED's real index for Chicken, rooster

The scene feed comment and the live test's canned document gave
"Chicken, rooster" index 365. In CED's AudioSet label list it is 99,
which is also what the realtime docs show.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(realtime): call the test event accessor

The scene-event tests range over a method instead of its returned slice.
Call the synchronized accessor so the OpenAI test package compiles.

Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(parakeet-cpp): pin parakeet.cpp master with sound events

mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and
#74 (the missing <algorithm> include that broke the image builds) are
on master now. Pin 6dea76a instead of the #75 PR head, and update the
header comment the bump bot reads.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(parakeet-cpp): pin parakeet.cpp with ced.cpp on main

parakeet.cpp #76 moved its ced.cpp submodule from the head of
localai-org/ced.cpp#3 (a branch-only commit) to ced.cpp main, where
#3 landed with an identical tree. Pin 623a968 so the image builds no
longer depend on that branch.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(transcription): carry speaker labels on words and streamed segments

A diarizing backend could label transcript segments, but two paths
dropped the label: TranscriptWord had no speaker field, so live
transcription words and word-level timestamps could not carry one, and
the stream=true transcript.text.done event left the speaker out of
its segments.

TranscriptWord gains an optional speaker (proto field 4, additive).
It flows through the live event and result mapping, the JSON word
output of the endpoint and the CLI, and transcript.text.done now
includes a segment's speaker when there is one. Empty labels are
omitted, so responses without diarization are unchanged.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
(cherry picked from commit 2f0049f979)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(importers): detect the parakeet.cpp diarization GGUF

The Nemotron-3-Diarization GGUFs are published in
mudler/parakeet-cpp-gguf as nemotron-3-diarization-<quant>.gguf. The
parakeet-cpp importer did not recognise that name, so a direct
`local-ai models import` of the file fell through to another importer.

A direct URL to the file now imports with the diarization usecase. A
repo import still picks ASR weights when the repo also ships the
diarization model, and falls back to the diarization weights only
when there are no others.

Ported from #12323.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): advertise diarization and sound detection for parakeet-cpp

The capability table listed parakeet-cpp as transcription only, though
the backend now answers Diarize (Nemotron-3-Diarization) and
SoundDetection (CED) depending on the model kind it loads.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): label transcript segments with the diarization companion

A diarization_model companion only fed live speaker events and
Diarize; /v1/audio/transcriptions ignored it.

With the companion attached and diarize=true (the OpenAI endpoint's
default), unary transcription now labels each segment with its
speaker and splits segments at speaker turns; with word timestamps
each word carries its speaker. The stream=true final result labels
each utterance with the speaker who said most of it. Both use the
checkpoint's own diarization over the whole clip, as NeMo's diarize()
does. Words take the speaker whose segments overlap them most, or the
nearest segment within 0.5 s, the same rule as parakeet.cpp's
speaker-attributed ASR.

Docs: the diarization_model row and a paragraph on transcript
speakers; Nemotron-3-Diarization handles up to 8 speakers.

Ported from #12323.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(realtime): speaker segments from committed-turn transcription

Speaker events reached a realtime session only from the live
semantic_vad path, which needs a cache-aware streaming transcription
model. Committed-turn transcription (server_vad, or any offline
model) always asked the backend for diarize=false and dropped the
segments' speakers.

pipeline.diarization (off by default) asks the transcription model for
speaker labels on each committed turn and emits every labelled segment
as a conversation.item.input_audio_transcription.segment event, with
its text, before the turn's completed event. It is opt-in because some
backends fail a diarization request they cannot serve.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add parakeet-cpp-realtime-scene-tdt

parakeet-cpp-realtime-scene pairs the streaming EOU model with the
diarization and CED companions; its speaker and sound events need a
cache-aware streaming model. This entry does the same with Parakeet
TDT 0.6B v3 (multilingual, offline) for realtime under server_vad:
set it as both transcription and sound_detection and turn on
pipeline.diarization, and each committed turn gets speaker segments
and sound tags from one parakeet-cpp backend.

Files and sha256 match the Hub and are shared with the existing TDT v3,
diarization and CED-Tiny entries. A real-model spec checks the
combination on a clip with two speakers and a rooster: A-B-A-B speaker
turns, and "Chicken, rooster" among the sound tags. The test loader
now binds the sound entry points like main.go.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add CED-Base variants of the parakeet-cpp scene models

parakeet-cpp-realtime-scene and parakeet-cpp-realtime-scene-tdt ship
with CED-Tiny. The -base variants use CED-Base (86M), which tags sounds
more confidently (on the rooster clip "Crowing" 0.65 against 0.49 for
Tiny).

Measured on CPU over a 37 s clip: the live diarization + sound stream
runs at 0.125 of real time with CED-Base against 0.103 with CED-Tiny,
because diarization dominates; sound detection per committed turn costs
0.031 against 0.005. The realtime docs list both and note that any CED
size works as sound_model.

Files and sha256 match the Hub and are shared with the existing
parakeet-cpp-ced-base entry. The TDT variant passes the real-model
scene spec with CED-Base (A-B-A-B speakers, "Chicken, rooster" found).

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): register pipeline.diarization in the config metadata

TestAllFieldsHaveRegistryEntries fails on the branch because the new
pipeline.diarization field has no registry entry. Add one so the model
editor shows it as a toggle next to the sound detection options.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-29 23:58:17 +02:00
localai-org-maint-botandmudler b93d111d9b chore: ⬆️ Update NVIDIA/NeMo-Speech.cpp to 0f706e43cf1fbc031bad1423e05460d3acaeaa1c (#12361)
⬆️ Update NVIDIA/NeMo-Speech.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-29 23:56:46 +02:00
localai-org-maint-botandmudler 7cadb5748a chore: ⬆️ Update localai-org/ced.cpp to b10237678d1c3b30c77d19f2e63f6c198c7f8d09 (#12346)
⬆️ Update localai-org/ced.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-29 18:17:24 +02:00
localai-org-maint-botandmudler 01db8f33c1 chore: ⬆️ Update CrispStrobe/CrispASR to 2cd383a926e3c37334e75eb5d8b8a85bd82ed22d (#12348)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-29 18:17:09 +02:00
localai-org-maint-botandmudler 73cb1c3fdf chore: ⬆️ Update ggml-org/whisper.cpp to 6e4ab854f67f743900934a703d5603419384c961 (#12349)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-29 18:16:56 +02:00
localai-org-maint-botandmudler e0ad36a7cf chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 53e6c2066150802ad3cd4b655b31c696e78e0019 (#12350)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-29 18:16:47 +02:00
localai-org-maint-botandmudler bbaa545e27 chore: ⬆️ Update ikawrakow/ik_llama.cpp to d741de5074cd424dd3ba7cfc4d9b7649f1eb0463 (#12351)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-29 18:16:35 +02:00
localai-org-maint-botandmudler f9d51ed421 chore: ⬆️ Update 0xShug0/audio.cpp to f825d1d1b92af309585aeb656b2a59c44fc603eb (#12343)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-29 12:29:08 +02:00
mudler-agentandEttore Di Giacinto b700da3eb3 fix(huggingface): list repos nested more than one directory deep (#12355)
The HuggingFace tree API returns each entry's path relative to the repo
root ("assets/plots", not "plots"). The recursive listing prefixed the
parent directory again, so it requested "assets/assets/plots", got a
404, and failed the whole listing. The importer then treated the URI as
a non-HF repo and no importer matched.

This broke the import of GGUF repos that keep per-quant subfolders next
to a nested assets tree, such as
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF.

The test mock now returns root-relative directory paths like the real
API and routes on the exact tree path, so a doubled path 404s.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-29 10:37:31 +02:00
localai-org-maint-botandmudler 2fe459ca5f chore(model-gallery): ⬆️ update checksum (#12342)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-29 08:16:50 +02:00
localai-org-maint-botandmudler 84a2b209a0 chore: ⬆️ Update ggml-org/llama.cpp to 4da6337767f973e2b4d0797e5b323d77d8565e4a (#12318)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 21:41:46 +02:00
localai-org-maint-botandmudler b77937f61c chore(website): refresh the counters (#12329)
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 17:41:59 +02:00
Ettore Di Giacinto d8571a8ee9 fix(failover): spill 429 admission rejections to the next target
#12113 changed admission control to reject with 429 instead of 503.
failoverWriter only held back responses with status >= 500, so a 429
rejection reached the client and the chain never spilled to its next
target.

Hold 429 as well. An admission rejection is still flagged and spills
without tripping the target. Any other 429 is not retryable, so it is
released to the client unchanged.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
2026-09-28 15:34:29 +00:00
Ettore Di Giacinto 9c156656bd Merge PR #12285: feat(failover): serve a model name from a chain of local and remote targets
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 15:32:30 +00:00
mudler-agentandEttore Di Giacinto 590512d9eb feat(distributed): report worker version and show models in node inspector (#12328)
* ci: bump Hugo from 0.146.3 to 0.166.0

The hugo-theme-relearn submodule was bumped to 9.1.x in #12096,
which requires Hugo >= 0.165.0. The pinned 0.146.3 broke the docs
site build with a template error in alias.html that could not
evaluate the Locale field on langs.Language.

Bump HUGO_VERSION to 0.166.0 (latest stable) to satisfy the
theme minimum and resolve the alias.html template error.

Assisted-by: nib:claude-sonnet-4.5 [bash] [read] [edit]

* feat(distributed): report worker version and show models in node inspector

Workers now send their LocalAI build version and git commit at
registration. The controller stores them on BackendNode and exposes
them through the existing node list/detail API responses.

The node inspector side pane now fetches and renders the list of
loaded models (name, state, in-flight) instead of showing only a
count, matching what the node detail page already displays.

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 13:13:31 +02:00
localai-org-maint-botandmudler 102fbe0c78 chore: ⬆️ Update leejet/stable-diffusion.cpp to 3f8527a46c54ecf4cb4ed6003da8e8982283c73c (#12317)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 13:10:28 +02:00
localai-org-maint-botandmudler 197311f033 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to ead199a2bc4c53a57cac90095ae049a111d9e98d (#12316)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 13:10:11 +02:00
localai-org-maint-botandmudler b449ad3828 chore: ⬆️ Update 0xShug0/audio.cpp to 77491a33c589c53ff18add050095cf35647c8213 (#12315)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 13:09:57 +02:00
localai-org-maint-botandmudler 7f139c8add chore: ⬆️ Update CrispStrobe/CrispASR to ec98831d0776ec8a16ccaf93955693eb7ecfbec3 (#12314)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 13:09:32 +02:00
localai-org-maint-botandmudler adbff0a44c chore: ⬆️ Update ikawrakow/ik_llama.cpp to ed27bf7ed25e637692e89cd341d802522a2cee8a (#12313)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 13:09:15 +02:00
Stefan Walcz bc1d9924de fix(llama-cpp): keep llama.cpp's default cache_ram instead of no limit (#12297)
grpc-server.cpp forced params.cache_ram_mib = -1 (no limit) since #7009.
Since v4.3 kv_unified and cache_idle_slots are on by default, so every
distinct prompt now leaves its slot KV state in the host-side prompt
cache, and without a limit the backend grows until the host runs out of
memory.

Measured on gfx1151 (Strix Halo, 128 GB), llama-cpp backend, one request
at a time, 100 distinct prompts of ~2000 characters plus a fixed system
prompt, max_tokens 200:

  model                        cache_ram     RSS loaded -> after 100
  gemma-4-26B-A4B (q8_0 KV)    -1 (default)  1.4 GB -> 25.3 GB
  Qwen3.6-35B-A3B (q8_0 KV)    -1 (default)  1.1 GB -> 19.5 GB
  gemma-4-26B-A4B              -1, same prompt 100x  1.4 GB -> 1.6 GB
  gemma-4-26B-A4B              4096          1.4 GB -> 5.4 GB (flat from
                                             request 20 on, same latency)
  Qwen3.6-35B-A3B              4096          1.1 GB -> 5.1 GB (flat)

The memory is not released when idle. In production a document
classification pass pushed the daily chat model to 34 GB RSS overnight.

Drop the override so llama.cpp's own default (8192 MiB) applies; the
cache_ram option still accepts -1 for users who want no limit. Update
both docs tables (the option reference and the prompt-cache table) and
note what -1 does.


Assisted-by: Claude:claude-opus-5-5
Assisted-by: Codex:GPT-6

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-09-28 11:37:24 +02:00
mudler-agentandEttore Di Giacinto 50c284fdcc fix: make the remaining VerifyPath checks effective (#12326)
utils.VerifyPath joins its argument onto the base path, so a path that
the caller already joined always passes. Several callers gave it joined
paths, and their checks could not fail:

- modeladmin (config view, patch, edit, pin and state): the config file
  path from the loader. A config loaded from outside the models
  directory (--models-config-file) could be pinned, and the pin wrote
  the outside file. The patch and state paths stopped later, in the
  mutation snapshot, with a different error.
- core/backend/tts.go: the model path joined onto the models path.
- The trellis2cpp and stablediffusion-ggml backends: option paths
  (*_path) joined onto the model path. A "../" value outside the model
  directory was accepted.

Add utils.VerifyResolvedPath for a full path. modeladmin and tts use
it. The backends now check the relative option value before they join
it. A rename in modeladmin checks the new relative name.

For models from a config file outside the models directory, the admin
API and web UI now return ErrPathNotTrusted for view, edit, pin, and
enable or disable. The docs describe this.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 10:04:22 +02:00
mudler-agentandEttore Di Giacinto bc01ef2350 fix(gallery): keep model deletion inside the models directory (#12324)
listModelFiles gave utils.VerifyPath paths that it had already joined
onto the models directory. VerifyPath joins its argument onto the base
again, so an absolute path always passes and none of the four checks
could fail. Model deletion then removed files outside the models
directory:

- A model name such as "../outside/victim" removed
  outside/victim.yaml. The in-process MCP delete_model tool passes the
  name from the tool call without a check.
- A gallery file that lists a files: entry with "../" removed that
  file.

listModelFiles now gives VerifyPath the relative names.

InTrustedRoot also looped forever when a relative path was outside a
relative root. filepath.Dir stops at "." for a relative path, and the
loop waited for "/". The loop now stops when Dir returns its input.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 09:28:04 +02:00
localai-org-maint-botandmudler 10f3cd8015 feat(swagger): update swagger (#12308)
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 08:35:19 +02:00
mudler-agentandEttore Di Giacinto 97ad8f1d70 fix(distributed): stage the files a model install declares (#12309)
* fix(distributed): stage every shard of a split GGUF

A split GGUF is configured by its first shard only. llama.cpp opens the
other "-0000N-of-0000M.gguf" files from the same directory by name. The
router staged only the configured path, so the worker received shard 1
and the load failed with "failed to load GGUF split".

The router now stages the remaining shards next to the first one. A
missing shard fails the load and names the file. The file count for
progress and the payload size also include all shards. The payload size
feeds the load deadline and the disk-headroom check. For a 111 GB model
whose first shard is 10 MB, both were sized for less than 1 GB.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): stage the files a model install declares

Replace the split GGUF file name matching with the model's own file
list. A gallery install or an import records every file of the model in
._gallery_<name>.yaml (files:), and a config can list more under
download_files:. The router now stages all of these files, not only the
files that the config's path fields name. This includes the other
shards of a split GGUF, which llama.cpp opens by name.

The application gives the router a resolver that reads the two lists.
The resolver looks up the files by model name when it stages them, so a
replica that the reconciler loads from saved load options gets the same
files. backend.proto does not change.

A declared file that is missing on the frontend is skipped with a
warning. The load deadline and the disk headroom check include the
declared files.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 08:34:56 +02:00
localai-org-maint-botandmudler 81348581da chore(model-gallery): ⬆️ update checksum (#12310)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-28 08:34:31 +02:00
Pratik Gandhi ea10f26f57 docs: fix 11 dead links in the documentation (#12320)
- compatibility table: voxtral.c, OuteTTS and VoxCPM repos live under
  antirez, edwko and OpenBMB
- distributed inferencing: llama.cpp RPC README moved to tools/rpc under ggml-org
- customize-model: the phi-2 example config moved to LocalAI-examples, and
  embedded/models was replaced by the gallery
- model-gallery: malformed URL; link the gallery index
- GPU acceleration: ROCm install guide moved
- integrations: Wave Terminal docs page moved to ai-presets

Assisted-by: Claude:claude-fable-5-1

Signed-off-by: Pratik Gandhi <travpreneur@gmail.com>
2026-09-28 08:33:15 +02:00
61f4f67b75 sglang backend: pass through thinking_budget + require_reasoning (#12193)
* sglang backend: pass through thinking_budget + require_reasoning

sglang's raw Engine.async_generate() API (which this backend calls
directly, bypassing sglang's own OpenAI server) supports a precise,
tokenizer-derived reasoning-length budget via
sampling_params["custom_params"]["thinking_budget"] plus
require_reasoning=True, gated behind --enable-strict-thinking. Neither
was reachable through LocalAI: this backend built sampling_params only
from a fixed field mapping (temperature, top_p, ...) with no custom_params
key, and never passed require_reasoning to async_generate at all.

- LoadModel now reads a model-level "thinking_budget" option (same
  mechanism as the existing tool_parser/reasoning_parser options), and
  _build_sampling_params adds it as custom_params.thinking_budget on
  every request when configured.
- _new_reasoning_parser already derives, from the rendered prompt, whether
  the model's chat template pre-opened a reasoning block (Qwen3-style
  templates append <think> to the prompt instead of letting the model
  emit it) -- the same signal sglang's own OpenAI server computes from
  per-template config to decide require_reasoning. This backend has no
  template manager, so it now returns that signal too and _predict
  forwards it to async_generate(require_reasoning=...).

Verified against production (NVFP4, sm_121, Qwen3.6-35B-A3B) via a raw
Engine.async_generate() call bypassing this backend: 301 reasoning
tokens against a 300-token budget, clean completion, ~27s. Not yet
verified through this backend's own gRPC path end-to-end (no local
CUDA/sglang environment available here) -- existing + new unit tests in
test.py cover the pure-Python merge/passthrough logic only.

Scope note: require_reasoning is derived only from the existing
prompt-suffix heuristic, not sglang's full per-template
_get_reasoning_from_request decision tree (minimax-m3/hunyuan special
cases etc.) -- this backend has no template manager to evaluate that
tree against, and the prompt-suffix check is the one heuristic already
validated in this file (test_reasoning_parser_forced_when_template_prefills_think_tag).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: honour a model-level reasoning_default

A model YAML can already carry "parameters: reasoning_effort:", but that
value only reaches this backend when a *caller* sets it per request (the Go
side turns it into Metadata["enable_thinking"]). As a model-level default it
is silently dropped: a config reading "reasoning_effort: none" still produces
full reasoning on every request, so the config says one thing and the model
does another.

That gap is expensive in practice. On a self-hosted Qwen3.6-35B-A3B the
reasoning phase consumed the entire max_tokens budget before any content was
produced - 90% of code completions came back empty at max_tokens=768, and the
server log filled with "backend produced only reasoning, retrying". The
config looked like reasoning was off the whole time.

This adds "reasoning_default:off" (or ":on") on the same model-level
options: mechanism as thinking_budget. A per-request value always wins; the
default only fills in when the request is silent.

Measured on the stack above (sglang 0.5.20, NVFP4, GB10/sm_121) after
applying it:
  default (nothing set)          -> 0 chars reasoning, 27 tokens
  "reasoning_effort": "none"     -> 0 chars reasoning, 27 tokens
  metadata enable_thinking=true  -> capped at the 512-token thinking_budget,
                                    541 tokens total, finish_reason stop

Tests: three cases added to backend/python/sglang/test.py covering the
default, per-request override in both directions, and the unconfigured case
(which must leave the template untouched).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: validate thinking_budget instead of crashing LoadModel

Addresses the review on this PR:

- `int(thinking_budget)` raised on values like "5000.0" or "abc" and took
  LoadModel down. The option is now parsed by _parse_thinking_budget():
  integral numbers in any spelling are accepted, anything else is ignored
  with a warning on stderr.
- Zero and negative budgets are ignored with a warning instead of being
  passed to sglang, where they have no defined meaning. Turning reasoning
  off is what reasoning_default:off is for.
- A load-time warning when thinking_budget is set but enable_strict_thinking
  is not in engine_args, since sglang then ignores the budget silently.
- Tests for integral spellings, unset, zero, negative, non-integer and the
  strict-thinking warning.

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): explain reasoning options

Document the reasoning budget, strict-thinking requirement, and
precedence of request metadata over the model-level default.

Also note that the budget has to stay well below max_tokens (otherwise
it never triggers and the reply can end up empty), and that
POST /models/reload or a backend-only restart does not pick up changed
options; LocalAI itself has to be restarted.

Assisted-by: Codex:GPT-6
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): clarify configuration reloads

Distinguish rereading model configuration from updating a running backend. Keep the full LocalAI restart recommendation for changed reasoning options.

Assisted-by: Codex:GPT-6

* sglang backend: only pass require_reasoning when sglang supports it

Engine.async_generate() gained the require_reasoning keyword in sglang
0.5.13 and takes no **kwargs. The CPU profile builds v0.5.11 from source
and the other profiles only set a >=0.5.11 floor, so passing the keyword
unconditionally made every request fail with TypeError. Detect support
once at import time, as the file already does for sampling_seed.

enable_strict_thinking first appears in sglang 0.5.12; fix the comment.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 04:49:37 +02:00