Commit Graph
8418 Commits
Author SHA1 Message Date
Aniruddh Krovvidi a3d555653c fix(model): stop listing a backend that exited on its own (#12497)
Assisted-by: Claude:claude-fable-5-1

Signed-off-by: Aniruddh Krovvidi <akrovvidi05@gmail.com>
2026-10-07 09:28:49 +02:00
localai-org-maint-botandmudler 5537f4b1ef chore: ⬆️ Update vllm-project/vllm cu130 wheel to 0.31.0 (#12502)
⬆️ Update vllm-project/vllm cu130 wheel

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:27:43 +02:00
localai-org-maint-botandmudler e38d9c9be7 chore: ⬆️ Update PrismML-Eng/llama.cpp to 6bfcd79a2d426abcd2b50e3c2d09ae2225e70a17 (#12503)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:27:23 +02:00
localai-org-maint-botandmudler d6238f60cd chore(model-gallery): ⬆️ update checksum (#12505)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:27:07 +02:00
localai-org-maint-botandmudler 0a9fb07574 chore: ⬆️ Update localai-org/voice-detect.cpp to cf9e1d5641c710bf7d6fc9642f9b4cf7d2bde9b6 (#12506)
⬆️ Update localai-org/voice-detect.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:26:51 +02:00
localai-org-maint-botandmudler b554fc55a8 chore: ⬆️ Update CrispStrobe/CrispASR to e79a671fce388b50d3e2eb5fced1581140a29594 (#12531)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:26:35 +02:00
localai-org-maint-botandmudler c73977f5b6 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 739839a5c922cf9663988912004912e9f0671d7d (#12532)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:26:16 +02:00
localai-org-maint-botandmudler a1b0648f19 chore: ⬆️ Update ggml-org/whisper.cpp to d1be6fde11ac6e0407606b4e42fe72d34add8037 (#12534)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:21:02 +02:00
localai-org-maint-botandmudler f201055812 chore: ⬆️ Update mudler/parakeet.cpp to 9a28a3c1f7fb7d89505646b4d437033e4163f6e1 (#12511)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:20:39 +02:00
localai-org-maint-botandmudler 6d1c8d7f71 chore: ⬆️ Update leejet/stable-diffusion.cpp to a1ded76da5818803fca97a3b433669ef727d32cf (#12533)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:20:16 +02:00
localai-org-maint-botandEttore Di Giacinto 217038d4fe fix(agents): share the collection list across frontends (#12536)
With the PostgreSQL vector engine, each frontend kept the collections it
had opened in an in-memory map and answered "collection not found" for
any other. A collection created through one frontend was unknown to the
others until they restarted, and each frontend listed a different set.

Wrap the in-process collections backend in distributed mode so that the
database is the source of truth:

- lists come from the registry,
- a lookup miss checks the registry before it returns 404, and opens a
  collection that exists there (once per name, even under concurrency),
- a cached collection that left the registry is dropped and closed, with
  a re-check at most every 5 seconds,
- create and reset publish an event on the existing collection
  invalidation subject, so other frontends re-check at once.

Without the postgres engine, or outside distributed mode, nothing changes.

Assisted-by: Claude:claude-sonnet-5-5 go

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 22:45:23 +02:00
e2faec2bf2 fix(distributed): free only the named model on model.unload (#12537)
The worker answered model.unload by calling Free on the first backend
process in its map, whatever model the request named. Every idle
scale-down or LRU eviction of one model could therefore empty another
model's backend on the same node. LocalAI still counted that model as
loaded, so its next request failed. A parakeet diarization model then
returned 501 "speaker profiles require a loaded speaker encoder" until
someone reloaded it by hand.

Resolve the target from the model name (all replicas), prefer an
address when the request carries one, and free nothing for an unknown
model.

Also let parakeet-cpp Diarize check the diarization model before the
speaker-profile capability. A backend with nothing loaded now answers
FailedPrecondition, which LocalAI treats as a stale replica and
reloads, instead of a final Unimplemented.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 22:44:25 +02:00
localai-org-maint-botandEttore Di Giacinto 73b4acdca7 feat(gallery): add parakeet-cpp-multilingual-diarization-speakers (#12535)
The English-only 110M ASR in parakeet-cpp-nemotron-3-diarization-asr-speakers
garbles other languages. Add a bundle that pairs Parakeet TDT 0.6B v3
(25 European languages) with the Nemotron-3-Diarization model and the
WeSpeaker ResNet34 encoder, so one /v1/audio/diarization call with
include_text and include_speaker_profiles returns turns, multilingual text
and one voice-print embedding per speaker.

Assisted-by: Claude Code:claude-sonnet-5-5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 22:33:50 +02:00
mudler-agentandEttore Di Giacinto 690f3994b0 feat(auth): let users pause and resume their API keys (#12521)
A key can now be paused until the owner resumes it, or until a given
time. A paused key is rejected by validation before last_used is
updated, and a pause time that has passed lifts the pause by itself.
Existing keys stay active.

PATCH /api/auth/api-keys/:id takes {"disabled": bool, "paused_until":
RFC 3339 string or null}. Only the key owner can change it, and a
paused_until in the past is rejected. The key list returns the pause
fields. The Account page gets a Pause and Resume button for each key
and a Paused badge that shows the resume time.

Assisted-by: Claude Code:claude-sonnet-5-5

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 18:25:21 +02:00
fmterrors 1cd96e496b chore: fix some function names in comment (#12517)
Signed-off-by: fmterrors <fmterrors@outlook.com>
2026-10-06 17:26:03 +02:00
localai-org-maint-botandmudler e34847123e chore: ⬆️ Update ikawrakow/ik_llama.cpp to 3d27f5bb6348c4d79edded63b03e9a93efb2f80e (#12504)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-06 08:57:29 +02:00
Ettore Di Giacinto 7467b51281 test(gallery): expect speaker_recognition on the bundles with a voice component
The parakeet-cpp backend now answers VoiceEmbed and VoiceVerify, so the
pinning test asserts the usecase on the entries that declare it instead of
its absence. Document voice_verify_threshold.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:48:20 +00:00
Ettore Di Giacinto 1f22b16092 chore(parakeet-cpp): bump parakeet.cpp to 2f9e8ea
parakeet_capi_speaker_embed_pcm and parakeet_capi_free_floats, which
VoiceEmbed binds, come from this commit.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:48:20 +00:00
Ettore Di Giacinto 20f3c72d58 feat(parakeet-cpp): add the voice_verify_threshold model option
VoiceVerify used a fixed distance of 0.5 when the request had none. Read
voice_verify_threshold, a distance in (0, 2), from the model options and
keep 0.5 as the default. The real-library spec now cuts clips from a
two-voice recording and checks the bundle embedding size, the encoder
identity, determinism and the same-voice versus different-voice distance.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:48:20 +00:00
Ettore Di Giacinto e9af44b90b docs(voice): document a parakeet-cpp bundle as the embedding model
Describe the bundle as an embedding model for /v1/voice/*, the realtime
voice_recognition stage, and the limits: identify and plain verify only,
a 256-dimension space shared with voice-detect-wespeaker-resnet34, and the
libparakeet.so symbol it needs.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:46:14 +00:00
Ettore Di Giacinto ce4099ea38 feat(gallery): declare speaker_recognition on parakeet-cpp bundles
The small and standard bundles have a voice component, and
parakeet-cpp-realtime-scene-speakers loads a speaker encoder, so they can
serve /v1/voice/* and the realtime voice_recognition stage. Add the usecase
to their known_usecases and pin it in the gallery tests.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:46:06 +00:00
Ettore Di Giacinto fa25493c9e feat(config): let parakeet-cpp declare the speaker_recognition usecase
parakeet-cpp now implements VoiceEmbed and VoiceVerify, so list them in
its backend capabilities together with the speaker_recognition usecase.
The usecase is not guessed, because most parakeet models have no
speaker encoder: a bundle declares it in known_usecases, which is
enough for /v1/voice/* and for voice_recognition.model.

Add a realtime pipeline test where one bundle model names the vad,
transcription, sound_detection and voice_recognition stages and the
gate resolves the speaker through the backend's VoiceEmbed.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:46:06 +00:00
Ettore Di Giacinto 3beb2db0c9 feat(parakeet-cpp): serve VoiceEmbed and VoiceVerify from the speaker encoder
Bind parakeet_capi_speaker_embed_pcm and parakeet_capi_free_floats with a
Dlsym probe and implement VoiceEmbed on the speaker context, so a model
with a speaker_model or a bundle speaker_component can serve the
realtime voice_recognition stage and /v1/voice/*. The response carries
the sha256 identity of the encoder weights.

VoiceVerify embeds both clips and compares them by cosine distance. It
refuses anti_spoofing because there is no such head.

A libparakeet.so without the symbols answers Unimplemented, and a model
without a speaker encoder answers FailedPrecondition.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:46:06 +00:00
Ettore Di Giacinto 04df10a082 docs(voice): document speaker naming through a bundle component
Cover speaker_component: and speaker_tag: in the audio-to-text option
tables and bundle section, add a section to the voice recognition page
that explains which registered voices a bundle component receives, and
update the speaker_name condition in the realtime and diarization pages.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:45:14 +00:00
Ettore Di Giacinto 1d117292b1 feat(voice): name registered speakers through a bundle speaker component
Naming read only speaker_model:, so a parakeet-cpp bundle that sets
speaker_component:voice never received the registered voices, in live
transcription or in diarization.

Resolve speaker_component: when speaker_model: is not set. A component
is not an encoder file, so no file-name tag is derived for it: voices
with a hash or family identity and untagged voices reach the backend,
and a voice with only a file-name tag matches through the new optional
speaker_tag: alias. speaker_model: still wins, including when it points
at the bundle file.

The bundle gallery entries do not declare speaker_recognition: that
usecase means the VoiceEmbed and VoiceVerify RPCs, which the backend
does not serve, and it selects the default model for /v1/voice/*. The
pinning test now says so.

Assisted-by: Claude Code:claude-sonnet-5-5 go golangci-lint
2026-10-05 23:45:14 +00:00
Ettore Di Giacinto a2c498359b fix(voice): keep one voice store per embedding dimension
The in-memory local-store rejects vectors of another size than the ones
it holds. With 192-value voices registered, adding a 256-value voice
failed, and identifying with a 256-value probe returned an error.

The registry now keeps one store per embedding dimension. The first
dimension seen uses the configured store name, so single-encoder
instances are unchanged. Later dimensions use "<name>-<dim>". Identify
searches only the store of the probe size and returns no match when no
voice of that size exists. Forget finds the store from the stored
embedding. A name may hold one voice per encoder.

Assisted-by: Claude Code:claude-sonnet-5-5 golangci-lint
2026-10-05 23:42:49 +00:00
localai-org-maint-botandEttore Di Giacinto 68c980f3cd chore(deps): bump localrecall to v0.6.6 (#12508)
* chore(deps): bump localrecall to v0.6.6

LocalRecall v0.6.6 closes the PostgreSQL connection pool when a
collection fails to open. Before, each failed collection create left
its pool open. With a failing embedding model, every retry leaked one
more pool until PostgreSQL refused new clients.

LocalAI gets LocalRecall through LocalAGI, so this raises the indirect
requirement directly instead of waiting for a LocalAGI bump.

The release also limits each collection pool to 4 connections by
default. POSTGRES_POOL_MAX_CONNS changes the limit.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

* docs(agents): document the PostgreSQL pool size limit

LocalRecall v0.6.6 limits the connection pool of each collection to 4
connections and reads POSTGRES_POOL_MAX_CONNS to change it. Document
the variable next to the other PostgreSQL settings of the embedded
store.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 00:51:53 +02:00
localai-org-maint-botandmudler a4acf1dd40 feat(swagger): update swagger (#12500)
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 22:36:03 +02:00
mudler-agentandEttore Di Giacinto 0f8d5640ba chore(model-gallery): aggregate reviewed model additions (#12501)
Collect the validated model-gallery entries from 48 open pull requests onto current master. Exclude #12325 because its variant groupings violate gallery invariants, and exclude #12353 because the image model is declared as an unsupported llama.cpp chat model.

Source PRs: #12303, #12311, #12321, #12322, #12327, #12330, #12332, #12334, #12338, #12340, #12352, #12354, #12357, #12359, #12360, #12367, #12369, #12370, #12371, #12376, #12393, #12394, #12396, #12398, #12400, #12409, #12422, #12423, #12429, #12431, #12432, #12434, #12435, #12444, #12448, #12450, #12454, #12457, #12464, #12468, #12473, #12476, #12478, #12483, #12488, #12489, #12492, #12494.

Assisted-by: nib:gpt-5.6-sol

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-05 22:35:35 +02:00
mudler-agentandEttore Di Giacinto 79a7631cc5 feat(parakeet-cpp): encoder fingerprint for speaker naming, VAD trim and word filter options, pin bump (#12491)
* chore(parakeet-cpp): bump parakeet.cpp to 2de154c

Brings in the speaker registry encoder fingerprint, the VAD segment trim
and the opt-in word filter, a fix for a per-call thread count that stayed
set on the process-wide backend after a Silero VAD pass, and bundle
components loaded from memory.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): encoder fingerprint for speaker naming, vad_trim and guard_* options

Speaker naming. A registered voice now carries the encoder that made it:
the embedding family (voicedetect:<arch>:<name>:<dim>) and the sha256 of
the weights. The backend reports the family of the loaded speaker model in
Status, voice enrollment from speaker_profiles stores it as encoder_family
(old entries load without it), and the registry sent to parakeet.cpp is
built with parakeet_capi_speaker_registry_add_embedding_fp. The library
then refuses a registry of another encoder family and the error names both
families; another quantization of the same family only warns. A voice with
only a weights hash gets the loaded family when the hashes are equal.

Voices without a fingerprint (registered from audio: libvoicedetect cannot
report one) keep the file-name rule and are used with a warning. The
library cannot mix them with fingerprinted voices in one registry, so a
request that has any uses the old registry for all. speaker_strict:true
drops them instead. A library without the symbols behaves as before.

Transcription. vad_trim (seconds, 0 keeps the whole cuts) goes through the
VAD options JSON, so it reaches /v1/vad and the segmenter. The guard_*
options guard_min_local_conf, guard_local_radius and guard_drop_punct_only
turn on the word filter through parakeet_capi_transcribe_path_json_with,
or through the segmenter with vad:true. They are off by default, bad
values fail the load, and a library without the symbol fails it with a
clear message. The dropped word count is logged at debug level.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-05 15:57:16 +02:00
localai-org-maint-botandlocalai-org-maint-bot 3d586cc3c5 fix(audio-cpp): forward voice reference transcripts (#11997)
Saved voices send ref_text, but Fish Audio requires reference_text.
Derive the canonical parameter while preserving explicit overrides.
Both TTS modes use the shared builder.

Add regression cases and document the parameter alias.

Assisted-by: Codex:gpt-6

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 11:05:43 +02:00
localai-org-maint-botandmudler e81e180ce2 chore(website): refresh the counters (#12490)
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 10:27:05 +02:00
localai-org-maint-botandmudler 05b20bd7ab chore: ⬆️ Update mudler/parakeet.cpp to 0cca477249ffb16c1623fb947d5bac0624961d41 (#12460)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 08:04:04 +02:00
localai-org-maint-botandmudler 0a72af4e95 chore: ⬆️ Update localai-org/voice-detect.cpp to bca46bcbc2fe68169c2d7c414e9b290a7cb89911 (#12486)
⬆️ Update localai-org/voice-detect.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 08:03:37 +02:00
localai-org-maint-botandmudler d79489adb3 chore: ⬆️ Update localai-org/ced.cpp to 736a4ee46a31d4ff38b41add65ce0f2cf1aaa05f (#12487)
⬆️ Update localai-org/ced.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 08:03:24 +02:00
localai-org-maint-botandmudler 7310887e26 chore(model-gallery): ⬆️ update checksum (#12485)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 01:40:12 +02:00
Stefan Walcz 5f5feea15a fix(agentpool): keep the conversation in the agent web chat (#12410)
* feat(agents): keep the web chat history between messages

The agent chat endpoint ran every message as a fresh job
(ag.Ask(WithText(message))), so a follow-up such as "add two days to
item 3 and recalculate" never saw the answer it referred to. Agents
then rebuilt their reply from scratch instead of changing it.

Use the agent's own conversation tracker, as the Telegram and Slack
connectors already do: send the earlier turns with the new message and
record successful answers. Failed, cancelled or empty runs are not
recorded, and the tracker drops a conversation after the agent's
last_message_duration of inactivity. The distributed (NATS) chat path is
unchanged.

Assisted-by: Claude:claude-opus-5-5 ginkgo
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>

* fix(agents): keep each web chat conversation's history separate

Review on this PR: the first version kept the history in the agent's
conversation tracker under one fixed key per agent. The web UI keeps
several conversations per agent (New Chat, switching, Clear) and did not
send which one a message belongs to, so the hidden history mixed them:
a new chat received the previous chat's turns, and Clear only cleared
the screen.

The client now sends the earlier turns of the conversation it is
showing as `history` with POST /api/agents/:name/chat, and the server
keeps no web chat history of its own. Each conversation only ever sees
its own turns; New Chat and Clear start without history. The server
uses only user and assistant turns with text, bounded to the most
recent 40 turns and 64,000 characters. `history` is optional, so
existing callers keep the previous behaviour (no history); the
distributed (NATS) path does not forward it yet.

Tests: two conversations of one agent stay apart, an empty history
(New Chat or Clear) starts fresh, system/tool/empty turns are dropped,
the bounds keep the most recent turns; the UI helper that builds the
history from the visible messages has node --test coverage. Docs:
features/agents.md describes `history` and the web UI behaviour.

Assisted-by: Claude:claude-opus-5-5 ginkgo
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>

---------

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-05 01:38:29 +02:00
Stefan Walcz 9b0b7da96e fix(rocm): point *_TENSILE_LIBPATH at the architecture folder (#12411)
Newer ROCm releases split the rocBLAS and hipBLASLt TensileLibrary data
into one folder per GPU architecture (lib/hipblaslt/library/gfx1151/...).
The run.sh of the ROCm-capable backends exports ROCBLAS_TENSILE_LIBPATH and
HIPBLASLT_TENSILE_LIBPATH as the parent folder, and both libraries only look
directly in that folder, so they miss every kernel:

  rocblaslt error: Cannot read ".../hipblaslt/library/TensileLibrary_lazy_gfx1151.dat"
  hipModuleLoad failed: .../hipblaslt/library/Kernels.so-000-gfx1151.hsaco

and fall back to slower code paths.

rocm_tensile_dir keeps the folder when it has files at the top level (the
older flat layout) or several architecture folders, and descends into the
architecture folder when the bundle carries exactly one.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-05 01:38:09 +02:00
localai-org-maint-botandlocalai-org-maint-bot e70c409a93 fix(kokoros): default optional status fields (#12447)
The speaker encoder metadata field makes the Rust status initializer
incomplete. Use protobuf defaults for absent metadata and memory fields.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 01:37:46 +02:00
localai-org-maint-botandlocalai-org-maint-bot 4a72b2ac2e test(cli): transfer socket descriptor ownership (#12465)
Duplicate the socket descriptor before handing it to the activation helper.
Close the original file to prevent its finalizer from closing a reused
coverage descriptor after the helper closes its own file wrapper.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 01:37:29 +02:00
leilei3167andlocalai-org-maint-bot d90f703bd9 fix(xsysinfo): read process VRAM without proc children (#12482)
* fix(xsysinfo): read process VRAM without proc children

Kernels without CONFIG_PROC_CHILDREN have no task children file, so
ProcessVRAM dropped every DRM reading. When that file is missing, walk
child processes from /proc/<pid>/stat ppid links instead. Other read
errors still drop the reading.

Fixes #12481

Signed-off-by: leilei3167 <imleilei123@gmail.com>

* docs(system): describe the proc children fallback

Document VRAM reporting on kernels without CONFIG_PROC_CHILDREN.

Assisted-by: Codex:GPT-6

---------

Signed-off-by: leilei3167 <imleilei123@gmail.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 01:37:12 +02:00
Plamen K. Kosseff 104c2fb440 fix(oci): stage image downloads in a configured directory, not in /tmp (#12280)
* fix(oci): stage image downloads in a configured directory, not in /tmp

Problem:
- the image tar and its compressed layers staged in os.TempDir()
- /tmp is commonly a RAM-backed tmpfs far smaller than a backend image
- big installs failed with a full /tmp or silently ate RAM

Change:
- staging path resolved like the other storage paths
  (LOCALAI_DOWNLOAD_STAGING_PATH, default ${basepath}/downloading),
  threaded to the extractor as a download option
- every download works in its own subdirectory holding its tar and
  layers: nothing shared between concurrent downloads, one removal
  cleans a download up, a crash leaves one self-contained orphan
- callers that pass no staging directory keep the OS temp behavior

Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>

* fix(oci): handle staging cleanup errors

Log staging cleanup failures to satisfy errcheck. Test extraction with an
unavailable OS temp directory and verify that staging is removed.

Assisted-by: Codex:GPT-6
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>

---------

Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
2026-10-05 01:27:45 +02:00
nexxtmobile.deandnexxtmobile.de 4d681b7f6d fix(realtime): detach VAD-commit transcription from barge-in cancellation (#12446)
* fix(realtime): keep VAD-commit transcription alive across barge-in, cancel it at teardown, order commits

Barge-in (new speech onset) cancels the turn's SourceVAD response context
(realtime_turncoord.go: respSink.cancel(SourceVAD)). The VAD commit body
runs under that same context, so an in-flight Whisper STT call was aborted
with 'context canceled' whenever the caller kept talking while the first
chunk was transcribing. The user's turn was lost: no transcript, no
LLM/TTS response.

v1 of this fix ran the transcription with context.WithoutCancel(ctx). The
review correctly pointed out two correctness gaps:

1. Teardown lost its cancellation. WithoutCancel detaches from every
   cancellation, so a transcription in flight at session close outlived
   the session and blocked respSink.shutdown (which joins the response
   goroutines) until the backend finished the job.
2. Out-of-order commits. Consecutive commits run in parallel goroutines,
   so a fast second transcription could append its user item before a
   slow first one: the conversation became [second, first] and the second
   response saw only [second].

Changes (core/http/endpoints/openai/):
- Session gains a session-lifetime context (sessionCtx), cancelled by
  conncoord's Teardown BEFORE respSink.shutdown joins the response
  goroutines. The transcription (and the voice-gate resolution) run under
  it: they survive barge-in (which cancels only the per-response context)
  but are cancelled with the session.
- Commit slots order the user-item appends in speech order:
  Session.nextCommitSlot() is claimed at commit issue time (VAD CommitTurn
  / client commit), a commit's item append waits on the previous slot's
  done (aborts on the session context), and every exit closes the slot so
  a failed or torn-down commit never blocks the next. Transcriptions stay
  parallel; only the appends are ordered.
- If the turn's response context was cancelled while the (detached)
  transcription ran — barge-in, superseded by a newer commit — the user
  item still commits (appendUserItem, split out of generateResponse) so
  the LLM context keeps the full user input, but no response is generated
  for the superseded turn; the newer speech triggers its own response on
  the complete history.
- Regression tests (realtime_commit_order_test.go) cover both review
  schedules — teardown during an in-flight transcription, and
  held-first/finished-second out-of-order completion — plus the
  barge-in-during-transcription item survival, driving the real commit
  path with a transcription double that honours context cancellation.
- docs/design/realtime-state-machines.md: implementation-status entry for
  the committed-turn pipeline (transcription lifetime + commit order).

Fixes #12445

Validated: builds, go vet clean, all openai specs + respcoord/turncoord/
conncoord suites pass under -race (incl. the 3 new regression specs).
Production A/B (call-center voice agent, SIP, silero-vad +
whisper-large-turbo + LLM + TTS, server_vad ~600 ms) on LocalAI v4.11.0:
unpatched — 'transcription_failed: context canceled', first part of the
utterance lost, agent answers only the remainder; patched — full
transcript committed, agent answers the complete utterance, barge-in
still cancels the in-flight assistant TTS response as intended, and
teardown cancels the in-flight transcription instead of waiting for the
backend.

Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>

* fix(realtime): release commit slots in order on every exit; share slot+issue boundary

Follow-up to the review of c8f991be (issue #12445): two schedules still
broke commit ordering.

1. A failed middle commit released later turns before earlier turns
   finished. slot.done closed on every return, but the wait on
   slot.prevDone happened only on the successful nonempty-transcript
   path. Hold transcription A, let B fail (or return an empty
   transcript / be rejected by the voice gate), then complete C: B
   closed its channel without waiting for A, so C appended and started
   its response without A (response history [third], final history
   [third first]). Fix: the slot now releases (done closes) only AFTER
   the predecessor has finished — on EVERY exit path, including errors,
   empty transcripts, gate rejections and teardown (the session context
   can still stop the wait, so teardown never blocks on a
   never-finishing predecessor). The success path keeps its append gate
   (wait before appending the user item); the deferred release gate
   enforces the same order on every other exit.

2. Slot order and response issue order could disagree between the two
   producers. The VAD CommitTurn and the client
   input_audio_buffer.commit reserved the slot and called
   respSink.issue separately; a pause between the two let the other
   producer reserve AND issue first, so the later issue superseded the
   EARLIER turn's response (response history [first], final history
   [first second], second turn un-answered). Fix: both producers now go
   through Session.issueCommit, which claims the slot and issues the
   body under one lock (commitOrderMu) — slot order == issue order.
   respSink.issue is non-blocking, so the lock never stalls
   VAD/barge-in handling.

Regression tests (realtime_commit_order_test.go) now drive the REAL
issue path — Session.issueCommit into the real responseSink/respcoord,
so coordinator supersession and the spawned response goroutines are
exercised — and cover: teardown during an in-flight transcription;
held-first/finished-second out-of-order completion; barge-in
(respSink.cancel) item survival; a FAILED middle commit; an EMPTY
middle commit; interleaved VAD/client producers in both directions.
The failed/empty middle specs fail deterministically without the
release gate (verified against the pre-fix code).

docs/design/realtime-state-machines.md: implementation-status entry
updated (append gate + release gate + shared issue boundary).

Validated: builds, go vet clean, all openai specs + respcoord/turncoord/
conncoord suites pass under -race (476 specs, incl. the 4 new ones).

Fixes #12445

Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>

---------

Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>
Co-authored-by: nexxtmobile.de <kai@nexxtmobile.de>
2026-10-05 01:27:40 +02:00
mudler-agentandEttore Di Giacinto d66383c0e2 docs(readme): refresh news through LocalAI 4.11 (#12484)
The news section stops at June despite newer published releases.
Summarize releases 4.5 through 4.11 and collapse existing news without
removing history. Keep the release index and blog links visible.

Assisted-by: ChatGPT:gpt-6-astra

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 22:24:54 +02:00
mudler-agentandEttore Di Giacinto 3185b6fcf8 feat(parakeet-cpp): load bundle GGUF files, bundle gallery entries, pin bump (#12479)
* feat(parakeet-cpp): load bundle GGUF files and use their components by role

A bundle GGUF holds several models (ASR, VAD, diarization, sound events,
speaker encoder) in one file, each with its own licence. Detect a bundle
at load through parakeet_capi_bundle_components_json and open components
with parakeet_capi_load_component. The three symbols are probed together,
so an older libparakeet.so still loads plain files as before.

The only ASR component is the primary model; bundle_asr:<name> picks one
when there are several. A Silero VAD component of the primary bundle is
loaded without an option and serves /v1/vad and vad:true. The diar, ced
and voice components load on request: diar_component, sound_component and
speaker_component, or a companion option (diarization_model, sound_model,
speaker_model, vad_model) that names a bundle, even the model file itself.
vad_component picks a VAD component and implies vad:true.

A role the bundle cannot fill fails the load with the component list, and
a diarization or sound request on a model without that role names the
bundle components. Every existing option and single-file model behaves as
before.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* chore(parakeet-cpp): bump parakeet.cpp to 781a973

Brings in the bundle GGUF format and its C-API (parakeet_capi_load_component,
parakeet_capi_bundle_components_json, parakeet_capi_load_error).

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): gallery entries for the bundle GGUF files, docs

Add parakeet-cpp-bundle-small (338 MB: Parakeet TDT+CTC 110M, Nemotron-3-
Diarization, CED-Small, WeSpeaker ResNet34-LM, Silero VAD), -standard
(1.1 GB, Parakeet TDT 0.6B v3 instead of the 110M model) and
-moondream-redux (215 MB: packed Redux and Silero VAD, CPU only). One
install serves transcription, VAD, diarization, sound events and speaker
naming through the component options. The existing single-purpose entries
stay.

A bundle has no single licence, so the entries use license: other and
state the licence and credit of each component in the description, with
the upstream inconsistency of the CED licence. The docs get a section on
bundles in audio-to-text with the entries, the roles, the options and the
licence notice, and pointers from the VAD, diarization and sound
classification pages. A gallery test checks the file names, checksums,
usecases and options.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 20:33:13 +02:00
mudler-agentandEttore Di Giacinto d3724eef50 fix(nodes): stage dedicated diarization audio (#12474)
Dedicated diarization forwards frontend audio paths to remote workers,
which cannot read those temporary files. Stage the input before the RPC
and release it afterward, following the transcription lifecycle.

Clone the request so staging does not change caller-owned data. Cover
input bytes, request fields, cleanup, and error propagation in tests.
Document distributed diarization staging on the existing feature page.

Assisted-by: nib:gpt-6-astra

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 13:14:00 +02:00
mudler-agentandEttore Di Giacinto ed4a3975be feat(parakeet-cpp): gallery entries for the VAD-only Moondream slices, pin bump (#12469)
* feat(parakeet-cpp): add gallery entries for the VAD-only Moondream slices

Add parakeet-cpp-vad-moondream-redux and parakeet-cpp-vad-moondream-ultra.
They install the VAD head of Moondream Redux and Ultra (Q8_0) as small
files of 10 MB and 6 MB, cut out of the full models without retraining,
for the VAD endpoint. The files cannot transcribe, and a transcription
request fails with a clear error.

The files load only with a parakeet.cpp build that has VAD-only GGUF
support (parakeet.cpp pull request 87). The backend pin must move to a
commit that includes it before these entries work in a released image.
The parakeet-cpp-vad entry keeps installing Silero.

The docs list the files with the size, load time and memory compared
with loading a whole model. A gallery test checks the usecase, the file
name and the checksum of each entry.

Assisted-by: Claude Code:claude-sonnet-5-5 [golangci-lint]

* chore(parakeet-cpp): bump parakeet.cpp to e53a253

Brings in the VAD-only GGUF loader.

Assisted-by: Claude Code:claude-sonnet-5-5 [git] [gh]

* docs(gallery): link the parakeet.cpp VAD docs instead of the merged PR

Assisted-by: Claude Code:claude-sonnet-5-5 [git]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 09:34:40 +02:00
mudler-agentandEttore Di Giacinto 99043b442c feat(router): route with native decision models (#12449)
* fix(schema): preserve SystemOne image inputs

Assisted-by: OpenAI

* test(schema): follow Ginkgo conventions for decision inputs

Assisted-by: OpenAI

* feat(llama-cpp): dispatch native decisions through Score

Upgrade the stock dependency and reconcile Score/TTS patches. Reuse native decision parsing, tasks, formatting and response-reader cleanup; preserve ordinary scoring admission and guard older dependencies.

Assisted-by: OpenAI

* refactor(systemone): share request and model validation

Assisted-by: OpenAI:gpt-5

* fix(systemone): preserve HTTP wire-byte validation limit

Keep structural validation separate from the serialized internal request bound so HTML escaping cannot reject valid HTTP payloads.

Assisted-by: OpenAI:gpt-5

* feat(systemone): bound images and account native decisions

Preserve public wire limits independently from router serialization. Reject unsupported NER images, map native request/capability errors, and stamp explicit usage once. Advertise decisions for stock llama-cpp.

Assisted-by: OpenAI

* fix(systemone): record usage on registered native route

Exercise real registration and billing with a mock native backend. Reject empty native responses, malformed image URLs, trailing JSON, and wire overflow including whitespace.

Assisted-by: OpenAI

* feat(router): add lazy native decision transport

Bind named models through internal ModelSystemOne calls with shared validation and bounded abandoned operations. Remove request and echoed-error contents from decision traces.

Assisted-by: OpenAI:gpt-5

* feat(router): classify overlapping policies with native decisions

Ask independent noul questions, validate probabilities and preserve first-superset routing. Wire the central factory with config-sensitive invalidation and cancellation-safe resolution. Document native framing and bounded operation limits.

Assisted-by: OpenAI:gpt-5

* feat(gallery): add pinned Julia-1 native decision model

Add a separate text-only llama-cpp Q8 entry with pinned Apache-2.0 source provenance and checksum. Installed using the gallery installer and exercised choice, score and noul on CPU.

Assisted-by: OpenAI

* test(router): verify native decisions through central factory

Add an opt-in real-model Ginkgo integration covering the native Go loader and C++ transport, token usage, independent overlapping labels, and candidate selection. Document owned-server execution and the intentionally non-quality threshold.

Assisted-by: Codex:gpt-5

* fix(llama-cpp): align upstream pin and preserve decision signatures

Advance to bed0a856 without losing the automated upstream bump. Detect full-request fill_task support at compile time and forward every question for Nimble framing while retaining the earlier native signature. Preserve reconciled SCORE/TTS patches; add standalone compatibility coverage.

Assisted-by: Codex:gpt-5

* feat(gallery): add native decision family defaults

Pin Laya, Kev-4B, lev, OpenJev and Nimble artifacts. Verify Laya/Kev/lev gallery installs and CPU contracts on both native pins; clearly mark OpenJev/Nimble runtime validation pending and their noncommercial licenses.

Assisted-by: OpenAI

* docs(decisions): clarify integrated Nimble prerequisite

Record the exact combined backend pin while retaining pending OpenJev and Nimble installation/runtime validation status.

Assisted-by: Codex:gpt-5

* fix(gallery): indent native decision model sequences

Match repository yamllint indentation for Laya, Kev, lev and OpenJev list fields. Parsed gallery data is unchanged; reproduce CI gallery lint failure before the whitespace-only fix and pass the same command afterward.

Assisted-by: Codex:gpt-5

* docs(decisions): record OpenJev and Nimble CPU validation

Record gallery installation, checksum/metadata verification and multiquestion native smoke results on bed0a856. Retain noncommercial and text-only limitations without accuracy or deterministic-output claims.

Assisted-by: OpenAI

* fix(ui): expose native Decisions router classifiers

Select classifier models using metadata-driven capability routing, retain tuned thresholds, and validate native decision selections before saving. Cover both native backends and create/save/reopen in the real React editor.

Assisted-by: Codex:gpt-5

* fix(router): exclude aliases from native decision discovery

Check the originally named config before advertising native Decisions eligibility. Retain target capability inheritance for ordinary generation aliases. Exercise the actual capabilities endpoint with native models on both backends, aliases, and disabled models.

Assisted-by: Codex:gpt-5

* feat(systemone): share bounded multimodal input validation

Preserve text wire limits while admitting bounded PNG/JPEG decision input. Share collection and header validation across internal and public callers and keep the native runner response budget independent.

Assisted-by: OpenAI:API-assistant

* fix(systemone): bound admission lifetimes and validate complete images

Retain shared admission leases through actual work completion, including abandoned internal operations. Decode bounded image pixels, cap public native responses before usage stamping, and preserve oversized malformed text status precedence.

Assisted-by: OpenAI:API-assistant

* fix(router): classify images before media fetching

Preserve ordered structured probes for native decisions. Defer OpenAI
media preparation until routing selects the served model, so rejected
decision URLs cannot trigger downloads before shared validation.

Guard direct image collection with context-aware shared admission.
Keep text classifiers and embedding caches from discarding image input.
Retain fail-closed classifier configuration and runtime fallback policy.

Add middleware, typed-content, admission, cancellation and cache tests.

Assisted-by: OpenAI:API-assistant

* fix(router): bound extraction before serialization

Check probe budgets before copying text or marshaling message state.
Count JSON escaping so oversized internal inputs fail before allocation.

Preserve typed Anthropic blocks through selected-model conversion and
fallback. Keep retry coverage in Ginkgo without global test registration.

Assisted-by: OpenAI

* fix(router): bound supported probe serialization

Arbitrary structs can bypass the probe budget through pointer marshalers,
string tags, and promoted fields. Accept concrete chat schema types and
plain JSON values instead of emulating arbitrary struct serialization.

Budget escaped direct prompts before marshaling so raw length cannot hide
serialized expansion. Preserve runtime fallback and reject oversized
input before invoking the decision runner.

Add Ginkgo allocation, boundary, and marshaler invocation regressions.
Six-package tests, three-package race tests, and full-T2 delta lint pass.

Assisted-by: OpenAI:GPT-5 golangci-lint

* feat(decisions): enable bounded OpenJev images

Validate native decision images before permissive media parsing and pixel
allocation. Require both decision image support and a vision projector;
missing or audio-only projectors cannot silently become text decisions.

Pin the OpenJev Q8 projector and document its license and disk footprint.
Add native safety tests, canonical limit parity, gallery and load-option
checks, and a reproducible CPU direct-RPC contrasting-image smoke.

Assisted-by: OpenAI:GPT-5

* fix(decisions): reject incomplete image streams

stb accepts corrupt PNG Adler checksums and truncated JPEG scans.
Use bounded zlib validation and strict libjpeg decoding before parsing.
Keep dimension and aggregate pixel checks ahead of decoder allocations.

Wire decoder dependencies into native builds and runtime packaging.
Add regressions for appended EOI and embedded marker bypasses.

Assisted-by: OpenAI:GPT-5

* fix(ci): gate native decision image validation

Run the decoder security tests outside the stdlib-only native suite.
Fetch vendor headers at the backend pin and provision decoder dependencies.
Gate Go limit parity and production CMake wiring without model downloads.

Assisted-by: OpenAI:GPT-5

* test(decisions): cover multimodal public API paths

Exercise shared image contracts through the registered HTTP routes and
external mock backend. Add opt-in cached gallery installation and real
OpenJev image decisions through SystemOne and both routing APIs.

Assisted-by: Codex:gpt-5

* test(decisions): assert isolation and cache bypass

Observe external RPC calls and compare complete classifier history.
Winner-only and cache-miss checks could hide dropped history or cache use.

Give real inference its own application and model directory so shared
backend mappings and loaded processes cannot affect mixed suite order.

Assisted-by: OpenAI:ChatGPT

* test(decisions): isolate fixture globals

Disable optional global services in the isolated HTTP fixture and register
cleanup before setup assertions. Verify meter provider identity survives
fixture creation and destruction.

Snapshot observed usage before assertions so failures cannot retain the
mutex. Require a successful usage stamp before checking error responses.

Assisted-by: Codex:gpt-5 golangci-lint

* fix(application): honor optional telemetry controls

Skip failover gauge registration when metrics are disabled. Register
against the application meter rather than looking up the global provider.

Allow embedders to retain the bounded routing log without billing stats.
Keep the existing default when stats are disabled. The isolated HTTP
fixture uses this option without losing its native router assertions.

Assisted-by: Codex:gpt-5 golangci-lint

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 09:34:21 +02:00
localai-org-maint-botandmudler f035746db9 docs: ⬆️ update docs version mudler/LocalAI (#12458)
⬆️ Update docs version mudler/LocalAI

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-04 08:48:14 +02:00
localai-org-maint-botandmudler e0e21de3e9 chore: ⬆️ Update CrispStrobe/CrispASR to 199de52d9068aeba366dd95d09a96bb45b2a7a14 (#12459)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-04 08:48:02 +02:00