Commit Graph
8426 Commits
Author SHA1 Message Date
localai-org-maint-botandmudler ec8d92b69d chore: ⬆️ Update CrispStrobe/CrispASR to 3f4ca9372869fe82a36f6266fa20050d8c35bfa8 (#12556)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-08 09:57:08 +02:00
localai-org-maint-botandlocalai-org-maint-bot f4a09b4b77 fix(vram): reject oversized GGUF metadata before allocation (#12560)
Upgrade gguf-parser-go to v0.26.3 so string lengths are checked against the remaining file size before allocation. Master AIO CI crashed in the background gallery warmer when v0.25.0 tried to allocate several terabytes; panic recovery cannot catch a fatal runtime OOM.

Check both overflow-sized and file-exceeding strings through the remote reader, and document the size-only estimate fallback.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-08 09:56:59 +02:00
mudler-agentandEttore Di Giacinto 895d50385f fix(distributed): converge model configs across frontends (#12558)
* fix(galleryop): announce model changes before the preload

After a gallery install or delete, the replica that ran it replaced its
config loader, then preloaded every installed model, and only then
published the models invalidation. The preload does remote lookups and
checksums for each model, so on a large models directory peers learned
about the change minutes after the originator listed it. When the
preload failed or the operation was cancelled, the event was never sent.

Publish the invalidation, and apply the delete lifecycle, as soon as
the loader holds the new set. The preload still runs afterwards with
its own error handling. Its failure is reported on the operation, but
it no longer rolls back a deletion that peers have already applied.

Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): resync model configs from the models directory

Frontends refresh their model configs only when a models invalidation
arrives on NATS. NATS keeps no history, so a frontend that is
disconnected when the message is published never applies the change.
It keeps serving the old config, for example an alias that points at
the previous model, until some later change happens to touch it.

Each frontend now reruns the peer reconcile against the shared models
directory after every NATS reconnect, and every
--model-config-resync-interval (default 30s) when a config file
changed. The pass names no model, so only models whose file changed
get a revision transition, and an unchanged directory costs one read
of the config files.

The reconcile replaced the whole loader with a parse of the models
directory, which dropped models loaded with --config-file and
published a deletion revision for them. Configs defined outside the
directory are now kept, both there and after a gallery install.

Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(nodes): stop stale frontends from retargeting alias rules

A scheduling rule keyed by an alias derives its target from the alias
mapping of the frontend that reads it. Every frontend keeps its own
copy of the model configs, so after an alias is repointed a frontend
that has not reloaded it still resolves the old target. Two frontends
then rewrote the rule's stored target_model against each other on
alternate reconciler ticks, and the outdated one scaled up the model
the alias used to point at.

The registry already records the accepted config revision of each
model. A frontend now derives a rule's target from its own alias
mapping only when its config revision for the rule's name matches
that record. Otherwise it keeps the stored target_model: it neither
writes the column nor reconciles replicas of the old target. The check
reads the database only for a rule whose stored and derived targets
differ. With no accepted revision on record, the old behaviour stays.

Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-08 09:56:12 +02:00
localai-org-maint-botandmudler 89a0955b5b chore(model-gallery): ⬆️ update checksum (#12557)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 23:09:59 +02:00
Plamen K. Kosseff c03acb0647 feat(models): report per-model disk usage in API and WebUI (#12551)
New admin endpoint GET /api/models/storage reports disk usage of the
installed models: each model's files and sizes, files shared between
models counted once, and configured files that are missing from disk.
The models WebUI page shows a usage summary, each model's file table,
and the full file list with per-file status.

Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
2026-10-07 22:03:48 +02:00
mudler-agent cd604b5c5b fix(distributed): bound, stop and cancel model loads with leases and worker operations (#12524)
Fence model load jobs by generation, lease them on the database clock, bound the work on the worker with operations and a process-group watchdog, and add one stop path with a load-cancel API. See the pull request for the design, the rolling upgrade notes and the test evidence.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-07 16:59:56 +02:00
localai-org-maint-botandlocalai-org-maint-bot c525ad16e8 fix(audio): record transcription usage and traces (#12546)
* fix(audio): record transcription usage and traces

Count successful transcription requests by model, including streams.
Keep token counts at zero because transcription exposes no token usage.

Capture multipart API trace metadata without reading uploaded audio.
Do not record failed transcription or client writes as successful usage.

Assisted-by: Codex:gpt-6-astra

* fix(audio): preserve aliases in streaming usage

Streaming transcription records the resolved target as its usage model.
Pass the requested name so JSON and SSE requests share the alias bucket.

Assisted-by: Codex:gpt-6-astra

* test(http): check multipart reader close errors

Assert successful reader cleanup to satisfy the errcheck CI gate.

Assisted-by: Codex:gpt-6

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-07 16:57:23 +02:00
localai-org-maint-botandEttore Di Giacinto 6b794651a4 feat(diarization): return sound events with include_sounds (#12544)
* feat(diarization): return sound events with include_sounds

A client that wants text, speakers, voice prints and sound events had to
make a diarization call and a separate sound call. Add an include_sounds
request field to /v1/audio/diarization that adds a sounds array of closed
events {start, end, label, confidence}, in seconds.

The parakeet-cpp backend runs a tagger-only scene stream over the clip,
the same stream and thresholds the live path uses, so a clip gives the
same events offline and live. A model with no sound_model companion, or a
backend that does not report sound events, fails with 501 and the stable
code include_sounds_unsupported instead of an empty list. The proto
carries sounds_included so an empty list still means "nothing heard".

The localai-proxy backend forwards the field. Swagger, docs and the
e2e mock backend are updated.

Assisted-by: Claude:claude-sonnet-5-5 [protoc swag go]

* feat(gallery): add parakeet-cpp-multilingual-diarization-speakers-sounds

Same as parakeet-cpp-multilingual-diarization-speakers (TDT 0.6B v3,
Nemotron-3-Diarization, WeSpeaker) plus a CED-Tiny sound_model, so one
model name serves /v1/audio/diarization with include_text,
include_speaker_profiles and include_sounds. It declares the
sound_classification usecase like the realtime scene entries.

Assisted-by: Claude:claude-sonnet-5-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-07 14:10:45 +02:00
Aniruddh Krovvidi a3d555653c fix(model): stop listing a backend that exited on its own (#12497)
Assisted-by: Claude:claude-fable-5-1

Signed-off-by: Aniruddh Krovvidi <akrovvidi05@gmail.com>
2026-10-07 09:28:49 +02:00
localai-org-maint-botandmudler 5537f4b1ef chore: ⬆️ Update vllm-project/vllm cu130 wheel to 0.31.0 (#12502)
⬆️ Update vllm-project/vllm cu130 wheel

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:27:43 +02:00
localai-org-maint-botandmudler e38d9c9be7 chore: ⬆️ Update PrismML-Eng/llama.cpp to 6bfcd79a2d426abcd2b50e3c2d09ae2225e70a17 (#12503)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:27:23 +02:00
localai-org-maint-botandmudler d6238f60cd chore(model-gallery): ⬆️ update checksum (#12505)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:27:07 +02:00
localai-org-maint-botandmudler 0a9fb07574 chore: ⬆️ Update localai-org/voice-detect.cpp to cf9e1d5641c710bf7d6fc9642f9b4cf7d2bde9b6 (#12506)
⬆️ Update localai-org/voice-detect.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:26:51 +02:00
localai-org-maint-botandmudler b554fc55a8 chore: ⬆️ Update CrispStrobe/CrispASR to e79a671fce388b50d3e2eb5fced1581140a29594 (#12531)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:26:35 +02:00
localai-org-maint-botandmudler c73977f5b6 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 739839a5c922cf9663988912004912e9f0671d7d (#12532)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:26:16 +02:00
localai-org-maint-botandmudler a1b0648f19 chore: ⬆️ Update ggml-org/whisper.cpp to d1be6fde11ac6e0407606b4e42fe72d34add8037 (#12534)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:21:02 +02:00
localai-org-maint-botandmudler f201055812 chore: ⬆️ Update mudler/parakeet.cpp to 9a28a3c1f7fb7d89505646b4d437033e4163f6e1 (#12511)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:20:39 +02:00
localai-org-maint-botandmudler 6d1c8d7f71 chore: ⬆️ Update leejet/stable-diffusion.cpp to a1ded76da5818803fca97a3b433669ef727d32cf (#12533)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:20:16 +02:00
localai-org-maint-botandEttore Di Giacinto 217038d4fe fix(agents): share the collection list across frontends (#12536)
With the PostgreSQL vector engine, each frontend kept the collections it
had opened in an in-memory map and answered "collection not found" for
any other. A collection created through one frontend was unknown to the
others until they restarted, and each frontend listed a different set.

Wrap the in-process collections backend in distributed mode so that the
database is the source of truth:

- lists come from the registry,
- a lookup miss checks the registry before it returns 404, and opens a
  collection that exists there (once per name, even under concurrency),
- a cached collection that left the registry is dropped and closed, with
  a re-check at most every 5 seconds,
- create and reset publish an event on the existing collection
  invalidation subject, so other frontends re-check at once.

Without the postgres engine, or outside distributed mode, nothing changes.

Assisted-by: Claude:claude-sonnet-5-5 go

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 22:45:23 +02:00
e2faec2bf2 fix(distributed): free only the named model on model.unload (#12537)
The worker answered model.unload by calling Free on the first backend
process in its map, whatever model the request named. Every idle
scale-down or LRU eviction of one model could therefore empty another
model's backend on the same node. LocalAI still counted that model as
loaded, so its next request failed. A parakeet diarization model then
returned 501 "speaker profiles require a loaded speaker encoder" until
someone reloaded it by hand.

Resolve the target from the model name (all replicas), prefer an
address when the request carries one, and free nothing for an unknown
model.

Also let parakeet-cpp Diarize check the diarization model before the
speaker-profile capability. A backend with nothing loaded now answers
FailedPrecondition, which LocalAI treats as a stale replica and
reloads, instead of a final Unimplemented.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 22:44:25 +02:00
localai-org-maint-botandEttore Di Giacinto 73b4acdca7 feat(gallery): add parakeet-cpp-multilingual-diarization-speakers (#12535)
The English-only 110M ASR in parakeet-cpp-nemotron-3-diarization-asr-speakers
garbles other languages. Add a bundle that pairs Parakeet TDT 0.6B v3
(25 European languages) with the Nemotron-3-Diarization model and the
WeSpeaker ResNet34 encoder, so one /v1/audio/diarization call with
include_text and include_speaker_profiles returns turns, multilingual text
and one voice-print embedding per speaker.

Assisted-by: Claude Code:claude-sonnet-5-5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 22:33:50 +02:00
mudler-agentandEttore Di Giacinto 690f3994b0 feat(auth): let users pause and resume their API keys (#12521)
A key can now be paused until the owner resumes it, or until a given
time. A paused key is rejected by validation before last_used is
updated, and a pause time that has passed lifts the pause by itself.
Existing keys stay active.

PATCH /api/auth/api-keys/:id takes {"disabled": bool, "paused_until":
RFC 3339 string or null}. Only the key owner can change it, and a
paused_until in the past is rejected. The key list returns the pause
fields. The Account page gets a Pause and Resume button for each key
and a Paused badge that shows the resume time.

Assisted-by: Claude Code:claude-sonnet-5-5

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 18:25:21 +02:00
fmterrors 1cd96e496b chore: fix some function names in comment (#12517)
Signed-off-by: fmterrors <fmterrors@outlook.com>
2026-10-06 17:26:03 +02:00
localai-org-maint-botandmudler e34847123e chore: ⬆️ Update ikawrakow/ik_llama.cpp to 3d27f5bb6348c4d79edded63b03e9a93efb2f80e (#12504)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-06 08:57:29 +02:00
Ettore Di Giacinto 7467b51281 test(gallery): expect speaker_recognition on the bundles with a voice component
The parakeet-cpp backend now answers VoiceEmbed and VoiceVerify, so the
pinning test asserts the usecase on the entries that declare it instead of
its absence. Document voice_verify_threshold.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:48:20 +00:00
Ettore Di Giacinto 1f22b16092 chore(parakeet-cpp): bump parakeet.cpp to 2f9e8ea
parakeet_capi_speaker_embed_pcm and parakeet_capi_free_floats, which
VoiceEmbed binds, come from this commit.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:48:20 +00:00
Ettore Di Giacinto 20f3c72d58 feat(parakeet-cpp): add the voice_verify_threshold model option
VoiceVerify used a fixed distance of 0.5 when the request had none. Read
voice_verify_threshold, a distance in (0, 2), from the model options and
keep 0.5 as the default. The real-library spec now cuts clips from a
two-voice recording and checks the bundle embedding size, the encoder
identity, determinism and the same-voice versus different-voice distance.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:48:20 +00:00
Ettore Di Giacinto e9af44b90b docs(voice): document a parakeet-cpp bundle as the embedding model
Describe the bundle as an embedding model for /v1/voice/*, the realtime
voice_recognition stage, and the limits: identify and plain verify only,
a 256-dimension space shared with voice-detect-wespeaker-resnet34, and the
libparakeet.so symbol it needs.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:46:14 +00:00
Ettore Di Giacinto ce4099ea38 feat(gallery): declare speaker_recognition on parakeet-cpp bundles
The small and standard bundles have a voice component, and
parakeet-cpp-realtime-scene-speakers loads a speaker encoder, so they can
serve /v1/voice/* and the realtime voice_recognition stage. Add the usecase
to their known_usecases and pin it in the gallery tests.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:46:06 +00:00
Ettore Di Giacinto fa25493c9e feat(config): let parakeet-cpp declare the speaker_recognition usecase
parakeet-cpp now implements VoiceEmbed and VoiceVerify, so list them in
its backend capabilities together with the speaker_recognition usecase.
The usecase is not guessed, because most parakeet models have no
speaker encoder: a bundle declares it in known_usecases, which is
enough for /v1/voice/* and for voice_recognition.model.

Add a realtime pipeline test where one bundle model names the vad,
transcription, sound_detection and voice_recognition stages and the
gate resolves the speaker through the backend's VoiceEmbed.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:46:06 +00:00
Ettore Di Giacinto 3beb2db0c9 feat(parakeet-cpp): serve VoiceEmbed and VoiceVerify from the speaker encoder
Bind parakeet_capi_speaker_embed_pcm and parakeet_capi_free_floats with a
Dlsym probe and implement VoiceEmbed on the speaker context, so a model
with a speaker_model or a bundle speaker_component can serve the
realtime voice_recognition stage and /v1/voice/*. The response carries
the sha256 identity of the encoder weights.

VoiceVerify embeds both clips and compares them by cosine distance. It
refuses anti_spoofing because there is no such head.

A libparakeet.so without the symbols answers Unimplemented, and a model
without a speaker encoder answers FailedPrecondition.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:46:06 +00:00
Ettore Di Giacinto 04df10a082 docs(voice): document speaker naming through a bundle component
Cover speaker_component: and speaker_tag: in the audio-to-text option
tables and bundle section, add a section to the voice recognition page
that explains which registered voices a bundle component receives, and
update the speaker_name condition in the realtime and diarization pages.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:45:14 +00:00
Ettore Di Giacinto 1d117292b1 feat(voice): name registered speakers through a bundle speaker component
Naming read only speaker_model:, so a parakeet-cpp bundle that sets
speaker_component:voice never received the registered voices, in live
transcription or in diarization.

Resolve speaker_component: when speaker_model: is not set. A component
is not an encoder file, so no file-name tag is derived for it: voices
with a hash or family identity and untagged voices reach the backend,
and a voice with only a file-name tag matches through the new optional
speaker_tag: alias. speaker_model: still wins, including when it points
at the bundle file.

The bundle gallery entries do not declare speaker_recognition: that
usecase means the VoiceEmbed and VoiceVerify RPCs, which the backend
does not serve, and it selects the default model for /v1/voice/*. The
pinning test now says so.

Assisted-by: Claude Code:claude-sonnet-5-5 go golangci-lint
2026-10-05 23:45:14 +00:00
Ettore Di Giacinto a2c498359b fix(voice): keep one voice store per embedding dimension
The in-memory local-store rejects vectors of another size than the ones
it holds. With 192-value voices registered, adding a 256-value voice
failed, and identifying with a 256-value probe returned an error.

The registry now keeps one store per embedding dimension. The first
dimension seen uses the configured store name, so single-encoder
instances are unchanged. Later dimensions use "<name>-<dim>". Identify
searches only the store of the probe size and returns no match when no
voice of that size exists. Forget finds the store from the stored
embedding. A name may hold one voice per encoder.

Assisted-by: Claude Code:claude-sonnet-5-5 golangci-lint
2026-10-05 23:42:49 +00:00
localai-org-maint-botandEttore Di Giacinto 68c980f3cd chore(deps): bump localrecall to v0.6.6 (#12508)
* chore(deps): bump localrecall to v0.6.6

LocalRecall v0.6.6 closes the PostgreSQL connection pool when a
collection fails to open. Before, each failed collection create left
its pool open. With a failing embedding model, every retry leaked one
more pool until PostgreSQL refused new clients.

LocalAI gets LocalRecall through LocalAGI, so this raises the indirect
requirement directly instead of waiting for a LocalAGI bump.

The release also limits each collection pool to 4 connections by
default. POSTGRES_POOL_MAX_CONNS changes the limit.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

* docs(agents): document the PostgreSQL pool size limit

LocalRecall v0.6.6 limits the connection pool of each collection to 4
connections and reads POSTGRES_POOL_MAX_CONNS to change it. Document
the variable next to the other PostgreSQL settings of the embedded
store.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-06 00:51:53 +02:00
localai-org-maint-botandmudler a4acf1dd40 feat(swagger): update swagger (#12500)
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 22:36:03 +02:00
mudler-agentandEttore Di Giacinto 0f8d5640ba chore(model-gallery): aggregate reviewed model additions (#12501)
Collect the validated model-gallery entries from 48 open pull requests onto current master. Exclude #12325 because its variant groupings violate gallery invariants, and exclude #12353 because the image model is declared as an unsupported llama.cpp chat model.

Source PRs: #12303, #12311, #12321, #12322, #12327, #12330, #12332, #12334, #12338, #12340, #12352, #12354, #12357, #12359, #12360, #12367, #12369, #12370, #12371, #12376, #12393, #12394, #12396, #12398, #12400, #12409, #12422, #12423, #12429, #12431, #12432, #12434, #12435, #12444, #12448, #12450, #12454, #12457, #12464, #12468, #12473, #12476, #12478, #12483, #12488, #12489, #12492, #12494.

Assisted-by: nib:gpt-5.6-sol

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-05 22:35:35 +02:00
mudler-agentandEttore Di Giacinto 79a7631cc5 feat(parakeet-cpp): encoder fingerprint for speaker naming, VAD trim and word filter options, pin bump (#12491)
* chore(parakeet-cpp): bump parakeet.cpp to 2de154c

Brings in the speaker registry encoder fingerprint, the VAD segment trim
and the opt-in word filter, a fix for a per-call thread count that stayed
set on the process-wide backend after a Silero VAD pass, and bundle
components loaded from memory.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): encoder fingerprint for speaker naming, vad_trim and guard_* options

Speaker naming. A registered voice now carries the encoder that made it:
the embedding family (voicedetect:<arch>:<name>:<dim>) and the sha256 of
the weights. The backend reports the family of the loaded speaker model in
Status, voice enrollment from speaker_profiles stores it as encoder_family
(old entries load without it), and the registry sent to parakeet.cpp is
built with parakeet_capi_speaker_registry_add_embedding_fp. The library
then refuses a registry of another encoder family and the error names both
families; another quantization of the same family only warns. A voice with
only a weights hash gets the loaded family when the hashes are equal.

Voices without a fingerprint (registered from audio: libvoicedetect cannot
report one) keep the file-name rule and are used with a warning. The
library cannot mix them with fingerprinted voices in one registry, so a
request that has any uses the old registry for all. speaker_strict:true
drops them instead. A library without the symbols behaves as before.

Transcription. vad_trim (seconds, 0 keeps the whole cuts) goes through the
VAD options JSON, so it reaches /v1/vad and the segmenter. The guard_*
options guard_min_local_conf, guard_local_radius and guard_drop_punct_only
turn on the word filter through parakeet_capi_transcribe_path_json_with,
or through the segmenter with vad:true. They are off by default, bad
values fail the load, and a library without the symbol fails it with a
clear message. The dropped word count is logged at debug level.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-05 15:57:16 +02:00
localai-org-maint-botandlocalai-org-maint-bot 3d586cc3c5 fix(audio-cpp): forward voice reference transcripts (#11997)
Saved voices send ref_text, but Fish Audio requires reference_text.
Derive the canonical parameter while preserving explicit overrides.
Both TTS modes use the shared builder.

Add regression cases and document the parameter alias.

Assisted-by: Codex:gpt-6

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 11:05:43 +02:00
localai-org-maint-botandmudler e81e180ce2 chore(website): refresh the counters (#12490)
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 10:27:05 +02:00
localai-org-maint-botandmudler 05b20bd7ab chore: ⬆️ Update mudler/parakeet.cpp to 0cca477249ffb16c1623fb947d5bac0624961d41 (#12460)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 08:04:04 +02:00
localai-org-maint-botandmudler 0a72af4e95 chore: ⬆️ Update localai-org/voice-detect.cpp to bca46bcbc2fe68169c2d7c414e9b290a7cb89911 (#12486)
⬆️ Update localai-org/voice-detect.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 08:03:37 +02:00
localai-org-maint-botandmudler d79489adb3 chore: ⬆️ Update localai-org/ced.cpp to 736a4ee46a31d4ff38b41add65ce0f2cf1aaa05f (#12487)
⬆️ Update localai-org/ced.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 08:03:24 +02:00
localai-org-maint-botandmudler 7310887e26 chore(model-gallery): ⬆️ update checksum (#12485)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 01:40:12 +02:00
Stefan Walcz 5f5feea15a fix(agentpool): keep the conversation in the agent web chat (#12410)
* feat(agents): keep the web chat history between messages

The agent chat endpoint ran every message as a fresh job
(ag.Ask(WithText(message))), so a follow-up such as "add two days to
item 3 and recalculate" never saw the answer it referred to. Agents
then rebuilt their reply from scratch instead of changing it.

Use the agent's own conversation tracker, as the Telegram and Slack
connectors already do: send the earlier turns with the new message and
record successful answers. Failed, cancelled or empty runs are not
recorded, and the tracker drops a conversation after the agent's
last_message_duration of inactivity. The distributed (NATS) chat path is
unchanged.

Assisted-by: Claude:claude-opus-5-5 ginkgo
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>

* fix(agents): keep each web chat conversation's history separate

Review on this PR: the first version kept the history in the agent's
conversation tracker under one fixed key per agent. The web UI keeps
several conversations per agent (New Chat, switching, Clear) and did not
send which one a message belongs to, so the hidden history mixed them:
a new chat received the previous chat's turns, and Clear only cleared
the screen.

The client now sends the earlier turns of the conversation it is
showing as `history` with POST /api/agents/:name/chat, and the server
keeps no web chat history of its own. Each conversation only ever sees
its own turns; New Chat and Clear start without history. The server
uses only user and assistant turns with text, bounded to the most
recent 40 turns and 64,000 characters. `history` is optional, so
existing callers keep the previous behaviour (no history); the
distributed (NATS) path does not forward it yet.

Tests: two conversations of one agent stay apart, an empty history
(New Chat or Clear) starts fresh, system/tool/empty turns are dropped,
the bounds keep the most recent turns; the UI helper that builds the
history from the visible messages has node --test coverage. Docs:
features/agents.md describes `history` and the web UI behaviour.

Assisted-by: Claude:claude-opus-5-5 ginkgo
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>

---------

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-05 01:38:29 +02:00
Stefan Walcz 9b0b7da96e fix(rocm): point *_TENSILE_LIBPATH at the architecture folder (#12411)
Newer ROCm releases split the rocBLAS and hipBLASLt TensileLibrary data
into one folder per GPU architecture (lib/hipblaslt/library/gfx1151/...).
The run.sh of the ROCm-capable backends exports ROCBLAS_TENSILE_LIBPATH and
HIPBLASLT_TENSILE_LIBPATH as the parent folder, and both libraries only look
directly in that folder, so they miss every kernel:

  rocblaslt error: Cannot read ".../hipblaslt/library/TensileLibrary_lazy_gfx1151.dat"
  hipModuleLoad failed: .../hipblaslt/library/Kernels.so-000-gfx1151.hsaco

and fall back to slower code paths.

rocm_tensile_dir keeps the folder when it has files at the top level (the
older flat layout) or several architecture folders, and descends into the
architecture folder when the bundle carries exactly one.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-05 01:38:09 +02:00
localai-org-maint-botandlocalai-org-maint-bot e70c409a93 fix(kokoros): default optional status fields (#12447)
The speaker encoder metadata field makes the Rust status initializer
incomplete. Use protobuf defaults for absent metadata and memory fields.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 01:37:46 +02:00
localai-org-maint-botandlocalai-org-maint-bot 4a72b2ac2e test(cli): transfer socket descriptor ownership (#12465)
Duplicate the socket descriptor before handing it to the activation helper.
Close the original file to prevent its finalizer from closing a reused
coverage descriptor after the helper closes its own file wrapper.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 01:37:29 +02:00
leilei3167andlocalai-org-maint-bot d90f703bd9 fix(xsysinfo): read process VRAM without proc children (#12482)
* fix(xsysinfo): read process VRAM without proc children

Kernels without CONFIG_PROC_CHILDREN have no task children file, so
ProcessVRAM dropped every DRM reading. When that file is missing, walk
child processes from /proc/<pid>/stat ppid links instead. Other read
errors still drop the reading.

Fixes #12481

Signed-off-by: leilei3167 <imleilei123@gmail.com>

* docs(system): describe the proc children fallback

Document VRAM reporting on kernels without CONFIG_PROC_CHILDREN.

Assisted-by: Codex:GPT-6

---------

Signed-off-by: leilei3167 <imleilei123@gmail.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 01:37:12 +02:00
Plamen K. Kosseff 104c2fb440 fix(oci): stage image downloads in a configured directory, not in /tmp (#12280)
* fix(oci): stage image downloads in a configured directory, not in /tmp

Problem:
- the image tar and its compressed layers staged in os.TempDir()
- /tmp is commonly a RAM-backed tmpfs far smaller than a backend image
- big installs failed with a full /tmp or silently ate RAM

Change:
- staging path resolved like the other storage paths
  (LOCALAI_DOWNLOAD_STAGING_PATH, default ${basepath}/downloading),
  threaded to the extractor as a download option
- every download works in its own subdirectory holding its tar and
  layers: nothing shared between concurrent downloads, one removal
  cleans a download up, a crash leaves one self-contained orphan
- callers that pass no staging directory keep the OS temp behavior

Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>

* fix(oci): handle staging cleanup errors

Log staging cleanup failures to satisfy errcheck. Test extraction with an
unavailable OS temp directory and verify that staging is removed.

Assisted-by: Codex:GPT-6
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>

---------

Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
2026-10-05 01:27:45 +02:00