Commit Graph
2273 Commits
Author SHA1 Message Date
localai-org-maint-botandEttore Di Giacinto 6b794651a4 feat(diarization): return sound events with include_sounds (#12544)
* feat(diarization): return sound events with include_sounds

A client that wants text, speakers, voice prints and sound events had to
make a diarization call and a separate sound call. Add an include_sounds
request field to /v1/audio/diarization that adds a sounds array of closed
events {start, end, label, confidence}, in seconds.

The parakeet-cpp backend runs a tagger-only scene stream over the clip,
the same stream and thresholds the live path uses, so a clip gives the
same events offline and live. A model with no sound_model companion, or a
backend that does not report sound events, fails with 501 and the stable
code include_sounds_unsupported instead of an empty list. The proto
carries sounds_included so an empty list still means "nothing heard".

The localai-proxy backend forwards the field. Swagger, docs and the
e2e mock backend are updated.

Assisted-by: Claude:claude-sonnet-5-5 [protoc swag go]

* feat(gallery): add parakeet-cpp-multilingual-diarization-speakers-sounds

Same as parakeet-cpp-multilingual-diarization-speakers (TDT 0.6B v3,
Nemotron-3-Diarization, WeSpeaker) plus a CED-Tiny sound_model, so one
model name serves /v1/audio/diarization with include_text,
include_speaker_profiles and include_sounds. It declares the
sound_classification usecase like the realtime scene entries.

Assisted-by: Claude:claude-sonnet-5-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-07 14:10:45 +02:00
localai-org-maint-botandmudler 5537f4b1ef chore: ⬆️ Update vllm-project/vllm cu130 wheel to 0.31.0 (#12502)
⬆️ Update vllm-project/vllm cu130 wheel

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:27:43 +02:00
localai-org-maint-botandmudler e38d9c9be7 chore: ⬆️ Update PrismML-Eng/llama.cpp to 6bfcd79a2d426abcd2b50e3c2d09ae2225e70a17 (#12503)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:27:23 +02:00
localai-org-maint-botandmudler 0a9fb07574 chore: ⬆️ Update localai-org/voice-detect.cpp to cf9e1d5641c710bf7d6fc9642f9b4cf7d2bde9b6 (#12506)
⬆️ Update localai-org/voice-detect.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:26:51 +02:00
localai-org-maint-botandmudler b554fc55a8 chore: ⬆️ Update CrispStrobe/CrispASR to e79a671fce388b50d3e2eb5fced1581140a29594 (#12531)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:26:35 +02:00
localai-org-maint-botandmudler c73977f5b6 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 739839a5c922cf9663988912004912e9f0671d7d (#12532)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:26:16 +02:00
localai-org-maint-botandmudler a1b0648f19 chore: ⬆️ Update ggml-org/whisper.cpp to d1be6fde11ac6e0407606b4e42fe72d34add8037 (#12534)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:21:02 +02:00
localai-org-maint-botandmudler f201055812 chore: ⬆️ Update mudler/parakeet.cpp to 9a28a3c1f7fb7d89505646b4d437033e4163f6e1 (#12511)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:20:39 +02:00
localai-org-maint-botandmudler 6d1c8d7f71 chore: ⬆️ Update leejet/stable-diffusion.cpp to a1ded76da5818803fca97a3b433669ef727d32cf (#12533)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-07 08:20:16 +02:00
e2faec2bf2 fix(distributed): free only the named model on model.unload (#12537)
The worker answered model.unload by calling Free on the first backend
process in its map, whatever model the request named. Every idle
scale-down or LRU eviction of one model could therefore empty another
model's backend on the same node. LocalAI still counted that model as
loaded, so its next request failed. A parakeet diarization model then
returned 501 "speaker profiles require a loaded speaker encoder" until
someone reloaded it by hand.

Resolve the target from the model name (all replicas), prefer an
address when the request carries one, and free nothing for an unknown
model.

Also let parakeet-cpp Diarize check the diarization model before the
speaker-profile capability. A backend with nothing loaded now answers
FailedPrecondition, which LocalAI treats as a stale replica and
reloads, instead of a final Unimplemented.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 22:44:25 +02:00
localai-org-maint-botandmudler e34847123e chore: ⬆️ Update ikawrakow/ik_llama.cpp to 3d27f5bb6348c4d79edded63b03e9a93efb2f80e (#12504)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-06 08:57:29 +02:00
Ettore Di Giacinto 1f22b16092 chore(parakeet-cpp): bump parakeet.cpp to 2f9e8ea
parakeet_capi_speaker_embed_pcm and parakeet_capi_free_floats, which
VoiceEmbed binds, come from this commit.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:48:20 +00:00
Ettore Di Giacinto 20f3c72d58 feat(parakeet-cpp): add the voice_verify_threshold model option
VoiceVerify used a fixed distance of 0.5 when the request had none. Read
voice_verify_threshold, a distance in (0, 2), from the model options and
keep 0.5 as the default. The real-library spec now cuts clips from a
two-voice recording and checks the bundle embedding size, the encoder
identity, determinism and the same-voice versus different-voice distance.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:48:20 +00:00
Ettore Di Giacinto 3beb2db0c9 feat(parakeet-cpp): serve VoiceEmbed and VoiceVerify from the speaker encoder
Bind parakeet_capi_speaker_embed_pcm and parakeet_capi_free_floats with a
Dlsym probe and implement VoiceEmbed on the speaker context, so a model
with a speaker_model or a bundle speaker_component can serve the
realtime voice_recognition stage and /v1/voice/*. The response carries
the sha256 identity of the encoder weights.

VoiceVerify embeds both clips and compares them by cosine distance. It
refuses anti_spoofing because there is no such head.

A libparakeet.so without the symbols answers Unimplemented, and a model
without a speaker encoder answers FailedPrecondition.

Assisted-by: Claude Code:claude-sonnet-5-5
2026-10-05 23:46:06 +00:00
mudler-agentandEttore Di Giacinto 79a7631cc5 feat(parakeet-cpp): encoder fingerprint for speaker naming, VAD trim and word filter options, pin bump (#12491)
* chore(parakeet-cpp): bump parakeet.cpp to 2de154c

Brings in the speaker registry encoder fingerprint, the VAD segment trim
and the opt-in word filter, a fix for a per-call thread count that stayed
set on the process-wide backend after a Silero VAD pass, and bundle
components loaded from memory.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): encoder fingerprint for speaker naming, vad_trim and guard_* options

Speaker naming. A registered voice now carries the encoder that made it:
the embedding family (voicedetect:<arch>:<name>:<dim>) and the sha256 of
the weights. The backend reports the family of the loaded speaker model in
Status, voice enrollment from speaker_profiles stores it as encoder_family
(old entries load without it), and the registry sent to parakeet.cpp is
built with parakeet_capi_speaker_registry_add_embedding_fp. The library
then refuses a registry of another encoder family and the error names both
families; another quantization of the same family only warns. A voice with
only a weights hash gets the loaded family when the hashes are equal.

Voices without a fingerprint (registered from audio: libvoicedetect cannot
report one) keep the file-name rule and are used with a warning. The
library cannot mix them with fingerprinted voices in one registry, so a
request that has any uses the old registry for all. speaker_strict:true
drops them instead. A library without the symbols behaves as before.

Transcription. vad_trim (seconds, 0 keeps the whole cuts) goes through the
VAD options JSON, so it reaches /v1/vad and the segmenter. The guard_*
options guard_min_local_conf, guard_local_radius and guard_drop_punct_only
turn on the word filter through parakeet_capi_transcribe_path_json_with,
or through the segmenter with vad:true. They are off by default, bad
values fail the load, and a library without the symbol fails it with a
clear message. The dropped word count is logged at debug level.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-05 15:57:16 +02:00
localai-org-maint-botandlocalai-org-maint-bot 3d586cc3c5 fix(audio-cpp): forward voice reference transcripts (#11997)
Saved voices send ref_text, but Fish Audio requires reference_text.
Derive the canonical parameter while preserving explicit overrides.
Both TTS modes use the shared builder.

Add regression cases and document the parameter alias.

Assisted-by: Codex:gpt-6

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 11:05:43 +02:00
localai-org-maint-botandmudler 05b20bd7ab chore: ⬆️ Update mudler/parakeet.cpp to 0cca477249ffb16c1623fb947d5bac0624961d41 (#12460)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 08:04:04 +02:00
localai-org-maint-botandmudler 0a72af4e95 chore: ⬆️ Update localai-org/voice-detect.cpp to bca46bcbc2fe68169c2d7c414e9b290a7cb89911 (#12486)
⬆️ Update localai-org/voice-detect.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 08:03:37 +02:00
localai-org-maint-botandmudler d79489adb3 chore: ⬆️ Update localai-org/ced.cpp to 736a4ee46a31d4ff38b41add65ce0f2cf1aaa05f (#12487)
⬆️ Update localai-org/ced.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-05 08:03:24 +02:00
Stefan Walcz 9b0b7da96e fix(rocm): point *_TENSILE_LIBPATH at the architecture folder (#12411)
Newer ROCm releases split the rocBLAS and hipBLASLt TensileLibrary data
into one folder per GPU architecture (lib/hipblaslt/library/gfx1151/...).
The run.sh of the ROCm-capable backends exports ROCBLAS_TENSILE_LIBPATH and
HIPBLASLT_TENSILE_LIBPATH as the parent folder, and both libraries only look
directly in that folder, so they miss every kernel:

  rocblaslt error: Cannot read ".../hipblaslt/library/TensileLibrary_lazy_gfx1151.dat"
  hipModuleLoad failed: .../hipblaslt/library/Kernels.so-000-gfx1151.hsaco

and fall back to slower code paths.

rocm_tensile_dir keeps the folder when it has files at the top level (the
older flat layout) or several architecture folders, and descends into the
architecture folder when the bundle carries exactly one.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-05 01:38:09 +02:00
localai-org-maint-botandlocalai-org-maint-bot e70c409a93 fix(kokoros): default optional status fields (#12447)
The speaker encoder metadata field makes the Rust status initializer
incomplete. Use protobuf defaults for absent metadata and memory fields.

Assisted-by: Codex:gpt-6

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-05 01:37:46 +02:00
mudler-agentandEttore Di Giacinto 3185b6fcf8 feat(parakeet-cpp): load bundle GGUF files, bundle gallery entries, pin bump (#12479)
* feat(parakeet-cpp): load bundle GGUF files and use their components by role

A bundle GGUF holds several models (ASR, VAD, diarization, sound events,
speaker encoder) in one file, each with its own licence. Detect a bundle
at load through parakeet_capi_bundle_components_json and open components
with parakeet_capi_load_component. The three symbols are probed together,
so an older libparakeet.so still loads plain files as before.

The only ASR component is the primary model; bundle_asr:<name> picks one
when there are several. A Silero VAD component of the primary bundle is
loaded without an option and serves /v1/vad and vad:true. The diar, ced
and voice components load on request: diar_component, sound_component and
speaker_component, or a companion option (diarization_model, sound_model,
speaker_model, vad_model) that names a bundle, even the model file itself.
vad_component picks a VAD component and implies vad:true.

A role the bundle cannot fill fails the load with the component list, and
a diarization or sound request on a model without that role names the
bundle components. Every existing option and single-file model behaves as
before.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* chore(parakeet-cpp): bump parakeet.cpp to 781a973

Brings in the bundle GGUF format and its C-API (parakeet_capi_load_component,
parakeet_capi_bundle_components_json, parakeet_capi_load_error).

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): gallery entries for the bundle GGUF files, docs

Add parakeet-cpp-bundle-small (338 MB: Parakeet TDT+CTC 110M, Nemotron-3-
Diarization, CED-Small, WeSpeaker ResNet34-LM, Silero VAD), -standard
(1.1 GB, Parakeet TDT 0.6B v3 instead of the 110M model) and
-moondream-redux (215 MB: packed Redux and Silero VAD, CPU only). One
install serves transcription, VAD, diarization, sound events and speaker
naming through the component options. The existing single-purpose entries
stay.

A bundle has no single licence, so the entries use license: other and
state the licence and credit of each component in the description, with
the upstream inconsistency of the CED licence. The docs get a section on
bundles in audio-to-text with the entries, the roles, the options and the
licence notice, and pointers from the VAD, diarization and sound
classification pages. A gallery test checks the file names, checksums,
usecases and options.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 20:33:13 +02:00
mudler-agentandEttore Di Giacinto ed4a3975be feat(parakeet-cpp): gallery entries for the VAD-only Moondream slices, pin bump (#12469)
* feat(parakeet-cpp): add gallery entries for the VAD-only Moondream slices

Add parakeet-cpp-vad-moondream-redux and parakeet-cpp-vad-moondream-ultra.
They install the VAD head of Moondream Redux and Ultra (Q8_0) as small
files of 10 MB and 6 MB, cut out of the full models without retraining,
for the VAD endpoint. The files cannot transcribe, and a transcription
request fails with a clear error.

The files load only with a parakeet.cpp build that has VAD-only GGUF
support (parakeet.cpp pull request 87). The backend pin must move to a
commit that includes it before these entries work in a released image.
The parakeet-cpp-vad entry keeps installing Silero.

The docs list the files with the size, load time and memory compared
with loading a whole model. A gallery test checks the usecase, the file
name and the checksum of each entry.

Assisted-by: Claude Code:claude-sonnet-5-5 [golangci-lint]

* chore(parakeet-cpp): bump parakeet.cpp to e53a253

Brings in the VAD-only GGUF loader.

Assisted-by: Claude Code:claude-sonnet-5-5 [git] [gh]

* docs(gallery): link the parakeet.cpp VAD docs instead of the merged PR

Assisted-by: Claude Code:claude-sonnet-5-5 [git]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 09:34:40 +02:00
mudler-agentandEttore Di Giacinto 99043b442c feat(router): route with native decision models (#12449)
* fix(schema): preserve SystemOne image inputs

Assisted-by: OpenAI

* test(schema): follow Ginkgo conventions for decision inputs

Assisted-by: OpenAI

* feat(llama-cpp): dispatch native decisions through Score

Upgrade the stock dependency and reconcile Score/TTS patches. Reuse native decision parsing, tasks, formatting and response-reader cleanup; preserve ordinary scoring admission and guard older dependencies.

Assisted-by: OpenAI

* refactor(systemone): share request and model validation

Assisted-by: OpenAI:gpt-5

* fix(systemone): preserve HTTP wire-byte validation limit

Keep structural validation separate from the serialized internal request bound so HTML escaping cannot reject valid HTTP payloads.

Assisted-by: OpenAI:gpt-5

* feat(systemone): bound images and account native decisions

Preserve public wire limits independently from router serialization. Reject unsupported NER images, map native request/capability errors, and stamp explicit usage once. Advertise decisions for stock llama-cpp.

Assisted-by: OpenAI

* fix(systemone): record usage on registered native route

Exercise real registration and billing with a mock native backend. Reject empty native responses, malformed image URLs, trailing JSON, and wire overflow including whitespace.

Assisted-by: OpenAI

* feat(router): add lazy native decision transport

Bind named models through internal ModelSystemOne calls with shared validation and bounded abandoned operations. Remove request and echoed-error contents from decision traces.

Assisted-by: OpenAI:gpt-5

* feat(router): classify overlapping policies with native decisions

Ask independent noul questions, validate probabilities and preserve first-superset routing. Wire the central factory with config-sensitive invalidation and cancellation-safe resolution. Document native framing and bounded operation limits.

Assisted-by: OpenAI:gpt-5

* feat(gallery): add pinned Julia-1 native decision model

Add a separate text-only llama-cpp Q8 entry with pinned Apache-2.0 source provenance and checksum. Installed using the gallery installer and exercised choice, score and noul on CPU.

Assisted-by: OpenAI

* test(router): verify native decisions through central factory

Add an opt-in real-model Ginkgo integration covering the native Go loader and C++ transport, token usage, independent overlapping labels, and candidate selection. Document owned-server execution and the intentionally non-quality threshold.

Assisted-by: Codex:gpt-5

* fix(llama-cpp): align upstream pin and preserve decision signatures

Advance to bed0a856 without losing the automated upstream bump. Detect full-request fill_task support at compile time and forward every question for Nimble framing while retaining the earlier native signature. Preserve reconciled SCORE/TTS patches; add standalone compatibility coverage.

Assisted-by: Codex:gpt-5

* feat(gallery): add native decision family defaults

Pin Laya, Kev-4B, lev, OpenJev and Nimble artifacts. Verify Laya/Kev/lev gallery installs and CPU contracts on both native pins; clearly mark OpenJev/Nimble runtime validation pending and their noncommercial licenses.

Assisted-by: OpenAI

* docs(decisions): clarify integrated Nimble prerequisite

Record the exact combined backend pin while retaining pending OpenJev and Nimble installation/runtime validation status.

Assisted-by: Codex:gpt-5

* fix(gallery): indent native decision model sequences

Match repository yamllint indentation for Laya, Kev, lev and OpenJev list fields. Parsed gallery data is unchanged; reproduce CI gallery lint failure before the whitespace-only fix and pass the same command afterward.

Assisted-by: Codex:gpt-5

* docs(decisions): record OpenJev and Nimble CPU validation

Record gallery installation, checksum/metadata verification and multiquestion native smoke results on bed0a856. Retain noncommercial and text-only limitations without accuracy or deterministic-output claims.

Assisted-by: OpenAI

* fix(ui): expose native Decisions router classifiers

Select classifier models using metadata-driven capability routing, retain tuned thresholds, and validate native decision selections before saving. Cover both native backends and create/save/reopen in the real React editor.

Assisted-by: Codex:gpt-5

* fix(router): exclude aliases from native decision discovery

Check the originally named config before advertising native Decisions eligibility. Retain target capability inheritance for ordinary generation aliases. Exercise the actual capabilities endpoint with native models on both backends, aliases, and disabled models.

Assisted-by: Codex:gpt-5

* feat(systemone): share bounded multimodal input validation

Preserve text wire limits while admitting bounded PNG/JPEG decision input. Share collection and header validation across internal and public callers and keep the native runner response budget independent.

Assisted-by: OpenAI:API-assistant

* fix(systemone): bound admission lifetimes and validate complete images

Retain shared admission leases through actual work completion, including abandoned internal operations. Decode bounded image pixels, cap public native responses before usage stamping, and preserve oversized malformed text status precedence.

Assisted-by: OpenAI:API-assistant

* fix(router): classify images before media fetching

Preserve ordered structured probes for native decisions. Defer OpenAI
media preparation until routing selects the served model, so rejected
decision URLs cannot trigger downloads before shared validation.

Guard direct image collection with context-aware shared admission.
Keep text classifiers and embedding caches from discarding image input.
Retain fail-closed classifier configuration and runtime fallback policy.

Add middleware, typed-content, admission, cancellation and cache tests.

Assisted-by: OpenAI:API-assistant

* fix(router): bound extraction before serialization

Check probe budgets before copying text or marshaling message state.
Count JSON escaping so oversized internal inputs fail before allocation.

Preserve typed Anthropic blocks through selected-model conversion and
fallback. Keep retry coverage in Ginkgo without global test registration.

Assisted-by: OpenAI

* fix(router): bound supported probe serialization

Arbitrary structs can bypass the probe budget through pointer marshalers,
string tags, and promoted fields. Accept concrete chat schema types and
plain JSON values instead of emulating arbitrary struct serialization.

Budget escaped direct prompts before marshaling so raw length cannot hide
serialized expansion. Preserve runtime fallback and reject oversized
input before invoking the decision runner.

Add Ginkgo allocation, boundary, and marshaler invocation regressions.
Six-package tests, three-package race tests, and full-T2 delta lint pass.

Assisted-by: OpenAI:GPT-5 golangci-lint

* feat(decisions): enable bounded OpenJev images

Validate native decision images before permissive media parsing and pixel
allocation. Require both decision image support and a vision projector;
missing or audio-only projectors cannot silently become text decisions.

Pin the OpenJev Q8 projector and document its license and disk footprint.
Add native safety tests, canonical limit parity, gallery and load-option
checks, and a reproducible CPU direct-RPC contrasting-image smoke.

Assisted-by: OpenAI:GPT-5

* fix(decisions): reject incomplete image streams

stb accepts corrupt PNG Adler checksums and truncated JPEG scans.
Use bounded zlib validation and strict libjpeg decoding before parsing.
Keep dimension and aggregate pixel checks ahead of decoder allocations.

Wire decoder dependencies into native builds and runtime packaging.
Add regressions for appended EOI and embedded marker bypasses.

Assisted-by: OpenAI:GPT-5

* fix(ci): gate native decision image validation

Run the decoder security tests outside the stdlib-only native suite.
Fetch vendor headers at the backend pin and provision decoder dependencies.
Gate Go limit parity and production CMake wiring without model downloads.

Assisted-by: OpenAI:GPT-5

* test(decisions): cover multimodal public API paths

Exercise shared image contracts through the registered HTTP routes and
external mock backend. Add opt-in cached gallery installation and real
OpenJev image decisions through SystemOne and both routing APIs.

Assisted-by: Codex:gpt-5

* test(decisions): assert isolation and cache bypass

Observe external RPC calls and compare complete classifier history.
Winner-only and cache-miss checks could hide dropped history or cache use.

Give real inference its own application and model directory so shared
backend mappings and loaded processes cannot affect mixed suite order.

Assisted-by: OpenAI:ChatGPT

* test(decisions): isolate fixture globals

Disable optional global services in the isolated HTTP fixture and register
cleanup before setup assertions. Verify meter provider identity survives
fixture creation and destruction.

Snapshot observed usage before assertions so failures cannot retain the
mutex. Require a successful usage stamp before checking error responses.

Assisted-by: Codex:gpt-5 golangci-lint

* fix(application): honor optional telemetry controls

Skip failover gauge registration when metrics are disabled. Register
against the application meter rather than looking up the global provider.

Allow embedders to retain the bounded routing log without billing stats.
Keep the existing default when stats are disabled. The isolated HTTP
fixture uses this option without losing its native router assertions.

Assisted-by: Codex:gpt-5 golangci-lint

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 09:34:21 +02:00
localai-org-maint-botandmudler e0e21de3e9 chore: ⬆️ Update CrispStrobe/CrispASR to 199de52d9068aeba366dd95d09a96bb45b2a7a14 (#12459)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-04 08:48:02 +02:00
mudler-agentandEttore Di Giacinto 6c718fdda7 feat(parakeet-cpp): VAD (Silero and the Moondream head), vad_model option and gallery entries (#12463)
* feat(parakeet-cpp): implement the VAD call and add the vad_model option

The backend now serves the VAD gRPC call (POST /vad and /v1/vad) with the
standalone VAD of libparakeet. It accepts a Silero VAD GGUF as the model
file, or an ASR model with a VAD head (Moondream Ultra and Redux). The
request audio is float32 PCM at 16 kHz; the response lists the speech
segments in seconds, like the silero-vad backend. A model with neither
fails the request with the library message.

The vad_threshold, vad_min_pause, vad_min_speech, vad_speech_pad and
vad_max_segment options tune the segmenter. Unset values keep the
defaults of the detector in use, and a bad value fails the load.

The vad_model option names a Silero GGUF, resolved against the models
directory like the other companion files. It lets any ASR model cut long
audio at pauses through parakeet_capi_transcribe_path_json_vad_with, and
it implies vad. vad:true alone still uses the model's own head.

The new symbols are probed like the existing optional ones. A library
without them still loads; the feature that needs one fails with a clear
message only when it is used.

Assisted-by: Claude:claude-sonnet-5-5 [go test]

* feat(gallery): add parakeet-cpp VAD entries and a v3 plus Silero example

Add VAD-only entries for the parakeet-cpp backend: the VAD heads of
Moondream Redux (packed, CPU) and Ultra (Q8_0), which share their files
with the existing ASR entries, and Silero VAD v6.2.3 as a GGUF (MIT,
Silero Team). The parakeet-cpp-vad entry installs Silero; it has no variants,
because variant ranking prefers the larger build that fits and these are
different detectors.

Add parakeet-cpp-tdt-0.6b-v3-silero-vad, a v3 entry that sets vad_model
so long audio is cut at pauses by Silero.

The Silero GGUF entries point at the intended Hugging Face URL of the
file; the existing silero-vad entries are unchanged. A test checks the
usecases, the shared files and the default entry and the vad_model reference.

Assisted-by: Claude:claude-sonnet-5-5 [go test]

* docs: describe parakeet-cpp VAD and the vad_model option

Document the VAD endpoint on the parakeet-cpp backend (Silero GGUF and
the VAD heads of Moondream Ultra and Redux), the vad_* tuning options,
and the vad_model option that lets an ASR model without a VAD head cut
long audio with Silero.

Assisted-by: Claude:claude-sonnet-5-5

* chore(parakeet-cpp): bump parakeet.cpp to 6165e3d

Pin the release that adds the standalone VAD (Ultra/Redux head and
Silero) and the C API calls the backend now uses.

Assisted-by: Claude:claude-sonnet-5-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 01:02:09 +02:00
mudler-agentandEttore Di Giacinto 00faf4ec26 feat(parakeet-cpp): Moondream Ultra and Redux gallery entries, VAD option, pin bump (#12455)
* chore(parakeet-cpp): bump parakeet.cpp to 11c1a0f

Picks up support for the Moondream Ultra and Redux models and the
VAD-segmented transcription entry point (C-API ABI is still 10).

Assisted-by: Claude:claude-sonnet-5-5

* feat(parakeet-cpp): add vad option for long-audio transcription

Models with a VAD head (Moondream Ultra and Redux) can cut long audio at
pauses. Setting vad:true in the model options routes offline
transcription through parakeet_capi_transcribe_path_json_vad. The symbol
is probed at startup like the other optional entry points, and vad:true
fails the load with a clear message when the library lacks it. A model
without a VAD head fails the request with the library's own message.
The option is off by default and does not affect streaming.

Also say in the load error that a packed ternary Redux model is CPU only,
because the library reports its refusal on a GPU backend through its log,
not through the C API.

Assisted-by: Claude:claude-sonnet-5-5

* feat(gallery): add Moondream Ultra and Redux for parakeet-cpp

Add five entries from the public parakeet-cpp GGUF repository: Ultra in
F16 and Q8_0, and Redux as packed ternary (CPU only, offline only) and
as dequantized F16 and Q8_0 (any backend). The entries enable vad:true so
long audio is cut at pauses. Checksums come from the repository's LFS
metadata. The weights are CC-BY-4.0.

Assisted-by: Claude:claude-sonnet-5-5

* docs(parakeet-cpp): document Moondream Ultra, Redux and the vad option

Assisted-by: Claude:claude-sonnet-5-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-03 22:08:12 +02:00
localai-org-maint-botandmudler 9e373dba33 chore: ⬆️ Update PrismML-Eng/llama.cpp to 2459f68b5c0eb26261fd5a81682004b93cd645ba (#12341)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-03 08:29:22 +02:00
localai-org-maint-botandmudler 37eb16eb1b chore: ⬆️ Update CrispStrobe/CrispASR to 966561aa596cfc653aa0e9885d44117fad9cca35 (#12437)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-03 08:29:08 +02:00
localai-org-maint-botandmudler d5f2f607b9 chore: ⬆️ Update ggml-org/whisper.cpp to 60c0be6ac8fa71b1a2ae2dd938a31a34a508e774 (#12439)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-03 08:25:40 +02:00
372b1f8983 feat(silero-vad): allow threshold/silence/pad via model options (#12430)
* feat(silero-vad): allow threshold/silence/pad via model options

Signed-off-by: anton ziderer <Antonziderer@mail.ru>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(silero-vad): ignore NaN thresholds

NaN passes the detector validation but prevents speech comparisons from succeeding.
Ignore it like malformed input and document the option validation.
Convert the option tests to Ginkgo and cover invalid overrides.

Assisted-by: Codex:gpt-6
Signed-off-by: anton ziderer <Antonziderer@mail.ru>

---------

Signed-off-by: anton ziderer <Antonziderer@mail.ru>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-02 23:59:49 +02:00
localai-org-maint-botandmudler 2c2da7ee0c chore: ⬆️ Update ikawrakow/ik_llama.cpp to 5f89bfc81268b4d56d2af63ccbed59de17c64c09 (#12440)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-02 23:57:20 +02:00
mudler-agentandEttore Di Giacinto 38a2aa45fe feat(gallery): add Nimble 9B and CLM decision models, bump vllm.cpp to a19294a9 (#12397)
* feat(gallery): add nimble-9b-vllm-cpp decision model

Add Bespoke Nimble 9B, converted for vllm.cpp and pinned to the weights
commit 52eead25 of mudler/Bespoke-Nimble-9B-vllm-cpp (HEAD only adds the
model card). It is a redistribution of bespokelabs/Bespoke-Nimble-9B
with the LoRA merged into Qwen3.5-9B; config.json names NimbleModel, so
no hf_overrides are needed.

The artifact sits under overrides, where the installer reads it. The
entry sets an 8192-token context, Nimble's own prompt limit, and a KV
pool of 1024 blocks of 32 tokens for 4 sequences (about 1 GiB at 32 KiB
per token for the 8 full-attention layers).

Installed with local-ai models install and served on CPU through the
vllm-cpp backend: the model card's billing request gives billing
(0.986), refund 0.998 and urgency 0.33. Peak resident memory was
18.4 GB, so the description asks for about 20 GB of free RAM.

List the entry in the decisions gallery table. CLM stays out of the
gallery: the pinned engine cannot load the published head layout.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

* feat(gallery): add clm-v0.1-8b-vllm-cpp and bump vllm.cpp for CLM

The CLM checkpoint on mudler/CLM-v0.1-8B-vllm-cpp stores the heads with
the reference's own tensor names (state_head.inp, hidden.N, norms.N,
out). The pinned vllm.cpp 96788348 still expects the old .0/.2/.4/.6
layout and refuses the load with "head.safetensors incomplete for
state_head". vllm.cpp a19294a9 matches the reference layout and adds the
converter that produced the upload, so move the pin there. The ABI stays
at v30.

Add the CLM entry, pinned to the weights commit 0d1903b1 (HEAD only adds
the model card), with a 4096-token context and a KV pool for 4 sequences
(about 2.25 GiB at 144 KiB per token for Qwen3-8B).

Installed with local-ai models install and served on CPU against a
libvllm built at a19294a9: the model card example (john works at google,
entity type) gives person 0.950, the same as the card. Peak resident
memory was 17.9 GB.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-02 09:41:31 +02:00
localai-org-maint-botandmudler a4cf94ecc5 chore: ⬆️ Update ggml-org/llama.cpp to a868c3e3c56657f7e8a6231190dbbe90e7dd86c0 (#12419)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-02 09:39:01 +02:00
localai-org-maint-botandmudler ba091ed3fa chore: ⬆️ Update ikawrakow/ik_llama.cpp to d9e286846d6f8232db48ec5c111a4ea3aea675ef (#12418)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-02 09:38:48 +02:00
Stefan Walcz c3bea567fe fix(llama-cpp): do not stream the error text as content on pre-stream failures (#12425)
When the first result of a streamed request is an error (for example a
prompt that exceeds the context), PredictStream wrote the error message
as a Reply and only then returned the error status. LocalAI treated that
Reply as the first token: it sent the assistant role chunk and the error
text as `content` on an HTTP 200 stream. Because a chunk had already been
written, the pre-stream HTTP error path from #12204 never triggered, so
streaming clients still got a 200 with the error as model output, while
the same request without streaming correctly returns a 400.

Return the error only as the gRPC status. The e2e backend suite gets a
`context_overflow` capability (enabled for llama-cpp) that streams an
over-long prompt and asserts an error status with no content.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-02 09:23:57 +02:00
Stefan Walcz 1056c62f4c fix(cloud-proxy): surface Anthropic refusals instead of empty replies (#12424)
When an Anthropic model declines to answer, the Messages API returns
HTTP 200 with `stop_reason: "refusal"` and empty `content`. The
translate mode mapped that to a normal reply with no content, so the
OpenAI-compatible response looked like a successful completion
(`finish_reason: "stop"`, empty message). Routers, agents and UIs could
not tell "the model declined" from "the model had nothing to say", and
no fallback was triggered.

Return an explicit error for `stop_reason: "refusal"` in both the
non-streaming path and the streaming path (`message_delta`). Regular
replies, including empty `end_turn` replies, are unchanged.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-02 09:22:49 +02:00
Stefan Walcz 985df46d24 fix(whisperx): keep the transcript when diarization fails, report real errors (#12427)
AudioTranscription caught every exception and returned an empty
TranscriptResult. A failed diarization step therefore discarded a
transcript that was already finished: with an HF token that has not
accepted the terms of the gated pyannote pipeline, the download fails
with 403 and every transcription came back as an empty text with
HTTP 200.

Diarization now degrades: if it fails, the transcript is returned
without speaker labels and the reason is logged. Any other failure
aborts the call with INTERNAL instead of pretending success with an
empty text.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-02 09:18:12 +02:00
Stefan Walcz 6aa7b9b871 fix(llama-cpp): let parallel:1 in the model options win over LLAMACPP_PARALLEL (#12426)
The environment fallback was applied whenever n_parallel was still 1
after option parsing. An explicit `parallel: 1` in the model YAML is
indistinguishable from the default that way, so it was replaced by
LLAMACPP_PARALLEL. The docs say options in the YAML take precedence
over environment variables; a single model could not be forced to one
slot while the global variable was set.

Track whether the options set the slot count and resolve it in a small
helper (parallel_params.h): option first, then LLAMACPP_PARALLEL, then
1. The helper gets a standalone unit test picked up by
`make test-backend-cpp`.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-10-02 09:17:44 +02:00
localai-org-maint-botandmudler dd19ee8912 chore: ⬆️ Update CrispStrobe/CrispASR to fdc3a0007d68f8d3905f20e9cd4d18b5f193e096 (#12413)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-02 08:15:31 +02:00
mudler-agentandEttore Di Giacinto 9eb5a9e61d feat(audio): remember speakers from diarization (#12414)
* feat(schema): validate portable speaker profiles

Add the versioned profile schema for explicit speaker enrollment.
Validate compatibility against separately supplied loaded-encoder metadata.
Reject unusable speakers, invalid vectors, and inconsistent clean spans.

This slice does not change HTTP routes, backend integration, or the UI.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet): export profiles with transcripts

Export opt-in speaker profiles and trusted encoder metadata.
Replay registrations by ID so duplicate display names keep independent
vectors.

Use one profile-capable diarization for slots, names, and clean spans.
Assign timestamped ASR words to those slots without a second diarization.
Preserve legacy opt-out and no-ASR behavior, and propagate failures.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(audio): enroll portable speaker profiles

Gate profile exports with voice-recognition permission and validate
registration against metadata from the loaded encoder. Preserve audio
enrollment and independent registrations with duplicate display names.

Exclude diarization and registration exchanges before API trace capture
so persisted traces cannot retain profile vectors or JSON audio.

Defer candidate dimensions to trusted loaded metadata. Sort candidates
by registration ID so incompatible profiles cannot suppress legacy voices
through registry iteration order. Keep portable identity checks closed
when trusted metadata is unavailable.

Test persisted traces, explicit slot zero, and selection through offline
and live transport. Document privacy and the ephemeral registry lifecycle.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): remember speakers from diarization

Add a Studio page for diarization and opt-in speaker profiles. Preview
clean intervals from the original recording before explicit registration.

Join profiles by raw speaker labels, preserve duplicate names, and relabel
turns only after a successful save. Discard stale results when the model
or recording changes. Share registration metadata with voice management
without storing vectors or recordings from this flow.

Document permissions and the global, ephemeral registry. Cover enrollment,
permissions, previews, and asynchronous races with mocked Playwright tests.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: clarify HTTP speaker enrollment support

Replace the stale enrollment limitation with the current HTTP workflow.
Distinguish native transport from explicit registration and link its docs.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(parakeet): pin merged speaker profile support

Use the merged commit from mudler/parakeet.cpp#80.
Its tree matches the previously accepted native pin.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: add diarization enrollment setup example

Connect the existing gallery modes to the speaker enrollment workflow.
Show installation, private profile export, explicit raw-slot registration,
and later recognition without another export.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(blog): explain diarization speaker profiles

Put the diarization walkthrough on the LocalAI website in the feature PR.
Cover the three gallery modes, explicit enrollment, and privacy limits.
Link setup instructions and keep availability conditional on feature support.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(blog): focus diarization on everyday use

Explain what users can do with recordings before the setup steps.
Replace the technical walkthrough with a short Studio guide and link
readers to the existing reference for model names and developer use.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(blog): lead with speaker capabilities

Present speaker recognition through everyday uses and a short UI flow.
Keep technical reference details in the existing documentation.

Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(diarization): satisfy Go lint checks

Avoid copying protobuf message state when extending backend status, check the multipart reader close result, and document the focused testing.T lint exemptions.

Assisted-by: nib:gpt-5.6-sol

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-02 08:14:00 +02:00
mudler-agent 53cd6c5716 fix(funasr): select Python 3.12 tokenizers
FunASR left Transformers and Hugging Face Hub unconstrained, which let uv backtrack to tokenizers 0.10.3 without Python 3.12 wheels. Keep Transformers on the supported 4.x range, including the Intel upgrade profile.\n\nAssisted-by: nib:gpt-5.6-sol\nSigned-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-01 20:48:44 +00:00
b9c975e71c chore: ⬆️ Update ggml-org/llama.cpp to a4d880fd5c7f88713ded6db9f0111893bd78afa6 (#12345)
* ⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(llama-cpp): migrate score batches to the new API

The pinned engine removes common_batch_add and the raw batch view.
Use common_batch entries and llama_process for score suffix decoding.
Read shared-prefix scores from the current common_batch view.

Validation: reproduce both compiler errors on the original patch.
The patched server context and complete grpc-server translation unit
pass g++ -std=c++17 -fsyntax-only with generated protobuf headers.

Assisted-by: Codex:gpt-6

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-01 08:40:58 +02:00
localai-org-maint-botandEttore Di Giacinto 2ae6cae70d feat(parakeet-cpp): name speakers from the shared voice registry (#12382)
* feat(voice): list registered voices and record which encoder made them

The voice registry could register, identify and forget but not list, and
it did not remember which speaker encoder produced an embedding. Add
Metadata.Model and Registry.List, answered from the index the store
registry already keeps for Forget. Needed so a backend can be given the
registered voices that match its own speaker encoder.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(voice): store the encoder model with a registered voice

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(voice): pick the registered voices that match a speaker model

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(proto): carry known voices and speaker names on diarize and live messages

Assisted-by: Claude:claude-haiku-4-5 [Claude Code]

* feat(diarization): name speakers from the voice registry

When a diarization model has a speaker_model option, the endpoint sends
the registered voices made by that encoder to the backend. The backend's
name and name_score come back as extra fields next to the normalized
SPEAKER_NN speaker, and the speakers summary carries the first name seen
for each speaker. RTTM output and results without names are unchanged.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(live): pass registered voices to a live session and surface speaker names

Live sessions now send the registered voices that match the model's
speaker_model to the backend, and each speaker segment carries the name
the backend matched. The realtime segment event gains an optional
speaker_name field.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): load a speaker model and build per-request voice registries

Adds the speaker bindings (ABI v9 and v10, probed separately), the
speaker_model, speaker_threshold and speaker_margin options, and a
per-request registry builder over the known voices.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): name the speakers in Diarize from the known voices

Diarize builds a per-request speaker registry from the known voices when a
speaker model is loaded, calls the named C functions, and puts each slot's
registered name and score on the segments. The registry is freed on every
path. A library without ABI 10 reports Unimplemented instead of dropping
the names.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(parakeet-cpp): name speakers in the live scene stream

The live scene stream now begins with a known-voice registry when a
speaker model is loaded and the live config carries voices, and each
closed speaker segment takes its slot's current name from the feed's
names map. A segment that closes before its slot is identified has an
empty name. The registry is freed after the stream, on every path.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* feat(gallery): speaker naming entries and docs for parakeet-cpp

Add three gallery entries that load the WeSpeaker ResNet34 speaker model
next to the diarization or realtime scene models, and document speaker
names in the voice recognition, diarization, audio to text and realtime
pages.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(parakeet-cpp): skip an unusable registered voice instead of failing the request

A registered voice with the wrong embedding size, or one the C side
refused, failed the whole diarization request, so one legacy voice broke
the model for every user. Skip such voices with a warning that does not
carry the voice name, and take the plain path when none is left.

Also map an exact 0 speaker threshold or margin to a tiny positive value,
since the C side reads 0 as "use the default", and fix a stale comment
about which contexts Free() walks.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* fix(diarization): warn once per model about voices from another encoder; document the privacy limit

The different-encoder warning fired on every request. Log it once per
feature and speaker model, then at debug level. Document that the global
voice registry lets any caller of a speaker_model model learn matching
names, and that skipped wrong-sized voices are logged.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

* chore(parakeet-cpp): bump parakeet.cpp to 8c8cec0 (C-API v10) and check speaker naming against the real library

The pin moves from 623a968 to 8c8cec0, which brings in everything merged
in parakeet.cpp since: the voice identification change (C-API v9, #78) and
raw-embedding enroll plus diarize-only speaker naming (C-API v10, #79).

New real-library specs (gated on PARAKEET_BACKEND_TEST_SPEAKER_MODEL,
_DIAR_MODEL, _WAV and, for the live path, _STREAM_MODEL) name the two
speakers of two_speakers.wav from a committed pair of WeSpeaker embeddings,
with the voices passed in reversed order. They also check that the float32
threshold reaches C through purego. The shared test loader now registers
the v9/v10 and scene symbols as main.go does.

The rebase onto origin/master had no conflicts.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-01 08:25:52 +02:00
localai-org-maint-botandmudler 8e8b23414d chore: ⬆️ Update ikawrakow/ik_llama.cpp to 32cddbfcefed93896a39c64e7c38c119de8682e6 (#12385)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-01 08:23:28 +02:00
localai-org-maint-botandmudler b8303f381e chore: ⬆️ Update localai-org/voice-detect.cpp to b74a896f47c6d04fcca0a962ff317528fd0b0019 (#12384)
⬆️ Update localai-org/voice-detect.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-01 08:23:15 +02:00
localai-org-maint-botandmudler e89d24916b chore: ⬆️ Update CrispStrobe/CrispASR to ba8c1ea667f30b1c0e32ef8574cee68d9f30bcf3 (#12386)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-01 08:23:04 +02:00
localai-org-maint-botandmudler ef2e9167f2 chore: ⬆️ Update mudler/parakeet.cpp to 8c8cec0c4564610a0a4b30a8a6f2ead15d1a76fb (#12387)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-01 08:22:52 +02:00
localai-org-maint-botandmudler 74e24098e4 chore: ⬆️ Update NVIDIA/NeMo-Speech.cpp to 4c101bc7113f49101a3e11d2c994c519f41939f6 (#12388)
⬆️ Update NVIDIA/NeMo-Speech.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-01 08:22:18 +02:00
localai-org-maint-botandmudler 57658f11f3 chore: ⬆️ Update 0xShug0/audio.cpp to 9a02e61326aaaf9d462b584ca5e0daba22c0abfc (#12389)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-10-01 08:22:05 +02:00