From 04df10a082487dd48f2c990f97a29e0ca9bcfd5e Mon Sep 17 00:00:00 2001 From: Ettore Di Giacinto Date: Mon, 5 Oct 2026 23:43:23 +0000 Subject: [PATCH] docs(voice): document speaker naming through a bundle component Cover speaker_component: and speaker_tag: in the audio-to-text option tables and bundle section, add a section to the voice recognition page that explains which registered voices a bundle component receives, and update the speaker_name condition in the realtime and diarization pages. Assisted-by: Claude Code:claude-sonnet-5-5 --- docs/content/features/audio-diarization.md | 2 +- docs/content/features/audio-to-text.md | 5 ++- docs/content/features/openai-realtime.md | 2 +- docs/content/features/voice-recognition.md | 45 +++++++++++++++++++++- 4 files changed, 49 insertions(+), 5 deletions(-) diff --git a/docs/content/features/audio-diarization.md b/docs/content/features/audio-diarization.md index 3d6b9c702..649da38e9 100644 --- a/docs/content/features/audio-diarization.md +++ b/docs/content/features/audio-diarization.md @@ -85,7 +85,7 @@ Adds per-speaker totals and (when the backend supports it and `include_text=true ### Speaker names -With a parakeet-cpp model that has a `speaker_model:` and voices registered through `/v1/voice/register`, segments whose speaker matches a registered voice gain `name` and `name_score` (the cosine similarity of the match), and the matching `speakers` entry gains `name`. Both fields are omitted for a speaker that was not identified, so an unnamed response looks exactly as before. `speaker` stays `SPEAKER_NN`, and RTTM output still uses `SPEAKER_NN`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription) for the setup and the limits. +With a parakeet-cpp model that has a `speaker_model:` (or a `speaker_component:`, for a bundle file such as `parakeet-cpp-bundle-standard`) and voices registered through `/v1/voice/register`, segments whose speaker matches a registered voice gain `name` and `name_score` (the cosine similarity of the match), and the matching `speakers` entry gains `name`. Both fields are omitted for a speaker that was not identified, so an unnamed response looks exactly as before. `speaker` stays `SPEAKER_NN`, and RTTM output still uses `SPEAKER_NN`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription) for the setup and the limits. ```json { diff --git a/docs/content/features/audio-to-text.md b/docs/content/features/audio-to-text.md index 98cf54a24..9dc13520e 100644 --- a/docs/content/features/audio-to-text.md +++ b/docs/content/features/audio-to-text.md @@ -202,7 +202,8 @@ The same backend also serves the `/v1/audio/diarization` and `/v1/audio/classifi | `diarization_model:` | an ASR model | a `speaker` on transcript segments (and words), and speaker segments during realtime live transcription | | `sound_model:` | an ASR model | sound events during realtime live transcription | | `diarization_latency:` | a model with a diarization companion | latency mode for the live speaker stream; default `low` | -| `speaker_model:` | a model with a diarization model | names registered speakers (see [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription)) | +| `speaker_model:` | a model with a diarization model | names registered speakers; a bundle can use `speaker_component:` instead (see [Bundle GGUF files](#bundle-gguf-files-several-models-in-one-file)) (see [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription)) | +| `speaker_tag:` | a model with `speaker_component` | extra encoder tag for registered voices that carry only a file-name tag (see [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-from-a-bundle)) | | `speaker_threshold:` | a model with `speaker_model` | distance (1 minus cosine similarity) under which a speaker is named, in (0, 2); default `0.5` | | `speaker_margin:` | a model with `speaker_model` | how much the best match must beat the runner-up, in [0, 1); default `0.05` | | `speaker_strict:` | a model with `speaker_model` | do not use registered voices that have no encoder fingerprint (see [Voice Recognition]({{% relref "voice-recognition" %}}#encoder-fingerprint)); default `false` | @@ -350,6 +351,8 @@ The component that each role uses, and the option that picks another one: | Sound events | the `ced` component, only when asked for | `sound_component:` | | Speaker naming | the `voice` component, only when asked for | `speaker_component:` (needs a diarization component) | +`speaker_component:` also names registered speakers: LocalAI sends the voices from `/v1/voice/register` to the bundle's speaker component, as it does for `speaker_model:`. This works in `/v1/audio/diarization` and in realtime live transcription, and it needs no `speaker_model:`. `speaker_threshold`, `speaker_margin` and `speaker_strict` apply too. Only voices that carry an encoder fingerprint (voices enrolled from `speaker_profiles`) and voices with no tag at all are used by default. A voice registered with a tag only matches through `speaker_tag:`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-from-a-bundle). + A `*_component` option without the matching companion option takes the component from the model file itself. The companion options (`diarization_model:`, `sound_model:`, `speaker_model:`, `vad_model:`, `asr_model:`) can also name a bundle file, even the same file as the model: the only component of the wanted kind is used, and the `*_component` option picks one when there are several. This YAML loads the same file for four roles: ```yaml diff --git a/docs/content/features/openai-realtime.md b/docs/content/features/openai-realtime.md index f66f498f2..017e62c7d 100644 --- a/docs/content/features/openai-realtime.md +++ b/docs/content/features/openai-realtime.md @@ -173,7 +173,7 @@ Each closed speaker segment emits a `conversation.item.input_audio_transcription } ``` -`speaker_name` is the name of a voice registered through `/v1/voice/register`, and is present only when the model has a `speaker_model:` and the speaker was identified. A segment that closes before its speaker is identified has none, and later segments of the same speaker do. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription). The segments of the offline path below carry no `speaker_name`. +`speaker_name` is the name of a voice registered through `/v1/voice/register`, and is present only when the model has a `speaker_model:` (or, for a parakeet-cpp bundle file, a `speaker_component:`) and the speaker was identified. A segment that closes before its speaker is identified has none, and later segments of the same speaker do. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription). The segments of the offline path below carry no `speaker_name`. Each sound event emits a `conversation.item.sound_detection` event with one tag and the detection window's `start`/`end`: diff --git a/docs/content/features/voice-recognition.md b/docs/content/features/voice-recognition.md index bbefd10fb..932ce0a82 100644 --- a/docs/content/features/voice-recognition.md +++ b/docs/content/features/voice-recognition.md @@ -229,6 +229,42 @@ this, speakers only carry labels such as `SPEAKER_00`. [Speaker Diarization]({{% relref "audio-diarization" %}}) for the response. +### Naming speakers from a bundle + +A parakeet-cpp bundle file holds the speaker encoder as a component, so its +config has `speaker_component:voice` and no `speaker_model:` (the +`parakeet-cpp-bundle-small` and `parakeet-cpp-bundle-standard` gallery +entries are set up like this). LocalAI treats `speaker_component:` as the +speaker encoder of the model when `speaker_model:` is not set. A +`speaker_model:` entry always wins, including one that points at the bundle +file itself. + +A bundle has no encoder file name, so LocalAI cannot compare file-name tags, +and it never guesses a tag from the bundle file name. For a bundle component +LocalAI sends the backend: + +- voices with an [encoder fingerprint](#encoder-fingerprint), whatever their + tag: the backend checks the fingerprint against the loaded component. + Voices enrolled from `speaker_profiles` are in this group. The WeSpeaker + component of the published bundles has the same weights as + `voice-detect-wespeaker-resnet34`, so a voice enrolled with one works with + the other; +- voices with no tag (registered before tags existed); +- voices whose tag equals the `speaker_tag:` option, if set. + +Other voices are ignored. To use a voice that carries only a file-name tag, +set the same tag on the model config. For voices registered through the +`voice-detect-wespeaker-resnet34` gallery model: + +```yaml +options: +- speaker_component:voice +- speaker_tag:voice-detect-wespeaker-resnet34.gguf +``` + +Use `speaker_tag:` only for voices made by the same weights as the bundle +component. The check by size alone cannot tell two encoders apart. + ### Which voices are used LocalAI sends the backend only the registered voices made by the same @@ -295,6 +331,8 @@ options). | Option | Default | Meaning | |---|---|---| | `speaker_model:` | none | speaker encoder GGUF; needs a diarization model (the primary one, or `diarization_model:`) | +| `speaker_component:` | none | speaker encoder component of a bundle file; used as the speaker encoder for naming when `speaker_model:` is not set (see [Naming speakers from a bundle](#naming-speakers-from-a-bundle)) | +| `speaker_tag:` | none | encoder tag that also counts as the loaded encoder, for voices that carry only a file-name tag; mainly for a bundle component | | `speaker_threshold:` | `0.5` | largest distance (1 minus cosine similarity, the unit `/v1/voice/identify` reports) at which a speaker is named; must be in (0, 2) | | `speaker_margin:` | `0.05` | the best match must beat the runner-up by this much, otherwise the speaker stays unnamed; must be in [0, 1) | | `speaker_strict:` | `false` | ignore registered voices that carry no [encoder fingerprint](#encoder-fingerprint); needs a libparakeet that exports `parakeet_capi_speaker_registry_set_strict` | @@ -307,7 +345,7 @@ speakers and makes fewer mistakes. - The voice registry is in memory and global. Registered names disappear when LocalAI restarts, and every user of the instance shares them. -- Anyone who is allowed to call a model with `speaker_model:` can learn which +- Anyone who is allowed to call a model with `speaker_model:` or `speaker_component:` can learn which registered names match their audio, and their audio is matched against voices registered by any user, because the voice registry is global. Restrict such models with the per-user model allowlist. @@ -320,7 +358,10 @@ speakers and makes fewer mistakes. - Accuracy was measured on one fixture (two read-speech voices). Check the threshold on your own audio. - The backend needs a libparakeet with C-API v10. With an older library a - model config that sets `speaker_model:` fails to load. + model config that sets `speaker_model:` or `speaker_component:` fails to load. +- A bundle does not serve `/v1/voice/register`, `/v1/voice/identify` or + `/v1/voice/verify` itself. Register voices with a voice-detect model (or + from `speaker_profiles`), then name them through the bundle. ## API reference