mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-09 22:54:42 -04:00
docs(voice): document speaker naming through a bundle component
Cover speaker_component: and speaker_tag: in the audio-to-text option tables and bundle section, add a section to the voice recognition page that explains which registered voices a bundle component receives, and update the speaker_name condition in the realtime and diarization pages. Assisted-by: Claude Code:claude-sonnet-5-5
This commit is contained in:
4 files changed
+49
-5
No files matched your search
@@ -85,7 +85,7 @@ Adds per-speaker totals and (when the backend supports it and `include_text=true
|
||||
|
||||
### Speaker names
|
||||
|
||||
With a parakeet-cpp model that has a `speaker_model:` and voices registered through `/v1/voice/register`, segments whose speaker matches a registered voice gain `name` and `name_score` (the cosine similarity of the match), and the matching `speakers` entry gains `name`. Both fields are omitted for a speaker that was not identified, so an unnamed response looks exactly as before. `speaker` stays `SPEAKER_NN`, and RTTM output still uses `SPEAKER_NN`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription) for the setup and the limits.
|
||||
With a parakeet-cpp model that has a `speaker_model:` (or a `speaker_component:`, for a bundle file such as `parakeet-cpp-bundle-standard`) and voices registered through `/v1/voice/register`, segments whose speaker matches a registered voice gain `name` and `name_score` (the cosine similarity of the match), and the matching `speakers` entry gains `name`. Both fields are omitted for a speaker that was not identified, so an unnamed response looks exactly as before. `speaker` stays `SPEAKER_NN`, and RTTM output still uses `SPEAKER_NN`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription) for the setup and the limits.
|
||||
|
||||
```json
|
||||
{
|
||||
|
||||
@@ -202,7 +202,8 @@ The same backend also serves the `/v1/audio/diarization` and `/v1/audio/classifi
|
||||
| `diarization_model:<path>` | an ASR model | a `speaker` on transcript segments (and words), and speaker segments during realtime live transcription |
|
||||
| `sound_model:<path>` | an ASR model | sound events during realtime live transcription |
|
||||
| `diarization_latency:<model\|low\|very_low\|ultra_low>` | a model with a diarization companion | latency mode for the live speaker stream; default `low` |
|
||||
| `speaker_model:<path>` | a model with a diarization model | names registered speakers (see [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription)) |
|
||||
| `speaker_model:<path>` | a model with a diarization model | names registered speakers; a bundle can use `speaker_component:<name>` instead (see [Bundle GGUF files](#bundle-gguf-files-several-models-in-one-file)) (see [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription)) |
|
||||
| `speaker_tag:<tag>` | a model with `speaker_component` | extra encoder tag for registered voices that carry only a file-name tag (see [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-from-a-bundle)) |
|
||||
| `speaker_threshold:<float>` | a model with `speaker_model` | distance (1 minus cosine similarity) under which a speaker is named, in (0, 2); default `0.5` |
|
||||
| `speaker_margin:<float>` | a model with `speaker_model` | how much the best match must beat the runner-up, in [0, 1); default `0.05` |
|
||||
| `speaker_strict:<bool>` | a model with `speaker_model` | do not use registered voices that have no encoder fingerprint (see [Voice Recognition]({{% relref "voice-recognition" %}}#encoder-fingerprint)); default `false` |
|
||||
@@ -350,6 +351,8 @@ The component that each role uses, and the option that picks another one:
|
||||
| Sound events | the `ced` component, only when asked for | `sound_component:<name>` |
|
||||
| Speaker naming | the `voice` component, only when asked for | `speaker_component:<name>` (needs a diarization component) |
|
||||
|
||||
`speaker_component:` also names registered speakers: LocalAI sends the voices from `/v1/voice/register` to the bundle's speaker component, as it does for `speaker_model:`. This works in `/v1/audio/diarization` and in realtime live transcription, and it needs no `speaker_model:`. `speaker_threshold`, `speaker_margin` and `speaker_strict` apply too. Only voices that carry an encoder fingerprint (voices enrolled from `speaker_profiles`) and voices with no tag at all are used by default. A voice registered with a tag only matches through `speaker_tag:`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-from-a-bundle).
|
||||
|
||||
A `*_component` option without the matching companion option takes the component from the model file itself. The companion options (`diarization_model:`, `sound_model:`, `speaker_model:`, `vad_model:`, `asr_model:`) can also name a bundle file, even the same file as the model: the only component of the wanted kind is used, and the `*_component` option picks one when there are several. This YAML loads the same file for four roles:
|
||||
|
||||
```yaml
|
||||
|
||||
@@ -173,7 +173,7 @@ Each closed speaker segment emits a `conversation.item.input_audio_transcription
|
||||
}
|
||||
```
|
||||
|
||||
`speaker_name` is the name of a voice registered through `/v1/voice/register`, and is present only when the model has a `speaker_model:` and the speaker was identified. A segment that closes before its speaker is identified has none, and later segments of the same speaker do. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription). The segments of the offline path below carry no `speaker_name`.
|
||||
`speaker_name` is the name of a voice registered through `/v1/voice/register`, and is present only when the model has a `speaker_model:` (or, for a parakeet-cpp bundle file, a `speaker_component:`) and the speaker was identified. A segment that closes before its speaker is identified has none, and later segments of the same speaker do. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription). The segments of the offline path below carry no `speaker_name`.
|
||||
|
||||
Each sound event emits a `conversation.item.sound_detection` event with one tag and the detection window's `start`/`end`:
|
||||
|
||||
|
||||
@@ -229,6 +229,42 @@ this, speakers only carry labels such as `SPEAKER_00`.
|
||||
[Speaker Diarization]({{% relref "audio-diarization" %}}) for the
|
||||
response.
|
||||
|
||||
### Naming speakers from a bundle
|
||||
|
||||
A parakeet-cpp bundle file holds the speaker encoder as a component, so its
|
||||
config has `speaker_component:voice` and no `speaker_model:` (the
|
||||
`parakeet-cpp-bundle-small` and `parakeet-cpp-bundle-standard` gallery
|
||||
entries are set up like this). LocalAI treats `speaker_component:` as the
|
||||
speaker encoder of the model when `speaker_model:` is not set. A
|
||||
`speaker_model:` entry always wins, including one that points at the bundle
|
||||
file itself.
|
||||
|
||||
A bundle has no encoder file name, so LocalAI cannot compare file-name tags,
|
||||
and it never guesses a tag from the bundle file name. For a bundle component
|
||||
LocalAI sends the backend:
|
||||
|
||||
- voices with an [encoder fingerprint](#encoder-fingerprint), whatever their
|
||||
tag: the backend checks the fingerprint against the loaded component.
|
||||
Voices enrolled from `speaker_profiles` are in this group. The WeSpeaker
|
||||
component of the published bundles has the same weights as
|
||||
`voice-detect-wespeaker-resnet34`, so a voice enrolled with one works with
|
||||
the other;
|
||||
- voices with no tag (registered before tags existed);
|
||||
- voices whose tag equals the `speaker_tag:` option, if set.
|
||||
|
||||
Other voices are ignored. To use a voice that carries only a file-name tag,
|
||||
set the same tag on the model config. For voices registered through the
|
||||
`voice-detect-wespeaker-resnet34` gallery model:
|
||||
|
||||
```yaml
|
||||
options:
|
||||
- speaker_component:voice
|
||||
- speaker_tag:voice-detect-wespeaker-resnet34.gguf
|
||||
```
|
||||
|
||||
Use `speaker_tag:` only for voices made by the same weights as the bundle
|
||||
component. The check by size alone cannot tell two encoders apart.
|
||||
|
||||
### Which voices are used
|
||||
|
||||
LocalAI sends the backend only the registered voices made by the same
|
||||
@@ -295,6 +331,8 @@ options).
|
||||
| Option | Default | Meaning |
|
||||
|---|---|---|
|
||||
| `speaker_model:<path>` | none | speaker encoder GGUF; needs a diarization model (the primary one, or `diarization_model:`) |
|
||||
| `speaker_component:<name>` | none | speaker encoder component of a bundle file; used as the speaker encoder for naming when `speaker_model:` is not set (see [Naming speakers from a bundle](#naming-speakers-from-a-bundle)) |
|
||||
| `speaker_tag:<tag>` | none | encoder tag that also counts as the loaded encoder, for voices that carry only a file-name tag; mainly for a bundle component |
|
||||
| `speaker_threshold:<float>` | `0.5` | largest distance (1 minus cosine similarity, the unit `/v1/voice/identify` reports) at which a speaker is named; must be in (0, 2) |
|
||||
| `speaker_margin:<float>` | `0.05` | the best match must beat the runner-up by this much, otherwise the speaker stays unnamed; must be in [0, 1) |
|
||||
| `speaker_strict:<bool>` | `false` | ignore registered voices that carry no [encoder fingerprint](#encoder-fingerprint); needs a libparakeet that exports `parakeet_capi_speaker_registry_set_strict` |
|
||||
@@ -307,7 +345,7 @@ speakers and makes fewer mistakes.
|
||||
|
||||
- The voice registry is in memory and global. Registered names disappear when
|
||||
LocalAI restarts, and every user of the instance shares them.
|
||||
- Anyone who is allowed to call a model with `speaker_model:` can learn which
|
||||
- Anyone who is allowed to call a model with `speaker_model:` or `speaker_component:` can learn which
|
||||
registered names match their audio, and their audio is matched against voices
|
||||
registered by any user, because the voice registry is global. Restrict such
|
||||
models with the per-user model allowlist.
|
||||
@@ -320,7 +358,10 @@ speakers and makes fewer mistakes.
|
||||
- Accuracy was measured on one fixture (two read-speech voices). Check the
|
||||
threshold on your own audio.
|
||||
- The backend needs a libparakeet with C-API v10. With an older library a
|
||||
model config that sets `speaker_model:` fails to load.
|
||||
model config that sets `speaker_model:` or `speaker_component:` fails to load.
|
||||
- A bundle does not serve `/v1/voice/register`, `/v1/voice/identify` or
|
||||
`/v1/voice/verify` itself. Register voices with a voice-detect model (or
|
||||
from `speaker_profiles`), then name them through the bundle.
|
||||
|
||||
## API reference
|
||||
|
||||
|
||||
Reference in new issue
Block a user