mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-05 04:24:39 -04:00
feat(parakeet-cpp): name speakers from the shared voice registry (#12382)
* feat(voice): list registered voices and record which encoder made them The voice registry could register, identify and forget but not list, and it did not remember which speaker encoder produced an embedding. Add Metadata.Model and Registry.List, answered from the index the store registry already keeps for Forget. Needed so a backend can be given the registered voices that match its own speaker encoder. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(voice): store the encoder model with a registered voice Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(voice): pick the registered voices that match a speaker model Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(proto): carry known voices and speaker names on diarize and live messages Assisted-by: Claude:claude-haiku-4-5 [Claude Code] * feat(diarization): name speakers from the voice registry When a diarization model has a speaker_model option, the endpoint sends the registered voices made by that encoder to the backend. The backend's name and name_score come back as extra fields next to the normalized SPEAKER_NN speaker, and the speakers summary carries the first name seen for each speaker. RTTM output and results without names are unchanged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(live): pass registered voices to a live session and surface speaker names Live sessions now send the registered voices that match the model's speaker_model to the backend, and each speaker segment carries the name the backend matched. The realtime segment event gains an optional speaker_name field. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): load a speaker model and build per-request voice registries Adds the speaker bindings (ABI v9 and v10, probed separately), the speaker_model, speaker_threshold and speaker_margin options, and a per-request registry builder over the known voices. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): name the speakers in Diarize from the known voices Diarize builds a per-request speaker registry from the known voices when a speaker model is loaded, calls the named C functions, and puts each slot's registered name and score on the segments. The registry is freed on every path. A library without ABI 10 reports Unimplemented instead of dropping the names. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): name speakers in the live scene stream The live scene stream now begins with a known-voice registry when a speaker model is loaded and the live config carries voices, and each closed speaker segment takes its slot's current name from the feed's names map. A segment that closes before its slot is identified has an empty name. The registry is freed after the stream, on every path. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(gallery): speaker naming entries and docs for parakeet-cpp Add three gallery entries that load the WeSpeaker ResNet34 speaker model next to the diarization or realtime scene models, and document speaker names in the voice recognition, diarization, audio to text and realtime pages. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(parakeet-cpp): skip an unusable registered voice instead of failing the request A registered voice with the wrong embedding size, or one the C side refused, failed the whole diarization request, so one legacy voice broke the model for every user. Skip such voices with a warning that does not carry the voice name, and take the plain path when none is left. Also map an exact 0 speaker threshold or margin to a tiny positive value, since the C side reads 0 as "use the default", and fix a stale comment about which contexts Free() walks. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(diarization): warn once per model about voices from another encoder; document the privacy limit The different-encoder warning fired on every request. Log it once per feature and speaker model, then at debug level. Document that the global voice registry lets any caller of a speaker_model model learn matching names, and that skipped wrong-sized voices are logged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * chore(parakeet-cpp): bump parakeet.cpp to 8c8cec0 (C-API v10) and check speaker naming against the real library The pin moves from 623a968 to 8c8cec0, which brings in everything merged in parakeet.cpp since: the voice identification change (C-API v9, #78) and raw-embedding enroll plus diarize-only speaker naming (C-API v10, #79). New real-library specs (gated on PARAKEET_BACKEND_TEST_SPEAKER_MODEL, _DIAR_MODEL, _WAV and, for the live path, _STREAM_MODEL) name the two speakers of two_speakers.wav from a committed pair of WeSpeaker embeddings, with the voices passed in reversed order. They also check that the float32 threshold reaches C through purego. The shared test loader now registers the v9/v10 and scene symbols as main.go does. The rebase onto origin/master had no conflicts. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
1 parent
51be7b48e5
commit
2ae6cae70d
48 files changed
+2244
-61
No files matched your search
@@ -79,6 +79,26 @@ Adds per-speaker totals and (when the backend supports it and `include_text=true
|
||||
}
|
||||
```
|
||||
|
||||
### Speaker names
|
||||
|
||||
With a parakeet-cpp model that has a `speaker_model:` and voices registered through `/v1/voice/register`, segments whose speaker matches a registered voice gain `name` and `name_score` (the cosine similarity of the match), and the matching `speakers` entry gains `name`. Both fields are omitted for a speaker that was not identified, so an unnamed response looks exactly as before. `speaker` stays `SPEAKER_NN`, and RTTM output still uses `SPEAKER_NN`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription) for the setup and the limits.
|
||||
|
||||
```json
|
||||
{
|
||||
"task": "diarize",
|
||||
"duration": 12.34,
|
||||
"num_speakers": 2,
|
||||
"segments": [
|
||||
{"id": 0, "speaker": "SPEAKER_00", "label": "0", "start": 0.00, "end": 2.34, "text": "Hello, world.", "name": "Alice", "name_score": 0.82},
|
||||
{"id": 1, "speaker": "SPEAKER_01", "label": "1", "start": 2.34, "end": 4.10, "text": "How are you?"}
|
||||
],
|
||||
"speakers": [
|
||||
{"id": "SPEAKER_00", "label": "0", "name": "Alice", "total_speech_duration": 5.6, "segment_count": 3},
|
||||
{"id": "SPEAKER_01", "label": "1", "total_speech_duration": 1.76, "segment_count": 1}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### Response - `rttm`
|
||||
|
||||
NIST RTTM, the standard interchange format used by `pyannote.metrics` / `dscore`:
|
||||
|
||||
@@ -200,9 +200,14 @@ The same backend also serves the `/v1/audio/diarization` and `/v1/audio/classifi
|
||||
| `diarization_model:<path>` | an ASR model | a `speaker` on transcript segments (and words), and speaker segments during realtime live transcription |
|
||||
| `sound_model:<path>` | an ASR model | sound events during realtime live transcription |
|
||||
| `diarization_latency:<model\|low\|very_low\|ultra_low>` | a model with a diarization companion | latency mode for the live speaker stream; default `low` |
|
||||
| `speaker_model:<path>` | a model with a diarization model | names registered speakers (see [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription)) |
|
||||
| `speaker_threshold:<float>` | a model with `speaker_model` | distance (1 minus cosine similarity) under which a speaker is named, in (0, 2); default `0.5` |
|
||||
| `speaker_margin:<float>` | a model with `speaker_model` | how much the best match must beat the runner-up, in [0, 1); default `0.05` |
|
||||
|
||||
With a `diarization_model` companion, `/v1/audio/transcriptions` labels each segment with its `speaker` (`"0"`, `"1"`, ... in order of first appearance) and splits segments where the speaker changes; with `timestamp_granularities[]=word` each word carries its speaker too. With `stream=true` the closing `transcript.text.done` event lists the segments with their speakers. Pass `-F diarize=false` to skip diarization for one request. The diarization GGUF can also be imported directly: `local-ai models import https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/nemotron-3-diarization-f16.gguf`.
|
||||
|
||||
`speaker_model:` needs libparakeet with C-API v10. A wrong setup fails at load time with one of these errors: `parakeet-cpp: speaker_model needs libparakeet.so ABI 10 (parakeet_capi_speaker_registry_add_embedding); the loaded library is older`, `parakeet-cpp: speaker_model needs a diarization model (the primary or diarization_model:)`, `parakeet-cpp: a speaker model cannot be the primary model; use it as speaker_model: next to a diarization model`, `parakeet-cpp: speaker_model "<path>" is a <kind> model, expected a speaker model` (the file is not a speaker encoder GGUF), or `parakeet-cpp: speaker_threshold "<value>" must be a distance in (0, 2) (1 minus cosine similarity)` / `parakeet-cpp: speaker_margin "<value>" must be a number in [0, 1)` for a bad number.
|
||||
|
||||
The loader rejects a companion whose role duplicates the primary's own (for example `asr_model:` on an already-ASR primary, or `sound_model:` on a CED primary), and rejects a companion GGUF that does not match the role its option names (for example `sound_model:` pointing at an ASR GGUF fails to load, naming the kind it expected). See [Speaker Diarization]({{% relref "audio-diarization" %}}) for the `Diarize` RPC and [Sound Classification]({{% relref "audio-classification" %}}) for `SoundDetection`, and [Realtime API]({{% relref "openai-realtime" %}}) for the live speaker/sound events emitted during a realtime session.
|
||||
|
||||
### Segment timestamps
|
||||
|
||||
@@ -166,12 +166,15 @@ Each closed speaker segment emits a `conversation.item.input_audio_transcription
|
||||
"item_id": "item_abc",
|
||||
"content_index": 0,
|
||||
"speaker": "0",
|
||||
"speaker_name": "Alice",
|
||||
"start": 1.92,
|
||||
"end": 4.10,
|
||||
"text": ""
|
||||
}
|
||||
```
|
||||
|
||||
`speaker_name` is the name of a voice registered through `/v1/voice/register`, and is present only when the model has a `speaker_model:` and the speaker was identified. A segment that closes before its speaker is identified has none, and later segments of the same speaker do. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription). The segments of the offline path below carry no `speaker_name`.
|
||||
|
||||
Each sound event emits a `conversation.item.sound_detection` event with one tag and the detection window's `start`/`end`:
|
||||
|
||||
```json
|
||||
|
||||
@@ -185,6 +185,85 @@ recognition - the voice-recognition HTTP API is designed to swap the
|
||||
backing store without changing the wire format.
|
||||
{{% /notice %}}
|
||||
|
||||
## Naming speakers in diarization and live transcription
|
||||
|
||||
The parakeet-cpp backend can put the names of registered voices on
|
||||
diarization results and on live transcription speaker segments. Without
|
||||
this, speakers only carry labels such as `SPEAKER_00`.
|
||||
|
||||
1. Register each voice with the WeSpeaker encoder. Install the model with
|
||||
`local-ai models install voice-detect-wespeaker-resnet34`, then call
|
||||
`/v1/voice/register` with `"model": "voice-detect-wespeaker-resnet34"`
|
||||
(see the [1:N workflow](#1n-identification-workflow-register--identify--forget)).
|
||||
2. Install one of the gallery models that loads the same encoder:
|
||||
`parakeet-cpp-nemotron-3-diarization-speakers` (diarization),
|
||||
`parakeet-cpp-nemotron-3-diarization-asr-speakers` (diarization with
|
||||
`include_text`) or `parakeet-cpp-realtime-scene-speakers` (live
|
||||
transcription). Each one adds
|
||||
`speaker_model:voice-detect-wespeaker-resnet34.gguf` to a
|
||||
parakeet-cpp model config.
|
||||
3. Call `/v1/audio/diarization` with that model. Matched segments gain a
|
||||
`name` and a `name_score`, and the matching entry in `speakers` gains a
|
||||
`name`. `speaker` stays `SPEAKER_NN`, and RTTM output is unchanged. See
|
||||
[Speaker Diarization]({{% relref "audio-diarization" %}}) for the
|
||||
response.
|
||||
|
||||
### Which voices are used
|
||||
|
||||
LocalAI sends the backend only the registered voices made by the same
|
||||
encoder as the model's `speaker_model:` file. Each registered voice is
|
||||
tagged with the name of the voice-detect model that made it, which by
|
||||
default is the GGUF file name (`voice-detect-wespeaker-resnet34.gguf` for the
|
||||
gallery entry). The tag must equal the base name of the `speaker_model:`
|
||||
file. Voices made with another encoder are ignored, and LocalAI logs a
|
||||
warning when that leaves no usable voice. Voices registered before the tag
|
||||
existed have no tag: they are used when their embedding size matches the
|
||||
tagged ones (or all of them, when no voice carries a matching tag). The
|
||||
backend skips a voice whose embedding size does not match the speaker model's,
|
||||
with a warning in the LocalAI log. Naming then falls back to the remaining
|
||||
voices, or to no names.
|
||||
|
||||
{{% notice warning %}}
|
||||
Do not set a `model_name:` option on the voice-detect model config. It
|
||||
replaces the default name, the voices are then tagged with it, and they no
|
||||
longer match the `speaker_model:` file. Keep the default name.
|
||||
{{% /notice %}}
|
||||
|
||||
### Options
|
||||
|
||||
These go in the `options:` list of the parakeet-cpp model config (see
|
||||
[Audio to Text]({{% relref "audio-to-text" %}}) for the other parakeet-cpp
|
||||
options).
|
||||
|
||||
| Option | Default | Meaning |
|
||||
|---|---|---|
|
||||
| `speaker_model:<path>` | none | speaker encoder GGUF; needs a diarization model (the primary one, or `diarization_model:`) |
|
||||
| `speaker_threshold:<float>` | `0.5` | largest distance (1 minus cosine similarity, the unit `/v1/voice/identify` reports) at which a speaker is named; must be in (0, 2) |
|
||||
| `speaker_margin:<float>` | `0.05` | the best match must beat the runner-up by this much, otherwise the speaker stays unnamed; must be in [0, 1) |
|
||||
|
||||
parakeet.cpp's measured starting values for `speaker_threshold` are 0.5 for
|
||||
WeSpeaker ResNet34 and CAM++, and 0.3 for ECAPA. A lower value names fewer
|
||||
speakers and makes fewer mistakes.
|
||||
|
||||
### Limits
|
||||
|
||||
- The voice registry is in memory and global. Registered names disappear when
|
||||
LocalAI restarts, and every user of the instance shares them.
|
||||
- Anyone who is allowed to call a model with `speaker_model:` can learn which
|
||||
registered names match their audio, and their audio is matched against voices
|
||||
registered by any user, because the voice registry is global. Restrict such
|
||||
models with the per-user model allowlist.
|
||||
- With `include_text=true` the names use the default threshold and margin:
|
||||
`speaker_threshold` and `speaker_margin` only apply to diarization without
|
||||
text.
|
||||
- In live transcription, a speaker segment that closes before its speaker
|
||||
is identified has no name. Later segments of that speaker do.
|
||||
- Overlapping speech is not resolved.
|
||||
- Accuracy was measured on one fixture (two read-speech voices). Check the
|
||||
threshold on your own audio.
|
||||
- The backend needs a libparakeet with C-API v10. With an older library a
|
||||
model config that sets `speaker_model:` fails to load.
|
||||
|
||||
## API reference
|
||||
|
||||
### `POST /v1/voice/verify` (1:1)
|
||||
|
||||
Reference in new issue
Block a user