docs(voice): document a parakeet-cpp bundle as the embedding model

Describe the bundle as an embedding model for /v1/voice/*, the realtime
voice_recognition stage, and the limits: identify and plain verify only,
a 256-dimension space shared with voice-detect-wespeaker-resnet34, and the
libparakeet.so symbol it needs.

Assisted-by: Claude Code:claude-sonnet-5-5
This commit is contained in:
Ettore Di Giacinto committed 2026-10-05 23:46:14 +00:00
1 parent ce4099ea38
commit e9af44b90b
3 files changed
+62

No files matched your search

+1
View File
@@ -350,6 +350,7 @@ The component that each role uses, and the option that picks another one:
| Diarization | the `diar` component, only when asked for | `diar_component:<name>` |
| Sound events | the `ced` component, only when asked for | `sound_component:<name>` |
| Speaker naming | the `voice` component, only when asked for | `speaker_component:<name>` (needs a diarization component) |
| Voice embedding and verification (`/v1/voice/*`, the realtime `voice_recognition` stage) | the `voice` component, only when asked for | `speaker_component:<name>`; declare `speaker_recognition` in `known_usecases`. Needs a libparakeet.so with `parakeet_capi_speaker_embed_pcm`. See [Voice Recognition]({{% relref "voice-recognition#a-parakeet-cpp-bundle-as-the-embedding-model" %}}) |
`speaker_component:` also names registered speakers: LocalAI sends the voices from `/v1/voice/register` to the bundle's speaker component, as it does for `speaker_model:`. This works in `/v1/audio/diarization` and in realtime live transcription, and it needs no `speaker_model:`. `speaker_threshold`, `speaker_margin` and `speaker_strict` apply too. Only voices that carry an encoder fingerprint (voices enrolled from `speaker_profiles`) and voices with no tag at all are used by default. A voice registered with a tag only matches through `speaker_tag:`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-from-a-bundle).
+27
View File
@@ -547,6 +547,33 @@ pipeline:
audio: /models/voices/bob.wav
```
### One bundle for every stage
A parakeet-cpp bundle can fill the speaker stage as well as the others, so one model serves
`vad`, `transcription`, `sound_detection` and `voice_recognition`. The bundle config must
list `speaker_recognition` in `known_usecases` (the gallery entries do), and the libparakeet.so
must export `parakeet_capi_speaker_embed_pcm`:
```yaml
name: my-realtime
pipeline:
vad: parakeet-cpp-bundle-small
transcription: parakeet-cpp-bundle-small
sound_detection: parakeet-cpp-bundle-small
llm: qwen
tts: kokoro
voice_recognition:
model: parakeet-cpp-bundle-small
mode: identify
threshold: 0.5 # WeSpeaker distance; see Voice Recognition
```
The bundle's encoder has the same weights as `voice-detect-wespeaker-resnet34`, so voices
registered with that model are matched. Embedding an utterance uses the same engine lock as
transcription, so it adds to the turn latency. Voices from a 192-dimension model (ECAPA-TDNN,
for example) cannot be matched by a bundle; see
[Voice Recognition]({{% relref "voice-recognition#a-parakeet-cpp-bundle-as-the-embedding-model" %}}).
### Identifying speakers without gating
To recognize who is speaking and surface it to the client and the LLM without ever rejecting a turn, set `enforce: false` and add an `identity` block. The `identity` block works with or without the gate; when it is set, the speaker is resolved on every turn even if `when: first`.
@@ -206,6 +206,40 @@ embedding size, so registering a voice of a new size no longer fails.
- Naming in diarization and live transcription reads the registry as a whole
and uses the voices that match the loaded encoder.
## A parakeet-cpp bundle as the embedding model
A parakeet-cpp model that has a speaker encoder can serve `/v1/voice/embed`,
`/v1/voice/verify`, `/v1/voice/register` and `/v1/voice/identify`, and the
`voice_recognition` stage of a realtime pipeline. That is every
[bundle]({{% relref "audio-to-text#bundle-gguf-files-several-models-in-one-file" %}}) with a `voice` component
(`parakeet-cpp-bundle-small` and `parakeet-cpp-bundle-standard`), and any
parakeet-cpp config with a `speaker_model:` or `speaker_component:` option, such as
`parakeet-cpp-realtime-scene-speakers`. Declare `speaker_recognition` in the
`known_usecases` of the config; the gallery entries above already do. Then use the model
name as `model`:
```bash
local-ai models install parakeet-cpp-bundle-small
curl http://localhost:8080/v1/voice/embed -H "Content-Type: application/json" \
-d '{"model": "parakeet-cpp-bundle-small", "audio": "https://example.com/clip.wav"}'
```
The `voice` component holds the same weights as `voice-detect-wespeaker-resnet34`
(256 dimensions), so a voice registered with one model also matches embeddings from the
other. Use a distance threshold near 0.5 for this encoder, as for `speaker_model:`
naming below. The response `model` field is the `sha256:` identity of the encoder
weights.
Only embedding, identification and plain verification are available. There is no
anti-spoofing head, so a verify request with `anti_spoofing: true` is refused, and
`/v1/voice/analyze` is not served by parakeet-cpp. The backend needs a libparakeet.so
that exports `parakeet_capi_speaker_embed_pcm`. With an older library these calls fail
with an `Unimplemented` error that names the missing symbol, and a model without a
speaker encoder fails with a `FailedPrecondition` error.
Voices of other embedding sizes stay in their own store, see
[Voices from different encoders](#voices-from-different-encoders).
## Naming speakers in diarization and live transcription
The parakeet-cpp backend can put the names of registered voices on