mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-10 07:47:29 -04:00
docs(voice): document a parakeet-cpp bundle as the embedding model
Describe the bundle as an embedding model for /v1/voice/*, the realtime voice_recognition stage, and the limits: identify and plain verify only, a 256-dimension space shared with voice-detect-wespeaker-resnet34, and the libparakeet.so symbol it needs. Assisted-by: Claude Code:claude-sonnet-5-5
This commit is contained in:
3 files changed
+62
No files matched your search
@@ -350,6 +350,7 @@ The component that each role uses, and the option that picks another one:
|
||||
| Diarization | the `diar` component, only when asked for | `diar_component:<name>` |
|
||||
| Sound events | the `ced` component, only when asked for | `sound_component:<name>` |
|
||||
| Speaker naming | the `voice` component, only when asked for | `speaker_component:<name>` (needs a diarization component) |
|
||||
| Voice embedding and verification (`/v1/voice/*`, the realtime `voice_recognition` stage) | the `voice` component, only when asked for | `speaker_component:<name>`; declare `speaker_recognition` in `known_usecases`. Needs a libparakeet.so with `parakeet_capi_speaker_embed_pcm`. See [Voice Recognition]({{% relref "voice-recognition#a-parakeet-cpp-bundle-as-the-embedding-model" %}}) |
|
||||
|
||||
`speaker_component:` also names registered speakers: LocalAI sends the voices from `/v1/voice/register` to the bundle's speaker component, as it does for `speaker_model:`. This works in `/v1/audio/diarization` and in realtime live transcription, and it needs no `speaker_model:`. `speaker_threshold`, `speaker_margin` and `speaker_strict` apply too. Only voices that carry an encoder fingerprint (voices enrolled from `speaker_profiles`) and voices with no tag at all are used by default. A voice registered with a tag only matches through `speaker_tag:`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-from-a-bundle).
|
||||
|
||||
|
||||
@@ -547,6 +547,33 @@ pipeline:
|
||||
audio: /models/voices/bob.wav
|
||||
```
|
||||
|
||||
### One bundle for every stage
|
||||
|
||||
A parakeet-cpp bundle can fill the speaker stage as well as the others, so one model serves
|
||||
`vad`, `transcription`, `sound_detection` and `voice_recognition`. The bundle config must
|
||||
list `speaker_recognition` in `known_usecases` (the gallery entries do), and the libparakeet.so
|
||||
must export `parakeet_capi_speaker_embed_pcm`:
|
||||
|
||||
```yaml
|
||||
name: my-realtime
|
||||
pipeline:
|
||||
vad: parakeet-cpp-bundle-small
|
||||
transcription: parakeet-cpp-bundle-small
|
||||
sound_detection: parakeet-cpp-bundle-small
|
||||
llm: qwen
|
||||
tts: kokoro
|
||||
voice_recognition:
|
||||
model: parakeet-cpp-bundle-small
|
||||
mode: identify
|
||||
threshold: 0.5 # WeSpeaker distance; see Voice Recognition
|
||||
```
|
||||
|
||||
The bundle's encoder has the same weights as `voice-detect-wespeaker-resnet34`, so voices
|
||||
registered with that model are matched. Embedding an utterance uses the same engine lock as
|
||||
transcription, so it adds to the turn latency. Voices from a 192-dimension model (ECAPA-TDNN,
|
||||
for example) cannot be matched by a bundle; see
|
||||
[Voice Recognition]({{% relref "voice-recognition#a-parakeet-cpp-bundle-as-the-embedding-model" %}}).
|
||||
|
||||
### Identifying speakers without gating
|
||||
|
||||
To recognize who is speaking and surface it to the client and the LLM without ever rejecting a turn, set `enforce: false` and add an `identity` block. The `identity` block works with or without the gate; when it is set, the speaker is resolved on every turn even if `when: first`.
|
||||
|
||||
@@ -206,6 +206,40 @@ embedding size, so registering a voice of a new size no longer fails.
|
||||
- Naming in diarization and live transcription reads the registry as a whole
|
||||
and uses the voices that match the loaded encoder.
|
||||
|
||||
## A parakeet-cpp bundle as the embedding model
|
||||
|
||||
A parakeet-cpp model that has a speaker encoder can serve `/v1/voice/embed`,
|
||||
`/v1/voice/verify`, `/v1/voice/register` and `/v1/voice/identify`, and the
|
||||
`voice_recognition` stage of a realtime pipeline. That is every
|
||||
[bundle]({{% relref "audio-to-text#bundle-gguf-files-several-models-in-one-file" %}}) with a `voice` component
|
||||
(`parakeet-cpp-bundle-small` and `parakeet-cpp-bundle-standard`), and any
|
||||
parakeet-cpp config with a `speaker_model:` or `speaker_component:` option, such as
|
||||
`parakeet-cpp-realtime-scene-speakers`. Declare `speaker_recognition` in the
|
||||
`known_usecases` of the config; the gallery entries above already do. Then use the model
|
||||
name as `model`:
|
||||
|
||||
```bash
|
||||
local-ai models install parakeet-cpp-bundle-small
|
||||
curl http://localhost:8080/v1/voice/embed -H "Content-Type: application/json" \
|
||||
-d '{"model": "parakeet-cpp-bundle-small", "audio": "https://example.com/clip.wav"}'
|
||||
```
|
||||
|
||||
The `voice` component holds the same weights as `voice-detect-wespeaker-resnet34`
|
||||
(256 dimensions), so a voice registered with one model also matches embeddings from the
|
||||
other. Use a distance threshold near 0.5 for this encoder, as for `speaker_model:`
|
||||
naming below. The response `model` field is the `sha256:` identity of the encoder
|
||||
weights.
|
||||
|
||||
Only embedding, identification and plain verification are available. There is no
|
||||
anti-spoofing head, so a verify request with `anti_spoofing: true` is refused, and
|
||||
`/v1/voice/analyze` is not served by parakeet-cpp. The backend needs a libparakeet.so
|
||||
that exports `parakeet_capi_speaker_embed_pcm`. With an older library these calls fail
|
||||
with an `Unimplemented` error that names the missing symbol, and a model without a
|
||||
speaker encoder fails with a `FailedPrecondition` error.
|
||||
|
||||
Voices of other embedding sizes stay in their own store, see
|
||||
[Voices from different encoders](#voices-from-different-encoders).
|
||||
|
||||
## Naming speakers in diarization and live transcription
|
||||
|
||||
The parakeet-cpp backend can put the names of registered voices on
|
||||
|
||||
Reference in new issue
Block a user