diff --git a/docs/content/features/audio-to-text.md b/docs/content/features/audio-to-text.md index 9dc13520e..de0eb3ec2 100644 --- a/docs/content/features/audio-to-text.md +++ b/docs/content/features/audio-to-text.md @@ -350,6 +350,7 @@ The component that each role uses, and the option that picks another one: | Diarization | the `diar` component, only when asked for | `diar_component:` | | Sound events | the `ced` component, only when asked for | `sound_component:` | | Speaker naming | the `voice` component, only when asked for | `speaker_component:` (needs a diarization component) | +| Voice embedding and verification (`/v1/voice/*`, the realtime `voice_recognition` stage) | the `voice` component, only when asked for | `speaker_component:`; declare `speaker_recognition` in `known_usecases`. Needs a libparakeet.so with `parakeet_capi_speaker_embed_pcm`. See [Voice Recognition]({{% relref "voice-recognition#a-parakeet-cpp-bundle-as-the-embedding-model" %}}) | `speaker_component:` also names registered speakers: LocalAI sends the voices from `/v1/voice/register` to the bundle's speaker component, as it does for `speaker_model:`. This works in `/v1/audio/diarization` and in realtime live transcription, and it needs no `speaker_model:`. `speaker_threshold`, `speaker_margin` and `speaker_strict` apply too. Only voices that carry an encoder fingerprint (voices enrolled from `speaker_profiles`) and voices with no tag at all are used by default. A voice registered with a tag only matches through `speaker_tag:`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-from-a-bundle). diff --git a/docs/content/features/openai-realtime.md b/docs/content/features/openai-realtime.md index 017e62c7d..f4ba114b2 100644 --- a/docs/content/features/openai-realtime.md +++ b/docs/content/features/openai-realtime.md @@ -547,6 +547,33 @@ pipeline: audio: /models/voices/bob.wav ``` +### One bundle for every stage + +A parakeet-cpp bundle can fill the speaker stage as well as the others, so one model serves +`vad`, `transcription`, `sound_detection` and `voice_recognition`. The bundle config must +list `speaker_recognition` in `known_usecases` (the gallery entries do), and the libparakeet.so +must export `parakeet_capi_speaker_embed_pcm`: + +```yaml +name: my-realtime +pipeline: + vad: parakeet-cpp-bundle-small + transcription: parakeet-cpp-bundle-small + sound_detection: parakeet-cpp-bundle-small + llm: qwen + tts: kokoro + voice_recognition: + model: parakeet-cpp-bundle-small + mode: identify + threshold: 0.5 # WeSpeaker distance; see Voice Recognition +``` + +The bundle's encoder has the same weights as `voice-detect-wespeaker-resnet34`, so voices +registered with that model are matched. Embedding an utterance uses the same engine lock as +transcription, so it adds to the turn latency. Voices from a 192-dimension model (ECAPA-TDNN, +for example) cannot be matched by a bundle; see +[Voice Recognition]({{% relref "voice-recognition#a-parakeet-cpp-bundle-as-the-embedding-model" %}}). + ### Identifying speakers without gating To recognize who is speaking and surface it to the client and the LLM without ever rejecting a turn, set `enforce: false` and add an `identity` block. The `identity` block works with or without the gate; when it is set, the speaker is resolved on every turn even if `when: first`. diff --git a/docs/content/features/voice-recognition.md b/docs/content/features/voice-recognition.md index 932ce0a82..625015b32 100644 --- a/docs/content/features/voice-recognition.md +++ b/docs/content/features/voice-recognition.md @@ -206,6 +206,40 @@ embedding size, so registering a voice of a new size no longer fails. - Naming in diarization and live transcription reads the registry as a whole and uses the voices that match the loaded encoder. +## A parakeet-cpp bundle as the embedding model + +A parakeet-cpp model that has a speaker encoder can serve `/v1/voice/embed`, +`/v1/voice/verify`, `/v1/voice/register` and `/v1/voice/identify`, and the +`voice_recognition` stage of a realtime pipeline. That is every +[bundle]({{% relref "audio-to-text#bundle-gguf-files-several-models-in-one-file" %}}) with a `voice` component +(`parakeet-cpp-bundle-small` and `parakeet-cpp-bundle-standard`), and any +parakeet-cpp config with a `speaker_model:` or `speaker_component:` option, such as +`parakeet-cpp-realtime-scene-speakers`. Declare `speaker_recognition` in the +`known_usecases` of the config; the gallery entries above already do. Then use the model +name as `model`: + +```bash +local-ai models install parakeet-cpp-bundle-small +curl http://localhost:8080/v1/voice/embed -H "Content-Type: application/json" \ + -d '{"model": "parakeet-cpp-bundle-small", "audio": "https://example.com/clip.wav"}' +``` + +The `voice` component holds the same weights as `voice-detect-wespeaker-resnet34` +(256 dimensions), so a voice registered with one model also matches embeddings from the +other. Use a distance threshold near 0.5 for this encoder, as for `speaker_model:` +naming below. The response `model` field is the `sha256:` identity of the encoder +weights. + +Only embedding, identification and plain verification are available. There is no +anti-spoofing head, so a verify request with `anti_spoofing: true` is refused, and +`/v1/voice/analyze` is not served by parakeet-cpp. The backend needs a libparakeet.so +that exports `parakeet_capi_speaker_embed_pcm`. With an older library these calls fail +with an `Unimplemented` error that names the missing symbol, and a model without a +speaker encoder fails with a `FailedPrecondition` error. + +Voices of other embedding sizes stay in their own store, see +[Voices from different encoders](#voices-from-different-encoders). + ## Naming speakers in diarization and live transcription The parakeet-cpp backend can put the names of registered voices on