* feat(voice): list registered voices and record which encoder made them The voice registry could register, identify and forget but not list, and it did not remember which speaker encoder produced an embedding. Add Metadata.Model and Registry.List, answered from the index the store registry already keeps for Forget. Needed so a backend can be given the registered voices that match its own speaker encoder. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(voice): store the encoder model with a registered voice Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(voice): pick the registered voices that match a speaker model Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(proto): carry known voices and speaker names on diarize and live messages Assisted-by: Claude:claude-haiku-4-5 [Claude Code] * feat(diarization): name speakers from the voice registry When a diarization model has a speaker_model option, the endpoint sends the registered voices made by that encoder to the backend. The backend's name and name_score come back as extra fields next to the normalized SPEAKER_NN speaker, and the speakers summary carries the first name seen for each speaker. RTTM output and results without names are unchanged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(live): pass registered voices to a live session and surface speaker names Live sessions now send the registered voices that match the model's speaker_model to the backend, and each speaker segment carries the name the backend matched. The realtime segment event gains an optional speaker_name field. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): load a speaker model and build per-request voice registries Adds the speaker bindings (ABI v9 and v10, probed separately), the speaker_model, speaker_threshold and speaker_margin options, and a per-request registry builder over the known voices. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): name the speakers in Diarize from the known voices Diarize builds a per-request speaker registry from the known voices when a speaker model is loaded, calls the named C functions, and puts each slot's registered name and score on the segments. The registry is freed on every path. A library without ABI 10 reports Unimplemented instead of dropping the names. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): name speakers in the live scene stream The live scene stream now begins with a known-voice registry when a speaker model is loaded and the live config carries voices, and each closed speaker segment takes its slot's current name from the feed's names map. A segment that closes before its slot is identified has an empty name. The registry is freed after the stream, on every path. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(gallery): speaker naming entries and docs for parakeet-cpp Add three gallery entries that load the WeSpeaker ResNet34 speaker model next to the diarization or realtime scene models, and document speaker names in the voice recognition, diarization, audio to text and realtime pages. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(parakeet-cpp): skip an unusable registered voice instead of failing the request A registered voice with the wrong embedding size, or one the C side refused, failed the whole diarization request, so one legacy voice broke the model for every user. Skip such voices with a warning that does not carry the voice name, and take the plain path when none is left. Also map an exact 0 speaker threshold or margin to a tiny positive value, since the C side reads 0 as "use the default", and fix a stale comment about which contexts Free() walks. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(diarization): warn once per model about voices from another encoder; document the privacy limit The different-encoder warning fired on every request. Log it once per feature and speaker model, then at debug level. Document that the global voice registry lets any caller of a speaker_model model learn matching names, and that skipped wrong-sized voices are logged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * chore(parakeet-cpp): bump parakeet.cpp to 8c8cec0 (C-API v10) and check speaker naming against the real library The pin moves from 623a968 to 8c8cec0, which brings in everything merged in parakeet.cpp since: the voice identification change (C-API v9, #78) and raw-embedding enroll plus diarize-only speaker naming (C-API v10, #79). New real-library specs (gated on PARAKEET_BACKEND_TEST_SPEAKER_MODEL, _DIAR_MODEL, _WAV and, for the live path, _STREAM_MODEL) name the two speakers of two_speakers.wav from a committed pair of WeSpeaker embeddings, with the voices passed in reversed order. They also check that the float32 threshold reaches C through purego. The shared test loader now registers the v9/v10 and scene symbols as main.go does. The rebase onto origin/master had no conflicts. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
15 KiB
+++ disableToc = false title = "Voice Recognition" weight = 36 url = "/features/voice-recognition/" +++
LocalAI supports voice (speaker) recognition: speaker verification (1:1), speaker identification (1:N) against a built-in vector store, speaker embedding, and demographic analysis (age / gender / emotion from voice).
The audio analog to Face Recognition,
served over the same /v1/voice/* HTTP API by two backends:
voice-detect(recommended, default). A standalone C++/ggml engine (voice-detect.cpp): no Python, no onnxruntime, no torch runtime. Each gallery entry is a single self-describing GGUF. This is the recommended option for new deployments.speaker-recognition(Python). The original SpeechBrain / ONNX backend. Still supported; see the Python backend below.
Both backends expose the identical wire format, so the API examples on
this page work with either - only the gallery entry name (the model
field) changes.
voice-detect (ggml) backend
The voice-detect backend reads the embedding (or analysis)
architecture (voicedetect.arch) directly from the GGUF metadata, so
installing a gallery entry is all that is needed to select an engine. It
drives the VoiceEmbed / VoiceVerify / VoiceAnalyze gRPC rpcs behind the
/v1/voice/{embed,verify,analyze,register,identify,forget} endpoints.
Gallery entries
| Gallery entry | Model | Embedding dim | License |
|---|---|---|---|
voice-detect-ecapa-tdnn |
SpeechBrain ECAPA-TDNN (VoxCeleb) | 192 | Apache 2.0 - commercial-safe |
voice-detect-wespeaker-resnet34 |
WeSpeaker ResNet34 (VoxCeleb) | 256 | CC-BY-4.0 |
voice-detect-eres2net |
3D-Speaker ERes2Net (VoxCeleb) | 192 | Apache 2.0 - commercial-safe |
voice-detect-campplus |
3D-Speaker CAM++ (VoxCeleb) | 192 | Apache 2.0 - commercial-safe |
voice-detect-emotion-wav2vec2 |
audEERING wav2vec2 (age / gender / emotion) | analyze head | CC-BY-NC-SA-4.0 - non-commercial |
The four speaker-recognition entries drive verify / embed / identify.
voice-detect-emotion-wav2vec2 is the analysis head behind
/v1/voice/analyze (continuous age estimate plus gender and emotion
class scores) and is non-commercial / research use only.
Quickstart
Install the default entry (recommended for copy-paste):
local-ai models install voice-detect-ecapa-tdnn
Verify that two audio clips were spoken by the same person:
curl -sX POST http://localhost:8080/v1/voice/verify \
-H "Content-Type: application/json" \
-d '{
"model": "voice-detect-ecapa-tdnn",
"audio1": "https://example.com/alice_1.wav",
"audio2": "https://example.com/alice_2.wav"
}'
Analyze age / gender / emotion (install the analyze entry first):
local-ai models install voice-detect-emotion-wav2vec2
curl -sX POST http://localhost:8080/v1/voice/analyze \
-H "Content-Type: application/json" \
-d '{"model": "voice-detect-emotion-wav2vec2", "audio": "https://example.com/alice.wav"}'
The 1:N register / identify / forget workflow and the rest of the API
are identical to the API reference below - just pass a
voice-detect-* model name. The default verify threshold is ~0.25 for
the ECAPA-TDNN / ERes2Net / CAM++ recognizers and ~0.30 for WeSpeaker
ResNet34.
speaker-recognition (Python) backend
The speaker-recognition backend follows the same two-engine pattern
under one image.
Engines
| Gallery entry | Model | Size | License |
|---|---|---|---|
speechbrain-ecapa-tdnn |
ECAPA-TDNN on VoxCeleb (SpeechBrain) | ~17 MB | Apache 2.0 - commercial-safe |
wespeaker-resnet34 |
WeSpeaker ResNet34 ONNX | ~26 MB | Apache 2.0 - commercial-safe |
Both entries are commercial-safe Apache-2.0. SpeechBrain is the
default - it's a lightweight pure-PyTorch checkpoint that auto-
downloads on first use. The wespeaker-resnet34 entry wires the
direct-ONNX path for CPU-only deployments that don't want the torch
runtime.
Quickstart
Install the default backend and model:
local-ai models install speechbrain-ecapa-tdnn
Verify that two audio clips were spoken by the same person:
curl -sX POST http://localhost:8080/v1/voice/verify \
-H "Content-Type: application/json" \
-d '{
"model": "speechbrain-ecapa-tdnn",
"audio1": "https://example.com/alice_1.wav",
"audio2": "https://example.com/alice_2.wav"
}'
Response:
{
"verified": true,
"distance": 0.18,
"threshold": 0.25,
"confidence": 28.0,
"model": "speechbrain-ecapa-tdnn",
"processing_time_ms": 340.0
}
1:N identification workflow (register → identify → forget)
Same flow as face recognition, same in-memory vector store under the hood.
-
Register known speakers:
curl -sX POST http://localhost:8080/v1/voice/register \ -H "Content-Type: application/json" \ -d '{ "model": "speechbrain-ecapa-tdnn", "name": "Alice", "audio": "https://example.com/alice.wav" }' # → {"id": "b2f...", "name": "Alice", "registered_at": "2026-04-22T..."} -
Identify an unknown probe:
curl -sX POST http://localhost:8080/v1/voice/identify \ -H "Content-Type: application/json" \ -d '{ "model": "speechbrain-ecapa-tdnn", "audio": "https://example.com/unknown.wav", "top_k": 5 }' # → {"matches": [{"id":"b2f...","name":"Alice","distance":0.19,"match":true,...}]} -
Remove a speaker by ID:
curl -sX POST http://localhost:8080/v1/voice/forget \ -d '{"id": "b2f..."}' # → 204 No Content
{{% notice warning %}} Storage caveat. The default vector store is in-memory. All registered speakers are lost when LocalAI restarts. Persistent storage (pgvector) is a tracked future enhancement shared with face recognition - the voice-recognition HTTP API is designed to swap the backing store without changing the wire format. {{% /notice %}}
Naming speakers in diarization and live transcription
The parakeet-cpp backend can put the names of registered voices on
diarization results and on live transcription speaker segments. Without
this, speakers only carry labels such as SPEAKER_00.
- Register each voice with the WeSpeaker encoder. Install the model with
local-ai models install voice-detect-wespeaker-resnet34, then call/v1/voice/registerwith"model": "voice-detect-wespeaker-resnet34"(see the 1:N workflow). - Install one of the gallery models that loads the same encoder:
parakeet-cpp-nemotron-3-diarization-speakers(diarization),parakeet-cpp-nemotron-3-diarization-asr-speakers(diarization withinclude_text) orparakeet-cpp-realtime-scene-speakers(live transcription). Each one addsspeaker_model:voice-detect-wespeaker-resnet34.ggufto a parakeet-cpp model config. - Call
/v1/audio/diarizationwith that model. Matched segments gain anameand aname_score, and the matching entry inspeakersgains aname.speakerstaysSPEAKER_NN, and RTTM output is unchanged. See [Speaker Diarization]({{% relref "audio-diarization" %}}) for the response.
Which voices are used
LocalAI sends the backend only the registered voices made by the same
encoder as the model's speaker_model: file. Each registered voice is
tagged with the name of the voice-detect model that made it, which by
default is the GGUF file name (voice-detect-wespeaker-resnet34.gguf for the
gallery entry). The tag must equal the base name of the speaker_model:
file. Voices made with another encoder are ignored, and LocalAI logs a
warning when that leaves no usable voice. Voices registered before the tag
existed have no tag: they are used when their embedding size matches the
tagged ones (or all of them, when no voice carries a matching tag). The
backend skips a voice whose embedding size does not match the speaker model's,
with a warning in the LocalAI log. Naming then falls back to the remaining
voices, or to no names.
{{% notice warning %}}
Do not set a model_name: option on the voice-detect model config. It
replaces the default name, the voices are then tagged with it, and they no
longer match the speaker_model: file. Keep the default name.
{{% /notice %}}
Options
These go in the options: list of the parakeet-cpp model config (see
[Audio to Text]({{% relref "audio-to-text" %}}) for the other parakeet-cpp
options).
| Option | Default | Meaning |
|---|---|---|
speaker_model:<path> |
none | speaker encoder GGUF; needs a diarization model (the primary one, or diarization_model:) |
speaker_threshold:<float> |
0.5 |
largest distance (1 minus cosine similarity, the unit /v1/voice/identify reports) at which a speaker is named; must be in (0, 2) |
speaker_margin:<float> |
0.05 |
the best match must beat the runner-up by this much, otherwise the speaker stays unnamed; must be in [0, 1) |
parakeet.cpp's measured starting values for speaker_threshold are 0.5 for
WeSpeaker ResNet34 and CAM++, and 0.3 for ECAPA. A lower value names fewer
speakers and makes fewer mistakes.
Limits
- The voice registry is in memory and global. Registered names disappear when LocalAI restarts, and every user of the instance shares them.
- Anyone who is allowed to call a model with
speaker_model:can learn which registered names match their audio, and their audio is matched against voices registered by any user, because the voice registry is global. Restrict such models with the per-user model allowlist. - With
include_text=truethe names use the default threshold and margin:speaker_thresholdandspeaker_marginonly apply to diarization without text. - In live transcription, a speaker segment that closes before its speaker is identified has no name. Later segments of that speaker do.
- Overlapping speech is not resolved.
- Accuracy was measured on one fixture (two read-speech voices). Check the threshold on your own audio.
- The backend needs a libparakeet with C-API v10. With an older library a
model config that sets
speaker_model:fails to load.
API reference
POST /v1/voice/verify (1:1)
| field | type | description |
|---|---|---|
model |
string | gallery entry name (e.g. speechbrain-ecapa-tdnn) |
audio1, audio2 |
string | URL, base64, or data-URI of an audio file |
threshold |
float, optional | cosine-distance cutoff; default 0.25 for ECAPA-TDNN |
anti_spoofing |
bool, optional | reserved - unused in the current release |
Returns verified, distance, threshold, confidence, model,
and processing_time_ms.
POST /v1/voice/analyze
Returns demographic attributes (age, gender, emotion) inferred from speech:
| field | type | description |
|---|---|---|
model |
string | gallery entry |
audio |
string | URL / base64 / data-URI |
actions |
string[] | subset of ["age","gender","emotion"]; empty = all supported |
Emotion is inferred from the SUPERB emotion-recognition checkpoint
(superb/wav2vec2-base-superb-er, Apache 2.0) - 4-way categorical
neutral / happy / angry / sad. The model auto-downloads on the first
analyze call.
Age and gender are opt-in: no standard-transformers checkpoint
with a clean classifier head is shipped as the default. The
high-accuracy Audeering age/gender model uses a custom multi-task
head that AutoModelForAudioClassification doesn't load safely
(the age weights are silently dropped and the classifier is
re-initialised with random values). To enable age/gender, set
age_gender_model:<repo> in the model YAML's options: pointing at
a checkpoint with a vanilla Wav2Vec2ForSequenceClassification
head. Override the emotion default similarly via emotion_model:.
Set either to an empty string to disable that head.
If a head fails to load (offline, disk full, transformers
missing), the engine degrades gracefully: it still returns the
attributes it could compute. When nothing can be computed the backend
returns 501 Unimplemented.
Analyze is supported by both speechbrain-ecapa-tdnn and
wespeaker-resnet34 - the speaker recognizer and the analysis head
are independent.
POST /v1/voice/register (1:N enrollment)
| field | type | description |
|---|---|---|
model |
string | voice recognition model |
audio |
string | speaker audio to enroll |
name |
string | human-readable label |
labels |
map[string]string, optional | arbitrary metadata |
store |
string, optional | vector store model; defaults to local-store |
Returns {id, name, registered_at}. The id is an opaque UUID used
by /v1/voice/identify and /v1/voice/forget.
POST /v1/voice/identify (1:N recognition)
| field | type | description |
|---|---|---|
model |
string | voice recognition model |
audio |
string | probe audio |
top_k |
int, optional | max matches to return; default 5 |
threshold |
float, optional | cosine-distance cutoff; default 0.25 |
store |
string, optional | vector store model |
Returns a list of matches sorted by ascending distance, each with
id, name, labels, distance, confidence, and match
(distance ≤ threshold).
POST /v1/voice/forget
| field | type | description |
|---|---|---|
id |
string | ID returned by /v1/voice/register |
Returns 204 No Content on success, 404 Not Found if the ID is
unknown.
POST /v1/voice/embed
Returns the L2-normalized speaker embedding vector.
| field | type | description |
|---|---|---|
model |
string | voice model |
audio |
string | URL / base64 / data-URI |
Returns {embedding: float[], dim: int, model: string}. Dimension
depends on the recognizer: 192 for ECAPA-TDNN, 256 for WeSpeaker
ResNet34.
Note: the OpenAI-compatible
/v1/embeddingsendpoint is intentionally text-only - it does nothing useful with audio input. Use/v1/voice/embedfor audio.
Audio input
Audio is materialised by the HTTP layer to a temporary WAV file before the gRPC call. All audio fields accept:
http:///https://URLs (downloaded server-side, subject toValidateExternalURLsafety checks).- Raw base64 (no prefix).
- Data URIs (
data:audio/wav;base64,...).
The backend itself always receives a filesystem path - the same convention the Whisper / Voxtral transcription backends use.
Threshold reference
| Recognizer | Cosine-distance threshold |
|---|---|
| ECAPA-TDNN (SpeechBrain, VoxCeleb) | ~0.25 |
| WeSpeaker ResNet34 | ~0.30 |
| 3D-Speaker ERes2Net | ~0.28 |
Pass threshold explicitly when switching recognizers - the per-model
default only applies when omitted.
Related features
- Face Recognition - the image analog; the two share a registry design.
- Audio to Text - transcription (Whisper, Voxtral, faster-whisper). Runs in addition to, not instead of, voice recognition.
- Stores - the generic vector store powering both the face and voice 1:N recognition pipelines.
- Embeddings - text-only OpenAI-compatible
embedding endpoint; for audio embeddings use
/v1/voice/embed.
