mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-11 00:02:49 -04:00
* feat(diarization): return sound events with include_sounds
A client that wants text, speakers, voice prints and sound events had to
make a diarization call and a separate sound call. Add an include_sounds
request field to /v1/audio/diarization that adds a sounds array of closed
events {start, end, label, confidence}, in seconds.
The parakeet-cpp backend runs a tagger-only scene stream over the clip,
the same stream and thresholds the live path uses, so a clip gives the
same events offline and live. A model with no sound_model companion, or a
backend that does not report sound events, fails with 501 and the stable
code include_sounds_unsupported instead of an empty list. The proto
carries sounds_included so an empty list still means "nothing heard".
The localai-proxy backend forwards the field. Swagger, docs and the
e2e mock backend are updated.
Assisted-by: Claude:claude-sonnet-5-5 [protoc swag go]
* feat(gallery): add parakeet-cpp-multilingual-diarization-speakers-sounds
Same as parakeet-cpp-multilingual-diarization-speakers (TDT 0.6B v3,
Nemotron-3-Diarization, WeSpeaker) plus a CED-Tiny sound_model, so one
model name serves /v1/audio/diarization with include_text,
include_speaker_profiles and include_sounds. It declares the
sound_classification usecase like the realtime scene entries.
Assisted-by: Claude:claude-sonnet-5-5
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
608 lines
26 KiB
Markdown
608 lines
26 KiB
Markdown
+++
|
|
disableToc = false
|
|
title = "Voice Recognition"
|
|
weight = 36
|
|
url = "/features/voice-recognition/"
|
|
+++
|
|
|
|

|
|
|
|
LocalAI supports voice (speaker) recognition: speaker verification
|
|
(1:1), speaker identification (1:N) against a built-in vector store,
|
|
speaker embedding, and demographic analysis (age / gender / emotion
|
|
from voice).
|
|
|
|
The audio analog to [Face Recognition](/features/face-recognition/),
|
|
served over the same `/v1/voice/*` HTTP API by two backends:
|
|
|
|
- **`voice-detect` (recommended, default).** A standalone C++/ggml
|
|
engine ([voice-detect.cpp](https://github.com/localai-org/voice-detect.cpp)):
|
|
no Python, no onnxruntime, no torch runtime. Each gallery entry is a
|
|
single self-describing GGUF. This is the recommended option for new
|
|
deployments.
|
|
- **`speaker-recognition` (Python).** The original SpeechBrain / ONNX
|
|
backend. Still supported; see [the Python backend](#speaker-recognition-python-backend)
|
|
below.
|
|
|
|
Both backends expose the identical wire format, so the API examples on
|
|
this page work with either - only the gallery entry name (the `model`
|
|
field) changes.
|
|
|
|
## voice-detect (ggml) backend
|
|
|
|
The `voice-detect` backend reads the embedding (or analysis)
|
|
architecture (`voicedetect.arch`) directly from the GGUF metadata, so
|
|
installing a gallery entry is all that is needed to select an engine. It
|
|
drives the VoiceEmbed / VoiceVerify / VoiceAnalyze gRPC rpcs behind the
|
|
`/v1/voice/{embed,verify,analyze,register,identify,forget}` endpoints.
|
|
|
|
### Gallery entries
|
|
|
|
| Gallery entry | Model | Embedding dim | License |
|
|
|---|---|---|---|
|
|
| `voice-detect-ecapa-tdnn` | SpeechBrain ECAPA-TDNN (VoxCeleb) | 192 | **Apache 2.0 - commercial-safe** |
|
|
| `voice-detect-wespeaker-resnet34` | WeSpeaker ResNet34 (VoxCeleb) | 256 | CC-BY-4.0 |
|
|
| `voice-detect-eres2net` | 3D-Speaker ERes2Net (VoxCeleb) | 192 | **Apache 2.0 - commercial-safe** |
|
|
| `voice-detect-campplus` | 3D-Speaker CAM++ (VoxCeleb) | 192 | **Apache 2.0 - commercial-safe** |
|
|
| `voice-detect-emotion-wav2vec2` | audEERING wav2vec2 (age / gender / emotion) | analyze head | **CC-BY-NC-SA-4.0 - non-commercial** |
|
|
|
|
The four speaker-recognition entries drive verify / embed / identify.
|
|
`voice-detect-emotion-wav2vec2` is the analysis head behind
|
|
`/v1/voice/analyze` (continuous age estimate plus gender and emotion
|
|
class scores) and is **non-commercial / research use only**.
|
|
|
|
### Quickstart
|
|
|
|
Install the default entry (recommended for copy-paste):
|
|
|
|
```bash
|
|
local-ai models install voice-detect-ecapa-tdnn
|
|
```
|
|
|
|
Verify that two audio clips were spoken by the same person:
|
|
|
|
```bash
|
|
curl -sX POST http://localhost:8080/v1/voice/verify \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "voice-detect-ecapa-tdnn",
|
|
"audio1": "https://example.com/alice_1.wav",
|
|
"audio2": "https://example.com/alice_2.wav"
|
|
}'
|
|
```
|
|
|
|
Analyze age / gender / emotion (install the analyze entry first):
|
|
|
|
```bash
|
|
local-ai models install voice-detect-emotion-wav2vec2
|
|
|
|
curl -sX POST http://localhost:8080/v1/voice/analyze \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"model": "voice-detect-emotion-wav2vec2", "audio": "https://example.com/alice.wav"}'
|
|
```
|
|
|
|
The 1:N register / identify / forget workflow and the rest of the API
|
|
are identical to the [API reference](#api-reference) below - just pass a
|
|
`voice-detect-*` model name. The default verify threshold is ~0.25 for
|
|
the ECAPA-TDNN / ERes2Net / CAM++ recognizers and ~0.30 for WeSpeaker
|
|
ResNet34.
|
|
|
|
## speaker-recognition (Python) backend
|
|
|
|
The `speaker-recognition` backend follows the same two-engine pattern
|
|
under one image.
|
|
|
|
### Engines
|
|
|
|
| Gallery entry | Model | Size | License |
|
|
|---|---|---|---|
|
|
| `speechbrain-ecapa-tdnn` | ECAPA-TDNN on VoxCeleb (SpeechBrain) | ~17 MB | **Apache 2.0 - commercial-safe** |
|
|
| `wespeaker-resnet34` | WeSpeaker ResNet34 ONNX | ~26 MB | **Apache 2.0 - commercial-safe** |
|
|
|
|
Both entries are commercial-safe Apache-2.0. SpeechBrain is the
|
|
default - it's a lightweight pure-PyTorch checkpoint that auto-
|
|
downloads on first use. The `wespeaker-resnet34` entry wires the
|
|
direct-ONNX path for CPU-only deployments that don't want the torch
|
|
runtime.
|
|
|
|
## Quickstart
|
|
|
|
Install the default backend and model:
|
|
|
|
```bash
|
|
local-ai models install speechbrain-ecapa-tdnn
|
|
```
|
|
|
|
Verify that two audio clips were spoken by the same person:
|
|
|
|
```bash
|
|
curl -sX POST http://localhost:8080/v1/voice/verify \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "speechbrain-ecapa-tdnn",
|
|
"audio1": "https://example.com/alice_1.wav",
|
|
"audio2": "https://example.com/alice_2.wav"
|
|
}'
|
|
```
|
|
|
|
Response:
|
|
|
|
```json
|
|
{
|
|
"verified": true,
|
|
"distance": 0.18,
|
|
"threshold": 0.25,
|
|
"confidence": 28.0,
|
|
"model": "speechbrain-ecapa-tdnn",
|
|
"processing_time_ms": 340.0
|
|
}
|
|
```
|
|
|
|
## 1:N identification workflow (register → identify → forget)
|
|
|
|
Same flow as face recognition, same in-memory vector store under the
|
|
hood.
|
|
|
|
1. Register known speakers:
|
|
|
|
```bash
|
|
curl -sX POST http://localhost:8080/v1/voice/register \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "speechbrain-ecapa-tdnn",
|
|
"name": "Alice",
|
|
"audio": "https://example.com/alice.wav"
|
|
}'
|
|
# → {"id": "b2f...", "name": "Alice", "registered_at": "2026-04-22T..."}
|
|
```
|
|
|
|
2. Identify an unknown probe:
|
|
|
|
```bash
|
|
curl -sX POST http://localhost:8080/v1/voice/identify \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"model": "speechbrain-ecapa-tdnn",
|
|
"audio": "https://example.com/unknown.wav",
|
|
"top_k": 5
|
|
}'
|
|
# → {"matches": [{"id":"b2f...","name":"Alice","distance":0.19,"match":true,...}]}
|
|
```
|
|
|
|
3. Remove a speaker by ID:
|
|
|
|
```bash
|
|
curl -sX POST http://localhost:8080/v1/voice/forget \
|
|
-d '{"id": "b2f..."}'
|
|
# → 204 No Content
|
|
```
|
|
|
|
{{% notice warning %}}
|
|
**Storage caveat.** The default vector store is in-memory. All
|
|
registered speakers are lost when LocalAI restarts. Persistent storage
|
|
(pgvector) is a tracked future enhancement shared with face
|
|
recognition - the voice-recognition HTTP API is designed to swap the
|
|
backing store without changing the wire format.
|
|
{{% /notice %}}
|
|
|
|
### Voices from different encoders
|
|
|
|
Voices from encoders with different embedding sizes can be registered on the
|
|
same instance, for example 192-value ECAPA-TDNN voices next to 256-value
|
|
WeSpeaker ResNet34 voices. LocalAI keeps one in-memory vector store per
|
|
embedding size, so registering a voice of a new size no longer fails.
|
|
|
|
- Identification compares the probe only with voices of the same size. Voices
|
|
from an encoder of another size are never candidates.
|
|
- If no voice of the probe size is registered, `/v1/voice/identify` returns an
|
|
empty `matches` list and the realtime voice gate reports an unknown speaker.
|
|
Neither returns an error.
|
|
- Two encoders can still give the same size (ECAPA and CAM++ both give 192
|
|
values). Those voices share a store and the encoder tag or identity checks
|
|
described below apply.
|
|
- A name is not unique. One name can hold one voice per encoder, and each
|
|
registration has its own ID. `/v1/voice/forget` removes the voice with that
|
|
ID only, so forget each ID to remove a person from every encoder.
|
|
- Naming in diarization and live transcription reads the registry as a whole
|
|
and uses the voices that match the loaded encoder.
|
|
|
|
## A parakeet-cpp bundle as the embedding model
|
|
|
|
A parakeet-cpp model that has a speaker encoder can serve `/v1/voice/embed`,
|
|
`/v1/voice/verify`, `/v1/voice/register` and `/v1/voice/identify`, and the
|
|
`voice_recognition` stage of a realtime pipeline. That is every
|
|
[bundle]({{% relref "audio-to-text#bundle-gguf-files-several-models-in-one-file" %}}) with a `voice` component
|
|
(`parakeet-cpp-bundle-small` and `parakeet-cpp-bundle-standard`), and any
|
|
parakeet-cpp config with a `speaker_model:` or `speaker_component:` option, such as
|
|
`parakeet-cpp-realtime-scene-speakers`. Declare `speaker_recognition` in the
|
|
`known_usecases` of the config; the gallery entries above already do. Then use the model
|
|
name as `model`:
|
|
|
|
```bash
|
|
local-ai models install parakeet-cpp-bundle-small
|
|
curl http://localhost:8080/v1/voice/embed -H "Content-Type: application/json" \
|
|
-d '{"model": "parakeet-cpp-bundle-small", "audio": "https://example.com/clip.wav"}'
|
|
```
|
|
|
|
The `voice` component holds the same weights as `voice-detect-wespeaker-resnet34`
|
|
(256 dimensions), so a voice registered with one model also matches embeddings from the
|
|
other. Use a distance threshold near 0.5 for this encoder, as for `speaker_model:`
|
|
naming below. The response `model` field is the `sha256:` identity of the encoder
|
|
weights.
|
|
|
|
`/v1/voice/verify` uses a distance threshold of 0.5 when the request has none. Set the
|
|
model option `voice_verify_threshold:<distance>` (a distance in (0, 2)) to change it.
|
|
|
|
Only embedding, identification and plain verification are available. There is no
|
|
anti-spoofing head, so a verify request with `anti_spoofing: true` is refused, and
|
|
`/v1/voice/analyze` is not served by parakeet-cpp. The backend needs a libparakeet.so
|
|
that exports `parakeet_capi_speaker_embed_pcm`. With an older library these calls fail
|
|
with an `Unimplemented` error that names the missing symbol, and a model without a
|
|
speaker encoder fails with a `FailedPrecondition` error.
|
|
|
|
Voices of other embedding sizes stay in their own store, see
|
|
[Voices from different encoders](#voices-from-different-encoders).
|
|
|
|
## Naming speakers in diarization and live transcription
|
|
|
|
The parakeet-cpp backend can put the names of registered voices on
|
|
diarization results and on live transcription speaker segments. Without
|
|
this, speakers only carry labels such as `SPEAKER_00`.
|
|
|
|
1. Register each voice with the WeSpeaker encoder. Install the model with
|
|
`local-ai models install voice-detect-wespeaker-resnet34`, then call
|
|
`/v1/voice/register` with `"model": "voice-detect-wespeaker-resnet34"`
|
|
(see the [1:N workflow](#1n-identification-workflow-register--identify--forget)).
|
|
2. Install one of the gallery models that loads the same encoder:
|
|
`parakeet-cpp-nemotron-3-diarization-speakers` (diarization),
|
|
`parakeet-cpp-nemotron-3-diarization-asr-speakers` (diarization with
|
|
`include_text`), `parakeet-cpp-multilingual-diarization-speakers-sounds`
|
|
(multilingual diarization with `include_text`, voice prints and sound events)
|
|
or `parakeet-cpp-realtime-scene-speakers` (live
|
|
transcription). Each one adds
|
|
`speaker_model:voice-detect-wespeaker-resnet34.gguf` to a
|
|
parakeet-cpp model config.
|
|
3. Call `/v1/audio/diarization` with that model. Matched segments gain a
|
|
`name` and a `name_score`, and the matching entry in `speakers` gains a
|
|
`name`. `speaker` stays `SPEAKER_NN`, and RTTM output is unchanged. See
|
|
[Speaker Diarization]({{% relref "audio-diarization" %}}) for the
|
|
response.
|
|
|
|
### Naming speakers from a bundle
|
|
|
|
A parakeet-cpp bundle file holds the speaker encoder as a component, so its
|
|
config has `speaker_component:voice` and no `speaker_model:` (the
|
|
`parakeet-cpp-bundle-small` and `parakeet-cpp-bundle-standard` gallery
|
|
entries are set up like this). LocalAI treats `speaker_component:` as the
|
|
speaker encoder of the model when `speaker_model:` is not set. A
|
|
`speaker_model:` entry always wins, including one that points at the bundle
|
|
file itself.
|
|
|
|
A bundle has no encoder file name, so LocalAI cannot compare file-name tags,
|
|
and it never guesses a tag from the bundle file name. For a bundle component
|
|
LocalAI sends the backend:
|
|
|
|
- voices with an [encoder fingerprint](#encoder-fingerprint), whatever their
|
|
tag: the backend checks the fingerprint against the loaded component.
|
|
Voices enrolled from `speaker_profiles` are in this group. The WeSpeaker
|
|
component of the published bundles has the same weights as
|
|
`voice-detect-wespeaker-resnet34`, so a voice enrolled with one works with
|
|
the other;
|
|
- voices with no tag (registered before tags existed);
|
|
- voices whose tag equals the `speaker_tag:` option, if set.
|
|
|
|
Other voices are ignored. To use a voice that carries only a file-name tag,
|
|
set the same tag on the model config. For voices registered through the
|
|
`voice-detect-wespeaker-resnet34` gallery model:
|
|
|
|
```yaml
|
|
options:
|
|
- speaker_component:voice
|
|
- speaker_tag:voice-detect-wespeaker-resnet34.gguf
|
|
```
|
|
|
|
Use `speaker_tag:` only for voices made by the same weights as the bundle
|
|
component. The check by size alone cannot tell two encoders apart.
|
|
|
|
### Which voices are used
|
|
|
|
LocalAI sends the backend only the registered voices made by the same
|
|
encoder as the model's `speaker_model:` file. Each registered voice is
|
|
tagged with the name of the voice-detect model that made it, which by
|
|
default is the GGUF file name (`voice-detect-wespeaker-resnet34.gguf` for the
|
|
gallery entry). The tag must equal the base name of the `speaker_model:`
|
|
file. Voices made with another encoder are ignored, and LocalAI logs a
|
|
warning when that leaves no usable voice. Voices registered before the tag
|
|
existed have no tag: they are used when their embedding size matches the
|
|
tagged ones (or all of them, when no voice carries a matching tag). The
|
|
backend skips a voice whose embedding size does not match the speaker model's,
|
|
with a warning in the LocalAI log. Naming then falls back to the remaining
|
|
voices, or to no names.
|
|
|
|
### Encoder fingerprint
|
|
|
|
Two encoders can give embeddings of the same size (ECAPA and CAM++ both give
|
|
192 values), so a size match does not prove the voices and the `speaker_model:`
|
|
file share an embedding space. A voice enrolled from `speaker_profiles` (see
|
|
[Speaker Diarization]({{% relref "audio-diarization" %}})) is stored with the
|
|
encoder that made it: its **weights** (`sha256:` of the encoder file, kept in
|
|
the voice's `model` field as before) and its **family**
|
|
(`voicedetect:<arch>:<name>:<dim>`, read from the encoder GGUF metadata and
|
|
stored as `encoder_family`). The parakeet-cpp backend builds the registry with
|
|
that fingerprint, and libparakeet checks it against the loaded `speaker_model:`
|
|
before it names anyone:
|
|
|
|
| Registered voices | Result |
|
|
|---|---|
|
|
| Same family, same weights | names are assigned |
|
|
| Same family, other weights (for example another quantization) | names are assigned, the library logs a warning |
|
|
| Another family, and no other usable voice | the request fails, and the error names both families |
|
|
| Another family, with usable voices of the right family | the other voices are left out, with a warning |
|
|
| No fingerprint | used as before, with a warning that the encoder is unverified |
|
|
|
|
A voice with only a weights identity takes the family of the loaded encoder
|
|
when the weights are the same file. A voice with a different weights hash and
|
|
no family is dropped, as before.
|
|
|
|
A voice registered from audio through the voice-detect backend has no
|
|
fingerprint: libvoicedetect reports no architecture or model name, so the
|
|
backend cannot tell the family, and only the file-name tag described above
|
|
applies. Such voices and fingerprinted voices cannot share one registry in the
|
|
library. When a request has any unfingerprinted voice, all of its voices are
|
|
used without the fingerprint check (the old behaviour). To get the check, enroll
|
|
every voice from `speaker_profiles`. With `speaker_strict:true` the backend
|
|
ignores unfingerprinted voices, and a request that has only those fails with
|
|
the library's message. The family is also reported in the internal backend
|
|
status next to the identity.
|
|
|
|
{{% notice warning %}}
|
|
Do not set a `model_name:` option on the voice-detect model config. It
|
|
replaces the default name, the voices are then tagged with it, and they no
|
|
longer match the `speaker_model:` file. Keep the default name.
|
|
{{% /notice %}}
|
|
|
|
### Options
|
|
|
|
These go in the `options:` list of the parakeet-cpp model config (see
|
|
[Audio to Text]({{% relref "audio-to-text" %}}) for the other parakeet-cpp
|
|
options).
|
|
|
|
| Option | Default | Meaning |
|
|
|---|---|---|
|
|
| `speaker_model:<path>` | none | speaker encoder GGUF; needs a diarization model (the primary one, or `diarization_model:`) |
|
|
| `speaker_component:<name>` | none | speaker encoder component of a bundle file; used as the speaker encoder for naming when `speaker_model:` is not set (see [Naming speakers from a bundle](#naming-speakers-from-a-bundle)) |
|
|
| `speaker_tag:<tag>` | none | encoder tag that also counts as the loaded encoder, for voices that carry only a file-name tag; mainly for a bundle component |
|
|
| `speaker_threshold:<float>` | `0.5` | largest distance (1 minus cosine similarity, the unit `/v1/voice/identify` reports) at which a speaker is named; must be in (0, 2) |
|
|
| `speaker_margin:<float>` | `0.05` | the best match must beat the runner-up by this much, otherwise the speaker stays unnamed; must be in [0, 1) |
|
|
| `speaker_strict:<bool>` | `false` | ignore registered voices that carry no [encoder fingerprint](#encoder-fingerprint); needs a libparakeet that exports `parakeet_capi_speaker_registry_set_strict` |
|
|
|
|
parakeet.cpp's measured starting values for `speaker_threshold` are 0.5 for
|
|
WeSpeaker ResNet34 and CAM++, and 0.3 for ECAPA. A lower value names fewer
|
|
speakers and makes fewer mistakes.
|
|
|
|
### Limits
|
|
|
|
- The voice registry is in memory and global. Registered names disappear when
|
|
LocalAI restarts, and every user of the instance shares them.
|
|
- Anyone who is allowed to call a model with `speaker_model:` or `speaker_component:` can learn which
|
|
registered names match their audio, and their audio is matched against voices
|
|
registered by any user, because the voice registry is global. Restrict such
|
|
models with the per-user model allowlist.
|
|
- With `include_text=true` the names use the default threshold and margin:
|
|
`speaker_threshold` and `speaker_margin` only apply to diarization without
|
|
text.
|
|
- In live transcription, a speaker segment that closes before its speaker
|
|
is identified has no name. Later segments of that speaker do.
|
|
- Overlapping speech is not resolved.
|
|
- Accuracy was measured on one fixture (two read-speech voices). Check the
|
|
threshold on your own audio.
|
|
- The backend needs a libparakeet with C-API v10. With an older library a
|
|
model config that sets `speaker_model:` or `speaker_component:` fails to load.
|
|
- A bundle does not serve `/v1/voice/register`, `/v1/voice/identify` or
|
|
`/v1/voice/verify` itself. Register voices with a voice-detect model (or
|
|
from `speaker_profiles`), then name them through the bundle.
|
|
|
|
## API reference
|
|
|
|
### `POST /v1/voice/verify` (1:1)
|
|
|
|
| field | type | description |
|
|
|---|---|---|
|
|
| `model` | string | gallery entry name (e.g. `speechbrain-ecapa-tdnn`) |
|
|
| `audio1`, `audio2` | string | URL, base64, or data-URI of an audio file |
|
|
| `threshold` | float, optional | cosine-distance cutoff; default 0.25 for ECAPA-TDNN |
|
|
| `anti_spoofing` | bool, optional | reserved - unused in the current release |
|
|
|
|
Returns `verified`, `distance`, `threshold`, `confidence`, `model`,
|
|
and `processing_time_ms`.
|
|
|
|
### `POST /v1/voice/analyze`
|
|
|
|
Returns demographic attributes (age, gender, emotion) inferred from
|
|
speech:
|
|
|
|
| field | type | description |
|
|
|---|---|---|
|
|
| `model` | string | gallery entry |
|
|
| `audio` | string | URL / base64 / data-URI |
|
|
| `actions` | string[] | subset of `["age","gender","emotion"]`; empty = all supported |
|
|
|
|
Emotion is inferred from the SUPERB emotion-recognition checkpoint
|
|
(`superb/wav2vec2-base-superb-er`, Apache 2.0) - 4-way categorical
|
|
neutral / happy / angry / sad. The model auto-downloads on the first
|
|
analyze call.
|
|
|
|
Age and gender are **opt-in**: no standard-transformers checkpoint
|
|
with a clean classifier head is shipped as the default. The
|
|
high-accuracy Audeering age/gender model uses a custom multi-task
|
|
head that `AutoModelForAudioClassification` doesn't load safely
|
|
(the age weights are silently dropped and the classifier is
|
|
re-initialised with random values). To enable age/gender, set
|
|
`age_gender_model:<repo>` in the model YAML's `options:` pointing at
|
|
a checkpoint with a vanilla `Wav2Vec2ForSequenceClassification`
|
|
head. Override the emotion default similarly via `emotion_model:`.
|
|
Set either to an empty string to disable that head.
|
|
|
|
If a head fails to load (offline, disk full, `transformers`
|
|
missing), the engine degrades gracefully: it still returns the
|
|
attributes it could compute. When nothing can be computed the backend
|
|
returns `501 Unimplemented`.
|
|
|
|
Analyze is supported by both `speechbrain-ecapa-tdnn` and
|
|
`wespeaker-resnet34` - the speaker recognizer and the analysis head
|
|
are independent.
|
|
|
|
### `POST /v1/voice/register` (1:N enrollment)
|
|
|
|
| field | type | description |
|
|
|---|---|---|
|
|
| `model` | string | voice recognition model |
|
|
| `audio` | string | speaker audio to enroll |
|
|
| `name` | string | human-readable label |
|
|
| `labels` | map[string]string, optional | arbitrary metadata |
|
|
| `store` | string, optional | vector store model; defaults to local-store |
|
|
|
|
Returns `{id, name, registered_at}`. The `id` is an opaque UUID used
|
|
by `/v1/voice/identify` and `/v1/voice/forget`.
|
|
|
|
### `POST /v1/voice/identify` (1:N recognition)
|
|
|
|
| field | type | description |
|
|
|---|---|---|
|
|
| `model` | string | voice recognition model |
|
|
| `audio` | string | probe audio |
|
|
| `top_k` | int, optional | max matches to return; default 5 |
|
|
| `threshold` | float, optional | cosine-distance cutoff; default 0.25 |
|
|
| `store` | string, optional | vector store model |
|
|
|
|
Returns a list of matches sorted by ascending distance, each with
|
|
`id`, `name`, `labels`, `distance`, `confidence`, and `match`
|
|
(`distance ≤ threshold`).
|
|
|
|
### `POST /v1/voice/forget`
|
|
|
|
| field | type | description |
|
|
|---|---|---|
|
|
| `id` | string | ID returned by `/v1/voice/register` |
|
|
|
|
Returns `204 No Content` on success, `404 Not Found` if the ID is
|
|
unknown.
|
|
|
|
### `POST /v1/voice/embed`
|
|
|
|
Returns the L2-normalized speaker embedding vector.
|
|
|
|
| field | type | description |
|
|
|---|---|---|
|
|
| `model` | string | voice model |
|
|
| `audio` | string | URL / base64 / data-URI |
|
|
|
|
Returns `{embedding: float[], dim: int, model: string}`. Dimension
|
|
depends on the recognizer: 192 for ECAPA-TDNN, 256 for WeSpeaker
|
|
ResNet34.
|
|
|
|
> **Note:** the OpenAI-compatible `/v1/embeddings` endpoint is
|
|
> intentionally text-only - it does nothing useful with audio input.
|
|
> Use `/v1/voice/embed` for audio.
|
|
|
|
## Audio input
|
|
|
|
Audio is materialised by the HTTP layer to a temporary WAV file
|
|
before the gRPC call. All audio fields accept:
|
|
|
|
- `http://` / `https://` URLs (downloaded server-side, subject to
|
|
`ValidateExternalURL` safety checks).
|
|
- Raw base64 (no prefix).
|
|
- Data URIs (`data:audio/wav;base64,...`).
|
|
|
|
The backend itself always receives a filesystem path - the same
|
|
convention the Whisper / Voxtral transcription backends use.
|
|
|
|
## Threshold reference
|
|
|
|
| Recognizer | Cosine-distance threshold |
|
|
|---|---|
|
|
| ECAPA-TDNN (SpeechBrain, VoxCeleb) | ~0.25 |
|
|
| WeSpeaker ResNet34 | ~0.30 |
|
|
| 3D-Speaker ERes2Net | ~0.28 |
|
|
|
|
Pass `threshold` explicitly when switching recognizers - the per-model
|
|
default only applies when omitted.
|
|
|
|
## Related features
|
|
|
|
- [Face Recognition](/features/face-recognition/) - the image analog;
|
|
the two share a registry design.
|
|
- [Audio to Text](/features/audio-to-text/) - transcription (Whisper,
|
|
Voxtral, faster-whisper). Runs in addition to, not instead of,
|
|
voice recognition.
|
|
- [Stores](/features/stores/) - the generic vector store powering
|
|
both the face and voice 1:N recognition pipelines.
|
|
- [Embeddings](/features/embeddings/) - text-only OpenAI-compatible
|
|
embedding endpoint; for audio embeddings use `/v1/voice/embed`.
|
|
|
|
## Portable profile registration
|
|
|
|
`POST /v1/voice/register` also accepts a JSON alternative to `audio`:
|
|
|
|
```javascript
|
|
// result is the parsed diarization response; slot is a selected raw speaker slot.
|
|
const request = {
|
|
model: "parakeet-diarization",
|
|
name: "Ada",
|
|
labels: {team: "research"},
|
|
speaker_slot: slot,
|
|
speaker_profiles: result.speaker_profiles
|
|
};
|
|
// POST JSON.stringify(request) with Content-Type: application/json.
|
|
```
|
|
|
|
Copy the complete `speaker_profiles` object returned by diarization unchanged.
|
|
Select `speaker_slot` explicitly, including for slot zero. It is the raw numeric
|
|
slot whose decimal string matches the diarization `label`, not a normalized
|
|
`SPEAKER_NN`, array index, or display name. `audio` and `speaker_profiles` are
|
|
mutually exclusive. `speaker_slot` without profiles is also invalid. Audio-only
|
|
registration keeps its existing JSON shape and behavior.
|
|
|
|
The server loads the requested, authorized model and obtains encoder identity and
|
|
dimension from backend metadata. It validates the complete profile export and
|
|
selects the requested usable slot. Missing slots, unavailable speech, unsupported
|
|
versions, non-finite/zero/wrong-size vectors and encoder mismatch return 400.
|
|
A backend without trusted encoder metadata returns 501. Success returns the
|
|
existing `{id, name, registered_at}` response.
|
|
|
|
Portable registrations store the **server-derived SHA-256 identity**, not a
|
|
caller-provided filename tag. Offline/live recognition admits these registrations
|
|
only when the loaded encoder has the same identity and dimension. Legacy audio
|
|
registrations retain their filename-tag compatibility rules. `/v1/voice/identify`
|
|
filters incompatible matches; a backend unable to report trusted identity cannot
|
|
match portable registrations, even when vector dimensions agree. Filtering can
|
|
return fewer than `top_k` results. The parakeet diarization model need not support
|
|
the separate audio-only VoiceEmbed RPC used by `/v1/voice/identify`.
|
|
|
|
Each successful enrollment inserts a new registration with its own ID and vector.
|
|
Duplicate display names do not merge embeddings or update an earlier enrollment.
|
|
There is no automatic enrollment or sample aggregation.
|
|
|
|
The recognition registry is **global, in-memory and per LocalAI instance**;
|
|
registrations are lost on restart and are not synchronized across frontends.
|
|
This is not durable “remembering” and not a per-user private address book. The
|
|
persistent `/api/voice-profiles` TTS-cloning feature is unrelated. Export and
|
|
registration use the existing voice-recognition permission, with existing model
|
|
access restrictions; permission does not establish biometric consent.
|
|
|
|
API tracing excludes the entire exchange for `/v1/audio/diarization`, its
|
|
`/audio/diarization` alias, and `/v1/voice/register` before capturing bodies.
|
|
This also protects JSON base64 audio when profile export is off. These routes
|
|
produce no in-memory or persisted API trace; other routes keep their existing
|
|
tracing behavior. External proxies and client logs must apply the same privacy
|
|
policy. Existing trace files from older versions are not retroactively scrubbed.
|
|
|
|
For offline and live diarization replay, registry tags never determine the
|
|
encoder dimension. LocalAI orders candidates by registration ID (tagged first),
|
|
then uses loaded encoder metadata to filter dimensions. Portable registrations
|
|
require an exact SHA-256 identity match as well. Older backends without trusted
|
|
metadata reject portable candidates and retain their native legacy dimension
|
|
checks. Identification filters compatibility after the store's `top_k` query;
|
|
incompatible results can crowd out compatible candidates within that window.
|