mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-06 04:52:31 -04:00
* feat(schema): validate portable speaker profiles Add the versioned profile schema for explicit speaker enrollment. Validate compatibility against separately supplied loaded-encoder metadata. Reject unusable speakers, invalid vectors, and inconsistent clean spans. This slice does not change HTTP routes, backend integration, or the UI. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(parakeet): export profiles with transcripts Export opt-in speaker profiles and trusted encoder metadata. Replay registrations by ID so duplicate display names keep independent vectors. Use one profile-capable diarization for slots, names, and clean spans. Assign timestamped ASR words to those slots without a second diarization. Preserve legacy opt-out and no-ASR behavior, and propagate failures. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(audio): enroll portable speaker profiles Gate profile exports with voice-recognition permission and validate registration against metadata from the loaded encoder. Preserve audio enrollment and independent registrations with duplicate display names. Exclude diarization and registration exchanges before API trace capture so persisted traces cannot retain profile vectors or JSON audio. Defer candidate dimensions to trusted loaded metadata. Sort candidates by registration ID so incompatible profiles cannot suppress legacy voices through registry iteration order. Keep portable identity checks closed when trusted metadata is unavailable. Test persisted traces, explicit slot zero, and selection through offline and live transport. Document privacy and the ephemeral registry lifecycle. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(ui): remember speakers from diarization Add a Studio page for diarization and opt-in speaker profiles. Preview clean intervals from the original recording before explicit registration. Join profiles by raw speaker labels, preserve duplicate names, and relabel turns only after a successful save. Discard stale results when the model or recording changes. Share registration metadata with voice management without storing vectors or recordings from this flow. Document permissions and the global, ephemeral registry. Cover enrollment, permissions, previews, and asynchronous races with mocked Playwright tests. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: clarify HTTP speaker enrollment support Replace the stale enrollment limitation with the current HTTP workflow. Distinguish native transport from explicit registration and link its docs. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(parakeet): pin merged speaker profile support Use the merged commit from mudler/parakeet.cpp#80. Its tree matches the previously accepted native pin. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add diarization enrollment setup example Connect the existing gallery modes to the speaker enrollment workflow. Show installation, private profile export, explicit raw-slot registration, and later recognition without another export. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): explain diarization speaker profiles Put the diarization walkthrough on the LocalAI website in the feature PR. Cover the three gallery modes, explicit enrollment, and privacy limits. Link setup instructions and keep availability conditional on feature support. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): focus diarization on everyday use Explain what users can do with recordings before the setup steps. Replace the technical walkthrough with a short Studio guide and link readers to the existing reference for model names and developer use. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): lead with speaker capabilities Present speaker recognition through everyday uses and a short UI flow. Keep technical reference details in the existing documentation. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(diarization): satisfy Go lint checks Avoid copying protobuf message state when extending backend status, check the multipart reader close result, and document the focused testing.T lint exemptions. Assisted-by: nib:gpt-5.6-sol Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
468 lines
23 KiB
Markdown
468 lines
23 KiB
Markdown
+++
|
|
disableToc = false
|
|
title = "Speaker Diarization"
|
|
weight = 33
|
|
url = "/features/audio-diarization/"
|
|
+++
|
|
|
|

|
|
|
|
Speaker diarization answers the question **"who spoke when?"** - given an audio clip with multiple speakers, it returns time-stamped segments labelled with a stable speaker ID (`SPEAKER_00`, `SPEAKER_01`, …).
|
|
|
|
LocalAI exposes this through the `/v1/audio/diarization` endpoint, modelled after `/v1/audio/transcriptions`. Five backends are supported today:
|
|
|
|
- **[sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx)** - pyannote-3.0 segmentation + a speaker-embedding extractor (3D-Speaker, NeMo, WeSpeaker) + fast clustering. Pure diarization - no transcription cost. Recommended when you only need speaker turns.
|
|
- **[vibevoice.cpp](https://github.com/microsoft/VibeVoice)** - produces speaker-labelled segments as a by-product of its long-form ASR pass, so you can optionally get a transcript per segment for free.
|
|
- **[NeMo-Speech.cpp](https://github.com/NVIDIA/NeMo-Speech.cpp)** - NVIDIA Sortformer, served standalone by the [NeMo-Speech.cpp backend]({{%relref "features/nemo-speech-cpp" %}}). It is end to end, so the speaker capacity is fixed by the checkpoint and the count hints are ignored. The same backend can instead put speaker tags on a transcript, by attaching a Sortformer model to an ASR one.
|
|
- **[audio.cpp](https://github.com/0xShug0/audio.cpp)** - the `sortformer_diar` family, served by the multi-modality [audio.cpp backend]({{%relref "features/audio-cpp" %}}).
|
|
- **[parakeet.cpp](https://github.com/mudler/parakeet.cpp)** - NVIDIA Nemotron-3-Diarization (Sortformer, up to 8 speakers), served standalone or paired with a Parakeet ASR model for per-segment text. See the [Audio to Text]({{% relref "audio-to-text" %}}) page for the parakeet-cpp option reference.
|
|
|
|
Because diarization is exposed as a regular OpenAI-compatible endpoint, any HTTP client works. There is no Python dependency on pyannote or NeMo on the consumer side.
|
|
|
|
## Endpoint
|
|
|
|
```
|
|
POST /v1/audio/diarization
|
|
Content-Type: multipart/form-data
|
|
```
|
|
|
|
| Field | Type | Description |
|
|
|-------|------|-------------|
|
|
| `file` | file (required) | audio file in any format `ffmpeg` accepts |
|
|
| `model` | string (required) | name of the diarization-capable model |
|
|
| `num_speakers` | int | exact speaker count when known (>0 forces; 0 = auto) |
|
|
| `min_speakers` | int | hint when auto-detecting |
|
|
| `max_speakers` | int | hint when auto-detecting |
|
|
| `clustering_threshold` | float | cosine distance threshold used when `num_speakers` is unknown |
|
|
| `min_duration_on` | float | discard segments shorter than this many seconds |
|
|
| `min_duration_off` | float | merge gaps shorter than this many seconds |
|
|
| `language` | string | only meaningful for backends that bundle ASR (e.g. vibevoice) |
|
|
| `include_text` | bool | when the backend can emit per-segment transcript for free, populate it |
|
|
| `response_format` | string | `json` (default), `verbose_json`, or `rttm` |
|
|
|
|
### Response - `json` (default)
|
|
|
|
Compact payload, no transcription, no per-speaker summary:
|
|
|
|
```json
|
|
{
|
|
"task": "diarize",
|
|
"duration": 12.34,
|
|
"num_speakers": 2,
|
|
"segments": [
|
|
{"id": 0, "speaker": "SPEAKER_00", "label": "0", "start": 0.00, "end": 2.34},
|
|
{"id": 1, "speaker": "SPEAKER_01", "label": "1", "start": 2.34, "end": 4.10}
|
|
]
|
|
}
|
|
```
|
|
|
|
`speaker` is the normalized, zero-padded label clients should display. `label` preserves the raw backend-emitted ID for clients that maintain their own speaker dictionary.
|
|
|
|
### Response - `verbose_json`
|
|
|
|
Adds per-speaker totals and (when the backend supports it and `include_text=true`) the per-segment transcript:
|
|
|
|
```json
|
|
{
|
|
"task": "diarize",
|
|
"duration": 12.34,
|
|
"language": "en",
|
|
"num_speakers": 2,
|
|
"segments": [
|
|
{"id": 0, "speaker": "SPEAKER_00", "label": "0", "start": 0.00, "end": 2.34, "text": "Hello, world."},
|
|
{"id": 1, "speaker": "SPEAKER_01", "label": "1", "start": 2.34, "end": 4.10, "text": "How are you?"}
|
|
],
|
|
"speakers": [
|
|
{"id": "SPEAKER_00", "label": "0", "total_speech_duration": 5.6, "segment_count": 3},
|
|
{"id": "SPEAKER_01", "label": "1", "total_speech_duration": 1.76, "segment_count": 1}
|
|
]
|
|
}
|
|
```
|
|
|
|
### Speaker names
|
|
|
|
With a parakeet-cpp model that has a `speaker_model:` and voices registered through `/v1/voice/register`, segments whose speaker matches a registered voice gain `name` and `name_score` (the cosine similarity of the match), and the matching `speakers` entry gains `name`. Both fields are omitted for a speaker that was not identified, so an unnamed response looks exactly as before. `speaker` stays `SPEAKER_NN`, and RTTM output still uses `SPEAKER_NN`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription) for the setup and the limits.
|
|
|
|
```json
|
|
{
|
|
"task": "diarize",
|
|
"duration": 12.34,
|
|
"num_speakers": 2,
|
|
"segments": [
|
|
{"id": 0, "speaker": "SPEAKER_00", "label": "0", "start": 0.00, "end": 2.34, "text": "Hello, world.", "name": "Alice", "name_score": 0.82},
|
|
{"id": 1, "speaker": "SPEAKER_01", "label": "1", "start": 2.34, "end": 4.10, "text": "How are you?"}
|
|
],
|
|
"speakers": [
|
|
{"id": "SPEAKER_00", "label": "0", "name": "Alice", "total_speech_duration": 5.6, "segment_count": 3},
|
|
{"id": "SPEAKER_01", "label": "1", "total_speech_duration": 1.76, "segment_count": 1}
|
|
]
|
|
}
|
|
```
|
|
|
|
### Response - `rttm`
|
|
|
|
NIST RTTM, the standard interchange format used by `pyannote.metrics` / `dscore`:
|
|
|
|
```
|
|
SPEAKER audio 1 0.000 2.340 <NA> <NA> SPEAKER_00 <NA> <NA>
|
|
SPEAKER audio 1 2.340 1.760 <NA> <NA> SPEAKER_01 <NA> <NA>
|
|
```
|
|
|
|
Returned as `Content-Type: text/plain; charset=utf-8`.
|
|
|
|
## Quick start
|
|
|
|
First install a diarization-capable model from the gallery. The example below uses `vibevoice-cpp-asr`, which serves the vibevoice.cpp backend and returns speaker-labelled segments (and, optionally, a transcript):
|
|
|
|
```bash
|
|
local-ai run vibevoice-cpp-asr
|
|
```
|
|
|
|
```bash
|
|
curl http://localhost:8080/v1/audio/diarization \
|
|
-H "Content-Type: multipart/form-data" \
|
|
-F file="@meeting.wav" \
|
|
-F model="vibevoice-cpp-asr" \
|
|
-F num_speakers=3
|
|
```
|
|
|
|
The sections below show how to configure the two supported backends by hand when you want full control over the segmentation and embedding models.
|
|
|
|
## Backend setup - sherpa-onnx (pure diarization)
|
|
|
|
Sherpa-onnx needs two ONNX models: pyannote segmentation and a speaker-embedding extractor. Place them under your LocalAI models directory and reference them from the YAML:
|
|
|
|
```yaml
|
|
name: pyannote-diarization
|
|
backend: sherpa-onnx
|
|
type: diarization
|
|
parameters:
|
|
model: sherpa-onnx-pyannote-segmentation-3-0/model.onnx
|
|
options:
|
|
- diarize.embedding_model=3dspeaker_speech_campplus_sv_zh-cn_16k-common.onnx
|
|
# Optional clustering knobs (per-call DiarizeRequest fields override these):
|
|
- diarize.threshold=0.5
|
|
- diarize.min_duration_on=0.3
|
|
- diarize.min_duration_off=0.5
|
|
known_usecases:
|
|
- FLAG_DIARIZATION
|
|
```
|
|
|
|
Both `model:` and `diarize.embedding_model=` are resolved relative to the LocalAI models directory.
|
|
|
|
## Backend setup - vibevoice.cpp (diarization + ASR)
|
|
|
|
vibevoice.cpp's ASR mode emits `[{Start, End, Speaker, Content}]` natively, so a single pass gives both diarization and transcription:
|
|
|
|
```yaml
|
|
name: vibevoice-diarize
|
|
backend: vibevoice-cpp
|
|
parameters:
|
|
model: vibevoice-asr.gguf
|
|
options:
|
|
- type=asr
|
|
- tokenizer=vibevoice-tokenizer.gguf
|
|
known_usecases:
|
|
- FLAG_DIARIZATION
|
|
- FLAG_TRANSCRIPT
|
|
```
|
|
|
|
Pass `include_text=true` on the request to populate the `text` field on each diarization segment.
|
|
|
|
```bash
|
|
curl http://localhost:8080/v1/audio/diarization \
|
|
-H "Content-Type: multipart/form-data" \
|
|
-F file="@interview.wav" \
|
|
-F model="vibevoice-diarize" \
|
|
-F include_text=true \
|
|
-F response_format=verbose_json
|
|
```
|
|
|
|
## Backend setup - parakeet-cpp (Nemotron-3-Diarization)
|
|
|
|
Choose an existing gallery entry for the output you need:
|
|
|
|
| Output | Gallery entry | Request options |
|
|
|---|---|---|
|
|
| Speaker turns only | `parakeet-cpp-nemotron-3-diarization` | Default options |
|
|
| Speaker turns and transcript | `parakeet-cpp-nemotron-3-diarization-asr` | `include_text=true`, `response_format=verbose_json` |
|
|
| Speaker turns, transcript, and identification | `parakeet-cpp-nemotron-3-diarization-asr-speakers` | Same transcript options; explicitly enroll voices for names |
|
|
|
|
The complete `-asr-speakers` entry downloads Nemotron-3-Diarization, Parakeet TDT+CTC 110M ASR, and the WeSpeaker ResNet34 speaker encoder.
|
|
It configures both `asr_model` and `speaker_model`; no custom gallery configuration is needed.
|
|
See [Remember speakers in the Web UI](#remember-speakers-in-the-web-ui) for installation and enrollment.
|
|
|
|
For manual configuration, this example pairs Sortformer with ASR:
|
|
|
|
```yaml
|
|
name: parakeet-diarize
|
|
backend: parakeet-cpp
|
|
parameters:
|
|
model: nemotron-3-diarization-q8_0.gguf
|
|
options:
|
|
- asr_model:tdt_ctc-110m-f16.gguf
|
|
known_usecases:
|
|
- diarization
|
|
```
|
|
|
|
Getting text on each segment needs both: an `asr_model` companion loaded on the model, and `include_text=true` on the request. With only one of the two, segments carry no text and no error is raised. Sortformer has a fixed speaker capacity and no clustering stage, so `num_speakers`, `min_speakers`, `max_speakers` and `clustering_threshold` are ignored (logged at debug); `min_duration_on` and `min_duration_off` are honored. Speaker labels are the decimal index the model assigned (`"0"`, `"1"`, …), or `"unknown"` when a segment has no diarized speaker.
|
|
|
|
```bash
|
|
curl http://localhost:8080/v1/audio/diarization \
|
|
-H "Content-Type: multipart/form-data" \
|
|
-F file="@meeting.wav" \
|
|
-F model="parakeet-diarize" \
|
|
-F include_text=true \
|
|
-F response_format=verbose_json
|
|
```
|
|
|
|
Sortformer clusters on voice-like characteristics, not on "is this a human". A loud non-speech sound with voice-like pitch and rhythm (a rooster crow, in one test clip) can come back as its own speaker segment alongside the real speakers. This is model behavior, not a bug in the LocalAI integration: treat an unexpected extra speaker as a hint the clip may contain a non-speech sound, and use [Sound Classification]({{% relref "audio-classification" %}}) to confirm what it is.
|
|
|
|
## Notes
|
|
|
|
- **Speaker identity across files**: speaker IDs (`SPEAKER_00`, `SPEAKER_01`, …) are local to each request. To track the same person across multiple recordings, combine `/v1/audio/diarization` with `/v1/voice/embed` (speaker embedding) and maintain your own embedding store.
|
|
- **Hints vs. forces**: `num_speakers` overrides clustering when set; `min_speakers` / `max_speakers` are advisory and only honored by backends that expose a range hint. vibevoice.cpp and parakeet-cpp (Sortformer) ignore them - the model picks the count itself.
|
|
- **Sample rate**: input is automatically converted to 16 kHz mono via ffmpeg before the backend sees it; sherpa-onnx pyannote-3.0 requires 16 kHz.
|
|
|
|
## See also
|
|
|
|
- [Sound Classification]({{% relref "audio-classification" %}}) - tag non-speech sound events (alarms, glass breaking, baby cry) in a clip.
|
|
|
|
### Backend profile transport
|
|
|
|
The parakeet backend supports opt-in speaker profile export through the internal
|
|
`DiarizeRequest.include_speaker_profiles` field. This native transport underpins
|
|
HTTP profile export and explicit enrollment through `POST /v1/voice/register`,
|
|
as described in [Portable speaker enrollment](#portable-speaker-enrollment) below.
|
|
It requires a configured `speaker_model` and a library
|
|
with `parakeet_capi_diarize_profiles_pcm_json`; an empty recognition registry
|
|
is supported. Export does not register anyone. With `include_text` and a loaded
|
|
ASR companion, one profile-capable diarization supplies all speaker slots,
|
|
profiles, names, and intervals. Timestamped ASR words are assigned to those
|
|
same slots; the backend does not run a second diarization. Either inference
|
|
failure fails the request. If no ASR companion is loaded, the existing fallback
|
|
applies: the response includes profiles and diarization segments without text.
|
|
A loaded ASR companion without the timestamped PCM API returns an explicit error.
|
|
|
|
`DiarizeResponse.speaker_profiles_json` carries the native version-1
|
|
`speaker_profiles` object, including original clean preview intervals and one
|
|
embedding per usable speaker. Normal requests retain their existing output.
|
|
Profile `speaker` values are raw native slot IDs. Match their decimal string to
|
|
segment `label` or speaker-summary `label`, not to normalized `SPEAKER_NN`,
|
|
array position, or display name. Slots can be sparse, and profile order can
|
|
differ from transcript order. Profiles retain their original clean intervals
|
|
even when transcript segments use word boundaries or duration filters.
|
|
These vectors are sensitive biometric data: callers must authorize export and
|
|
explicit enrollment separately.
|
|
|
|
The internal backend Status response supplies `speaker_encoder`, derived from
|
|
the loaded encoder's SHA-256 identity and dimension. Enrollment code must use
|
|
`backend.ModelSpeakerEncoder` with server-selected model configuration and
|
|
validate profiles against that result, never against caller-provided metadata.
|
|
Unavailable metadata or unsupported export fails closed. Renaming a GGUF does
|
|
not change its identity; modifying or quantizing its bytes does.
|
|
|
|
Recognition replay carries registration IDs separately from display names.
|
|
Distinct IDs with the same display name remain independent native entries,
|
|
and both offline and realtime matches are translated back to display names.
|
|
Legacy transport clients without IDs retain name-keyed behavior. The native
|
|
registry's aggregation defaults are unchanged. LocalAI's recognition registry
|
|
remains global and in-memory; this adds neither persistence nor automatic
|
|
registration and is unrelated to persistent TTS voice cloning.
|
|
|
|
## Portable speaker enrollment
|
|
|
|
Profile-capable parakeet models can export one biometric embedding per discovered
|
|
speaker, including when the recognition registry is empty. Export is opt-in:
|
|
|
|
```bash
|
|
curl http://localhost:8080/v1/audio/diarization \
|
|
-F model=parakeet-diarization -F file=@conversation.wav \
|
|
-F include_speaker_profiles=true -F include_text=true \
|
|
-F response_format=verbose_json
|
|
```
|
|
|
|
The `/audio/diarization` alias has the same protection. With user authentication,
|
|
export additionally requires the **voice-recognition** permission. Existing model
|
|
access controls still apply. Without opt-in, `speaker_profiles` is omitted.
|
|
Both `json` and `verbose_json` support profiles; `rttm` with profiles returns 400.
|
|
`include_text=true` retains supported transcripts in either JSON format.
|
|
Unsupported profile backends return 501 rather than silently omitting profiles.
|
|
|
|
Alternatively send `Content-Type: application/json`:
|
|
|
|
```json
|
|
{
|
|
"model": "parakeet-diarization",
|
|
"file": "<raw base64 audio bytes>",
|
|
"include_speaker_profiles": true,
|
|
"include_text": true,
|
|
"response_format": "verbose_json"
|
|
}
|
|
```
|
|
|
|
The `speaker_profiles` response object contains `version: 1`,
|
|
`encoder: {"identity": "sha256:<64 lowercase hex digits>", "dimension": N}`,
|
|
and `speakers`. Each speaker contains:
|
|
|
|
- `speaker`: the raw numeric speaker slot;
|
|
- `clean_duration`: retained clean speech in seconds;
|
|
- `intervals`: `{start, end}` ranges in seconds in the original recording;
|
|
- `unavailable_reason`: null for usable profiles, otherwise a reason string;
|
|
- `embedding`: one vector for a usable speaker, omitted when unavailable.
|
|
|
|
**UI association:** convert each profile's numeric `speaker` to a decimal string
|
|
and match segment/summary `label`. Do not use `SPEAKER_NN`, array position, or
|
|
human name. Slots may be sparse and out of order; display names may repeat.
|
|
Preview `intervals` against the original audio, not separated audio. Disable
|
|
saving unavailable profiles. Enrollment is explicit, never automatic; only
|
|
relabel after a successful registration response. See
|
|
[portable voice registration](/features/voice-recognition/#portable-profile-registration).
|
|
|
|
Profiles are sensitive, unsigned biometric data, not proof of identity or consent.
|
|
Do not log their vectors. Obtain the speaker's consent before enrollment.
|
|
|
|
API tracing excludes the entire exchange for `/v1/audio/diarization`, its
|
|
`/audio/diarization` alias, and `/v1/voice/register` before capturing bodies.
|
|
This also protects JSON base64 audio when profile export is off. These routes
|
|
produce no in-memory or persisted API trace; other routes keep their existing
|
|
tracing behavior. External proxies and client logs must apply the same privacy
|
|
policy. Existing trace files from older versions are not retroactively scrubbed.
|
|
|
|
## Remember speakers in the Web UI
|
|
|
|
Use a LocalAI build with portable enrollment support and a profile-capable `parakeet-cpp` backend.
|
|
The backend needs the profile APIs from merged upstream commit
|
|
[`bee7c14`](https://github.com/mudler/parakeet.cpp/commit/bee7c14dfcc23613df58176c59a40459e7b47095) or a compatible later build.
|
|
Installing the model weights alone does not update an older backend.
|
|
|
|
1. Open **Models → Explore** and search for `parakeet-cpp-nemotron-3-diarization-asr-speakers`.
|
|
2. Select **Install** and wait for installation to complete. Check **Operate → Activity** for progress or errors.
|
|
3. Open **Studio → Diarization** (or `/app/diarization`). Select that model and upload your recording.
|
|
|
|
Obtain the speaker's consent before enrollment. To remember a speaker from that recording:
|
|
|
|
1. Select **Prepare speakers to remember**, then select **Diarize**. This
|
|
requests profiles, transcript text, and speaker summaries. Use a
|
|
profile-capable parakeet-cpp model configured with a speaker encoder.
|
|
2. In **Speakers**, select **Preview 1**, **Preview 2**, or another available
|
|
interval to listen to clean speech from the original recording. Playback
|
|
stops at the end of that interval. **Stop preview** stops it earlier.
|
|
Your browser must support the recording's audio format.
|
|
3. For an unknown speaker, select **Name and remember**. Enter a name and
|
|
select **Remember**. No second recording or audio upload is needed.
|
|
4. After the server confirms registration, the name appears on all turns for
|
|
that speaker. A failed save keeps the entered name so you can retry.
|
|
|
|
Upload another recording and select **Diarize** to match remembered voices.
|
|
You can turn off **Prepare speakers to remember**; recognition does not require another profile export.
|
|
Matches show their names; unmatched speakers keep their speaker labels.
|
|
With preparation off, the UI requests speaker turns without transcript text.
|
|
Use the API example below to request text without exporting profiles.
|
|
|
|
Speakers without a usable profile cannot be
|
|
remembered; try longer speech without overlapping speakers. Duplicate names
|
|
are allowed: each save creates a separate registration, not a merged voice.
|
|
Changing the model or recording clears the current results and save dialog.
|
|
A save already sent to the server can still complete, but cannot rename turns
|
|
in a different recording.
|
|
|
|
The page requires the **Audio Diarization** permission and access to the selected model. Preparing profiles and
|
|
remembering speakers additionally require **Voice Recognition**. Users without
|
|
that permission can still run normal diarization. If the backend does not
|
|
support profiles, the page reports an error: choose a compatible model or
|
|
turn off **Prepare speakers to remember**. It does not silently retry without
|
|
profiles.
|
|
|
|
{{% notice warning %}}
|
|
Remembered voices are shared globally on this server and are lost when it
|
|
restarts. Nothing is enrolled automatically. The browser stores only the new
|
|
registration's ID, name, and registration time for the existing voice
|
|
management list, not its embedding or recording. That list is local to the
|
|
browser and is not a durable server registry.
|
|
{{% /notice %}}
|
|
|
|
Use **Manage remembered voices**, then the **Enrollment** tab, to see or
|
|
remove registrations saved in this browser. Clean-clip voice enrollment stays
|
|
available there and does not require diarization.
|
|
|
|
### API example: install, export, and remember
|
|
|
|
This example uses the same complete gallery entry and requires `curl` and `jq`.
|
|
The commands assume a local server without authentication.
|
|
If authentication is enabled, add `-H "Authorization: Bearer <key>"` to every request using your authorized key.
|
|
Keep keys out of shared scripts, logs, and shell history; see [Authentication]({{% relref "authentication" %}}).
|
|
Installation requires model-management access; inference and enrollment require the permissions described above.
|
|
|
|
Install the model if it is not already installed:
|
|
|
|
```bash
|
|
LOCALAI=http://localhost:8080
|
|
MODEL=parakeet-cpp-nemotron-3-diarization-asr-speakers
|
|
curl --fail-with-body "$LOCALAI/models/apply" \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"id":"localai@parakeet-cpp-nemotron-3-diarization-asr-speakers"}'
|
|
```
|
|
|
|
Installation is asynchronous. Wait for successful completion in **Operate → Activity** before continuing.
|
|
API clients can query the returned job `status` URL; see the [model gallery API]({{% relref "model-gallery" %}}).
|
|
|
|
{{% notice warning %}}
|
|
Exported profiles contain biometric vectors. Obtain consent before enrollment.
|
|
Keep the recording, response, and registration files private. Do not log or share their contents.
|
|
Use a new private directory so existing files cannot retain broader permissions. Delete these files when no longer needed.
|
|
{{% /notice %}}
|
|
|
|
Export profiles and transcript text from your recording, keeping the complete JSON response:
|
|
|
|
```bash
|
|
umask 077
|
|
WORK=$(mktemp -d)
|
|
curl --fail-with-body "$LOCALAI/v1/audio/diarization" \
|
|
-F "model=$MODEL" -F file=@conversation.wav \
|
|
-F include_text=true -F include_speaker_profiles=true \
|
|
-F response_format=verbose_json > "$WORK/diarization.json"
|
|
|
|
# Inspect raw slots, clean intervals, and transcript labels without printing vectors.
|
|
jq '.speaker_profiles.speakers[] | {speaker, clean_duration, intervals, unavailable_reason}' \
|
|
"$WORK/diarization.json"
|
|
jq '.segments[] | {label, start, end, text}' "$WORK/diarization.json"
|
|
```
|
|
|
|
Choose a usable raw `speaker` slot whose decimal string matches the intended segment `label`.
|
|
Listen to its `intervals` in the original recording before assigning a name.
|
|
Do not select by array position, normalized `SPEAKER_NN`, or display name.
|
|
If `unavailable_reason` indicates insufficient speech, try a longer recording without overlapping speakers.
|
|
|
|
Replace `0` below with your chosen raw slot. Zero is valid, but does not mean “the first array element.”
|
|
Keep the complete `speaker_profiles` object unchanged:
|
|
|
|
```bash
|
|
SLOT=0
|
|
NAME=Ada
|
|
jq --arg model "$MODEL" --arg name "$NAME" --argjson slot "$SLOT" \
|
|
'{model: $model, name: $name, speaker_slot: $slot, speaker_profiles: .speaker_profiles}' \
|
|
"$WORK/diarization.json" > "$WORK/register.json"
|
|
curl --fail-with-body "$LOCALAI/v1/voice/register" \
|
|
-H 'Content-Type: application/json' \
|
|
--data-binary @"$WORK/register.json"
|
|
```
|
|
|
|
After successful registration, submit another recording with the same model:
|
|
|
|
```bash
|
|
curl --fail-with-body "$LOCALAI/v1/audio/diarization" \
|
|
-F "model=$MODEL" -F file=@next-conversation.wav \
|
|
-F include_text=true -F response_format=verbose_json > "$WORK/next.json"
|
|
jq '.segments[] | {label, name, start, end, text}' "$WORK/next.json"
|
|
|
|
# Remove private example outputs when no longer needed.
|
|
rm -f "$WORK/diarization.json" "$WORK/register.json" "$WORK/next.json"
|
|
rmdir "$WORK"
|
|
```
|
|
|
|
Matching speakers can now carry `name`, even though this request omits `include_speaker_profiles`.
|
|
Keep `include_text=true` and `verbose_json` when you want transcript text.
|
|
Recognition is not proof of identity. Registrations remain global and disappear on server restart.
|
|
See [portable profile registration](/features/voice-recognition/#portable-profile-registration) for encoder compatibility and validation rules.
|