chore: merge master into distributed test branch

Preserve multipart language hints alongside JSON diarization and speaker
profiles when resolving the endpoint conflict.

Assisted-by: Codex:gpt-6
This commit is contained in:
localai-org-maint-bot committed 2026-10-02 08:05:12 +00:00
commit ba75324ae4
104 files changed
+5603 -291

No files matched your search

+271 -1
View File
@@ -79,6 +79,26 @@ Adds per-speaker totals and (when the backend supports it and `include_text=true
}
```
### Speaker names
With a parakeet-cpp model that has a `speaker_model:` and voices registered through `/v1/voice/register`, segments whose speaker matches a registered voice gain `name` and `name_score` (the cosine similarity of the match), and the matching `speakers` entry gains `name`. Both fields are omitted for a speaker that was not identified, so an unnamed response looks exactly as before. `speaker` stays `SPEAKER_NN`, and RTTM output still uses `SPEAKER_NN`. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription) for the setup and the limits.
```json
{
"task": "diarize",
"duration": 12.34,
"num_speakers": 2,
"segments": [
{"id": 0, "speaker": "SPEAKER_00", "label": "0", "start": 0.00, "end": 2.34, "text": "Hello, world.", "name": "Alice", "name_score": 0.82},
{"id": 1, "speaker": "SPEAKER_01", "label": "1", "start": 2.34, "end": 4.10, "text": "How are you?"}
],
"speakers": [
{"id": "SPEAKER_00", "label": "0", "name": "Alice", "total_speech_duration": 5.6, "segment_count": 3},
{"id": "SPEAKER_01", "label": "1", "total_speech_duration": 1.76, "segment_count": 1}
]
}
```
### Response - `rttm`
NIST RTTM, the standard interchange format used by `pyannote.metrics` / `dscore`:
@@ -160,7 +180,19 @@ curl http://localhost:8080/v1/audio/diarization \
## Backend setup - parakeet-cpp (Nemotron-3-Diarization)
Nemotron-3-Diarization is Sortformer, served standalone or paired with a Parakeet ASR model. Install `parakeet-cpp-nemotron-3-diarization` from the gallery for diarization only, or `parakeet-cpp-nemotron-3-diarization-asr` for the same model paired with `parakeet-cpp-tdt_ctc-110m` through the `asr_model` option:
Choose an existing gallery entry for the output you need:
| Output | Gallery entry | Request options |
|---|---|---|
| Speaker turns only | `parakeet-cpp-nemotron-3-diarization` | Default options |
| Speaker turns and transcript | `parakeet-cpp-nemotron-3-diarization-asr` | `include_text=true`, `response_format=verbose_json` |
| Speaker turns, transcript, and identification | `parakeet-cpp-nemotron-3-diarization-asr-speakers` | Same transcript options; explicitly enroll voices for names |
The complete `-asr-speakers` entry downloads Nemotron-3-Diarization, Parakeet TDT+CTC 110M ASR, and the WeSpeaker ResNet34 speaker encoder.
It configures both `asr_model` and `speaker_model`; no custom gallery configuration is needed.
See [Remember speakers in the Web UI](#remember-speakers-in-the-web-ui) for installation and enrollment.
For manual configuration, this example pairs Sortformer with ASR:
```yaml
name: parakeet-diarize
@@ -195,3 +227,241 @@ Sortformer clusters on voice-like characteristics, not on "is this a human". A l
## See also
- [Sound Classification]({{% relref "audio-classification" %}}) - tag non-speech sound events (alarms, glass breaking, baby cry) in a clip.
### Backend profile transport
The parakeet backend supports opt-in speaker profile export through the internal
`DiarizeRequest.include_speaker_profiles` field. This native transport underpins
HTTP profile export and explicit enrollment through `POST /v1/voice/register`,
as described in [Portable speaker enrollment](#portable-speaker-enrollment) below.
It requires a configured `speaker_model` and a library
with `parakeet_capi_diarize_profiles_pcm_json`; an empty recognition registry
is supported. Export does not register anyone. With `include_text` and a loaded
ASR companion, one profile-capable diarization supplies all speaker slots,
profiles, names, and intervals. Timestamped ASR words are assigned to those
same slots; the backend does not run a second diarization. Either inference
failure fails the request. If no ASR companion is loaded, the existing fallback
applies: the response includes profiles and diarization segments without text.
A loaded ASR companion without the timestamped PCM API returns an explicit error.
`DiarizeResponse.speaker_profiles_json` carries the native version-1
`speaker_profiles` object, including original clean preview intervals and one
embedding per usable speaker. Normal requests retain their existing output.
Profile `speaker` values are raw native slot IDs. Match their decimal string to
segment `label` or speaker-summary `label`, not to normalized `SPEAKER_NN`,
array position, or display name. Slots can be sparse, and profile order can
differ from transcript order. Profiles retain their original clean intervals
even when transcript segments use word boundaries or duration filters.
These vectors are sensitive biometric data: callers must authorize export and
explicit enrollment separately.
The internal backend Status response supplies `speaker_encoder`, derived from
the loaded encoder's SHA-256 identity and dimension. Enrollment code must use
`backend.ModelSpeakerEncoder` with server-selected model configuration and
validate profiles against that result, never against caller-provided metadata.
Unavailable metadata or unsupported export fails closed. Renaming a GGUF does
not change its identity; modifying or quantizing its bytes does.
Recognition replay carries registration IDs separately from display names.
Distinct IDs with the same display name remain independent native entries,
and both offline and realtime matches are translated back to display names.
Legacy transport clients without IDs retain name-keyed behavior. The native
registry's aggregation defaults are unchanged. LocalAI's recognition registry
remains global and in-memory; this adds neither persistence nor automatic
registration and is unrelated to persistent TTS voice cloning.
## Portable speaker enrollment
Profile-capable parakeet models can export one biometric embedding per discovered
speaker, including when the recognition registry is empty. Export is opt-in:
```bash
curl http://localhost:8080/v1/audio/diarization \
-F model=parakeet-diarization -F file=@conversation.wav \
-F include_speaker_profiles=true -F include_text=true \
-F response_format=verbose_json
```
The `/audio/diarization` alias has the same protection. With user authentication,
export additionally requires the **voice-recognition** permission. Existing model
access controls still apply. Without opt-in, `speaker_profiles` is omitted.
Both `json` and `verbose_json` support profiles; `rttm` with profiles returns 400.
`include_text=true` retains supported transcripts in either JSON format.
Unsupported profile backends return 501 rather than silently omitting profiles.
Alternatively send `Content-Type: application/json`:
```json
{
"model": "parakeet-diarization",
"file": "<raw base64 audio bytes>",
"include_speaker_profiles": true,
"include_text": true,
"response_format": "verbose_json"
}
```
The `speaker_profiles` response object contains `version: 1`,
`encoder: {"identity": "sha256:<64 lowercase hex digits>", "dimension": N}`,
and `speakers`. Each speaker contains:
- `speaker`: the raw numeric speaker slot;
- `clean_duration`: retained clean speech in seconds;
- `intervals`: `{start, end}` ranges in seconds in the original recording;
- `unavailable_reason`: null for usable profiles, otherwise a reason string;
- `embedding`: one vector for a usable speaker, omitted when unavailable.
**UI association:** convert each profile's numeric `speaker` to a decimal string
and match segment/summary `label`. Do not use `SPEAKER_NN`, array position, or
human name. Slots may be sparse and out of order; display names may repeat.
Preview `intervals` against the original audio, not separated audio. Disable
saving unavailable profiles. Enrollment is explicit, never automatic; only
relabel after a successful registration response. See
[portable voice registration](/features/voice-recognition/#portable-profile-registration).
Profiles are sensitive, unsigned biometric data, not proof of identity or consent.
Do not log their vectors. Obtain the speaker's consent before enrollment.
API tracing excludes the entire exchange for `/v1/audio/diarization`, its
`/audio/diarization` alias, and `/v1/voice/register` before capturing bodies.
This also protects JSON base64 audio when profile export is off. These routes
produce no in-memory or persisted API trace; other routes keep their existing
tracing behavior. External proxies and client logs must apply the same privacy
policy. Existing trace files from older versions are not retroactively scrubbed.
## Remember speakers in the Web UI
Use a LocalAI build with portable enrollment support and a profile-capable `parakeet-cpp` backend.
The backend needs the profile APIs from merged upstream commit
[`bee7c14`](https://github.com/mudler/parakeet.cpp/commit/bee7c14dfcc23613df58176c59a40459e7b47095) or a compatible later build.
Installing the model weights alone does not update an older backend.
1. Open **Models → Explore** and search for `parakeet-cpp-nemotron-3-diarization-asr-speakers`.
2. Select **Install** and wait for installation to complete. Check **Operate → Activity** for progress or errors.
3. Open **Studio → Diarization** (or `/app/diarization`). Select that model and upload your recording.
Obtain the speaker's consent before enrollment. To remember a speaker from that recording:
1. Select **Prepare speakers to remember**, then select **Diarize**. This
requests profiles, transcript text, and speaker summaries. Use a
profile-capable parakeet-cpp model configured with a speaker encoder.
2. In **Speakers**, select **Preview 1**, **Preview 2**, or another available
interval to listen to clean speech from the original recording. Playback
stops at the end of that interval. **Stop preview** stops it earlier.
Your browser must support the recording's audio format.
3. For an unknown speaker, select **Name and remember**. Enter a name and
select **Remember**. No second recording or audio upload is needed.
4. After the server confirms registration, the name appears on all turns for
that speaker. A failed save keeps the entered name so you can retry.
Upload another recording and select **Diarize** to match remembered voices.
You can turn off **Prepare speakers to remember**; recognition does not require another profile export.
Matches show their names; unmatched speakers keep their speaker labels.
With preparation off, the UI requests speaker turns without transcript text.
Use the API example below to request text without exporting profiles.
Speakers without a usable profile cannot be
remembered; try longer speech without overlapping speakers. Duplicate names
are allowed: each save creates a separate registration, not a merged voice.
Changing the model or recording clears the current results and save dialog.
A save already sent to the server can still complete, but cannot rename turns
in a different recording.
The page requires the **Audio Diarization** permission and access to the selected model. Preparing profiles and
remembering speakers additionally require **Voice Recognition**. Users without
that permission can still run normal diarization. If the backend does not
support profiles, the page reports an error: choose a compatible model or
turn off **Prepare speakers to remember**. It does not silently retry without
profiles.
{{% notice warning %}}
Remembered voices are shared globally on this server and are lost when it
restarts. Nothing is enrolled automatically. The browser stores only the new
registration's ID, name, and registration time for the existing voice
management list, not its embedding or recording. That list is local to the
browser and is not a durable server registry.
{{% /notice %}}
Use **Manage remembered voices**, then the **Enrollment** tab, to see or
remove registrations saved in this browser. Clean-clip voice enrollment stays
available there and does not require diarization.
### API example: install, export, and remember
This example uses the same complete gallery entry and requires `curl` and `jq`.
The commands assume a local server without authentication.
If authentication is enabled, add `-H "Authorization: Bearer <key>"` to every request using your authorized key.
Keep keys out of shared scripts, logs, and shell history; see [Authentication]({{% relref "authentication" %}}).
Installation requires model-management access; inference and enrollment require the permissions described above.
Install the model if it is not already installed:
```bash
LOCALAI=http://localhost:8080
MODEL=parakeet-cpp-nemotron-3-diarization-asr-speakers
curl --fail-with-body "$LOCALAI/models/apply" \
-H 'Content-Type: application/json' \
-d '{"id":"localai@parakeet-cpp-nemotron-3-diarization-asr-speakers"}'
```
Installation is asynchronous. Wait for successful completion in **Operate → Activity** before continuing.
API clients can query the returned job `status` URL; see the [model gallery API]({{% relref "model-gallery" %}}).
{{% notice warning %}}
Exported profiles contain biometric vectors. Obtain consent before enrollment.
Keep the recording, response, and registration files private. Do not log or share their contents.
Use a new private directory so existing files cannot retain broader permissions. Delete these files when no longer needed.
{{% /notice %}}
Export profiles and transcript text from your recording, keeping the complete JSON response:
```bash
umask 077
WORK=$(mktemp -d)
curl --fail-with-body "$LOCALAI/v1/audio/diarization" \
-F "model=$MODEL" -F file=@conversation.wav \
-F include_text=true -F include_speaker_profiles=true \
-F response_format=verbose_json > "$WORK/diarization.json"
# Inspect raw slots, clean intervals, and transcript labels without printing vectors.
jq '.speaker_profiles.speakers[] | {speaker, clean_duration, intervals, unavailable_reason}' \
"$WORK/diarization.json"
jq '.segments[] | {label, start, end, text}' "$WORK/diarization.json"
```
Choose a usable raw `speaker` slot whose decimal string matches the intended segment `label`.
Listen to its `intervals` in the original recording before assigning a name.
Do not select by array position, normalized `SPEAKER_NN`, or display name.
If `unavailable_reason` indicates insufficient speech, try a longer recording without overlapping speakers.
Replace `0` below with your chosen raw slot. Zero is valid, but does not mean “the first array element.”
Keep the complete `speaker_profiles` object unchanged:
```bash
SLOT=0
NAME=Ada
jq --arg model "$MODEL" --arg name "$NAME" --argjson slot "$SLOT" \
'{model: $model, name: $name, speaker_slot: $slot, speaker_profiles: .speaker_profiles}' \
"$WORK/diarization.json" > "$WORK/register.json"
curl --fail-with-body "$LOCALAI/v1/voice/register" \
-H 'Content-Type: application/json' \
--data-binary @"$WORK/register.json"
```
After successful registration, submit another recording with the same model:
```bash
curl --fail-with-body "$LOCALAI/v1/audio/diarization" \
-F "model=$MODEL" -F file=@next-conversation.wav \
-F include_text=true -F response_format=verbose_json > "$WORK/next.json"
jq '.segments[] | {label, name, start, end, text}' "$WORK/next.json"
# Remove private example outputs when no longer needed.
rm -f "$WORK/diarization.json" "$WORK/register.json" "$WORK/next.json"
rmdir "$WORK"
```
Matching speakers can now carry `name`, even though this request omits `include_speaker_profiles`.
Keep `include_text=true` and `verbose_json` when you want transcript text.
Recognition is not proof of identity. Registrations remain global and disappear on server restart.
See [portable profile registration](/features/voice-recognition/#portable-profile-registration) for encoder compatibility and validation rules.
+7
View File
@@ -112,6 +112,8 @@ In addition to `file` and `model`, the endpoint accepts the following multipart
| `stream` | When `true`, the endpoint emits an SSE stream of `transcript.text.delta` events followed by a final `transcript.text.done` event. |
| `diarize` | LocalAI extension - speaker diarization. WhisperX requires `HF_TOKEN`; requests fail with `FailedPrecondition` when it is missing. |
If speaker diarization fails after transcription succeeded, the WhisperX backend logs the error and returns the transcript without speaker labels. Other transcription failures return an error instead of an empty transcript. Diarization still requires `HF_TOKEN`.
The response body for `verbose_json` includes `text`, `language`, `duration`, and `segments[]` (with `speaker` populated when diarization is enabled).
## Streaming transcriptions
@@ -200,9 +202,14 @@ The same backend also serves the `/v1/audio/diarization` and `/v1/audio/classifi
| `diarization_model:<path>` | an ASR model | a `speaker` on transcript segments (and words), and speaker segments during realtime live transcription |
| `sound_model:<path>` | an ASR model | sound events during realtime live transcription |
| `diarization_latency:<model\|low\|very_low\|ultra_low>` | a model with a diarization companion | latency mode for the live speaker stream; default `low` |
| `speaker_model:<path>` | a model with a diarization model | names registered speakers (see [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription)) |
| `speaker_threshold:<float>` | a model with `speaker_model` | distance (1 minus cosine similarity) under which a speaker is named, in (0, 2); default `0.5` |
| `speaker_margin:<float>` | a model with `speaker_model` | how much the best match must beat the runner-up, in [0, 1); default `0.05` |
With a `diarization_model` companion, `/v1/audio/transcriptions` labels each segment with its `speaker` (`"0"`, `"1"`, ... in order of first appearance) and splits segments where the speaker changes; with `timestamp_granularities[]=word` each word carries its speaker too. With `stream=true` the closing `transcript.text.done` event lists the segments with their speakers. Pass `-F diarize=false` to skip diarization for one request. The diarization GGUF can also be imported directly: `local-ai models import https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/nemotron-3-diarization-f16.gguf`.
`speaker_model:` needs libparakeet with C-API v10. A wrong setup fails at load time with one of these errors: `parakeet-cpp: speaker_model needs libparakeet.so ABI 10 (parakeet_capi_speaker_registry_add_embedding); the loaded library is older`, `parakeet-cpp: speaker_model needs a diarization model (the primary or diarization_model:)`, `parakeet-cpp: a speaker model cannot be the primary model; use it as speaker_model: next to a diarization model`, `parakeet-cpp: speaker_model "<path>" is a <kind> model, expected a speaker model` (the file is not a speaker encoder GGUF), or `parakeet-cpp: speaker_threshold "<value>" must be a distance in (0, 2) (1 minus cosine similarity)` / `parakeet-cpp: speaker_margin "<value>" must be a number in [0, 1)` for a bad number.
The loader rejects a companion whose role duplicates the primary's own (for example `asr_model:` on an already-ASR primary, or `sound_model:` on a CED primary), and rejects a companion GGUF that does not match the role its option names (for example `sound_model:` pointing at an ASR GGUF fails to load, naming the kind it expected). See [Speaker Diarization]({{% relref "audio-diarization" %}}) for the `Diarize` RPC and [Sound Classification]({{% relref "audio-classification" %}}) for `SoundDetection`, and [Realtime API]({{% relref "openai-realtime" %}}) for the live speaker/sound events emitted during a realtime session.
### Segment timestamps
+11 -3
View File
@@ -108,11 +108,19 @@ Install one from the gallery and filter on the `decisions` tag:
| `tev1-4b-vllm-cpp` | Tev1 4B | Autoregressive Qwen3.5-4B fine-tune that answers with an option letter, about 9.3 GB |
| `tev1-0.8b-vllm-cpp` | Tev1 0.8B | Autoregressive Qwen3.5-0.8B fine-tune that answers with an option letter, about 1.8 GB |
| `kev-0.8b-vllm-cpp` | kev 0.8B | Qwen3.5-0.8B-Base with a merged LoRA and a PointerHead readout, converted for vllm.cpp only, about 1.53 GB |
| `nimble-9b-vllm-cpp` | Bespoke Nimble 9B | Qwen3.5-9B with the Nimble LoRA merged, reads the answer-letter logits, converted for vllm.cpp only, about 19.3 GB |
| `clm-v0.1-8b-vllm-cpp` | CLM v0.1 8B | Bi-encoder: Qwen3-8B backbone with state and action heads, answers by cosine similarity, converted for vllm.cpp only, about 16.5 GB |
The engine, [vllm.cpp]({{% relref "features/vllm-cpp" %}}), also supports the
CLM and xor decision models. Those checkpoints need a conversion step, so
they are not gallery entries yet. The kev entry installs a checkpoint that was
already converted with the vllm.cpp `convert-kev.py` script.
xor decision model. That checkpoint needs a conversion step, so it is not a
gallery entry yet. The kev, Nimble and CLM entries install checkpoints that
were already converted with the vllm.cpp `convert-kev.py`, `convert-nimble.py`
and `convert-clm.py` scripts.
Nimble refuses a question with more than 26 choices (the upstream release
allows 255). On CPU it needs about 20 GB of free RAM. For CLM, put the
question in `instructions`: the state head reads the state followed by the
instructions. On CPU it needs about 19 GB of free RAM.
Tev1 is an autoregressive decision model. The engine answers each question by
scoring the option letters, so its `confidence` is the entropy measure Ollama
+21
View File
@@ -39,6 +39,27 @@ Both views use the same model selection and store the view, search, filter, and
selection in the URL. Installing from Explore does not move you away from the
catalog; the entry updates in place when the operation finishes.
## Cyber-Ornith 1.5 9B
Cyber-Ornith 1.5 is a Qwen3.5 fine-tune for security auditing, terminal tasks, and tool use.
The gallery includes Q4_K_M and Q6_K GGUF builds for text chat with llama.cpp.
Both use the embedded chat template and a 32,768-token context by default.
Install with automatic variant selection:
```bash
local-ai models install cyber-ornith-1.5-9b-obliterated
```
To select Q6_K explicitly:
```bash
local-ai models install cyber-ornith-1.5-9b-obliterated --variant cyber-ornith-1.5-9b-obliterated-q6
```
See the [model card](https://huggingface.co/DuoNeural/Cyber-Ornith-1.5-9B-OBLITERATED)
and [GGUF downloads](https://huggingface.co/mradermacher/Cyber-Ornith-1.5-9B-OBLITERATED-i1-GGUF).
## Cyber-Tiel-Coder
Install `cyber-tiel-coder-35b-a3b-q4-mtp` for coding and image chat with llama.cpp.
+3
View File
@@ -166,12 +166,15 @@ Each closed speaker segment emits a `conversation.item.input_audio_transcription
"item_id": "item_abc",
"content_index": 0,
"speaker": "0",
"speaker_name": "Alice",
"start": 1.92,
"end": 4.10,
"text": ""
}
```
`speaker_name` is the name of a voice registered through `/v1/voice/register`, and is present only when the model has a `speaker_model:` and the speaker was identified. A segment that closes before its speaker is identified has none, and later segments of the same speaker do. See [Voice Recognition]({{% relref "voice-recognition" %}}#naming-speakers-in-diarization-and-live-transcription). The segments of the offline path below carry no `speaker_name`.
Each sound event emits a `conversation.item.sound_detection` event with one tag and the detection window's `start`/`end`:
```json
+4
View File
@@ -33,6 +33,8 @@ Available additional parameters: `top_p`, `top_k`, `max_tokens`
Reasoning models return their thinking in the `reasoning` field. When a model reasons and calls a tool in the same turn, see [Interleaved Thinking with Tool Calls]({{%relref "features/interleaved-thinking" %}}).
When `stream: true` is set and the llama.cpp backend fails before the first chunk, for example because the prompt exceeds the context size, the request fails with an HTTP error. The error message is not streamed as assistant content. An error after streaming has started is reported inside the stream.
### Edit completions
https://platform.openai.com/docs/api-reference/edits
@@ -627,6 +629,8 @@ options:
**Note:** The `parallel` option can also be set via the `LLAMACPP_PARALLEL` environment variable, and `grpc_servers` can be set via the `LLAMACPP_GRPC_SERVERS` environment variable. Options specified in the YAML file take precedence over environment variables.
An explicit `parallel: 1` (or `n_parallel: 1`) in the model options takes precedence over `LLAMACPP_PARALLEL`, like any other value. The environment variable is only used when neither option is set; if it is missing or not a number, the backend uses one slot.
##### Hardware auto-tuning (and how to override it)
On a detected GPU, LocalAI fills a few performance-relevant defaults the model config leaves unset - a larger physical batch on NVIDIA Blackwell, and a VRAM-scaled `parallel` slot count for concurrent serving. Both are gated on **per-device** VRAM at the model's context: when a large context already fills a single card (e.g. a 27B model with a 200k context across 2×16 GiB), the batch boost and the extra parallel slots are suppressed so they can't tip the tighter GPU into CUDA out-of-memory.
+144
View File
@@ -185,6 +185,85 @@ recognition - the voice-recognition HTTP API is designed to swap the
backing store without changing the wire format.
{{% /notice %}}
## Naming speakers in diarization and live transcription
The parakeet-cpp backend can put the names of registered voices on
diarization results and on live transcription speaker segments. Without
this, speakers only carry labels such as `SPEAKER_00`.
1. Register each voice with the WeSpeaker encoder. Install the model with
`local-ai models install voice-detect-wespeaker-resnet34`, then call
`/v1/voice/register` with `"model": "voice-detect-wespeaker-resnet34"`
(see the [1:N workflow](#1n-identification-workflow-register--identify--forget)).
2. Install one of the gallery models that loads the same encoder:
`parakeet-cpp-nemotron-3-diarization-speakers` (diarization),
`parakeet-cpp-nemotron-3-diarization-asr-speakers` (diarization with
`include_text`) or `parakeet-cpp-realtime-scene-speakers` (live
transcription). Each one adds
`speaker_model:voice-detect-wespeaker-resnet34.gguf` to a
parakeet-cpp model config.
3. Call `/v1/audio/diarization` with that model. Matched segments gain a
`name` and a `name_score`, and the matching entry in `speakers` gains a
`name`. `speaker` stays `SPEAKER_NN`, and RTTM output is unchanged. See
[Speaker Diarization]({{% relref "audio-diarization" %}}) for the
response.
### Which voices are used
LocalAI sends the backend only the registered voices made by the same
encoder as the model's `speaker_model:` file. Each registered voice is
tagged with the name of the voice-detect model that made it, which by
default is the GGUF file name (`voice-detect-wespeaker-resnet34.gguf` for the
gallery entry). The tag must equal the base name of the `speaker_model:`
file. Voices made with another encoder are ignored, and LocalAI logs a
warning when that leaves no usable voice. Voices registered before the tag
existed have no tag: they are used when their embedding size matches the
tagged ones (or all of them, when no voice carries a matching tag). The
backend skips a voice whose embedding size does not match the speaker model's,
with a warning in the LocalAI log. Naming then falls back to the remaining
voices, or to no names.
{{% notice warning %}}
Do not set a `model_name:` option on the voice-detect model config. It
replaces the default name, the voices are then tagged with it, and they no
longer match the `speaker_model:` file. Keep the default name.
{{% /notice %}}
### Options
These go in the `options:` list of the parakeet-cpp model config (see
[Audio to Text]({{% relref "audio-to-text" %}}) for the other parakeet-cpp
options).
| Option | Default | Meaning |
|---|---|---|
| `speaker_model:<path>` | none | speaker encoder GGUF; needs a diarization model (the primary one, or `diarization_model:`) |
| `speaker_threshold:<float>` | `0.5` | largest distance (1 minus cosine similarity, the unit `/v1/voice/identify` reports) at which a speaker is named; must be in (0, 2) |
| `speaker_margin:<float>` | `0.05` | the best match must beat the runner-up by this much, otherwise the speaker stays unnamed; must be in [0, 1) |
parakeet.cpp's measured starting values for `speaker_threshold` are 0.5 for
WeSpeaker ResNet34 and CAM++, and 0.3 for ECAPA. A lower value names fewer
speakers and makes fewer mistakes.
### Limits
- The voice registry is in memory and global. Registered names disappear when
LocalAI restarts, and every user of the instance shares them.
- Anyone who is allowed to call a model with `speaker_model:` can learn which
registered names match their audio, and their audio is matched against voices
registered by any user, because the voice registry is global. Restrict such
models with the per-user model allowlist.
- With `include_text=true` the names use the default threshold and margin:
`speaker_threshold` and `speaker_margin` only apply to diarization without
text.
- In live transcription, a speaker segment that closes before its speaker
is identified has no name. Later segments of that speaker do.
- Overlapping speech is not resolved.
- Accuracy was measured on one fixture (two read-speech voices). Check the
threshold on your own audio.
- The backend needs a libparakeet with C-API v10. With an older library a
model config that sets `speaker_model:` fails to load.
## API reference
### `POST /v1/voice/verify` (1:1)
@@ -323,3 +402,68 @@ default only applies when omitted.
both the face and voice 1:N recognition pipelines.
- [Embeddings](/features/embeddings/) - text-only OpenAI-compatible
embedding endpoint; for audio embeddings use `/v1/voice/embed`.
## Portable profile registration
`POST /v1/voice/register` also accepts a JSON alternative to `audio`:
```javascript
// result is the parsed diarization response; slot is a selected raw speaker slot.
const request = {
model: "parakeet-diarization",
name: "Ada",
labels: {team: "research"},
speaker_slot: slot,
speaker_profiles: result.speaker_profiles
};
// POST JSON.stringify(request) with Content-Type: application/json.
```
Copy the complete `speaker_profiles` object returned by diarization unchanged.
Select `speaker_slot` explicitly, including for slot zero. It is the raw numeric
slot whose decimal string matches the diarization `label`, not a normalized
`SPEAKER_NN`, array index, or display name. `audio` and `speaker_profiles` are
mutually exclusive. `speaker_slot` without profiles is also invalid. Audio-only
registration keeps its existing JSON shape and behavior.
The server loads the requested, authorized model and obtains encoder identity and
dimension from backend metadata. It validates the complete profile export and
selects the requested usable slot. Missing slots, unavailable speech, unsupported
versions, non-finite/zero/wrong-size vectors and encoder mismatch return 400.
A backend without trusted encoder metadata returns 501. Success returns the
existing `{id, name, registered_at}` response.
Portable registrations store the **server-derived SHA-256 identity**, not a
caller-provided filename tag. Offline/live recognition admits these registrations
only when the loaded encoder has the same identity and dimension. Legacy audio
registrations retain their filename-tag compatibility rules. `/v1/voice/identify`
filters incompatible matches; a backend unable to report trusted identity cannot
match portable registrations, even when vector dimensions agree. Filtering can
return fewer than `top_k` results. The parakeet diarization model need not support
the separate audio-only VoiceEmbed RPC used by `/v1/voice/identify`.
Each successful enrollment inserts a new registration with its own ID and vector.
Duplicate display names do not merge embeddings or update an earlier enrollment.
There is no automatic enrollment or sample aggregation.
The recognition registry is **global, in-memory and per LocalAI instance**;
registrations are lost on restart and are not synchronized across frontends.
This is not durable “remembering” and not a per-user private address book. The
persistent `/api/voice-profiles` TTS-cloning feature is unrelated. Export and
registration use the existing voice-recognition permission, with existing model
access restrictions; permission does not establish biometric consent.
API tracing excludes the entire exchange for `/v1/audio/diarization`, its
`/audio/diarization` alias, and `/v1/voice/register` before capturing bodies.
This also protects JSON base64 audio when profile export is off. These routes
produce no in-memory or persisted API trace; other routes keep their existing
tracing behavior. External proxies and client logs must apply the same privacy
policy. Existing trace files from older versions are not retroactively scrubbed.
For offline and live diarization replay, registry tags never determine the
encoder dimension. LocalAI orders candidates by registration ID (tagged first),
then uses loaded encoder metadata to filter dimensions. Portable registrations
require an exact SHA-256 identity match as well. Older backends without trusted
metadata reject portable candidates and retain their native legacy dimension
checks. Identification filters compatibility after the store's `top_k` query;
incompatible results can crowd out compatible candidates within that window.
+3 -3
View File
@@ -74,9 +74,9 @@ availability may lag upstream releases.
### Chat Bots
- [Discord bot](https://github.com/mudler/LocalAGI/tree/main/examples/discord)
- [Slack bot](https://github.com/mudler/LocalAGI/tree/main/examples/slack)
- [Telegram bot](https://github.com/mudler/LocalAI/tree/master/examples/telegram-bot)
- [Discord bot](https://github.com/mudler/LocalAI-examples/tree/main/discord-bot)
- [Slack bot](https://github.com/mudler/LocalAI-examples/tree/main/slack-bot)
- [Telegram bot](https://github.com/mudler/LocalAI-examples/tree/main/telegram-bot)
- [Hellper (Telegram)](https://github.com/JackBekket/Hellper)
### Home Automation
@@ -123,6 +123,8 @@ curl -X POST http://localhost:8080/backend/shutdown \
Returns `200 OK` with the shutdown confirmation message on success.
Stopping a backend removes its watchdog timers and eviction state. A timeout from a stopped backend does not shut down a replacement at a different address.
## Error Responses
| Status Code | Description |
+2
View File
@@ -198,6 +198,8 @@ image blocks, and per-request usage tokens are dropped through the
internal `Predict()` signature. Use passthrough mode when your clients need
the upstream's full feature set.
In translate mode, an Anthropic response with `stop_reason: "refusal"` is returned to the client as an error instead of an empty successful reply, for both non-streaming and streaming requests. A streamed response may already have delivered partial content when the refusal arrives. Responses that end normally (`end_turn`) are unaffected, even when their content is empty.
#### Anthropic prompt caching
`proxy.cache_prompt: true` makes the translator add Anthropic