Files
LocalAI/docs/content/features/openai-realtime.md
T
2fa36e7147 feat(parakeet-cpp): speaker diarization, sound detection and live scene events (#12335)
* feat(parakeet-cpp): load diarization and CED models and companions

Repin PARAKEET_VERSION to parakeet.cpp PR #75's head, which adds
parakeet_capi_model_kind (ABI v8). Bind the new diarization, sound
event and combined scene stream C symbols through the same
purego.Dlsym probe pattern already used for the batched JSON entry
point, so the backend still loads against an older libparakeet.so.

Load now classifies the loaded GGUF by role (ASR, diarization or
sound) via parakeet_capi_model_kind and can load up to two companion
models from Options[] (asr_model:, diarization_model:, sound_model:,
paths resolved against opts.ModelPath), verifying each companion's
kind and freeing every context opened so far on any failure. Free
releases the primary and every companion. AudioTranscription now
names the loaded role when it is not ASR instead of a generic model
not loaded error. The dynamic batcher starts only when an ASR context
ends up loaded, primary or companion.

This is groundwork only: the Diarize and SoundDetection RPCs and the
live scene stream that actually use these new roles land in later
commits.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): reset role fields on a failed companion load

loadRoles' freeLoaded only released the C contexts it had opened; it
left ctxPtr/diarCtx/tagCtx and companions pointing at those now-freed
contexts, so a later Free() on the same instance would double-free.
Zero all four alongside the CppFree calls.

Also route AudioTranscriptionStream and AudioTranscriptionLive through
notASRError when ctxPtr is unset but a diarization or sound model is
loaded, matching AudioTranscription: both used to return the generic
model-not-loaded error instead of naming the loaded role.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): add speaker diarization

Implement the Diarize RPC for the parakeet-cpp Go backend, wired to
Nemotron-3-Diarization through libparakeet.so's diarization C-API.

Plain diarization uses parakeet_capi_diarize_pcm; when include_text is
set and an ASR companion is loaded, parakeet_capi_transcribe_and_
diarize_json fills each segment's text instead. Speaker labels are the
decimal index, or "unknown" for -1 (no diarized speaker overlaps).
min_duration_off merges same-speaker segments across a short gap
before min_duration_on drops the segments still too short, then ids
are renumbered. num_speakers/min_speakers/max_speakers/clustering_
threshold have no Sortformer equivalent and are logged at debug
instead of rejected.

Verified against the real Nemotron-3-Diarization + parakeet-tdt_ctc-
110m checkpoints on the two_speakers.wav fixture: correct A-B-A-B
speaker segmentation and matching speaker-attributed transcripts.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): add sound event detection

Wire the SoundDetection RPC to the CED tagger context (p.tagCtx)
loaded by Task 1's role classification. It runs the whole clip
through a one-shot parakeet_capi_sound_stream_* session (window
10s, hop 10s, top_k set to the tagger's class count so every
drained window carries a full score list), averages each class's
score across the drained windows, sorts descending, then applies
the request's threshold and top_k (0 keeps every class).

No tagCtx returns FailedPrecondition; a libparakeet.so missing the
sound_stream symbols returns Unimplemented. Every C call runs under
engineMu, and the stream is always freed, even when a feed or drain
call fails partway through.

Verified against a real ced-tiny-q8_0.gguf on the rooster.wav demo
clip: "Chicken, rooster" tops the list at score 0.91.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): cancel sound detection mid-feed, shrink the lock

SoundDetection now checks ctx before each 10 s feed slice (mirroring
driver.go's feedSlices) and returns Canceled if the caller gave up,
so a long clip can be interrupted instead of feeding to completion
regardless. The stream is still freed on every path, cancellation
included.

Also narrow engineMu to the C calls: the drained JSON document is
now decoded after the lock is released, splitting soundStreamScores
into a locked soundStreamDrain (opts, begin, feed, drain, free) and
an unlocked json.Unmarshal.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): stream speaker and sound events during live transcription

Add two additive proto fields, LiveSpeakerSegment and LiveSoundEvent,
repeated on TranscriptLiveResponse. When a diarization or sound
companion model is loaded, AudioTranscriptionLive now runs a no-ASR
scene stream (parakeet_capi_scene_stream_begin) beside the ASR
streaming session, feeding it the same PCM slices and forwarding any
closed speaker or sound events alongside the matching ASR delta, or
on their own when a slice has no ASR output.

The scene stream is freed and reopened on a mid-stream Config reset,
flushed with is_last before the closing FinalResult, and degrades
gracefully (a warning, not an error) when begin or a later feed call
fails, so live transcription keeps working ASR-only. Existing live
behavior is unchanged when no companion is configured, and no scene
C call is made in that case.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): keep scene events off the ASR critical path in live

Emit each slice's ASR result right after the ASR feed, before the
scene feed for that slice runs, so a companion diarization/sound
model never adds scene compute latency in front of the delta or
<EOU> that drives realtime turn detection. Closed speakers/sounds go
out afterward as their own response, so a slice with both now
produces two responses, ASR first. The live feed log line now
reports ASR and scene wall time separately.

Re-check the diarization/sound contexts a scene stream was begun
with against the live contexts before every feed, under the same
lock: Free() can race between an ASR feed and the matching scene
feed and free the model the stream borrows. A mismatch now returns
without touching the C side. Freeing the stream itself stays
unconditional; the scene stream's destructor only releases its own
buffers and never touches the borrowed contexts.

Also recover a panicking stub inside the live test goroutine instead
of crashing the test binary, and reset the live decode-lag tracker on
a mid-stream config reset, matching what its own comment already
promised.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(realtime): surface live speaker and sound events

Carry the backend's closed speaker segments and sound events
(TranscriptLiveResponse fields 7/8) through LiveTranscriptionEvent
as LiveSpeakerSegment/LiveSoundEvent (nanoseconds mapped to
seconds), and forward them from the semantic_vad live path.

Each speaker segment emits
conversation.item.input_audio_transcription.segment with speaker,
start, end and empty text under the turn's item id. Each sound
event emits conversation.item.sound_detection with one tag
(label, score = peak, index) and the event's new optional
start/end seconds fields, omitted when unset so the existing
unary/windowed sound-detection path is unaffected.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(realtime): keep start/end on a zero-second transcription segment

ConversationItemInputAudioTranscriptionSegmentEvent.Start/End used
omitempty, so a speaker segment starting at 0.0s dropped its
"start" key. Nothing emitted this event before the live scene-event
path, so drop omitempty: the segment always carries real times.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(gallery): add parakeet-cpp diarization, CED and realtime scene models

Add gallery entries for the new parakeet-cpp capabilities: standalone
Nemotron-3-Diarization, the same paired with the Parakeet TDT+CTC
110M ASR model for speaker-attributed text, CED-Tiny and CED-Base
sound classifiers, and a realtime scene bundle combining the
streaming EOU ASR model with diarization and sound companions.

SHA256 taken from the Hub API; licenses from each model card
(openmdw-1.1 for Nemotron-3-Diarization, apache-2.0 for CED,
cc-by-4.0 for the Parakeet ASR models).

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: document parakeet-cpp diarization, sound detection and live scene events

Cover the new parakeet-cpp capabilities across the feature pages:
Nemotron-3-Diarization as a diarization backend (with and without
speaker text, the ignored speaker-count hints, the Sortformer
voice-like-sound quirk), CED as a sound classification backend, the
asr_model/diarization_model/sound_model/diarization_latency companion
options, and the realtime live speaker/sound events (event shapes,
the speech-turn-only limitation, and using this or
pipeline.sound_detection but not both).

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): correct the realtime-scene license and wording nits

parakeet-cpp-realtime-scene mistakenly copied cc-by-4.0 from the
existing realtime_eou_120m-v1 entry; the model card lists the NVIDIA
open model license instead. Switch to the gallery's usual spelling
for that license and keep the diarization/CED licenses called out in
the description.

Also: audio-diarization.md now says getting per-segment text needs
both an asr_model companion and include_text=true on the request, and
audio-to-text.md's option table reads "Use on" (a pairing the loader
does not enforce) instead of "Allowed on".

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): reject a companion role that duplicates the primary's

loadRoles let a companion option (asr_model:/diarization_model:/
sound_model:) assign into a role field the primary already occupied,
for example asr_model: on an already-ASR primary. The companion's
context silently overwrote ctxPtr/diarCtx/tagCtx, and Free() only
walks those three fields, so the original primary context was never
freed again.

Reject a companion whose role the primary already holds before its
GGUF is even loaded, freeing everything loadRoles opened so far, the
same way a wrong-kind companion is already rejected.

Also warn, rather than silently fall through, when
parakeet_capi_model_kind reports PARAKEET_MODEL_KIND_NONE for a
successfully loaded primary; the primary is still treated as ASR,
matching today's behavior.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): cap live scene sound score retention

sceneBegin started the live diarization/sound companion stream with
the C API's default sound options, whose top_k keeps 5 scores per
window forever until drained. The live scene path never drains sound
scores (only the offline SoundDetection RPC does, with its own fresh
stream), so this window queue on the C side grew for the whole
session's lifetime.

Set opts.Sound.TopK = 0 before starting the scene stream: this
disables score retention while leaving sound event detection (onset/
offset), which the live path actually consumes, unaffected.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): merge diarization segments per speaker, harden Diarize

mergeCloseSegments only compared neighbors in the single start-sorted
segment list, so two same-speaker segments never merged once another
speaker's turn fell between them (A, B, A): the short B segment broke
the adjacency the merge relied on. Group segments by speaker first,
merge within each speaker's own start-ordered run, then re-sort the
result by start so interleaved speakers come back out in timeline
order.

Also harden Diarize's entry points the same way streamFeedDoc/
sceneFeed already are: diarizeCall re-checks p.diarCtx (and, on the
include_text path, p.ctxPtr) under engineMu right before the C call,
so a Free() racing between Diarize's own checks and the lock can no
longer reach the C side with a freed context. When the include_text
call returns NULL, last_error is now read from both contexts and
whichever came back non-empty is reported, since either side of the
pairing can be the one that failed. A WAV decode failure is reported
as InvalidArgument instead of an unwrapped/untyped error.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): harden SoundDetection's engine checks

soundStreamDrain ran every C call under engineMu but never re-checked
p.tagCtx there, so a Free() racing between SoundDetection's own
tagCtx==0 check and this lock could still reach the C side with a
freed context. Re-check p.tagCtx under the lock and return
ModelNotLoaded when it was cleared, mirroring diarizeCall's own
re-check. A WAV decode failure is now reported as InvalidArgument
instead of an unwrapped/untyped error.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(parakeet-cpp): cover a mid-session scene feed failure

feedSlicesScene already degrades gracefully when a scene feed call
fails mid-session: it frees the broken stream and carries the ASR-only
session forward. Add a spec covering that path end to end: the scene
stream is freed exactly once, later audio slices still produce ASR
responses, and no speaker/sound events appear before or after the
failure.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: fix the parakeet-cpp companion role table and realtime scene docs

audio-to-text.md's companion option table read "Use on" with a note
that the loader did not enforce the pairing; it now rejects a
companion whose role duplicates the primary's, so restore the
"Allowed on" wording and describe the real enforcement.

openai-realtime.md's live speaker/sound section claimed a mid-stream
session.update resets the companion stream and that it flushes on
session close; neither happens, since the realtime core opens one
live stream (and so one scene stream) per speech turn and closes it
at that turn's commit, with no mid-stream Config in between. Document
that lifecycle instead, state precisely that start/end are seconds
from the start of the turn's own audio, and note that the diarization
model starts a fresh session every turn, so a speaker index is only
meaningful within one turn. The example sound tag ("Rooster", index
17) did not match any real CED label; index 17 in ced-tiny-q8_0.gguf
is "Baby laughter". Replaced with "Chicken, rooster" at its real
index, 99.

Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(parakeet-cpp): use CED's real index for Chicken, rooster

The scene feed comment and the live test's canned document gave
"Chicken, rooster" index 365. In CED's AudioSet label list it is 99,
which is also what the realtime docs show.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(realtime): call the test event accessor

The scene-event tests range over a method instead of its returned slice.
Call the synchronized accessor so the OpenAI test package compiles.

Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(parakeet-cpp): pin parakeet.cpp master with sound events

mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and
#74 (the missing <algorithm> include that broke the image builds) are
on master now. Pin 6dea76a instead of the #75 PR head, and update the
header comment the bump bot reads.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(parakeet-cpp): pin parakeet.cpp with ced.cpp on main

parakeet.cpp #76 moved its ced.cpp submodule from the head of
localai-org/ced.cpp#3 (a branch-only commit) to ced.cpp main, where
#3 landed with an identical tree. Pin 623a968 so the image builds no
longer depend on that branch.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(transcription): carry speaker labels on words and streamed segments

A diarizing backend could label transcript segments, but two paths
dropped the label: TranscriptWord had no speaker field, so live
transcription words and word-level timestamps could not carry one, and
the stream=true transcript.text.done event left the speaker out of
its segments.

TranscriptWord gains an optional speaker (proto field 4, additive).
It flows through the live event and result mapping, the JSON word
output of the endpoint and the CLI, and transcript.text.done now
includes a segment's speaker when there is one. Empty labels are
omitted, so responses without diarization are unchanged.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
(cherry picked from commit 2f0049f979)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(importers): detect the parakeet.cpp diarization GGUF

The Nemotron-3-Diarization GGUFs are published in
mudler/parakeet-cpp-gguf as nemotron-3-diarization-<quant>.gguf. The
parakeet-cpp importer did not recognise that name, so a direct
`local-ai models import` of the file fell through to another importer.

A direct URL to the file now imports with the diarization usecase. A
repo import still picks ASR weights when the repo also ships the
diarization model, and falls back to the diarization weights only
when there are no others.

Ported from #12323.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): advertise diarization and sound detection for parakeet-cpp

The capability table listed parakeet-cpp as transcription only, though
the backend now answers Diarize (Nemotron-3-Diarization) and
SoundDetection (CED) depending on the model kind it loads.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(parakeet-cpp): label transcript segments with the diarization companion

A diarization_model companion only fed live speaker events and
Diarize; /v1/audio/transcriptions ignored it.

With the companion attached and diarize=true (the OpenAI endpoint's
default), unary transcription now labels each segment with its
speaker and splits segments at speaker turns; with word timestamps
each word carries its speaker. The stream=true final result labels
each utterance with the speaker who said most of it. Both use the
checkpoint's own diarization over the whole clip, as NeMo's diarize()
does. Words take the speaker whose segments overlap them most, or the
nearest segment within 0.5 s, the same rule as parakeet.cpp's
speaker-attributed ASR.

Docs: the diarization_model row and a paragraph on transcript
speakers; Nemotron-3-Diarization handles up to 8 speakers.

Ported from #12323.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(realtime): speaker segments from committed-turn transcription

Speaker events reached a realtime session only from the live
semantic_vad path, which needs a cache-aware streaming transcription
model. Committed-turn transcription (server_vad, or any offline
model) always asked the backend for diarize=false and dropped the
segments' speakers.

pipeline.diarization (off by default) asks the transcription model for
speaker labels on each committed turn and emits every labelled segment
as a conversation.item.input_audio_transcription.segment event, with
its text, before the turn's completed event. It is opt-in because some
backends fail a diarization request they cannot serve.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add parakeet-cpp-realtime-scene-tdt

parakeet-cpp-realtime-scene pairs the streaming EOU model with the
diarization and CED companions; its speaker and sound events need a
cache-aware streaming model. This entry does the same with Parakeet
TDT 0.6B v3 (multilingual, offline) for realtime under server_vad:
set it as both transcription and sound_detection and turn on
pipeline.diarization, and each committed turn gets speaker segments
and sound tags from one parakeet-cpp backend.

Files and sha256 match the Hub and are shared with the existing TDT v3,
diarization and CED-Tiny entries. A real-model spec checks the
combination on a clip with two speakers and a rooster: A-B-A-B speaker
turns, and "Chicken, rooster" among the sound tags. The test loader
now binds the sound entry points like main.go.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add CED-Base variants of the parakeet-cpp scene models

parakeet-cpp-realtime-scene and parakeet-cpp-realtime-scene-tdt ship
with CED-Tiny. The -base variants use CED-Base (86M), which tags sounds
more confidently (on the rooster clip "Crowing" 0.65 against 0.49 for
Tiny).

Measured on CPU over a 37 s clip: the live diarization + sound stream
runs at 0.125 of real time with CED-Base against 0.103 with CED-Tiny,
because diarization dominates; sound detection per committed turn costs
0.031 against 0.005. The realtime docs list both and note that any CED
size works as sound_model.

Files and sha256 match the Hub and are shared with the existing
parakeet-cpp-ced-base entry. The TDT variant passes the real-model
scene spec with CED-Base (A-B-A-B speakers, "Chicken, rooster" found).

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): register pipeline.diarization in the config metadata

TestAllFieldsHaveRegistryEntries fails on the branch because the new
pipeline.diarization field has no registry entry. Add one so the model
editor shows it as a toggle next to the sound detection options.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-29 23:58:17 +02:00

40 KiB


title: "Realtime API" weight: 37

The realtime voice loop: VAD to STT to LLM to TTS, over WebSocket or WebRTC

LocalAI supports the OpenAI Realtime API which enables low-latency, multi-modal conversations (voice and text) over WebSocket.

To use the Realtime API, you need to configure a pipeline model that defines the components for Voice Activity Detection (VAD), Transcription (STT), Language Model (LLM), and Text-to-Speech (TTS).

Configuration

Create a model configuration file (e.g., gpt-realtime.yaml) in your models directory. For a complete reference of configuration options, see [Model Configuration]({{%relref "advanced/model-configuration" %}}).

name: gpt-realtime
pipeline:
  vad: silero-vad-ggml
  transcription: whisper-large-turbo
  llm: qwen3-4b
  tts: tts-1

This configuration links the following components:

  • vad: The Voice Activity Detection model (e.g., silero-vad-ggml) to detect when the user is speaking.
  • transcription: The Speech-to-Text model (e.g., whisper-large-turbo) to transcribe user audio.
  • llm: The Large Language Model (e.g., qwen3-4b) to generate responses.
  • tts: The Text-to-Speech model (e.g., tts-1) to synthesize the audio response.

Make sure all referenced models (silero-vad-ggml, whisper-large-turbo, qwen3-4b, tts-1) are also installed or defined in your LocalAI instance.

A pipeline stage can name a [failover chain]({{%relref "features/model-failover" %}}); the stage then switches targets without closing the session.

Streaming the pipeline

By default each stage runs to completion before the next begins: the whole utterance is transcribed, the full LLM reply is generated, then it is synthesized. Each stage can instead be streamed incrementally, which lowers the time-to-first-audio of a turn:

name: gpt-realtime
pipeline:
  vad: silero-vad-ggml
  transcription: whisper-large-turbo
  llm: qwen3-4b
  tts: tts-1
  streaming:
    llm: true             # stream LLM tokens as transcript deltas
    tts: true             # emit audio deltas per synthesized chunk
    transcription: true   # stream transcript text deltas of the user's speech
    clause_chunking: true # synthesize each clause as soon as it completes
  • streaming.tts: emit a response.output_audio.delta per audio chunk the TTS backend produces (requires a backend that supports streaming synthesis), instead of one delta for the whole utterance. Falls back to a single unary delta otherwise.
  • streaming.transcription: stream conversation.item.input_audio_transcription.delta events as the transcript is produced (requires a transcription backend that supports streaming).
  • streaming.llm: stream the LLM reply token-by-token as response.output_audio_transcript.delta events. The full reply is buffered and synthesized once it is complete - streamed as audio chunks when streaming.tts is enabled (and the TTS backend supports it), otherwise as a single unary delta. Reasoning/thinking is always stripped from the spoken transcript. Tool calls are supported while streaming when the LLM uses its tokenizer template (use_tokenizer_template: true): the backend's autoparser then delivers content and tool calls separately, so the spoken transcript never leaks tool-call tokens. Grammar-based function calling keeps the buffered path.
  • streaming.clause_chunking: instead of buffering the whole reply before TTS, split it into speakable clauses and synthesize each as soon as it completes, lowering the time-to-first-audio. The splitter is script-aware: it uses Unicode sentence segmentation (so it handles CJK 。!? with no whitespace), CJK clause punctuation (,、;:), and Thai/Lao spaces - it does not rely on whitespace sentence boundaries, so it works for languages such as Chinese, Japanese and Thai where the old per-sentence approach degraded to whole-message buffering. Requires streaming.llm; scripts that genuinely need a dictionary (e.g. Khmer, Burmese) simply stay buffered until a space or end-of-message. Off by default.

All streaming flags are off by default, so existing pipelines are unaffected.

Model warm-up (cold start)

Without warm-up the pipeline's models are loaded into memory only on first use within a session: the VAD on the first audio chunk, transcription at the first end-of-speech, the LLM on the first reply, and TTS on the first spoken output. On a cold session this staggers a load delay across those first few interactions - and a model that fails to load (missing weights, wrong backend, out of memory) only fails part-way through the first turn.

To avoid that, LocalAI warms the pipeline by default: it loads the VAD, transcription, LLM and TTS backends into memory before the session is announced, and the session start blocks until they are all ready. The loads run concurrently, so the wait is the slowest single model, not the sum. This means:

  • The first turn pays no cold-start cost - every backend is already resident.
  • Model-load errors surface at session start. If any stage fails to load, the session is not started and the client receives a model_load_error instead of session.created, so a broken pipeline fails fast and visibly rather than mid-call.

Set disable_warmup: true to restore the lazy "load on first use" behavior - session start no longer waits on loading and load errors surface on the first turn instead. Useful if you want idle sessions to avoid holding model memory they may never use:

name: gpt-realtime
pipeline:
  vad: silero-vad-ggml
  transcription: whisper-large-turbo
  llm: qwen3-4b
  tts: tts-1
  disable_warmup: true   # lazily load each model on first use instead of at session start

Pre-loading a pipeline on demand

Warm-up only fires when a realtime session opens. To load a pipeline into memory ahead of time - e.g. to warm it right after boot, or when running with disable_warmup: true - POST the model name to the admin-only /backend/load endpoint. For a pipeline model it loads every configured sub-model (VAD, transcription, LLM, TTS, sound_detection, voice_recognition) concurrently:

curl -X POST http://localhost:8080/backend/load \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-realtime"}'

The endpoint is not realtime-specific - it pre-loads any model. See [Backend Monitor]({{%relref "operations/backend-monitor" %}}) for the full request/response reference (it is the inverse of /backend/shutdown).

Turn detection

Turn detection decides when the user has finished speaking and the pipeline should respond. Two modes are supported, matching the OpenAI session schema:

  • server_vad (default): silence-based. The VAD model watches the audio and the turn commits after silence_duration_ms (default 500 ms) of silence. Simple and model-agnostic, but a fixed silence window must trade interrupting mid-sentence pauses against sluggish responses.
  • semantic_vad: model-driven. The transcription model itself signals end-of-utterance and the silence window becomes dynamic: short right after the model emits its end-of-utterance token, much longer when it does not - so pausing to think no longer gets cut off, while finished sentences get a fast response.

semantic_vad requires a transcription model that emits an end-of-utterance token over a cache-aware streaming decode - currently parakeet-cpp-realtime_eou_120m-v1 (the model is trained to distinguish "paused, expecting a reply" from "paused mid-thought"). The realtime pipeline feeds it the microphone audio live while the user speaks. With any other transcription backend the session degrades gracefully to silence-only detection using the eagerness timeout below (a warning is logged once). The model also emits a distinct end-of-backchannel token (<EOB>) for short acknowledgments like "uh-huh": those are transcribed but never treated as the user yielding the turn.

Sessions can opt in via session.update (turn_detection: {"type": "semantic_vad", "eagerness": "medium"}), or the pipeline can set a server-side default so clients need no changes:

name: gpt-realtime
pipeline:
  vad: silero-vad-ggml
  transcription: parakeet-cpp-realtime_eou_120m-v1
  llm: qwen3-4b
  tts: tts-1
  turn_detection:
    type: semantic_vad   # default for sessions on this model (server_vad if unset)
    eagerness: medium    # low | medium | high | auto (auto == medium)
    retranscribe: false  # see below
    # vad_window_sec: 6  # widen the per-tick VAD scan window (see below)

A client session.update still overrides type and eagerness per session.

VAD scan window: each turn-detection tick the VAD rescans only the most recent slice of buffered audio — sized automatically from the commit silence threshold (the server_vad silence window, or the semantic eagerness fallback) plus a warm-up margin, so long turns cost the same per tick as short ones. vad_window_sec widens the window if needed; values below the automatic floor are ignored. Buffered turn audio is retained for at most 90 s — a turn that genuinely never pauses (continuous speech or a noise source the VAD keeps classifying as speech) keeps only its most recent 90 s for the commit-time batch transcription (the semantic live stream is unaffected — it already consumed the audio incrementally).

Eagerness sets the fallback silence window used when no end-of-utterance token was seen (the model missed it, or the user genuinely trails off): low waits 8 s, medium/auto 4 s, high 2 s - the same max-timeout semantics OpenAI documents. After the token is seen, the turn commits on the next VAD tick (~300 ms).

Live captions: while the user speaks, semantic_vad streams conversation.item.input_audio_transcription.delta events under the item id the commit will later reuse, so clients can render the words as they are recognized. The completed event at commit carries the authoritative transcript and replaces the partial text (with retranscribe: true it may differ from the captions); a turn discarded before commit emits conversation.item.input_audio_transcription.failed so clients can retract its captions.

retranscribe (server-side only, semantic_vad only) cross-checks the streaming decode against a batch decode at commit time:

  • false (default): the transcript accumulated from the live stream is used as-is - the model runs once per utterance and the LLM starts immediately at commit.
  • true: the committed audio is re-transcribed offline. If the batch decode also ends with the end-of-utterance token the turn proceeds (using the batch transcript); if it does not, the commit is cancelled and the session keeps listening - treating the streaming token as a false positive. Both transcripts are compared and logged, which makes this mode a useful diagnostic for how well the streaming and batch decodes align, at the cost of one extra decode per turn.

Live speaker and sound events (parakeet-cpp)

When the semantic_vad transcription model is a parakeet-cpp model loaded with a diarization_model and/or sound_model companion (see [Audio to Text]({{% relref "audio-to-text" %}})), the realtime session also streams speaker and sound events while a turn is live, alongside the transcript deltas. Nothing needs to change on the client: unrecognized event types are ignored by standard OpenAI Realtime clients.

The transcription model, with its companions:

name: parakeet-realtime-scene
backend: parakeet-cpp
parameters:
  model: realtime_eou_120m-v1-f16.gguf
options:
  - diarization_model:nemotron-3-diarization-q8_0.gguf
  - sound_model:ced-tiny-q8_0.gguf

The realtime pipeline that uses it:

name: gpt-realtime
pipeline:
  vad: silero-vad-ggml
  transcription: parakeet-realtime-scene
  llm: qwen3-4b
  tts: tts-1
  turn_detection:
    type: semantic_vad

Each closed speaker segment emits a conversation.item.input_audio_transcription.segment event under the turn's item id, with an empty text (the event exists to carry the speaker boundary, not a transcript - the transcript still comes from the ordinary delta/completed events):

{
  "type": "conversation.item.input_audio_transcription.segment",
  "item_id": "item_abc",
  "content_index": 0,
  "speaker": "0",
  "start": 1.92,
  "end": 4.10,
  "text": ""
}

Each sound event emits a conversation.item.sound_detection event with one tag and the detection window's start/end:

{
  "type": "conversation.item.sound_detection",
  "item_id": "item_abc",
  "content_index": 0,
  "detections": [{"label": "Chicken, rooster", "score": 0.91, "index": 99}],
  "start": 24.0,
  "end": 30.0
}

The start/end on both event types are seconds measured from the start of the current turn's own audio, not the session or the WebSocket connection - the same base the streamed transcript words use.

The companion stream is opened fresh for each speech turn, alongside that turn's ASR live session, and closed when the turn commits: whatever it had not yet emitted is drained and sent at that point. Because the diarization model runs a brand new session every turn, its speaker indices are scoped to the turn too - "speaker": "0" in one turn and "speaker": "0" in the next are not guaranteed to be the same person, even within the same conversation.

score is the peak score seen for that tag while the sound was live, not an average.

Limitation: under semantic_vad, live transcription (and so this companion stream) only runs during speech turns - it does not see audio between turns. A sound that happens while nobody is speaking is not detected this way. If you need sound events independent of speech turns, use the pipeline's sound_detection model instead (see [Sound Classification]({{% relref "audio-classification" %}})), which classifies each VAD-committed utterance on its own. Use one or the other, not both, on the same session - they overlap in purpose and would emit sound detections twice.

Speaker and sound events with an offline model (Parakeet TDT v3)

The live events above need a cache-aware streaming transcription model. An offline model such as Parakeet TDT 0.6B v3 (25 languages) runs under server_vad instead: each VAD-committed turn is transcribed as a whole. The gallery model parakeet-cpp-realtime-scene-tdt bundles it with Nemotron-3-Diarization and CED-Tiny, so one parakeet-cpp backend handles transcription, speakers and sounds. Point both transcription and sound_detection at it and turn on diarization:

name: gpt-realtime-scene
pipeline:
  vad: silero-vad-ggml
  transcription: parakeet-cpp-realtime-scene-tdt
  sound_detection: parakeet-cpp-realtime-scene-tdt
  diarization: true
  llm: qwen3-4b
  tts: tts-1

pipeline.diarization asks the transcription model for speaker labels on each committed turn and emits every labelled segment as a conversation.item.input_audio_transcription.segment event before the turn's completed event. Unlike the live path, these segments carry their text:

{
  "type": "conversation.item.input_audio_transcription.segment",
  "item_id": "item_abc",
  "content_index": 0,
  "id": "seg_1",
  "speaker": "1",
  "start": 6.85,
  "end": 10.82,
  "text": "Well, I don't wish to see it any more, observed Phoebe, turning away her eyes."
}

sound_detection classifies the same committed audio and emits one conversation.item.sound_detection event per turn (see [Sound Classification]({{% relref "audio-classification" %}})). As on the live path, times are relative to the turn's audio and speaker labels are only consistent within a turn. pipeline.diarization is off by default: it needs a transcription model that diarizes (parakeet-cpp with a diarization_model companion), and some other backends fail a diarization request they cannot serve.

Choosing the sound model

Both scene models ship with CED-Tiny, the cheapest to run all the time. parakeet-cpp-realtime-scene-base and parakeet-cpp-realtime-scene-tdt-base are the same pipelines with CED-Base (86M, the largest CED), which tags sounds more confidently. Any CED GGUF from mudler/ced-gguf (tiny, mini, small, base) works as sound_model. Measured on CPU (Ryzen 9 9950X3D) over a 37 s clip with two speakers and a rooster, as a fraction of real time:

CED-Tiny CED-Base
Live scene stream (diarization low + sound), EOU path 0.103 0.125
Sound detection per committed turn, TDT path 0.005 0.031

The EOU model's own ASR stream adds 0.016. Diarization dominates the live cost, so CED-Base keeps the live path about 7x faster than real time.

Disabling thinking

For reasoning models, you can force the pipeline LLM's thinking off without editing the LLM model config:

pipeline:
  llm: qwen3-4b
  disable_thinking: true   # maps to enable_thinking=false for the realtime LLM

This is applied only to the realtime session's copy of the LLM config, so it does not affect other users of the same model. Leave it unset to use the LLM model config's own reasoning settings.

Conversation compaction (long sessions on CPU)

By default a realtime session feeds only the last max_history_items turns to the LLM; older turns are dropped and forgotten. On CPU, long calls also grow expensive as the prompt fills with verbatim history. Enable compaction to instead fold older turns into a rolling summary, so long calls stay cheap without losing earlier context.

Compaction works with two numbers:

  • max_history_items is the live window - the recent turns kept verbatim in the prompt.
  • compaction.trigger_items is the high-water mark - let the buffer grow to here, then summarize the overflow (everything above max_history_items) into a rolling memory and evict it. It must be greater than max_history_items; if it is not, it is clamped up.

The gap between the two controls how often summarization runs: a summary call fires roughly every (trigger_items - max_history_items) turns (here, about every 6 turns).

pipeline:
  max_history_items: 6        # live window - recent turns kept verbatim
  compaction:
    enabled: true
    trigger_items: 12         # summarize overflow back down to max_history_items
    summary_model: ""         # optional: a small model for the summary (CPU); default = pipeline LLM
    max_summary_tokens: 512

{{% notice tip %}} On CPU, set summary_model to a small, fast model so compaction never competes with the conversation LLM for compute. Left empty, the pipeline's own LLM produces the summary. {{% /notice %}}

Clients can also manage history directly via the now-supported conversation.item.delete, conversation.item.truncate, and input_audio_buffer.clear realtime events.

Classifier mode (LocalAI extension)

On hardware that can afford prompt processing but not token generation — a Raspberry Pi running a small LLM, for example — a realtime session can replace autoregressive generation with prefill-only classification: you register a fixed list of options, each user turn is scored against them with the Score primitive (a single forward pass, no decode), and the winning option's canned reply is spoken and/or its canned tool call is emitted. On llama-cpp, scoring runs through the same server slot the LLM uses, so the conversation prefix stays KV-cached across turns, and all options are scored together in one batched decode: the shared prefix (prompt plus the options' common token prefix) is processed once, then each option's unique tail rides a forked sequence in a single forward pass. A warm turn costs roughly one pass over the new words plus one small batch over the option tails, independent of the option count.

Enable it in the pipeline config:

name: drone-pi
pipeline:
  vad: silero-vad-sherpa
  transcription: parakeet-cpp-realtime_eou_120m-v1
  llm: lfm2.5-1.2b-instruct       # scores AND (if asked) generates
  tts: vits-piper-en_US-amy-sherpa
  classifier:
    enabled: true
    threshold: 0.85               # softmax floor the winner must clear (see note below)
    fallback:
      mode: reply                 # none | reply | generate
      reply: "Say again?"
    options:
      - id: up
        description: the user asks the drone to move or fly up/higher
        reply: Going up.
        tool:
          name: move
          arguments: {direction: up}
      - id: greeting
        description: the user greets the assistant
        reply: Hello, ready to fly.

Or per session / per response from the client (the field is additive — OpenAI clients simply never send it):

{"type": "session.update", "session": {"type": "realtime", "localai_classifier": {
  "enabled": true, "threshold": 0.85,
  "options": [{"id": "up", "description": "...", "reply": "Going up.",
               "tool": {"name": "move", "arguments": {"direction": "up"}}}],
  "fallback": {"mode": "reply", "reply": "Say again?"}
}}}

Like tools, the option list is replaced wholesale on each update. A response.create may carry its own localai_classifier to override the session for one response — {"enabled": false} runs normal generation once.

How a classified response behaves:

  • The winner's reply is emitted through the ordinary response events (spoken via TTS, or response.output_text.* in text-only mode), and its tool (if any) is emitted as a standard function_call item with exactly the configured arguments — the client executes it and reports back with conversation.item.create as usual. In classifier mode you typically should not send a follow-up response.create after the tool output: the canned reply already acknowledged the command, and the follow-up would classify a tool-output turn.
  • Every classified response also emits a localai.classifier.result event carrying the full softmax distribution, the chosen option id (empty when the fallback applied), the threshold, and the scoring latency — useful for visualizing confidence in a client UI.
  • A committed turn whose transcript is empty (the VAD fired on noise and the ASR heard no words) is never scored — an empty prompt produces a confidently arbitrary winner. The fallback applies directly: the result event carries an empty scores list, and generate mode falls through to generation.
  • Wake-word gating: set address: {names: ["drone"], mode: ignore} and the assistant only acts on turns that mention one of the names ("Drone go up", not just "go up"). The check is a deterministic case-insensitive whole-word match on the latest transcript — deliberately not model-based: scoring cannot detect a missing name (a 1.2B scorer rates "go up" as addressed with p≈1.0 even with a dedicated addressing stage), while a literal match is exact and free. Unaddressed turns skip scoring entirely (ambient conversation costs nothing) and emit a result event with empty scores and fallback: "not_addressed"; mode: ignore completes the response silently, mode: reply speaks reply.
  • When no option clears threshold, fallback.mode decides: none completes the response with no output, reply speaks the canned fallback reply, and generate falls through to normal autoregressive generation (slow on weak hardware, but always available since the same model config serves both paths). Set the threshold high: with the default raw normalization a confident in-list pick lands near 1.0, while an out-of-list request spreads its probability across the options — measured on a 1.2B scorer, in-list utterances scored ≥0.97 and out-of-list ones peaked around 0.8, so a floor of ~0.85 separates them. A low threshold (say 0.35) practically never falls back. Also keep each option's description narrowly scoped: a catch-all clause like "…or asks for help" turns that option into a magnet for every request the model cannot map, defeating the fallback.
  • Agentic follow-up turns (after a server-side assistant tool executes) always use generation — the option list describes user intents, not tool outputs.

Knobs that matter for latency and accuracy: keep option descriptions short (they all go into the scoring system prompt) and the option count small. By default only the latest user message is scored — earlier turns echo option names (the canned replies especially) and empirically make small scoring models re-choose the previous option regardless of the new command. history_items: N opts the trailing N conversation messages back in (role-labeled); only do that with a scorer large enough to weigh the context. normalization: mean divides each option's joint log-prob by its token count — useful when option ids have very different lengths. The scoring model needs a Go-side chat template (template.chat / template.chat_message); without one the scoring prompt falls back to a generic ChatML envelope, which may be off-distribution for the model. Use classifier.model to score on a different config than the pipeline LLM (rarely needed).

Argument slots (hybrid classify-then-complete)

A canned tool call can leave holes for the model to fill: declare slots on the option's tool and reference them as "{{name}}" in the arguments template. When the option wins, LocalAI runs a short grammar-constrained completion that continues the exact scoring prompt (so the llama.cpp prompt cache is already warm) with the chosen route JSON re-opened at the first slot — only the value tokens are free; everything else is pinned by the grammar. Number slots substitute unquoted, enum/string slots inside their quotes.

tool:
  name: move_drone
  arguments:
    direction: forward
    distance: "{{distance}}"
    units: "{{units}}"
  slots:
    - name: distance
      type: number            # number | enum | string
    - name: units
      type: enum
      values: [m, meters, ft, feet]
      default: m              # used if inference fails outright
      hint: assume m when the user gives no units

The option's spoken reply can reference the same placeholders — reply: "Going forward {{distance}} {{units}}." — and the filled values are spliced in as plain text before the reply is emitted, so what the assistant says confirms what it inferred. Reply placeholders are optional (one that names no slot stays literal).

Slot declarations (and hints) are appended to the option's description in the shared system prompt, so they also inform scoring and cost no extra per-turn tokens. The localai.classifier.result event carries the final arguments and a fill_latency_ms. On an inference failure the slots' defaults apply (and template the reply); if any slot lacks a default the response fails (or falls through to generation with fallback.mode: generate). Slot filling requires completion in the scoring model's known_usecases alongside score.

Registering an option list (pipeline seed or session.update) prewarms the scoring prompt in the background: one throwaway score prefills the option-list prompt and plants a state checkpoint exactly at the per-turn probe boundary. Every scoring call also declares that boundary (the probe-invariant prompt prefix) to the backend, which checkpoints there on each prefill — on hybrid/recurrent models (which cannot rewind their state arbitrarily, only restore checkpoints) this is what keeps every turn at probe-size cost instead of a full option-list re-prefill. A client that swaps option lists at runtime (e.g. voice-switched command modes) pays nothing on the first turn after a swap: the rewarm hides behind the acknowledgement reply, and it runs on every registration deliberately — if the list's slot was evicted, the rewarm is exactly the re-prefill the next turn would otherwise pay in the foreground. Swapping between several lists? Give the scoring model a slot per list and make the prefix routing selective: options: [parallel:3, sps:0.5] (the default slot-similarity threshold of 0.1 funnels different lists onto one slot — they share enough prompt structure to clear it).

The concrete scoring model must declare score in known_usecases. A single llama.cpp model can serve ordinary inference and classification concurrently by declaring multiple use cases, for example known_usecases: [chat, completion, score]; LocalAI reserves the scoring slots only when score is present. Scoring also requires the unified KV cache, which is enabled by default, so a score-enabled model cannot set kv_unified:false.

Transports

The Realtime API supports two transports: WebSocket and WebRTC.

WebSocket

Connect to the WebSocket endpoint:

ws://localhost:8080/v1/realtime?model=gpt-realtime

Audio is sent and received as raw PCM in the WebSocket messages, following the OpenAI Realtime API protocol.

WebRTC

The WebRTC transport enables browser-based voice conversations with lower latency. OpenAI-compatible clients can send a raw SDP offer and select the model with the query parameter:

POST http://localhost:8080/v1/realtime/calls?model=gpt-realtime
Content-Type: application/sdp

<SDP offer body>

The response has the application/sdp content type and contains the bare SDP answer.

The unified OpenAI interface is also supported. Send multipart/form-data with an sdp field that contains the offer and a JSON session field. LocalAI reads the model from the session object:

curl http://localhost:8080/v1/realtime/calls \
  -F "sdp=<offer.sdp;type=application/sdp" \
  -F 'session={"type":"realtime","model":"gpt-realtime"};type=application/json'

LocalAI also accepts its original JSON request format for compatibility. A JSON request contains top-level sdp and model fields and receives a JSON response with sdp and session_id fields.

Opus backend requirement

WebRTC uses the Opus audio codec for encoding and decoding audio on RTP tracks. The opus backend must be installed for WebRTC to work. Install it from the Backends page in the web UI, or from the backend gallery with the API:

curl -X POST http://localhost:8080/backends/apply \
  -H "Content-Type: application/json" \
  -d '{"id": "opus"}'

For a local binary installation, you can instead use the CLI:

local-ai backends install opus

Or set the EXTERNAL_GRPC_BACKENDS environment variable if running a local build:

EXTERNAL_GRPC_BACKENDS=opus:/path/to/backend/go/opus/opus

The opus backend is loaded automatically when a WebRTC session starts. It does not require any model configuration file - just the backend binary.

WebRTC behind Docker host networking or NAT

By default pion gathers a host ICE candidate for every local interface. Under Docker host networking that includes bridge addresses (docker0/veth, 172.x) that a remote browser cannot route to: the call typically connects on a good candidate and then drops a few seconds later when ICE consent checks fail on the unreachable ones. Two settings let you advertise only the reachable address:

# Advertise these IPs as the host ICE candidates (e.g. the host's LAN IP)
LOCALAI_WEBRTC_NAT_1TO1_IPS=192.168.1.10

# ...or restrict ICE gathering to specific interfaces
LOCALAI_WEBRTC_ICE_INTERFACES=eth0

{{% notice tip %}} For a browser on another LAN machine talking to LocalAI in a host-networked container, set LOCALAI_WEBRTC_NAT_1TO1_IPS to the host's LAN IP. This is the most reliable fix for WebRTC connections that establish and then drop. {{% /notice %}}

Fixed WebRTC UDP port

By default, each WebRTC peer connection uses an ephemeral UDP port. To route all realtime WebRTC ICE traffic through one shared port, start LocalAI with --web-rtc-udp-port 3478 or set LOCALAI_WEBRTC_UDP_PORT=3478.

When running in a container, publish the same port with the UDP protocol:

docker run -p 8080:8080 -p 3478:3478/udp \
  -e LOCALAI_WEBRTC_UDP_PORT=3478 localai/localai:latest

Allow the selected UDP port through the host and network firewalls. If LocalAI cannot bind it, WebRTC signaling requests return an HTTP 500 error describing the bind failure.

Protocol

The API follows the OpenAI Realtime API protocol for handling sessions, audio buffers, and conversation items.

Response modalities (text-only sessions)

By default a realtime session responds with audio plus a transcript. To make the model respond with text only, set the output modalities on either the session or an individual response:

{"type": "session.update", "session": {"output_modalities": ["text"]}}
{"type": "response.create", "response": {"output_modalities": ["text"]}}

output_modalities is the GA field name. For compatibility, LocalAI also accepts the legacy Realtime beta field name modalities as an alias (a lot of community sample code still sends modalities: ["text"]):

{"type": "session.update", "session": {"modalities": ["text"]}}

The GA output_modalities wins when both are present. A response-level value overrides the session-level one, and when neither is set the session falls back to ["audio"].

Out-of-band responses and metadata

A response can be created outside the default conversation by setting conversation to none. The reply is not added to the conversation history, which is what makes it usable for a side channel — answering a chat message or a webhook while a spoken conversation is in progress.

Because such a response arrives on the same socket as everything else, attach metadata to correlate it. LocalAI echoes the map back verbatim on both response.created and response.done:

{"type": "response.create", "response": {
  "conversation": "none",
  "output_modalities": ["text"],
  "metadata": {"client_run": "abc123"},
  "input": [{"type": "message", "role": "user",
             "content": [{"type": "input_text", "text": "Is the oven still on?"}]}]
}}
{"type": "response.done", "response": {
  "id": "resp_...", "status": "completed",
  "metadata": {"client_run": "abc123"}
}}

Keys are strings up to 64 characters, values up to 512. A response created without metadata omits the field rather than sending an empty object.

Gating a realtime pipeline with voice recognition

A pipeline realtime model can require speaker verification before it responds. Add a voice_recognition block under pipeline. When present, each committed utterance is verified against authorized speakers; unauthorized utterances are dropped before the LLM runs (no LLM call, no tool execution, no TTS). The session stays open.

The same block also drives two optional, independent behaviors: an authorization gate (enforce) and speaker surfacing/personalization (identity). Set enforce: false to keep recognizing the speaker without ever rejecting a turn.

name: my-realtime
pipeline:
  vad: silero-vad
  transcription: whisper
  llm: qwen
  tts: kokoro
  voice_recognition:
    model: speaker-recognition   # the speaker-recognition backend model
    mode: identify               # "identify" (registry) or "verify" (references)
    threshold: 0.25              # cosine distance; <= passes
    enforce: true                # authorization gate (default true)
    when: every                  # "every" (default) or "first"
    on_reject: drop_event        # "drop_event" (default) or "drop_silent"
    anti_spoofing: false         # optional liveness check (verify mode)

    # identify mode: authorized registry identities (multiple persons)
    allow:
      names: ["alice", "bob"]    # match registered speaker names
      labels: ["family"]         # OR any identity carrying this label
      # empty allow = any registered speaker within threshold passes

    # verify mode: reference speakers (multiple persons)
    references:
      - name: alice
        audio: /models/voices/alice.wav
      - name: bob
        audio: /models/voices/bob.wav

Identifying speakers without gating

To recognize who is speaking and surface it to the client and the LLM without ever rejecting a turn, set enforce: false and add an identity block. The identity block works with or without the gate; when it is set, the speaker is resolved on every turn even if when: first.

name: my-realtime
pipeline:
  vad: silero-vad
  transcription: whisper
  llm: qwen
  tts: kokoro
  voice_recognition:
    model: speaker-recognition
    mode: identify
    threshold: 0.25
    # Authorization gate. Defaults to enforcing (rejects unauthorized speakers).
    # Set enforce:false to identify the speaker WITHOUT rejecting anyone.
    enforce: false
    when: every
    # Surface the recognized speaker to the client and the LLM. Works with or
    # without enforce; when set, identity is resolved on every turn even if
    # when:first.
    identity:
      announce: true            # emit the conversation.item.speaker event
      announce_unknown: false   # also emit it when there is no confident match
      personalize: true         # tell the LLM who is speaking
      inject_name: true         # set the per-message OpenAI name field
      inject_system_note: true  # append a "current speaker" line to the system message
      note_unknown: false       # append a "speaker is unknown" note when unidentified
Field Meaning
model Speaker-recognition backend model name.
mode identify matches against speakers registered via /v1/voice/register; verify matches against the references audios.
threshold Maximum cosine distance that still counts as a match (default ~0.25).
enforce Authorization gate. true (or omitted) rejects unauthorized speakers (the gating behavior above). false resolves and surfaces the speaker without ever dropping a turn.
when every verifies each utterance; first verifies once then trusts the session. When an identity block is set, the speaker is still resolved on every turn even with first.
on_reject drop_event drops and emits a speaker_not_authorized error event; drop_silent drops quietly.
anti_spoofing Verify mode only: runs the backend liveness check (slower).
allow.names / allow.labels identify mode: which registry identities are authorized. Empty = any registered speaker.
references verify mode: authorized reference speakers; the utterance passes if it matches any.
identity.announce Emit the conversation.item.speaker event to the client (see below).
identity.announce_unknown Also emit that event when there is no confident match. By default the event is emitted only on a match.
identity.personalize Inform the LLM who is speaking.
identity.inject_name Set the per-message OpenAI name field on each user turn.
identity.inject_system_note Append a The current speaker is <Name>. line to the system message.
identity.note_unknown When unidentified, append The current speaker is unknown. (lets the model ask who it is talking to).

identify mode requires the voice registry (speakers registered through /v1/voice/register). verify mode needs no registry: reference audios are embedded once at model load.

The conversation.item.speaker event

When identity.announce is enabled, the server emits a conversation.item.speaker event after the user conversation item, naming the recognized speaker:

{
  "type": "conversation.item.speaker",
  "item_id": "item_abc",
  "speaker": { "name": "Jeremy", "id": "spk_1", "labels": { "role": "owner" }, "confidence": 92.0, "distance": 0.1, "matched": true }
}

confidence is a 0-100 score, distance is the cosine distance, and matched is true when a confident match was found. labels carries any labels attached to the registered speaker (identify mode); it is omitted when the speaker has none. The name and id fields are omitted when empty. By default the event is emitted only on a match; set identity.announce_unknown: true to also emit it (with matched: false) when no speaker is identified.

This event is a LocalAI extension to the OpenAI Realtime API and is server-emitted only. Standard OpenAI Realtime clients ignore event types they do not recognize, so enabling it is non-breaking.

Examples

  • Realtime voice assistant demo (Go): a minimal Go client for the Realtime (WebSocket) API with a full talk-back voice loop and an example tool call. Ships a docker compose setup that brings up a realtime-capable LocalAI for you.
  • Realtime voice assistant example (Python): thin-client architecture (Silero VAD on the client, heavy lifting on LocalAI), suited to running the client on a Raspberry Pi.