nexxtmobile.deandnexxtmobile.de 4d681b7f6d fix(realtime): detach VAD-commit transcription from barge-in cancellation (#12446)
* fix(realtime): keep VAD-commit transcription alive across barge-in, cancel it at teardown, order commits

Barge-in (new speech onset) cancels the turn's SourceVAD response context
(realtime_turncoord.go: respSink.cancel(SourceVAD)). The VAD commit body
runs under that same context, so an in-flight Whisper STT call was aborted
with 'context canceled' whenever the caller kept talking while the first
chunk was transcribing. The user's turn was lost: no transcript, no
LLM/TTS response.

v1 of this fix ran the transcription with context.WithoutCancel(ctx). The
review correctly pointed out two correctness gaps:

1. Teardown lost its cancellation. WithoutCancel detaches from every
   cancellation, so a transcription in flight at session close outlived
   the session and blocked respSink.shutdown (which joins the response
   goroutines) until the backend finished the job.
2. Out-of-order commits. Consecutive commits run in parallel goroutines,
   so a fast second transcription could append its user item before a
   slow first one: the conversation became [second, first] and the second
   response saw only [second].

Changes (core/http/endpoints/openai/):
- Session gains a session-lifetime context (sessionCtx), cancelled by
  conncoord's Teardown BEFORE respSink.shutdown joins the response
  goroutines. The transcription (and the voice-gate resolution) run under
  it: they survive barge-in (which cancels only the per-response context)
  but are cancelled with the session.
- Commit slots order the user-item appends in speech order:
  Session.nextCommitSlot() is claimed at commit issue time (VAD CommitTurn
  / client commit), a commit's item append waits on the previous slot's
  done (aborts on the session context), and every exit closes the slot so
  a failed or torn-down commit never blocks the next. Transcriptions stay
  parallel; only the appends are ordered.
- If the turn's response context was cancelled while the (detached)
  transcription ran — barge-in, superseded by a newer commit — the user
  item still commits (appendUserItem, split out of generateResponse) so
  the LLM context keeps the full user input, but no response is generated
  for the superseded turn; the newer speech triggers its own response on
  the complete history.
- Regression tests (realtime_commit_order_test.go) cover both review
  schedules — teardown during an in-flight transcription, and
  held-first/finished-second out-of-order completion — plus the
  barge-in-during-transcription item survival, driving the real commit
  path with a transcription double that honours context cancellation.
- docs/design/realtime-state-machines.md: implementation-status entry for
  the committed-turn pipeline (transcription lifetime + commit order).

Fixes #12445

Validated: builds, go vet clean, all openai specs + respcoord/turncoord/
conncoord suites pass under -race (incl. the 3 new regression specs).
Production A/B (call-center voice agent, SIP, silero-vad +
whisper-large-turbo + LLM + TTS, server_vad ~600 ms) on LocalAI v4.11.0:
unpatched — 'transcription_failed: context canceled', first part of the
utterance lost, agent answers only the remainder; patched — full
transcript committed, agent answers the complete utterance, barge-in
still cancels the in-flight assistant TTS response as intended, and
teardown cancels the in-flight transcription instead of waiting for the
backend.

Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>

* fix(realtime): release commit slots in order on every exit; share slot+issue boundary

Follow-up to the review of c8f991be (issue #12445): two schedules still
broke commit ordering.

1. A failed middle commit released later turns before earlier turns
   finished. slot.done closed on every return, but the wait on
   slot.prevDone happened only on the successful nonempty-transcript
   path. Hold transcription A, let B fail (or return an empty
   transcript / be rejected by the voice gate), then complete C: B
   closed its channel without waiting for A, so C appended and started
   its response without A (response history [third], final history
   [third first]). Fix: the slot now releases (done closes) only AFTER
   the predecessor has finished — on EVERY exit path, including errors,
   empty transcripts, gate rejections and teardown (the session context
   can still stop the wait, so teardown never blocks on a
   never-finishing predecessor). The success path keeps its append gate
   (wait before appending the user item); the deferred release gate
   enforces the same order on every other exit.

2. Slot order and response issue order could disagree between the two
   producers. The VAD CommitTurn and the client
   input_audio_buffer.commit reserved the slot and called
   respSink.issue separately; a pause between the two let the other
   producer reserve AND issue first, so the later issue superseded the
   EARLIER turn's response (response history [first], final history
   [first second], second turn un-answered). Fix: both producers now go
   through Session.issueCommit, which claims the slot and issues the
   body under one lock (commitOrderMu) — slot order == issue order.
   respSink.issue is non-blocking, so the lock never stalls
   VAD/barge-in handling.

Regression tests (realtime_commit_order_test.go) now drive the REAL
issue path — Session.issueCommit into the real responseSink/respcoord,
so coordinator supersession and the spawned response goroutines are
exercised — and cover: teardown during an in-flight transcription;
held-first/finished-second out-of-order completion; barge-in
(respSink.cancel) item survival; a FAILED middle commit; an EMPTY
middle commit; interleaved VAD/client producers in both directions.
The failed/empty middle specs fail deterministically without the
release gate (verified against the pre-fix code).

docs/design/realtime-state-machines.md: implementation-status entry
updated (append gate + release gate + shared issue boundary).

Validated: builds, go vet clean, all openai specs + respcoord/turncoord/
conncoord suites pass under -race (476 specs, incl. the 4 new ones).

Fixes #12445

Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>

---------

Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>
Co-authored-by: nexxtmobile.de <kai@nexxtmobile.de>
2026-10-05 01:27:40 +02:00
2026-04-08 19:23:16 +02:00
2025-02-15 18:17:15 +01:00
2023-05-04 15:01:29 +02:00




LocalAI License

Follow LocalAI_API Join LocalAI Discord Community

mudler%2FLocalAI | Trendshift

Deutsch | Español | français | 日本語 | 한국어 | Português | Русский | 中文

LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.

A small core, not a bundle. Each backend wraps a best-in-class engine (llama.cpp, vLLM, whisper.cpp, stable-diffusion, MLX...) in its own image, pulled only when a model needs it. You install nothing you don't use.

  • Composable by design: backends are separate and pulled on demand, so you install only what your model needs
  • Open and extensible: load any model, or build your own backend in any language against an open interface
  • Drop-in API compatibility: OpenAI, Anthropic, and ElevenLabs APIs across every backend
  • Any model, any modality: LLMs, vision, voice, image, and video behind one API
  • Any hardware: NVIDIA, AMD, Intel, Apple Silicon, Vulkan, or CPU-only
  • Multi-user ready: API key auth, user quotas, role-based access
  • Built-in AI agents: autonomous agents with tool use, RAG, MCP, and skills
  • Privacy-first: your data never leaves your infrastructure

A small LocalAI core with backends (llama.cpp, vLLM, MLX, whisper.cpp, stable-diffusion, kokoro, parakeet.cpp...) plugged in as separate on-demand images

Created by Ettore Di Giacinto and maintained by the LocalAI team.

📖 Documentation | 💬 Discord | 💻 Quickstart | 🖼️ Models | ❓FAQ

Guided tour

https://github.com/user-attachments/assets/08cbb692-57da-48f7-963d-2e7b43883c18

Click to see more!

User and auth

https://github.com/user-attachments/assets/228fa9ad-81a3-4d43-bfb9-31557e14a36c

Agents

https://github.com/user-attachments/assets/6270b331-e21d-4087-a540-6290006b381a

Usage metrics per user

https://github.com/user-attachments/assets/cbb03379-23b4-4e3d-bd26-d152f057007f

Fine-tuning and Quantization

https://github.com/user-attachments/assets/5ba4ace9-d3df-4795-b7d4-b0b404ea71ee

WebRTC

https://github.com/user-attachments/assets/ed88e34c-fed3-4b83-8a67-4716a9feeb7b

Quickstart

macOS

Download LocalAI for macOS

Note: The DMG is not signed by Apple. After installing, run: sudo xattr -d com.apple.quarantine /Applications/LocalAI.app. See #6268 for details.

Containers (Docker, podman, ...)

Already ran LocalAI before? Use docker start -i local-ai to restart an existing container.

CPU only:

docker run -ti --name local-ai -p 8080:8080 localai/localai:latest

NVIDIA GPU:

# CUDA 13
docker run -ti --name local-ai -p 8080:8080 --gpus all localai/localai:latest-gpu-nvidia-cuda-13

# CUDA 12
docker run -ti --name local-ai -p 8080:8080 --gpus all localai/localai:latest-gpu-nvidia-cuda-12

# NVIDIA Jetson ARM64 (CUDA 12, for AGX Orin and similar)
docker run -ti --name local-ai -p 8080:8080 --gpus all localai/localai:latest-nvidia-l4t-arm64

# NVIDIA Jetson ARM64 (CUDA 13, for DGX Spark)
docker run -ti --name local-ai -p 8080:8080 --gpus all localai/localai:latest-nvidia-l4t-arm64-cuda-13

AMD GPU (ROCm):

docker run -ti --name local-ai -p 8080:8080 --device=/dev/kfd --device=/dev/dri --group-add=video localai/localai:latest-gpu-hipblas

Intel GPU (oneAPI):

docker run -ti --name local-ai -p 8080:8080 --device=/dev/dri/card1 --device=/dev/dri/renderD128 localai/localai:latest-gpu-intel

Vulkan GPU:

docker run -ti --name local-ai -p 8080:8080 localai/localai:latest-gpu-vulkan

Loading models

# From the model gallery (see available models with `local-ai models list` or at https://models.localai.io)
local-ai run llama-3.2-1b-instruct:q4_k_m
# From Huggingface
local-ai run huggingface://TheBloke/phi-2-GGUF/phi-2.Q8_0.gguf
# From the Ollama OCI registry
local-ai run ollama://gemma:2b
# From a YAML config
local-ai run https://gist.githubusercontent.com/.../phi-2.yaml
# From a standard OCI registry (e.g., Docker Hub)
local-ai run oci://localai/phi-2:latest

To work with a running LocalAI server from the terminal, start the built-in agent from another shell. It answers questions, reads your files and runs commands on your machine, asking you to approve anything that changes state. Inside a session, /models lists installed models and /model <name> switches between them. See the Terminal agent docs.

# Terminal 1
local-ai run llama-3.2-1b-instruct:q4_k_m

# Terminal 2
local-ai chat --model llama-3.2-1b-instruct:q4_k_m

Automatic Backend Detection: LocalAI automatically detects your GPU capabilities and downloads the appropriate backend. For advanced options, see GPU Acceleration.

For more details, see the Getting Started guide.

Latest News

  • October 2026: LocalAI 4.11.0 — Audio scenes with transcription, speaker labels, and sound detection. Also the Decisions API (/v1/systemone), model failover chains, signed OCI galleries, and Kimodo text-to-animation. Release notes
  • September 2026: LocalAI 4.10.0 — A fleet operations dashboard, one credentials file for private model sources, and local-ai benchmark for measuring model latency and throughput. Release notes
  • August 2026: LocalAI 4.9.0 — Deny-by-default authentication, opt-in chat context compression, unified model and backend management, and MiniMax-H3 video generation with audio. Release notes
  • August 2026: LocalAI 4.8.0 — Introducing vllm.cpp (alpha development builds), image-to-3D generation with a built-in viewer, and audio.cpp for speech, transcription, VAD, diarization, source separation and sound generation. Release notes
  • July 2026: LocalAI 4.7.0 — Managed voice cloning profiles, local video and talking-avatar generation, and interleaved reasoning with tool calls. Release notes
  • July 2026: LocalAI 4.6.0 — AMD ROCm fixes, predictable realtime pipeline warmup, conversation forking, and more reliable distributed model loading. Release notes
  • June 2026: LocalAI 4.5.0 — Native depth estimation and sound-event detection, NER-based PII filtering, and speaker-aware realtime conversations with automatic history compaction. Release notes
Older news

For older news and full release notes, see GitHub Releases and the blog.

Features

Supported Backends & Acceleration

LocalAI supports 60+ backends including llama.cpp, vLLM, SGLang, transformers, whisper.cpp, diffusers, MLX, MLX-VLM, and many more. Hardware acceleration is available for NVIDIA (CUDA 12/13), AMD (ROCm), Intel (oneAPI/SYCL), Apple Silicon (Metal), Vulkan, and NVIDIA Jetson (L4T). All backends can be installed on-the-fly from the Backend Gallery.

See the full Backend & Model Compatibility Table and GPU Acceleration guide.

Backends built by us

Most backends wrap a best-in-class upstream engine. A handful of them are native C/C++/GGML engines (no Python at inference) developed and maintained by the LocalAI project itself:

Backend What it does
vllm.cpp From-scratch C++20 port of vLLM for text generation: paged KV cache, continuous batching, prefix caching, safetensors + GGUF loading, engine-enforced structured output, on CPU, CUDA, Metal and Vulkan. Also serves MiniMax-H3 joint video+audio generation
parakeet.cpp C++/GGML port of NVIDIA NeMo Parakeet ASR (tdt/ctc/rnnt/hybrid), with cache-aware streaming transcription
moss-transcribe.cpp C++/GGML port of OpenMOSS MOSS-Transcribe-Diarize: joint long-form transcription, speaker diarization and timestamping in a single pass
moss-tts.cpp C++/GGML port of the OpenMOSS MOSS-TTS family: text-to-speech (MOSS-TTS-Local v1.5, 48 kHz stereo) with reference-audio voice cloning, through the MOSS-Audio-Tokenizer neural codec
magpie-tts.cpp C++/GGML port of NVIDIA's Magpie TTS Multilingual 357M: 22.05 kHz mono text-to-speech in 5 voices and 9+ languages, with the NanoCodec neural codec and tokenizer/G2P embedded in a single GGUF
ced.cpp C++/GGML port of the CED audio-tagging models: sound-event classification (527-class AudioSet) over REST and the realtime API for live recognition
voice-detect.cpp Speaker recognition and voice analysis (ECAPA-TDNN, WeSpeaker, ERes2Net, CAM++, wav2vec2 age/gender/emotion), replacing the Python speaker-recognition backend
voxtral-tts.c Mistral Voxtral-4B-TTS text-to-speech in pure C: 20 preset voices across 9 languages, 24 kHz WAV output, no dependencies beyond libc
vibevoice.cpp Native port of Microsoft VibeVoice for TTS (voice cloning) and long-form ASR with speaker diarization
rf-detr.cpp Native RF-DETR object detection and instance segmentation
locate-anything.cpp Open-vocabulary object detection and visual grounding (LocateAnything-3B)
depth-anything.cpp Depth Anything 3 monocular metric depth + camera pose estimation
face-detect.cpp Face detection, recognition, demographics and anti-spoofing (SCRFD/ArcFace, YuNet/SFace), replacing the Python insightface backend
free-splatter.cpp Pose-free 3D reconstruction (FreeSplatter): turns a handful of plain photos into 3D Gaussians, no camera poses or GPU required
trellis2.cpp C++/GGML port of Microsoft TRELLIS.2: single-image to textured 3D mesh (GLB with PBR materials)
kimodo.cpp C++/GGML text-to-motion on CPU and Vulkan, exported as animated skeleton GLB
privacy-filter.cpp Standalone GGML PII/NER token-classification engine powering LocalAI's PII redaction tier
LocalVQE Joint acoustic echo cancellation, noise suppression, and dereverberation
local-store Local-first vector database for embeddings (shipped in-tree)

We also maintain apex-quant, a per-tensor, per-layer quantization recipe for Mixture-of-Experts models that exploits their structural sparsity to produce GGUFs matching or beating Q8_0 quality - and they run out of the box on stock llama.cpp.

Resources

Team

LocalAI is maintained by a small team of humans, together with the wider community of contributors.

A huge thank you to everyone who contributes code, reviews PRs, files issues, and helps users in Discord — LocalAI is a community-driven project and wouldn't exist without you. See the full contributors list.

Citation

If you utilize this repository, data in a downstream project, please consider citing it with:

@misc{localai,
  author = {Ettore Di Giacinto},
  title = {LocalAI: The free, Open source OpenAI alternative},
  year = {2023},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/go-skynet/LocalAI}},

Sponsors

Do you find LocalAI useful?

Support the project by becoming a backer or sponsor. Your logo will show up here with a link to your website.

A huge thank you to our generous sponsors who support this project covering CI expenses, and our Sponsor list:

Past sponsors


Individual sponsors

A special thanks to individual sponsors, a full list is on GitHub and buymeacoffee. Special shout out to drikster80 for being generous. Thank you everyone!

License

LocalAI is a community-driven project created by Ettore Di Giacinto and maintained by the LocalAI team.

MIT - Author Ettore Di Giacinto mudler@localai.io

Acknowledgements

LocalAI couldn't have been built without the help of great software already available from the community. Thank you!

Contributors

This is a community project, a special thanks to our contributors!

S
Description
No description provided
Readme MIT
180 MiB
0 Stars 1 Watchers 0 Forks
Languages
Go 66.6%
JavaScript 15.2%
Python 4.4%
C++ 4.4%
HTML 3.1%
Other 6.2%