mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-10 15:52:29 -04:00
Barge-in (new speech onset) cancels the turn's SourceVAD response context (realtime_turncoord.go: respSink.cancel(SourceVAD)). The VAD commit body runs under that same context, so an in-flight Whisper STT call was aborted with 'context canceled' whenever the caller kept talking while the first chunk was transcribing. The user's turn was lost: no transcript, no LLM/TTS response. v1 of this fix ran the transcription with context.WithoutCancel(ctx). The review correctly pointed out two correctness gaps: 1. Teardown lost its cancellation. WithoutCancel detaches from every cancellation, so a transcription in flight at session close outlived the session and blocked respSink.shutdown (which joins the response goroutines) until the backend finished the job. 2. Out-of-order commits. Consecutive commits run in parallel goroutines, so a fast second transcription could append its user item before a slow first one: the conversation became [second, first] and the second response saw only [second]. Changes (core/http/endpoints/openai/): - Session gains a session-lifetime context (sessionCtx), cancelled by conncoord's Teardown BEFORE respSink.shutdown joins the response goroutines. The transcription (and the voice-gate resolution) run under it: they survive barge-in (which cancels only the per-response context) but are cancelled with the session. - Commit slots order the user-item appends in speech order: Session.nextCommitSlot() is claimed at commit issue time (VAD CommitTurn / client commit), a commit's item append waits on the previous slot's done (aborts on the session context), and every exit closes the slot so a failed or torn-down commit never blocks the next. Transcriptions stay parallel; only the appends are ordered. - If the turn's response context was cancelled while the (detached) transcription ran — barge-in, superseded by a newer commit — the user item still commits (appendUserItem, split out of generateResponse) so the LLM context keeps the full user input, but no response is generated for the superseded turn; the newer speech triggers its own response on the complete history. - Regression tests (realtime_commit_order_test.go) cover both review schedules — teardown during an in-flight transcription, and held-first/finished-second out-of-order completion — plus the barge-in-during-transcription item survival, driving the real commit path with a transcription double that honours context cancellation. - docs/design/realtime-state-machines.md: implementation-status entry for the committed-turn pipeline (transcription lifetime + commit order). Fixes #12445 Validated: builds, go vet clean, all openai specs + respcoord/turncoord/ conncoord suites pass under -race (incl. the 3 new regression specs). Production A/B (call-center voice agent, SIP, silero-vad + whisper-large-turbo + LLM + TTS, server_vad ~600 ms) on LocalAI v4.11.0: unpatched — 'transcription_failed: context canceled', first part of the utterance lost, agent answers only the remainder; patched — full transcript committed, agent answers the complete utterance, barge-in still cancels the in-flight assistant TTS response as intended, and teardown cancels the in-flight transcription instead of waiting for the backend. Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>