mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-10 15:52:29 -04:00
* fix(realtime): keep VAD-commit transcription alive across barge-in, cancel it at teardown, order commits
Barge-in (new speech onset) cancels the turn's SourceVAD response context
(realtime_turncoord.go: respSink.cancel(SourceVAD)). The VAD commit body
runs under that same context, so an in-flight Whisper STT call was aborted
with 'context canceled' whenever the caller kept talking while the first
chunk was transcribing. The user's turn was lost: no transcript, no
LLM/TTS response.
v1 of this fix ran the transcription with context.WithoutCancel(ctx). The
review correctly pointed out two correctness gaps:
1. Teardown lost its cancellation. WithoutCancel detaches from every
cancellation, so a transcription in flight at session close outlived
the session and blocked respSink.shutdown (which joins the response
goroutines) until the backend finished the job.
2. Out-of-order commits. Consecutive commits run in parallel goroutines,
so a fast second transcription could append its user item before a
slow first one: the conversation became [second, first] and the second
response saw only [second].
Changes (core/http/endpoints/openai/):
- Session gains a session-lifetime context (sessionCtx), cancelled by
conncoord's Teardown BEFORE respSink.shutdown joins the response
goroutines. The transcription (and the voice-gate resolution) run under
it: they survive barge-in (which cancels only the per-response context)
but are cancelled with the session.
- Commit slots order the user-item appends in speech order:
Session.nextCommitSlot() is claimed at commit issue time (VAD CommitTurn
/ client commit), a commit's item append waits on the previous slot's
done (aborts on the session context), and every exit closes the slot so
a failed or torn-down commit never blocks the next. Transcriptions stay
parallel; only the appends are ordered.
- If the turn's response context was cancelled while the (detached)
transcription ran — barge-in, superseded by a newer commit — the user
item still commits (appendUserItem, split out of generateResponse) so
the LLM context keeps the full user input, but no response is generated
for the superseded turn; the newer speech triggers its own response on
the complete history.
- Regression tests (realtime_commit_order_test.go) cover both review
schedules — teardown during an in-flight transcription, and
held-first/finished-second out-of-order completion — plus the
barge-in-during-transcription item survival, driving the real commit
path with a transcription double that honours context cancellation.
- docs/design/realtime-state-machines.md: implementation-status entry for
the committed-turn pipeline (transcription lifetime + commit order).
Fixes #12445
Validated: builds, go vet clean, all openai specs + respcoord/turncoord/
conncoord suites pass under -race (incl. the 3 new regression specs).
Production A/B (call-center voice agent, SIP, silero-vad +
whisper-large-turbo + LLM + TTS, server_vad ~600 ms) on LocalAI v4.11.0:
unpatched — 'transcription_failed: context canceled', first part of the
utterance lost, agent answers only the remainder; patched — full
transcript committed, agent answers the complete utterance, barge-in
still cancels the in-flight assistant TTS response as intended, and
teardown cancels the in-flight transcription instead of waiting for the
backend.
Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>
* fix(realtime): release commit slots in order on every exit; share slot+issue boundary
Follow-up to the review of c8f991be (issue #12445): two schedules still
broke commit ordering.
1. A failed middle commit released later turns before earlier turns
finished. slot.done closed on every return, but the wait on
slot.prevDone happened only on the successful nonempty-transcript
path. Hold transcription A, let B fail (or return an empty
transcript / be rejected by the voice gate), then complete C: B
closed its channel without waiting for A, so C appended and started
its response without A (response history [third], final history
[third first]). Fix: the slot now releases (done closes) only AFTER
the predecessor has finished — on EVERY exit path, including errors,
empty transcripts, gate rejections and teardown (the session context
can still stop the wait, so teardown never blocks on a
never-finishing predecessor). The success path keeps its append gate
(wait before appending the user item); the deferred release gate
enforces the same order on every other exit.
2. Slot order and response issue order could disagree between the two
producers. The VAD CommitTurn and the client
input_audio_buffer.commit reserved the slot and called
respSink.issue separately; a pause between the two let the other
producer reserve AND issue first, so the later issue superseded the
EARLIER turn's response (response history [first], final history
[first second], second turn un-answered). Fix: both producers now go
through Session.issueCommit, which claims the slot and issues the
body under one lock (commitOrderMu) — slot order == issue order.
respSink.issue is non-blocking, so the lock never stalls
VAD/barge-in handling.
Regression tests (realtime_commit_order_test.go) now drive the REAL
issue path — Session.issueCommit into the real responseSink/respcoord,
so coordinator supersession and the spawned response goroutines are
exercised — and cover: teardown during an in-flight transcription;
held-first/finished-second out-of-order completion; barge-in
(respSink.cancel) item survival; a FAILED middle commit; an EMPTY
middle commit; interleaved VAD/client producers in both directions.
The failed/empty middle specs fail deterministically without the
release gate (verified against the pre-fix code).
docs/design/realtime-state-machines.md: implementation-status entry
updated (append gate + release gate + shared issue boundary).
Validated: builds, go vet clean, all openai specs + respcoord/turncoord/
conncoord suites pass under -race (476 specs, incl. the 4 new ones).
Fixes #12445
Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>
---------
Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>
Co-authored-by: nexxtmobile.de <kai@nexxtmobile.de>