mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-03 11:34:35 -04:00
4bc0e92a3dea09dca9d9b87f343f795818bcf2fb
477
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4bc0e92a3d |
chore: merge master into distributed transport PR
Bring the distributed branch onto current master before the CI fix. Assisted-by: Codex:gpt-6 |
||
|
|
5613572f38 |
feat(systemone): validate requests and align the docs with Ollama's contract
All three routes now validate the request before it reaches a model: body size (413 over 64 KiB), state, question count, blank ids, option and level counts, and noul criteria keys. Forwarded decision requests skipped this before, so a malformed question surfaced as a backend error. The docs claimed the wire shape matches Ollama's. Field names and question types do; confidence, error shape, keep_alive and state rendering differ, and the docs now say so. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
70ce62901f |
refactor: name the capability decisions instead of systemone
The usecase describes what a model can do, and the category is the Decisions API. SystemOne stays as the wire contract: the /v1/systemone routes, the Score RPC question_type and the swagger tag are unchanged. The usecase, flag, auth feature, UI label, gallery tags and docs page are now decisions. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
b3d65fd538 |
fix(systemone): route NER models to the NER path and refuse decision models on permute and separate
vllm_decide refuses NER architectures and the NER entry point refuses decision architectures, so each model kind 500ed on half of the routes. A token_classify model now goes to the NER path on /v1/systemone, and /permute and /separate return 400 for decision models. Docs and instructions state which kind serves which route. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
b4852d62d2 |
feat(systemone): register the decisions API on auth and instructions
Adds a default-on systemone route feature for the three /v1/systemone routes and an /api/instructions area for them. No MCP tool is added: the endpoints run inference and are not admin install/edit actions, and the route-map test still passes. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
e03cf8dac8 |
feat(systemone): refuse models that do not declare the usecase
A chat-only model now gets a 400 naming known_usecases: [systemone] instead of a backend error. Configs declaring no usecases and token_classify models stay allowed so existing laya and GLiNER setups keep working. Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
2fa36e7147 |
feat(parakeet-cpp): speaker diarization, sound detection and live scene events (#12335)
* feat(parakeet-cpp): load diarization and CED models and companions
Repin PARAKEET_VERSION to parakeet.cpp PR #75's head, which adds
parakeet_capi_model_kind (ABI v8). Bind the new diarization, sound
event and combined scene stream C symbols through the same
purego.Dlsym probe pattern already used for the batched JSON entry
point, so the backend still loads against an older libparakeet.so.
Load now classifies the loaded GGUF by role (ASR, diarization or
sound) via parakeet_capi_model_kind and can load up to two companion
models from Options[] (asr_model:, diarization_model:, sound_model:,
paths resolved against opts.ModelPath), verifying each companion's
kind and freeing every context opened so far on any failure. Free
releases the primary and every companion. AudioTranscription now
names the loaded role when it is not ASR instead of a generic model
not loaded error. The dynamic batcher starts only when an ASR context
ends up loaded, primary or companion.
This is groundwork only: the Diarize and SoundDetection RPCs and the
live scene stream that actually use these new roles land in later
commits.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): reset role fields on a failed companion load
loadRoles' freeLoaded only released the C contexts it had opened; it
left ctxPtr/diarCtx/tagCtx and companions pointing at those now-freed
contexts, so a later Free() on the same instance would double-free.
Zero all four alongside the CppFree calls.
Also route AudioTranscriptionStream and AudioTranscriptionLive through
notASRError when ctxPtr is unset but a diarization or sound model is
loaded, matching AudioTranscription: both used to return the generic
model-not-loaded error instead of naming the loaded role.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): add speaker diarization
Implement the Diarize RPC for the parakeet-cpp Go backend, wired to
Nemotron-3-Diarization through libparakeet.so's diarization C-API.
Plain diarization uses parakeet_capi_diarize_pcm; when include_text is
set and an ASR companion is loaded, parakeet_capi_transcribe_and_
diarize_json fills each segment's text instead. Speaker labels are the
decimal index, or "unknown" for -1 (no diarized speaker overlaps).
min_duration_off merges same-speaker segments across a short gap
before min_duration_on drops the segments still too short, then ids
are renumbered. num_speakers/min_speakers/max_speakers/clustering_
threshold have no Sortformer equivalent and are logged at debug
instead of rejected.
Verified against the real Nemotron-3-Diarization + parakeet-tdt_ctc-
110m checkpoints on the two_speakers.wav fixture: correct A-B-A-B
speaker segmentation and matching speaker-attributed transcripts.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): add sound event detection
Wire the SoundDetection RPC to the CED tagger context (p.tagCtx)
loaded by Task 1's role classification. It runs the whole clip
through a one-shot parakeet_capi_sound_stream_* session (window
10s, hop 10s, top_k set to the tagger's class count so every
drained window carries a full score list), averages each class's
score across the drained windows, sorts descending, then applies
the request's threshold and top_k (0 keeps every class).
No tagCtx returns FailedPrecondition; a libparakeet.so missing the
sound_stream symbols returns Unimplemented. Every C call runs under
engineMu, and the stream is always freed, even when a feed or drain
call fails partway through.
Verified against a real ced-tiny-q8_0.gguf on the rooster.wav demo
clip: "Chicken, rooster" tops the list at score 0.91.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): cancel sound detection mid-feed, shrink the lock
SoundDetection now checks ctx before each 10 s feed slice (mirroring
driver.go's feedSlices) and returns Canceled if the caller gave up,
so a long clip can be interrupted instead of feeding to completion
regardless. The stream is still freed on every path, cancellation
included.
Also narrow engineMu to the C calls: the drained JSON document is
now decoded after the lock is released, splitting soundStreamScores
into a locked soundStreamDrain (opts, begin, feed, drain, free) and
an unlocked json.Unmarshal.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): stream speaker and sound events during live transcription
Add two additive proto fields, LiveSpeakerSegment and LiveSoundEvent,
repeated on TranscriptLiveResponse. When a diarization or sound
companion model is loaded, AudioTranscriptionLive now runs a no-ASR
scene stream (parakeet_capi_scene_stream_begin) beside the ASR
streaming session, feeding it the same PCM slices and forwarding any
closed speaker or sound events alongside the matching ASR delta, or
on their own when a slice has no ASR output.
The scene stream is freed and reopened on a mid-stream Config reset,
flushed with is_last before the closing FinalResult, and degrades
gracefully (a warning, not an error) when begin or a later feed call
fails, so live transcription keeps working ASR-only. Existing live
behavior is unchanged when no companion is configured, and no scene
C call is made in that case.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): keep scene events off the ASR critical path in live
Emit each slice's ASR result right after the ASR feed, before the
scene feed for that slice runs, so a companion diarization/sound
model never adds scene compute latency in front of the delta or
<EOU> that drives realtime turn detection. Closed speakers/sounds go
out afterward as their own response, so a slice with both now
produces two responses, ASR first. The live feed log line now
reports ASR and scene wall time separately.
Re-check the diarization/sound contexts a scene stream was begun
with against the live contexts before every feed, under the same
lock: Free() can race between an ASR feed and the matching scene
feed and free the model the stream borrows. A mismatch now returns
without touching the C side. Freeing the stream itself stays
unconditional; the scene stream's destructor only releases its own
buffers and never touches the borrowed contexts.
Also recover a panicking stub inside the live test goroutine instead
of crashing the test binary, and reset the live decode-lag tracker on
a mid-stream config reset, matching what its own comment already
promised.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(realtime): surface live speaker and sound events
Carry the backend's closed speaker segments and sound events
(TranscriptLiveResponse fields 7/8) through LiveTranscriptionEvent
as LiveSpeakerSegment/LiveSoundEvent (nanoseconds mapped to
seconds), and forward them from the semantic_vad live path.
Each speaker segment emits
conversation.item.input_audio_transcription.segment with speaker,
start, end and empty text under the turn's item id. Each sound
event emits conversation.item.sound_detection with one tag
(label, score = peak, index) and the event's new optional
start/end seconds fields, omitted when unset so the existing
unary/windowed sound-detection path is unaffected.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(realtime): keep start/end on a zero-second transcription segment
ConversationItemInputAudioTranscriptionSegmentEvent.Start/End used
omitempty, so a speaker segment starting at 0.0s dropped its
"start" key. Nothing emitted this event before the live scene-event
path, so drop omitempty: the segment always carries real times.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(gallery): add parakeet-cpp diarization, CED and realtime scene models
Add gallery entries for the new parakeet-cpp capabilities: standalone
Nemotron-3-Diarization, the same paired with the Parakeet TDT+CTC
110M ASR model for speaker-attributed text, CED-Tiny and CED-Base
sound classifiers, and a realtime scene bundle combining the
streaming EOU ASR model with diarization and sound companions.
SHA256 taken from the Hub API; licenses from each model card
(openmdw-1.1 for Nemotron-3-Diarization, apache-2.0 for CED,
cc-by-4.0 for the Parakeet ASR models).
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs: document parakeet-cpp diarization, sound detection and live scene events
Cover the new parakeet-cpp capabilities across the feature pages:
Nemotron-3-Diarization as a diarization backend (with and without
speaker text, the ignored speaker-count hints, the Sortformer
voice-like-sound quirk), CED as a sound classification backend, the
asr_model/diarization_model/sound_model/diarization_latency companion
options, and the realtime live speaker/sound events (event shapes,
the speech-turn-only limitation, and using this or
pipeline.sound_detection but not both).
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(gallery): correct the realtime-scene license and wording nits
parakeet-cpp-realtime-scene mistakenly copied cc-by-4.0 from the
existing realtime_eou_120m-v1 entry; the model card lists the NVIDIA
open model license instead. Switch to the gallery's usual spelling
for that license and keep the diarization/CED licenses called out in
the description.
Also: audio-diarization.md now says getting per-segment text needs
both an asr_model companion and include_text=true on the request, and
audio-to-text.md's option table reads "Use on" (a pairing the loader
does not enforce) instead of "Allowed on".
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): reject a companion role that duplicates the primary's
loadRoles let a companion option (asr_model:/diarization_model:/
sound_model:) assign into a role field the primary already occupied,
for example asr_model: on an already-ASR primary. The companion's
context silently overwrote ctxPtr/diarCtx/tagCtx, and Free() only
walks those three fields, so the original primary context was never
freed again.
Reject a companion whose role the primary already holds before its
GGUF is even loaded, freeing everything loadRoles opened so far, the
same way a wrong-kind companion is already rejected.
Also warn, rather than silently fall through, when
parakeet_capi_model_kind reports PARAKEET_MODEL_KIND_NONE for a
successfully loaded primary; the primary is still treated as ASR,
matching today's behavior.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): cap live scene sound score retention
sceneBegin started the live diarization/sound companion stream with
the C API's default sound options, whose top_k keeps 5 scores per
window forever until drained. The live scene path never drains sound
scores (only the offline SoundDetection RPC does, with its own fresh
stream), so this window queue on the C side grew for the whole
session's lifetime.
Set opts.Sound.TopK = 0 before starting the scene stream: this
disables score retention while leaving sound event detection (onset/
offset), which the live path actually consumes, unaffected.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): merge diarization segments per speaker, harden Diarize
mergeCloseSegments only compared neighbors in the single start-sorted
segment list, so two same-speaker segments never merged once another
speaker's turn fell between them (A, B, A): the short B segment broke
the adjacency the merge relied on. Group segments by speaker first,
merge within each speaker's own start-ordered run, then re-sort the
result by start so interleaved speakers come back out in timeline
order.
Also harden Diarize's entry points the same way streamFeedDoc/
sceneFeed already are: diarizeCall re-checks p.diarCtx (and, on the
include_text path, p.ctxPtr) under engineMu right before the C call,
so a Free() racing between Diarize's own checks and the lock can no
longer reach the C side with a freed context. When the include_text
call returns NULL, last_error is now read from both contexts and
whichever came back non-empty is reported, since either side of the
pairing can be the one that failed. A WAV decode failure is reported
as InvalidArgument instead of an unwrapped/untyped error.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): harden SoundDetection's engine checks
soundStreamDrain ran every C call under engineMu but never re-checked
p.tagCtx there, so a Free() racing between SoundDetection's own
tagCtx==0 check and this lock could still reach the C side with a
freed context. Re-check p.tagCtx under the lock and return
ModelNotLoaded when it was cleared, mirroring diarizeCall's own
re-check. A WAV decode failure is now reported as InvalidArgument
instead of an unwrapped/untyped error.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* test(parakeet-cpp): cover a mid-session scene feed failure
feedSlicesScene already degrades gracefully when a scene feed call
fails mid-session: it frees the broken stream and carries the ASR-only
session forward. Add a spec covering that path end to end: the scene
stream is freed exactly once, later audio slices still produce ASR
responses, and no speaker/sound events appear before or after the
failure.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs: fix the parakeet-cpp companion role table and realtime scene docs
audio-to-text.md's companion option table read "Use on" with a note
that the loader did not enforce the pairing; it now rejects a
companion whose role duplicates the primary's, so restore the
"Allowed on" wording and describe the real enforcement.
openai-realtime.md's live speaker/sound section claimed a mid-stream
session.update resets the companion stream and that it flushes on
session close; neither happens, since the realtime core opens one
live stream (and so one scene stream) per speech turn and closes it
at that turn's commit, with no mid-stream Config in between. Document
that lifecycle instead, state precisely that start/end are seconds
from the start of the turn's own audio, and note that the diarization
model starts a fresh session every turn, so a speaker index is only
meaningful within one turn. The example sound tag ("Rooster", index
17) did not match any real CED label; index 17 in ced-tiny-q8_0.gguf
is "Baby laughter". Replaced with "Chicken, rooster" at its real
index, 99.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): use CED's real index for Chicken, rooster
The scene feed comment and the live test's canned document gave
"Chicken, rooster" index 365. In CED's AudioSet label list it is 99,
which is also what the realtime docs show.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(realtime): call the test event accessor
The scene-event tests range over a method instead of its returned slice.
Call the synchronized accessor so the OpenAI test package compiles.
Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(parakeet-cpp): pin parakeet.cpp master with sound events
mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and
#74 (the missing <algorithm> include that broke the image builds) are
on master now. Pin 6dea76a instead of the #75 PR head, and update the
header comment the bump bot reads.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(parakeet-cpp): pin parakeet.cpp with ced.cpp on main
parakeet.cpp #76 moved its ced.cpp submodule from the head of
localai-org/ced.cpp#3 (a branch-only commit) to ced.cpp main, where
#3 landed with an identical tree. Pin 623a968 so the image builds no
longer depend on that branch.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(transcription): carry speaker labels on words and streamed segments
A diarizing backend could label transcript segments, but two paths
dropped the label: TranscriptWord had no speaker field, so live
transcription words and word-level timestamps could not carry one, and
the stream=true transcript.text.done event left the speaker out of
its segments.
TranscriptWord gains an optional speaker (proto field 4, additive).
It flows through the live event and result mapping, the JSON word
output of the endpoint and the CLI, and transcript.text.done now
includes a segment's speaker when there is one. Empty labels are
omitted, so responses without diarization are unchanged.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
(cherry picked from commit
|
||
|
|
c4ab987e3e |
chore: merge master into distributed transport PR
Preserve failover support alongside the PostgreSQL broadcast carrier. Assisted-by: Codex:gpt-6 golangci-lint |
||
|
|
9c156656bd |
Merge PR #12285: feat(failover): serve a model name from a chain of local and remote targets
Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
4151ceda85 |
chore: merge master into distributed transport PR
Keep the version-reporting import from master and omit the unused sanitize import after the transport changes. Assisted-by: Codex:GPT-6 |
||
|
|
590512d9eb |
feat(distributed): report worker version and show models in node inspector (#12328)
* ci: bump Hugo from 0.146.3 to 0.166.0 The hugo-theme-relearn submodule was bumped to 9.1.x in #12096, which requires Hugo >= 0.165.0. The pinned 0.146.3 broke the docs site build with a template error in alias.html that could not evaluate the Locale field on langs.Language. Bump HUGO_VERSION to 0.166.0 (latest stable) to satisfy the theme minimum and resolve the alias.html template error. Assisted-by: nib:claude-sonnet-4.5 [bash] [read] [edit] * feat(distributed): report worker version and show models in node inspector Workers now send their LocalAI build version and git commit at registration. The controller stores them on BackendNode and exposes them through the existing node list/detail API responses. The node inspector side pane now fetches and renders the list of loaded models (name, state, in-flight) instead of showing only a count, matching what the node detail page already displays. --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
733eeda123 |
chore: merge master into distributed transport PR
Keep the newer SQLite dependency from master to resolve the conflict. Assisted-by: Codex:gpt-6 |
||
|
|
f82efdb43b |
fix(models): fallback to application config default context size in /v1/models/capabilities (#12202) (#12216)
* fix(models): fallback to application config default context size (#12202) Honor appConfig.ContextSize in /v1/models/capabilities when model context_size is unset. * docs(models): explain context size fallback Describe the application default used by capability discovery and preserve the distinction between total context and per-request limits. Assisted-by: Codex:GPT-6 * fix(models): apply the default context size only when context_size is unset The request path applies the application default context size only when a model leaves context_size unset. An explicit 0 or -1 falls through to the backend fallback. The capabilities endpoint now does the same, so it reports the value the backend uses. Assisted-by: Claude:claude-opus-5-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
eec532704c |
fix(functions): honor function_arguments_key when building the tool grammar (#11677)
* fix(functions): honor function_arguments_key when building the tool grammar
All four call sites of `Functions.ToJSONStructure(name, args string)` pass
`FunctionsConfig.FunctionNameKey` as *both* arguments, so
`FunctionArgumentsKey` never reaches the grammar generator.
`ToJSONStructure` writes the two properties into the same map:
property[nameKey] = FunctionName{Const: function.Name}
property[argsKey] = Argument{...}
When `nameKey == argsKey` the second assignment overwrites the first, so a
model configured with `function_name_key` gets a grammar carrying only the
arguments object -- the `{"const": "<function name>"}` constraint is gone and
the grammar can no longer express which function was called.
With `function_name_key: function`, the generated property set collapses from
{"function": {"const": "get_weather"}, "arguments": {...}}
to
{"function": {"type": "object", "properties": {...}}}
Setting only `function_arguments_key` is equally broken in the other
direction: the grammar keeps emitting `arguments` while `ParseFunctionCall`
(pkg/functions/parse.go) looks up the configured key, so the parsed call comes
back with its arguments empty.
The default configuration is unaffected -- with both keys empty
`ToJSONStructure` falls back to `name`/`arguments` for both parameters, which
is why this went unnoticed.
The existing `ToJSONStructure()` unit test already calls the helper with two
distinct keys, so only the call sites were wrong. Extend that test with a case
that keeps both custom keys distinct and asserts the two properties survive.
Signed-off-by: Anai-Guo <antai12232931@outlook.com>
* test(functions): cover configured grammar keys
Route grammar construction through FunctionsConfig so the regression test
covers the key wiring used by every endpoint.
Assisted-by: Codex:gpt-5
* chore: empty commit to trigger workflow approval
Signed-off-by: Tai An <antai12232931@outlook.com>
---------
Signed-off-by: Anai-Guo <antai12232931@outlook.com>
Signed-off-by: Tai An <antai12232931@outlook.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
|
||
|
|
afebecc63d |
fix(ollama): report on-disk size for /api/tags and /api/ps (#11989)
Hardcoding size/size_vram as 0 made Ollama clients treat loaded models as free. Prefer ModelFileName+ModelPath Stat when available, and omit size_vram (and size) when the value is unknown instead of emitting literal zeros. Resolve each listed model by its stored ID so tagged variants use their own weights. Fixes #11969 Signed-off-by: lei_lei <imleilei123@gmail.com> |
||
|
|
f20cc16033 |
fix(openai): reject an unknown audio response_format as a 400 before the backend runs
Transcription and diarization checked response_format only after the backend had run, and returned a plain error for an unknown value. Failover counts a plain error as a target failure, so one request with a bad response_format ran the backend on every target of a chain and tripped all of them. Check the format before the backend runs and answer 400. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code] |
||
|
|
dae9a431e8 |
Merge remote-tracking branch 'origin/master' into feat/failover-chains
Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code] |
||
|
|
f154bd990a |
feat(system): report per-model DRM VRAM (#12026)
* feat(system): report per-model DRM VRAM Expose optional resident device memory for local backend process trees. Deduplicate DRM clients and omit unsupported or incomplete readings. Document accounting limits and preserve a measured zero in JSON. Closes #11970. Assisted-by: Codex:gpt-6 * fix(system): document trusted procfs reads Scope G304 annotations to paths built from the fixed procfs root, integer process IDs, and kernel directory entries. These reads accept no user-controlled path components. Assisted-by: Codex:GPT-6 gosec --------- Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
5794495a37 |
fix(responses): preserve streamed output items (#12048)
Keep each message and reasoning item at its announced output index. Include the answer in completed responses with reasoning or fallback function calls, and retain reasoning supplied through backend deltas. Add regression coverage for stream indices, final output, plain text, and automatic tool parsing. Assisted-by: Codex:GPT-6 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
0e52bb657e |
fix(responses): wait for complete JSON tool calls (#12001)
Partial JSON parsing heals a name-only chunk into a tool call. The stream emits that call with empty arguments and skips later chunks. Require complete JSON before emitting terminal tool-call events. Preserve complete calls before an unfinished trailing call, and count only actual tool calls. Add split-chunk regression tests and docs. Refs #11635. The non-streaming report remains unconfirmed. Assisted-by: Codex:GPT-6 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
27ceda46b0 |
refactor: name the realtime pipeline stages with constants
Review asked for constants instead of the stage literals ("vad",
"transcription", "llm", "tts", "sound_detection") passed to resolveStage,
stageCall and isChainStage, so the uses can be cross-checked. Add
PipelineStage* constants next to the Pipeline type in core/config: the
names match its yaml keys, and core/backend (preload roles) and the openai
realtime endpoint both need them.
Use them in realtime_model.go (stage routing and preload roles),
realtime.go (the voice_recognition preload role) and core/backend
preload.go. model_failover events take their stage from the stageChains
keys, so they now carry the constants too. The failover tests use the
constants for inputs and keep literal wire values in their event
assertions.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
14c692098e |
fix(openai): answer 400 for a missing or malformed audio upload
TranscriptEndpoint, DiarizationEndpoint and SoundClassificationEndpoint
returned the raw c.FormFile("file") error. Echo turns a non-HTTPError into
a 500, so a request with no multipart boundary or no file field looked like
a server fault. This became visible in tests/e2e once the suite registered
a transcription-capable model (lp-transcription) at runtime: the
default-model middleware then resolves a model, the request reaches the
handler, and "should return mocked transcription" got a 500 in some spec
orders (seed 1790493709).
Read the upload through a small uploadedFile helper that maps any
FormFile failure to 400 with the field name and parser reason. Server-side
failures after that (temp dir, file create, copy) stay 500. The image and
LocalAI upload endpoints already return 400 here.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
5ed67be0bf |
fix(failover): skip a target that answers HTTP 501 instead of failing the chain
The non-OpenAI endpoints (depth, detection, face_*, voice_*, images, video, 3d) map a backend's gRPC Unimplemented to an echo 501 without the gRPC status. The retry loop did not see a capability gap, and IsRetryable is false for 501, so the client got 501 and the next target was never tried. Treat a returned or written 501 as a capability gap. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
e62854c340 |
fix(failover): pass credential lookup from CLI
Pass the API key environment lookup through ApplicationConfig to satisfy core configuration lint. Keep credential resolution dynamic and exclude the callback from serialization. Handle the five close results reported by errcheck. Assisted-by: Codex:gpt-6 golangci-lint Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
132bb216a1 |
fix(failover): resolve chains in transcription and sound-only sessions
Transcription-only and sound-detection-only realtime sessions passed a chain config straight to the model loader. It has no backend, so the loader fell back to greedy backend auto-detection: slow, and ending in an unhelpful error. Sound-only sessions are a main use of chains. The stage routing of the full pipeline moves into a stageRouter that both realtime model kinds embed. Every stage resolves to the chain's active target at build time and goes through the failover plan per call. The session sends failover events for any model with chain stages, and restarts them when a transcription session.update swaps the model. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
2ecbaab010 |
fix(failover): send remote targets their own upstream model
A chain request reached a cloud-proxy target with the client's model, the chain name, whenever the target set no upstream_model: passthrough forwards the body's model and translate falls back to it. The upstream answered 404, which neither retries nor trips, while the liveness probe, which checks the target's own name, kept passing. PrepareTarget now sets the upstream model of a remote target to proxy.upstream_model or the target name, the same name the probe uses. The request pipeline and realtime chain stages both call it. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
191d2d6301 |
feat(failover): add MCP tools to list chains and pin targets
Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
ea55e2ffc3 |
feat(failover): switch realtime pipeline stages per call
A stage that names a chain is resolved on every call, so a switch keeps the session and its conversation. Clients get localai.model.failover events at session start and on every switch. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
2a1213f1e2 |
feat(failover): expose chain status, pins and events over REST and SSE
Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
f7c0bb5f4b |
feat(config): validate failover chain targets across configs
Reject chains whose targets are missing or are chains, at load and on create or edit, and warn when the targets share no usecase. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
7e14d196e1 |
fix(distributed): follow current backend protocol
Master added Animate3D, negative_prompt, and context_size after this branch diverged. The old suite did not exercise those paths, and Kokoros no longer implemented the generated service trait. Extend binary conformance across the tunnel owner and peer relay. Allow long development versions so rebased binaries can register in PostgreSQL. Clear the security findings introduced by the branch's new code. Assisted-by: Codex:GPT-5 [apply_patch] [exec_command] |
||
|
|
f49e94875f |
fix(distributed): harden binary conformance
Keep inpainting request inputs in private ephemeral storage and publish only the completed image. Extend the real-binary feature matrix with canonical remeshing and typed protocol assertions, and preserve multipart diarization language hints. Assisted-by: Codex:GPT-5 [apply_patch] [exec_command] |
||
|
|
fc31e69373 |
test(distributed): complete binary backend conformance
Exercise the remaining public backend routes through real frontend and worker binaries, and keep fixture staging limits scoped to this conformance deployment. Fix inpainting artifacts so returned generated-image URLs resolve through the configured static mount. Assisted-by: Codex:GPT-5 [apply_patch] [exec_command] |
||
|
|
2e63b9822a |
fix(distributed): preserve master protocol extensions
Carry prefix-cache pressure and reported residency over PostgreSQL, keep truthful backend stop replies, and retain capacity-aware S3 staging release over the worker control tunnel after the rebase. Adapt the overlapping master tests to the tunnel-native protocol and cover cross-replica cache events and departed-node cleanup. Assisted-by: Codex:GPT-5 [apply_patch] [exec_command] |
||
|
|
17e8d5b7d8 |
test(openai): drop the audio snapshot nothing reads
The mutex round the realtime transport double added a snapshot accessor for each
recorded slice. Only the event one has a caller, so make lint refuses the build:
realtime_doubles_test.go:64:25: func (*fakeTransport).recordedAudio is unused (unused)
No spec has ever read the audio log, before the mutex or after it, so the
accessor is deleted rather than nolinted and the struct comment says where the
next one comes from. audioLog stays written, because a double that silently
discarded what a coordinator sent it would be a different double.
Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
230b39eca3 |
test: fix the three data races that made -race runs noisy
None was introduced by this branch and all three are in test code, which is what made them survive: every suite passed on every run and only the race detector said otherwise. A known-failing -race run is worse than a noisy one, because a real race raised by production code lands in the same report and is read as one of these. galleryop: gatedModelManager guarded the recorded names and not the gate channel itself. A spec frees the parked worker by closing the gate and installing a fresh one, on the spec goroutine, while the worker goroutine reads the field to park on it. The channel is now read and replaced under the same mutex, and cleanup closes idempotently. pkg/model: two specs swapped xlog's package logger to capture output and swapped it back on cleanup. xlog.SetLogger writes an unsynchronised global, so the restore raced with the backend process watcher, which logs while a process is stopping; the captured bytes.Buffer was written by that goroutine and read by an Eventually at the same time. SetLogger is now called once for the whole test binary, from init, before a goroutine exists to race with, and a spec swaps the DESTINATION under a mutex through a routing slog.Handler. Per-spec level filtering is preserved deliberately: one of these specs asserts that a debug emission is filtered OUT and would pass vacuously against a handler that recorded everything. openai: fakeTransport appended to its event and audio logs from the response and turn coordinators' goroutines while a spec ranged over them. Both are behind a mutex and are read through snapshot accessors; the fields are renamed so a raw read from another spec file does not compile. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
d72b591a78 |
feat(cluster): make a peer prove which replica it is
GET /api/cluster/peer authenticated with the deployment's shared registration token and took the dialling replica's id from ?id= on trust. Every worker holds that token, so anything holding it could open a peer link as any replica: relay through it to every worker tunnel that replica owns, displace a real replica's inbound link by declaring its id, and point the roughly 31 GiB per-session receive window at one replica. Validating the id against the instances table does not fix this, because the attack declares a real replica's id. So the route now checks two credentials and needs both. The shared token still says the dialler belongs to this deployment; a new per-replica credential says which replica it is. The credential follows the per-node worker credential rather than inventing a second mechanism: crypto/rand.Text, stored only as a hex SHA-256, compared in constant time, with no fallback to the shared token. It differs in the stronger direction. A worker's credential is minted by the frontend and handed over once; a replica writes its own instances row, so it mints its own secret, publishes only the hash in the same statement that publishes its address, and never sends the plaintext anywhere but the peer dial. A peer that presents no credential is refused, not waved through. An old replica and an attacker holding the shared token send the same request, so accepting the first accepts the second; there is no safe downgrade here, only a quiet one. The refusal is made loud instead, on both sides, naming the upgrade rather than the network. On the documented frontend-first order a new replica still dials an old one; an old replica cannot dial a new one, which costs relayed requests that land on a not-yet-restarted replica and surfaces as no route, never as absence. A rejected peer gets its own sentinel, ErrPeerRejected, whose unwrap chain carries ErrPeerUnreachable as well and no absence sentinel at all. Keeping the older sentinel means no existing consumer changes behaviour; the cause stays out of the chain, so absence cannot escape through it and nothing can read an authorization failure as a worker that went away. One consequence beyond the fix: a replica with no advertised address has no instances row, so it now cannot dial out either. It was already unreachable inward. The startup error and the docs say so. Registry.Register, NewMembership, NewPeerPool, PeerHandler and RegisterClusterRoutes all gained required arguments, so the identity cannot be dropped without a compile failure. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
eb1656ed9c |
chore(distributed): take the nats-io modules out of the build
Distributed mode has not dialled a message broker since the control plane moved onto the workers' own outward tunnels and every fan-out family moved onto PostgreSQL LISTEN/NOTIFY. What was left was the dependency itself, and the code that existed only to feed it. Dropped from go.mod: nats-io/jwt/v2, nats-io/nats.go, nats-io/nkeys, nats-io/nuid and testcontainers-go/modules/nats, along with the fourteen indirect requires that only the NATS testcontainer pulled in. go.sum carries no nats line either, so the removal is not the partial kind where the require goes and the checksum stays. Deleted with them: pkg/natsauth in full, the broker client's remaining options and TLS files, the per-node JWT minting on both the register and the approve path, and the natsauth.Config parameter threaded through the node routes. The credential manager is renamed and stripped rather than deleted, because it still holds the tunnel token that every re-registration rotates. The bus flags stay accepted and ignored, and are now hidden, on every command that had them, so an existing unit file, compose file or Helm values file still starts on the day of the upgrade. What is not kept is the validation that REQUIRED one: a distributed frontend started with no bus URL is no longer fatal. The TLS paths lose type:"existingfile" deliberately, so a certificate deleted along with the broker cannot fail a startup. One operator-visible behaviour change: --nats-require-auth no longer makes an agent worker wait through admin approval. Ask for that wait with --distributed-require-auth, which already implied it. It is documented in the migration section and pinned from both sides. A deployment now needs PostgreSQL and the frontends' own HTTP listener, and nothing else. coverage-baseline.txt moves from 54.2 to 62.0. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
eb7ed9b910 |
refactor(distributed): delete MessagingClient and shrink the NATS client to fan-out
Nothing in the tree publishes, subscribes, queue-subscribes or requests through
the MessagingClient interface any more, so it is deleted rather than shrunk to
Broadcaster: two exported names for one method set in one package is an
invitation for the next author to pick whichever the surrounding file already
imported.
$ grep -rn 'messaging\.MessagingClient' --include='*.go' .
core/services/syncstate/syncstate.go:54: // It is messaging.Broadcaster rather than messaging.MessagingClient because
(one hit, a comment; no live referent. The naive grep in the plan also matches
prose and the local test type names fakeMessagingClient and
countingMessagingClient, so it can never be empty.)
*messaging.Client is shrunk to exactly Broadcaster plus its own lifecycle.
QueueSubscribe, QueueSubscribeReply, SubscribeReply, Request, Conn and the
package helpers QueueSubscribeJSON and RequestJSON go with it; none had a
production caller. Deleting the methods rather than only the call sites is what
makes putting a family back on this carrier a build error instead of a line that
compiles, publishes successfully, and is delivered onto a carrier the deployment
is being taken off. Conn is in that list because while it existed every other
name was one c.Conn().X() away; the flush-and-verdict that its real consumers
needed is now ConfirmRoundTrip, which keeps the NATS JWT permission specs armed.
The client, its options and its TLS plumbing are NOT deleted, and both processes
stay on the bus. agent.<name>.cancel is the one fan-out family that could not
move: its only subscriber is the agent worker, which has no database and cannot
join the PostgreSQL carrier at all, so a cancel published there would reach no
worker and be reported as sent. The frontend passes the client to
newFanoutBridges as its cancelCarrier and the worker subscribes on it, so
--nats-url stays required on agent-worker. Both go with the tunnel cancel verb.
The struct field is renamed Nats -> CancelCarrier to say what it is for, and
agentpool loses the messaging.Publisher it held only to be non-nil: it never
published on it, and it was gating whether a frontend runs agents distributed or
in an in-process pool. Retiring the bus would have flipped every replica back to
the in-process pool silently. The gate now reads the agent store, which is the
dependency the mode actually requires.
Also deletes four subject builders with no production publisher
(SubjectFineTuneProgress, SubjectFineTuneCancel, SubjectCacheInvalidateSkills,
SubjectCacheInvalidateCollection), the queue and request/reply halves of the
shared test double, and the e2e specs that were their only callers. Every
surviving subject is now pinned to its exact literal, because a subject is a
cross-version wire format and a rename that looks internal stops half a fleet
hearing the other half.
Docs: distributed-mode.md and cli-reference.md no longer claim NATS carries the
agent-worker job subjects, the frontend's cross-replica events, or an agent
worker's real work.
Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
978f7922bd |
feat(distributed): move the last nine fan-out families onto PostgreSQL
Gallery progress and cancel, the operation cache's start and end, the model
and backend cache invalidations, staging progress, and the prefix cache's
observations and invalidations now travel on the LISTEN/NOTIFY carrier. No
subject is published or subscribed on messaging.Client anywhere in the tree,
which is what makes retiring that package a deletion rather than a migration:
$ grep -rn 'natsClient\.Publish\|nats\.Publish\|\.Nats\.Publish\|QueueSubscribe\|SubscribeReply\|\.Request(' \
--include='*.go' core/ pkg/ | grep -v _test \
| grep -v 'c\.Request()\|ctx\.Request()\|Request()\.Context' \
| grep -v 'core/services/testutil/fakebus.go'
core/services/messaging/client.go:168,170,172,227,234,236,250,252,254,268,269,287
core/services/messaging/interfaces.go:21,22,23
Every remaining hit is inside core/services/messaging itself. The production
reads of the NATS client are now three, all of them the documented agent-worker
exception: Close on shutdown, the agent pool's publisher, and the agent-cancel
carrier passed to newFanoutBridges.
Prefix-cache observations publish like every other family rather than through a
method that refuses a message too large for a notification. The plan proposed
such a refusal on the reasoning that a long prompt makes a chain of thousands of
entries; ExtractChain caps a chain at Config.MaxDepth blocks, MaxDepth is a
constant with no operator knob, and the chain reaching Sync.Observe has one
source, the router's own extraction hook. A worst-case observation is a few
kilobytes against an 8000-byte cap, so the hot-path spill the refusal was
designed to avoid cannot occur, and shipping it would have added the programme's
only deliberate message drop to guard a condition that cannot arise. pgbus gains
FitsInline instead, a predicate that shares one size decision with Publish and
decides nothing, and core/application refuses at startup to wire a prefix cache
whose configured depth would put every observation over the cap.
The carrier choice is no longer stated at four sites. StagingTracker.SetPublisher
and SubscribeBroadcasts become one SetBroadcaster, so a tracker that publishes
where its peers are not listening cannot be spelled; prefixcache.Sync gains
SubscribeBroadcasts, which reads the carrier it publishes on; and the gallery
service and the operation cache are wired by methods on DistributedServices that
name no carrier at all, so the NATS client beside it cannot be handed over.
OpCache.SetMessagingClient and GalleryService.SetNATSClient are renamed to
SetBroadcaster so a missed call site fails to compile.
Two pre-existing defects that the two-real-carrier specs surfaced are fixed. A
progress tick published before a cancel and delivered after it cleared Cancelled
and left the operation reading as still running on that replica; mergeStatus now
drops a stale tick rather than merging it. GetStatus and GetAllStatus handed out
the stored OpStatus pointer while the broadcast subscribers mutated it in place,
so an /api/operations response could be marshalled mid-write; both now copy.
Assisted-by: Claude Opus 5 [claude-code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
6e2be0fd5a |
refactor(distributed): move job and agent fan-out onto the PostgreSQL carrier
Five of the six families whose subscriber is an open HTTP response rather than a process-lifetime cache now travel on pgbus: jobs.<id>.progress, jobs.<id>.result, jobs.<id>.cancel, agent.<name>.events.<user> and responses.<id>.cancel. Both ends of each move together, so there is no state where a publisher is on one carrier and its subscriber on the other. agent.<name>.cancel does NOT move, and the plan was wrong about why. Its only subscriber in the tree is the agent worker, which has no database and so cannot join the PostgreSQL carrier at all. Publishing that cancel on pgbus would have lost every cancel of a worker-run agent while returning nil, which reports a cancel that reached nobody as a cancel that was sent. EventBridge now names its cancel carrier separately, a frontend replica sets it to the carrier the worker reads, and it stays there until a cancel rides the worker's tunnel like every other verb addressed to a worker. The carrier drops at 256 rather than blocking, which is not safe on its own for a result: a lost result has no successor message. It is not the only path. The claiming replica persists the terminal line before it releases the claim, and an open progress stream re-reads the job row once after subscribing and then periodically, so a dropped terminal broadcast costs promptness and never the answer. Both per-request subscriptions close in a defer instead of on one return path, and pgbus grows Subscribers() so the leak they would otherwise cause can be asserted. It has no other symptom: only the first subscriber of a channel issues a LISTEN, so a leaked filter just adds one closure per notification for every stream the replica has ever served. Subscribe now issues its LISTEN before it registers, which makes that count a readiness signal rather than a figure to compare against itself. Two rules that were stated at several sites and pinned at none are now one each. The re-broadcaster is built beside the dispatcher and the bridge and handed to the dispatch loop, so no line is left that can point it at a carrier nobody subscribes to while every spec stays green. The set of statuses a job never leaves is one exported set that the SSE bridge and the store both read. The last hand-written subject filter in production code became messaging.SubjectAgentEventsWildcard. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
affad9c06d |
feat(distributed): carry the state.*.delta families on PostgreSQL
syncstate.Config held one carrier field typed as the NATS client, so a pgbus.Bus could not be handed to a SyncedMap at all: it satisfies messaging.Broadcaster and not MessagingClient. The durable re-hydration path built for the responses map therefore had a NATS-only consumer and nothing in the build said so. The field becomes Bus messaging.Broadcaster, SubscribeJSON moves to its own file and relaxes its parameter to Broadcaster, and the four adopters fan out over PostgreSQL LISTEN/NOTIFY: fine-tune jobs, quantization jobs, agent tasks with their per-tenant children, and Open Responses metadata. A new spec proves it on a real database, over two Bus instances on two pinned listener connections: a Set and a Delete carry, a payload past the 8000-byte notification cap comes back byte identical through the spill row, two families sharing one LISTEN channel stay separate, and a terminated listener re-hydrates a row written while it was gone. The five sites that each chose a carrier for an adopter are collapsed into one DistributedServices.Broadcast() accessor. Five field reads were five chances to leave one family on NATS with nothing failing, because messaging.Client satisfies Broadcaster too. The accessor also refuses to hand out a nil pgbus.Bus wrapped in a non-nil interface, which every adopter would read as "broadcast" and dereference on the first Set. SetTaskSyncNATS and SetJobSyncNATS are renamed to SetTaskSyncBus and SetJobSyncBus so a missed wiring site fails to compile. The response metadata table gains a retention of its own, defaulting to 24 hours. It inherited the Open Responses store TTL, which defaults to 0 meaning no expiration. Zero is defensible for a map that dies with the process and is not for a table: the table grew for the life of the deployment and a restarting replica re-hydrated every response the cluster had ever created. A row that names its own expiry is still judged on that column alone, and "this row is dead" now has one SQL spelling that PurgeExpired deletes by and ListUnexpired is the negation of, so a hydrate cannot resurrect what a sweep has already retired. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
1bf2b3f3d9 |
fix(distributed): give one tenant's agent tasks a subject of their own
Every AgentJobService built its tasks SyncedMap with the name "agent.tasks", and there is one service per user. So every tenant published on and subscribed to the same subject, state.agent-tasks.delta, and SyncedMap.apply scopes nothing: a task tenant A created was written into tenant B's in-memory map on every replica, and ListTasks reads that map. Nothing repaired it short of a process restart. The subject now carries the tenant in a token of its own, state.<name>.<tenant>.delta. Four tokens where the unscoped builder makes three, deliberately: SubjectMatches compares token count before anything else, so a tenant's subject and the cluster-wide one cannot cross-match, and neither can two tenants. Putting the tenant inside the name token would not do that, because the sanitizer folds '.' to '-' and the only filter that could then span tenants is state.*.delta, which spans every other family too. The rule is stated once. subscribeFilters calls publishSubject rather than restating the subject, so a map cannot end up publishing scoped and subscribing unscoped, which would leak exactly as before while every publish assertion passed. The one case that decides on its own is the cluster-wide administrative view: it hydrates from every tenant's rows, so it also takes the per-tenant wildcard, or it would be stale the moment any tenant wrote. A tenant hydrates from its own rows and applies only its own deltas. PerTenant defaults to false, so finetune, quantization and the responses store keep the subject they have. The second half of the same defect was the delete. taskStoreAdapter.Delete called DeleteTask(id) and JobStore deleted by primary key with no user predicate, reachable from DELETE /api/agent/tasks/:id, which takes the id off the URL. A tenant who learned another tenant's task id destroyed that tenant's row. The user id now travels with the id and lands as a user_id predicate. Empty stays the administrative any-owner scope, the same thing an empty id already means for ListTasks and ListJobs. A foreign delete removes nothing and returns no error: not yours and not there are the same answer to the caller, and neither is a store failure. SetUserID rebuilds the tasks map for the same reason SetTaskSyncNATS does. GetJobs happens to set the user id first, nothing enforced it, and with the order reversed the map would be built with an empty tenant and put that user's tasks back on the cluster-wide subject. Both halves predate this programme; they are surfaced here rather than caused. Neither is fully closed for a deployment with the agent pool off, where the task routes are still served by one cluster-wide service that every authenticated caller shares; that is a separate gap and it is documented. testutil.FakeBus grew a real defect this was the first change to trip: Unsubscribe matched on the filter string, so with two subscribers on one filter, closing one deafened the other. Subscriptions now carry an id. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
4b15161ffc |
feat(distributed): give responses.metadata something to re-hydrate from
The responses.metadata SyncedMap had no durable Store, so its reconnect re-hydrate replaced nothing. That was survivable while responses converged through deltas on a broker that mostly stayed up. It is not survivable on a carrier whose listener is one pinned PostgreSQL session: every response created while the subscription was down stays invisible on that replica forever, and the symptom is a 404 from one replica and a 200 from another for the same response_id. State that must survive a gap now lives in a response_metadata table, and the notification only says it changed. The map writes through on a Set and reads the table on hydrate, on reconnect and on reconcile, so the gap closes instead of becoming permanent. The row carries the whole projection as JSON rather than one column per field. A column-per-field schema would be a second definition of what a peer may act on, and the two would drift the first time syncedResponse gained a field: the map would broadcast the new field and hydrate without it, so a replica that had reconnected would serve a different response body from one that had not, with nothing failing anywhere. Only PayloadJSON is ever decoded; owner_replica and owner are indexed copies for an operator reading the table by hand. A missing row and an unreachable database are different facts. Every store and adapter method returns a driver failure as an error and never as an empty result, and syncstate replaces nothing when its source errors, so an outage leaves the map holding what it had rather than blanking it into a cluster-wide 404. Liveness is the database's clock, spelled expires_at IS NULL OR expires_at > now(), because every replica hydrating from this table must agree on which rows are live and a Go-side cutoff makes that a property of whichever process asked. The test container shares the host clock, so no behavioural spec can tell the two apart; the statement shape is pinned instead. The constructor refuses a non-PostgreSQL handle, because an unguarded now() on the single-binary path reads as a missing migration. A ticker sweeps expired rows every five minutes on each replica, and Close waits for it rather than racing it. Note that the sweep removes nothing while LOCALAI_OPEN_RESPONSES_STORE_TTL is 0, which is the default: with no TTL nothing ever expires and the table grows for the life of the deployment. The docs say so plainly. EnableDistributed takes the store positionally and last, so a call site that forgets it fails to compile rather than silently restoring the deltas-only map this change exists to replace. A nil store there is refused by name: it is reached only from the distributed branch of route registration, so it is a wiring bug and not a deployment shape. What still never leaves the owning replica is unchanged: the resume buffer and the CancelFunc. The write-through is one row per response state change, not one per generated token. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
e3040c7e3c |
feat(distributed): make MCP execution and discovery a selection
mcp.tools.execute and mcp.discovery were the only NATS subjects that combined a queue group with a reply, and no carrier in this design provides both. They never needed one: a queue group is a way of choosing a subscriber, and choosing is a query. The frontend now lists the approved, non-draining agent nodes, asks the node_connections table in one joined statement which of those tunnels a live replica holds, prefers one this replica holds so the call skips the relay hop, and issues an ordinary control RPC on the path task 4 already mounted. A peer-held tunnel is reached through the relay. That is a choice a broker's hidden balancing could not make. The selection reads presence and nothing else. It is filtered only on node type and on the two statuses an operator controls, never on a health verdict written on another clock, because refusing a worker that is connected and answering is the same defect as picking one that is gone. An empty fleet answers ErrNoAgentWorker, which is deliberately neither ErrWorkerUnroutable nor anything cluster.IsWorkerAnswer accepts: nothing was asked of any worker, so no reap guard may act on it. A reply carrying an Error is the worker's own answer and is returned unchanged; it is never offered to a second worker, which would turn "this MCP server rejected your arguments" into "the fleet is broken" and could run a tool twice. A call that never reached a worker is retried against a different pick, at most three times, and whatever error is finally returned is returned unwrapped so its identity survives the loop. MCP prompts and resources now answer 501 in distributed mode instead of an empty 200. They are served only from sessions the frontend holds, and in distributed mode it holds none. That gap predates the removal of the bus and is not closed by it; this only stops it being silent. Agent workers keep every other subject, including nodes.<id>.backend.stop. Their minted JWT loses the two MCP subjects and keeps a non-empty allow list, because NATS reads an empty one as no restriction at all. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
125e7080ce |
feat(distributed): give agent workers a tunnel of their own
Phase 2 gated agent nodes out of tunnel credentials at the mint site. That was right while nothing dialled into an agent worker: a credential would have replaced nothing, and the gate was structural rather than a second check that could drift. It is wrong now that the frontend needs to reach an agent worker by RPC. attachTunnelToken mints for backend and agent nodes and CLEARS for anything else, through one tunnelEligible predicate rather than two conditions that can be widened separately. ConnectHandler still never reads NodeType, so an empty hash is still what refuses an ineligible node. An agent worker now starts a loopback control server behind the same bearer check a backend worker uses, and holds one tunnel whose only stream tag is http: it runs no backend processes, so the grpc tag has nothing to route to and is not offered. Its MCP tool, MCP discovery and backend.stop verbs are served from ONE implementation reached by both the bus and the tunnel, so a frontend cannot get different bytes depending on which carrier delivered. The tunnel is an ADDITION. --nats-url is still required, and agent jobs, MCP execution, MCP CI jobs and nodes.<id>.backend.stop all still travel on the bus. Absence semantics are unchanged. An agent node now has a real node_connections row whose departure ages past the grace, so the node type check in HealthMonitor.tunnelDeparted stopped being an optimisation and became the rule; its comment says so, and the spec that pins it is shown red under a mutation that deletes the check. The scheduler needed no change: every placement query already filters node_type = backend, so an agent node never reaches nodeMayTakeWork. Shared rules moved to one site each. The request bounds, the POST-only check and the unknown-path 404 live in workerctl and are called by both worker packages; the bearer check that guards every extra route is one function in core/services/nodes used by both server constructors. workerctl.AllPaths splits into BackendPaths and AgentPaths, with AllPaths as their deduped union, because a backend worker does not mount the agent verbs and asserting otherwise would fail a correct worker. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
94b248666a |
feat(distributed): read worker absence from the database, not from a bus timeout
The scheduler decided whether a worker had gone away from nats.ErrNoResponders: one frontend's observation that nobody answered IT within a request budget. Two replicas asking in the same moment could disagree and demote each other's workers, and a worker re-homing its tunnel between replicas looked identical to one that had died. SmartRouter now reads cluster.Presence instead. Only PresenceGone -- no live replica holds the tunnel AND the departure has outlived the reconnect grace -- excludes a node from placement, and it is a fact every replica reads identically from the database. PresenceReconnecting, PresenceUnknown and a failed presence query are all non-verdicts and place work as normal: excluding on a database hiccup would cost the fleet its capacity for a reason that has nothing to do with any worker. nodeAnswersOnBus is deleted. It excluded on a sentinel no control RPC can produce, so it decided nothing while PingNode cost a relayed round trip per scheduling decision to feed it. PingNode goes with it, from the adapter and from NodeCommandSender. isRequestTimeout drops nats.ErrTimeout: every verb this adapter sends now travels over the worker's tunnel. The predicate is named nodeMayTakeWork rather than nodeHasRoute. "Route" is ErrWorkerUnroutable in this package, the condition nobody may act on; PresenceGone is the one a scheduler may. Spelling them the same way is the collapse this work exists to prevent. Also folds in ReapStale's return rename: it counts connection rows CLEARED, never rows deleted, and reading it as a delete count would make a worker that is re-dialling right now look forgotten. The spec pinning that a message merely quoting "nats: timeout" is not a timeout was scripting a SUCCESSFUL reply carrying the phrase, which comes back with a nil error and never reaches the classifier. Restoring the string match left it green. It now scripts a 5xx whose body carries the phrase, and asserts that the phrase reaches the classifier as a precondition. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
e8d2563abe |
fix(cluster): stop a late request frame reading as the worker's verdict
Making a worker's refusal reaping evidence created a defect one layer along, at the producer. The worker refused a ReadStreamRequest failure with ErrStreamRequestInvalid and its own comment said "Includes the deadline above expiring", which was harmless while every refusal reached the frontend as "no route" and became a reap the moment one of them did not. So a request frame that had merely not ARRIVED yet was reported as a non-transient verdict about a backend. It is reachable on the relay path, which carries most production traffic: the worker's header timer starts when the OWNING replica opens the stream, while the frame is written by the DIALLING replica only after the relay's acceptance travels back to it, so a whole peer-link round trip runs inside that window, on a link this design deliberately loads with multi-gigabyte artifacts beside token streams. For a long-deadline caller the endpoint is ConnectionEvictingClient, which stops the model across the fleet. It also falsified the "neither clears on its own" argument that licensed the reap. There is now a fourth refusal, ErrStreamNotServed, for what a worker could not serve for a reason of its OWN. It is deliberately outside IsWorkerAnswer, so it reaches a consumer under the no-route umbrella and reaps nothing, which is the same treatment an unrecognised code already gets. Four producers move onto it: a request frame that timed out (a malformed one stays a verdict, because that is a frontend bug no retry fixes), both SetReadDeadline failures, which are facts about the stream and not about a target nothing has dialled yet, and WriteStreamRefusal's default for a reason nobody classified. classifyServiceFailure keeps ErrStreamTargetUnavailable as its default on purpose: inverting it would make errno enumeration the single point of failure for the reap, and a miss there is a row nothing can ever delete. What it gains is a deny-list of two causes that are provably this worker's own clock or its own context. Also: - The read-site caller-deadline guard in the handshake was unpinned: the existing seam spends the budget before the handshake starts, so only the write could ever fail. A spec whose deadline falls between the request and the reply pins it, and each guard now reddens on its own. - The documented worker-first failure line omitted the JSON error envelope the old frontend returns, so an operator grepping it found nothing. - The peer-link disclosure names the aimable per-session receive window in all four places, and LastDialErrorOf records why a third consumer must go through IsWorkerAnswer rather than roll its own list. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
118eea6931 |
fix(distributed): let a worker's own refusal be evidence about its backend
A worker that refuses a stream has answered, and cluster.Dial keeps the three tunnelproto sentinels out of the ErrNoRoute umbrella precisely so a consumer can act on that. No consumer did. Since workers stopped listening, a backend process that crashed on a healthy worker is no longer a dead listener's codes.Unavailable: the worker refuses the stream with ErrStreamTargetUnavailable, gRPC flattens it into Unavailable anyway, and nodes.unroutable reported the whole thing as "this frontend has no route". Every reap path then answered ProbeUnknown and left the row, so the replica slot never freed and at the default MaxReplicasPerModel=1 the only cleanup left was LRU eviction of models that were working. isWorkerAnswer is exported as cluster.IsWorkerAnswer, so the errors the dialer keeps out of the umbrella are by construction the errors the consumers treat as the worker answering. nodes.unroutable and pkg/model's transportFailure both use it; ConnectionEvictingClient, the site reached during inference, goes through transportFailure rather than asking the transport directly. A reply code this frontend does not recognise is still not an answer, so a newer worker's vocabulary costs a retry and not a replica. The reap guards keep the allow-list rather than requiring ErrNoRoute: an unrecognised dial error must mean "no route", never "the backend is gone". Also in this final pass over the branch: - Docs: recommend upgrading FRONTENDS first, with the symptom of each order. Workers-first fails now that a 4xx registration is a verdict rather than an outage, so an old frontend's "address is required for backend workers" makes each restarted worker exit and drains the fleet a node per restart. - Docs: LOCALAI_WORKER_TUNNEL=false is a fatal startup error, not a degraded mode, in both places that described it; and a frontend rollback needs every worker restarted, because re-registration force-clears the address columns. - A replica with no advertised address now says so every five minutes and names the workers only it can reach, instead of one startup warning for a cost paid for the life of the process. - callerRanOut's rule now holds at all three siblings, so an expired caller deadline stops reading as a broken tunnel; probeHealth's withdrawn reason for using the raw client is corrected; the dead DoOrCached is deleted and its coverage kept on DoOrCachedResult; sweepLeakedInFlight enumerates the outcomes that reach it. - The peer route's self-declared id is recorded as a phase-3 deferral, in the handler, in the isolation claim it narrows, and in the operator docs. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
778c97aaf7 |
fix(cluster): stop blaming a peer for the caller's own expired deadline
Review round 1 on the end-to-end proof. Zero blocking items, eleven non-blocking, and three of them turned out to be production defects rather than notes on the report. The one that matters is a misclassification the phase is built to prevent. A dial carries the caller's deadline down to the socket, so when the budget runs out the socket's timer fires and the error travels back up through the WebSocket handshake and the multiplexer. The context's cancellation is a separate timer whose func the scheduler has to run before ctx.Err() stops returning nil, and nothing orders the two. Under contention the socket's error is back in PeerPool.Open first, ctx.Err() reads nil, and a peer that is listening and healthy is reported as ErrPeerUnreachable to a caller that simply ran out of time. An unreachable peer is a fact a caller may act on and an expired deadline is not, and core/services/nodes routes around a replica it is told is unreachable. callerRanOut answers that question in one place: ctx.Err() when it is set, and otherwise the wall clock against the caller's own deadline. That is sound because it is the same instant the socket compared itself against, so if the socket's timer fired this comparison is past it too. The ambiguous instant resolves towards the caller, which is the direction that never blames a peer. The spec that caught it, peerlink_test.go's "blames the caller's deadline", was red in three of seven -race runs and had been since Task 5, which is often enough to read as noise and is why single-run verification never saw it. Rather than leave the proof to a coin flip, a second spec makes the window deterministic: Open is handed a context whose deadline has passed and whose cancellation has not been delivered, against an address nothing is listening on, so the dial fails for real. It reddens without the fix. The peer link's yamux windows were applied to one end only. A receive window is advertised by the side that RECEIVES, so configuring the dialler alone tunes exactly one direction, and the direction left on the 256 KiB default is the one that carries a relayed model artifact INTO the replica that owns the worker's tunnel. That is the largest thing the link ever moves and it is the direction the load measurement exercises: the review read it as flowing toward the dialler and it does not. PeerLinkConfig is now exported and used on both ends. Measured, same box, 128 MiB staged through the relay against the same transfer without one: the relayed path cost 1.6x to 2.0x the direct path's transfer window before, and 1.06x to 1.25x after. The SSRF reachability spec could be fooled into reporting an SSRF that did not happen. It bound the victim on 127.0.0.2 at an ephemeral port and required 127.0.0.1 at the same port to refuse, so any other spec in the run holding that number made the dial succeed; red one run in seven, green five of five in isolation. It now picks from below the kernel's ephemeral range, the same fix the harness got for the adjacent-port collision. The rest are the specs and the report saying what they mean. Scenario 1's advertisement assertion could not tell "the worker advertises nothing" from "the JSON key moved", which matters because removing the advertisement is the change it covers. It was green against a renamed key. The roster now keeps the raw key set beside the decoded fields and the spec requires both keys present before reading them as empty. Scenario 4's refusal-body check was a four-way disjunction admitting bare "tunnel", "not connected" and "unroutable". Those alternatives were inert and each would be satisfied by refusals that say nothing about routing, in the one assertion the whole negative control rests on. It is "no route" alone. The head-of-line gate bounded the worst probe by the whole transfer window, which admits about eightfold degradation and loosens as the box slows. It is now half the window, plus a scale-free ratio against the worst probe under the SAME cold load with nothing to transfer, which is the control that isolates the transfer from the load. Not tighter than that, and the reason is measured rather than cautious: under a concurrent -race suite the worst relayed probe reached a fifth of its window, so a quarter-window gate would have had 1.2x of margin, and a spec that fails one run in three is worse than no spec. The report entry printed p90 and p99 off samples of twenty, where both land on the same element and p99 often lands on the max, so one number appeared three times under three names. A quantile is now printed only when the sample can separate it. Two claims in the report were wrong and are withdrawn rather than softened. Scenario 2's race is closed by the trailing re-read of the owner, not by the pre-assertion the report credited: a move to the non-owner mid-request would serve directly and still return 200, and only the trailing read reddens on it. And "the median request is unchanged" holds on this box and not on the reviewer's, where the relayed median rises up to 82% and p99 up to 3.5x. What survives on both is structural: the worst probe is a small fraction of the window in which bytes are moving, so the session interleaves rather than serialising. Sharing a session with a bulk transfer costs latency; it does not cost service. The disk footprint note undercounted, and the reviewer lost a run to a full disk on this box, so it is worth having right: two bulk models seeded into two frontends and staged to the worker is about 768 MiB, not 512 MiB. Left alone deliberately: the worker's backend port allocator still hands out ports without checking they are free, and its default range still overlaps the kernel's ephemeral range. It is confirmed, it is out of scope here, and it is being tracked as a named follow-up rather than fixed under an e2e task. Assisted-by: Claude Opus 5 [claude-code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |