Upgrade gguf-parser-go to v0.26.3 so string lengths are checked against the remaining file size before allocation. Master AIO CI crashed in the background gallery warmer when v0.25.0 tried to allocate several terabytes; panic recovery cannot catch a fatal runtime OOM.
Check both overflow-sized and file-exceeding strings through the remote reader, and document the size-only estimate fallback.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* fix(galleryop): announce model changes before the preload
After a gallery install or delete, the replica that ran it replaced its
config loader, then preloaded every installed model, and only then
published the models invalidation. The preload does remote lookups and
checksums for each model, so on a large models directory peers learned
about the change minutes after the originator listed it. When the
preload failed or the operation was cancelled, the event was never sent.
Publish the invalidation, and apply the delete lifecycle, as soon as
the loader holds the new set. The preload still runs afterwards with
its own error handling. Its failure is reported on the operation, but
it no longer rolls back a deletion that peers have already applied.
Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(distributed): resync model configs from the models directory
Frontends refresh their model configs only when a models invalidation
arrives on NATS. NATS keeps no history, so a frontend that is
disconnected when the message is published never applies the change.
It keeps serving the old config, for example an alias that points at
the previous model, until some later change happens to touch it.
Each frontend now reruns the peer reconcile against the shared models
directory after every NATS reconnect, and every
--model-config-resync-interval (default 30s) when a config file
changed. The pass names no model, so only models whose file changed
get a revision transition, and an unchanged directory costs one read
of the config files.
The reconcile replaced the whole loader with a parse of the models
directory, which dropped models loaded with --config-file and
published a deletion revision for them. Configs defined outside the
directory are now kept, both there and after a gallery install.
Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(nodes): stop stale frontends from retargeting alias rules
A scheduling rule keyed by an alias derives its target from the alias
mapping of the frontend that reads it. Every frontend keeps its own
copy of the model configs, so after an alias is repointed a frontend
that has not reloaded it still resolves the old target. Two frontends
then rewrote the rule's stored target_model against each other on
alternate reconciler ticks, and the outdated one scaled up the model
the alias used to point at.
The registry already records the accepted config revision of each
model. A frontend now derives a rule's target from its own alias
mapping only when its config revision for the rule's name matches
that record. Otherwise it keeps the stored target_model: it neither
writes the column nor reconciles replicas of the old target. The check
reads the database only for a rule whose stored and derived targets
differ. With no accepted revision on record, the old behaviour stays.
Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
New admin endpoint GET /api/models/storage reports disk usage of the
installed models: each model's files and sizes, files shared between
models counted once, and configured files that are missing from disk.
The models WebUI page shows a usage summary, each model's file table,
and the full file list with per-file status.
Assisted-by: Claude Code:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
Fence model load jobs by generation, lease them on the database clock, bound the work on the worker with operations and a process-group watchdog, and add one stop path with a load-cancel API. See the pull request for the design, the rolling upgrade notes and the test evidence.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(audio): record transcription usage and traces
Count successful transcription requests by model, including streams.
Keep token counts at zero because transcription exposes no token usage.
Capture multipart API trace metadata without reading uploaded audio.
Do not record failed transcription or client writes as successful usage.
Assisted-by: Codex:gpt-6-astra
* fix(audio): preserve aliases in streaming usage
Streaming transcription records the resolved target as its usage model.
Pass the requested name so JSON and SSE requests share the alias bucket.
Assisted-by: Codex:gpt-6-astra
* test(http): check multipart reader close errors
Assert successful reader cleanup to satisfy the errcheck CI gate.
Assisted-by: Codex:gpt-6
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(diarization): return sound events with include_sounds
A client that wants text, speakers, voice prints and sound events had to
make a diarization call and a separate sound call. Add an include_sounds
request field to /v1/audio/diarization that adds a sounds array of closed
events {start, end, label, confidence}, in seconds.
The parakeet-cpp backend runs a tagger-only scene stream over the clip,
the same stream and thresholds the live path uses, so a clip gives the
same events offline and live. A model with no sound_model companion, or a
backend that does not report sound events, fails with 501 and the stable
code include_sounds_unsupported instead of an empty list. The proto
carries sounds_included so an empty list still means "nothing heard".
The localai-proxy backend forwards the field. Swagger, docs and the
e2e mock backend are updated.
Assisted-by: Claude:claude-sonnet-5-5 [protoc swag go]
* feat(gallery): add parakeet-cpp-multilingual-diarization-speakers-sounds
Same as parakeet-cpp-multilingual-diarization-speakers (TDT 0.6B v3,
Nemotron-3-Diarization, WeSpeaker) plus a CED-Tiny sound_model, so one
model name serves /v1/audio/diarization with include_text,
include_speaker_profiles and include_sounds. It declares the
sound_classification usecase like the realtime scene entries.
Assisted-by: Claude:claude-sonnet-5-5
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
With the PostgreSQL vector engine, each frontend kept the collections it
had opened in an in-memory map and answered "collection not found" for
any other. A collection created through one frontend was unknown to the
others until they restarted, and each frontend listed a different set.
Wrap the in-process collections backend in distributed mode so that the
database is the source of truth:
- lists come from the registry,
- a lookup miss checks the registry before it returns 404, and opens a
collection that exists there (once per name, even under concurrency),
- a cached collection that left the registry is dropped and closed, with
a re-check at most every 5 seconds,
- create and reset publish an event on the existing collection
invalidation subject, so other frontends re-check at once.
Without the postgres engine, or outside distributed mode, nothing changes.
Assisted-by: Claude:claude-sonnet-5-5 go
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
The worker answered model.unload by calling Free on the first backend
process in its map, whatever model the request named. Every idle
scale-down or LRU eviction of one model could therefore empty another
model's backend on the same node. LocalAI still counted that model as
loaded, so its next request failed. A parakeet diarization model then
returned 501 "speaker profiles require a loaded speaker encoder" until
someone reloaded it by hand.
Resolve the target from the model name (all replicas), prefer an
address when the request carries one, and free nothing for an unknown
model.
Also let parakeet-cpp Diarize check the diarization model before the
speaker-profile capability. A backend with nothing loaded now answers
FailedPrecondition, which LocalAI treats as a stale replica and
reloads, instead of a final Unimplemented.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
The English-only 110M ASR in parakeet-cpp-nemotron-3-diarization-asr-speakers
garbles other languages. Add a bundle that pairs Parakeet TDT 0.6B v3
(25 European languages) with the Nemotron-3-Diarization model and the
WeSpeaker ResNet34 encoder, so one /v1/audio/diarization call with
include_text and include_speaker_profiles returns turns, multilingual text
and one voice-print embedding per speaker.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
A key can now be paused until the owner resumes it, or until a given
time. A paused key is rejected by validation before last_used is
updated, and a pause time that has passed lifts the pause by itself.
Existing keys stay active.
PATCH /api/auth/api-keys/:id takes {"disabled": bool, "paused_until":
RFC 3339 string or null}. Only the key owner can change it, and a
paused_until in the past is rejected. The key list returns the pause
fields. The Account page gets a Pause and Resume button for each key
and a Paused badge that shows the resume time.
Assisted-by: Claude Code:claude-sonnet-5-5
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
The parakeet-cpp backend now answers VoiceEmbed and VoiceVerify, so the
pinning test asserts the usecase on the entries that declare it instead of
its absence. Document voice_verify_threshold.
Assisted-by: Claude Code:claude-sonnet-5-5
parakeet_capi_speaker_embed_pcm and parakeet_capi_free_floats, which
VoiceEmbed binds, come from this commit.
Assisted-by: Claude Code:claude-sonnet-5-5
VoiceVerify used a fixed distance of 0.5 when the request had none. Read
voice_verify_threshold, a distance in (0, 2), from the model options and
keep 0.5 as the default. The real-library spec now cuts clips from a
two-voice recording and checks the bundle embedding size, the encoder
identity, determinism and the same-voice versus different-voice distance.
Assisted-by: Claude Code:claude-sonnet-5-5
Describe the bundle as an embedding model for /v1/voice/*, the realtime
voice_recognition stage, and the limits: identify and plain verify only,
a 256-dimension space shared with voice-detect-wespeaker-resnet34, and the
libparakeet.so symbol it needs.
Assisted-by: Claude Code:claude-sonnet-5-5
The small and standard bundles have a voice component, and
parakeet-cpp-realtime-scene-speakers loads a speaker encoder, so they can
serve /v1/voice/* and the realtime voice_recognition stage. Add the usecase
to their known_usecases and pin it in the gallery tests.
Assisted-by: Claude Code:claude-sonnet-5-5
parakeet-cpp now implements VoiceEmbed and VoiceVerify, so list them in
its backend capabilities together with the speaker_recognition usecase.
The usecase is not guessed, because most parakeet models have no
speaker encoder: a bundle declares it in known_usecases, which is
enough for /v1/voice/* and for voice_recognition.model.
Add a realtime pipeline test where one bundle model names the vad,
transcription, sound_detection and voice_recognition stages and the
gate resolves the speaker through the backend's VoiceEmbed.
Assisted-by: Claude Code:claude-sonnet-5-5
Bind parakeet_capi_speaker_embed_pcm and parakeet_capi_free_floats with a
Dlsym probe and implement VoiceEmbed on the speaker context, so a model
with a speaker_model or a bundle speaker_component can serve the
realtime voice_recognition stage and /v1/voice/*. The response carries
the sha256 identity of the encoder weights.
VoiceVerify embeds both clips and compares them by cosine distance. It
refuses anti_spoofing because there is no such head.
A libparakeet.so without the symbols answers Unimplemented, and a model
without a speaker encoder answers FailedPrecondition.
Assisted-by: Claude Code:claude-sonnet-5-5
Cover speaker_component: and speaker_tag: in the audio-to-text option
tables and bundle section, add a section to the voice recognition page
that explains which registered voices a bundle component receives, and
update the speaker_name condition in the realtime and diarization pages.
Assisted-by: Claude Code:claude-sonnet-5-5
Naming read only speaker_model:, so a parakeet-cpp bundle that sets
speaker_component:voice never received the registered voices, in live
transcription or in diarization.
Resolve speaker_component: when speaker_model: is not set. A component
is not an encoder file, so no file-name tag is derived for it: voices
with a hash or family identity and untagged voices reach the backend,
and a voice with only a file-name tag matches through the new optional
speaker_tag: alias. speaker_model: still wins, including when it points
at the bundle file.
The bundle gallery entries do not declare speaker_recognition: that
usecase means the VoiceEmbed and VoiceVerify RPCs, which the backend
does not serve, and it selects the default model for /v1/voice/*. The
pinning test now says so.
Assisted-by: Claude Code:claude-sonnet-5-5 go golangci-lint
The in-memory local-store rejects vectors of another size than the ones
it holds. With 192-value voices registered, adding a 256-value voice
failed, and identifying with a 256-value probe returned an error.
The registry now keeps one store per embedding dimension. The first
dimension seen uses the configured store name, so single-encoder
instances are unchanged. Later dimensions use "<name>-<dim>". Identify
searches only the store of the probe size and returns no match when no
voice of that size exists. Forget finds the store from the stored
embedding. A name may hold one voice per encoder.
Assisted-by: Claude Code:claude-sonnet-5-5 golangci-lint
* chore(deps): bump localrecall to v0.6.6
LocalRecall v0.6.6 closes the PostgreSQL connection pool when a
collection fails to open. Before, each failed collection create left
its pool open. With a failing embedding model, every retry leaked one
more pool until PostgreSQL refused new clients.
LocalAI gets LocalRecall through LocalAGI, so this raises the indirect
requirement directly instead of waiting for a LocalAGI bump.
The release also limits each collection pool to 4 connections by
default. POSTGRES_POOL_MAX_CONNS changes the limit.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
* docs(agents): document the PostgreSQL pool size limit
LocalRecall v0.6.6 limits the connection pool of each collection to 4
connections and reads POSTGRES_POOL_MAX_CONNS to change it. Document
the variable next to the other PostgreSQL settings of the embedded
store.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* chore(parakeet-cpp): bump parakeet.cpp to 2de154c
Brings in the speaker registry encoder fingerprint, the VAD segment trim
and the opt-in word filter, a fix for a per-call thread count that stayed
set on the process-wide backend after a Silero VAD pass, and bundle
components loaded from memory.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* feat(parakeet-cpp): encoder fingerprint for speaker naming, vad_trim and guard_* options
Speaker naming. A registered voice now carries the encoder that made it:
the embedding family (voicedetect:<arch>:<name>:<dim>) and the sha256 of
the weights. The backend reports the family of the loaded speaker model in
Status, voice enrollment from speaker_profiles stores it as encoder_family
(old entries load without it), and the registry sent to parakeet.cpp is
built with parakeet_capi_speaker_registry_add_embedding_fp. The library
then refuses a registry of another encoder family and the error names both
families; another quantization of the same family only warns. A voice with
only a weights hash gets the loaded family when the hashes are equal.
Voices without a fingerprint (registered from audio: libvoicedetect cannot
report one) keep the file-name rule and are used with a warning. The
library cannot mix them with fingerprinted voices in one registry, so a
request that has any uses the old registry for all. speaker_strict:true
drops them instead. A library without the symbols behaves as before.
Transcription. vad_trim (seconds, 0 keeps the whole cuts) goes through the
VAD options JSON, so it reaches /v1/vad and the segmenter. The guard_*
options guard_min_local_conf, guard_local_radius and guard_drop_punct_only
turn on the word filter through parakeet_capi_transcribe_path_json_with,
or through the segmenter with vad:true. They are off by default, bad
values fail the load, and a library without the symbol fails it with a
clear message. The dropped word count is logged at debug level.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Saved voices send ref_text, but Fish Audio requires reference_text.
Derive the canonical parameter while preserving explicit overrides.
Both TTS modes use the shared builder.
Add regression cases and document the parameter alias.
Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* feat(agents): keep the web chat history between messages
The agent chat endpoint ran every message as a fresh job
(ag.Ask(WithText(message))), so a follow-up such as "add two days to
item 3 and recalculate" never saw the answer it referred to. Agents
then rebuilt their reply from scratch instead of changing it.
Use the agent's own conversation tracker, as the Telegram and Slack
connectors already do: send the earlier turns with the new message and
record successful answers. Failed, cancelled or empty runs are not
recorded, and the tracker drops a conversation after the agent's
last_message_duration of inactivity. The distributed (NATS) chat path is
unchanged.
Assisted-by: Claude:claude-opus-5-5 ginkgo
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
* fix(agents): keep each web chat conversation's history separate
Review on this PR: the first version kept the history in the agent's
conversation tracker under one fixed key per agent. The web UI keeps
several conversations per agent (New Chat, switching, Clear) and did not
send which one a message belongs to, so the hidden history mixed them:
a new chat received the previous chat's turns, and Clear only cleared
the screen.
The client now sends the earlier turns of the conversation it is
showing as `history` with POST /api/agents/:name/chat, and the server
keeps no web chat history of its own. Each conversation only ever sees
its own turns; New Chat and Clear start without history. The server
uses only user and assistant turns with text, bounded to the most
recent 40 turns and 64,000 characters. `history` is optional, so
existing callers keep the previous behaviour (no history); the
distributed (NATS) path does not forward it yet.
Tests: two conversations of one agent stay apart, an empty history
(New Chat or Clear) starts fresh, system/tool/empty turns are dropped,
the bounds keep the most recent turns; the UI helper that builds the
history from the visible messages has node --test coverage. Docs:
features/agents.md describes `history` and the web UI behaviour.
Assisted-by: Claude:claude-opus-5-5 ginkgo
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
---------
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
Newer ROCm releases split the rocBLAS and hipBLASLt TensileLibrary data
into one folder per GPU architecture (lib/hipblaslt/library/gfx1151/...).
The run.sh of the ROCm-capable backends exports ROCBLAS_TENSILE_LIBPATH and
HIPBLASLT_TENSILE_LIBPATH as the parent folder, and both libraries only look
directly in that folder, so they miss every kernel:
rocblaslt error: Cannot read ".../hipblaslt/library/TensileLibrary_lazy_gfx1151.dat"
hipModuleLoad failed: .../hipblaslt/library/Kernels.so-000-gfx1151.hsaco
and fall back to slower code paths.
rocm_tensile_dir keeps the folder when it has files at the top level (the
older flat layout) or several architecture folders, and descends into the
architecture folder when the bundle carries exactly one.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
The speaker encoder metadata field makes the Rust status initializer
incomplete. Use protobuf defaults for absent metadata and memory fields.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Duplicate the socket descriptor before handing it to the activation helper.
Close the original file to prevent its finalizer from closing a reused
coverage descriptor after the helper closes its own file wrapper.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* fix(xsysinfo): read process VRAM without proc children
Kernels without CONFIG_PROC_CHILDREN have no task children file, so
ProcessVRAM dropped every DRM reading. When that file is missing, walk
child processes from /proc/<pid>/stat ppid links instead. Other read
errors still drop the reading.
Fixes#12481
Signed-off-by: leilei3167 <imleilei123@gmail.com>
* docs(system): describe the proc children fallback
Document VRAM reporting on kernels without CONFIG_PROC_CHILDREN.
Assisted-by: Codex:GPT-6
---------
Signed-off-by: leilei3167 <imleilei123@gmail.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* fix(oci): stage image downloads in a configured directory, not in /tmp
Problem:
- the image tar and its compressed layers staged in os.TempDir()
- /tmp is commonly a RAM-backed tmpfs far smaller than a backend image
- big installs failed with a full /tmp or silently ate RAM
Change:
- staging path resolved like the other storage paths
(LOCALAI_DOWNLOAD_STAGING_PATH, default ${basepath}/downloading),
threaded to the extractor as a download option
- every download works in its own subdirectory holding its tar and
layers: nothing shared between concurrent downloads, one removal
cleans a download up, a crash leaves one self-contained orphan
- callers that pass no staging directory keep the OS temp behavior
Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
* fix(oci): handle staging cleanup errors
Log staging cleanup failures to satisfy errcheck. Test extraction with an
unavailable OS temp directory and verify that staging is removed.
Assisted-by: Codex:GPT-6
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
---------
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>