mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-09 22:54:42 -04:00
e81e180ce2d48ba923b62334d4ca70a2fd9481dd
8387
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e81e180ce2 |
chore(website): refresh the counters (#12490)
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
05b20bd7ab |
chore: ⬆️ Update mudler/parakeet.cpp to 0cca477249ffb16c1623fb947d5bac0624961d41 (#12460)
⬆️ Update mudler/parakeet.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
0a72af4e95 |
chore: ⬆️ Update localai-org/voice-detect.cpp to bca46bcbc2fe68169c2d7c414e9b290a7cb89911 (#12486)
⬆️ Update localai-org/voice-detect.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
d79489adb3 |
chore: ⬆️ Update localai-org/ced.cpp to 736a4ee46a31d4ff38b41add65ce0f2cf1aaa05f (#12487)
⬆️ Update localai-org/ced.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
7310887e26 |
chore(model-gallery): ⬆️ update checksum (#12485)
⬆️ Checksum updates in gallery/index.yaml Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
5f5feea15a |
fix(agentpool): keep the conversation in the agent web chat (#12410)
* feat(agents): keep the web chat history between messages The agent chat endpoint ran every message as a fresh job (ag.Ask(WithText(message))), so a follow-up such as "add two days to item 3 and recalculate" never saw the answer it referred to. Agents then rebuilt their reply from scratch instead of changing it. Use the agent's own conversation tracker, as the Telegram and Slack connectors already do: send the earlier turns with the new message and record successful answers. Failed, cancelled or empty runs are not recorded, and the tracker drops a conversation after the agent's last_message_duration of inactivity. The distributed (NATS) chat path is unchanged. Assisted-by: Claude:claude-opus-5-5 ginkgo Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> * fix(agents): keep each web chat conversation's history separate Review on this PR: the first version kept the history in the agent's conversation tracker under one fixed key per agent. The web UI keeps several conversations per agent (New Chat, switching, Clear) and did not send which one a message belongs to, so the hidden history mixed them: a new chat received the previous chat's turns, and Clear only cleared the screen. The client now sends the earlier turns of the conversation it is showing as `history` with POST /api/agents/:name/chat, and the server keeps no web chat history of its own. Each conversation only ever sees its own turns; New Chat and Clear start without history. The server uses only user and assistant turns with text, bounded to the most recent 40 turns and 64,000 characters. `history` is optional, so existing callers keep the previous behaviour (no history); the distributed (NATS) path does not forward it yet. Tests: two conversations of one agent stay apart, an empty history (New Chat or Clear) starts fresh, system/tool/empty turns are dropped, the bounds keep the most recent turns; the UI helper that builds the history from the visible messages has node --test coverage. Docs: features/agents.md describes `history` and the web UI behaviour. Assisted-by: Claude:claude-opus-5-5 ginkgo Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> --------- Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> |
||
|
|
9b0b7da96e |
fix(rocm): point *_TENSILE_LIBPATH at the architecture folder (#12411)
Newer ROCm releases split the rocBLAS and hipBLASLt TensileLibrary data into one folder per GPU architecture (lib/hipblaslt/library/gfx1151/...). The run.sh of the ROCm-capable backends exports ROCBLAS_TENSILE_LIBPATH and HIPBLASLT_TENSILE_LIBPATH as the parent folder, and both libraries only look directly in that folder, so they miss every kernel: rocblaslt error: Cannot read ".../hipblaslt/library/TensileLibrary_lazy_gfx1151.dat" hipModuleLoad failed: .../hipblaslt/library/Kernels.so-000-gfx1151.hsaco and fall back to slower code paths. rocm_tensile_dir keeps the folder when it has files at the top level (the older flat layout) or several architecture folders, and descends into the architecture folder when the bundle carries exactly one. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> |
||
|
|
e70c409a93 |
fix(kokoros): default optional status fields (#12447)
The speaker encoder metadata field makes the Rust status initializer incomplete. Use protobuf defaults for absent metadata and memory fields. Assisted-by: Codex:gpt-6 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
4a72b2ac2e |
test(cli): transfer socket descriptor ownership (#12465)
Duplicate the socket descriptor before handing it to the activation helper. Close the original file to prevent its finalizer from closing a reused coverage descriptor after the helper closes its own file wrapper. Assisted-by: Codex:gpt-6 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
d90f703bd9 |
fix(xsysinfo): read process VRAM without proc children (#12482)
* fix(xsysinfo): read process VRAM without proc children Kernels without CONFIG_PROC_CHILDREN have no task children file, so ProcessVRAM dropped every DRM reading. When that file is missing, walk child processes from /proc/<pid>/stat ppid links instead. Other read errors still drop the reading. Fixes #12481 Signed-off-by: leilei3167 <imleilei123@gmail.com> * docs(system): describe the proc children fallback Document VRAM reporting on kernels without CONFIG_PROC_CHILDREN. Assisted-by: Codex:GPT-6 --------- Signed-off-by: leilei3167 <imleilei123@gmail.com> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
104c2fb440 |
fix(oci): stage image downloads in a configured directory, not in /tmp (#12280)
* fix(oci): stage image downloads in a configured directory, not in /tmp
Problem:
- the image tar and its compressed layers staged in os.TempDir()
- /tmp is commonly a RAM-backed tmpfs far smaller than a backend image
- big installs failed with a full /tmp or silently ate RAM
Change:
- staging path resolved like the other storage paths
(LOCALAI_DOWNLOAD_STAGING_PATH, default ${basepath}/downloading),
threaded to the extractor as a download option
- every download works in its own subdirectory holding its tar and
layers: nothing shared between concurrent downloads, one removal
cleans a download up, a crash leaves one self-contained orphan
- callers that pass no staging directory keep the OS temp behavior
Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
* fix(oci): handle staging cleanup errors
Log staging cleanup failures to satisfy errcheck. Test extraction with an
unavailable OS temp directory and verify that staging is removed.
Assisted-by: Codex:GPT-6
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
---------
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
|
||
|
|
4d681b7f6d |
fix(realtime): detach VAD-commit transcription from barge-in cancellation (#12446)
* fix(realtime): keep VAD-commit transcription alive across barge-in, cancel it at teardown, order commits
Barge-in (new speech onset) cancels the turn's SourceVAD response context
(realtime_turncoord.go: respSink.cancel(SourceVAD)). The VAD commit body
runs under that same context, so an in-flight Whisper STT call was aborted
with 'context canceled' whenever the caller kept talking while the first
chunk was transcribing. The user's turn was lost: no transcript, no
LLM/TTS response.
v1 of this fix ran the transcription with context.WithoutCancel(ctx). The
review correctly pointed out two correctness gaps:
1. Teardown lost its cancellation. WithoutCancel detaches from every
cancellation, so a transcription in flight at session close outlived
the session and blocked respSink.shutdown (which joins the response
goroutines) until the backend finished the job.
2. Out-of-order commits. Consecutive commits run in parallel goroutines,
so a fast second transcription could append its user item before a
slow first one: the conversation became [second, first] and the second
response saw only [second].
Changes (core/http/endpoints/openai/):
- Session gains a session-lifetime context (sessionCtx), cancelled by
conncoord's Teardown BEFORE respSink.shutdown joins the response
goroutines. The transcription (and the voice-gate resolution) run under
it: they survive barge-in (which cancels only the per-response context)
but are cancelled with the session.
- Commit slots order the user-item appends in speech order:
Session.nextCommitSlot() is claimed at commit issue time (VAD CommitTurn
/ client commit), a commit's item append waits on the previous slot's
done (aborts on the session context), and every exit closes the slot so
a failed or torn-down commit never blocks the next. Transcriptions stay
parallel; only the appends are ordered.
- If the turn's response context was cancelled while the (detached)
transcription ran — barge-in, superseded by a newer commit — the user
item still commits (appendUserItem, split out of generateResponse) so
the LLM context keeps the full user input, but no response is generated
for the superseded turn; the newer speech triggers its own response on
the complete history.
- Regression tests (realtime_commit_order_test.go) cover both review
schedules — teardown during an in-flight transcription, and
held-first/finished-second out-of-order completion — plus the
barge-in-during-transcription item survival, driving the real commit
path with a transcription double that honours context cancellation.
- docs/design/realtime-state-machines.md: implementation-status entry for
the committed-turn pipeline (transcription lifetime + commit order).
Fixes #12445
Validated: builds, go vet clean, all openai specs + respcoord/turncoord/
conncoord suites pass under -race (incl. the 3 new regression specs).
Production A/B (call-center voice agent, SIP, silero-vad +
whisper-large-turbo + LLM + TTS, server_vad ~600 ms) on LocalAI v4.11.0:
unpatched — 'transcription_failed: context canceled', first part of the
utterance lost, agent answers only the remainder; patched — full
transcript committed, agent answers the complete utterance, barge-in
still cancels the in-flight assistant TTS response as intended, and
teardown cancels the in-flight transcription instead of waiting for the
backend.
Signed-off-by: nexxtmobile.de <kai@nexxtmobile.de>
* fix(realtime): release commit slots in order on every exit; share slot+issue boundary
Follow-up to the review of
|
||
|
|
d66383c0e2 |
docs(readme): refresh news through LocalAI 4.11 (#12484)
The news section stops at June despite newer published releases. Summarize releases 4.5 through 4.11 and collapse existing news without removing history. Keep the release index and blog links visible. Assisted-by: ChatGPT:gpt-6-astra Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
3185b6fcf8 |
feat(parakeet-cpp): load bundle GGUF files, bundle gallery entries, pin bump (#12479)
* feat(parakeet-cpp): load bundle GGUF files and use their components by role A bundle GGUF holds several models (ASR, VAD, diarization, sound events, speaker encoder) in one file, each with its own licence. Detect a bundle at load through parakeet_capi_bundle_components_json and open components with parakeet_capi_load_component. The three symbols are probed together, so an older libparakeet.so still loads plain files as before. The only ASR component is the primary model; bundle_asr:<name> picks one when there are several. A Silero VAD component of the primary bundle is loaded without an option and serves /v1/vad and vad:true. The diar, ced and voice components load on request: diar_component, sound_component and speaker_component, or a companion option (diarization_model, sound_model, speaker_model, vad_model) that names a bundle, even the model file itself. vad_component picks a VAD component and implies vad:true. A role the bundle cannot fill fails the load with the component list, and a diarization or sound request on a model without that role names the bundle components. Every existing option and single-file model behaves as before. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * chore(parakeet-cpp): bump parakeet.cpp to 781a973 Brings in the bundle GGUF format and its C-API (parakeet_capi_load_component, parakeet_capi_bundle_components_json, parakeet_capi_load_error). Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): gallery entries for the bundle GGUF files, docs Add parakeet-cpp-bundle-small (338 MB: Parakeet TDT+CTC 110M, Nemotron-3- Diarization, CED-Small, WeSpeaker ResNet34-LM, Silero VAD), -standard (1.1 GB, Parakeet TDT 0.6B v3 instead of the 110M model) and -moondream-redux (215 MB: packed Redux and Silero VAD, CPU only). One install serves transcription, VAD, diarization, sound events and speaker naming through the component options. The existing single-purpose entries stay. A bundle has no single licence, so the entries use license: other and state the licence and credit of each component in the description, with the upstream inconsistency of the CED licence. The docs get a section on bundles in audio-to-text with the entries, the roles, the options and the licence notice, and pointers from the VAD, diarization and sound classification pages. A gallery test checks the file names, checksums, usecases and options. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
d3724eef50 |
fix(nodes): stage dedicated diarization audio (#12474)
Dedicated diarization forwards frontend audio paths to remote workers, which cannot read those temporary files. Stage the input before the RPC and release it afterward, following the transcription lifecycle. Clone the request so staging does not change caller-owned data. Cover input bytes, request fields, cleanup, and error propagation in tests. Document distributed diarization staging on the existing feature page. Assisted-by: nib:gpt-6-astra Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
ed4a3975be |
feat(parakeet-cpp): gallery entries for the VAD-only Moondream slices, pin bump (#12469)
* feat(parakeet-cpp): add gallery entries for the VAD-only Moondream slices Add parakeet-cpp-vad-moondream-redux and parakeet-cpp-vad-moondream-ultra. They install the VAD head of Moondream Redux and Ultra (Q8_0) as small files of 10 MB and 6 MB, cut out of the full models without retraining, for the VAD endpoint. The files cannot transcribe, and a transcription request fails with a clear error. The files load only with a parakeet.cpp build that has VAD-only GGUF support (parakeet.cpp pull request 87). The backend pin must move to a commit that includes it before these entries work in a released image. The parakeet-cpp-vad entry keeps installing Silero. The docs list the files with the size, load time and memory compared with loading a whole model. A gallery test checks the usecase, the file name and the checksum of each entry. Assisted-by: Claude Code:claude-sonnet-5-5 [golangci-lint] * chore(parakeet-cpp): bump parakeet.cpp to e53a253 Brings in the VAD-only GGUF loader. Assisted-by: Claude Code:claude-sonnet-5-5 [git] [gh] * docs(gallery): link the parakeet.cpp VAD docs instead of the merged PR Assisted-by: Claude Code:claude-sonnet-5-5 [git] --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
99043b442c |
feat(router): route with native decision models (#12449)
* fix(schema): preserve SystemOne image inputs Assisted-by: OpenAI * test(schema): follow Ginkgo conventions for decision inputs Assisted-by: OpenAI * feat(llama-cpp): dispatch native decisions through Score Upgrade the stock dependency and reconcile Score/TTS patches. Reuse native decision parsing, tasks, formatting and response-reader cleanup; preserve ordinary scoring admission and guard older dependencies. Assisted-by: OpenAI * refactor(systemone): share request and model validation Assisted-by: OpenAI:gpt-5 * fix(systemone): preserve HTTP wire-byte validation limit Keep structural validation separate from the serialized internal request bound so HTML escaping cannot reject valid HTTP payloads. Assisted-by: OpenAI:gpt-5 * feat(systemone): bound images and account native decisions Preserve public wire limits independently from router serialization. Reject unsupported NER images, map native request/capability errors, and stamp explicit usage once. Advertise decisions for stock llama-cpp. Assisted-by: OpenAI * fix(systemone): record usage on registered native route Exercise real registration and billing with a mock native backend. Reject empty native responses, malformed image URLs, trailing JSON, and wire overflow including whitespace. Assisted-by: OpenAI * feat(router): add lazy native decision transport Bind named models through internal ModelSystemOne calls with shared validation and bounded abandoned operations. Remove request and echoed-error contents from decision traces. Assisted-by: OpenAI:gpt-5 * feat(router): classify overlapping policies with native decisions Ask independent noul questions, validate probabilities and preserve first-superset routing. Wire the central factory with config-sensitive invalidation and cancellation-safe resolution. Document native framing and bounded operation limits. Assisted-by: OpenAI:gpt-5 * feat(gallery): add pinned Julia-1 native decision model Add a separate text-only llama-cpp Q8 entry with pinned Apache-2.0 source provenance and checksum. Installed using the gallery installer and exercised choice, score and noul on CPU. Assisted-by: OpenAI * test(router): verify native decisions through central factory Add an opt-in real-model Ginkgo integration covering the native Go loader and C++ transport, token usage, independent overlapping labels, and candidate selection. Document owned-server execution and the intentionally non-quality threshold. Assisted-by: Codex:gpt-5 * fix(llama-cpp): align upstream pin and preserve decision signatures Advance to bed0a856 without losing the automated upstream bump. Detect full-request fill_task support at compile time and forward every question for Nimble framing while retaining the earlier native signature. Preserve reconciled SCORE/TTS patches; add standalone compatibility coverage. Assisted-by: Codex:gpt-5 * feat(gallery): add native decision family defaults Pin Laya, Kev-4B, lev, OpenJev and Nimble artifacts. Verify Laya/Kev/lev gallery installs and CPU contracts on both native pins; clearly mark OpenJev/Nimble runtime validation pending and their noncommercial licenses. Assisted-by: OpenAI * docs(decisions): clarify integrated Nimble prerequisite Record the exact combined backend pin while retaining pending OpenJev and Nimble installation/runtime validation status. Assisted-by: Codex:gpt-5 * fix(gallery): indent native decision model sequences Match repository yamllint indentation for Laya, Kev, lev and OpenJev list fields. Parsed gallery data is unchanged; reproduce CI gallery lint failure before the whitespace-only fix and pass the same command afterward. Assisted-by: Codex:gpt-5 * docs(decisions): record OpenJev and Nimble CPU validation Record gallery installation, checksum/metadata verification and multiquestion native smoke results on bed0a856. Retain noncommercial and text-only limitations without accuracy or deterministic-output claims. Assisted-by: OpenAI * fix(ui): expose native Decisions router classifiers Select classifier models using metadata-driven capability routing, retain tuned thresholds, and validate native decision selections before saving. Cover both native backends and create/save/reopen in the real React editor. Assisted-by: Codex:gpt-5 * fix(router): exclude aliases from native decision discovery Check the originally named config before advertising native Decisions eligibility. Retain target capability inheritance for ordinary generation aliases. Exercise the actual capabilities endpoint with native models on both backends, aliases, and disabled models. Assisted-by: Codex:gpt-5 * feat(systemone): share bounded multimodal input validation Preserve text wire limits while admitting bounded PNG/JPEG decision input. Share collection and header validation across internal and public callers and keep the native runner response budget independent. Assisted-by: OpenAI:API-assistant * fix(systemone): bound admission lifetimes and validate complete images Retain shared admission leases through actual work completion, including abandoned internal operations. Decode bounded image pixels, cap public native responses before usage stamping, and preserve oversized malformed text status precedence. Assisted-by: OpenAI:API-assistant * fix(router): classify images before media fetching Preserve ordered structured probes for native decisions. Defer OpenAI media preparation until routing selects the served model, so rejected decision URLs cannot trigger downloads before shared validation. Guard direct image collection with context-aware shared admission. Keep text classifiers and embedding caches from discarding image input. Retain fail-closed classifier configuration and runtime fallback policy. Add middleware, typed-content, admission, cancellation and cache tests. Assisted-by: OpenAI:API-assistant * fix(router): bound extraction before serialization Check probe budgets before copying text or marshaling message state. Count JSON escaping so oversized internal inputs fail before allocation. Preserve typed Anthropic blocks through selected-model conversion and fallback. Keep retry coverage in Ginkgo without global test registration. Assisted-by: OpenAI * fix(router): bound supported probe serialization Arbitrary structs can bypass the probe budget through pointer marshalers, string tags, and promoted fields. Accept concrete chat schema types and plain JSON values instead of emulating arbitrary struct serialization. Budget escaped direct prompts before marshaling so raw length cannot hide serialized expansion. Preserve runtime fallback and reject oversized input before invoking the decision runner. Add Ginkgo allocation, boundary, and marshaler invocation regressions. Six-package tests, three-package race tests, and full-T2 delta lint pass. Assisted-by: OpenAI:GPT-5 golangci-lint * feat(decisions): enable bounded OpenJev images Validate native decision images before permissive media parsing and pixel allocation. Require both decision image support and a vision projector; missing or audio-only projectors cannot silently become text decisions. Pin the OpenJev Q8 projector and document its license and disk footprint. Add native safety tests, canonical limit parity, gallery and load-option checks, and a reproducible CPU direct-RPC contrasting-image smoke. Assisted-by: OpenAI:GPT-5 * fix(decisions): reject incomplete image streams stb accepts corrupt PNG Adler checksums and truncated JPEG scans. Use bounded zlib validation and strict libjpeg decoding before parsing. Keep dimension and aggregate pixel checks ahead of decoder allocations. Wire decoder dependencies into native builds and runtime packaging. Add regressions for appended EOI and embedded marker bypasses. Assisted-by: OpenAI:GPT-5 * fix(ci): gate native decision image validation Run the decoder security tests outside the stdlib-only native suite. Fetch vendor headers at the backend pin and provision decoder dependencies. Gate Go limit parity and production CMake wiring without model downloads. Assisted-by: OpenAI:GPT-5 * test(decisions): cover multimodal public API paths Exercise shared image contracts through the registered HTTP routes and external mock backend. Add opt-in cached gallery installation and real OpenJev image decisions through SystemOne and both routing APIs. Assisted-by: Codex:gpt-5 * test(decisions): assert isolation and cache bypass Observe external RPC calls and compare complete classifier history. Winner-only and cache-miss checks could hide dropped history or cache use. Give real inference its own application and model directory so shared backend mappings and loaded processes cannot affect mixed suite order. Assisted-by: OpenAI:ChatGPT * test(decisions): isolate fixture globals Disable optional global services in the isolated HTTP fixture and register cleanup before setup assertions. Verify meter provider identity survives fixture creation and destruction. Snapshot observed usage before assertions so failures cannot retain the mutex. Require a successful usage stamp before checking error responses. Assisted-by: Codex:gpt-5 golangci-lint * fix(application): honor optional telemetry controls Skip failover gauge registration when metrics are disabled. Register against the application meter rather than looking up the global provider. Allow embedders to retain the bounded routing log without billing stats. Keep the existing default when stats are disabled. The isolated HTTP fixture uses this option without losing its native router assertions. Assisted-by: Codex:gpt-5 golangci-lint --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
f035746db9 |
docs: ⬆️ update docs version mudler/LocalAI (#12458)
⬆️ Update docs version mudler/LocalAI Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
e0e21de3e9 |
chore: ⬆️ Update CrispStrobe/CrispASR to 199de52d9068aeba366dd95d09a96bb45b2a7a14 (#12459)
⬆️ Update CrispStrobe/CrispASR Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
c99414019f |
chore(model-gallery): ⬆️ update checksum (#12462)
⬆️ Checksum updates in gallery/index.yaml Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
def5a0085e |
fix(packaging): skip Linux GPU libraries on Darwin (#12466)
Return before initializing Linux library helpers on macOS. The system Bash lacks associative arrays, and Metal uses system frameworks. Assisted-by: Codex:gpt-6 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
ed1fbc5674 |
fix(ci): list all release digest artifacts (#12467)
Large release runs exceed the artifact API limit of 1,000 items. Use authenticated public API pagination and set its limit from the run artifact count so manifest jobs can find their existing digests. Assisted-by: Codex:gpt-6 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
6c718fdda7 |
feat(parakeet-cpp): VAD (Silero and the Moondream head), vad_model option and gallery entries (#12463)
* feat(parakeet-cpp): implement the VAD call and add the vad_model option The backend now serves the VAD gRPC call (POST /vad and /v1/vad) with the standalone VAD of libparakeet. It accepts a Silero VAD GGUF as the model file, or an ASR model with a VAD head (Moondream Ultra and Redux). The request audio is float32 PCM at 16 kHz; the response lists the speech segments in seconds, like the silero-vad backend. A model with neither fails the request with the library message. The vad_threshold, vad_min_pause, vad_min_speech, vad_speech_pad and vad_max_segment options tune the segmenter. Unset values keep the defaults of the detector in use, and a bad value fails the load. The vad_model option names a Silero GGUF, resolved against the models directory like the other companion files. It lets any ASR model cut long audio at pauses through parakeet_capi_transcribe_path_json_vad_with, and it implies vad. vad:true alone still uses the model's own head. The new symbols are probed like the existing optional ones. A library without them still loads; the feature that needs one fails with a clear message only when it is used. Assisted-by: Claude:claude-sonnet-5-5 [go test] * feat(gallery): add parakeet-cpp VAD entries and a v3 plus Silero example Add VAD-only entries for the parakeet-cpp backend: the VAD heads of Moondream Redux (packed, CPU) and Ultra (Q8_0), which share their files with the existing ASR entries, and Silero VAD v6.2.3 as a GGUF (MIT, Silero Team). The parakeet-cpp-vad entry installs Silero; it has no variants, because variant ranking prefers the larger build that fits and these are different detectors. Add parakeet-cpp-tdt-0.6b-v3-silero-vad, a v3 entry that sets vad_model so long audio is cut at pauses by Silero. The Silero GGUF entries point at the intended Hugging Face URL of the file; the existing silero-vad entries are unchanged. A test checks the usecases, the shared files and the default entry and the vad_model reference. Assisted-by: Claude:claude-sonnet-5-5 [go test] * docs: describe parakeet-cpp VAD and the vad_model option Document the VAD endpoint on the parakeet-cpp backend (Silero GGUF and the VAD heads of Moondream Ultra and Redux), the vad_* tuning options, and the vad_model option that lets an ASR model without a VAD head cut long audio with Silero. Assisted-by: Claude:claude-sonnet-5-5 * chore(parakeet-cpp): bump parakeet.cpp to 6165e3d Pin the release that adds the standalone VAD (Ultra/Redux head and Silero) and the C API calls the backend now uses. Assisted-by: Claude:claude-sonnet-5-5 --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
00faf4ec26 |
feat(parakeet-cpp): Moondream Ultra and Redux gallery entries, VAD option, pin bump (#12455)
* chore(parakeet-cpp): bump parakeet.cpp to 11c1a0f Picks up support for the Moondream Ultra and Redux models and the VAD-segmented transcription entry point (C-API ABI is still 10). Assisted-by: Claude:claude-sonnet-5-5 * feat(parakeet-cpp): add vad option for long-audio transcription Models with a VAD head (Moondream Ultra and Redux) can cut long audio at pauses. Setting vad:true in the model options routes offline transcription through parakeet_capi_transcribe_path_json_vad. The symbol is probed at startup like the other optional entry points, and vad:true fails the load with a clear message when the library lacks it. A model without a VAD head fails the request with the library's own message. The option is off by default and does not affect streaming. Also say in the load error that a packed ternary Redux model is CPU only, because the library reports its refusal on a GPU backend through its log, not through the C API. Assisted-by: Claude:claude-sonnet-5-5 * feat(gallery): add Moondream Ultra and Redux for parakeet-cpp Add five entries from the public parakeet-cpp GGUF repository: Ultra in F16 and Q8_0, and Redux as packed ternary (CPU only, offline only) and as dequantized F16 and Q8_0 (any backend). The entries enable vad:true so long audio is cut at pauses. Checksums come from the repository's LFS metadata. The weights are CC-BY-4.0. Assisted-by: Claude:claude-sonnet-5-5 * docs(parakeet-cpp): document Moondream Ultra, Redux and the vad option Assisted-by: Claude:claude-sonnet-5-5 --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
9e373dba33 |
chore: ⬆️ Update PrismML-Eng/llama.cpp to 2459f68b5c0eb26261fd5a81682004b93cd645ba (#12341)
⬆️ Update PrismML-Eng/llama.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
37eb16eb1b |
chore: ⬆️ Update CrispStrobe/CrispASR to 966561aa596cfc653aa0e9885d44117fad9cca35 (#12437)
⬆️ Update CrispStrobe/CrispASR Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
d5f2f607b9 |
chore: ⬆️ Update ggml-org/whisper.cpp to 60c0be6ac8fa71b1a2ae2dd938a31a34a508e774 (#12439)
⬆️ Update ggml-org/whisper.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
372b1f8983 |
feat(silero-vad): allow threshold/silence/pad via model options (#12430)
* feat(silero-vad): allow threshold/silence/pad via model options Signed-off-by: anton ziderer <Antonziderer@mail.ru> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(silero-vad): ignore NaN thresholds NaN passes the detector validation but prevents speech comparisons from succeeding. Ignore it like malformed input and document the option validation. Convert the option tests to Ginkgo and cover invalid overrides. Assisted-by: Codex:gpt-6 Signed-off-by: anton ziderer <Antonziderer@mail.ru> --------- Signed-off-by: anton ziderer <Antonziderer@mail.ru> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
f52cc1b96a |
feat(swagger): update swagger (#12433)
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
e4fa051ee3 |
refactor(distributed): put the NATS-only paths behind interfaces (#12395)
* feat(messaging): add shared subject rules Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(messaging): cover BroadcastRoots, ControlRoots and SubjectRoot Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(messaging): add Broadcaster and enforce subject rules in every carrier Broadcaster is the fan-out half of MessagingClient. The NATS client and the in-memory FakeBus now refuse a subject outside the served roots and any wildcard other than a whole single token, and FakeBus shares MatchSubject instead of its own copy. FakeBus Unsubscribe now removes its own subscription instead of the first one with the same subject. A shared conformance suite in messagingtest runs against both carriers. The distributed e2e specs that used invented test.* subjects, and the one that subscribed with a > filter, now use subjects from subjects.go. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: depend on Broadcaster where only publish and subscribe are used Narrowed to messaging.Broadcaster: nodes/staging_progress.go, nodes/install_progress_publisher.go, galleryop/operation.go, galleryop/service.go, agentpool/user_services.go, agentpool/agent_jobs.go, openresponses/store.go, openresponses/sync.go, syncstate/syncstate.go, finetune/service.go, quantization/service.go and failover/distsync/distsync.go. SubscribeJSON now takes a Broadcaster because it only calls Subscribe, which lets the narrowed consumers use it. Stayed wide: worker/supervisor.go, because its client field also serves the SubscribeReply handlers in worker/lifecycle.go. The request/reply, queue and wiring files (nodes/unloader.go, nodes/file_stager_s3.go, jobs/dispatcher.go, agents/dispatcher.go, agents/events.go, worker/file_staging.go, cli/agent_worker.go, http/app.go) are unchanged by design. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): name the no-route condition and confine the carrier error Consumers matched nats.ErrNoResponders, which names an absence, to demote a node. They now match ErrNoRoute, the control path maps the carrier's failure onto it, and timeouts and worker refusals are pinned as not being no-route. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(nodes): state which FileStager implementations return ErrNoRoute Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): build backend clients through one node-aware seam Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: describe the distributed transport seams Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: correct comments that overclaim after the seams refactor Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(agent-worker): refuse an unserved LOCALAI_AGENT_SUBJECT at startup The messaging client now refuses a subject whose root no carrier serves. An agent worker started with a custom LOCALAI_AGENT_SUBJECT such as tenant-a.agent.execute used to start and then wait on a subject the frontend never publishes to. After the subject rules landed it exited at subscribe time with an error that did not name the setting. Behaviour change: the worker now checks LOCALAI_AGENT_SUBJECT before it registers or connects, and exits with an error that names the variable and says to use a served subject under the agent root, for example agent.execute. The served roots are not widened: a custom root was never delivered by the frontend, and a wider set would reopen the drift the subject rules exist to close. The flag help and the agent worker docs state the constraint. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(nodes): pin the reactions to ErrNoRoute Three callers react to ErrNoRoute and had no spec: the reconciler's upgrade drain falls back to the legacy forced install, the reconciler marks the node unhealthy when a pending op has no route, and the backend-op fan-out marks the node unhealthy. Each spec drives the real caller with a scripted no-responders reply and reads the result from the registry or the recorded requests. A fourth spec pins the other side: a pending op that times out leaves the node healthy and only counts the attempt, so mapping timeouts onto ErrNoRoute would fail here. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(messaging): pin client subject checks and fail the carrier suite in CI Add specs that call Publish, Request, Subscribe, QueueSubscribe, SubscribeReply and QueueSubscribeReply on a client with no connection. Each call must return ErrUnservedSubject for bogus.thing and ErrUnsupportedWildcard for jobs.>. This proves that the subject check runs before the connection is used, and needs no server. The NATS conformance suite is the only check that runs the subject rules against a real carrier. Before this change it skipped without output when Docker was missing. Now it fails when CI is set, so a Linux runner without Docker cannot hide it. It still skips on local runs and on macOS CI, which has no Docker. Add SubjectNodeBackendInstallProgress to the list of constructors that must build served subjects, and ask contributors to extend the list. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: state what ErrNoRoute may change, and group the distributed guides The seams note said MarkUnhealthy was the only state change allowed on ErrNoRoute. A pending backend op still records the failed attempt, counts toward the reconciler's retry limit and is dead-lettered after the maximum attempts. The note now says that MarkUnhealthy is the only change to the node's own state, and that the per-op accounting is not a verdict about the node. The note also documents that the NATS conformance run fails under CI when Docker is missing. The distributed-seams row moves next to the distributed-state row in the topics table. The liveness ping spec header now says no route is a reason to skip the worker, not proof that the worker is gone. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): give the backend client factory the node id Mechanical: the method gains a nodeID parameter and the eight test fakes are updated. No behaviour change. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): drop the optional node-aware factory The node id is now in the main method, so the optional interface and its helper had no behaviour of their own. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): dial backend probes through the client factory Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): dial workers' file servers through a per-node dialer Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(http): proxy backend logs through the per-node worker dialer The admin backend-logs proxy (list, lines and the WebSocket stream) now reaches a worker through the same per-node dialer as the HTTP file stager, so every frontend-to-worker dial goes through one seam. The shared direct dialer keeps alive for 15s where the proxy used 30s. Harmless for requests bounded at 15s. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(http): keep the backend-logs proxy independent of the admin connection The proxy request had no context before the dialer change and is bounded only by its 15s timeout. Keep it that way. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: move the worker control payloads to workerctl Mechanical move of the request and reply structs, the install progress event and the file payloads out of messaging. The verbs no longer belong to one carrier. No alias is left behind. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(worker): serve the lifecycle verbs through a controlServer The worker registers one handler per verb and a NATS server maps each verb to its subject. Registration errors now name the verb. node.stop is served with SubscribeReply, which is identical on the wire because the handler never replies. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(worker): report install progress through the control sink Install and upgrade now emit download progress through the sink the control server hands them. The debounce and the terminal flush stay in the handler path, built over that sink by the new nodes.NewDebouncedInstallProgressSink, which replaces NewDebouncedInstallProgressPublisher. The subject and payload on the wire are unchanged. The supervisor no longer holds the bus, and installFn and upgradeFn let specs drive both verbs without a gallery. The malformed-request log lines are restored for install, upgrade, backend.delete, model.unload, model.stop and model.delete, with the reply bytes unchanged. The signal adapter is renamed noReply, which also lets worker.go import os/signal without an alias again. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(worker): serve the file-staging verbs through a controlServer An empty list-dir answer is now {} rather than {"files":null}, because the typed reply omits an empty Files slice. The frontend decodes both to a nil slice in nodes/file_stager_s3.go. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(messaging): add WorkQueue and the NATS producer This is the producer side of the competing-consumer seam. The work kinds map one to one to today's subjects and queue groups: task to jobs.new and mcp-ci to jobs.mcp-ci.new (both in group workers), agent-run to agent.execute (group agent-workers). Enqueue publishes the payload as Publish does today, with one JSON marshal. FakeBus now records queue groups and keeps reply handlers so later specs can pin and drive them. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(messaging): add the NATS WorkConsumer An in-flight limit of one runs the handler inline on the delivery goroutine, as the MCP CI consumer does today. Any other limit spawns per delivery, as the agent consumer does. Queue groups are unchanged. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: publish queued work through WorkQueue The job dispatcher, the agent pool and the agent scheduler enqueue through messaging.WorkQueue; the NATS implementation publishes to the same subjects as before. DistributedServices builds the queue next to the NATS client and hands it to the dispatcher and the agent pool, whose distributed mode switch now reads a non-nil WorkQueue. The unused AgentPoolService.SetNATSClient is removed. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: consume queued work through WorkConsumer The agent dispatcher and the MCP CI consumer register through messaging.WorkConsumer. The NATS implementation keeps the inline one-at-a-time model for MCP CI and the per-delivery model for agent runs. handleMCPCIJob reports on the events publisher the carrier hands it instead of a captured client. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: delete the consumers nothing in production reached jobs.new has a producer and no production consumer, and the agent dispatcher's Dispatch was only called from tests. Publishing jobs.new is unchanged. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(mcp): send MCP requests to agent workers through AgentControl Timeouts still honour only the deadline, not cancellation, exactly as today. The NATS no-responders error maps to ErrNoRoute and a timeout does not. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(agent-worker): serve MCP requests and backend.stop through agentRPCServer The agent worker's MCP tool and discovery reply subscriptions and its backend stop listener move behind an unexported agentRPCServer interface, served on NATS by nodes.NATSAgentRPCServer. The handlers become typed mcp.ToolHandler and mcp.DiscoveryHandler values that answer every failure with a reply carrying Error. Queue group (agent-workers), inline execution on the delivery goroutine, the background handler context, the unmarshal error reply texts and the reply-less backend stop subscription are unchanged. The backend stop handler takes the decoded backend name, so it can still close that backend's MCP sessions. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(messaging): remove helpers that only tests used BroadcastRoots, ControlRoots and SubjectRoot had no production caller. The roots spec now asserts every served root through ValidateSubject instead. MatchSubject moves back into the test support package, the only place that used it, with its table. NATSAgentRPCServer drops the subscription list it stored and never read, and NewNATSAgentRPCServer gets a doc comment. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(mcp): round trip the agent RPC server over a real NATS server One spec sends a tool request and a discovery request through NATSAgentControl to NATSAgentRPCServer and checks that the handlers see the decoded requests and the replies come back. It also puts an undecodable body on the tool subject and checks the server answers with an unmarshal error instead of leaving the requester to time out. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: describe the distributed transport seams The developer note now lists the final seams: fan-out, queues, both halves of the control verbs and of agent RPC, and the dial. It records the open items a second carrier has to handle. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test: pin the in-flight limit each queue consumer asks for The work queue specs pin what Consume does for a given limit, but nothing pinned which limit each production consumer passes. Changing the agent worker's MCP CI limit from 1 to 0 would have let MCP CI jobs run concurrently on each worker with every test green. Move the MCP CI Consume call into startMCPCIConsumer with the same wiring and pin that it asks for (WorkMCPCI, 1). Pin that NATSDispatcher.Start asks for (WorkAgentRun, maxConcurrent) for several limits. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: remove helpers the branch left without a caller SubjectJobCancelWildcard lost its last subscriber when the frontend stopped listening on jobs.*.cancel; the NATS permissions and conformance suite spell the subject out, so nothing reads the constant. decodeBackendStopRequest returned a stopAll flag that production dropped and only a test read. decodeBackendStop is now the single decoder with the same semantics: an empty body is stop-all, an empty Backend is stop-all, malformed JSON is an error. stopBackends still derives stop-all from Backend, so no reply changes. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(messaging): keep an explicitly empty agent queue a plain subscription Before the work queue seam the agent worker passed LOCALAI_AGENT_QUEUE straight to QueueSubscribe, so an explicitly empty value made a plain subscription and every agent worker ran every agent run. WithAgentRunRoute replaced an empty queue with agent-workers, which silently changed that. Keep the queue as given once the option is applied. An empty subject still falls back to agent.execute, since it never had a meaning of its own. The flag default stays agent-workers, so only an explicitly empty value reaches this. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: correct comments and record the PR B notes Fix the recordingFactory comment (it also records the parallel flag), document that a negative maxInFlight is unbounded and that Unsubscribe from a handler deadlocks, and say a permanently undecodable payload returns nil. Record controlHandler's undecodable return as a kept exception, and add the second carrier notes to the developer note: the reconciler has no ClientFactory option, the logs proxy honours HTTP_PROXY, verbs one carrier serves need an opt-out, terminal replies come from the result event, and agent runs publish through the NATS-bound EventBridge, which is not an additive change. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
776d9cc30b |
chore(model-gallery): ⬆️ update checksum (#12441)
⬆️ Checksum updates in gallery/index.yaml Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
2c2da7ee0c |
chore: ⬆️ Update ikawrakow/ik_llama.cpp to 5f89bfc81268b4d56d2af63ccbed59de17c64c09 (#12440)
⬆️ Update ikawrakow/ik_llama.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
c65d64e766 |
docs: publish the LocalAI 4.11 release post (#12442)
* docs: publish the LocalAI 4.11 release post Highlight speaker enrollment, model failover, and single-host operations with short UI recordings. Assisted-by: nib:gpt-5.6-sol Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: improve the LocalAI 4.11 release post Match the direct feature-post style and replace empty UI recordings with populated speaker and failover examples. Assisted-by: nib:gpt-5.6-sol Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
58830f7ac5 |
docs: introduce decision models on the blog (#12428)
Explain decision-model uses and the current gallery choices. Link the setup guide and include a bounded API example with access requirements and model-confidence caveats. Assisted-by: OpenAI:undisclosed [shell] Co-authored-by: Ettore Di Giacinto <mudler@localai.io>v4.11.0 |
||
|
|
38a2aa45fe |
feat(gallery): add Nimble 9B and CLM decision models, bump vllm.cpp to a19294a9 (#12397)
* feat(gallery): add nimble-9b-vllm-cpp decision model Add Bespoke Nimble 9B, converted for vllm.cpp and pinned to the weights commit 52eead25 of mudler/Bespoke-Nimble-9B-vllm-cpp (HEAD only adds the model card). It is a redistribution of bespokelabs/Bespoke-Nimble-9B with the LoRA merged into Qwen3.5-9B; config.json names NimbleModel, so no hf_overrides are needed. The artifact sits under overrides, where the installer reads it. The entry sets an 8192-token context, Nimble's own prompt limit, and a KV pool of 1024 blocks of 32 tokens for 4 sequences (about 1 GiB at 32 KiB per token for the 8 full-attention layers). Installed with local-ai models install and served on CPU through the vllm-cpp backend: the model card's billing request gives billing (0.986), refund 0.998 and urgency 0.33. Peak resident memory was 18.4 GB, so the description asks for about 20 GB of free RAM. List the entry in the decisions gallery table. CLM stays out of the gallery: the pinned engine cannot load the published head layout. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-sonnet-5-5 * feat(gallery): add clm-v0.1-8b-vllm-cpp and bump vllm.cpp for CLM The CLM checkpoint on mudler/CLM-v0.1-8B-vllm-cpp stores the heads with the reference's own tensor names (state_head.inp, hidden.N, norms.N, out). The pinned vllm.cpp 96788348 still expects the old .0/.2/.4/.6 layout and refuses the load with "head.safetensors incomplete for state_head". vllm.cpp a19294a9 matches the reference layout and adds the converter that produced the upload, so move the pin there. The ABI stays at v30. Add the CLM entry, pinned to the weights commit 0d1903b1 (HEAD only adds the model card), with a 4096-token context and a KV pool for 4 sequences (about 2.25 GiB at 144 KiB per token for Qwen3-8B). Installed with local-ai models install and served on CPU against a libvllm built at a19294a9: the model card example (john works at google, entity type) gives person 0.950, the same as the card. Peak resident memory was 17.9 GB. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-sonnet-5-5 --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
5c28659fd9 |
chore(model-gallery): ⬆️ update checksum (#12415)
⬆️ Checksum updates in gallery/index.yaml Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
a4cf94ecc5 |
chore: ⬆️ Update ggml-org/llama.cpp to a868c3e3c56657f7e8a6231190dbbe90e7dd86c0 (#12419)
⬆️ Update ggml-org/llama.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
ba091ed3fa |
chore: ⬆️ Update ikawrakow/ik_llama.cpp to d9e286846d6f8232db48ec5c111a4ea3aea675ef (#12418)
⬆️ Update ikawrakow/ik_llama.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
6aac3646be |
chore: ⬆️ Update mudler/parakeet.cpp to bee7c14dfcc23613df58176c59a40459e7b47095 (#12420)
⬆️ Update mudler/parakeet.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
c3bea567fe |
fix(llama-cpp): do not stream the error text as content on pre-stream failures (#12425)
When the first result of a streamed request is an error (for example a prompt that exceeds the context), PredictStream wrote the error message as a Reply and only then returned the error status. LocalAI treated that Reply as the first token: it sent the assistant role chunk and the error text as `content` on an HTTP 200 stream. Because a chunk had already been written, the pre-stream HTTP error path from #12204 never triggered, so streaming clients still got a 200 with the error as model output, while the same request without streaming correctly returns a 400. Return the error only as the gRPC status. The e2e backend suite gets a `context_overflow` capability (enabled for llama-cpp) that streams an over-long prompt and asserts an error status with no content. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> |
||
|
|
1056c62f4c |
fix(cloud-proxy): surface Anthropic refusals instead of empty replies (#12424)
When an Anthropic model declines to answer, the Messages API returns HTTP 200 with `stop_reason: "refusal"` and empty `content`. The translate mode mapped that to a normal reply with no content, so the OpenAI-compatible response looked like a successful completion (`finish_reason: "stop"`, empty message). Routers, agents and UIs could not tell "the model declined" from "the model had nothing to say", and no fallback was triggered. Return an explicit error for `stop_reason: "refusal"` in both the non-streaming path and the streaming path (`message_delta`). Regular replies, including empty `end_turn` replies, are unchanged. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> |
||
|
|
985df46d24 |
fix(whisperx): keep the transcript when diarization fails, report real errors (#12427)
AudioTranscription caught every exception and returned an empty TranscriptResult. A failed diarization step therefore discarded a transcript that was already finished: with an HF token that has not accepted the terms of the gated pyannote pipeline, the download fails with 403 and every transcription came back as an empty text with HTTP 200. Diarization now degrades: if it fails, the transcript is returned without speaker labels and the reason is logged. Any other failure aborts the call with INTERNAL instead of pretending success with an empty text. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> |
||
|
|
6aa7b9b871 |
fix(llama-cpp): let parallel:1 in the model options win over LLAMACPP_PARALLEL (#12426)
The environment fallback was applied whenever n_parallel was still 1 after option parsing. An explicit `parallel: 1` in the model YAML is indistinguishable from the default that way, so it was replaced by LLAMACPP_PARALLEL. The docs say options in the YAML take precedence over environment variables; a single model could not be forced to one slot while the global variable was set. Track whether the options set the slot count and resolve it in a small helper (parallel_params.h): option first, then LLAMACPP_PARALLEL, then 1. The helper gets a standalone unit test picked up by `make test-backend-cpp`. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de> |
||
|
|
dd19ee8912 |
chore: ⬆️ Update CrispStrobe/CrispASR to fdc3a0007d68f8d3905f20e9cd4d18b5f193e096 (#12413)
⬆️ Update CrispStrobe/CrispASR Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
9eb5a9e61d |
feat(audio): remember speakers from diarization (#12414)
* feat(schema): validate portable speaker profiles Add the versioned profile schema for explicit speaker enrollment. Validate compatibility against separately supplied loaded-encoder metadata. Reject unusable speakers, invalid vectors, and inconsistent clean spans. This slice does not change HTTP routes, backend integration, or the UI. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(parakeet): export profiles with transcripts Export opt-in speaker profiles and trusted encoder metadata. Replay registrations by ID so duplicate display names keep independent vectors. Use one profile-capable diarization for slots, names, and clean spans. Assign timestamped ASR words to those slots without a second diarization. Preserve legacy opt-out and no-ASR behavior, and propagate failures. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(audio): enroll portable speaker profiles Gate profile exports with voice-recognition permission and validate registration against metadata from the loaded encoder. Preserve audio enrollment and independent registrations with duplicate display names. Exclude diarization and registration exchanges before API trace capture so persisted traces cannot retain profile vectors or JSON audio. Defer candidate dimensions to trusted loaded metadata. Sort candidates by registration ID so incompatible profiles cannot suppress legacy voices through registry iteration order. Keep portable identity checks closed when trusted metadata is unavailable. Test persisted traces, explicit slot zero, and selection through offline and live transport. Document privacy and the ephemeral registry lifecycle. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(ui): remember speakers from diarization Add a Studio page for diarization and opt-in speaker profiles. Preview clean intervals from the original recording before explicit registration. Join profiles by raw speaker labels, preserve duplicate names, and relabel turns only after a successful save. Discard stale results when the model or recording changes. Share registration metadata with voice management without storing vectors or recordings from this flow. Document permissions and the global, ephemeral registry. Cover enrollment, permissions, previews, and asynchronous races with mocked Playwright tests. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: clarify HTTP speaker enrollment support Replace the stale enrollment limitation with the current HTTP workflow. Distinguish native transport from explicit registration and link its docs. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(parakeet): pin merged speaker profile support Use the merged commit from mudler/parakeet.cpp#80. Its tree matches the previously accepted native pin. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: add diarization enrollment setup example Connect the existing gallery modes to the speaker enrollment workflow. Show installation, private profile export, explicit raw-slot registration, and later recognition without another export. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): explain diarization speaker profiles Put the diarization walkthrough on the LocalAI website in the feature PR. Cover the three gallery modes, explicit enrollment, and privacy limits. Link setup instructions and keep availability conditional on feature support. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): focus diarization on everyday use Explain what users can do with recordings before the setup steps. Replace the technical walkthrough with a short Studio guide and link readers to the existing reference for model names and developer use. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(blog): lead with speaker capabilities Present speaker recognition through everyday uses and a short UI flow. Keep technical reference details in the existing documentation. Assisted-by: OpenAI:unknown Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(diarization): satisfy Go lint checks Avoid copying protobuf message state when extending backend status, check the multipart reader close result, and document the focused testing.T lint exemptions. Assisted-by: nib:gpt-5.6-sol Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
53cd6c5716 |
fix(funasr): select Python 3.12 tokenizers
FunASR left Transformers and Hugging Face Hub unconstrained, which let uv backtrack to tokenizers 0.10.3 without Python 3.12 wheels. Keep Transformers on the supported 4.x range, including the Intel upgrade profile.\n\nAssisted-by: nib:gpt-5.6-sol\nSigned-off-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
668802d6ca |
fix(watchdog): ignore stale backend evictions (#12333)
Remove watchdog tracking when backends stop, crash, or fail to start. Validate eviction addresses under the model lifecycle lock so delayed shutdowns cannot stop a replacement backend. Preserve replacement size estimates when removing an old address. Add lifecycle regression tests and document shutdown behavior. Fixes #12331 Assisted-by: Codex:gpt-6 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
b9c975e71c |
chore: ⬆️ Update ggml-org/llama.cpp to a4d880fd5c7f88713ded6db9f0111893bd78afa6 (#12345)
* ⬆️ Update ggml-org/llama.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> * fix(llama-cpp): migrate score batches to the new API The pinned engine removes common_batch_add and the raw batch view. Use common_batch entries and llama_process for score suffix decoding. Read shared-prefix scores from the current common_batch view. Validation: reproduce both compiler errors on the original patch. The patched server context and complete grpc-server translation unit pass g++ -std=c++17 -fsyntax-only with generated protobuf headers. Assisted-by: Codex:gpt-6 --------- Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
3d0311e640 |
docs: fix Discord, Slack and Telegram example links (#12392)
The chat bot examples moved to mudler/LocalAI-examples; the old paths in LocalAGI and LocalAI no longer exist. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Pratik Gandhi <travpreneur@gmail.com> |
||
|
|
2ae6cae70d |
feat(parakeet-cpp): name speakers from the shared voice registry (#12382)
* feat(voice): list registered voices and record which encoder made them The voice registry could register, identify and forget but not list, and it did not remember which speaker encoder produced an embedding. Add Metadata.Model and Registry.List, answered from the index the store registry already keeps for Forget. Needed so a backend can be given the registered voices that match its own speaker encoder. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(voice): store the encoder model with a registered voice Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(voice): pick the registered voices that match a speaker model Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(proto): carry known voices and speaker names on diarize and live messages Assisted-by: Claude:claude-haiku-4-5 [Claude Code] * feat(diarization): name speakers from the voice registry When a diarization model has a speaker_model option, the endpoint sends the registered voices made by that encoder to the backend. The backend's name and name_score come back as extra fields next to the normalized SPEAKER_NN speaker, and the speakers summary carries the first name seen for each speaker. RTTM output and results without names are unchanged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(live): pass registered voices to a live session and surface speaker names Live sessions now send the registered voices that match the model's speaker_model to the backend, and each speaker segment carries the name the backend matched. The realtime segment event gains an optional speaker_name field. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): load a speaker model and build per-request voice registries Adds the speaker bindings (ABI v9 and v10, probed separately), the speaker_model, speaker_threshold and speaker_margin options, and a per-request registry builder over the known voices. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): name the speakers in Diarize from the known voices Diarize builds a per-request speaker registry from the known voices when a speaker model is loaded, calls the named C functions, and puts each slot's registered name and score on the segments. The registry is freed on every path. A library without ABI 10 reports Unimplemented instead of dropping the names. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): name speakers in the live scene stream The live scene stream now begins with a known-voice registry when a speaker model is loaded and the live config carries voices, and each closed speaker segment takes its slot's current name from the feed's names map. A segment that closes before its slot is identified has an empty name. The registry is freed after the stream, on every path. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(gallery): speaker naming entries and docs for parakeet-cpp Add three gallery entries that load the WeSpeaker ResNet34 speaker model next to the diarization or realtime scene models, and document speaker names in the voice recognition, diarization, audio to text and realtime pages. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(parakeet-cpp): skip an unusable registered voice instead of failing the request A registered voice with the wrong embedding size, or one the C side refused, failed the whole diarization request, so one legacy voice broke the model for every user. Skip such voices with a warning that does not carry the voice name, and take the plain path when none is left. Also map an exact 0 speaker threshold or margin to a tiny positive value, since the C side reads 0 as "use the default", and fix a stale comment about which contexts Free() walks. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(diarization): warn once per model about voices from another encoder; document the privacy limit The different-encoder warning fired on every request. Log it once per feature and speaker model, then at debug level. Document that the global voice registry lets any caller of a speaker_model model learn matching names, and that skipped wrong-sized voices are logged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * chore(parakeet-cpp): bump parakeet.cpp to 8c8cec0 (C-API v10) and check speaker naming against the real library The pin moves from 623a968 to 8c8cec0, which brings in everything merged in parakeet.cpp since: the voice identification change (C-API v9, #78) and raw-embedding enroll plus diarize-only speaker naming (C-API v10, #79). New real-library specs (gated on PARAKEET_BACKEND_TEST_SPEAKER_MODEL, _DIAR_MODEL, _WAV and, for the live path, _STREAM_MODEL) name the two speakers of two_speakers.wav from a committed pair of WeSpeaker embeddings, with the voices passed in reversed order. They also check that the float32 threshold reaches C through purego. The shared test loader now registers the v9/v10 and scene symbols as main.go does. The rebase onto origin/master had no conflicts. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |