Commit Graph
7948 Commits
Author SHA1 Message Date
dependabot[bot] 40c4d5990e chore(deps): bump scipy in /backend/python/transformers
Bumps [scipy](https://github.com/scipy/scipy) from 1.15.1 to 1.18.0.
- [Release notes](https://github.com/scipy/scipy/releases)
- [Commits](https://github.com/scipy/scipy/compare/v1.15.1...v1.18.0)

---
updated-dependencies:
- dependency-name: scipy
  dependency-version: 1.18.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-11 23:06:18 +00:00
localai-org-maint-bot 7dc617c7c1 fix(transformers): use Python 3.12
SciPy 1.18 requires Python 3.12 or newer. Keep this backend on a
compatible portable Python for Linux and macOS builds.

Assisted-by: Codex:gpt-5.6 [Codex]
(cherry picked from commit 70bf6d4a3a)
2026-09-11 23:05:33 +00:00
81041d8b1b chore(deps): bump LocalAGI and localrecall to v0.6.5 (#11985)
* chore(deps): bump github.com/mudler/localrecall to v0.6.5

Picks up two Postgres engine fixes: the RRF fusion no longer scores
every hybrid-search candidate 0 through integer division, and the
search_vector text config is no longer pinned to 'simple' for the life
of the process after one transient lookup failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANmTgdYikmzBq67jVJ1CgX

* chore(deps): bump github.com/mudler/LocalAGI to d93d478

Picks up mudler/LocalAGI#493, which bumps localrecall to v0.6.5 there
too, so the direct pin in this module and the version arriving through
LocalAGI agree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ANmTgdYikmzBq67jVJ1CgX

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-12 01:05:13 +02:00
dependabot[bot] 1f3d89fc4a chore(deps): update transformers requirement
Updates the requirements on [transformers](https://github.com/huggingface/transformers) to permit the latest version.
- [Release notes](https://github.com/huggingface/transformers/releases)
- [Commits](https://github.com/huggingface/transformers/compare/v5.15.0...v5.15.1)

---
updated-dependencies:
- dependency-name: transformers
  dependency-version: 5.15.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-11 22:51:12 +00:00
Ettore Di Giacinto c9acb41903 fix(gallery): use tokenizer templates for Gemma
Let the model-provided tokenizer template format Gemma conversations instead of maintaining a shared inline prompt template.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 22:44:39 +00:00
Ettore Di Giacinto db12122329 fix(reasoning): ignore preclosed prompt markers
Do not seed streaming reasoning state when the latest prompt thinking marker is already followed by its matching closing marker. This keeps direct Gemma 4 output in content when its template disables thinking with a preclosed channel.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 22:44:39 +00:00
Ettore Di Giacinto bbd5488871 fix(distributed): make backend.stop report what it stopped
backend.stop was the one lifecycle subject a worker never answered. The
controller published and returned nil as soon as the local publish
succeeded, so a stop that killed nothing, and a stop that failed
outright, were indistinguishable from one that worked.

The unload endpoint calls model.unload and then StopBackend. Only the
first is acknowledged, so the endpoint answered 200 while the backend
kept running and held its VRAM, and its own "backend stop failed" branch
could never run. The worker logged the failure and nobody saw it.

Give the subject a reply. The worker now enumerates the process keys it
terminated and reports any per-process error, so StopBackend fails when
the stop failed. Resolving to nothing stays a success: stopping a backend
that is not running leaves the caller in the state it asked for, and
eviction paths stop already-gone models routinely. The empty list is what
says nothing matched, and ReportsStoppedProcesses is what makes that
emptiness trustworthy, the same way BackendDeleteReply handles it.

A worker built before this reply still receives the request and still
stops the backend, it only stays silent, so a timeout degrades to the old
assumption rather than failing every stop on a fleet mid-upgrade. Only
silence degrades: a transport error is still reported, because
UnloadRemoteModel skips its registry cleanup for a node it could not
reach and needs to keep hearing about that.

Assisted-by: Claude:claude-opus-5 golangci-lint
2026-09-11 22:44:11 +00:00
Ettore Di Giacinto 51c0a44bca gallery: apply PR #11574
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:43:46 +00:00
Ettore Di Giacinto d385a0bf8e gallery: apply PR #11494
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:43:42 +00:00
Ettore Di Giacinto a398c559b3 fix(gallery): persist inference defaults under parameters
Gallery installs merged family defaults at the YAML root and only re-marshaled them on the artifact path. Persist the defaults in the loader-visible parameters map for every install path while preserving authored overrides.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 22:43:10 +00:00
Ettore Di Giacinto 14103f1d78 gallery: apply PR #11983
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:56 +00:00
Ettore Di Giacinto 81f6898b69 gallery: apply PR #11960
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:52 +00:00
Ettore Di Giacinto c99d0d44cd gallery: apply PR #11958
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:48 +00:00
Ettore Di Giacinto ada207679a gallery: apply PR #11951
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:43 +00:00
Ettore Di Giacinto f528bf07dd gallery: apply PR #11949
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:38 +00:00
Ettore Di Giacinto c7ff1aabf0 gallery: apply PR #11947
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:33 +00:00
Ettore Di Giacinto 8b003539db gallery: apply PR #11946
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:29 +00:00
Ettore Di Giacinto d95b5d787f gallery: apply PR #11943
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:24 +00:00
Ettore Di Giacinto 9f100e9062 gallery: apply PR #11930
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:19 +00:00
Ettore Di Giacinto 08a87788b4 gallery: apply PR #11927
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:15 +00:00
Ettore Di Giacinto ab191912e9 gallery: apply PR #11926
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:11 +00:00
Ettore Di Giacinto 16fb9e81bd gallery: apply PR #11923
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:06 +00:00
Ettore Di Giacinto 075d0d3fa4 gallery: apply PR #11922
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:37:00 +00:00
Ettore Di Giacinto 9401844c6b gallery: apply PR #11913
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:36:55 +00:00
Ettore Di Giacinto 43cf7ed77f gallery: apply PR #11909
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:36:50 +00:00
Ettore Di Giacinto 29347bad67 gallery: apply PR #11905
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:36:46 +00:00
Ettore Di Giacinto 5a62ed1614 gallery: apply PR #11900
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:36:40 +00:00
Ettore Di Giacinto b9902148ae gallery: apply PR #11873
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:36:35 +00:00
Ettore Di Giacinto d81d7f821e gallery: apply PR #11841
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:36:30 +00:00
Ettore Di Giacinto c70e18392e gallery: apply PR #11832
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:36:26 +00:00
Ettore Di Giacinto 9e86521698 gallery: apply PR #11705
Assisted-by: localai-org-maint-bot:glm5.2 [gh]
2026-09-11 22:36:21 +00:00
Ettore Di Giacintoandxingyifeng df70ff311d feat: add FunASR speech recognition backend (#10090)
Adds FunASR/SenseVoice as a Python backend for speech-to-text with
support for CPU, CUDA 12/13, ROCm, Intel SYCL, L4T, and Apple MPS.

Co-authored-by: xingyifeng <xingyifeng@users.noreply.github.com>
2026-09-11 22:06:26 +00:00
Ettore Di Giacinto 6983477a71 fix(agent-ui): keep chat open for status
Agent Status replaced the chat route and unmounted its EventSource. Any response still in flight could then disappear from the conversation.\n\nOpen status in a separate tab so the chat keeps its live connection until the response completes.\n\nAssisted-by: Codex:gpt-5 [eslint]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:56:08 +00:00
Ettore Di Giacinto db09452d54 feat(tts): support multi-reference personalities
Saved profiles previously resolved to one audio path and transcript, so
cloning backends could not use several examples of one personality.

Store ordered audio and transcript pairs while preserving the legacy
first-reference fields. Fish Speech and audio.cpp receive all pairs,
including on distributed workers. Other backends retain their
single-reference behavior.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:56:08 +00:00
Ettore Di Giacinto ddda0436a3 fix(ui): exclude navigation from the TTS mock
The API mock also matched navigation to /app/tts and returned a WAV download instead of the React page. Let non-POST requests reach the test server.

Assisted-by: Codex:gpt-5 [Playwright]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:56:02 +00:00
Ettore Di Giacinto 1bcbff586b fix(ui): target the TTS input in history tests
TTS instructions add a second textarea to the page. Target the speech input by its placeholder so the history test does not depend on the page having one textarea.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:56:02 +00:00
Ettore Di Giacinto 919b5c96fa feat(ui): add per-request TTS instructions
Let studio users guide speech delivery for backends that support request instructions. Blank guidance stays out of requests and media history.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:56:02 +00:00
Ettore Di Giacinto d55474a149 feat(audio): list available TTS voices
Clients cannot discover the named voices that an installed TTS model accepts without consulting backend-specific documentation. Expose voice metadata through the audio API and let custom model configs declare their own catalog.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:55:21 +00:00
Ettore Di Giacinto 21e81434b3 fix(llama-cpp): stage model load diagnostics
The generated gRPC source tree omitted the new header and test. Every llama.cpp-derived backend therefore failed when grpc-server.cpp included the missing header.

Assisted-by: Codex:gpt-5.6 [systematic-debugging]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:55:14 +00:00
Ettore Di Giacinto a6519182db fix(llama-cpp): explain tensor count mismatch
llama.cpp reports the same tensor-count error for unsupported model layouts and damaged GGUF files. Add a focused hint so operators can update the backend or verify the model without losing the upstream diagnostic.

Assisted-by: Codex:gpt-5.6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:55:14 +00:00
Ettore Di Giacinto cce7a9e441 fix(fish-speech): resolve relocated source
The editable install records the backend build path, which does not exist after LocalAI relocates the packaged backend. Add the runtime source directory to PYTHONPATH so inference modules remain importable.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:55:10 +00:00
Ettore Di Giacinto f5eb09f329 fix(audio-cpp): bundle rocRoller for ROCm
ROCm 7.2 links rocBLAS consumers to librocroller.so.1. Add that runtime family to the ROCm bundle so packaged backends resolve the dependency.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:53:43 +00:00
Ettore Di Giacinto bfcf1e37ec fix(audio-cpp): install rocBLAS headers
The HIP build reaches ggml configuration and requires the rocBLAS CMake package. Install its development package with the existing hipBLAS dependency.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:53:43 +00:00
Ettore Di Giacinto db0b00a5ac fix(audio-cpp): normalize HIP target list
audio.cpp forwards GPU_TARGETS to CMake as a semicolon-delimited list. The comma-delimited LocalAI value was treated as one invalid HIP architecture during configuration.

Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:53:43 +00:00
Ettore Di Giacinto 70be2c49a1 fix(audio-cpp): install hipBLAS headers
The ROCm builder lacks the CMake package metadata that ggml requires. Install the development package only for hipBLAS builds.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:53:43 +00:00
Ettore Di Giacinto 878b99384f feat(audio-cpp): add ROCm backend image
The pinned audio.cpp revision supports HIP, but LocalAI neither builds a ROCm image nor accepts its backend option. AMD hosts therefore fall back to the CPU image.

Build and publish the HIP variant, connect it to AMD capability selection, and accept both upstream HIP names.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:53:43 +00:00
Ettore Di Giacinto a15780858e feat(diffusers): add AudioLDM2 generation
Expose diffusers audio pipelines through the existing sound-generation RPC. AudioLDM2 can now return PCM WAV output from the model gallery without a separate backend.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:53:37 +00:00
Ettore Di Giacinto 783556bc93 feat(prefixcache): index reported KV residency
Add a NATS event contract and exact-residency provider for backend KV cache reports. Keep guessed request observations as the default routing source while maintaining the reported index for future producers.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:53:33 +00:00
Ettore Di Giacinto f884d4b456 fix(mlx-video): use supported Python on Darwin
The pinned mlx-video package requires Python 3.11 or newer, while an empty PYTHON_VERSION selected the backend helper default of 3.10. Pin the available 3.11.13 portable runtime for the Darwin package build.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:52:41 +00:00
Ettore Di Giacinto 21df4a120a feat(mlx): add Apple Silicon video backend
Add a Darwin-only MLX-Video backend for LTX-2 and converted Wan checkpoints, expose it through the existing video API, and wire packaging, discovery, tests, docs, and an example.

Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-11 21:52:41 +00:00