Compare commits

...

214 Commits

Author SHA1 Message Date
localai-org-maint-bot
79cc9e7678 fix(fish-speech): use CUDA toolkit ptxas
Prefer an explicitly configured Triton assembler, otherwise use the executable ptxas from CUDA_HOME so torch.compile can target GPU architectures newer than Triton bundled tooling.

Assisted-by: Codex:gpt-5
2026-08-04 16:06:58 +00:00
localai-org-maint-bot
8f52437c81 fix(gallery): describe Genesis Hermes model accurately (#11342)
Replace copied HauhauCS base-model text with metadata for the actual Genesis Hermes V6 artifact and link its upstream base model.

Assisted-by: Codex:gpt-5 [Hugging Face]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-04 17:47:32 +02:00
mudler's LocalAI [bot]
cd516452dd fix(rocm): stop building the ggml CPU variant matrix for hipblas llama.cpp (#11346)
No -gpu-rocm-hipblas-llama-cpp image has been published since 2026-08-01.
Every build since has been killed by GitHub at exactly its 6h job limit:

    job 91830652349  cancelled  6h00m   (2026-08-04)
    job 91763226161  cancelled  6h00m   (2026-08-03)
    job 91466626154  cancelled  6h00m   (2026-08-02 full matrix)

The registry shows the damage: master-gpu-rocm-hipblas-llama-cpp last
built 2026-08-01 05:53, latest-gpu-rocm-hipblas-llama-cpp 2026-07-15,
against master-cpu-llama-cpp which is current.

Same cause as #11321, different mechanism. Since #11255 every x86 GPU
image also builds ggml's CPU_ALL_VARIANTS matrix. SYCL died because icpx
stalls on one translation unit; ROCm dies on volume. hipcc compiles the
HIP kernels once per entry in AMDGPU_TARGETS, and that list is eleven
architectures (gfx908, gfx90a, gfx942, gfx950, gfx1030, gfx1100, gfx1101,
gfx1102, gfx1151, gfx1200, gfx1201). The CPU matrix lands on top of that.

The numbers are unambiguous. The same job took 2h27m in the 2026-07-26
full matrix, before #11255. #11255 merged 2026-08-01 07:26, an hour and a
half after the last image was published, and it has been 6h00m ever since.
The tail of the last run shows it 61% through ggml-hip at the 83 minute
mark, still building HIP template instances.

Route hipblas to the portable fallback, exactly as #11321 did for SYCL and
for the same practical reason: it is what these images shipped before
#11255, and run.sh already prefers *-cpu-all when present and falls back
otherwise. Expected to restore the 2h27m build with room to spare.

Not fixed here: the CPU variant matrix is genuinely wanted on ROCm for
partial offload. Getting it needs the build to fit in 6h, which means
trimming AMDGPU_TARGETS or splitting the job per architecture. Both are
larger changes than unbreaking the image, and neither should ride along
with a build that is currently not shipping at all.

Verified: make test-build-scripts passes, including the extended
llama-cpp-build-target_test.sh. bonsai is unaffected (own compile script,
ROCm builds in 1h52m) and turboquant has no hipblas variant.


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-04 15:55:58 +02:00
mudler's LocalAI [bot]
3f0db2a9c2 feat(vllm-cpp): enable and vendor the MLX GEMM provider on darwin/metal (#11137)
* feat(vllm-cpp): enable and vendor the MLX GEMM provider on darwin/metal

The darwin vllm-cpp image built the Metal backend with vllm.cpp's native MSL
GEMM only. vllm.cpp also ships an optional MLX provider for the dense GEMM,
kept OFF upstream because it costs a ~19 MB libmlx.dylib plus a ~105 MB
mlx.metallib, on the stated position that it must earn that cost by
measurement.

Measured on an Apple M4 (16 GiB, macOS 26.5.2) it does. One binary, arms
toggled with VT_OP_PROVIDER_DISABLE=mlx so there is no build-difference
confound, Qwen3-1.7B-bf16 p=512 g=128, 2 reps, arm order alternated per rep:

  B=1   5.79 vs 3.08 agg tok/s (1.88x)   TTFT 3.32 s vs 7.68 s
  B=8   25.70 vs 13.69 (1.88x)           TTFT 13.95 s vs 34.38 s
  B=16  38.65 vs 17.69 (2.19x)           TTFT 18.33 s vs 54.48 s

Peak RSS is unchanged (6.65 to 7.50 GB in both arms) and the output is
bit-identical: vllm.cpp's three-way parity test measures mlx-vs-msl NMSE of 0
on all six shapes, and mlx-vs-cpu equal to msl-vs-cpu, against a 5e-4 bar. MLX
serves the dense GEMM alone; paged attention stays vllm.cpp's own kernel
because MLX has no paged-KV primitive. Full disposition, including the
INDICATIVE status and the isolation actually achieved, is in vllm.cpp
docs/BENCHMARKS.md "MLX GEMM provider A/B on Apple M4".

Build: MLX comes from the pinned prebuilt pip wheel (MLX_VERSION, default
0.29.3) into a venv under the backend dir. Building MLX from source needs
`xcrun metal`, i.e. a full Xcode the macOS runners do not have, while the wheel
ships include/, lib/libmlx.dylib and the compiled metallib ready to link. The
install is a stamp FILE rather than a phony target, because a phony
prerequisite is always newer than libvllm and would re-link it every
invocation. VLLM_CPP_MLX=off restores the previous Metal build.

Packaging vendors libmlx.dylib, mlx.metallib and MLX's MIT license into
package/lib/. Three things this had to get right, each verified on the M4
before it was written rather than after:

  1. libvllm.dylib links @rpath/libmlx.dylib and its build-time LC_RPATH points
     inside the build venv, a path no user has. Every build rpath is deleted
     and replaced with @loader_path/lib.
  2. MLX loads its metallib from beside its OWN dylib, so both files must land
     in the same directory or every Metal op fails with "Failed to load the
     default metallib".
  3. install_name_tool invalidates the code signature and macOS refuses to load
     an arm64 image with a stale one, so the patched library is re-signed
     ad-hoc.

Verified end to end on the M4 by building through this Makefile and running the
packaged artifact: `DYLD_PRINT_LIBRARIES` resolves libmlx from package/lib/,
`codesign -v` passes, no build-venv path survives in the load commands, and a
real generation runs with the provider selected (op=65 selected=mlx) and zero
metallib failures. A missing rpath now fails the build instead of the user's
first inference.

Cost: the darwin vllm-cpp image grows by about 124 MB.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]

* fix(vllm-cpp): default the MLX GEMM provider OFF on darwin

This branch opened with VLLM_CPP_MLX=on, justified by an A/B that measured the
MLX provider at 1.88x to 2.19x against the native MSL GEMM. That measurement was
correct when taken and is now stale: vllm.cpp's own Metal kernels have improved
several-fold since, through mma prefill attention, a vectorised decode V
accumulation, vectorised attention staging, a fused qk-norm-RoPE preamble and a
simdgroup-per-row softmax. The native path MLX was compared against no longer
exists.

Re-measured on the same Apple M4, in the same binary, with the arms toggled by
VT_OP_PROVIDER_DISABLE=mlx, on Qwen3-1.7B-bf16 warm at p=512 g=128:

  MLX provider ON   prefill TTFT 1370 ms   warm throughput 11.98 tok/s
  MLX provider OFF  prefill TTFT 1400 ms   warm throughput 22.06 tok/s

Shipping the previous default would have halved Apple Silicon throughput.

MLX's steel GEMM is still about 20% faster than ours in isolation, but the
provider pays a per-op mx::eval synchronisation plus an output memcpy, because it
cannot write into our buffer. Across prefill's roughly 112 GEMMs that overhead
leaves a 2% gain; on decode, where the same synchronisation is paid once per
matmul per token, it costs 46%. The option is kept for prefill-dominated
workloads, where the margin is small but real.

The README section is rewritten rather than patched: it previously presented the
stale table as the reason for the default, so leaving it in place would have made
the new default look arbitrary.

Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vllm-cpp): bump vllm.cpp and default MLX ON, gated to prefill

Bumps VLLM_CPP_VERSION from 9e1c9025 to eec09bed and turns VLLM_CPP_MLX back on.
These two must move together, which is why they are one commit.

Upstream now shape-gates the MLX provider to prefill: it declines m < 2, which is
exactly the decode GEMV. MLX's steel GEMM wins prefill, 524.5 ms of TTFT against
602 for the native path, but loses decode badly because the provider pays an
mx::eval synchronisation and an output memcpy on every call while decode makes
about 112 calls per token. Ungated it does both; gated it does only the good half.

Measured on an Apple M4 with Qwen3-1.7B-bf16 warm at p=512 g=128:

  MLX gated to prefill (pin >= 89c46aeb)   TTFT 524.5 ms   24.40 tok/s, 99.1% of MLX-LM
  MLX ungated (older pins)                 TTFT 537 ms     12.7 tok/s
  MLX off                                  TTFT 602 ms     23.9 tok/s

This branch briefly defaulted the provider off, which was the correct call for an
ungated provider at the old pin. The gate is what makes on correct again, so the
pin and the flag are coupled: rolling VLLM_CPP_VERSION back before 89c46aeb while
leaving MLX on would select the middle row and roughly halve throughput. Both the
Makefile comment and the README state that dependency explicitly.

The bump also brings six Metal kernels landed upstream since the old pin — mma
prefill attention, a vectorised decode V accumulation, vectorised attention
staging, a fused qk-norm-RoPE preamble, a simdgroup-per-row softmax and a
simdgroup-per-head preamble — which take the non-MLX Metal path from 89.4% to
96.4% of MLX-LM on their own.

One caveat, recorded in the README: MLX's GEMM is not bit-identical to the native
kernel, so an MLX build produces a different greedy sequence than a non-MLX build.
That is a property of the provider rather than of the gate and predates this
packaging.

Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(vllm-cpp): correct the MLX-gated figure to 97.6%, from 99.1%

The previous commit quoted 99.1% of MLX-LM for the prefill-gated MLX build. That
figure divided by a two-run MLX-LM baseline, 27.135 and 27.744 generation tok/s
averaged to 27.44. Re-measured interleaved with ours over four ABBA blocks,
MLX-LM's decode is 27.848 with a 0.34% spread across six runs, so the 27.135 was
an outlier and averaging it in overstated us by roughly 1.5 points.

Corrected: the gated configuration is 24.37 tok/s, or 97.6% of MLX-LM, and the
MLX-off build is 23.9 tok/s or 95.9%. Prefill TTFT is unchanged at 524.5 ms
against MLX-LM's 532.6, so we remain about 1.5% faster there.

Nothing else changes. MLX still wins prefill and loses decode, the shape gate is
still the right disposition, and the pin and the flag are still coupled. The gate
is worth about 1.7 points over the MLX-off build rather than 2.7.

Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): pin MLX gate from mainline

The previous pin was a merge commit from the experimental C ABI v9 branch. Pin the same MLX prefill gate on upstream main so the backend build does not pull unrelated ABI v9 work into every platform variant.

Assisted-by: Codex:gpt-5 [systematic-debugging]

* fix(vllm-cpp): restore backend build portability

Keep the current master pin when enabling MLX so every backend variant builds against the known-good vllm.cpp revision. Suppress Apple clang’s GNU constant-folding diagnostic for Objective-C++ Metal compilation only, since upstream treats warnings as errors.

Assisted-by: Codex:gpt-5 [systematic-debugging]

* fix(vllm-cpp): demote MLX header VLA warning

MLX 0.29.3 headers trigger Apple clang's gnu-folding-constant diagnostic in the Objective-C++ provider. Keep the diagnostic visible while exempting only it from vllm.cpp's global warnings-as-errors policy.

Assisted-by: Codex:gpt-5 [systematic-debugging]

* fix(vllm-cpp): suppress MLX header VLA warning

Target-level Objective-C++ -Werror is appended after the directory flags, so a no-error demotion is re-promoted. Disable this single warning for the MLX header while keeping every other warning fatal.

Assisted-by: Codex:gpt-5 [systematic-debugging]

* fix(vllm-cpp): pin source-scoped MLX warning fix

Move the AppleClang warning exception into vllm.cpp where its target warning policy is defined, and pin LocalAI to that source-scoped fix.

Assisted-by: Codex:gpt-5

* fix(vllm-cpp): pin effective MLX warning suppression

The source-scoped no-error flag was overridden by the target warning policy. Pin the companion vllm.cpp change that disables only the MLX header diagnostic for its Objective-C++ translation unit.

Assisted-by: Codex:gpt-5

* fix(vllm-cpp): pin diagnostic pragma fix

Pin the companion vllm.cpp correction that scopes the AppleClang folding warning suppression inside the MLX translation unit, after command-line warning policy.

Assisted-by: Codex:gpt-5 [systematic-debugging]

* fix(vllm-cpp): pin remaining Darwin build fixes

Advance the MLX-enabled backend to the vllm.cpp revision already validated by the dependency update branch. This includes the feature guards and AppleClang pragma boundary needed by the Darwin build.

Assisted-by: Codex:gpt-5 [systematic-debugging]

* fix(vllm-cpp): pin MLX system dependency boundary

Pin the companion vllm.cpp change that models MLX as an imported system dependency, keeping third-party header diagnostics out of the project's warnings-as-errors policy while retaining fatal warnings for project sources.

Assisted-by: Codex:gpt-5 [Codex]

* fix(vllm-cpp): pin scoped MLX warning guard

Advance vllm.cpp to the companion fix that keeps MLX headers on a SYSTEM dependency and scopes AppleClang folding-constant suppression to the external includes.

Assisted-by: Codex:gpt-5 [systematic-debugging] [test-driven-development]

* fix(vllm-cpp): use available MLX wheel

MLX 0.29.3 is no longer available to the Darwin runner, so the backend build stopped before CMake. Pin the first available compatible wheel and keep the documented default in sync.

Assisted-by: Codex:gpt-5

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-04 15:41:39 +02:00
mudler's LocalAI [bot]
137dfcf15a chore: ⬆️ Update antirez/ds4 to b7e9f0091139999b6c070a57590c447c5741da5c (#11333)
* ⬆️ Update antirez/ds4

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(ds4): link upstream CUDA MMQ objects

The updated ds4 CUDA object now calls into the vendored MMQ implementation. Build and link those objects into both the gRPC server and distributed worker.

Assisted-by: Codex:gpt-5 [Codex]

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-04 15:27:03 +02:00
localai-org-maint-bot
750ab91b2b test(advisorylock): replace fixed sleeps with signals (#11343)
Wait for observable loop events instead of budgeting hundreds of milliseconds for scheduler timing. Keep a short bounded overlap observation for the two-leader exclusion check.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-04 15:04:29 +02:00
dependabot[bot]
08598a8611 chore(deps): bump the npm_and_yarn group across 1 directory with 5 updates (#11341)
Bumps the npm_and_yarn group with 5 updates in the /core/http/react-ui directory:

| Package | From | To |
| --- | --- | --- |
| [hono](https://github.com/honojs/hono) | `4.12.25` | `4.12.34` |
| [@hono/node-server](https://github.com/honojs/node-server) | `1.19.14` | `2.1.0` |
| [fast-uri](https://github.com/fastify/fast-uri) | `3.1.4` | `3.1.5` |
| [ip-address](https://github.com/beaugunderson/ip-address) | `10.2.0` | `10.4.0` |
| [undici](https://github.com/nodejs/undici) | `7.28.0` | `7.29.0` |



Updates `hono` from 4.12.25 to 4.12.34
- [Release notes](https://github.com/honojs/hono/releases)
- [Commits](https://github.com/honojs/hono/compare/v4.12.25...v4.12.34)

Updates `@hono/node-server` from 1.19.14 to 2.1.0
- [Release notes](https://github.com/honojs/node-server/releases)
- [Commits](https://github.com/honojs/node-server/compare/v1.19.14...v2.1.0)

Updates `fast-uri` from 3.1.4 to 3.1.5
- [Release notes](https://github.com/fastify/fast-uri/releases)
- [Commits](https://github.com/fastify/fast-uri/compare/v3.1.4...v3.1.5)

Updates `ip-address` from 10.2.0 to 10.4.0
- [Release notes](https://github.com/beaugunderson/ip-address/releases)
- [Commits](https://github.com/beaugunderson/ip-address/compare/v10.2.0...v10.4.0)

Updates `undici` from 7.28.0 to 7.29.0
- [Release notes](https://github.com/nodejs/undici/releases)
- [Commits](https://github.com/nodejs/undici/compare/v7.28.0...v7.29.0)

---
updated-dependencies:
- dependency-name: hono
  dependency-version: 4.12.34
  dependency-type: direct:production
  dependency-group: npm_and_yarn
- dependency-name: "@hono/node-server"
  dependency-version: 2.1.0
  dependency-type: indirect
  dependency-group: npm_and_yarn
- dependency-name: fast-uri
  dependency-version: 3.1.5
  dependency-type: indirect
  dependency-group: npm_and_yarn
- dependency-name: ip-address
  dependency-version: 10.4.0
  dependency-type: indirect
  dependency-group: npm_and_yarn
- dependency-name: undici
  dependency-version: 7.29.0
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-04 12:13:49 +02:00
mudler's LocalAI [bot]
211aa0a536 chore: ⬆️ Update mudler/vllm.cpp to a42b8187caff02c570c28e19e4dc2b1d7f55ed14 (#11174)
⬆️ Update mudler/vllm.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-04 08:17:41 +02:00
mudler's LocalAI [bot]
c86b3b207b chore: ⬆️ Update ikawrakow/ik_llama.cpp to 60389410a1ff01f9d37dcc6261db33b3183bdea2 (#11331)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-04 08:17:14 +02:00
mudler's LocalAI [bot]
62316e52a9 chore: ⬆️ Update 0xShug0/audio.cpp to 4e3aea2fd99aeaa5924e71c51eb2793846045332 (#11332)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-04 08:17:02 +02:00
mudler's LocalAI [bot]
3090101156 chore: ⬆️ Update CrispStrobe/CrispASR to fe3caf8e363b27572dbdd1a9d37083f25e6decda (#11334)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-04 08:16:49 +02:00
mudler's LocalAI [bot]
8b667cd1ce chore: ⬆️ Update ggml-org/whisper.cpp to 64d57d3df5c8dacee098577257edcaa154bf5ef3 (#11326)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-04 08:16:36 +02:00
dependabot[bot]
93fe086798 chore(deps): bump the npm_and_yarn group across 1 directory with 2 updates (#11338)
Bumps the npm_and_yarn group with 2 updates in the /core/http/react-ui directory: [@hono/node-server](https://github.com/honojs/node-server) and [brace-expansion](https://github.com/juliangruber/brace-expansion).


Updates `@hono/node-server` from 1.19.14 to 2.0.12
- [Release notes](https://github.com/honojs/node-server/releases)
- [Commits](https://github.com/honojs/node-server/compare/v1.19.14...v2.0.12)

Updates `brace-expansion` from 1.1.12 to 1.1.18
- [Release notes](https://github.com/juliangruber/brace-expansion/releases)
- [Commits](https://github.com/juliangruber/brace-expansion/compare/v1.1.12...v1.1.18)

---
updated-dependencies:
- dependency-name: "@hono/node-server"
  dependency-version: 2.0.12
  dependency-type: indirect
  dependency-group: npm_and_yarn
- dependency-name: brace-expansion
  dependency-version: 1.1.18
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-04 08:16:22 +02:00
mudler's LocalAI [bot]
2e14511fe2 docs(blog): add release write-ups for 3.10 through 4.3 (#11330)
The blog has a deep post for 4.8 and a history post that covers the earlier
releases at summary altitude, but nothing in between. These five fill that
gap in the same shape as what-landed-in-localai-4-8: what the release was
for, runnable examples, and the limits that apply.

Every endpoint, CLI flag, env var and gallery entry is verified against the
matching release tag rather than taken from the release notes. That caught
two paths the published 3.10.0 notes got wrong: tracing is /api/traces, not
/api/v1/trace, and a stored response is fetched from /v1/responses/:id, not
/api/v1/responses/{response_id}.


Assisted-by: Claude Code:claude-opus-5[1m]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-04 00:13:45 +02:00
mudler's LocalAI [bot]
88fdda6211 chore(model-gallery): ⬆️ update checksum (#11327)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-03 23:17:12 +02:00
mudler's LocalAI [bot]
f447faf08d chore: ⬆️ Update ggml-org/llama.cpp to 221f0f6356efe2260023208365705ec5d5a7c8f5 (#11303)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-03 23:03:39 +02:00
mudler's LocalAI [bot]
6e7c0a4df8 blog, website: edit out the AI writing tells readers called out on HN (#11324)
* blog: rewrite the engines post without the AI tells

The HN thread on this post (item 49125065) spent most of its comments on the
writing rather than the engines. Readers quoted specific lines back as tells.
This is the same post with the same numbers, edited against the updated
no-ai-slop skill.

Every figure, table and link is unchanged, except that "27% of the memory"
is now the underlying 363 MB against 1328 MB from the table.

Two substantive framing fixes, both from the reply draft in
hn-reply-engines-post.md:

- vllm.cpp is no longer implied to be a speed win. The table is a tie, the
  result is the install size, and the post now says so before a reader has to
  work it out and post about it.
- Added one line on the language mix. Readers took the C++/Python/Go tree as
  incoherence rather than as a Go core with per-ecosystem backends.

Cut throughout: the ledger metaphor ("what those ports buy", "not paid for in
throughput"), unearned framing ("the honest reading is", "has nothing to do
with"), the shape summary ("that is the general shape of these wins"),
confident deference ("people who are better at those models than we are"),
self-grading numbers ("a good result for a 66 MiB binary"), verbless
comparisons, three of the four exactness idioms, and the aphoristic headings
and verdicts. The double-tricolon summary is one plain clause now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* blog, website: same anti-slop sweep over the rest of the site

One-by-one pass over the other four posts and the site templates, with the
same rules used on the engines post. All figures, tables, links and PR
numbers are unchanged everywhere; the edits are to prose only.

apex-moe-quantization: ledger metaphors were the main issue, eight uses of
buy/cost/pay/spend for things that are not money. Also "the honest reading
is", "that is the comparison that matters", and two section-ending aphorisms
("Size is a speed knob as much as a memory knob", "Q6_K is the ceiling worth
paying for").

localai-since-march-2023: light touch, this one already reads like a person.
Removed "the curve is not the point", a "not the feature list, but the four
decisions" contrast, and two "X is what made / is the piece that" forms.

parakeet-cpp-asr-on-cpu: six exactness idioms across one post, "byte for
byte" twice, "character for character" twice, "byte-identical" twice and
"bit-identical" once, including in the title. Down to one, kept where the
precision is load-bearing. Also the "what end-of-utterance detection buys
you" heading and the "we say so rather than averaging it away" flex.

what-landed-in-localai-4-8: no changes. It is dense, flat and ends every
section on a PR number or a plain fact, which is the shape the other posts
should look like.

Site templates: "Most backends wrap somebody else's engine. These do not."
was the same contrast the engines post opened with. Also "Not a degraded mode
that technically runs", "A port only ships once it matches the original",
"Speed is the part we then go and win ... not a marketing run", and the last
"byte for byte" on the landing page.

Hugo builds clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* website: it is eighteen engines, not nineteen

Three places said nineteen: the /engines/ page description, the JUL 2026
timeline entry on the landing page, and the header comment in
data/engines.yaml.

Eighteen is right, confirmed two ways. The "Backends built by us" table in
the README has exactly 18 rows, and data/engines.yaml has 19 entries of which
one is apex-quant, which is a quantization recipe rather than an engine. The
two lists otherwise match name for name.

The yaml comment is the likely origin: it read "the nineteen native engines
the LocalAI team wrote, and the one quantization recipe that feeds them",
which counts apex-quant twice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 23:03:25 +02:00
mudler's LocalAI [bot]
e2311045d3 fix(mcp): drop the duplicated scheduling methods on stubClient (#11323)
master does not compile:

    vet: core/http/endpoints/mcp/localai_assistant_test.go:157:19:
    method stubClient.ListScheduling already declared at
    core/http/endpoints/mcp/localai_assistant_test.go:87:19

Two fixes for the same breakage landed. The four Scheduling methods were
already present at lines 87-99, in interface order after ListNodes, by
the time #11318 merged; #11318 appended its own copy after
GetRouterDecisions. The two blocks sit in different parts of the file, so
git merged both without a conflict and nothing flagged it.

Remove the appended copy and keep the one in interface order. Pure
deletion, no behaviour change.

Verified: go vet clean on ./core/http/endpoints/mcp/, and
go test ./core/http/endpoints/mcp/ passes.


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-03 22:53:04 +02:00
Ettore Di Giacinto
6bdb04ab5d docs: point the News page at the blog instead of a stale highlights list
The News page kept a hand-maintained "Highlights" list that had drifted:
it was missing all of 2025, duplicated the README's own news list, and
linked /features/middleware/ for a page that lives at operations/.

Both of its jobs already have owners. website/content/blog/ carries the
release write-ups and engineering notes, and GitHub Releases carries the
full changelog. Replace the list with a pointer at those two, so there is
one place to update instead of three.

The page keeps its url and front matter, so /docs/basics/news/ and the
root /basics/news/ redirect that .github/ci/gen-redirects.sh generates
both keep resolving.

Also drop the two contributor instructions in .agents that told authors
to add a whats-new.md bullet per feature: announcing a capability is the
release blog post's job, per .agents/preparing-a-release.md.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Write] [Bash]
2026-08-03 20:29:50 +00:00
mudler's LocalAI [bot]
bd076376be fix(ci): install the Go the module asks for when building the site (#11322)
Deploy site to GitHub Pages failed on five of the last eight master
pushes, always in the build job before Hugo runs:

    Setup go version spec 1.22
    ...
    go: downloading go1.26.0 (linux/amd64)
    go: download go1.26.0: golang.org/toolchain@v0.0.1-go1.26.0.linux-amd64:
        Get "https://proxy.golang.org/...": connect: network is unreachable
    ##[error]Command failed: go env GOPATH

The workflow pinned setup-go to 1.22 while go.mod declares go 1.26.0, so
the `go run ./.github/ci/modelslist.go` step that generates the gallery
page had to fetch the real toolchain from proxy.golang.org first. That
fetch is not reliably reachable from the runner, which is why the deploy
alternated between passing and failing rather than failing outright.

Track go.mod instead of a literal. The version the module needs is then
installed directly and there is no toolchain download to fail.

This matters beyond CI noise: the docs and the site, including the
release blog post, ship through this workflow.

Scoped deliberately to gh-pages, the workflow with the observed failure.
test-extra.yml pins 1.25.4 in a dozen places and is below go.mod for the
same reason, so those jobs also download a toolchain, but they are
currently green and rewriting twelve pins on a hunch risks more than it
fixes. Worth a follow-up.


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-03 19:18:56 +02:00
localai-org-maint-bot
d28ccf32b5 gallery: add Qwen3.6 14B FableVibes variants (#11317)
Add Q4_K_M and Q8_0 llama.cpp entries with the shared Q8_0 multimodal projector.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-03 19:02:10 +02:00
mudler's LocalAI [bot]
95bd59d78e fix(mcp): teach the assistant test stub the scheduling methods (#11318)
#11228 added ListScheduling, GetScheduling, SetScheduling and
DeleteScheduling to localaitools.LocalAIClient but did not update
stubClient, the hand-written test double in the mcp endpoints package.
The package therefore fails to typecheck, which takes out both lint and
tests on master:

    cannot use stubClient{} as localaitools.LocalAIClient value in
    argument to h.Initialize: stubClient does not implement
    localaitools.LocalAIClient (missing method DeleteScheduling)

Red on 8f74f74b, fd4ec083 and 8a68f357; green on cd62e8ff, the commit
before.

Add the four methods with the same inert bodies the rest of the stub
uses. The real implementations are covered in the localaitools suites;
this double only exists so the holder can be constructed.


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-03 19:01:40 +02:00
mudler's LocalAI [bot]
1741df0bf1 fix(ui): scale the chrome audit's timeout to the number of routes it walks (#11319)
chrome-audit.spec.js walks 25 routes in a single test, and has been the
UI E2E suite's failure on 5 of the last 6 master runs. It always dies the
same way, at the 30s per-test default:

    Test timeout of 30000ms exceeded.
    Error: page.waitForTimeout: Test timeout of 30000ms exceeded.
      19 |     await page.goto(route)
    > 20 |     await page.waitForTimeout(400)

The spec is new in 5cb0c1a8; the commit before it was green, and every
run since has been red on this file.

The failure is cumulative rather than one bad route. Across those runs
the clock runs out at line 19, 20 or 21 depending on where the loop
happens to be, and the timeout lands on waitForTimeout rather than on
goto, which is what running out of budget looks like as opposed to a
navigation that hangs. 30s over 25 routes is ~1.2s each, including a
deliberate 400ms settle, so there is very little headroom to begin with.

Give the test a budget proportional to its work: six seconds a route.
That absorbs a slow runner and still fails promptly if a route genuinely
hangs.

Verified: the spec passes on the current UI in 12.2s solo, and the full
suite passes 418 at 8 workers locally. What I could NOT do is reproduce
the CI timeout on this machine, which has 20 cores against the runner's
2 to 4; under synthetic CPU load it still finished in 13.5s. So the fix
is argued from the CI signature and the arithmetic, not from a local
repro, and the proof is this spec going green on the hosted runner.

Note test.setTimeout() has to be called inside the test body. At module
scope Playwright rejects it with "test.setTimeout() can only be called
from a test".


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-03 19:01:06 +02:00
mudler's LocalAI [bot]
b6d2e94153 fix(sglang): bound cuda-tile below the 1.6 prereleases (#11320)
Every CUDA sglang image failed in the 2026-08-02 full-matrix rebuild:
-gpu-nvidia-cuda-12-sglang, -gpu-nvidia-cuda-13-sglang and
-nvidia-l4t-cuda-13-arm64-sglang, all with the same build error.

    Building cuda-tile==1.6.0rc3
    x Failed to build `cuda-tile==1.6.0rc3`
      ModuleNotFoundError: No module named 'wheel_stub'
    hint: `cuda-tile` (v1.6.0rc3) was included because `sglang` (v0.5.16)
          depends on `flashinfer-python` (v0.6.14) which depends on `cuda-tile`

This is the failure mode requirements-cublas1{2,3}-after.txt already
carries an nvidia-modelopt bound for, arriving through a different
package. install.sh passes a global --prerelease=allow, which is
load-bearing for flash-attn-4, so an unbounded dependency resolves to a
prerelease; cuda-tile 1.6.0rc3's build backend imports wheel_stub without
declaring it in build-system.requires; --no-build-isolation means nothing
provides it, and the build dies.

Nothing in this repo changed. cuda-tile published 1.6.0rc1 and rc3 and
the weekly cron picked them up, which is the drift that job exists to
catch.

Bound the one package rather than dropping the global flag, matching the
existing precedent. 1.5.0 is the newest stable release, so <1.6 takes the
last good one. l4t13 gets the same bound: it installs plain sglang rather
than sglang[all], but flashinfer-python is a dependency of both.

NOT VERIFIED LOCALLY: reproducing this needs a CUDA docker build, which
this machine cannot run. The diagnosis is from the CI log and the
resolver's own hint, and the change follows a fix already proven in these
same files. CI on this PR is the check that matters.


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-03 19:00:28 +02:00
mudler's LocalAI [bot]
a0f7faaa2a fix(sycl): stop building the ggml CPU variant matrix with icpx (#11321)
Since #11255 and #11276 every GPU image also builds ggml's CPU_ALL_VARIANTS
matrix, so a partial offload uses the host's SIMD kernels. That works
everywhere except SYCL, where the Makefile compiles the whole tree with
icpx -fsycl: icpx never finishes ggml-cpu/arch/x86/repack.cpp at
-march=sapphirerapids. In run 30765516644 both sycl_f16 and sycl_f32 stopped
at that translation unit and sat there for 5h30m with a single compile in
flight until GitHub killed the job at its 6h limit, and turboquant's f16 job
lost its runner outright. gcc compiles the same file in seconds in the vulkan
and CPU jobs of the same run, so the CPU variant matrix is only unbuildable
under icpx.

Route SYCL back to the portable fallback binary, which is what these images
shipped before #11255. run.sh already prefers *-cpu-all when present and falls
back otherwise, so nothing else has to change.


Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-03 19:00:00 +02:00
localai-org-maint-bot
133c546c3f feat(api): add text moderation endpoint (#11316)
* feat(api): add text moderation endpoint

Add an OpenAI-compatible /v1/moderations endpoint backed by constrained local text generation. Register its auth and discovery surfaces, document the text-only MVP, and cover response shaping and access control.

Assisted-by: Codex:gpt-5

* test(mcp): update assistant client stub

Keep the LocalAI Assistant holder test stub aligned with the scheduling methods added to LocalAIClient so repository-wide type checking succeeds.\n\nAssisted-by: Codex:gpt-5

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-03 18:03:46 +02:00
Pete
8a68f3571c feat(api): add POST /v1/images/upscale endpoint (#10227)
* feat(api): add POST /v1/images/upscale endpoint

Add a new image upscaling endpoint that accepts a source image and
returns an upscaled version. Supports selectable upscaler models
(e.g. realesrgan) and a configurable scale factor (2x or 4x).

- backend.proto: add UpscaleImage RPC and UpscaleImageRequest message
- pkg/grpc: implement UpscaleImage in Backend interface, client, server
  and embed shim
- core/backend/upscale.go: new backend helper (mirrors ImageGeneration)
- core/http/endpoints/openai/upscale.go: new multipart/form-data handler
- core/http/routes/openai.go: register POST /v1/images/upscale
- core/http/auth/features.go: gate upscale routes under FeatureImages
- backend/python/diffusers/backend.py: implement UpscaleImage — uses
  diffusers upscale pipeline when loaded, falls back to Lanczos resize

* fix(grpc): add UpscaleImage stub to Base backend

All Go backends embedding Base now satisfy the AIModel interface
without needing to implement UpscaleImage explicitly.

* fix(images): complete upscale endpoint integration

Store generated upscales under the served images directory, validate scale factors, document and advertise the endpoint, and add a functional Stable Diffusion x4 gallery model.

Assisted-by: Codex:gpt-5

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-03 15:27:22 +02:00
localai-org-maint-bot
fd4ec083b9 feat(downloads): add resume-safe pause action (#11222)
Give gallery operations distinct pause and cancel paths. Pause preserves partial download data so reinstalling the same model or backend resumes through HTTP Range, while cancel keeps its destructive semantics. Surface the action in the Activity UI and document the API behavior.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-03 15:25:23 +02:00
Owen Adirah
8f74f74b10 feat(mcp): expose scheduling admin tools (#11228)
* feat(mcp): add scheduling client contracts

Assisted-by: Hephaestus:openai/gpt-5.5
Signed-off-by: Owen Adirah <owenadira@gmail.com>

* feat(mcp): add scheduling HTTP client support

Assisted-by: Hephaestus:openai/gpt-5.5
Signed-off-by: Owen Adirah <owenadira@gmail.com>

* feat(mcp): add in-process scheduling stubs

Assisted-by: Hephaestus:openai/gpt-5.5
Signed-off-by: Owen Adirah <owenadira@gmail.com>

* feat(mcp): register scheduling tools

Assisted-by: Hephaestus:openai/gpt-5.5
Signed-off-by: Owen Adirah <owenadira@gmail.com>

* test(mcp): map scheduling tools to REST routes

Assisted-by: Hephaestus:openai/gpt-5.5
Signed-off-by: Owen Adirah <owenadira@gmail.com>

* docs(mcp): document scheduling assistant tools

Assisted-by: Hephaestus:openai/gpt-5.5
Signed-off-by: Owen Adirah <owenadira@gmail.com>

* fix(mcp): wire in-process scheduling

Use an explicit MCP scheduling DTO and route in-process scheduling calls through the distributed node registry so the embedded assistant matches the REST scheduling surface.

Assisted-by: Hephaestus:openai/gpt-5.5
Signed-off-by: Owen Adirah <owenadira@gmail.com>

* fix(mcp): narrow scheduling dto

Assisted-by: Hephaestus:openai/gpt-5.5 [opencode]
Signed-off-by: Owen Adirah <owenadira@gmail.com>

---------

Signed-off-by: Owen Adirah <owenadira@gmail.com>
2026-08-03 15:24:29 +02:00
localai-org-maint-bot
cd62e8ff18 gallery: add Nemotron 3 embedding models (#11314)
Add multilingual 1B and 8B Q4_K_M GGUF embedding entries and link them as variants for automatic memory-aware selection.

Assisted-by: Codex:gpt-5 [web]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-03 15:23:08 +02:00
localai-org-maint-bot
af98e76f84 fix(gallery): remove broken DeepSeek V4 0731 entry (#11313)
fix(gallery): repair DeepSeek V4 0731 entry

Use the official single-file ggml-org MXFP4 artifact with its verified SHA256 and route it through llama.cpp instead of treating an unsloth repository page as a ds4 model file.

Assisted-by: Codex:gpt-5 [Hugging Face API]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-03 13:25:19 +02:00
localai-org-maint-bot
7f9ffd9f54 gallery: add AMD Instella MoE 16B variants (#11308)
Add Q4_K_M and Q8_0 GGUF builds for the trending Instella-MoE-16B-A3B-Think model, with host-selectable variant metadata and verified Hugging Face LFS hashes.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-03 12:16:29 +02:00
mudler's LocalAI [bot]
5cb0c1a872 feat(ui): close the gap between the shipped UI and the design mocks (#11307)
* feat(ui): give the Operate overview real numbers and traces a latency shape

First two items from a component-by-component comparison against the mocks.
The pattern that audit found: everything newly built matched, everything
pre-existing got the palette but not the layout, and an "absent rather than
empty" rule hid most of the overview exactly when someone was looking at an
idle installation.

**The headline grid is always rendered**, including at zero, with a fourth cell
for host memory. Hiding it removed the page's structure precisely when it was
most likely to be read, and "0 failed" is information — an absent panel is not.
The quiet case is now said in a line underneath instead of by showing nothing.

**The sections state counts** rather than listing their destinations: backends,
models, updates and running operations instead of the words "Usage and traces".
That needed installed backend and model counts in the summary context, which
are two more cheap reads on the poll that was already running.

**Traces rows carry latency as a bar as well as a figure**, scaled against the
slowest request currently in view and turning amber past two seconds. The table
had no latency column at all — the number was buried in the expanded detail, so
the shape of the tail was invisible while scanning. Scaling against the view
rather than an absolute ceiling is deliberate: what matters when reading a page
of traces is which of these are the outliers, and an absolute scale flattens
every row on a fast installation into nothing.

Full e2e suite: 409 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): name the engine on Home's resident models, and add jump-back-in

Third item from the mock comparison.

The mock showed each resident model with the engine serving it. /system carried
only the id, so the audit recorded this as blocked on a server field — but the
config loader is already in scope where that response is built, so it is one
lookup. SysInfoModel gains an optional `backend`, resolved from the model's
config and omitted rather than guessed when there is none (a loose file, or a
config since removed). Home renders the column blank in that case; the test
pins both halves of that.

Memory per model stays out. It is not one lookup — it would mean asking each
backend process — and inventing a number beside a real one is worse than
leaving the column off.

"Jump back in" is the block the mock had and Home did not. The quick-links row
above it is a set of first-run actions; these are the three places someone
returns to, each stated with what it currently holds rather than as a bare
label.

Go: routes suite passes. Full e2e suite: 412 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): rank the recommended models as lanes instead of equal cards

The hardware recommendations were a grid of equally-weighted cards. The list is
already sorted by fit, and a grid throws that order away: three cards side by
side say "pick one", when the page has actually formed an opinion about which
one.

They are lanes now, read top to bottom in fit order, with the leader carrying
the single amber "Best fit" label and the rest marked "Also fits". One opinion
per page — the alternatives are alternatives, not runners-up each worth their
own colour, which is how a strip of coloured badges ends up meaning nothing.

Below 720px the size and VRAM columns drop and the lane keeps the name and the
install action, which are the two things a narrow screen needs.

The existing panel spec moves off .rec-models-item onto .lane rather than being
deleted; dismissal, collapse, keyboard operation and install all still pass
unchanged, and there is a new assertion that exactly one row is called out.

Full e2e suite: 413 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* fix(ui): drop capsule chips app-wide, and un-break the empty voice library

**Pills are gone.** A capsule radius reads as a tag floating on the surface,
which fights a system whose structure is hairlines and square corners — and
with chips on Discover, Host, Activity and the biometrics pages, "some pages
have pills" was the real inconsistency rather than any one page.

Sixteen selectors move to the small radius: filter buttons, tab pills, activity
and biometrics chips, file and count badges, the jump-to-latest control, the
nav badge. Round *buttons* keep their circle — .lightbox__nav and
.home-send-btn are circles, not capsules — as do every progress track, status
dot and avatar, which are round because they are round, not because they are
tags.

**The empty voice library was unusable.** `.voice-library-empty` sets
min-height: 430px, border: 0 and background: transparent — a description of the
empty PANEL — and it had been attached to the action instead. The create button
was therefore a 430px transparent box that pushed itself out of the panel and
could not be seen. Moved onto the container it describes, which now centres its
action rather than letting it fall off the bottom. Same class-mangling shape as
the Agents header fixed earlier.

Full e2e suite: 416 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): put Host's headline figures on the shared hairline strip

Host had shadowed, clickable StatCards above a page that already has a rail, a
pane and a tab bar — a second dashboard language on one screen, and a different
one again from the figures inside its own detail pane.

The Operate overview's figure grid is generalised into a shared `.stat-strip`
and Host adopts it, so the two pages read as one system: same cell, same figure
scale, same tone vocabulary, and the same hairline grid the split-view StatGrid
already uses. The cells stay clickable and still route into the tab and filter
they describe, because a count is worth more when it is also the way to the
thing counted.

Tone is spent only where the number means something — running and updates when
non-zero — since a strip where every cell is coloured has no emphasis left.

Two bugs made on the way, both now covered:

- The first version put `<button>` elements inside a `<dl>` with `<dt>`/`<dd>`
  inside the buttons. Neither is valid, the browser re-parents both, and the
  cells collapsed. These cells are a set of controls, so a plain container of
  buttons is also the honest markup.
- Even correct, the strip rendered 2px tall: `.page--app` is a flex column
  whose split view takes flex:1, so a child with no intrinsic minimum is shrunk
  away. The old cards survived only because `.stat-card` carried
  min-height:96px. The strip now declines to shrink, with a test pinning it.

The stat-card specs are retargeted rather than deleted: they were written to
guard a class collision on a page that no longer uses cards, so they now guard
the strip's labels and its height.

Full e2e suite: 417 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): make Backends notices an edge rather than a filled card

The install and upgrade banners were tinted cards with a full border. A filled
panel makes every notice shout at the weight of an error, which is how notices
stop being read — and Backends shows one on most visits, so it was shouting
routinely.

They are now a hairline with a coloured left edge, the same treatment the
Operate overview gives rows that want a decision, so "this needs you" looks the
same wherever it appears. Counts in the notice take the monospace tabular
figures the rest of the console uses.

Also drops the last inline style on the page, and refreshes the inline-style
baseline, which has read 624 against a real count since #11288 landed. The gate
exits 0 either way, so nothing was failing — but a baseline 86 above the truth
would have let that many inline styles back in unnoticed. Now at 538, which
tightens the ratchet rather than loosening it.

The spec creates the upgrade it asserts on rather than skipping when the mock
has no notice: a test that skips is a test that proves nothing.

Full e2e suite: 418 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): finish the mock parity list, and stop hiding the recommendations

The last two items from the audit, plus a correction.

**Discover's use-case shelf is lanes.** These are a list of ways in, read in
order; a grid of equal cards asks the reader to compare them, which is not the
choice on offer.

**The request panel reaches every generator.** Video, 3D, Sound and Audio FX
join Images and Speech, so each one teaches its own endpoint rather than two of
six doing it. Audio FX records the fields that shape the request rather than
the bytes, since its payload is multipart.

**Recommendations no longer collapse themselves.** They were folded away by
default once anything was installed. That is the page's one opinion about this
host, and an opinion hidden by default is one the reader never gets. Someone
who disagrees can still collapse it and that choice is remembered — the
difference is that we no longer make it for them. Three specs asserted the old
default and now assert the new one.

The use-case heading also sat a line's width from the text it introduces, so
the two read as one paragraph. It has air under it now, and the shelf is
separated from the recommendations above it.

Two tests removed rather than kept: a generator loop whose only real assertion
was `expect(endpoint.length).toBeGreaterThan(0)`, and an earlier card-gap guard
that could only skip. A test that cannot fail is worse than no test, because it
reads as coverage.

Full e2e suite: 418 passed, 4 skipped. Inline styles at baseline.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* fix(ui): make the Host figures legible and give the strip its spacing back

Three defects introduced by the Host redesign, all found by looking at the
running app rather than by the suite.

**The figures were invisible.** "Running now" and "Updates available" rendered
pure black on the dark ground. Two causes compounding: the `--muted` tone alias
never landed, because the source rule has extra spaces before its brace and the
exact-match edit missed it silently; and a `<button>` does not inherit colour,
so with no tone rule the value fell back to the user agent's `buttontext`.
Both fixed, and a test now fails on any figure computing to pure black.

**The strip sat flush against the resources panel.** `.stat-strip` declares
`margin: 0 0 ...` and is declared later in the file than `.manage-summary`, so
the shorthand quietly won and the top margin became zero. Raised to
`.stat-strip.manage-summary` so it beats the shorthand on specificity rather
than on declaration order, which is the kind of thing that breaks again the
next time a rule moves.

**Discover's use-case heading had a doubled gap.** `.zero-pane` is a flex
column that already separates its children; adding a margin on top of the gap
stacked the two. The margin is gone and the heading keeps only its own breathing
room.

Full e2e suite: 420 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): make Studio's tabs path segments rather than a query parameter

`/app/studio?tab=images` reads like a filter applied to a page. It is
navigation: a different generator, with its own state and its own deep link. It
is now `/app/studio/images`, with the overview at `/app/studio`.

Legacy `?tab=` links are redirected once to the path form, replacing the
history entry so Back does not bounce between two spellings of the same place.
Bookmarks and older links keep working and land on the canonical URL rather
than a second version of it, which is the part worth having a test for.

The nine `?tab=` references were all in specs, none in docs, so the migration
is contained. They move to paths, and a new spec pins the redirect.

Full e2e suite: 421 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): make the hardware recommendation a section, not a dismissable card

It was a bordered card with a collapse control and a close button, sitting
inside a pane that is otherwise hairline sections. Two problems: it read as
something bolted onto the page rather than part of it, and treating it as an
interruption to be shut is the wrong frame for the one thing the page has to
say about the machine it is running on.

It is now a plain section with the same heading treatment as the shelves below
it. The collapse state, the dismissal, their storage keys and the legacy key
read for backwards compatibility all go with it, along with the installedCount
prop that existed only to pick a default collapse.

Five specs described behaviour that no longer exists and are removed rather
than adjusted — collapsing, dismissing, persistence of both, and the toggle's
keyboard handling. One new spec asserts the replacement contract: no control
with aria-expanded, no dismiss, and no card border.

Full e2e suite: 416 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* fix(ui): restore every stripped icon and every default-chrome control

You reported two broken icons. They were not two: an earlier automated edit had
stripped the `fa-*` class from twenty `<i>` elements across eleven pages, and an
`<i>` with no icon class renders nothing at all. Settings' save button, the
voice-profile back link, and eighteen others — agent row actions, task and job
buttons, import and create actions — were all drawing empty space.

Each is restored from its own context rather than a blanket icon: the agent row
gets pause/play, pen, comments, file-export and trash; the fine-tune toggle
swaps plus for xmark as it opens; the P2P documentation link gets the
external-link glyph.

The same edit left controls without their classes. Fine-tune's "Import config"
was rendering in the browser's own chrome, and `.p2p-cmd__copy` set a border
but no background, so it fell back to `buttonface` — a pale grey chip on a dark
command block. FineTune's "New job" also had its icon classes folded into the
button's className, the same mangling already fixed on the Agents header.

Rather than fix the reported two and wait for the next report, this adds a
standing audit: twenty-five routes are walked and the test fails on any visible
control rendering with user-agent chrome, or any `<i>` without an `fa-*` class.
It found the three remaining cases after the first sweep, and it is the reason
the next one cannot ship quietly.

Full e2e suite: 417 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-03 12:16:09 +02:00
mudler's LocalAI [bot]
cd890b6a26 chore: ⬆️ Update leejet/stable-diffusion.cpp to db99efdd6d2a43c7937fd55b3359206c680a75b0 (#11299)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-03 08:40:23 +02:00
mudler's LocalAI [bot]
1c0380ad44 chore: ⬆️ Update 0xShug0/audio.cpp to 5a8312ef7b8aa7cf14e9a24ac568cabd8725d68a (#11302)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-03 08:36:53 +02:00
mudler's LocalAI [bot]
cb6e4d4391 chore: ⬆️ Update CrispStrobe/CrispASR to fcb79282a6bc52e13d858026c42b24fb6e63c97a (#11304)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-03 08:35:38 +02:00
localai-org-maint-bot
cba54c5ea1 gallery: add grug-27b GGUF variants (#11311)
Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-03 08:35:16 +02:00
mudler's LocalAI [bot]
f951419207 chore(model-gallery): propose variant groupings for review (#11312)
chore(model-gallery): propose variant groupings

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-03 08:34:21 +02:00
mudler's LocalAI [bot]
58ea2f5d79 feat(ui): give Operate and Studio a front door, and fix two layout regressions (#11305)
* feat(ui): give Operate a front door and fold six rail groups into four

Opening Operate ran firstVisiblePath() and landed on Backends, because
Backends happens to be written first in operateConsole.groups. The section
that should answer "is anything wrong" opened on a package manager, and
nothing was reported until you visited it.

Adds /app/operate. Its one irreplaceable block is "Needs attention", which
is empty when nothing is wrong and says so in a line rather than rendering a
reassuring green panel. It collects stale backends, failed operations and
unhealthy nodes. Everything else on the page is a summary you could already
assemble by visiting four others.

The rail regroups from six headings to four: Inference and Activity were both
"the runtime right now", Access and System were both administration. No
destination is removed and no gate changes, so isConsoleItemVisible and
consolePaths are untouched. Overview leads the first group, which is what
makes firstVisiblePath() return it without knowing it exists.

Rail entries now carry a signal beside the label. This does not replace the
sidebar badge and is not built as if it does: the badge stays on the
always-visible sidebar entry for the reason recorded in Sidebar.jsx, that the
rail exists only on Operate routes and can be collapsed. The signals are
orientation while inside Operate, so they are aria-hidden and nothing urgent
depends on them alone.

OperateSummaryContext polls once for the whole console, following
OperationsContext, which exists because per-consumer setInterval against one
endpoint was the defect it fixed. It is mounted by ConsoleLayout for the
Operate console only, so "poll only while in Operate" needs no route check.
Built on usePolling, so it pauses on a hidden tab. Operations are read from
OperationsContext rather than polled a second time, and each source degrades
to no-signal on its own so one dead endpoint cannot blank the rest. It reads
the cached GET /api/backends/upgrades and never the POST that forces a real
registry check.

Traces and Usage get no signal yet: /api/traces returns the list, so a count
would mean fetching every trace to render one number. A counts endpoint is
the honest fix and is scoped separately.

Full e2e suite green (369 passed, 4 skipped), including a render-smoke entry
for the new route.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): open Studio on what this machine can actually make

Studio was a tab strip over six generators that opened on Images, which was
never a decision, only the first entry in BASE_TABS. Nothing said which
modalities this installation could run, so the way to learn that video had no
model was to pick the tab and find an empty select.

Adds an overview tab and makes it the fallback. Explicit tabs still win, so
existing deep links keep working; anything unrecognised or gated now lands on
the overview rather than Images.

Each tab carries a dot: filled when an installed model advertises that
modality, hollow when nothing serves it. That is the feature in one detail,
turning the strip from navigation into a report of what the machine can do
before anything is clicked. The dot is aria-hidden because the overview states
the same facts in words and the dots change as models load.

Two kinds of unavailable, which had to stop looking alike:
  - switched off, via a permission: no tab and no lane, unchanged
  - available with no model: a lane, and a route to installing one

Studio now owns one MODALITIES table so the tab strip and the overview cannot
disagree about what exists, and calls useModels() once, unfiltered, grouping in
the browser. useModels(capability) fetches the whole list and filters locally,
so a hook per modality would have been six identical requests to
/api/models/capabilities on every mount. There is a test for that.

Recent outputs read every localStorage store through a new
readAllMediaHistory(), which avoids mounting five hooks that carry save timers
the overview has no use for. 3D is read separately through use3DHistory rather
than folded in: its entries are GLB blobs in IndexedDB, so they cannot come
from the same synchronous read.

Typical cost is the median of this machine's own history, not a guess, and
renders as a dash when there is nothing to go on.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* fix(ui): stop the stat cards and the console rail breaking on small screens

Two unrelated causes behind one report that /app/manage looks wrong when the
window is narrow.

The stat cards were being laid out by the wrong rule at every width. Two
different components both claimed `.stat-grid`: the dashboard card strip that
holds .stat-card children, and the detail-pane StatGrid the split views
introduced further down App.css. Being later, the second won every shared
property, so the cards got its 120px columns and its 1px hairline gap in place
of their own 180px columns and spacing-md. Four cards were packed onto a row
that fits two, labels wrapped to three lines and clipped, and the icon crowded
the value. Renamed the strip to `.stat-cards`, after the children it actually
holds, which also removes the mismatch of a `.stat-grid` container full of
`.stat-card`s. The split-view component keeps `.stat-grid` and its BEM parts.

The expanded console rail had no bounded height. Thirteen destinations stacked
in one column is taller than a phone, so opening the menu pushed the page's own
heading past the fold: the menu replaced the page rather than annotating it.
Capped at 55vh with internal scrolling below 768px, so the content behind stays
reachable.

Both are asserted on behaviour rather than markup: no stat-card label may be
clipped, the card gap must not be the detail pane's hairline, and expanding the
rail must leave the page heading on screen.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): retemper the palette to localai.io and add the lane primitive

The token half of the style transfer, plus the shared list idiom the two
overviews had each grown their own copy of.

theme.css moves from Nord to the website's palette, variable names preserved so
every consumer moves with it: ground #13171f -> #0d1117, accent frost cyan
#88c0d0 -> action blue #4f8cff, success sage -> mint #56d6a4, warning -> the
amber #f1b95d the site spends only on the thing asking for a decision. Eyebrows
go mint. Dividers become an opaque #29384a hairline rather than alpha over a
varying surface, which is what makes stacked surfaces read crisply on the site.

Light is derived, not inverted. The site ships one theme and never had to
answer this, but the app does: blue darkens to #2f62d8, mint to #0d8b60 and
amber to #8a5d0b, all clearing 4.5:1 on a cool paper ground, where the
dark-mode values sit near 2:1. Same three roles, different values.

Three files restate the palette because CSS variables cannot reach them:
cmTheme.js (the whole CodeMirror theme), VoiceVisualizer and WaveformPlayer
(canvas). Left alone they would have quietly kept the app half-Nord.

The `.lane` primitive replaces the near-identical row CSS that OperateOverview
and StudioOverview had each written: a full-bleed row on a hairline that insets
on hover, with no card and no shadow. Callers supply only the column template.
Both pages now use it, along with `.lane-head` for section rhythm and a
`.page-pad` container for top-level pages outside a console shell — without
which Studio sat flush against the sidebar with its eyebrow clipped.

Studio's tab strip wraps rather than running off the edge at narrow widths.

Full e2e suite: 386 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): put Home's resident models on lanes and give the footer one line

Home's status line was three chips saying a thing was true. It now reports
figures: how many models are resident, how many nodes are healthy, what share
of memory is in use, set in tabular monospace so the digits line up. A chip
answers whether; a figure answers how much, which is what someone opening the
page at a glance is after.

Resident models move from status chips to lanes, with the id set in a new
`.lane__name--id` because an id is something you might type or paste and the UI
face makes it read as a label. /api/system-information carries only the id, so
there is deliberately no backend or memory column: inventing one would mean a
server change this does not make.

The footer was three centred rows and cost the bottom sixth of every page for
chrome. It is one line now, version left and links right, wrapping to centred
when the viewport is too narrow to hold both. Every link it had, it keeps.

Full e2e suite: 392 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* fix(ui): correct three contrast failures and stop a guaranteed-404 poll

A contrast audit of the new palette found three values below WCAG AA, one of
which the previous commit message claimed was fine:

- White on the #4f8cff button is 3.22:1, which is large-text only. The website
  does exactly this, but a button label in an app is not large text, so the
  label goes to dark ink at 5.88:1. Light mode keeps white, which is 5.44:1 on
  its darker blue.
- Light-mode success was 4.08:1 on paper, not the 4.5 claimed. Darkened to
  #0a734f, 5.56:1.
- Nord red was already 4.28:1 on raised surfaces, a pre-existing miss carried
  over unexamined. Lifted to #c96f78, 5.02:1.

Lanes gain the two states they were missing: a 44px target on coarse pointers,
matching what EntityRail already does so the two list idioms feel the same
under a thumb, and a reduced-motion variant that keeps the background feedback
while dropping the hover inset, which is a position change.

The Operate summary no longer asks for /api/nodes on a single-node install. The
cluster API answers 503 when distributed mode is off, so it was a guaranteed
miss every fifteen seconds; it is now gated on useDistributedMode, the same
condition the rail already uses for the Nodes entry.

Full e2e suite: 392 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* fix(ui): restore the gap between overview blocks, and stop claiming zero nodes

Two defects a design review surfaced.

`.lane-head:first-child { margin-top: 0 }` was meant to stop the first block on
a page carrying a top margin. But every <section> makes its lane-head a first
child, so the reset applied to all of them and the gap between blocks vanished:
"Sections" sat flush against the attention row above it. The header supplies its
own bottom margin, so a uniform top margin is correct everywhere.

The Cluster summary read "0 nodes" on a single-node install, which looks like a
fault when the cluster API is simply switched off. It now says "Single node".

Full e2e suite: 392 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): open dark by default, and stop clipping the collapsed sidebar footer

Dark is the identity rather than a preference: localai.io ships one theme and
it is this one, so an install should look like LocalAI before anyone has chosen
anything. The OS setting no longer selects light on first load. The toggle
still does, and a stored choice wins forever after, which the tests assert
both ways.

The collapsed sidebar footer stacked its controls but kept the expanded row's
inline padding, so their edges were clipped against the 51px rail.

Full e2e suite: 394 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(api): count traces server-side and give the Operate overview real totals

The overview's headline block had no source. /api/traces returns the trace
list, so "37 errors in 24h" meant fetching every buffered exchange to count it
in the browser — waste that grows with the buffer, to produce three integers.

Adds GET /api/traces/summary: totals, failures, p95 and a bucketed series for
sparklines, over a window that defaults to 24 hours and is capped at a week.

Deliberate calls, each with a spec:
- A 4xx is the caller getting it wrong, not the installation being unhealthy,
  so only 5xx and transport errors count as failures.
- p95 is a nearest-rank percentile rather than the slowest request, which is
  what a max would report and what makes latency panels lie.
- Buckets are oldest-first so a sparkline reads left to right, and the slice is
  never nil: nil serialises as null and breaks .map() on the other side, which
  is a silent runtime error rather than an empty chart.
- Exchanges outside the window are not counted at all.

The route is registered before /api/traces/:id so "summary" is not captured as
a trace ID.

On the client, Traces and Usage gain the rail signals they were shipped
without, the Observability section summary now states counts instead of listing
its destinations, and an installation that has served nothing says so rather
than showing three zeroes dressed as telemetry.

Sparkline is a bare stroke with an emphasised endpoint and no axes: the figure
above it already states the value, so its only job is the shape.

Go: 185 middleware specs pass. Full e2e suite: 396 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* fix(ui): stop the memory chart calling a trade-off an error

The VRAM-by-context chart rendered any build over the limit in error red, and
escalated the verdict to the error tone as soon as two context sizes crossed
it. But an over-limit build still installs — #11288 keeps a test on exactly
that — so red overstates what is happening. A model that fits at 32k and not
64k is a trade-off, not a fault.

Over-limit bars and the limit line now use the warning tone, which is the
constraint colour used everywhere else in this branch: know what you are doing,
not you may not. The error tone is reserved for "fits nowhere", where the model
genuinely cannot run on this host.

Full e2e suite: 397 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): give the new surfaces orchestrated motion

Uses the reveal system already in the codebase rather than adding a library:
pageReveal, .reveal-stagger and staggerStyle() were built for exactly this, and
anime.js would be ~17KB duplicating four lines of CSS for list reveals.

The overview's headline figures, attention rows and section lanes stagger in,
as do Studio's modality lanes and recent outputs, so a page assembles in the
order it is read instead of appearing all at once.

Two additions beyond stagger. Rail signals transition on opacity when a poll
lands, so a number changing reads as an update rather than a jump cut, and it
stays on the compositor so it cannot reflow the rail. The attention block
animates its left edge in — the one thing on the page that should announce
itself, and on the border rather than the text so nothing moves under a reader.

Both are dropped entirely under prefers-reduced-motion, alongside the lane
hover inset already handled.

Full e2e suite: 397 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): put the generators on a hairline field stack and record the request

The workbench treatment from the mocks, applied where it costs least: both
changes land on shared surfaces, so all six generators get them at once rather
than drifting apart page by page.

The control column stops being a shadowed card of boxed groups and becomes a
hairline field stack — the panel is the page's left half, not an object
floating on it — with uppercase micro-labels matching the eyebrow treatment
used elsewhere. Because .media-controls is shared, Images, Video, 3D, Speech,
Sound and Audio FX all move together.

RequestPanel shows the request the form actually built, with a copy-as-curl.
LocalAI is API-first and Studio is the best place in the app to teach its own
endpoints: the form stops being a black box, and a result worth keeping can be
reproduced from a shell without reverse-engineering which fields the page sent.
It records what was sent rather than what the form currently holds, and renders
nothing until a request has been made — a panel describing a request nobody
made is a tutorial, not a record. Wired into Images and Speech.

Full e2e suite: 401 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): make Chat a transcript instead of a bubble thread

Rounded, filled, asymmetric bubbles fight a system built on hairlines, and they
carry the speaker in shape and side rather than in words. The assistant side
had already given up its bubble; this finishes the job.

Both roles now run full width down one column, separated by a rule, each with a
mono role label. The user turn keeps a left edge in the action tone so the two
are still told apart at a glance, without a fill or a corner radius. The
avatars go: the accent and the label carry the speaker, so the glyph was
decoration once neither side had a bubble.

Saying who is speaking in words rather than in geometry is also what survives
being read aloud, printed, or looked at by someone who cannot pick the sides
apart by colour.

Full e2e suite: 404 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* feat(ui): dress the API reference in LocalAI's palette

The Swagger page was the last surface still shipping in someone else's colours,
which is conspicuous now that everything it links from is dark.

Swagger UI has no theming hook, so rather than fork it we serve our own index
ahead of the library's wildcard and restate the palette over its stylesheet.
The library's own bundle and assets are still what load, so a swagger-ui
upgrade cannot silently break the page — this is a skin, not a fork.

Two things needed real care. Swagger tints the entire operation row per method
via .opblock.opblock-post and friends, so the palette had to match that
specificity rather than reach for !important; the method now lives on one edge
instead of washing across the row, because a page where every row is a status
colour has no status colour left. And the filled method chip put white on pale
green, which was the least readable thing on the page — it is an outlined mono
chip now, carrying the method in its border and text.

Palette values are copied from theme.css rather than referenced: this page is
served by Go and never sees the app's CSS. The comment says so, and says to
keep them in step.

Go: routes and middleware suites pass. Full e2e suite: 405 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

* fix(ui): make tall split-view pages reachable, repair the Agents header, scale titles

Three things found by actually using the app rather than measuring it.

**Host was unusable.** The shell above a split view is overflow:hidden so the
document cannot grow, which left anything taller than the viewport simply
unreachable — and Host stacks a resources card, four stat cards and a tab bar
above its split, so the bottom of the pane fell off at every window height with
nothing to scroll. Every sweep I ran for this was horizontal, which is why it
kept coming back clean.

The page now scrolls inside the pinned shell. The pane keeps its own scroller:
letting it grow instead pushes the document taller and stretches the rail to
match, which is the regression e2e/discover-height.spec.js exists to catch, and
which the first version of this fix duly caused.

**The Agents header controls were unstyled** — "Create Agent" was rendering
with the browser's default chrome. The markup had been mangled at some point:
six unrelated classes merged into one string on the link, and the label and
button left with none at all and empty icons. Repaired, with the inline flex
replaced by a shared .header-actions class.

**Page titles take the editorial scale from the site**: larger, tracked at
-0.04em, on a line height near 1, so a two-word title reads as a statement
rather than a label. The typeface is unchanged — DESIGN.md keeps the existing
type system — so the whole difference is scale, tracking and leading, which is
where the site gets its voice from. This was the biggest reason the running app
still did not look like the mocks.

Full e2e suite: 404 passed, 4 skipped.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-03 00:16:08 +02:00
mudler's LocalAI [bot]
b89b0f73e5 chore(model-gallery): ⬆️ update checksum (#11306)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-02 22:59:45 +02:00
mudler's LocalAI [bot]
45cd47cb99 chore: ⬆️ Update ikawrakow/ik_llama.cpp to cb9147fd0d9c08a9a84eee5ac405a73f4e10e3e1 (#11300)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-02 22:59:17 +02:00
mudler's LocalAI [bot]
1aa97381f3 perf(gallery): warm variant descriptions alongside VRAM estimates (#11297)
Follow-up to #11288, which warmed the VRAM estimate caches at startup and left
the variant picker paying its own way.

Describing an entry's variants probes the weight files of every build it
offers, so the first time a model is opened costs 1.2-1.9s against a cold
cache. That is the same cost as an estimate wearing a different hat, and it
lands in the same caches underneath, so it belongs in the same pass rather than
in a second mechanism.

The warm-up now describes variants for the entries it walks. Entries that
declare none cost nothing: the call is gated on HasVariants rather than
attempted and discarded. The host resolve env is derived once for the run,
since it describes the machine rather than the entry.

Failure handling matches the estimate half. An entry whose variants cannot be
described is logged at debug and skipped, and the estimate for that same entry
is unaffected, because neither half is allowed to fail the other.

Measured against a live instance with 1,595 models, first ever call to
/api/models/variants/:id after a cold boot:

  before   1.2-1.9s
  after    2ms

The warm-up's own cost barely moves: 3m0s to 3m19s for 300 entries, of which
40 declared variants. It stays bounded by the same knobs, and
LOCALAI_VRAM_WARM_LIMIT=0 still turns the whole thing off.


Assisted-by: Claude:claude-opus-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-02 19:42:54 +02:00
mudler's LocalAI [bot]
74b7ea2829 feat(ui): replace the gallery and inventory tables with a rail and a detail pane (#11288)
* feat(ui): rename the Install Models nav entry to Discover

"Install Models" named the action rather than the destination, and it was
the only multi-word entry in a rail of one-word ones (Home, Chat, Studio,
Talk, Build, Operate). A bare "Models" was the obvious fix but it collides
with the installed-models view under Host, which is a different page for a
different job.

"Discover" keeps the rhythm and says what the page is for. The icon moves
from a download arrow to a compass for the same reason: the page is browsed
before it is installed from.

Translated in all seven locales rather than left to fall back, so a locale
switch does not leave the entry in English.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* feat(ui): replace the gallery table with a rail and a detail pane

The eight-column table was not the real problem; the click-to-expand row
underneath it was. Variants, files and a VRAM estimate never fitted inside a
<tr>, so they were pushed into a drawer that could hold one model at a time,
could not be linked to, and had no room to say anything useful.

The gallery is now a rail to scan and a pane that answers. The pane has two
states and no third: with nothing selected it is the discovery page, and with
a model selected it is that model's detail. Selection lives in the URL, so a
model is linkable and Back steps out of the detail instead of off the page.

The rail groups by capability while browsing and flattens to results the
moment a term is typed. That is a rule rather than a toggle: once someone has
said what they are looking for, the buckets are between them and the answer,
and making the user choose would be handing them our problem.

The detail pane plots VRAM against context length with the host's own limit
drawn across it. This is new information, not a restyle. A single number
invites "so will it run?", and the honest answer is usually "yes, up to a 32k
context", which is a shape rather than a number. The estimates were already
fetched for every context size, so it costs no new request. Backends that
take no context length say so instead of being given a meaningless chart, and
a host with no GPU gets no chart at all rather than bars with nothing to
compare against.

The split-button variant menu goes with the actions column. The pane lists
every build with its backend, quantization, size, fit and a details
disclosure, each installable, which is what the dropdown was a cramped
substitute for. Its tests move onto that list; the three contracts it alone
carried (fetch-once caching, the loading state, an unfit build staying
installable) are backfilled against the pane.

RecommendedModels moves inside the pane, where it has the width to argue for
a model instead of listing one, and keeps its own dismissal and collapse.

Rail entries carry no description. Two lines is the budget and the second is
better spent on whether the thing will run; the stripped-Markdown contract
moves to the pane's lede, tooltip included.

e2e: 123 passing across models-gallery, navigation, recommended-panel,
model-artifact-operation, operations-strip and page-render-smoke. Inline
styles in Models.jsx drop from 82 to 41.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* refactor(ui): extract the split view into shared components

Discover shipped its rail, pane and detail header as private functions inside
Models.jsx. Backends and Host have the same defect and want the same shape, so
leaving them there guarantees three rails that drift.

SplitView, EntityRail, DetailHeader and StatGrid now live under
components/split/. EntityRail is deliberately data-driven: a surface maps its
own entity onto { id, name, icon, meta, stripe, groupId } and keeps its
vocabulary to itself, which is what stops the rail learning about models,
backends and loaded state all at once.

The CSS moves with it. What was .discover__rail is .entity-rail, .discover__
pane is .split-view__pane and so on, because a class named after one page is a
lie on the next two. Only what is genuinely Discover's stays behind the old
prefix: the shelves, the hero and the VRAM-by-context chart.

Two additions the shared rail needs and Discover did not: a state stripe, for
surfaces read by condition before they are read by name, and an empty label.
Discover passes neither.

No behaviour change. e2e 100 passing across models-gallery, navigation and
models-recommended-panel.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* feat(ui): put the backend gallery on the split view

Same defect as the model gallery, so the same shape: a seven-column table over
a click-to-expand row that was the only place the repository, licence, tags and
links could go.

The rail groups backends by the use case they serve, sharing Discover's
taxonomy on purpose: a backend is the runtime a use case needs, so "vision"
ought to mean the same thing one level down. It flattens on a query for the
same reason it does on Discover.

The zero state is the one real departure. A backend's fitness is not free
memory, it is the accelerator and platform it was built for, so the pane leads
with what this host is, then what is not installed yet, then whether anything
installed has gone stale. The table listed 37 runtimes and left "which of these
can even run here" entirely to the reader.

Distribution moves into the pane, which is the one thing a row could never
carry: which nodes hold a copy and which do not, with the install-on-more
control next to it rather than squeezed against a chip.

The distributed and target-node action logic is unchanged, including the guard
that keeps a hardware-specific build off the fan-out path. The split-button
popover loses its per-row anchoring because there are no rows; one pane, one
anchor.

Selection lives in ?backend=, preserving the ?target= scope rather than
clobbering it.

e2e: 139 passing across models-gallery, navigation, backends-management,
models-recommended-panel, nodes-per-node-backend-actions, page-render-smoke,
operations-strip and model-artifact-operation. The backends spec gains six
split-view tests; its three description-cell tests move onto the pane lede.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* feat(ui): put the Host inventory on the split view

The last of the three surfaces, and the one that is not a catalog. Both tabs
had the same click-to-expand row, so the shell transfers; what does not
transfer is the zero state, because there is nothing to discover in your own
inventory.

With nothing selected the pane reports what is happening: how many models are
loaded, what failed, what has an update, and which models are holding VRAM
right now. Every number was already on the page. None of them had been
assembled into one statement, so "what is going on" was a question the tabs
could not answer however long you looked at them.

The rail buckets by state rather than capability - Running, Idle, Disabled for
models; Update available, Installed for backends - which is the opposite of the
galleries and deliberately so: nobody opens Host wondering which of their
models does vision. Entries carry a state stripe for the same reason.

Load and Stop are promoted out of the kebab, because that is what an operator
came for; the rest stays behind the menu rather than diluting it. Adopted,
pinned and alias badges follow the model into the pane: they are facts about
the thing, not about its state, and the rail line is spent on state.

Deliberately NOT done: folding the two tabs into one rail, as the mock had it.
It costs five URL parameters, the manage-tab localStorage key and the
stat-card shortcuts, all of which are live deep-links today. The tabs stay as
the group selector; merging them is a follow-up with its own migration.

e2e: full suite 355 passing. New host-split-view spec; alias-template,
manage-logs-link, manage-action-menu-position and model-editor-back-nav move
off `.table` and the row kebab onto the rail and the pane.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* polish(ui): accessibility and consistency pass over the three split views

Findings from a pass over what the previous four commits actually shipped,
rather than what they were supposed to.

The rail was not a listbox. ARIA lets a listbox contain options and groups,
and nothing else, but each group's collapse control is a button that has to
sit inside the scroller with the entries it folds. It is now a labelled group
of buttons, which is the honest description; selection is announced with
aria-current and the arrow keys are unaffected.

Every entry was its own tab stop, so tabbing past a forty-entry rail to reach
the pane took forty keystrokes. Roving tabindex makes the rail one stop, and
arrowing now moves focus with the selection instead of leaving it behind on an
entry Tab can no longer reach.

The rail rounds its corners with overflow:hidden, which was clipping the focus
ring off the first and last entries entirely. Inset outlines fix it.

A 30px row is fine under a mouse and too small under a thumb, so coarse
pointers get a 44px target without costing density on a desktop.

One slot said three different things: "9 models loaded" on Discover, "12
loaded" on Backends, "3 of 9" on Host. All three lists are a page of a larger
set, so all three now say so the same way.

Also removed: an emptyLabel prop on EntityRail that nothing passed, its dead
CSS rule, and MODELS_COLSPAN and ResourceRowDesc, which died with the tables.

e2e: full suite 355 passing.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(ui): correct three defects only a real gallery exposed

Running the branch against a live instance with 1,595 models and 1,017
backends, rather than against mocked fixtures, surfaced three things the e2e
suite could not.

Grouping did nothing. The rails matched on the use-case keys the filter chips
send (`chat`, `tts`, `transcript`), but those are a server-side vocabulary the
handler maps onto entries. What entries actually carry is free-form and
inconsistent: models come back tagged `llm`, `gguf`, `vision`, `coding`, and
backends `LLM`, `text-to-text`, `audio-transcription`. Nothing matched, so
every model landed in "Everything else" and the feature was decorative.

Grouping now lives in utils/entityGroups.js, shared by both galleries, matching
case-insensitively against the vocabulary the API really uses, with the entry's
backend as a fallback signal - a backend named `whisper` is a speech backend
whatever its tags say. Order is specific before general and that is
load-bearing: a vision model is tagged `llm` too, so testing text first would
swallow it.

The zero state claimed GPU memory on a machine with no GPU. The resources
endpoint reports system RAM in the same field when gpu_count is 0, so the hero
read "84.4 GB of GPU memory" next to the recommendations panel correctly
saying "No GPU detected". The number was never wrong, only its label; it now
says system memory unless a GPU is actually present.

The page title still said "Install Models" under a nav entry saying Discover.

Also: the keyboard test named the model it expected to arrive at, which made it
a hostage of the grouping table and broke the moment the buckets were fixed. It
now asserts that the selection moves and returns.

e2e: full suite 355 passing.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(ui): the filters and the rail were fighting over the same job

Four things you find odd on Discover, and they turn out to be one mistake seen
from four sides.

The rail grouped the current page. The listing is paginated at nine rows, so
those bucket headers described nine entries out of 1,595, and turning a page
reshuffled the sections under the reader. The structure was never stable
because it was computed over the wrong set.

The chips were redundant for the same reason, seen from the other side. They
send tag= and filter all 1,595 server-side. The rail grouped nine of them
client-side by the same axis. Two controls for one job, and the weaker one was
the one this branch added, so it goes. Grouping stays only on Host, where the
list is complete, local, and bucketed by state rather than capability.

The search bar felt odd because it sat in a full-width band while the thing it
narrowed was a 290px rail below and to the left. The whole band now lives in
the rail column: search, backend, use cases, refinements, then the list it
narrows. One column to say what you want, one to show what you got. Nineteen
chips do not fit at that width, so they fold into a disclosure that states the
selection. A disclosure and not a popover, deliberately: picking use cases is
multi-select and interleaves with the backend select and the toggles below,
and a popover dismisses itself the moment you touch either.

The header held two counts and two buttons at arm's length from all of it. The
counts were the third statement of the same number on one screen, after the
rail's "9 of 1,247" and the pane's own headline, so they go. The buttons move
into the pane's zero state, which is the surface that answers "what do I do
here".

Also: the two first-run empty states wore .loading-center, which is
display:flex in the default row direction because it exists to centre one
spinner. With four children that put the icon, the heading, the sentence and
the buttons on a single line with no gap. They are now a proper full-height
empty state.

e2e: full suite 353 passing. Grouping tests are replaced by ones asserting the
rail stays flat; chip tests open the disclosure first; two filter-layout tests
that asserted the old three-band arrangement now assert the column.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* polish(ui): make Discover a full-height view, group the chips, name the refinements

Four things, all of them the same complaint: the page read as a document with
controls scattered on it rather than as one view.

The header is fused. A title block with its own padding, a subtitle and two
counts made the split view look like an attachment to a document that happened
to sit below it. It is now a slim bar carrying the title, the count and the two
page-level actions, and the split fills the rest of the window. Rail and pane
scroll independently, so the filters and the pane's headline stay put while a
long list moves under them.

The chips group. Nineteen in a flat row is a lot to scan even behind a
disclosure, and they already belong to the four families the rest of the UI
speaks, so they are bucketed by those. "All" sits on its own above them without
a heading, because it is a reset rather than a use case.

The refinements stop looking dumped. When the band became a column they were
three controls left where they landed; they now read as a named section with
one control per row.

The zero state suggests again. It had decayed into a "Browsing / 9 of 1,247 /
select a model" line that restated the count for the third time on one screen.
It now offers the four use cases as tiles that set the filter, which is the
shelf idea from the mock without inventing curation or paying for a second
fetch.

Two bugs found by looking at it rather than at the tests: the disclosure was
clamped to 190px, which cut it off partway through its third section so two of
the five never appeared at all; and the creation actions rendered twice, once
in the new bar and once in the pane hero a few pixels away.

e2e: full suite 353 passing. The chip-row test now holds its contract across
the per-family rows rather than a single one, and additionally asserts every
family is present and non-empty.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(ui): pin the split view's height so a long detail scrolls the pane

Selecting a model with a long description grew the whole page and dragged the
rail down with it, which is the opposite of what "full height" was supposed to
buy.

The flex chain was right and the ceiling was missing. .app-layout and
.main-content are min-height:100dvh, which is a floor: flex distributes free
space but nothing caps growth, so a pane taller than the viewport expanded the
column, the document scrolled, and the rail stretched to match. height:100% on
the pane then resolved against an auto-height parent and did nothing.

The chat route already solves this by pinning .main-content to 100dvh. The
same treatment now applies to any route containing a .page--app, selected with
:has() so the shell does not have to learn which pages happen to be split
views. Below the stacking breakpoint the pin is lifted, because two stacked
halves in two short scrollers is worse than a page that scrolls.

Measured on a live instance: document height stays at the viewport across
selection (950px either side) and the pane overflows internally instead.

Adds discover-height.spec.js, which asserts the page height and the rail height
are unchanged by selection and that the pane is the thing that scrolls. The
existing specs could not have caught this: they mock short descriptions, and
the bug only appears when the pane has more content than the viewport holds.

e2e: full suite 355 passing.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* feat(ui): give Backends and Host the full-height view, and fix the Update button

Backends now matches Discover: the header fuses into a slim bar carrying the
title, the count and the page-level actions, the filters move into the rail
column where they narrow the rail and nothing else, and the split fills the
window. Its seven chips fit at rail width, so unlike Discover's nineteen they
need no disclosure. Host gets the bar and the height; its resource monitor,
summary cards and tabs stay above the split, because those are read once while
the rail and the pane are worked in.

Two things the height change surfaced.

The console layout is a flex row with align-items:flex-start, so its body sizes
to content. Right for the pages it was built for, wrong for a split view, which
needs a ceiling to scroll inside: without it the Backends rail ran past the
viewport and over the footer. Pinned with :has() so only split-view routes are
affected.

The filters vanished when nothing matched. Both galleries swapped the whole
shell for an empty state, which took the search box and the chips with it, so
the page said "try adjusting your search or filters" while offering neither.
The shell now stays and the empty state moves into the pane.

Also fixes the Update control on Host, which had no className at all and
rendered as bare text, next to a status span that had picked up btn classes and
two copies of `fas` and so rendered as a button you cannot press. They have
swapped appearances back.

e2e: full suite 355 passing. The render-smoke selector learns .view-bar__title,
since the pages it checks no longer all use PageHeader.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(ui): keep the view mounted while searching, and bring rail grouping back

Searching replaced the whole view with a loader. The search box lives in the
rail column, so every debounced refetch unmounted the field being typed into
and dropped its focus with it. The list, the filters and the pane went too.

The shell now stays and the rail says it is busy: a sweep bar under its header
and the stale list dimmed, so the eye knows the answer is being replaced
without losing its place. A cold start still gets the skeleton, because there
is nothing to keep.

The condition for that is "nothing has loaded yet", not "the list is empty".
Those differ exactly when someone is editing a query that matched nothing, and
getting it wrong there would unmount the view on the keystroke after a
no-results search - the worst possible moment.

Grouping comes back on both galleries. It was removed because nine rows could
not fill five buckets, so a page turn rebuilt the rail's whole structure. That
was a symptom of the page size rather than of grouping: the rail now asks for
30 rows instead of 9 (Backends 60 instead of 21), which is enough for the
sections to read as structure and turns five times fewer pages. The order of
the sections is fixed, so what changes between pages is membership, not
arrangement.

Grouped while browsing, flat while searching, as before: once a term is typed
the buckets stand between the reader and the answer.

Also gives GalleryLoader a class and a testid instead of six inline style
declarations on a bare div, which is why nothing could select it.

e2e: full suite 359 passing, including a new spec asserting the search box
keeps its focus and its value across a refetch, and that a cold start still
shows the skeleton.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* perf(gallery): stop invalidating the VRAM estimate caches on every request

Searching or turning a page felt slow. It was not the search and not the
listing: /api/models answers in 3-9ms. It was the VRAM estimate, which the
gallery asks for once per row, and which took ~2.3s every single time however
often the same model was asked about.

pkg/vram already caches what makes that expensive - the remote content-length
probes, the GGUF metadata reads and the HF repo sizes. Those caches key on a
gallery generation counter, and AvailableGalleryModelsCached triggered a
background refresh on every call, with each refresh bumping the counter. One
page view is one listing request plus thirty estimate requests, each of which
re-read the gallery and started another refresh, so the generation moved
constantly and every cache entry was stale before it could ever be read. The
caches were dead in production.

Three changes, each doing one thing:

A refresh interval. The cached list is still served immediately; this only
decides how often re-fetching from upstream is worth starting. Five minutes,
as a package variable so tests can drive it without waiting.

A generation bump only when the gallery actually changed. An unchanged gallery
re-fetched on schedule must not throw away work that is still valid, which is
the difference between an estimate costing nothing and costing a network round
trip.

A separate "loaded" flag. The cache engaged on `cached != nil`, so a gallery
that legitimately holds nothing read as never-loaded and took the blocking path
on every call, bumping the generation each time. Found by the test for the
interval, which could not pass while this was true.

Measured against a live instance with 1,595 models:

  one estimate, repeated     2.3s  -> 2ms
  a page of 30, in parallel  10s   -> 0.04s

A first, genuinely unseen model still costs its remote probe. That is inherent;
what changed is that it is now paid once per model per gallery version rather
than once per request.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* perf(gallery): warm VRAM estimates at startup, and stop the UI waiting on them

Two halves of the same complaint: the gallery stalls on VRAM estimation.

Server side, the estimates are now warmed in the background at startup.
Estimating an entry nobody has asked about costs a remote probe of its weight
files, and the gallery needs one per row, so the first visitor was paying for
the whole page. The warm-up walks the gallery in the order the UI lists it, so
the first page is ready before anyone reaches it.

It is bounded and it never blocks: 300 entries at 4 at a time by default, on
its own goroutine, stopping with the server's context. Warming the whole
gallery would be thousands of probes on every boot, which is rude to the
upstream and slow to finish; warming nothing leaves the first page paying two
seconds a row. Anything past the limit still warms itself on first view.
LOCALAI_VRAM_WARM_LIMIT=0 turns it off for an air-gapped host,
LOCALAI_VRAM_WARM_CONCURRENCY=1 slows it for a metered link.

Client side, the page no longer waits on estimates it does not need yet. It
fired one request per row at once; a browser allows about six connections per
host, so thirty estimates took every slot and the request behind a click - the
variant list, an install - queued behind work nobody asked for. That is the
freeze: the list was already usable, and the UI was busy fetching sizes. Four
at a time leaves room for the interactive request to overtake, and a row whose
estimate is still in flight says "sizing…" rather than leaving a blank where a
number will appear.

buildEstimateInput moves to core/gallery as EstimateInput, since the handler
and the warmer both need it.

Measured against 1,595 models, from a cold boot:

  page 1, 30 estimates in parallel   10s -> 0.04s
  full warm-up (299 of 300 entries)  3m, in the background

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* chore: untrack data/.local_user_id and ignore the runtime data dir

`local-ai run` writes its instance state under ./data when started from the
repo root, which is exactly what a contributor testing a build does. The
identity file ended up committed on this branch by a `git add -A` while
verifying the gallery changes against a live instance.

Anchored, so it matches the runtime directory at the repo root and not a
`data` directory nested inside some package.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-02 19:28:36 +02:00
mudler's LocalAI [bot]
8a80830f33 chore: ⬆️ Update ggml-org/llama.cpp to a7a6d0d269c896218b6c78e0933bd6a17519d3f6 (#11283)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-02 18:15:49 +02:00
localai-org-maint-bot
7621939028 gallery: add Qwythos 27B variants (#11292)
Add the recommended Q4_K_M build and an MTP-enabled variant with the shared vision projector. Tag the existing Qwythos 9B MTP entry so serving-feature ranking recognizes it.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-02 18:14:40 +02:00
localai-org-maint-bot
cff69a05bf gallery: add Qwen3.6 27B Q8 variant (#11293)
Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-02 18:14:08 +02:00
localai-org-maint-bot
896b4b6785 gallery: add VibeVoice ASR BitNet variants (#11296)
Add the recommended TQ2 build and a smaller aggressive quantization for the CrispASR backend.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-02 18:13:49 +02:00
mudler's LocalAI [bot]
0990be35b7 chore: ⬆️ Update PrismML-Eng/llama.cpp to 9ca265a57f85f2117942490f421f64a226dd9847 (#11280)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-02 17:55:32 +02:00
mudler's LocalAI [bot]
d0119bf62c feat(chat): local-ai chat is now a terminal agent (#11291)
* chore(deps): bump cogito to v0.11 ahead of the nib harness

nib is the agent harness that becomes 'local-ai chat'. It requires cogito
v0.11, so pull that bump forward on its own: minimal version selection would
apply it to LocalAI anyway, and both repos use cogito and cogito/clients.
Landing it separately keeps the harness change reviewable.

nib itself is not pinned yet. Nothing in LocalAI imports it, and 'go mod
tidy' runs as a goreleaser before-hook in CI, so an unimported require line
does not survive. It lands with its first importer.

No LocalAI call site needed a change. Both cogito.WithMaxAttempts callers
guard the argument above zero, so v0.11's new clamp is unreachable, and
LocalAI's Multimedia values implement only URL(), so v0.11's new
TypedMultimedia routing treats them as images exactly as v0.10 did.

Binary size (cmd/local-ai): 200,301,381 -> 200,336,045 bytes (+34,664).
A throwaway probe that links nib measured 201,042,243 bytes (+740,862 over
the pre-change baseline).

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(chat): resolve and seed the agent state directory

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): write the agent config atomically and tighten its modes

Replacing config.yaml in place truncated it first, so an interrupted write
would have destroyed the api_key nib keeps in the same file. Stage through a
sibling temp file and rename over the target instead, and match nib's 0700
directory mode.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(chat): probe the endpoint and classify failures

Probe lists what a LocalAI endpoint advertises and separates the two
failures that need different advice: nothing listening, and rejected
credentials.

go-openai reports a rejected key as one of two concrete types depending
on the error body, and both occur against a real LocalAI. The normal
error handler sends an OpenAI error envelope, which arrives as
*openai.APIError; the opaque-errors handler replies with a bare status
and no body, which arrives as *openai.RequestError. Classifying on only
one of them misses half the cases, so the status is read from either.

A cancelled probe is not reported as an unreachable server, because it
learned nothing about the endpoint, and neither is a reply that could
not be parsed, because something did answer. Both would otherwise send
the user off to start a server that may already be running.

The model list is returned verbatim and in server order. LocalAI lists
whatever it finds in the models directory, including stray archives and
dotfiles, and deciding which advertised ids are real belongs to whoever
presents them.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(chat): resolve the model from flag, config, or the server

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(chat): pin that model resolution sorts a copy of the caller's slice

The sort spec asserted only on what the chooser was offered, so replacing the
defensive copy with an in-place sort of req.Available still passed all 37
specs. Assert the input slice's order after the call, so the guarantee cannot
be dropped silently by a later refactor.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(chat): offer to start a server when none is reachable

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): bound the server wait and pin readiness and stop semantics

Set cmd.WaitDelay so a backend subprocess holding the child's stderr pipe
cannot block cmd.Wait forever, which would leave exited unclosed, burn the
whole shutdown grace on a clean exit, and leak the waiter goroutine.

Two test gaps closed alongside it: the readiness spec now counts polls, so
treating 503 as ready is observable, and Stop's single-interrupt contract is
pinned by giving StartedServer interrupt/kill hooks that a spec can count.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(chat): drive Stop through one process interface, hide exec plumbing

Two independent interrupt/kill func fields plus a nil check admitted wirings
no test could distinguish: the pair swapped, so a SIGKILL would strand the
backends SIGINT exists to let local-ai run clean up, or kill left nil, so a
wedged server never escalates. One two-method interface that *os.Process
already satisfies leaves nothing to swap and nothing to nil.

Also translate exec.ErrWaitDelay, whose text names an os/exec struct field,
into what the user can act on. os/exec only substitutes that sentinel when the
process exited without an error of its own, so no exit status is swallowed.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(chat): replace the REPL with the built-in agent

local-ai chat is now the nib agent harness compiled into the binary: tool
use behind an approval gate, sub-agents, MCP, plugins, and skills, all
auto-configured against the local server.

The REPL goes with it. Its model listing and its 401 classifier were
duplicates of the ones Probe now owns, and the classifier was the version
that misreads a bare 401 with no OpenAI error envelope, so keeping either
would leave the package with two divergent answers to the same question.

github.com/mudler/nib lands in go.mod in this commit rather than earlier:
go mod tidy runs as a goreleaser before-hook on every PR, so a require
line with no importer is stripped before it reaches CI.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(chat): split the pre-agent phase out of Run and pin it

Everything before the handoff is testable and nothing after it is: once
app.Run owns the terminal there is no seam left. prepare draws that line,
takes interactivity as a parameter so the prompts can be driven over a
pipe, and hands Run the state dir, the model, and any server it started.

The questions move onto one prompter that owns its buffered reader. A
fresh bufio.Reader per question reads ahead and discards what it buffered,
so the model choice typed behind an answer to "start a server?" was lost
and the next question saw EOF.

choose answers with a list index and refuses an empty offer, so a value
that was never on the list cannot reach ResolveModel, which persists it
and starts every later run against it.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): bound each server check with a deadline

Nothing bounded the model listing, so pointing chat at an address that
accepts the connection and then never replies left the user with no
output and no offer to start a server.

The budget is context.WithTimeout rather than a cancel plus a timer.
Probe deliberately refuses to call an endpoint unreachable on a
context.Canceled, since a caller who gave up learned nothing about the
server, and only honours a deadline. A cancel-based budget therefore
expires as the one error that suppresses ErrUnreachable, exactly for the
hung servers the offer exists to rescue.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): tell the user when their model choice cannot be saved

The choice is meant to be asked for once. When saving it fails the user
is silently asked again on the next run, and the only trace was an
xlog.Warn: the agent runs at log level error, and a --log-level=error run
swallows it entirely.

ModelRequest gains Notify for exactly this class of problem, one that is
worth telling the user about but not worth failing over, and the chat
wiring points it at the same writer the question was asked on.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): stop a session's server when the process is signalled

A server started for the session is stopped by a deferred call, and a
signal skips deferred calls: a SIGTERM between the spawn and the exit
left 'local-ai run' reparented to init with nothing left that knew to
shut it down. Ctrl+C was already safe, but only incidentally, because the
child shares this process' foreground process group.

A signal handler rather than Pdeathsig on the child. Pdeathsig is
Linux-only and, in Go, is delivered when the OS thread that forked exits
rather than when the process does, so it can fire on a healthy parent.
Setpgid would break the Ctrl+C that works today by taking the child out
of the foreground group.

SIGHUP joins SIGINT and SIGTERM: a terminal program whose terminal is
gone has nobody left to talk to. The same context is what cancels the
agent, which nib leaves to its embedder.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): only skip the server checks for work that stays local

Two argument shapes were classified wrongly. Every 'mcp ...' invocation
counted as management, so 'local-ai chat mcp --stdio', which serves the
agent over MCP and needs a model like any other session, was handed an
empty one. And --init, whose shell snippet a user pastes into an rc file
long before any server exists, went the other way: it demanded a running
server to print a static string.

The mcp split is asked of nib's own IsMCPManageSubcommand rather than
restated here, so a verb added upstream cannot drift out of this list.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): exit with the agent's status instead of reporting it twice

nib writes what went wrong to stderr and returns nothing but an exit
code, so returning that error unchanged had main log "Error running the
application error=exit status 1" underneath the message the user had just
read. The refusal to render the full-screen interface into a pipe is the
one they meet in practice: it names --cli, and burying that hides the fix.

ExitCodeError says "already reported, exit with this status". main
honours it and prints nothing more, so a piped or redirected chat still
fails a script the way it should.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* style(chat): route interactive chatter through one writer helper

The prompts and notices all write to a terminal, where a failed write is
not worth failing the session over and the read that follows the question
reports the real problem. say says that once instead of five discarded
error returns.

The command's one-line help comes along: chat is no longer "an
interactive chat session".

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(chat): record why the agent gets this process' streams

Injecting them is what makes nib refuse to draw its full-screen interface
into a pipe and name --cli, instead of rendering onto a terminal the
caller may not own. The tradeoff is worth stating where the wiring is.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): stop the session's server on cancellation, not on the way out

The deferred Stop is reached only if the agent returns, and cancelling
the context does not make it: nib hands the TUI to bubbletea without the
context, so what actually unwinds a running session today is bubbletea's
own SIGINT and SIGTERM handler. SIGHUP has no such backstop, and
registering for it removed the default disposition that used to end the
process outright, so kill -HUP left a live TUI with a cancelled context
and the started server still running.

runSession watches the context alongside the agent and stops the server
the moment it is cancelled, so the guarantee no longer depends on what
the agent does with cancellation. Stop is idempotent, so the deferred
call stays correct and free.

The doc comment on shutdownContext described the mechanism it was
supposed to work by rather than the one that does. Corrected, bubbletea's
handler included.

ResolveModel now checks the chooser's answer against what it offered.
The shipped chooser answers by list index and cannot be wrong, but
ModelChooser is exported, the answer is persisted, and every later run
starts against it, so the invariant belongs at the consumer.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(chat): bump nib to v0.5.1

v0.5.1 carries four fixes that matter to 'local-ai chat':

- --init now names the embedder's command, so the emitted widget invokes
  'local-ai chat' rather than a bare 'nib' the user does not have.
- A piped CLI session that succeeds exits 0 instead of failing with EOF.
- EOF at a tool-approval prompt denies the call rather than approving it,
  and the session exits 3 (app.ExitCodeApprovalNoInput) so a script can tell
  "answered" from "refused to act" without reading stdout. Read-only tools
  are unaffected and still run. ExitStatus already unwraps app.ExitError,
  so the code propagates with no change here.
- RunTUI passes the context to bubbletea and gives up bubbletea's own signal
  handler, which makes shutdownContext the single owner of the signal and
  stops a SIGHUP leaving a wedged TUI behind.

Verified against a live server on 127.0.0.1:8080: the three --init shells,
a piped prompt exiting 0, a denied 'touch' that left no file and exited 3,
a read-only 'ls' that still ran and exited 0, and a SIGHUP that unwound a
TUI running under a pty.

Two comment blocks in run.go described the old TUI behavior and are now
wrong, so they are corrected in the same change. No behavior change: both
shutdownContext and runSession are untouched, and stopping the server on
cancellation is still worth keeping independent of how promptly nib unwinds.

One known gap, not addressed here. The widget --init now emits runs
'output=$(local-ai chat --height 50%)', and runAgent injects Stdout
unconditionally, so under $(...) nib refuses the TUI for a non-terminal
stream. This is the cost the runAgent comment already anticipated, now that
the snippets no longer hardcode standalone nib. Ctrl+Space should not be
documented until that is decided.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): let nib own stdout, so the Ctrl+Space widget works

The widget 'local-ai chat --init' emits runs
'output=$(local-ai chat --height 50%)', which puts a pipe on stdout by
construction. runAgent injected os.Stdout unconditionally, and nib refuses
every mode but --cli when a stream it was handed is not a terminal, so
Ctrl+Space printed "Re-run with --cli to use the injected streams" and
inserted nothing. Verified against a pty before and after.

nib reads a nil stream as "not injected" and falls back to the process
stream, which is how an embedder asks for nib's own behavior. That is what
stdout needs: the interface renders on /dev/tty but writes the chosen
command to stdout even when stdout is a pipe, and that write is the whole
of the shell-capture idiom.

Stdin is deliberately left injected. A piped or redirected stdin really is
ignored by the interface, so the refusal is the honest answer there, and it
is the one users meet: 'echo q | local-ai chat' still says to re-run with
--cli, once, exit 1. Nilling stdin the way stdout is nilled would delete
that silently. Stderr is not gated by nib at all and is unchanged.

One case does change and cannot be kept: 'local-ai chat > out.txt' from a
terminal no longer refuses, because it is indistinguishable from the
widget. It renders on /dev/tty and writes the capture line to the file,
which is what standalone nib does.

The app.Options literal moves into agentOptions so the decision is
reachable from a spec rather than being a detail of a function that takes
the terminal. Both sides of the asymmetry are pinned: reinstating
'Stdout: opts.Out' fails "hands nib nothing for the process stdout", and
nilling stdin fails "hands the process stdin over".

Also rewrites the last comments describing the pre-v0.5.1 behavior.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(chat): say what the stream refusal actually keys on

Two comments still called it the refusal to render the interface "into a
pipe". That was true when both stdin and stdout were injected, but a pipe on
stdout no longer refuses, so the wording now points at precisely the case
that was un-refused to make Ctrl+Space work. Only a stdin that cannot be
read triggers it, and both comments now say so and name the command a user
meets it with, 'echo q | local-ai chat'.

The agentOptions doc also said a "file a caller chose" stays injected and
refused, which reads as though 'local-ai chat > out.txt' still refuses. It
does not: a shell redirect arrives as os.Stdout and is nil-ed like the
widget's pipe, because the two differ only in being a regular file rather
than a FIFO and nib's gate does not look at that. What stays injected is a
writer an in-process caller chose for itself. Says that now, in the doc and
in the spec comment that had the same ambiguity.

Comments only. No behavior change.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: local-ai chat is now the built-in terminal agent

`local-ai chat` was a plain chat prompt and is now an agent that runs
shell commands behind an approval gate, so the pages that described a
REPL were wrong rather than merely thin.

Adds a Terminal agent feature page at /features/terminal-agent covering
the approval gate, piped runs and their exit codes, Ctrl+Space, model
resolution, state directory, and the pass-through management commands
(including the `--yes` caveat that leaves a plugin installed but
disabled in a script).

The three-way "looking for something else" notice becomes four-way and
moves into an agentic-routing shortcode. Four hand-kept copies of the
same paragraph is what produced the drift the new page would otherwise
have added to; the shortcode takes `current=` so each page still marks
itself, and errors the build on a name that is not one of the four.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* website: the agent is in the binary, not a second install

The nib section sold a separate tool you also install, with a GitHub
link as the only way in, which is now the wrong order: the agent ships
compiled into local-ai, and the standalone binary is the second reason
to care rather than the first.

Leads with `local-ai chat`, keeps nib as the SSH-anywhere story, and
adds a docs CTA pointing at the new Terminal agent page. id="nib" is
left alone because localai.io/#nib is linked from outside.

The two credits on the demo clip named nib as the thing that drove the
machine; they now credit the agent in LocalAI, which is the same agent.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* website: fix the exit keys, the plugin warning, and the redirect gap

Three claims on the chat-agent pages that the code does not back.

try-it-out told readers to press Ctrl+D. nib has no Ctrl+D handler: the
full-screen interface quits on Esc or Ctrl+C, and Ctrl+D is only an exit
in --cli, where it arrives as ordinary tty EOF. That sentence had
replaced the removed /exit and /quit text, so the page was left with no
working way to leave a session. Document both modes, since they differ.

The plugin warning said nothing tells you the install stopped short. It
does: the command prints that the plugin was left disabled. What it does
not do is say so in its exit code, which is 0 either way. That is the
part a script cannot work around, and it is the reason to pass --yes.
Overstating it in the paragraph that gives the advice only makes the
advice easier to dismiss.

Redirecting stdout no longer refuses; the interface goes to /dev/tty and
only the yanked command reaches the file. It is what lets the Ctrl+Space
widget capture a command at all, since a redirect and out=$(...) are the
same thing to the stream gate. It was documented nowhere. A non-terminal
stdin is still refused, and the new text says which of the two it is.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): make the CLI flags outrank the agent config file

local-ai chat routed --endpoint, --model, --api-key, --trace-dir and --yolo
through nib's app.Options.Defaults. Defaults are seeds: they sit beneath the
config file, so the file silently undoes them. That made the flags accepted and
inert, and not in an edge case, since EnsureStateDir writes base_url on the
first run and the interactive picker writes model, so from the second run on
the file carried a value for both.

Observed against a live server: with base_url: http://127.0.0.1:9999/v1 in the
config and --endpoint http://127.0.0.1:8080 on the command line, the probe hit
8080 and every agent turn posted to 9999. With model: gemma-4-e2b-it-qat-q4_0
in the config, --model lfm2.5-8b-a1b was ignored on the wire.

nib v0.6.0 adds app.Options.Overrides, applied above the config file and above
the bare environment block. Move the whole block there: all five values are
decisions this invocation already made on the user's behalf, and a flag the
config file can undo is not a flag. Nothing is left in Defaults, because
LocalAI's one genuine seed, the initial base_url, is written into the config
file by EnsureStateDir rather than handed to nib.

Two limits come with the channel and are documented on agentOptions rather than
worked around. An override can only raise a field, since nib cannot tell "set
to the zero value" from "not set", so --yolo can turn approval off but nothing
on the command line turns it back on over an approval_mode: auto in the file.
And nib's own NIB_TRACE_DIR and NIB_YOLO are resolved after the config load and
still outrank these, deliberately, upstream.

The existing spec pinned that the right values reach app.Options, which they
always did, which is exactly why it could not see nib discarding them. The new
specs resolve the config the way app.Run resolves it, against a real config
file that disagrees with every flag, and one asserts Defaults stays empty.

docs/content/features/terminal-agent.md already documented --model as winning
over the saved model; that was false before this change and is true now, so no
docs edit was needed.

Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(chat): document intentional config file read

Assisted-by: Codex:gpt-5 [gosec]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-02 09:23:26 +02:00
mudler's LocalAI [bot]
359bd4850d docs(blog): bring the 4.8 release post up to the final changelog (#11287)
The post was written against the first draft of the release notes, when the
cycle stood at 214 PRs over thirteen days. It closed at 321 PRs over eighteen
days, and three of the larger user-facing changes landed after it was written.

- Correct the counts throughout: 321 PRs, eighteen days, 24 contributors
  (11 first-time), gallery 1,221 to 1,505.
- Add sections for the three new capabilities: 3D generation as a modality
  (Generate3D, FLAG_3D, /v1/3d/generations, trellis2cpp), audio.cpp serving
  six audio endpoints from one process, and the operations bar becoming the
  Activity page.
- Cover the two further hardening fixes (tar hardlink escape, cyclic $ref
  stack overflow) alongside the TRL one.
- Note the Valkey store, systemd socket activation, persistent trace history,
  in-place chat edits, the self-contained SYCL backend and the site split.
- Group the new-engine sections together rather than splitting them across
  the operational ones.

Embeds the existing vllm-race and magpie clips, and adds a 3D generation clip
cut from the demo recording to the conventions in .agents/preparing-a-release.md
(no audio track, 14s, named for the feature). blog.css styled figure img but
not figure video, so a clip in a post rendered outside the card; both selectors
now share the rule.


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-02 00:37:19 +02:00
mudler's LocalAI [bot]
a49f115b0d chore: ⬆️ Update ikawrakow/ik_llama.cpp to 0be97a7a5ad113f33e08729261649ccea2cdc5ff (#11282)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 23:53:14 +02:00
mudler's LocalAI [bot]
0d6b38e709 chore: ⬆️ Update 0xShug0/audio.cpp to 545e29a6f2fde24298cb3b0f07baab4352987ac9 (#11281)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 23:52:58 +02:00
mudler's LocalAI [bot]
bef30732cd chore: ⬆️ Update CrispStrobe/CrispASR to 66ac7843e319b588f5410051c575affd19424fb3 (#11279)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 23:52:44 +02:00
mudler's LocalAI [bot]
5b7ca31bd1 chore(model-gallery): ⬆️ update checksum (#11285)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 23:52:32 +02:00
mudler's LocalAI [bot]
21ecc799e5 fix(qwen3-tts-cpp): hold qwentts.cpp at 35ebe537, upstream master hangs in synthesis (#11286)
tests-qwen3-tts-cpp has been failing on master since 2026-07-31. The suite
loads every component fine and then stops: TTS() never returns from the
native call, so a job that takes ~5 minutes runs into the 20 minute Go test
timeout instead.

    goroutine 74 [syscall, 19 minutes]:
    github.com/ebitengine/purego.RegisterFunc.func4
    qwen3-tts-cpp.(*Qwen3TtsCpp).TTS  goqwen3ttscpp.go:154
    qwen3-tts-cpp.init.func2.4        e2e_test.go:90

Not a flake: reproduced on master and again on an explicit re-run.

Bisected across this cycle's eight qwentts.cpp bumps by their own check:
10832, 10850, 10902, 10964, 11006, 11039 and 11127 all pass in ~5 minutes;
11241 (abab6b3) fails at 1h58m. That PR was merged with this check already
red, which is how the hang reached master.

35ebe537..abab6b3 is three upstream commits, and the only functional one is
26dd8adb, "predictor: unroll the frame into one cgraph and sample in standard
ops", which is consistent with a generation loop that never reaches its stop
condition.

Hold the pin at the last known-good commit. The bump entry is commented out
rather than left in place, because it tracks upstream master and would put
the hang straight back on the next nightly run. Both spots carry a pointer to
the other so the hold is discoverable, and restoring it is uncommenting four
lines once upstream is fixed.


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-01 23:40:35 +02:00
localai-org-maint-bot
9fe1165f61 fix(turboquant): retain CPU variants in GPU builds (#11276)
Select the CPU_ALL_VARIANTS target for x86 GPU images so partial offload uses runtime-selected host kernels. Keep GPU arm64 builds on the portable fallback until their toolchains consistently provide gcc-14.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-01 16:06:36 +02:00
mudler's LocalAI [bot]
ad2be8a856 chore: ⬆️ Update ggml-org/llama.cpp to 876a4321163249c43ca4e986818fab5ab081f282 (#11177)
* ⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(llama-cpp): drop merged MiniMax-M3 patch

The bumped llama.cpp revision includes the MiniMax-M3 parser and template detection, so the carried patch now rejects during backend preparation. Remove the obsolete patch while retaining the independent score-task patch.

Assisted-by: Codex:gpt-5 [systematic-debugging]

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-01 16:06:13 +02:00
localai-org-maint-bot
3c02d2aa4d gallery: add Inkling Small GGUF variants (#11273)
Add Q4_K_M and IQ2_M sharded llama.cpp entries with the BF16 multimodal projector.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-01 14:14:59 +02:00
localai-org-maint-bot
c0a9c42771 gallery: add Fara1.5 9B GGUF variants (#11271)
Add the new 9B Fara computer-use model alongside its existing 27B sibling, with Q4_K_M and Q8_0 llama.cpp variants plus the required vision projector.

Assisted-by: Codex:gpt-5 [web]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-01 11:49:26 +02:00
localai-org-maint-bot
cedcbf97a9 fix(llama-cpp): retain CPU variants in GPU builds (#11255)
Build the runtime CPU variant set alongside x86 GPU backends so partial offload uses the host's SIMD kernels instead of the scalar fallback. Keep arm64 GPU images on the portable binary until their builders consistently provide gcc-14.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-01 09:26:23 +02:00
mudler's LocalAI [bot]
7e4a60c701 chore: ⬆️ Update TheTom/llama-cpp-turboquant to 8a891f4b566efdbd3cea92fafee3227a0a267683 (#11258)
⬆️ Update TheTom/llama-cpp-turboquant

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 09:25:51 +02:00
Zelys
76927ccde3 fix(utils): reject tar hardlinks that escape the extraction root (#11266)
* fix(utils): reject tar hardlinks that escape the extraction root

ExtractArchive pre-scans archive members and rejects symlinks, but tar
hardlink entries carry a regular file mode and so pass that check.
Header.Linkname was never validated, so an archive could create a link
to a path outside the destination directory.

Validate Linkname with the same path check already applied to member
names. Hardlinks that resolve inside the extraction root still extract,
so ordinary archives are unaffected.

pkg/oci/image.go already resolves tar.TypeLink targets before using
them; this brings the archive extraction path in line with it.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Zelys-DFKH <zelys@dfkhelper.com>

* test(utils): cover hardlink overwrite and in-root hardlinks

The existing hardlink test names a link target two levels above the
extraction root, so its final assertion checked a path the link never
resolved to and could not fail. Point the target one level up instead,
at the path that assertion already names.

Add two cases. The first uses a .tar.gz, where ExtractArchive binds a
Tar config with OverwriteExisting set, and follows the link entry with a
regular entry of the same name. Before the fix that pair linked to a
file outside the root and then truncated it through the link, which the
plain .tar case does not reach. The second extracts a hardlink whose
target is an earlier member of the same archive, covering the claim that
ordinary archives are unaffected.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Zelys-DFKH <zelys@dfkhelper.com>

---------

Signed-off-by: Zelys-DFKH <zelys@dfkhelper.com>
2026-08-01 09:25:35 +02:00
localai-org-maint-bot
fca7ab2df4 fix(gallery): correct Nanbeige 4.2 artifacts (#11269)
Use the case-sensitive Hugging Face filenames and refresh the linked SHA256 values for both gallery variants.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-01 09:13:59 +02:00
mudler's LocalAI [bot]
04764bbe89 chore(model gallery): 🤖 add 1 new models via gallery agent (#11268)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 09:13:25 +02:00
localai-org-maint-bot
2f3dd404b5 feat(import): route MLX TTS models to mlx-audio (#11267)
Detect text-to-speech MLX repositories during model import and emit a TTS-ready mlx-audio configuration. Expose mlx-audio in the backend preference dropdown for repositories without complete metadata.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-01 09:12:54 +02:00
mudler's LocalAI [bot]
a7440f032d chore: ⬆️ Update PrismML-Eng/llama.cpp to 4dd165625bb6c020285eec8b342af25cf60233dd (#11259)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 09:11:48 +02:00
mudler's LocalAI [bot]
4a6cd227a3 chore: ⬆️ Update 0xShug0/audio.cpp to f78227c52736a4792a50aa3f82ead7e7385c891b (#11261)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 01:23:41 +02:00
mudler's LocalAI [bot]
740d8684b5 chore: ⬆️ Update ggml-org/whisper.cpp to 2ca53bb45e38748d07b310eeb36245a7157ac882 (#11263)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 01:23:29 +02:00
mudler's LocalAI [bot]
cb432c4c99 chore: ⬆️ Update CrispStrobe/CrispASR to b5211ac635489049ee8ce86a82d69faa18e8d8da (#11264)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 01:23:18 +02:00
mudler's LocalAI [bot]
a4cd387100 chore: ⬆️ Update localai-org/rf-detr.cpp to 98d0f381b832ef08a608b65c7dd78db066ed8b9a (#11260)
⬆️ Update localai-org/rf-detr.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-01 00:49:42 +02:00
Dimitris Karakasilis
c089caf320 feat(sycl): make the intel llama.cpp backend self-contained on any host (#10991)
* feat(sycl): make the intel llama.cpp backend self-contained on any host

The SYCL backend shipped an incomplete oneAPI runtime AND relied on a
host-provided GPU driver, so it only ran inside the build container. On a
bare host it died with "libze_loader.so.1 / libdnnl.so.3: cannot open
shared object file", and even with the host's Intel driver installed it
SIGSEGV'd during SYCL init when the host driver was built against a newer
glibc than the backend's bundled loader (rolling-release distros).

package_intel_libs now bundles the complete, coherent oneAPI runtime
(the missing MKL ILP64 / sycl_blas / tbb_thread + oneDNN + the dlopen'd
UR adapters, plus a sweep of the backend binaries' own direct deps) and
the Intel GPU userspace driver (libze_intel_gpu + libigdrcl + IGC + gmm)
with its OpenCL ICD manifest, mirroring how package_vulkan_libs bundles
Mesa. run.sh points the Level Zero and OpenCL loaders at the bundled
driver, and install-base-deps.sh installs it in the SYCL build image.
Bundling the driver is safe across kernels because it talks to the host
i915/xe via the stable DRM UAPI (unlike NVIDIA's kernel-locked
userspace).

Validated on Arch (glibc 2.43, i915): the backend loads and runs on an
Iris Xe with no host Intel packages installed.

Assisted-by: Claude:claude-opus-4-8

Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me>

* fix(sycl): install a driver that exists, and let the user choose their own

The driver install added earlier in this branch asked apt for
intel-level-zero-gpu, which is not a package in Ubuntu 24.04. apt fails
outright on an unknown name, so neither driver was installed, nothing was there
to copy, and the images carried no driver at all.

It now comes from Intel's own repository, which has 25.18 for this Ubuntu
release, against 23.43 from late 2023 in the Ubuntu archive. The archive driver
does not know any card released since, so a machine with a recent Intel GPU
would end up carrying a driver that cannot drive it. Anything that goes wrong
during that install fails the build on purpose: an unreachable repository is a
passing problem that a retry fixes, while quietly carrying a different driver,
or none, is a difference nobody would notice until a user reports an idle GPU.

run.sh used to overwrite whatever driver the user had chosen. Level Zero uses
only the driver it is given, so on a machine with a card too new for the
carried driver, the GPU would go unused with no way back. Both that setting and
the OpenCL one are now left alone when already set, and the docs say how to
point a backend at the machine's own driver.

The OpenCL setting also used to be applied whenever the backend held a driver
list, even when the driver it named had not been copied, which leaves OpenCL
with nothing instead of falling back to the machine's own driver. It now
requires the copied driver to be present, and the packaging leaves out the list
entry of any driver it did not copy. The oneAPI images list a processor-only
OpenCL library, which was being carried with nothing behind it.

Two more corrections in the packaging. The scan for libraries a program is
linked against only looked at files named llama-cpp-*, so turboquant and bonsai,
which are also built for Intel GPUs, were left with the incomplete set of
libraries this branch set out to fix; it now looks at every program in the
directory. And a build that should carry a driver but ends up without one now
says so, which is what a stale prebuilt base image looks like: such a backend
still runs on a machine that has its own driver, so nothing fails and the only
other symptom is a user reporting an idle GPU.

Backends now also ask the driver to report how much graphics memory is free,
without which llama.cpp reads zero on an integrated GPU, since such a chip
shares the system memory instead of having its own. turboquant and bonsai get
the same run.sh handling as llama.cpp.

The driver is only carried by the builds that start through run.sh, because
run.sh is what points Level Zero and OpenCL at it. The Python backends for
Intel GPUs start differently and would never load it, so they keep using the
machine's own driver rather than carrying several hundred megabytes they cannot
use.

Checked in a container on Ubuntu 24.04: the install brings driver 25.18 with
the files where the packaging expects them, an unreachable repository fails the
build, and the copied set resolves on its own once the machine's Intel packages
are moved away.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me>

* fix(ci): rebuild every Linux backend when the GPU packaging script changes

scripts/build/package-gpu-libs.sh decides which GPU libraries end up inside an
image. The filter that builds the backend matrix listed it as an input of the
Python images only, so changing it rebuilt no Go and no C++ backend, even
though those run it from their own package.sh. A packaging fix aimed at the
Intel llama.cpp backend could merge and reach no image, which is the same
failure this rule was written to prevent.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me>

* fix(sycl): carry only the driver Level Zero uses, not the OpenCL one

llama.cpp reaches an Intel GPU through Level Zero, which hands the driver
programs that are already compiled and so needs only the back end of the
graphics compiler. The OpenCL driver can be handed source code instead, so it
needs the compiler's front end as well, and that arrives with its own copy of
clang. Carrying it cost about 139 MB in every backend built for Intel GPUs, and
took the carried set from 123 MB to 261 MB.

Nothing here takes that path. No LocalAI code selects an OpenCL device, each
backend image holds one backend, and the documentation never described OpenCL
as a way to run models: the only mentions are a stale clblas row in the
BUILD_TYPE table, for a llama.cpp backend that no longer exists and that no
build matrix entry uses, and the sycl-ls troubleshooting hint. Before this
branch the packaging carried the OpenCL loader and adapter but no driver, so
the path could not work in a released image either. There is nobody to keep
working.

The driver list that OpenCL reads is no longer carried, and run.sh no longer
sets OCL_ICD_VENDORS, so OpenCL inside a container keeps using whatever the
image provides rather than being pointed at a directory with no driver in it.

Checked in a container against the real 25.18 driver: the carried set is 123 MB
with nothing unresolved, and Level Zero still reports the GPU with the
machine's own Intel packages moved out of the way. Neither the Level Zero
driver nor the compiler back end names the front end or clang among the
libraries it opens by name, so the leaner set is complete for this path.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me>

---------

Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-31 23:39:53 +02:00
localai-org-maint-bot
9584377a50 feat(chat): edit saved conversation messages (#11189)
* feat(chat): edit saved conversation messages

Add inline edit, save, and cancel controls for stored user and assistant messages without triggering inference. Preserve structured message attachments and cancel edits when streaming starts.

Assisted-by: Codex:gpt-5

* test(chat): preserve seeded conversation on reload

The saved-message edit test reloads the page to verify persistence, but its init script was replacing localStorage with the original fixture on every navigation. Seed only an empty store so reloads exercise the data written by the application.

Assisted-by: Codex:gpt-5 [Codex]

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-31 23:38:03 +02:00
mudler's LocalAI [bot]
51c9cc1934 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 3f53a059024039358e9fef75b5dc0c99dbcb40f9 (#11262)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-31 23:37:41 +02:00
localai-org-maint-bot
22e401b43d docs: fix local Hugo working directory (#11183)
Direct repository-root users to the supported make docs target and document the equivalent direct Hugo invocation from docs/.

Fixes #10062

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-31 23:37:27 +02:00
mudler's LocalAI [bot]
11403f4797 chore(model-gallery): ⬆️ update checksum (#11265)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-31 23:36:46 +02:00
Ettore Di Giacinto
aa5a9c483a fix(website): connect the runtime, the engines and APEX into one thread
The page reads as a list of features with nothing joining them, so two
things did not land.

The engines section never said these are the backends LocalAI loads. The
runtime section describes a core that pulls each engine in on demand, and
the engines section describes engines written from scratch, and nothing on
the page connected the two sentences. Readers were taking parakeet.cpp and
the rest for unrelated side projects by the same people. The lede now says
whose backends they are before it says anything else.

APEX was used as a known term on first appearance, in a section that opened
onto a benchmark table. Nothing said what it is or why it follows the
engines. It now opens by placing itself in the stack: the engine decides how
fast a model runs, the weights decide whether it runs at all, and APEX is
the second of those. Then the numbers.

Also drops "Most backends wrap somebody else's engine. These do not", which
is the machine-written antithesis shape, and fixes a list that broke its own
parallel halfway through.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m]
2026-07-31 21:28:21 +00:00
Ettore Di Giacinto
4b3978dcba chore(website): derive the counters from data, refresh them weekly
The star, fork, contributor and release counts were typed into the templates
by hand, so they only moved when somebody remembered. They had already
drifted: stars read 48,042 against 48,067, forks 4,314 against 4,320, and
contributors 224 against 225.

They move to website/data/stats.yaml, which .github/ci/refresh-site-counters.sh
rewrites from the GitHub API, run weekly by a new workflow. The contributors
and releases endpoints never report a total, so the script asks for one item
per page and reads the count out of the Link header. It refuses to write a
zero or a non-number, which is what a rate-limited or failed call looks like,
and the workflow commits only when a number actually moved. The Discord count
has no API behind it, so the script reads the existing value back and carries
it through.

The engine count was wrong in a second way. The hero said 18, the section
heading said "Eighteen engines", the timeline said "Nineteen engines of our
own", and the /engines/ page derived 19 from the data file. All of them now
derive from that same file, so they cannot disagree again.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m]
2026-07-31 21:23:54 +00:00
localai-org-maint-bot
3f4e446adc gallery: add Qwopus3.6 27B Fusion variants (#11257)
Add Q4_K_M and Q8_0 llama.cpp entries for the newly released Qwopus3.6-27B Fusion reasoning and coding merge, with MTP enabled.

Assisted-by: Codex:gpt-5 [Hugging Face API]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-31 23:05:02 +02:00
Ettore Di Giacinto
e6b235baf2 fix(website): wrap the timeline, run integrations as a reel, fix blog cards
The timeline set six 15rem columns in a flex row with overflow-x:auto, which
needs 90rem and so scrolled sideways on any normal laptop. It is a wrapping
grid now, and the rule that carries the dots moves from the container onto
each item so a wrapped row still gets a line above it. Column gap is zero and
the items carry their own right padding, so the rule stays continuous.

Integrations move from a card grid to a reel. Any single integration is a weak
signal and the whole moving line is the strong one, so the count is doing the
argument. It pauses on hover and on keyboard focus, since the names are links.
The list grows from 8 to 26: Open WebUI, Dify, LibreChat, RAGFlow, Continue,
big-AGI, Nextcloud, Frigate, promptfoo, Mods, TypingMind, baibot, k8sgpt-
operator and others. Each was admitted only after opening that project's own
repository or docs and reading the line that names LocalAI. The ones that
failed that test are listed in the data file so nobody re-adds them.

The blog cards were hand-written, which is how one of them came to advertise
"Porting vLLM to C++", a post that does not exist, and how all three linked to
the blog index instead of an article. They range over the posts now.

The section intro used the "a changelog tells you what moved, these posts show
you what it does" shape, which is the standard machine-written antithesis. It
states what the posts contain instead, including the perplexity regression
that APEX costs, because publishing the price is the actual claim.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m]
2026-07-31 14:55:16 +00:00
Ettore Di Giacinto
0bedc75921 fix(website): rewrite the ecosystem band, drop two false coverage links
The band led with three sentences of hedging and printed a commit count
next to each employer, so a one-commit entry beside a large name read as
weakness rather than as the modest, true claim it was. It now opens on the
contributor count, sets the employers as a sentence instead of a pill wall,
and keeps the caveat to one line. The counts stay in ecosystem.yaml, since
they are the provenance for the list and anyone re-checking it needs them.

Two "coverage" cards were not about this project. The modelslab.com piece
reviews Frikallo/parakeet.cpp, an unrelated project of the same name, and
the snailtext.app benchmark measures Parakeet through ONNX Runtime without
mentioning LocalAI at all. Both are removed, along with the contributor
card that duplicated the band's opening line.

Press was four posts from one vendor, which read as the whole of the
coverage rather than one enthusiastic outlet. SUSE collapses to a single
series entry, and Pulumi, Semaphore and Spectro Cloud join it. Each was
opened and checked against the project before being added. K8sGPT and
LlamaIndex join the integrations; both document LocalAI as a backend.

The quotes move above the lists so the section opens on its strongest
line, which is somebody else's. The hero gains a GitHub call to action,
the APEX collection link was returning 404 and is corrected, and the
footer no longer describes the site as a design mock.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5[1m]
2026-07-31 14:24:26 +00:00
Ettore Di Giacinto
dad4d5956a ci: move lint back to hosted runners, arc image has no make/gcc
Follow-up to 42541dd4f, which routed lint to arc-runner-set. Both of its
jobs failed there in one second (run 30637392862): the runner image has
git, curl, unzip, tar, ldd and python3, but not make, and build-scripts
additionally needs gcc because the packaging-script tests compile a
throwaway binary and inspect it with ldd.

golangci-lint needs make twice over: `make protogen-go` (which also wants
curl + unzip to fetch protoc) and `make lint` itself. So both jobs go back
to ubuntu-latest.

gh-pages.yml stays on arc-runner-set and is unaffected: it uses no make and
no C toolchain, and setup-go / actions-hugo fetch their own toolchains.

The preflight steps stay. They cost about a second on the hosted pool, and
they are what turned this into a one-second named failure instead of an
opaque one midway through a build. When the runner image gains make + gcc,
re-routing is one runs-on line per job. Any such re-route must stay
push-only: lint also runs on pull_request, and fork PRs execute untrusted
code that must not reach a persistent self-hosted runner.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
2026-07-31 14:11:45 +00:00
Ettore Di Giacinto
42541dd4f6 ci: route site deploy and lint to the self-hosted runner
The GitHub-hosted runner pool is shared per ACCOUNT, not per repo, so a
burst in one repo starves every other. On 2026-07-31 it reached zero
scheduled jobs for 35 consecutive minutes with 39 jobs queued, while
arc-runner-set completed 12 jobs without interruption across the same
window. Actions was healthy globally (other public repos were scheduling
normally), so this is an account-level throttle we cannot fix from inside
the workflows, only route around.

Site publishing and lint are small, run on nearly every commit, and gain
nothing from waiting behind a saturated hosted queue, so both move to
arc-runner-set, the label already proven in generate_intel_image.yaml.

lint.yml is routed for PUSH ONLY, and this is the important part: that
workflow also triggers on pull_request, and a fork PR executes untrusted
contributor code. Running that on a persistent self-hosted runner would be
a real compromise vector, so anything that is not a push to mudler/LocalAI
stays on the ephemeral hosted pool. gh-pages.yml needs no such clause: it
triggers only on push-to-master and workflow_dispatch, so it never runs
pull-request code. Both carry a repository guard so forks, which have no
such runner label, fall back to hosted instead of queueing forever.

Neither workflow uses sudo or apt, and both fetch their own toolchains via
setup-go / actions-hugo. A self-hosted image can still be leaner than the
hosted one, so each lint job opens with a preflight that names the missing
tool (curl/unzip/make for protoc and lint; gcc/ldd/python3 for the
packaging-script tests) rather than failing opaquely mid-build. Reverting
is one runs-on expression per job.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
2026-07-31 14:08:49 +00:00
Ankit Aglawe
fb54d0faab gallery: add Parable Claude-Fable-5 agent-trace models (3B/4B/8B) (#10930) 2026-07-31 13:10:48 +02:00
localai-org-maint-bot
314a824039 gallery: consolidate POCKET-35B variants (#11249)
Keep the canonical POCKET-35B family, add its missing Q3_K_M build, and remove the duplicate artifact entries introduced by overlapping gallery additions.

Assisted-by: Codex:gpt-5 [web]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-31 13:07:46 +02:00
mudler's LocalAI [bot]
b60b01d783 chore: ⬆️ Update 0xShug0/audio.cpp to f32876cfb45732dd4f43264e9104d229e95b0bc3 (#11233)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-31 10:01:44 +02:00
mudler's LocalAI [bot]
5d461ec7d2 chore: ⬆️ Update CrispStrobe/CrispASR to 677e95d0e60010f10636c3a0b1ba215b38a4a943 (#11234)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-31 10:01:28 +02:00
mudler's LocalAI [bot]
25a8a73b35 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 9992f6b515ee63c7d6f7beee6b8414b0a6d1dd43 (#11235)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-31 10:01:14 +02:00
mudler's LocalAI [bot]
e356315f9c feat(swagger): update swagger (#11236)
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-31 10:00:59 +02:00
mudler's LocalAI [bot]
4076b32d42 chore: ⬆️ Update leejet/stable-diffusion.cpp to e31a86ce9110b11a98bd5990c329093244c2d1e3 (#11237)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-31 10:00:41 +02:00
mudler's LocalAI [bot]
d2be530d14 chore: ⬆️ Update ggml-org/whisper.cpp to 4523d0ce373ee4b2176b3251fff29fd4864fcf38 (#11240)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-31 10:00:28 +02:00
mudler's LocalAI [bot]
735420c216 chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to abab6b3bf317cfa1b788efce1d25f4f9239395ad (#11241)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-31 10:00:16 +02:00
localai-org-maint-bot
5e98f898db gallery: add Antares 1B GGUF variants (#11246)
feat(gallery): add Antares 1B GGUF variants

Add Q4_K_M and Q8_0 builds of the Granite 4.0-based security agent, linked as selectable gallery variants.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-31 09:20:52 +02:00
localai-org-maint-bot
f01589d98b fix(oci): identify signature verification requests (#11244)
Signed backend verification performs separate registry requests for manifests and referrers. Reuse LocalAI's version-aware User-Agent there so the full install flow is attributable to LocalAI.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-31 09:20:26 +02:00
mudler's LocalAI [bot]
daab94134c feat(website): add an ecosystem band, and an ADOPTERS file to back it (#11248)
Adds the "who turns up around this project" section, split into three lists
because the evidence behind each one is a different strength and collapsing
them into a single logo wall would overclaim.

  Contributors   21 companies whose engineers have commits here. Evidence is
                 the commit history plus the employer on that person's public
                 GitHub profile, so it is a claim about the person. Commit
                 counts are shown next to each name, including the ones that
                 are a single patch, because hiding that would be the whole
                 problem.
  Integrations   six projects that reference LocalAI in their own repository
                 or documentation, which anyone can verify without asking us.
  Press          four SUSE Communities articles about running LocalAI.

Names are set in type rather than fetched as logos. A logo reads as
endorsement, and a one-line typo fix from somebody who happens to work at a
large company does not support that, quite apart from what their trademark
policy says about it.

ADOPTERS.md is the mechanism for the stronger claim. An organisation that
wants to be listed as a user opens a pull request adding itself, which is both
the evidence and the permission, and is publicly auditable afterwards. The
file says plainly what the website does and does not claim, so the next person
to ask "can we add some big names" has the answer in the repository.


Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] [WebSearch]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-31 09:18:51 +02:00
mudler's LocalAI [bot]
94d5affcea feat(website): split the site, move docs to /docs, add a landing page (#11243)
* feat(website): split the site, move docs to /docs, add a landing page

The Hugo docs site has always been localai.io itself, which left nowhere to
explain what LocalAI is or show what the team builds. This adds a separate
marketing site at the root and moves the documentation under /docs/.

Docs:
  The existing site keeps its content tree and its Relearn theme, and now
  builds with baseURL <root>/docs/. Its _index.md, which held a hand written
  landing page, becomes a real documentation home.

  Every previously published URL keeps working. GitHub Pages has no server
  side rewrites, so .github/ci/gen-redirects.sh walks the built docs output
  and leaves a meta refresh plus a canonical link at each old root path. It
  covers bare .html files too, which is what keeps /gallery.html alive, and
  it never overwrites a path the marketing site already owns.

Website:
  A second Hugo site under website/ with its own layouts and no external
  theme, so the marketing side does not have to fight Relearn's home rooted
  menu and asset pipeline. CI builds both and merges them into one Pages
  artifact.

  The design is derived from the project logo rather than invented: the navy
  of the triangle, the cyan of the llama, the purple of the speed bars. Those
  offset bars became the motion signature. The background renders a real
  depth-anything.cpp depth map as contour lines and switches to a
  locate-anything.cpp style detection overlay over the engines section.

  Also included: an /engines/ index driven entirely by data/engines.yaml, a
  /blog/ section with five posts written from the release notes and the
  engine benchmark suites, install.sh and a Kubernetes manifest since the
  site advertises both, and a rule in .agents/ that release preparation now
  includes a blog post and demo clips.

Every figure on the site is derived from the repository or the GitHub API,
not from memory. Correcting them against their sources found one error in
README.md: voxtral-tts.c is text to speech, not speech to text.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write] [Agent]

* feat(website): add a star history chart, rewrite the history post in first person

The history post read like a changelog written by a committee. It is now in
Ettore's voice, first person, with the admissions left in.

The numbers paragraph in particular read like a directory listing. It now says
what the figures mean rather than which file they came from.

Adds an interactive star history chart, built from the GitHub stargazers API
rather than embedded from a third party, so the page makes no external request
and cannot break when someone else's service is down. The four releases the
post is organised around are marked on the curve, and the labels stack into
rows because three of them land within two months of each other.

The API stops paginating at 40,000 items, so the curve is measured up to
December 2025 and the segment from there to today's total is drawn dashed,
labelled as an estimate in the caption and in the tooltip. It is a straight
line between two known points, and the chart says so rather than implying it
is data.

Also drops "marketing site" from the README heading and everywhere else it
appeared, and calls it the main site instead.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-31 09:00:56 +02:00
mudler's LocalAI [bot]
b13c429b3b chore(model-gallery): ⬆️ update checksum (#11239)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-30 23:53:35 +02:00
mudler's LocalAI [bot]
82487e68f3 fix(grammars): restore backslash escaping in llama31 grammar fixture (#11242)
PR #11041 rewrote the testllama31inputResult1 fixture and un-escaped the
backslashes inside its Go raw-string literal, turning `[^"\\]` into
`[^"\]` and `["\\/bfnrt]` into `["\/bfnrt]`. The fixture is compared
line-by-line against the grammar built from PRIMITIVE_RULES in
bnf_rules.go, which is unchanged and still emits the doubled form, so
"generates a valid grammar from JSON schema" fails on every platform.

Restore the four fixture lines to their pre-#11041 form. The cyclic $ref
and depth specs added by that PR are untouched.

The regression reached master because only the DCO check ever reported
on #11041; its test runs were cancelled during the CI purge.


Assisted-by: Claude Code:claude-opus-5[1m] [Bash] [Edit]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-30 23:53:23 +02:00
localai-org-maint-bot
5c2099f031 gallery: add POCKET GGUF model families (#11231)
Add the Qwen3.5-MoE POCKET-35B and Gemma 4 POCKET-26B releases with Q4, Q2, and compact IQ1 variants where available. Verify every artifact hash against its Hugging Face linked etag.

Assisted-by: Codex:gpt-5 [web]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-30 23:17:50 +02:00
mudler's LocalAI [bot]
6394909557 refactor(ui): move inline styles onto the design system, add an inline-style gate (#11238)
* refactor(ui): add layout/text primitives and per-page CSS blocks

The React UI already shipped a design system (tokens, form grids, data
tables, stat cards, callouts) that the pages largely bypassed: ~2,000
`style={{ ... }}` literals across src/pages and src/components. Each one is
a spacing or colour decision made locally, so no two pages share a rhythm,
which is the main reason the app reads as unfinished rather than as one
product.

Two additions, both to App.css:

  - A small semantic primitive layer: .stack / .hstack for vertical and
    horizontal rhythm, .text-note / .text-sub / .text-meta / .text-mono for
    the text roles the pages kept re-deriving, .tone-* + .icon-chip for
    semantically tinted icons, plus scale-locked spacing and size steps.
    Deliberately short and semantic, not a utility framework: the size and
    spacing classes exist mainly so that an OFF-scale value stays an inline
    style and therefore stays visible.

  - Named blocks for the shapes fifteen pages actually have (.p2p-diagram,
    .usage-tile, .tr-code, .mw-badge, .set-rail, ...), so those shapes are
    defined once instead of per call site.

Two findings worth recording. The type scale is xs 0.6875 / sm 0.8125 /
base 0.875, and 0.75rem was in use roughly 100 times without being on it
(along with 0.7, 0.85, 1.1 and 0.625rem); all now snap to the nearest step.
And there were thirteen distinct table column widths across the app where
three or four would do; they are pulled into .col-w-* so the ladder is
visible in one place, ready to normalise separately.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]

* refactor(ui): move repeated inline styles onto shared classes

Six passes over src/pages and src/components, each matching a whole
`style={{ ... }}` attribute exactly so the swap is provably equivalent:

  - the identical "waiting for the first response" wrapper, repeated
    verbatim in 21 files
  - text roles (size + colour combinations) onto .text-note / .text-meta /
    .text-sub / .text-mono
  - semantic colours, type-scale steps, .panel-title, .list-row
  - spacing and weight steps, .stack / .hstack rows
  - table column widths, small pills, chart legend swatches
  - the remaining shapes appearing three or more times

One class of bug is worth calling out, because it is what a careless
style-to-class conversion produces and it is invisible to every check we
run. Adding `className="x"` to an element that already had a className
leaves TWO className attributes; JSX keeps the last and silently drops the
first, so `<i className={icon} className="text-xs" />` loses its icon while
passing eslint, `vite build` and the full Playwright suite. 112 of these
were introduced and repaired here. The gate added in a later commit fails
on them.

No visual change is intended beyond snapping off-scale font sizes onto the
type scale. Verified after every pass: eslint 0 errors, vite build passes,
Playwright page-render-smoke + navigation 22/22.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]

* refactor(ui): convert fifteen pages onto named classes, add inline-style gate

Full per-page conversions, each one reading the page, naming the shapes it
actually has, and leaving inline only what is computed at runtime:

  P2P 205->0   Traces 50->0   ConfigFieldRenderer 23->0   NodeDetail 29->1
  ModelEditor 25->1   NodeInstallPicker 29->1   FineTune 175->2
  Nodes 33->2   ImportModel 35->3   Talk 43->4   AgentJobDetails 27->4
  Usage 97->6   Settings 24->6   Middleware 54->8   Backends 87->45

Every remainder is genuinely dynamic: a data-driven badge colour, a
`width: ${pct}%`, a tooltip's coordinates.

Naming the shapes made reuse fall out on its own. Nodes reuses the P2P
setup shapes (both present the same "no workers yet, here is how to add
one" flow) and Model Editor reuses the Settings section rail, which was
byte-identical. Two shared *style objects* also turned out to be classes
wearing a costume and were deleted: `monoCell` in Usage and `hintStyle` in
ImportModel.

Some fixes fell out of the conversion. The evaluation toggle on the
fine-tuning page was a hand-rolled div that rendered as a clipped circle;
it is the existing Toggle component now. That page's empty state pointed at
a "New Job" button that was scrolled off the top of the page, and now
carries its own call to action. And `.input--file` is added at the system
level rather than as a local hack, because ImageGen and VideoGen truncate
their file inputs the same way today.

scripts/inline-style-gate.mjs is the ratchet that keeps this from
regressing. It does not forbid inline styles; it fails when the total goes
UP (same discipline as the coverage baseline) and when an element carries
two className attributes. eslint would catch the latter via
react/jsx-props-no-duplicate-props, but that needs eslint-plugin-react,
which this project does not depend on, so the check lives next to the tool
that causes the problem.

  npm run lint:inline-styles          # check against baseline
  npm run lint:inline-styles:report   # per-file counts, worst first
  npm run lint:inline-styles:write    # refresh after converting

Net: 2,061 -> 611 inline styles. eslint 0 errors, vite build passes,
Playwright page-render-smoke + navigation 22/22, gate green on both checks.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Bash] [Edit] [Write]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-30 23:17:06 +02:00
mudler's LocalAI [bot]
704b87dc8e chore: ⬆️ Update leejet/stable-diffusion.cpp to e92e86fb11b3028ac9edaf63d93709801d106b12 (#11206)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-30 16:55:42 +02:00
Leoy
632c4b6db2 refactor(backends): extract package-system-libs.sh from 31 package.sh (#11095)
refactor(backends): extract shared package-system-libs.sh from package.sh

The arch-detect-and-copy-system-libs block (Darwin rpath / x86_64 / aarch64
loader + libc/libstdc++/libgcc_s/libm/libgomp/libdl/librt/libpthread) was
inlined verbatim in 31 backend package.sh scripts. Extract it into a single
sourced scripts/build/package-system-libs.sh, the CPU-side counterpart to
scripts/build/package-gpu-libs.sh and its sourcing contract.

Consolidating the copies fixes three drift classes that had crept in:
  - libgcc_s.so.1 and libstdc++.so.6 were listed twice in 9 backends
    (acestep-cpp, crispasr, moss-tts-cpp, omnivoice-cpp, piper,
    qwen3-tts-cpp, silero-vad, stablediffusion-ggml, whisper); the shared
    script copies each once.
  - libgomp.so.1 was omitted from opus. OpenMP consumers dlopen it rather
    than link it, so the missing copy only failed at runtime; the shared
    script always includes it.
  - the Darwin @loader_path/lib rpath was applied only in piper and
    silero-vad; both now pass their packaged binary to the shared script,
    preserving that behavior. Every other backend passes an empty binary
    path so no rpath is added, preserving its current behavior.

Each backend's pre/post packaging steps (binary copy, run.sh, ldd closure
walks, ggml variant bundling, espeak/OpenBLAS extras, the ds4 validate step)
are preserved verbatim; only the inline if/elif/else arch block is replaced
by a single source line.

Signed-off-by: supermario_leo <leo.stack@outlook.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-30 16:30:49 +02:00
localai-org-maint-bot
68a0460681 fix(oci): install backends on filesystems without symlinks (#11166)
* fix(backends): fall back to copying links when the filesystem rejects symlinks (#10890)

Backend installation extracts the OCI image tar via containerd's
archive.Apply, which calls os.Symlink directly. On filesystems that do
not support symlinks (notably CIFS/SMB mounts, commonly used to back the
/backends volume) the syscall fails with "operation not supported" and
the whole install aborts, leaving an empty backend directory. The CUDA
llama.cpp image trips this on the libcublas.so -> libcublas.so.12.x
symlink.

When archive.Apply fails with a link-unsupported error, reset the
staging directory and re-extract with a pure-Go walker that still
attempts real symlinks/hardlinks first and degrades to copying the link
target's contents in place when the filesystem rejects them.
mutate.Extract already flattened the layers, so the tar carries no
whiteouts to interpret. Link copies are deferred to a second pass so
forward references resolve.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:opus-4.8 [Claude Code]

* fix(oci): check deferred Close in copyFilePreservingMode (errcheck)

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:opus-4.8 [Claude Code]

* fix(oci): reject path-traversal tar entries in the link-copy fallback

safeJoin sanitized "../.." entries by clamping them under root instead of
rejecting them, so a malicious entry was silently redirected rather than
refused. Join without the leading-slash trick and reject any entry whose
cleaned path resolves outside root; absolute link targets are still mapped
under root (image-root relative) rather than escaping.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:opus-4.8 [Claude Code]

* fix(oci): silence gosec on the validated link-copy file ops

Use hdr.FileInfo().Mode() instead of converting the int64 tar mode to
os.FileMode (removes two G115 overflow findings), and annotate the tar
extraction file operations with justified #nosec comments: every path is
validated by safeJoin against the extraction root before use (G304/G305).

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:opus-4.8 [Claude Code]

* fix(oci): build extraction image from downloaded layers

Avoid appending downloaded layers to the original remote-backed image, which duplicates the layer stack and reopens the source during extraction. Building from an empty image preserves the flattened whiteout semantics while keeping extraction local.

Assisted-by: Codex:gpt-5

* fix(oci): materialize chained links in dependency order

Retry deferred link copies until their targets exist so soname chains work on filesystems without symlink support. Document that copied links can increase backend storage usage on CIFS and SMB mounts.

Assisted-by: Codex:gpt-5

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-30 16:23:42 +02:00
Adira
ef724a3c9d feat(api): add /v1/detokenize endpoint (#9620)
* feat(api): add /v1/detokenize endpoint

Closes #1649.

Mirror of the existing /v1/tokenize path, requested by @benniekiss in
the issue thread for "complete API workflow" use cases that need to
turn token IDs back into text without local processing.

- Add Detokenize gRPC RPC with DetokenizeRequest{tokens} /
  DetokenizeResponse{content} messages.
- Implement in the llama.cpp backend using common_token_to_piece, the
  same primitive TokenizeString already uses internally.
- Other backends inherit the default Unimplemented from base.Base, in
  line with how Detect, Rerank, etc. are gated per-backend.
- Wire up the Go gRPC interface, server, client, and in-process embed
  wrapper alongside their TokenizeString counterparts.
- Add the schema types, ModelDetokenize wrapper, HTTP handler, route
  registration, RouteFeatureRegistry entry (gated by FeatureTokenize so
  no new feature flag is needed), and the discovery map entry under
  ai_functions.
- Regenerated swagger reflects the new endpoint and types.
- Update authentication.md to list /v1/detokenize alongside /v1/tokenize.

Assisted-by: Claude:claude-opus-4-7
Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com>

* test(e2e): add mock backend tests for /v1/detokenize

Add Detokenize to the mock gRPC backend and wire up two e2e tests in
the MockBackend suite: one that posts known token IDs and asserts a
non-empty content response, and a round-trip that tokenizes first then
detokenizes the returned IDs.

Addresses reviewer feedback on #9620.

Assisted-by: Claude:claude-sonnet-4-6
Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com>

* fix(kokoros): implement detokenize in the Rust backend service

The Detokenize RPC added in this PR grows the tonic-generated Backend
trait. Unlike the other languages there is nothing to inherit a default
from — Rust trait impls must list every method — so
backend/rust/kokoros failed to compile:

  error[E0046]: not all trait items implemented, missing: `detokenize`
    --> src/service.rs:72:1
  72 | impl Backend for KokorosService {

Go backends pick up the Unimplemented default from base.Base, and the
generated C++/Python servicer bases default to UNIMPLEMENTED, which is
why the Rust backend was the only one that broke. kokoros is the sole
Rust crate in the tree, so this is the full extent of the fallout.

Return Status::unimplemented("Not supported"), matching how this same
file already gates tokenize_string and ~20 other unsupported RPCs.

Fixes the tests-kokoros and backend-jobs-singlearch-4 (-cpu-kokoros)
failures on the previous head.

Assisted-by: Claude:claude-opus-5 cargo
Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com>

---------

Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-30 16:01:47 +02:00
mudler's LocalAI [bot]
965180581b chore: ⬆️ Update ggml-org/whisper.cpp to a630b35c6fc02c8879f751ec3f39a61327f01dc7 (#11205)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-30 15:58:24 +02:00
mudler's LocalAI [bot]
a740a25934 fix(ci): skip the master image rebuild for commits no image can see (#11223)
On 2026-07-30, 12 of the 23 queued runs of this workflow were commits like "add
1 new model to gallery" or a docs fix, each rebuilding all 18 container images.
That was roughly 216 queued jobs producing byte-identical output, in a queue
holding 1071 jobs with an oldest entry two days old.

Verified against the shipped Dockerfile before assuming it: the final stage
copies only entrypoint.sh, healthcheck.sh and the local-ai binary, there is no
go:embed of gallery/ or docs/, and the gallery is fetched at runtime from
github:mudler/LocalAI/gallery/index.yaml@master. A gallery-only commit produces
an identical image, and the gallery change reaches users through GitHub whether
or not an image is rebuilt, so nothing is delayed by skipping.

Add a `changes` job that decides once whether the push can affect an image; the
other 11 jobs take `needs: changes` and an `if:` on its output.

A job gate rather than paths-ignore on the trigger, for two reasons that both
fail silently if got wrong:

  - paths-ignore on `push` also applies to tag pushes, and a tag created on an
    existing commit carries an empty commits list. That would skip the release
    image build with no failure anywhere. The gate short-circuits to build for
    refs/tags/*, and for a base commit that is missing, zero or unresolvable --
    the same run-everything posture the backend matrix filter takes for a
    truncated diff.
  - the merge jobs use `if: ${{ !cancelled() && ... }}`, and !cancelled() is
    true when a dependency is skipped, so they need the gate named explicitly
    or they would try to merge manifest lists for images never built.

Checked the decision logic against real commits from the queue: the two
gallery/docs commits resolve to build=false, the two code commits to build=true,
and all three fallback paths (tag, zero base, unresolvable base) to build=true.


Assisted-by: Claude:opus-5 [claude-code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-30 15:58:05 +02:00
localai-org-maint-bot
df7c946be6 docs: clarify persistent container storage (#11190)
Document all stateful container paths, explain upgrade behavior and UnRAID mappings, and correct the obsolete troubleshooting mount target.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-30 15:55:30 +02:00
Owen Adirah
1659365059 docs: add reverse proxy timeout guidance (#11195)
docs: clarify reverse proxy bulk job guidance

Mention ingress controllers as another place to configure equivalent upstream response timeouts, and include an example private LocalAI URL for trusted bulk jobs.

Assisted-by: Hephaestus:openai/gpt-5.5 [opencode]

Signed-off-by: Owen Adirah <owenadira@gmail.com>
2026-07-30 15:54:25 +02:00
localai-org-maint-bot
5e541894df fix(llama-cpp): preserve GPU layers during option passthrough (#11193)
Stage the negative GPU-layer sentinels expected by the upstream argument parser, then restore LocalAI resolved values unless a passthrough flag explicitly overrides them. This avoids the parser assertion that terminated the backend for any generic option.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-30 15:54:05 +02:00
mudler's LocalAI [bot]
79e88581c8 chore: ⬆️ Update localai-org/trellis2cpp to 2f3e6e26edbbaaf8ce93d092f16f46968a366a6a (#11208)
⬆️ Update localai-org/trellis2cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-30 15:52:58 +02:00
mudler's LocalAI [bot]
052ccb5e00 chore: ⬆️ Update mudler/parakeet.cpp to 1bfbebfaaf493866f49597cd3b7901959d395c60 (#11209)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-30 15:29:22 +02:00
mudler's LocalAI [bot]
a1620579c7 fix(ci): skip the core image and release build on backend-only diffs (#11224)
Version-pin bumps dominate PR volume: 48 update/* PRs in the week to
2026-07-30, from 16 pins, each a two-line diff. bump_deps.yaml runs a 28-entry
matrix daily and opens one PR per moved pin; CRISPASR and the gallery checksum
produced one every day, ik-llama-cpp six in seven days. Almost all of them edit
nothing but a single backend/*/<name>/Makefile.

Neither image-pr.yml nor build-test.yaml can observe such a change. `make build`
is `go build ./cmd/local-ai`, GoReleaser builds that plus ./cmd/launcher, and
the core image's final stage ships only entrypoint.sh, healthcheck.sh and the
binary. The per-backend trees are copied into the builder but nothing in them
reaches the output. That is 7 + 3 jobs per bump PR that cannot fail for a reason
the diff caused, roughly 410 jobs a week.

Add backend/{cpp,go,python}/** to the paths-ignore of those two workflows. The
inputs that do reach the binary are deliberately outside those prefixes and so
still trigger a full run: backend/backend.proto (protogen-go), go.mod/go.sum
(the go mod tidy before-hook), and backend/Dockerfile.* .

Not applied to the workflows that genuinely read that tree:

  test.yml      TEST_PATHS names ./backend/go/{cloud-proxy,local-store,
                valkey-store}/...
  lint.yml      .golangci.yml carries backend/-scoped rules
  tests-e2e.yml the e2e suite drives real backends over gRPC
  backend_pr.yml its whole job is rebuilding the changed backend

Simulated against the change shapes that occur in this repo. Pin bumps and
python requirement bumps skip; backend.proto, go.mod, a core Go edit, a
backend/Dockerfile edit and any mixed diff all still run, since paths-ignore
skips only when every changed file matches.


Assisted-by: Claude:opus-5 [claude-code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-30 15:13:53 +02:00
mudler's LocalAI [bot]
aa5a17d452 chore: ⬆️ Update CrispStrobe/CrispASR to 4e863bae52aa76a875e4aca57db54ae6d4145c5c (#11207)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-30 15:01:56 +02:00
mudler's LocalAI [bot]
d0809121e7 chore(model gallery): 🤖 add 1 new models via gallery agent (#11213)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-30 14:40:11 +02:00
mudler's LocalAI [bot]
13cab3706c docs(ci): correct the claim that BuildKit exports ccache mounts to the registry (#11220)
.agents/ci-caching.md stated that `cache-to: type=registry,mode=max` "exports
the cache mount data into the registry cache, so subsequent builds restore it".
BuildKit does not do that. A `--mount=type=cache` lives in the builder's local
state and is not part of a registry cache export, and every CI job gets a fresh
runner with a fresh builder, so /root/.ccache starts empty on every build.

The compile script already prints `ccache -s` after a `ccache -z`, so the
evidence was sitting in the logs:

  job 89766266951  llama-cpp cublas-13  6369s  0 / 889 hits (and 0 / 1778)
  job 89766267281  llama-cpp hipblas    8160s  0 / 537 hits
  job 90210828110  llama-cpp cublas-12  5673s  0 / 813 hits

The first two are the control. Their commit, 90355cd44, changed exactly one
file: backend/go/magpie-tts-cpp/Makefile, nowhere near llama.cpp. The engine
source was byte-identical to the previous build, which is the case this section
claims ccache serves, and the hit rate was still 0.00%. A restored-but-stale
cache would show partial hits; 0-of-N is an empty cache.

So llama-cpp, ik-llama-cpp, turboquant, bonsai, ds4 and privacy-filter pay the
ccache wrapper overhead and get nothing back, and multi-hour C++ builds
recompile identical translation units every time.

Record this rather than silently extending it. Wiring the same mount into
Dockerfile.golang, which covers 215 of the 434 matrix entries, measured 18%
faster locally on a rebuild after a source edit with a 71.5% hit rate, but only
because that test reused a single builder across both builds. In CI it would be
a no-op. The note spells out what would actually work (ccache remote_storage or
sccache with a real backend, or round-tripping the cache dir through
actions/cache) and what each costs.

Also correct the composite-actions list, which still named test.yml as a
free-disk-space consumer after #11219 removed it.


Assisted-by: Claude:opus-5 [claude-code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-30 14:39:58 +02:00
mudler's LocalAI [bot]
cbc916aec2 fix(ci): build the native engine in a layer the registry cache can restore (#11221)
Dockerfile.golang builds 215 of the 434 matrix entries: every ggml/C++ engine
wrapped in Go. Each of those Makefiles clones an upstream repo at a pinned SHA
and compiles it once per SIMD variant (depth-anything-cpp builds four: avx,
avx2, avx512, fallback), and those variant targets depend only on the clone.
They cannot observe a change anywhere else in the LocalAI tree.

The compile sat below `COPY . /LocalAI`, so any edit anywhere invalidated it and
recompiled C++ that had not changed. Move it above that COPY, behind a copy of
only the backend's own directory.

This lands the compile in the part of the image the registry cache already
restores. Measured on two real CI builds of this Dockerfile (jobs 90551029008
and 90551028904, both fresh runners): 13 of 17 layers CACHED from
quay.io/go-skynet/ci-cache. The uncached tail is exactly `COPY . /LocalAI`, the
git-config RUN and the build RUN. Putting the engine above the COPY moves it
from the uncached tail into the cached region.

Local measurement, depth-anything-cpp CPU, rebuild after editing a Go file
outside the backend:

  master          78s, 216 C++ objects compiled
  this change     25s,   0 C++ objects compiled   (67% faster)

Note what this deliberately is not. An earlier attempt wired a
--mount=type=cache ccache into the same RUN. BuildKit does not export cache
mounts to a registry cache, so that measured well locally and is a no-op in CI
(see the ccache section of .agents/ci-caching.md). This change relies only on
ordinary layer caching, which the 13-of-17 figure above shows already works
here.

The layer copies the backend's whole directory rather than just the Makefile:
the CMake targets also need CMakeLists.txt and the file list differs per
backend. The cost is that editing a backend's own Go sources invalidates its
engine layer. The expensive cases are unaffected, since a shared-build-input or
backend.proto change, the weekly full-matrix cron and a tag push all rebuild
every backend while touching none of their directories.

Scoped to one backend for now: only depth-anything-cpp gains the `engine`
target. The other 27 fall through the `make -n engine` guard and build exactly
as before, verified against local-store and silero-vad. Rolling the target out
to the remaining 12 backends that define VARIANT_TARGETS is mechanical once this
is confirmed against the registry cache on master.

One caveat on merge: inserting layers shifts the cache keys, so the first build
of each entry after this lands is a full miss. It pays for itself on the second.


Assisted-by: Claude:opus-5 [claude-code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-30 14:39:13 +02:00
localai-org-maint-bot
1189c6825b feat(tracing): persist bounded trace histories (#11203)
* feat(tracing): persist bounded trace histories

Retain API and backend traces below the data path, restore them at initialization, and serialize clears with asynchronous consumers.

Assisted-by: Codex:gpt-5

* fix(tracing): satisfy persistence security checks

Document why persisted filenames cannot escape the trace directory and explicitly ignore the best-effort temporary-file cleanup result.

Assisted-by: Codex:gpt-5 [golangci-lint]

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-30 12:20:01 +02:00
mudler's LocalAI [bot]
cb417464e5 chore(model gallery): 🤖 add 1 new models via gallery agent (#11194)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-30 12:18:11 +02:00
mudler's LocalAI [bot]
c4617265b9 fix(ci): trim two pieces of per-PR work that buy nothing (#11219)
Measured over the week to 2026-07-30, 97% of CI wall-clock is queueing and 3%
is execution: a median 5-hour queue against a 4-20 minute median job. With the
queue saturated, throughput is concurrency divided by service time, so cutting
execution time raises the drain rate directly. Two steps stood out as paying
nothing for what they cost.

test.yml: drop the free-disk-space step (~3.1min per run, ~22 h/week). That
action exists to make room for docker buildx layers and this job runs no buildx
step. It was also sized for a `make test` that downloaded multi-GB GGUF/whisper
fixtures and built llama-cpp/whisper/stablediffusion-ggml; the test-suite reorg
moved all of that into tests/e2e-backends and tests/e2e-aio, as the Makefile
test target already records. Its tool-cache:true wipe was additionally deleting
/opt/hostedtoolcache, forcing setup-go and setup-node to re-download toolchains
that ship preinstalled on the runner.

build-test.yaml: build only the host target on pull_request. The three-platform
cross-compile (linux/amd64, linux/arm64, darwin/arm64) is the bulk of that job's
~6.6min median, ~47 h/week, and nothing consumes a PR's binaries. goreleaser's
--single-target still runs every before-hook (protogen-go, react-ui, go mod
tidy), so the "is the release build broken" signal is unchanged. master pushes
and tags keep building all three.

Also record why the Linux Go workflows pass cache: false to actions/setup-go,
since it reads as an oversight and is not. Set up Go has a median of 11 seconds
on those runners, so there is nothing to win, and the repo already sits at
GitHub's 10 GB Actions cache ceiling with 31 entries, where each setup-go entry
is 222-375 MB on Linux and up to 1.4 GB on macOS. Re-enabling it would evict
something that is earning its space.


Assisted-by: Claude:opus-5 [claude-code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-30 12:17:14 +02:00
mudler's LocalAI [bot]
47097041ff fix(vllm): apply Options[] engine flags before engine init (#11147)
fix(vllm): apply Options[] engine flags before engine init (#11130)

CLI-style flags in a model's `options:` array (`--quantization:gptq_marlin`,
`--enable-prefix-caching`, `--kv-cache-dtype:fp8_e5m2`) were discarded: the
backend only ever read `tool_parser`/`reasoning_parser` out of Options[], and
did so *after* `AsyncLLMEngine.from_engine_args()`, where nothing it set could
still reach the engine.

Map `--` prefixed options onto the AsyncEngineArgs dataclass before the engine
is constructed. Names are normalized the way vLLM's CLI spells them
(`--enable-prefix-caching` -> `enable_prefix_caching`), values are coerced to
the target field's type (bare flag -> True for booleans), and unknown or
uncoercible flags warn and are skipped instead of failing the load, since
Options[] is a bag shared with backend-level settings. Field types come from
the annotation's base so `Literal["auto", "float16"]` (vLLM's dtype) is not
mistaken for a float.

Precedence is typed proto fields -> `options:` -> `engine_args:`. The
production engine_args defaults seeded in hooks_vllm.go therefore skip any key
the user already set as an option, otherwise the later engine_args pass would
silently override it. Parser lookups now accept both spellings, so
`--reasoning-parser:qwen3` selects LocalAI's parser as well.

The helper's tests are stdlib-only and run in the lint workflow's
dependency-light job via `make test-python-helpers`.


Assisted-by: Claude:claude-opus-5 golangci-lint

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-30 12:13:00 +02:00
mudler's LocalAI [bot]
9c85cacfe3 feat(audio-cpp): add the audio.cpp native backend (#11141)
* backend(audio-cpp): add the native build scaffold

Links 0xShug0/audio.cpp engine_runtime through its public framework headers
and serves Health/Status. Model loading and the audio RPCs follow.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): keep the build-tree rpath at $ORIGIN

Upstream sets CMAKE_BUILD_WITH_INSTALL_RPATH in its own directory scope, so
CMake was appending its build-tree library dir to our target and baking an
absolute build-host path into the shipped binary. Set BUILD_WITH_INSTALL_RPATH
on the target so a package that forgets to bundle libggml*.so fails on the
build machine too, instead of only on a user's box.

Also document why EXCLUDE_FROM_ALL must stay on the add_subdirectory call,
correct the claim that Ubuntu ships no gRPC CMake config, stop the pin comment
from repeating the assignment token that bump_deps.sh rewrites, and make
test-engine fail rather than pass when no test is registered.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): parse namespaced model options

Splits option entries on the first colon so path values survive, and routes
load./session. prefixes to the upstream load and session option maps.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): reject out-of-range numeric model options

std::atoi is undefined once the digits exceed long and in practice wraps, so
device:2147483648 was accepted and handed the ggml backend selector a device
index of -2147483648 from a function whose error text promises a non-negative
integer. Parse with strtol and reject on ERANGE, on a value above INT_MAX, and
on any unconsumed trailing input. The error strings are unchanged.

Name the whole entry in the unknown-key error too: an entry like ':value' has
an empty key and left the user nothing to grep for in their YAML.

Tests look keys up through a helper instead of map::at, so a prefix off-by-one
fails one named check rather than aborting the binary and skipping the rest of
the suite, and cover the overflow, negative and non-numeric paths.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): route LocalAI RPCs onto audio.cpp tasks

Task-major resolution over the family's advertised capability set, with the
voice-reference and instructions signals selecting cloning and voice design,
and a streaming-to-offline fallback for server-streaming transcription only.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): use upstream's 'spk' task name and pin the preference order

The SpeakerRecognition short name was 'spkrec', which audio.cpp neither prints
nor parses; a name copied out of audio.cpp was rejected and a pinned 'spkrec'
would not survive the engine boundary. Emit 'spk', keep 'spkrec' as an
input-only alias, and correct the known-tasks lists.

Three assertions were vacuous because their fixtures advertised a single task,
so reversing a preference order or dropping the RPC name and the attempted
pairs from the capability error all passed. Give them fixtures that can tell
the orderings apart.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): convert sample, time and PCM units

Integer nanosecond conversion so 44.1 kHz stays exact, float seconds for the
VAD and diarization messages, and saturating s16le encode so an overshooting
sample cannot wrap to the opposite sign.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): harden seconds_to_samples against NaN and overflow

seconds_to_samples is the one entry point fed by untrusted-shaped input: a
float-seconds timestamp off the wire, or a boundary from a model that diverged.
Its guard covered only the low side, so NaN and out-of-range values fell through
to an undefined double-to-int64 cast and came back as INT64_MIN. A hugely
negative sample index used later as an offset or a length is a wild pointer
rather than merely a wrong timestamp. Reject NaN with the !(x > 0) form and
saturate before the cast.

Also round instead of truncating there. These functions exist to cross the float
seconds boundary the VAD and diarize messages use, and truncation lost a sample
about half the time on the samples-to-seconds-and-back round trip, starting at
n=1.

Pin the decode scale at INT16_MIN, pin nanosecond truncation on a nonzero
fraction, and record why the clamp argument order in f32_to_s16le is
load-bearing for NaN.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): map NaN PCM samples to silence explicitly

f32_to_s16le relied on std::min argument order to keep a NaN sample away from
std::lround, whose result is unspecified for NaN. That was too subtle to rest on
a comment, and the comment was itself wrong: it warned against a spelling that
the outer std::max already catches, while three real spellings leak, including
std::clamp, which is the idiomatic C++17 way to write the same clamp and so the
likeliest future edit.

Divert NaN before the clamp and encode it as 0. A NaN sample rendered as a
full-scale click is worse audio than a dropped one, and this unit converts audio
that may have originated off the wire.

Pin it with an exact-value check rather than a range check, since all three
outcomes the plausible spellings produce are finite and inside full scale, plus
an invalid-operation check that fails unless the NaN is diverted before any
ordered comparison. That second check is what catches modernizing the clamp and
dropping the guard together.

Also bound the seconds round-trip comment, which claimed unconditionally what
holds only below roughly 2^23 samples, and document NaN, saturation and that
bound in the header.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): assemble transcripts from runtime spans

The top-level transcript text is TaskResult.text_output verbatim. audio.cpp
carries text nowhere else: speech_segments, speaker_turns and word_timestamps
hold spans and labels only, so deriving the text from them empties the
transcript for any producer that omits word timing, VibeVoice diarized ASR
included. Fixtures cover every observed producer shape.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): keep a nested speaker turn's own label

A segment sourced from speaker_turns re-derived its speaker by greatest
overlap. A turn's overlap with its own span is the largest possible, so a turn
nested inside another speaker's turn could only tie with the container, and the
tie went to whichever came first. sortformer_diar binarizes each speaker
independently and sorts by start sample, so the container always comes first
and the interjecting speaker was silently erased from DiarizeSegment.speaker.
choose_segment_spans now carries the label out with the span.

Also pins the nearest-segment fallback against measuring from either endpoint
or from segment position, which a trailing-only stray word could not do, and
exercises the empty-word guard in join_words. Two fixtures that pin a rule but
do not mirror any pinned family are relabelled defensive.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): serialize runs with a wedge-aware guard

audio.cpp sessions are not reentrant and a wedged CUDA call cannot be
cancelled, so a plain mutex would pile every worker thread behind a stuck GPU.
Callers waiting past the configured bound, or arriving while the holder has
already overrun it, fail fast instead.

A caller that queues behind a healthy run deliberately does not stamp the
clock: only the thread that takes the lock does. Stamping on arrival would
restart the wedge clock on every request and hide a stuck run from everyone
behind it, which is the pile-up this guard exists to prevent.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): serialize inference through an InferenceLane

One audio.cpp model is loaded per backend process and its sessions are not
reentrant, so concurrent gRPC handlers have to take turns. Serialization alone
is not enough: a wedged GPU call cannot be cancelled from userspace, so an
unbounded queue behind one stuck run would swallow every gRPC worker thread
until the process is useless.

InferenceLane gives handlers a lane with room for one runner. LaneEntry occupies
it for a scope and gives it back on every exit, including an exception, and is
the only way to take the lane at all: occupy/vacate are private with LaneEntry
as the sole friend, so a caller cannot acquire without holding something that
releases. LaneEntry is immovable on purpose, because a moved-from entry would
have to stop releasing while the lane still recorded it as occupied.

A caller either waits indefinitely or brings a millisecond budget. A bounded
caller that cannot get in fails instead of waiting on, and a bounded caller
whose budget is already shorter than the age of the run in the lane fails
immediately, which is what stops a queue forming behind a wedged run. The two
failures carry different text: one names the wait it exhausted, the other states
the measured age of the run without claiming to know why it is long, since a
short budget meeting a legitimately long run lands there too.

The run's age is stamped only after acquisition. A waiter that published itself
as holder would restart the measurement and hide a genuinely stuck holder from
every caller behind it.

Budget negotiation and the overrun decision are pure functions taking their
inputs explicitly, so both are covered without threads or sleeping. The
per-model ceiling arrives as an int of milliseconds; a request may tighten it
and may never loosen it.

Replaces the previous run_guard unit, which was a derivative of an
Apache-2.0 file upstream and could not stay in an MIT tree. Written from a
behaviour contract with no reference to the removed code.

Tests: 65 checks, standard library only, single translation unit, clean under
-Wall -Wextra. Mutation tested at 23/23 killed; two of those mutants exposed
missing coverage and the tests were extended until they died. ThreadSanitizer
clean.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): make the B10 test able to fail, and document LaneEntry

Review of the previous commit found the B10 test could not fail for the reason
it was named. It aged the in-flight run to about 120 ms and then tried two
budgets, 30 ms and 60 ms, both under that age, so both callers took the
fail-fast path. "The two failure modes do not share one message" was comparing
two fail-fast messages that differ only in the budget they print, and the
timeout path was never reached. The second budget is now 400 ms, well over the
run's age, so that caller queues and times out, and a new check asserts which
path each caller took instead of inferring it from inequality. A mutant that
makes the fail-fast path emit the timeout message previously died only on B4 and
B8 checks; it now also dies on B10.

Comment-only changes elsewhere. LaneEntry now says it is not reentrant and does
not detect reentrancy: a second entry on a thread that already holds the lane
surfaces as LaneUnavailable with a positive budget, but parks silently in
unbounded mode, which matters because a handler may hold one across a whole
stream. The immovability note now names the shapes that work, an optional
emplaced in place or a unique_ptr, rather than saying to hold the entry
indirectly without saying how; all three documented forms were compiled before
being written down, which is how the note came to say that an optional of an
immovable type cannot itself be returned.

The header's explanation of why fail-fast exists is reworded. Two clauses traced
back to a specification written after reading the Apache-2.0 upstream header,
and while that was judged de minimis, this unit was rewritten precisely to carry
no upstream expression at all.

The margin table in the report was also wrong about which wall-clock margins are
load-sensitive: there are four, not one, and the tightest is the B3 arrival
check, which is now flagged at the call site. No margin value changed and none
moved across 65 runs.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): gate model loading on the audio.cpp family

Refuses any GGUF without an audiocpp.model_spec.family key and any non-GGUF
path without an explicit family option, so the model loader's greedy backend
probe cannot bind an unrelated llama.cpp GGUF to this backend (#9287).

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): load models and cache sessions per task

Loads one ILoadedVoiceModel and creates an IVoiceTaskSession lazily per
(task, mode), so the same model serves both the unary and streaming RPCs.
LoadModel derives the family from GGUF metadata or an explicit option and
fails with INVALID_ARGUMENT otherwise, so a failed load is a gRPC error the
backend probe can see.

audiocpp_backend::Task mirrors engine::runtime::VoiceTaskKind positionally,
and drift there is silent: every unit still compiles and every test still
passes while the backend runs a different task. Two mechanisms pin it. The
static_asserts in loaded_model.cpp catch an insertion or a reorder, and
-Werror=switch on that one file turns an appended upstream enumerator into a
build failure rather than a warning in a 600 file log.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): stop aborting the process on SIGTERM

The signal handler called grpc::Server::Shutdown directly. Shutdown takes an
absl::Mutex, which is not async-signal-safe: the handler can interrupt a thread
already holding that mutex, and abseil's deadlock detector responds by aborting.
Every SIGTERM therefore ended in exit 134 and a 'dying due to potential
deadlock' stack rather than a drained shutdown.

The handler now sets a lock-free atomic and returns. Server::Wait moves to a
helper thread so the main thread can poll that flag and call Shutdown itself,
outside any signal context. A condition variable would not have helped, because
notifying one from a handler is not async-signal-safe either.

SIGTERM and SIGINT both exit 0 with no stack trace, where both previously
exited 134.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): correct the status, lifetime and state contracts of LoadedModel

An environment fault during session creation was reported as UNIMPLEMENTED. A
missing libggml-cpu-*.so surfaced to the client as 'family silero_vad advertises
vad/offline but refused to create the session: Failed to initialize CPU
backend', which tells LocalAI the model cannot do this and must never be
retried, and sends an operator hunting a capability bug instead of a packaging
one. A throw from create_task_session is now a plain runtime_error, so it maps
to INTERNAL. Only a null return, where the family genuinely declined, stays a
CapabilityError.

The model.'s task: option was parsed and then dropped: it lived in a local that
died at the end of LoadModel and had no route to RequestShape::pinned_task.
LoadedModel now keeps it and exposes pinned_task().

The global model becomes a shared_ptr reached through snapshot(). An audio RPC
runs for seconds and cannot hold g_model_mu for its duration, so under a
unique_ptr a Free arriving mid-request would destroy the model underneath it.
Handlers now take a counted reference and whichever finishes last does the
teardown, outside the lock.

session_for documents the streaming state contract rather than resetting the
session itself. Resetting on a cache hit was tried first and is not possible:
silero_vad throws 'session prepare() must be called before Silero VAD reset()',
so it would turn an ordinary second fetch into a hard error. start_stream's base
implementation is already a reset, so a caller that runs prepare then
start_stream per stream gets a clean session; a probe against the bundled
silero_vad confirms an identical replay when it does and a carried-over stream
when it does not.

Also: an unknown backend: name is rejected before the model loads rather than
after; MainGPU is parsed instead of passed through std::atoi, which turned
'gpu1' into device 0 silently; and device carries a device_set flag, because 0
is both the default and a real device index, so MainGPU was overriding an
explicit device:0 that the neighbouring threads: handling promises will win.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): serve the VAD and Diarize RPCs

Both emit float seconds, converted from the runtime's sample-index spans, and
both take a counted reference to the loaded model through snapshot() and hold it
for the whole call: a Free arriving mid-request drops only the global's
reference, so whichever request finishes last destroys the model instead of one
of them running on freed weights. An AddressSanitizer build reproduces exactly
that heap-use-after-free inside ggml_vec_dot_f32 when the handler keeps a raw
pointer instead, which is why the shape is what it is.

The inference lane is taken before session_for, not after. session_for reads and
writes an unsynchronised session cache and the offline run calls prepare(),
which mutates the session, so both belong inside the lane.

Diarize routes before it reads the input file, so a family that cannot diarize
at all says so rather than complaining about the audio first. Its per-segment
text stays empty because audio.cpp's SpeakerTurn carries a span and a speaker
label only, and nested or overlapping turns are passed through untouched: a
sortformer turn inside another speaker's turn is correct output for overlapped
speech, and LocalAI is overlap-tolerant downstream. Duration counts frames
rather than floats, so a stereo input does not report twice its length.

Verified end to end against upstream's bundled silero_vad, which needs no
download, using the bundled 16 kHz speech asset: a synthetic tone returns
nothing, correctly, because silero detects speech and a sine is not speech.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): enforce ModelIdentity on VAD and Diarize

audio-cpp was the only C++ backend without the model-identity guard, and no
later task in the plan added it. pkg/grpc/server.go enforces checkModelIdentity
on exactly these two RPCs, for the reason #10952 records: in distributed mode a
worker can recycle a stopped backend's gRPC port for another model's backend,
and the controller's liveness-only probe cannot tell a stale cached route from a
live one. Without this guard a stale route gets a different model's VAD or
diarization answer back with a 200.

The loaded identity lives on LoadedModel rather than in a separate global, which
is where this differs from llama-cpp. A handler holding the model through
snapshot() then necessarily judges against the identity that model was loaded
with, and a concurrent reload cannot swap one without the other. The refusal is
NOT_FOUND carrying the verbatim grpcerrors.ModelMismatchSentinel substring.

session_for and run_offline now take a const LaneEntry & proof-of-holding
parameter. The rule that both must run under the inference lane was prose, which
is exactly how the plan came to specify the inverted order; it is now a compile
error. Restoring the inverted order fails to build rather than racing on an
unsynchronised session map with a mutating prepare().

Diarize's speaker-hint comment claimed the dropped hints were "not a silent
failure". From the caller's side that is what they are, and backend.proto
documents num_speakers as forcing, so the comment now says plainly that the
forwarding is dead for sortformer and that the family which lands must either
honour num_speakers or refuse it. read_audio_file inspects the error_code from
exists(), so an unsearchable parent directory no longer reports as a missing
file. The VAD handler records the stimulus that actually works, since silero
correctly ignores synthetic tones and the next task would otherwise rediscover
that.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): make the lane and identity guards structural

Two hardenings ahead of the eleven handlers still to be written, both of which
get harder to retrofit later.

The lane proof-of-holding parameter was a const reference, which binds to a
temporary, so session_for(rpc, shape, model->acquire(0)) compiled. Each such
temporary dies at the end of its own full-expression, releasing the lane between
two calls that must share one: precisely the split the parameter exists to
prevent, and the form a future author is most likely to reach for because it
reads as tidy. A non-const reference requires an lvalue, so the temporary form
now fails to compile while the named-local handlers build unchanged. The header
comment no longer implies the check is total either: it proves a lane was taken,
not that it is this model's lane.

The identity check was two lines each handler had to remember, with nothing
failing if a new one forgot them and no C++ equivalent of
model_identity_modalities_test.go to notice. snapshot() becomes
snapshot_unchecked(), whose only legitimate caller is Status, since HealthMessage
carries no ModelIdentity. Handlers go through snapshot_for(), which takes the
counted reference, refuses when nothing is loaded, and runs the identity check
before anything can route. Every handler already has to call something to obtain
the model, so the guarded call is now the shortest path and skipping it means
deliberately typing snapshot_unchecked. A convention that has to be remembered
can rot; this cannot.

Verified: the temporary-argument and inverted-order forms each fail to compile
with the expected diagnostic, the real handlers build, and bypassing the guard in
Diarize alone turns the identity test red on that RPC while VAD stays green.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): serve the AudioTranscription RPC

Adds result_map, the engine-to-proto boundary, and wires the offline
transcription RPC.

The handler branches on the ROUTED task: for Asr the request's prompt is
whisper-style decoding context and becomes a request option, for Alignment
the same field IS the transcript to align and becomes the text input.
Routing has already decided which.

The result text is TaskResult.text_output verbatim and is never derived
from the segments. audio.cpp carries transcript text in text_output and
nowhere else, so deriving it returns an empty transcript for every
producer that reports segments without word timing. transcript_assembly
already enforces that; this commit's job is not to undo it at the proto
boundary, and result_map_ctest pins it there.

read_audio_file now takes the sample rate the caller needs. Both file-fed
speech handlers ask for 16 kHz mono, for two reasons: silero_vad and
sortformer_diar refuse anything else outright, which turned an ordinary
44.1 kHz upload into INTERNAL, and nemotron_asr emits word timestamps in
its own 16 kHz feature domain whatever the input was, so only a 16 kHz
buffer makes the emitted nanoseconds right. Zero keeps the file's native
rate and channels, which is what source separation will need.

LoadedModel::check_can_serve answers a capability refusal before the lane
is taken and before the input file is read. Routing is a pure read of the
immutable capabilities, so a model that cannot serve an RPC no longer
waits out somebody else's run to say so. VAD and Diarize use it too.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): stop linking sentencepiece's vendored protobuf

engine_runtime links sentencepiece, whose default SPM_PROTOBUF_PROVIDER
builds the protobuf-lite 3.14.0 sources it vendors. The generated
backend.pb.cc is built against the toolchain's protobuf 3.21.12. Both
ended up in the binary: 476 google::protobuf:: symbols came from the
archive, 278 of them also defined by libprotobuf.so, and the archive won,
because once ld pulls a member in for sentencepiece's own code every
reference binds to the definitions that member carries.

The visible symptom is one function.
ParseContext::ParseMessage(MessageLite*, const char*) is what a generated
_InternalParse calls for a submessage field and for nothing else, so flat
messages parsed and nested ones did not: a TranscriptResult carrying
segments serialized to correct bytes that the same process could not read
back, and TranscriptLiveRequest, a oneof of submessages, could not have
been parsed at all. Underneath that, 3.21 generated code was running 3.14
arena, ArenaStringPtr and ExtensionSet code.

-Wl,--exclude-libs does not fix it. It makes those symbols LOCAL in
.dynsym and the parse still fails, because the binding was decided at
static link time and no visibility flag revisits it.

Setting SPM_PROTOBUF_PROVIDER to "package" before add_subdirectory points
sentencepiece at the protobuf the generated code was already built
against. Zero google::protobuf:: definitions remain in the executable
afterwards, every nested message round trips, and citrinet_asr, which
parses a SentencePiece ModelProto at load time and would break first if
this were wrong, still tokenizes and transcribes correctly.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): fix the segment text a transcription response is built from

Segment text is not decoration. core/http/endpoints/openai/transcription.go
routes response_format text, srt, vtt and lrc through
schema.TranscriptionResponse, which builds the entire body out of
Segments[].Text and never reads the top-level text. So for those four
formats the segment text IS the response.

nemotron_asr emits one word_timestamp per SentencePiece token, and the
word boundary is carried as a LEADING SPACE on the piece ("So", "me",
" call"). join_words inserted a space unconditionally, so
response_format=text returned "So me  call   me  na ture ," while the
correct sentence sat unread in the top-level field. The separator is now
chosen from the words themselves: whole words are space-joined, subword
pieces are concatenated, and one leading space anywhere selects the
latter. Concatenating the real nemotron pieces reproduces text_output
exactly, verified end to end.

This does not touch the top-level text, which is still text_output
verbatim. The rule that forbids deriving the transcript from the segments
is about the direction segments -> text; segment text has no source other
than its words.

Two smaller corrections in the same area:

timestamp_granularities ["word"] set only "word_timestamps", a key no
family in the pinned upstream reads. It now sets "return_timestamps",
which qwen3_asr does read and which both runs its forced aligner and
shortens its chunk window, so asking for word granularity no longer
silently returns nothing.

The request-option comment claimed more than it delivered. prompt,
translate and temperature are read by no ASR family, and are forwarded
only so a family adopting them works unchanged; the comment now says so
per key, and gives TranscriptRequest.diarize the same explicit treatment
threads already had.

Also: the shipping target now carries -Wall -Wextra -Wpedantic, which it
never did, so "the build is clean" starts meaning something; and
fill_transcript_result no longer swallows a null response pointer, since
answering OK with an empty transcript is the one failure mode this unit
exists to prevent.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): serve the AudioTransform RPC

Covers voice conversion, singing voice conversion, speech to speech and source
separation, the four tasks LocalAI's AudioTransform can represent.

AudioTransformResult carries one dst while htdemucs and mel_band_roformer
produce several named stems from a single run, so inference runs ONCE, every
stem is written as a sibling file <dst-stem>.<name>.<ext>, and params[stem]
selects which one dst receives, defaulting to vocals and falling back to the
first output. An unknown stem name is INVALID_ARGUMENT listing the real stem
names rather than a silent substitution, and the selection happens before the
first write so a refused request leaves no files behind. params[stem] is
consumed here and is not forwarded into the engine's request options.

The stem decision lives in stem_selection, which is stdlib only and therefore
tested by backend/cpp/run-unit-tests.sh. It also validates the names, because
they come from the model (htdemucs reads them from the GGUF's config.sources)
and each becomes a component of a path this backend writes: a name carrying a
path separator would escape the caller's output directory, and two stems
sharing a name would silently overwrite one another.

Both files are read at their native rate and channel count. Separation forces
it, since demucs and roformer refuse any rate but 44.1 kHz and lose the stereo
image that separates a centred vocal from a wide mix. The conversion families
all resample internally (seed_vc, vevo2, miocodec, chatterbox were each
checked), so passing the file through unchanged is also strictly better than
band limiting it to 16 kHz first.

Verified end to end against htdemucs f16 on a 44.1 kHz stereo mix: four stems
plus dst, dst byte identical to the selected stem, params[stem] selecting a
different one, an unknown stem refused with no files written, and mono input
preserved as mono output. Also against miocodec for the single output path,
where params[stem] is refused rather than ignored.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): refuse an impossible stem early, and stop blaming the caller for a failed write

Four fixes from the first review of the AudioTransform RPC.

check_can_serve now returns the resolved route, so params[stem] on a route that
is not source separation is refused from the route instead of after a full
inference: 11 ms rather than the 4.5 s a miocodec conversion costs, and far
worse on seed_vc or vevo2. The post-run refusal stays as the backstop for a
separation-routed family that returns no stems anyway. The typo'd-stem-name
case still needs the run, since no framework header publishes the stem names
before one.

Stem names carrying control bytes are refused. GGUF strings are length prefixed
and demucs reads its sources from JSON, so an embedded NUL survives to here:
two names differing only after the NUL are distinct std::strings, so the
duplicate check passes them, and then path::c_str() truncates both and they
open the same file. That is exactly the silent overwrite the duplicate check
exists to prevent, with the .wav lost as well.

A failed write is now INTERNAL rather than INVALID_ARGUMENT. The destination is
LocalAI's own generated-content directory, not anything the caller named, so a
full disk or a permission fault there is a server fault and is worth retrying,
which is the opposite of what INVALID_ARGUMENT tells a client. An empty output
path stays INVALID_ARGUMENT.

Two comment corrections and one clarification: the separators' required rate is
their checkpoint's declared samplerate rather than a hardcoded 44100, seed_vc
resamples with soxr and falls back to sinc-hann, and the "no files left behind"
guarantee covers a refused request, not a write that fails partway through the
loop.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(audio-transform): stop folding every upload to 16 kHz mono, and name the separation stems

Two defects that made source separation unusable through LocalAI's own API,
even though the backend served it correctly over gRPC.

/audio/transform normalized every upload to 16 kHz mono s16 through
utils.AudioToWav, with no way past it. htdemucs and mel_band_roformer refuse
any rate but their checkpoint's own and separate a centred vocal from a wide
mix using the stereo image, so every separation request through the HTTP API
died with "HTDemucs prepare() sample rate mismatch: expected 44100, got 16000"
while the same call over gRPC worked. The fold is not wrong, it is
backend-specific: LocalVQE's echo cancellation genuinely wants 16 kHz mono and
needs the reference in the same shape. So it becomes a declaration,
BackendCapability.AudioTransformInputMono16k, set for localvqe and for nothing
else. A backend that declares nothing gets its upload unchanged, which means no
backend has to opt in to work. utils.AudioToWavPreservingShape is the
non-folding conversion: a 16-bit PCM WAV passes through byte for byte at any
rate and channel count, anything else is transcoded to WAV with its rate and
channel layout kept.

The other defect is that the run-once stem design bought nothing. A separation
backend writes every stem beside dst from one inference, but AudioTransformResult
carried only dst, so the other three were files no caller could find and a
caller wanting all four had to run four separations. AudioTransformResult grows
a repeated AudioTransformStem, the backend fills it, core/backend validates that
each path really is inside the generated-content directory it handed over, and
the endpoint publishes them as an X-Audio-Stems JSON header beside the existing
X-Audio-Input-Url. JSON because a stem name is the model's own string and could
contain any separator a hand-rolled format would use.

Verified end to end through the HTTP endpoint with htdemucs f16 on a 44.1 kHz
stereo file: 200 with a 44.1 kHz stereo body, all four stems named and fetchable
through /generated-audio/, body byte identical to the selected stem, and
params[stem]=drums returning a different one. The same upload sent to a model
whose backend is localvqe still reaches the backend as 16 kHz mono, confirmed
both by the engine's own rate refusal and by the persisted input file.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(audio-transform): reject extensible WAV from the passthrough, escape stem URLs, convert stems with dst

Four fixes from the second review, plus one bug they made visible.

isPCM16Wav tested only the bit depth, and go-audio's IsValidFile never looks at
the format tag, so a 16-bit WAVE_FORMAT_EXTENSIBLE (0xFFFE) upload was passed
through untouched where the old fold would have transcoded it. audio.cpp's WAV
reader accepts 16-bit only when the tag is 1, so such a file died with
"unsupported WAV encoding". Extensible is what many DAWs and Windows tools
write and music files are this endpoint's new headline input, so it is a
first-contact failure rather than a corner. The check now requires tag 1, with a
spec that fails against the old implementation.

Stem URLs are percent-escaped. A stem name is the model's own string and legally
contains a space, a '#', a '?' or a '%'; an unescaped '#' truncates the URL
before the request is even sent. The name field keeps the raw name.

sample_rate and response_format are applied to the stems as well as to dst.
Applying beat documenting: dst IS one of those stems, so leaving them alone
broke the "dst duplicates the selected stem" invariant the whole design rests
on, and both conversions are no-ops when unset. A stem whose conversion fails is
dropped from the header rather than advertised in the wrong shape.

Verifying that turned up why it had never been noticed: the two fields were
never bound at all. The request arrives as multipart/form-data and echo's binder
falls back to the FIELD NAME without a form tag, matching only
case-insensitively, so "SampleRate" never matched "sample_rate" and "Format"
never matched "response_format". Both were documented in the endpoint table and
silently ignored. Two form tags fix it, and with them the conversion is
observable end to end.

Docs: audio-transform.md now documents what LocalAI does to an upload before the
backend sees it, which backend gets the 16 kHz mono fold and why, params[stem],
and the X-Audio-Stems header with a worked example.

Also records the known limitation that the fold lookup is on the bare backend
name, so pinned variants (vulkan-localvqe) do not match, and points at
IsLlamaCppBackend as the suffix-tolerant precedent.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): serve the TTS and SoundGeneration RPCs

TTSRequest.voice is treated as a speaker reference clip when it names an
existing regular file, which makes routing prefer VoiceCloning, and as a named
preset otherwise, in which case it travels as VoiceReference::cached_voice_id.
Both the clip and SoundGenerationRequest.src are read at the file's own rate and
channel count: upstream's own CLI and server do exactly that, every consuming
family resamples internally and mostly with a better resampler than ours, and
ace_step and stable_audio resample their input per channel, so a downmix here
would delete the stereo image they are built to consume.

The request builders live in their own unit rather than in grpc-server.cpp's
anonymous namespace so they can be tested; grpc-server.cpp has a main() and
cannot be linked into a test binary. The option keys are the whole point of
these functions, so each one was grepped against the pinned upstream and the
accounting is written down beside it. instructions maps to "instruct", which is
what upstream's own server maps the OpenAI field to and what qwen3_tts and
omnivoice read, and to "caption" for irodori_tts; the style tag is spelled
"instruct" too, because "instructions" is looked up nowhere. duration maps to
"duration_seconds", read by all three generation families, with the proto's own
name kept only as a forward-tolerant alias. Keys that no family reads say so.

Both handlers answer a capability refusal before taking the lane and before any
file read, so a model that cannot synthesise does not queue behind somebody
else's run to be told no.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): stop emitting an empty style language, and name the missing clip

StyleCondition::language was set whenever has_language() was true, with no
!empty() guard, while the language option twelve lines below had one.
core/backend/tts.go sets Language unconditionally, so has_language() is true on
every request LocalAI sends and carries "" when the caller named none. An
engaged-but-empty style language is worse than an absent one: supertonic reads
text_input->language behind its own !empty() guard and then overrides it from
style->language with no guard at all, so "" replaced its "en" default and its
tokenizer threw "invalid Supertonic language: ". Every /v1/audio/speech request
that set instructions and no language would have been an INTERNAL against a
supertonic model. A plain request never saw it, because the style condition only
exists when instructions are non-empty, which is why the chatterbox end to end
run did not catch it.

TTS also stops discarding the Route that check_can_serve already returns. A
family routed to voice cloning without a reference clip used to be refused from
inside its own prepare(), which meant an INTERNAL naming neither the RPC nor the
field to set; chatterbox advertises clon and no tts, so that was every
preset-only request to it. It is now an INVALID_ARGUMENT naming
TTSRequest.voice, answered in about 4 ms, and it cannot misfire because
has_voice_reference is what selected cloning in the first place. Reading
CapabilitySet::supports_speaker_reference to generalise this stays a follow-up.

The src read carries a written caveat rather than a family blocklist, because
ace_step's editing routes legitimately need src: setting src on a stable_audio
model corrupts the heap and aborts the process in the pinned upstream, and the
only thing keeping that off the network is that
schema.ElevenLabsSoundGenerationRequest has no field for it. Nobody reading that
Go schema would know why, so the reason is recorded where the field is read.

build_tts_shape is extracted so TTSStream cannot describe the same request
differently, and it arrived untested: two mutations of it survived until a
test_tts_shape case was added.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): serve the TTSStream and AudioTranscriptionStream RPCs

TTSStream leads with a streaming WAV header carrying 0xFFFFFFFF sizes, matching
the convention backend/go/vibevoice-cpp established, so an HTTP client can start
playback before the full PCM exists. Its chunks are read from
StreamEvent::named_audio_outputs and not audio_output: supertonic, omnivoice and
voxcpm2 all put their streamed audio there and leave audio_output empty until
the very end, so reading the obvious field yields a stream with no audio in it.
The finish_stream result is the family's own merged whole rather than a tail, so
it is emitted only when nothing was streamed.

Streaming transcription sends incremental deltas and degrades to a single delta
plus the final result on families that offer no streaming ASR, which is the same
message sequence with fewer deltas. The four streaming ASR families disagree on
what partial_text means: nemotron_asr, vibevoice_asr and higgs_audio_stt report
incremental fragments while voxtral_realtime reports the whole hypothesis and
reports it twice, so the reconciliation lives in one tested unit rather than in
the handler. nemotron_asr reports only through the stream event sink, and only
from inside finalize, so the audio driver installs one and clears it again
before returning: the session is cached and a sink left holding the caller's
frame is a use after free waiting for the next stream.

begin_stream is now the only implementation of the streaming state obligation,
prepare then start_stream. Streaming sessions are cached, and what clears the
previous stream is start_stream's reset; a family override that dropped it would
break every call site with no compile error, so there is one call site.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): keep streaming deltas on UTF-8 boundaries, refuse dtypes that abort

TranscriptStreamResponse.delta is a proto3 string, whose wire format requires
valid UTF-8. voxtral_realtime reports its hypothesis as a concatenation of raw
token BYTES (tokenizer_text.cpp:171-183), so the cumulative difference between
two consecutive reports is eventually a lone continuation byte, and the C++
runtime serializes that with only a logged warning while the Go runtime refuses
to unmarshal it: the client loses the remaining deltas AND the final_result.
Measured on a trace of a non-ASCII sentence, 11 of 31 messages failed to
unmarshal and every accented character was lost. TranscriptDeltaTracker now
holds back an incomplete trailing sequence and merges it into the next fragment;
reconcile flushes it, which it always can because the final text is complete.
The same trace now unmarshals in full with zero failures.

A streaming buffer whose float count is not a whole number of frames is refused
rather than truncated. The integer division dropped the tail floats from the fed
audio and therefore from the transcript, with no diagnostic; vibevoice_asr
refuses the same thing from the other side of the call.

A supertonic GGUF whose weights are not f32 is refused at load. It reaches
ggml_concat with mismatched operand types and ggml_abort takes the whole backend
process down on the first request, so nothing downstream can report it: the
model loads, then every request kills the process. Attributed rather than
assumed, the unary TTS path aborts identically, and upstream records that
package as untested. The refusal names the orig package and says what to run
before deleting the guard.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): stop a repeated lead byte from orphaning the next delta

The first UTF-8 fix closed the cumulative half only. Rule 2 discards a fragment
the known text already starts with, and when that fragment is the LEAD BYTE of a
new character it looks exactly like a repeat of an older character beginning
with the same byte. It was discarded rather than held, its continuation bytes
then arrived alone and began the next delta, and utf8_complete_prefix_length
only ever inspected the trailing sequence, so a delta invalid at the FRONT went
out whole. Through a real Go proto.Unmarshal the review's four-character repro
gave 3 deltas, 2 unmarshal failures and a lost transcript.

Reachable from the incremental families, not only from voxtral: nemotron_asr's
decoder cuts at a byte offset and vibevoice_asr's common_prefix_size compares
bytes, so both split characters. Measured over 30,000 randomized incremental
traces, 53.28% of Japanese traces and 9.52% of French ones carried at least one
delta the Go runtime refuses.

Two changes. Rule 2 no longer judges a fragment that ends mid-character, so the
lead byte is held instead of swallowed and the character survives intact; the
cost is a few duplicated bytes in a shrinking cumulative report, which no pinned
family produces. release() additionally drops leading orphan continuation bytes,
so no delta can begin mid-character whatever the rules above it decide. Losing a
byte keeps the stream alive; emitting one ends the RPC and takes the
final_result with it.

Post-fix all 60,000 traces produce zero unmarshal failures, and the cumulative
streams plus both pure-ASCII incremental streams are byte-identical to the
previous commit, so nothing changed for the families already working.

The weight-dtype allow list moves to family_gate, where it is stdlib-only and
pinned by a test rather than only by a comment. Two comment citations corrected.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): read only an exact repeat as a repeat, not any prefix

Rule 2 discarded any partial the known text merely started with. For a cumulative
family that is a duplicate; for an incremental family it is an ordinary short
fragment that happens to coincide with the start of the transcript, and it was
dropped, silently corrupting the text. Pure ASCII, no multi-byte character
anywhere: the fragments "pure ", "ascii ", "trans", "c", "ri", "p", "t" left the
client holding "pure ascii transcrit". Over 5,000 randomized traces per
transcript, 9.50% of pure-ASCII and 29.12% of French traces ended with the client
holding something other than final_result.text, with a 200 and no diagnostic.
Both incremental families emit fragments that small routinely, since nemotron_asr
cuts at a byte offset and vibevoice_asr at a common prefix.

Narrowing rule 2 to an exact repeat drives that to zero on all six transcripts
and changes no cumulative stream at all: 30,000 randomized cumulative traces are
byte-identical to the previous commit.

What rule 2 guarded was established from upstream rather than from its own
comment. The only duplicate any pinned family produces is voxtral_realtime's,
where process_available_stream_chunks feeds each event to the sink from inside
its loop and returns the last of the batch, so that event arrives twice with
byte-equal text. A duplicate is an exact repeat, so equality still covers it. The
case given up is a cumulative report that SHRINKS, which no pinned family can
produce: voxtral decodes a token vector that is only push_back'ed and cleared by
reset(), so within a stream it can only grow.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): serve the AudioTranscriptionLive RPC

The one bidirectional stream this backend serves. The client sends a
TranscriptLiveConfig, then TranscriptLiveAudio frames; the server acknowledges
with ready, emits deltas as the audio arrives, and sends final_result once the
read side closes. There is no offline fallback: live transcription has to
consume audio incrementally, so a family with no streaming ASR is refused
rather than served a batch run, which is what this RPC's Streaming-only
mode_candidates list already says.

The driver is a new sibling of run_streaming_audio, run_streaming_live, because
the audio does not exist yet: instead of slicing a buffer it pulls frames from
the caller until the read side closes. It installs the same ScopedStreamSink in
the same order, which is not optional, since nemotron_asr returns a bare event
from process_audio_chunk and reports every partial through the sink from inside
finalize(). It buffers the wire's frames up to the family's own preferred window
rather than feeding whatever size the client's audio callback produced, and it
does not call finish_stream at all when no audio arrived, because nemotron_asr
throws "finalize requires streamed audio" and an empty transcript is the
truthful answer to transcribing nothing.

Three things the handler had to get right and one it cannot:

  - The audio contract. A live request carries no samples, but nemotron_asr's
    streaming prepare() throws without an audio contract, and
    build_preparation_request derives it from TaskRequest::audio_input, so that
    field is an EMPTY buffer holding only the rate and the channel count.
  - 16 kHz or a refusal. The families express their spans in their own 16 kHz
    feature domain whatever the input was, and live frames cannot be resampled
    on the way in the way a file can, so an 8 kHz session would return
    timestamps 2x off with a 200. core/backend hardcodes 16000 anyway.
  - A mid-stream Config is refused. backend.proto calls it a decoder reset, but
    deltas already on the wire cannot be retracted, so a reset would leave the
    final text contradicting the transcript the client assembled. Ignoring the
    message would hand a client that believes it reset the decoder a transcript
    that silently continues the audio it thought it discarded.
  - The stale-route identity check cannot run here: TranscriptLiveRequest
    carries no ModelIdentity in either arm of its oneof, so snapshot_for does
    not instantiate for it. snapshot_unchecked's comment now names that as a
    second legitimate class of caller and says the fix is a proto change.

eou and eob stay false. They exist for cache-aware models that emit
end-of-utterance and end-of-backchannel tokens; audio.cpp's StreamEvent has no
equivalent signal, and a client uses eou to decide the speaker yielded the turn,
so a guess inferred from silence cuts people off mid-sentence.

The lane is held for the whole stream, which is as long as the user keeps
talking: the streaming session is stateful and cached, so a concurrent run would
interleave two callers' audio and corrupt both transcripts.

Verified against nemotron_asr over a real connection with a 14 s WAV in
512-sample frames: ready first, 59 incremental deltas with no repeated prefix,
concat(deltas) equal to final_result.text, word timestamps in nanoseconds, eou
and eob false. citrinet_asr answers UNIMPLEMENTED naming the family and listing
asr/offline. A config followed by a close returns an empty final_result rather
than hanging, and a first message that is not a config is INVALID_ARGUMENT. Two
concurrent streams both return the complete transcript.

Two cleanups on lines Task 12 touched, folded in. The DtypeAllowList terminator
is now asserted at compile time: the reported out-of-bounds read did not exist,
the single entry does terminate, but the loops have no other bound and any edit
that widened an entry would walk off the end. And the dtype guard now
short-circuits on "is there a table entry" through a new predicate rather than
on the emptiness of the description string, which would have skipped the check
on an entry with an empty allow list, i.e. on precisely the entry that refuses
every dtype.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): bound the lane a live stream can hold

AudioTranscriptionLive holds the model's inference lane for the whole stream,
which is correct (the streaming session is stateful and a concurrent run would
interleave two callers' audio) and newly dangerous. Every other RPC holds the
lane across compute, or across a write to a slow reader, and both of those
terminate on their own. A live stream instead blocks in a client-driven read,
and a peer that goes silent WITHOUT closing the stream never terminates
anything: the lane stays taken and every other request against that model queues
behind a client that stopped speaking.

live_watchdog is a one-shot idle timer that ends the stream when no frame has
arrived inside a window. It is standard library only, so it is unit tested
without an engine. gRPC's synchronous Read has no timeout and cannot be given
one, so the only way to unblock it is ServerContext::TryCancel, which decides
the wire status itself: the client sees CANCELLED rather than the
DEADLINE_EXCEEDED the handler returns, the reason is logged, and the lane coming
back is the point. When it fires the read loop throws rather than reporting
end-of-input, so the driver does not go on to finalize a decode nobody is
waiting for.

It is armed only after the lane is taken and disarmed as soon as the read side
closes, and both ends matter. Arming earlier would cover acquire(), which
legitimately blocks while another live stream runs, so a queued caller would be
cancelled for waiting its turn. Disarming later would cover our own decode,
where a window overrun is not a peer going quiet and cancelling would throw away
the transcript the client is waiting for.

The window is the new live_idle_timeout_ms option, 30 s by default, 0 meaning no
limit. core/http/endpoints/openai/realtime.go drives a 300 ms ticker and feeds
every tick that produced new audio while a turn is open, so 30 s of silence is a
hundred ticks that delivered nothing. It is also longer than any pause a speaker
takes mid-utterance, which is the case that must never be cut off, and
backend.proto lets one stream span many utterances, so a client that pauses
longer between them raises the option rather than discovering it.

Two smaller corrections in the same handler:

  - check_can_serve now runs BEFORE the sample rate check.
    pkg/grpc/grpcerrors/errors.go degrades to the file path on UNIMPLEMENTED and
    on nothing else, so a live-incapable model asked at a wrong rate was
    answering INVALID_ARGUMENT and costing the caller its fallback.
  - a negative sample rate is refused instead of silently becoming 16000. Zero
    still means 16000, which is what the proto documents; -1 is malformed rather
    than absent and gets the same refusal every other bad rate gets.

And one thing recorded rather than changed, at the handler: "live" here means
incremental INPUT, not low latency, and with the pinned families it does not yet
mean incremental OUTPUT either. nemotron_asr's process_audio_chunk only appends
to its buffer, so its whole decode and every delta happen inside finalize(),
after the client closes its send side. The policy-window buffering is inert for
that family and matters only for vibevoice_asr and higgs_audio_stt.

Verified on the wire with live_idle_timeout_ms:3000. A silent client acked at
371 ms and was cancelled at 3.371 s; a second live stream opened one second
later received its ack 2.37 s in, i.e. at the instant the first was cancelled,
and then transcribed successfully on the same cached session. Without the
watchdog it would still be waiting. Re-ran the live transcription (ready first,
59 incremental deltas, concat equal to the final text, word timestamps in
nanoseconds, eou and eob false), the citrinet refusal at both a right and a
wrong rate (UNIMPLEMENTED either way now), and Task 12's AudioTranscriptionStream
on nemotron_asr, which is unchanged.

Mutation testing the watchdog found a weakness in its own test: the destructor
test slept past the window inside the watched scope, so a destructor that
DETACHED the thread instead of joining it passed unnoticed. The test now uses a
window longer than the scope, which kills that mutant, and says why.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): refuse the unsupported RPCs with a reason

AudioEncode, AudioDecode, AudioTransformStream, AudioToAudioStream and
VoiceEmbed have no counterpart in audio.cpp's VoiceTaskKind. Each now returns
UNIMPLEMENTED naming the loaded family, what that family does support, and the
upstream limitation, instead of the generated base class's bare status. The
reasons live in a table in capability_routing.cpp so they are data rather than
literals copied into five handlers, and so a test can assert every one of them.

The five claims this was planned against were re-read at the pinned upstream
e800d435d130dc776baf6f3e6129bb62b1495c89, and one did not hold. "audio.cpp
streams tts and asr only" is false: silero_vad advertises vad with
RunMode::Streaming. The refusal stands on the narrower claim that survives, that
no family advertises streaming for any task AudioTransform routes to, and a test
asserts the refuted wording does not come back.

VoiceEmbed is the one refusal whose request carries a ModelIdentity, so it runs
the #10952 check before answering: a stale route must get NOT_FOUND and the
router's sentinel, not "audio.cpp cannot embed speakers" about a model that is
not loaded here. It cannot use snapshot_for, whose no-model branch would tell
the caller to load a model when no model can help, so it takes the reference
through snapshot_unchecked and checks identity itself. That function's comment
now names three classes of caller instead of two.

The two bidirectional surfaces refuse without reading their stream, verified
with a client that writes a config and eight frames first and gets the status
rather than hanging.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): correct the vevo2 clause, and assert the absences

Review found a false clause in the AudioToAudioStream refusal. It said s2s is
"offline voice conversion ... which converts one clip into another speaker's
voice", which is true of miocodec and false of vevo2: vevo2's s2s route is
`editing` and only `editing` (default_route_for_task and route_matches_task in
src/models/vevo2/session.cpp), documented as "Edit source speech into new target
text while using the target voice" and requiring --target-text, so it rewrites
what was said. vevo2's voice conversion is its separate vc task. It now reads
"offline clip-to-clip processing against a target voice, declared only by
miocodec (voice conversion) and vevo2 (speech editing)", and a test asserts the
miscast cannot come back. The conclusion is unchanged: neither family converses.

That defect was undetectable on the wire, since vevo2 does not load here, which
is the argument for upstream_absence_ctest.cpp. It links engine_runtime purely
to interrogate make_default_registry() and asserts the five premises the refusal
reasons rest on: no codec task kind, no family advertising spk, no streaming for
sep/vc/svc/s2s, miocodec advertising exactly vc and s2s, and s2s advertised by
exactly miocodec and vevo2. The last two are exact sets, so an addition fails
here rather than leaving a message stale. A positive control proves the registry
is populated and the query works before any absence is believed, and every
assertion has a reproduced negative control. This turns an AUDIO_CPP_VERSION
bump from "remember to re-read five prose paragraphs" into a test failure.

unsupported_surface now switches over UnsupportedRpc with no default label, so
-Wswitch reports a sixth enumerator added without a row at build time; the
runtime bounds guard it replaces is deleted.

The AudioTransformStream reason had a true premise and an overreaching
conclusion: an offline sep family could be buffered into a stream, as other
LocalAI backends do. It now says this backend declines to offer a buffered
offline call in disguise, rather than implying impossibility.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): make the missing-switch-case diagnostic fatal

unsupported_surface() switches UnsupportedRpc onto the table row that explains
it, with no default label, so -Wswitch reports an enumerator nobody handled. As
a warning that is not enough: adding a sixth enumerator and building the shipping
target gives exit 0, a binary and one warning, and the trailing
`return surfaces[0];` then answers the new RPC with AudioEncode's codec reason.
That is a confident, specific and false statement about audio.cpp on the wire, on
the one code path whose entire job is to be truthful about what this backend
cannot do, and it is worse than the runtime fallback it replaced, which at least
named itself as a bug in this file.

capability_routing.cpp therefore joins loaded_model.cpp on the existing
-Werror=switch pin, whose comment already made this argument for the engine enum.
The comment now covers both files. The pin stays per-file rather than
project-wide because upstream's own ace_step/vae_decoder.cpp has unhandled
-Wswitch cases of its own.

Verified: a sixth enumerator now fails `make grpc-server` with exit 2 and no
binary; appending a 14th VoiceTaskKind upstream still fails loaded_model.cpp, so
the two pins fire independently; both reverted clean.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): package the backend image

Bundles the dependency closure for the from-scratch image, the dlopened ggml
CPU-variant shared objects that ldd cannot see, and upstream's bundled
silero_vad and marblenet_vad assets so VAD works with no download.

The bundled loader sits in the package ROOT rather than at lib/ld.so. run.sh
execs it, which makes /proc/self/exe name the loader, and this backend has two
consumers of that path: ggml discovers the libggml-cpu-*.so by listing
dirname(/proc/self/exe), and resolve_model_path expands bundled:<name> under the
same directory. Rooting the loader makes the binary, the ggml objects and
assets/ share the one directory all three resolution mechanisms agree on.
llama-cpp's lib/ld.so layout would need assets/ moved into lib/ as well.

The image builds against apt gRPC and protobuf, like Dockerfile.ds4 and unlike
Dockerfile.privacy-filter. The from-source gRPC that install-base-deps.sh and
the base-grpc-* images supply vendors protobuf 26, which pulls abseil into
message_lite.h; with SPM_PROTOBUF_PROVIDER=package that collides with
sentencepiece's vendored mini-abseil and every absl::internal reference becomes
ambiguous. Noble's protobuf 3.21.12 predates the abseil dependency and is the
pair every earlier verification of this backend ran against.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): exempt the driver libraries from the packaging gate

package.sh already left libcuda.so* and libnvidia-* to the host when copying,
because the driver has to match the kernel module on whatever host runs the
image, but the validation gate had no matching exemption. With BUILD_TYPE=cublas
ggml is static and links CUDA::cuda_driver, so grpc-server carries DT_NEEDED
libcuda.so.1 and the gate would have rejected the very absence the copy loop
created, failing every cublas build in CI. One regex now feeds both.

Building a control for that found a second defect: ld.so --list refuses to trace
an object with an unresolvable dependency at all, exiting 127 without emitting a
per-library line, so the "=> not found" rule was dead code and no exemption could
have applied to it. The gate now traces with LD_TRACE_LOADED_OBJECTS and
LD_LIBRARY_PATH, which reports the missing name and exits 0, and which is also
what run.sh does at run time.

Adds a layout assertion so a future move of the loader into lib/ fails the build
instead of shipping a package that resolves bundled: models into lib/assets and
finds no ggml CPU backend, and records for Task 16 that the Darwin script must
not be a straight copy of privacy-filter-darwin.sh, which never calls package.sh
and would silently drop assets/.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): register the backend with CI and the gallery

Adds the five Linux matrix entries (cpu amd64/arm64 sharing a tag-suffix so the
manifest merge fires, cuda 12, cuda 13, vulkan), the path-filter case that keeps
later PRs touching backend/cpp/audio-cpp/ from getting zero CI jobs, the
bump-bot entry pointing at the AUDIO_CPP_VERSION pin in the backend Makefile,
the gallery meta plus its -development variant and the image entries for every
variant, and the Makefile docker-build wiring.

The matrix entries carry base-image only, with no builder-base-image, unlike
the llama-cpp and privacy-filter blocks they sit next to. The prebuilt
quay.io/go-skynet/ci-cache:base-grpc-* images ship a from-source gRPC whose
protobuf v26 depends on abseil, and this backend's sentencepiece is built with
SPM_PROTOBUF_PROVIDER=package, so it sees real abseil's absl::lts_20240116::
internal alongside its own vendored plain absl::internal and every
absl::internal:: reference becomes ambiguous. Building against base-grpc-amd64
fails at sentencepiece-static.dir/error.cc.o with "reference to 'internal' is
ambiguous". Dockerfile.audio-cpp installs apt's gRPC/protobuf 3.21.12 itself,
which is also the pair every unit and end-to-end run of this backend has been
verified against, and the CUDA toolkit therefore has to come from base-image.

No Darwin matrix entry and no metal gallery entries: the Metal build needs
scripts/build/audio-cpp-darwin.sh, a backends/audio-cpp-darwin make target and
a routing step in backend_build_darwin.yml, none of which exist yet, so an
entry added now would be routed to build-darwin-go-backend and look for
backend/go/audio-cpp/. The inferBackendPathDarwin case and the
DARWIN_BESPOKE_BUILDERS membership are in place, inert, so that adding the
entry later is a one-line change that cannot be claimed by the generic Go path.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): pin the CUDA architectures, drop the vulkan variant

Upstream sets CUDA_ARCHITECTURES to `native` on the engine_runtime target
whenever CMAKE_CUDA_ARCHITECTURES is unset at root scope, and docs/build/
linux.md says so outright. ggml's own default does not rescue it: it
list(APPEND)s in the ggml subdirectory scope, which never reaches the root
scope where the engine_runtime property is decided. No CI runner has a GPU for
`native` to enumerate, so both cublas entries would have gone red on the very
commit that first turns a CUDA build on.

Pin the list in backend/cpp/audio-cpp/Makefile, selected by CUDA_MAJOR_VERSION,
which Dockerfile.audio-cpp now forwards from the CI build-arg it was previously
discarding. The values are copied from ggml's own version guards rather than
invented, so engine_runtime and ggml compile for the same set: CUDA 12 keeps the
Maxwell/Pascal/Volta virtual archs and stops at 120a-real, CUDA 13 drops them
and adds 121a-real. The `a` suffix is used rather than `f` because the latter
needs CMake 3.31.8 and Ubuntu Noble ships 3.28.3. Verified by driving CMake
3.28.3's own CUDA architecture validator over both lists, with 120f-virtual as
the rejected control.

Drop the vulkan matrix entry, its two gallery entries, the vulkan capability
key on both metas and the Vulkan tag. Every other vulkan backend gets its Mesa
ICD drivers from .docker/install-base-deps.sh, which package-gpu-libs.sh then
bundles; Dockerfile.audio-cpp calls neither and installs only libvulkan-dev and
glslc, so the image would ship a Vulkan loader that finds no GPU. No CI job runs
a vulkan image against real hardware, so that would have passed green and failed
in users' hands. BUILD_TYPE=vulkan stays supported for local builds.

Also note on the cublas entries that cuda-major-version now selects the
architecture list and that cuda-minor-version and the base-image tag encode the
same toolkit, and correct the stale entry counts on matrixEntryKey.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): build for Darwin Metal

Bespoke C++ Darwin path like ds4 and privacy-filter: an includeDarwin matrix
entry, a backends/audio-cpp-darwin make target, a gated workflow step, and the
metal image entries plus metal/metal-darwin-arm64 capability keys in the
backend gallery.

The build script deliberately does NOT reassemble the package the way
privacy-filter-darwin.sh does. It runs the backend's own `make package` and
copies the result, so the Darwin package keeps the root-level layout the Linux
one has: grpc-server, run.sh, the ggml objects and assets/ in one directory,
with lib/ for the dylib closure. Hand-assembling would drop assets/, and
assets/ is what makes the bundled: model paths resolve with nothing downloaded.
The dylib walk is a full transitive closure rather than the single level ds4
and llama-cpp do, because Homebrew's grpc++ pulls libgrpc, abseil, upb, cares
and OpenSSL that grpc-server does not link itself, and a level-1 walk ships a
package that only works on a machine that already has Homebrew grpc.

Two fixes folded in, both in the backend Makefile:

  - an EMPTY CUDA_MAJOR_VERSION fell through to the CUDA 12 architecture list,
    which contains 120a-real and so needs nvcc >= 12.8. A local
    BUILD_TYPE=cublas build on a 12.0-12.7 host failed to compile where
    upstream's documented default (native) worked. EMPTY now maps to native,
    12 and 13 keep their lists, and any other non-empty value is an error on
    cublas builds. CI always passes a major, so CI is unaffected.

  - the Darwin branch now points CMake at Homebrew's keg-only libomp. AppleClang
    ships no OpenMP runtime and nothing is symlinked into /opt/homebrew, so
    FindOpenMP finds neither the library nor the header, and audio.cpp calls
    find_package(OpenMP REQUIRED) whenever ENGINE_ENABLE_OPENMP is on. Without
    the hint the macOS build would have died at configure time. If the keg is
    absent the build disables OpenMP instead of failing.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): make the Darwin fallbacks loud and the rpath walk complete

Review follow-up on the Darwin Metal build.

The OpenMP fallback was silent. If brew --prefix libomp ever comes back empty,
CI produced a green Metal package with 108 #pragma omp directives across ~30
files compiled out, and clang says nothing about an ignored omp pragma without
-Wsource-uses-openmp, so the only trace was one absent flag inside a set -x
cmake line. That regression would have been blamed on Metal. It now warns.

The @rpath arm of the dylib walk had no live candidate when it was written, on
the reasoning that a Metal build links ggml statically. The OpenMP fix in the
same commit made libomp.dylib one, and whether Homebrew records it as an
absolute opt path or as @rpath/libomp.dylib is not observable from Linux. The
walk now expands @rpath, @loader_path and @executable_path against the object's
own LC_RPATH entries, and only fails when nothing on disk answers, printing the
rpath list with the error so a failure on a machine nobody can attach to
explains itself.

Also: ADDITIONAL_LIBS now go through the closure rather than a bare cp, so they
are deduplicated and their own dependencies bundled; build/darwin/lib is
created explicitly instead of relying on package.sh pre-creating it; the libomp
probe uses nested ifneq rather than $(and ...), which needs GNU make 3.81 and
would otherwise expand empty and take the OFF branch on an older make; and
-DOpenMP_ROOT is quoted like its CUDA sibling.

Verified with a Linux harness that runs the script verbatim against a stubbed
otool: a level-2 transitive dep, an @rpath dep reachable only through LC_RPATH,
and an ADDITIONAL_LIBS dep are all bundled, a dependency cycle terminates,
system libraries are skipped, the packaged tree has assets/ at the root beside
grpc-server with the dylibs in lib/, and both failure paths exit non-zero.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): make bundled: reachable from a model YAML

resolve_model_path() tested the bundled: prefix on `candidate`, which prefers
ModelFile and falls back to Model. LocalAI fills ModelFile by joining ModelPath
onto the configured model string (pkg/model/loader.go, LoadModelWithFile), and
only sets it from a managed artifact otherwise, so a model YAML saying
`model: bundled:silero_vad` arrives as ModelFile "/models/bundled:silero_vad"
and Model "bundled:silero_vad". The prefix therefore never matched through the
normal load path: it matched only for a hand-written LoadModel call that left
ModelFile empty, which is exactly how task 15 verified it, and every model YAML
using the form failed with "model path does not exist:
/models/bundled:silero_vad".

Both fields are now checked, Model first, so the zero-download VAD path the
package ships assets for is reachable the way it is documented. A caller that
puts the form in ModelFile still works, so task 15's verification stands.

Compiled clean; the runtime check could not run on this host, whose system
libprotobuf/libre2 have gone missing (the pre-existing grpc-server binary no
longer resolves its libraries either), so it wants a container run.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): advertise the backend and document its options

Registers audio-cpp as preference-only in /backends/known: the family lives in
GGUF metadata that an importer cannot read from a remote repo, and one repo
hosts thirty families, so there is no honest auto-detect signal. Modality is a
single string and the import form chips on a fixed key set, so it registers as
tts with the other modalities named in the description rather than under an
invented key the UI would bucket as "other".

Adds a features page covering the option namespacing, the routing table per
endpoint, the RPCs this backend declines and why, the bundled VAD path, the
separation stem behaviour, and the family gotchas (supertonic needs the orig
package; chatterbox advertises cloning and no plain tts; nemotron_asr defers
its whole decode to finalize so live transcription emits nothing until the
client half-closes, unlike higgs_audio_stt and voxtral_realtime). Every option
name and family capability in it was read off the pinned upstream checkout.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(audio-cpp): test resolve_model_path, and correct the family names

The bundled: fix in 842443cd7 shipped without a test, which is how the bug got
there: task 15 verified the form with a hand-written LoadModel that left
ModelFile empty, and that is the one shape the server never produces. Four cases
in streaming_driver_ctest, which already links loaded_model.cpp, pin the
PRODUCTION shapes instead. The first fails against the pre-fix source (returns
the joined /models/bundled:silero_vad); the other three are the branches the
bundled: lookup now runs in front of and must fall through for.

Three family names in the docs were the source directory rather than the
registered family, on pages whose whole argument is that these names cannot be
guessed: demucs is htdemucs (demucs/loader.cpp:22), roformer is
mel_band_roformer (roformer/assets.h:15), and moss is TWO families,
moss_tts_local and moss_tts_nano. The hyphenated ASR names are underscored to
match, here and in the compatibility table.

The supertonic dtype note claimed more than the evidence carries. The f16 abort
is a local observation, identical through TTS and TTSStream; upstream's
docs/gguf.md leaves the 16-bit column untested and records q8_0 as "No
(unsupported weight dtype)", which says unusable rather than fatal. Both are
still refused, because the allow list is what the family can run. Corrected in
family_gate.h, family_gate.cpp and the docs together, since the docs inherited
the wording from the code.

The importers tripwire says in the file that it is a tripwire: it exercises no
audio-cpp behaviour, and the registration assertion lives in backend_test.go.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* gallery: add audio.cpp models covering every served RPC

One representative model per RPC group of the audio-cpp backend, plus the two
bundled VAD models, which need no download at all because the assets ship
inside the backend package.

Every hash was computed with sha256sum on the downloaded file. Quantizations
come from upstream's tested-status table in docs/gguf.md rather than a default
of q8_0: supertonic ships the orig package (its q8_0 is recorded as an
unsupported weight dtype and its f16 aborts in ggml_concat), and nemotron_asr
and htdemucs ship f16 because their q8_0 builds are recorded with drift while
16-bit is a clean pass.

Diarization and separation use the diarization and audio_transform usecases,
not transcript: /v1/audio/diarization and /audio/transform filter the default
model on FLAG_DIARIZATION and FLAG_AUDIO_TRANSFORM respectively, so a
transcript flag would have hidden both models from their own endpoints. The
forced aligner sets parameters.language, which the transcription endpoint uses
as the fallback when no language form field is sent, because the family
requires both a transcript and a language.

All ten entries were run twice: once against the raw gRPC server, and once
installed with local-ai models install and called through the HTTP endpoint.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* gallery: correct the audio.cpp entries' licenses

Swept all ten entries against the real upstream named in audio.cpp's
tools/model_manager.py rather than against the audio.cpp repo's own license.
Three were wrong:

  supertonic   apache-2.0 -> openrail      weights come from
                                           mlx-community/supertonic-3-mlx, and
                                           both it and Supertone/supertonic are
                                           openrail
  citrinet     apache-2.0 -> other         pulled from NGC
                                           nvidia/nemo/stt_en_citrinet_256,
                                           governed by the NGC Terms of Use
  sortformer   other -> cc-by-nc-4.0       nvidia/diar_sortformer_4spk-v1 is
                                           CC BY-NC 4.0, and the gallery already
                                           uses that exact string, so there is no
                                           reason to obscure a non-commercial bar

The license field is one word, so citrinet and sortformer also gained a
sentence saying why they are restricted. The other seven were confirmed
correct against their sources.

Also drops an unverified claim from the nemotron description. It said the
model drives the realtime transcription session; that endpoint actually calls
TranscribeStream, and the live RPC reaches LocalAI only through
realtime_semantic_vad.go. Neither path was exercised here, so the description
now states only the two calls that were.

MarbleNet gains the NeMo upstream under urls: for parity with silero.

No sha256, quantization, usecase or model choice changed.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(audio-transform): bound sample_rate, keep same-named uploads apart

Four defects the whole-branch review found on the Go side, plus two comment
corrections.

sample_rate is a disk-exhaustion hazard. The branch added the `form:` tag that
makes the field bind for the first time, so the resample path went from dead to
live, and utils.AudioResample interpolates the int straight into ffmpeg's -ar
with no bound. Measured with ffmpeg 7: -ar 999999999 on a 0.01 s clip writes
20 MB and exits 0, which scales linearly to the reported 3.9 GB for one second,
into a GeneratedContentDir nothing sweeps, and convertStems repeats it once per
separation stem. Clamped to 8000..192000 in the handler, before the temp dir and
before the model is touched, and rejected with a 400 outside it.

The low end was reported as "a 0-byte file". It is not: -ar 1 writes a 78-byte
header with no audio behind it, whose declared data size still claims 70 bytes,
so go-audio parses it as a 35 SECOND file and a size check does not see it. The
guard therefore compares the declared data chunk against the bytes actually on
disk, and AudioResample now fails rather than returning a WAV carrying nothing.

Both parts of a transform request land in one temp dir, and the raw copy was
named only after the client's basename, so `-F audio=@mic/clip.wav
-F reference=@loopback/clip.wav` wrote "raw-clip.wav" twice. Since
AudioToWavPreservingShape hardlinks an already-PCM16 WAV rather than copying it,
the reference part's os.Create truncated the inode audio.wav pointed at: mic and
reference came out identical, which makes an echo canceller null everything and
return near-silence with a 200. The raw copy now carries the form field name.

audio-cpp had no BackendCapabilities entry, so VoiceCloningForModel returned nil
before it ever consulted the model's tts.voice_cloning override and every
`voice: "profile:<id>"` request was refused with a 400, on a backend that ships
audio-cpp-chatterbox whose family serves cloning and not plain TTS. Registered
with its RPCs, usecases and the reference-audio contract, and deliberately
without the 16 kHz mono fold, which its separation families cannot survive.

GetBackendCapability was exact-match only, so every pinned gallery variant read
as an unknown backend: vulkan-localvqe lost the 16 kHz mono fold that used to be
unconditional and started failing inside LocalVQE, and the usecase gate does not
stand in for it because BuildFilteredFirstAvailableDefaultModel returns early
once the client names a model. Lookup now falls back to the meta name by
stripping the gallery's hardware prefix and release-channel suffix, exact match
first so nothing can be shadowed. Same class as #10945.

Also corrected: the AudioTransformRequest comment claimed echo's binder falls
back to the field name, which it does not in either direction (bindData binds
ONLY tagged fields and `continue`s otherwise; `model` arrives from
setModelNameFromRequest's c.FormValue). And the stable_audio `src` heap
corruption caveat now lives on ElevenLabsSoundGenerationRequest, where the Go
developer who would add the field can see it, instead of only in C++.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(audio-cpp): refuse a task pin the RPC cannot serve, and stop empty frames holding the lane

The model's `task:` option is copied into the request shape by all nine
handlers, which is correct, but resolve_route then replaced the RPC's candidate
list with the pin WHOLESALE and never asked whether the pin was something that
RPC routes to. One pin therefore bled across all nine surfaces, and because the
family still supported the pinned task the result was a wrong 200 rather than an
error. Reproduced live: nemotron with task:asr made Vad return 200 with zero
segments after a full ASR decode, so 14 seconds of speech was reported as
silence, and Diarize did the same; silero_vad with task:vad made
AudioTranscription return 200 with empty text and four segments whose spans were
VAD segments, which combined with response_format in {text,srt,vtt,lrc} building
the body solely from Segments[].Text yields a well formed SRT of four timed
EMPTY cues. It also contradicted the documented contract, that a family which
cannot serve a request is refused rather than rerouted.

A pin is now checked against the RPC's admissible task set before it is adopted,
and the refusal names both the pin and the RPC. The set is derived from
task_candidates with every shape flag set rather than restated, so a task added
to an RPC's candidates cannot become inadmissible by omission. Every legitimate
pin survives, and the test asserts all fifteen of them alongside the eight
crossings that must not.

The live watchdog was defeated by empty frames. idle.touch() ran on ANY message,
before the has_audio and pcm.empty() filters, so a peer writing unset-oneof or
zero-length frames faster than the window held the lane indefinitely while
feeding the decoder nothing. There is one lane per model and one model per
process, so that is a single client denying the whole backend, which is what the
watchdog exists to prevent, and the thrown text already said "no audio frame
arrived". The touch moved below the filters, which are now a named predicate so
the distinction is testable rather than a call order nobody can see.

Three comments corrected against measurement rather than reasoning:

- CMakeLists claimed zero google::protobuf:: definitions remain in the
  executable. nm -C --defined-only reports 2515, and that is expected: they are
  generated code, sentencepiece::ModelProto's own _InternalParse among them. The
  claim that holds, and the one the ABI fix is actually about, is that no
  vendored protobuf RUNTIME is linked and ParseContext::ParseMessage is
  UNDEFINED in the executable, resolving to libprotobuf.so.
- refuse_cloning_without_a_clip's "cannot misfire" paragraph had its reasoning
  backwards. Routing picks VoiceCloning as the FALLBACK when there is no clip,
  which is the case being caught; chatterbox, which ships in the gallery,
  advertises clon and no tts at all, so every voice-less request lands there.
- audio_units read "2.1 min at 96 kHz" for index 11289602, which is 1.96 min.
  2.1 min is 96 kHz's OWN first failure at 12288002. Both were remeasured and
  the note is now a per-rate table.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* build(audio-cpp): exclude the upstream checkout from the C++ gate, harden the darwin walk

run-unit-tests.sh pruned */llama.cpp/* but not */audio.cpp/*. It is safe today
only by luck: upstream's 44 tests all put "test" at the FRONT of the filename
(17 test-*.cpp, 27 test_*.cpp, zero *_test.cpp), so the glob misses every one of
them, and nothing enforces that. This gate runs on every PR for every backend
and compiles each match as a standalone translation unit with nothing but
nlohmann/json on the include path, so the day upstream adds or renames one test
the gate goes red repo-wide on an Apache-2.0 file nobody here wrote.

audio-cpp-darwin.sh now logs the raw otool -L output and the parsed LC_RPATH
list unconditionally, before the walk. Both awk filters in that script assume a
column layout nobody working on this can observe, since it runs only on the CI
Mac, and a green first Darwin run proves nothing about the assumption: an awk
that silently matched nothing yields an empty dependency list, which reads
exactly like "no non-system dependencies" and packages happily. Both filters
otherwise feed process substitutions, so their input never reached the log.

It also lists every symlink in the package and fails on one that cannot resolve
inside the image. A dangling link does not fail anything else here, because
every assertion tests with -e, which follows links; it fails at dlopen on a
user's Mac. Links are NOT banned outright, which the review suggested but which
would break the libggml.dylib -> libggml.0.dylib chain the `cp -a` above exists
to preserve. What is banned is a link that resolves on the build host and will
not resolve in the image: a broken one, or an absolute one pointing outside the
package.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* style(audio-cpp): drop em dashes from the audio-cpp capability entry

Follow-up to a84b3c4b9, no behaviour change.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): key the voice-cloning model rule on the resolved backend

Making GetBackendCapability strip the gallery hardware prefix and release
channel fixed pinned variants of /audio/transform, but VoiceCloningForModel
kept keying its per-backend switch on the caller's spelling. A pinned name
therefore resolved the capability by stripping and then missed every case in
the switch, falling through to the permissive default: cuda12-vibevoice-cpp
advertised voice cloning for the realtime 0.5B model, metal-coqui for
tacotron2, cuda12-crispasr for a pure ASR model, cpu-qwen3-tts-cpp for
CustomVoice. Each of those is a model that cannot clone, so /v1/audio/speech
accepted a profile: voice it had to fail on inside the backend rather than
rejecting it with a 400, and the UI advertised the capability too.

resolveBackendCapability now returns the key the entry was found under, and
callers that branch on backend identity use that key instead of the name they
were handed. The exact-match-first order is unchanged, so a backend genuinely
registered under a variant-looking name still keys on its own name.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* gallery(audio-cpp): declare audio_transform on the chatterbox entry

Chatterbox advertises VoiceCloning AND VoiceConversion
(src/models/chatterbox), and the entry's own description already said so, but
known_usecases listed only tts. /audio/transform selects its default model by
FLAG_AUDIO_TRANSFORM, so voice conversion was reachable only by naming the
model explicitly and was invisible to every usecase-driven surface. It is the
one audio.cpp task with a shipped gallery model and no way to find it.

Verified against the real model rather than inferred from the capability list:
AudioTransform with chatterbox-q8_0, speech as audio_path and a speaker clip
as reference_path, returns a 5.08 s 24 kHz mono WAV at -25.5 dB mean and zero
stems, which is the single-output shape voice conversion should have.

The description now says which endpoint reaches that half and warns that
installing this next to a source-separation model gives /audio/transform two
candidates, so the model should be named rather than defaulted.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* gallery(audio-cpp): add voice-design and singing-voice-conversion entries

Two of the three audio.cpp task kinds that had no gallery model now have one.
Both were driven end to end against the real weights through the backend
before being written, not inferred from the capability tables.

audio-cpp-irodori-voicedesign covers vdes. TTS carrying `instructions` routes
to the vdes task, so the voice is described in words rather than supplied as a
clip. Verified: "a calm elderly woman speaking slowly with a warm, gentle
tone" over an 8.76 s 48 kHz mono render at -16.8 dB mean, and a closed-loop
citrinet pass recovers the sentence with the accent drift expected from a
Japanese-first model read by an English recogniser.

audio-cpp-seedvc-singing covers svc, and pins task:svc because nothing else
can reach it. seed_vc advertises svc and ordinary voice conversion, no request
signal means "this input is singing", and auto-routing resolves the tie to
voice conversion every time. Verified with the pin: 5.04 s 44.1 kHz output
whose closed-loop citrinet transcription is exact.

s2s deliberately has no entry, and the reason is not effort. miocodec is the
only upstream family whose speech-to-speech route needs no text, and it
returned audio with correct duration and level but no recoverable speech in
four independent attempts: the stale build, v2 q8_0, v2 orig (the variant
upstream records as a clean Pass), both tasks, and matched 44.1 kHz inputs on
both sides. vevo2's route refuses with "Vevo2 text/prosody route requires
text_input or target_text", and session.cpp:897 fills target_text only from
request.text_input, which AudioTransform has no field to carry. The same
vevo2 weights convert voice correctly through the default route with an exact
ASR round trip, so the model and the plumbing are both healthy; it is the s2s
route specifically that this RPC cannot express.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* backend(audio-cpp): carry transform text through params, add the s2s entry

AudioTransform is audio-in / audio-out and its proto message has no text
field, but not every task it routes to is audio-only. vevo2's speech-to-speech
route is a text and prosody route: session.cpp:897 fills refs.target_text from
request.text_input and nowhere else, and the run refuses without one with
"Vevo2 text/prosody route requires text_input or target_text". The params map
is the only channel this RPC has that reaches the engine, so the text travels
through it and apply_transform_text_input unpacks it after the params have
been copied into task.options.

Before this, s2s was not awkward to reach through /audio/transform, it was
unreachable, and it was the last audio.cpp task kind with a real model and no
way to get to it.

target_text is canonical and text is its alias, the order vevo2's own option
table declares them in, so a request setting both gets the canonical one
rather than whichever the map happened to store first. An empty value falls
through to the next candidate instead of ending the search. language rides
along only when a text was found: on its own it conditions nothing, and
manufacturing a text_input for it would route a plain separation request
carrying a language hint through the text path. The keys are left in
task.options rather than erased, because vevo2's loader advertises target_text
as a request option and a family reading it there keeps working.

Nine tests, all confirmed failing on behaviour against a stub that returned
false before the implementation was written. Verified end to end afterwards:
vevo2-q8_0 with task:s2s and params[text] returns a 5.12 s 24 kHz output whose
closed-loop citrinet transcription is exact, and htdemucs separation with no
text param still returns its four stems, with and without params[stem].

audio-cpp-vevo2-speech-to-speech ships that route. Every audio.cpp task kind
with a loadable family now has a gallery entry; spk remains the only gap and
has no family upstream at all.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* docs(audio-cpp): document params[text] and the pinned transform tasks

The text channel and the two task pins are both invisible from the endpoint
contract alone: nothing in the AudioTransform form tells a reader that a
speech-to-speech model needs the line it is resynthesising, and nothing says
that asking for singing voice conversion without task:svc silently gets plain
voice conversion instead. Both are the kind of thing a user only discovers
from a refusal or, worse, from output that looks right and is not.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(utils): annotate the two G304 sites this branch introduced

gosec flags os.Open on a variable path, and both new call sites in ffmpeg.go
are its alerts on this PR. Neither is reachable by an outside caller: isPCM16Wav
opens the exact path it is about to hand ffmpeg as input, which in the upload
path is a server-created temp file named from path.Base of the client name so
no traversal survives, and wavAudioBytes opens AudioResample's own dst, a name
this package derives from src and has just had ffmpeg write.

Annotated in the repo's existing style rather than restructured, with the
reason spelled out, because a bare suppression is worth nothing to the next
reader. The three other G304 sites in this file, in passthroughWAV and
isTargetWav, predate the branch and are left untouched.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(audio-cpp): build arm64 with gcc-14 for the armv9.2 SME variants

The arm64 CPU image failed to build:

  cc1: error: invalid feature modifier 'sme' in
       '-march=armv9.2-a+dotprod+fp16+sve+i8mm+sve2+sme'

ggml's CPU_ALL_VARIANTS table includes armv9.2 variants compiled with +sme, and
Ubuntu Noble's default gcc-13 rejects that feature modifier. Every entry in the
table has to compile even though a host only ever dlopens the one its own CPU
supports, so a single unbuildable variant fails the whole image. gcc-14 accepts
it, which is exactly the fix llama-cpp already carries in
.docker/llama-cpp-compile.sh; this is the same problem reached by a different
Dockerfile.

Applied to every arm64 BUILD_TYPE rather than to the CPU one alone, and that
differs from llama-cpp on purpose. llama-cpp needs it only for its pure-CPU
image because its GPU builds run llama-cpp-fallback, which builds no variant
table. This backend's Makefile turns ENGINE_ENABLE_CPU_ALL_VARIANTS on for
every non-Darwin build, GPU included, so an arm64 GPU image would hit the
identical error. The matrix has no arm64 GPU entry today, which is precisely
why gating on an empty BUILD_TYPE would leave the trap armed for whoever adds
the first one.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-30 12:11:56 +02:00
mudler's LocalAI [bot]
7d8e0bac18 fix(model): deterministic, type-filtered backend auto-detection (#9287) (#10286)
* fix(model): deterministic, file-type-filtered backend auto-detect (#9287)

When a model config declares no explicit `backend:`, Load() fell into a
trial loop built by ranging the external-backends Go map (random order)
with no filtering, returning the first backend whose gRPC LoadModel
succeeded. An unrelated installed backend - e.g. the "opus" audio codec -
could therefore win a GGUF/LLM model load, so a model that should run on
llama.cpp wrongly tried to use opus.

Extract the candidate selection into a pure, testable function
SelectAutoLoadBackends that:

  - sorts the candidate list deterministically (no more map-order
    nondeterminism), and
  - for a `.gguf` model, filters to LLM-capable backends (via
    core/config.BackendCapabilities) and puts llama-cpp first, so an
    incompatible audio/codec/image backend can never win the trial loop.

If filtering would leave zero candidates, the full sorted set is returned
unchanged, so a previously-loadable model is never made unloadable.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(model): break core/config <-> pkg/model import cycle in backend auto-detect

The #9287 auto-detect change made pkg/model/autoload.go import core/config
for the backend capability table. core/config already imports pkg/model
(runtime_settings_registry.go uses model.DefaultWatchdogInterval), so this
closed a core/config -> pkg/model -> core/config import cycle and broke the
build and golangci-lint.

Invert the dependency so the lower-level pkg/model no longer imports the
higher-level core/config. pkg/model exposes RegisterLLMCapableBackendFunc and
uses the registered predicate; core/config (which owns the capability table)
registers it from an init(). The deterministic, GGUF-type-filtered selection
behaviour is unchanged. When the predicate is unwired the GGUF filter is
skipped, preserving the existing zero-candidate fallback.

The unit test now injects a fake capability predicate so SelectAutoLoadBackends
is exercised independently of the core/config table.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:opus-4.8 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-30 12:07:47 +02:00
Tai An
37f2087f97 fix(grammars): reject cyclic $ref in JSON-schema grammar to prevent stack-overflow crash (#11020) (#11041)
* fix(grammars): reject cyclic $ref in JSON-schema grammar to prevent stack-overflow crash

JSONSchemaConverter.visit resolved $ref entries by recursively calling
itself with no cycle detection. A client-supplied grammar_json_functions
schema whose $defs contains a self- or mutually-referential $ref (e.g.
{"A": {"$ref": "#/$defs/A"}}) made visit recurse until the goroutine
stack was exhausted, producing a fatal "stack overflow" that kills the
whole process rather than failing the single request. The schema is
converted synchronously in the /v1/chat/completions handler before any
backend call, so this is an unauthenticated remote crash. Fixes #11020.

Track the $ref targets currently on the recursion stack and error out
when one is re-entered, while popping after each descent so sibling
(non-cyclic) reuse of the same $ref is still allowed.

Signed-off-by: Tai An <antai12232931@outlook.com>

* fix(grammars): add a bounded recursion depth and cover llama31 $ref cycles

Addresses the review on #11041. The stack-set approach catches cyclic
$ref chains, but a deeply nested yet acyclic client schema (thousands of
nested arrays/objects) can still recurse through visit until the
goroutine stack is exhausted, which is the same unauthenticated remote
crash surface as #11020.

- Add a bounded depth counter to JSONSchemaConverter.visit (incremented
  with a defer-based cleanup, capped at maxSchemaDepth = 256, far above
  any realistic schema) so an over-deep schema fails the request with an
  ordinary error instead of crashing the process.
- Apply the same cyclic-$ref guard and depth bound to
  LLama31SchemaConverter.visit, the other production grammar entry point
  named in #11020, which previously had no cycle detection at all.
- Regression tests: a deeply nested acyclic schema is rejected while a
  moderately nested one still builds, plus direct/indirect $ref cycle
  and depth tests for the llama31 converter.

Signed-off-by: Tai An <antai12232931@outlook.com>

* test(grammars): make llama31 cycle fixtures valid function-call shapes

The two new llama31 $ref-cycle specs asserted on "cyclic $ref" but the
converter requires each top-level oneOf alternative to carry its
function-name property before descending, so both fixtures failed
earlier with "no function name found in the schema" and never reached
the cycle guard.

Give each fixture a valid llama31 shape: construct the converter with
NewLLama31SchemaConverter("function"), put "function": {"const": "test"}
on the top-level alternative, and hang the cyclic $ref under an
arguments property, so all 29 grammar specs pass and the assertions
genuinely observe the cyclic $ref error.

Signed-off-by: Tai An <antai12232931@outlook.com>

---------

Signed-off-by: Tai An <antai12232931@outlook.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-30 12:06:17 +02:00
mudler's LocalAI [bot]
c5ae41a29d chore: ⬆️ Update ikawrakow/ik_llama.cpp to 6647db9c27760044950fd6f99060456ae3d15df3 (#11204)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-30 11:23:57 +02:00
mudler's LocalAI [bot]
d225e15f0f fix(ci): skip the image and Go PR workflows on content they cannot see (#11218)
backend_pr.yml and test-extra.yml already filter themselves, so a gallery-only
or docs-only PR costs them about one job each. The image and Go workflows had no
filter of any kind, so a one-line gallery/index.yaml edit queued 20 jobs: 7
container image builds, 3 GoReleaser/darwin launcher builds, 3 unit test jobs, 2
golangci-lint, 1 e2e, 1 yamllint, plus the 3 that correctly stop after their
detect step. A docs-only PR queued the same.

This matters more than the job count suggests. Measured over the week to
2026-07-30, 97% of CI wall-clock is queueing and 3% is execution: a median
5-hour queue against a 4-20 minute median job. Cutting job count is the only
lever that shortens feedback time. The volume is there to cut, too: 13
gallery-only PRs merged that week with 10 open at once, and 78 of the 137 PRs
opened were bot-generated.

Add paths-ignore for gallery/**, docs/**, examples/** and **/*.md to the
pull_request trigger of image-pr.yml, build-test.yaml and tests-e2e.yml, and
add gallery/** to lint.yml, which already excluded the rest. That drops 13 of
the 20 jobs. None of the four can observe such a diff: gallery metadata is
parsed at runtime and never copied into an image, docs and markdown never enter
one at all, GoReleaser and the launcher take no such input, the e2e suite drives
backends over gRPC directly, and golangci-lint runs new-from-merge-base so a
diff with no touched Go lines is a no-op. The build-test exclusion also frees
macOS capacity, which is the scarcest runner class.

The two checks that do validate the gallery are deliberately left alone.
test.yml still runs core/gallery/variants_lint_test.go, which reads the real
gallery/index.yaml and asserts the index invariants, and yaml-check.yml still
lints the syntax.

paths-ignore skips a run only when every changed file matches, so a PR touching
the gallery and Go code still runs everything. master carries no branch
protection and no rulesets, so a skipped workflow reports no status and nothing
waits on it; .agents/ci-caching.md records that constraint for whenever required
status checks are introduced.

image.yml on master push is left unfiltered on purpose: skipping it would stop
the master and latest tags being republished for a gallery commit, which is a
publishing decision rather than a cost one.


Assisted-by: Claude:opus-5 [claude-code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-30 11:22:44 +02:00
localai-org-maint-bot
d6d9f899d6 gallery: add Nanbeige4.2 3B GGUF variants (#11170)
* gallery: add Nanbeige4.2 3B GGUF variants

Add Q4_K_M and Q8_0 builds of the compact Nanbeige4.2 agentic and reasoning model, grouped as install-time variants.

Assisted-by: Codex:gpt-5

* gallery: simplify Nanbeige4.2 model name

Apply the maintainer-requested canonical model name while retaining the quantization variants under the entry.\n\nAssisted-by: Codex:gpt-5

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-30 11:07:32 +02:00
localai-org-maint-bot
c5a7d394a5 gallery: add Laguna XS 2.1 GGUF variants (#11202)
Add the official Q4_K_M build and seven APEX quality and size variants for llama.cpp.

Assisted-by: Codex:gpt-5 [Hugging Face API]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-30 11:05:32 +02:00
localai-org-maint-bot
4b917936ef gallery: add Mellum2 Instruct GGUF variants (#11211)
Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-30 11:04:17 +02:00
mudler's LocalAI [bot]
aaec1d695e chore(model-gallery): ⬆️ update checksum (#11210)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-30 10:33:59 +02:00
Dimitris Karakasilis
c6c347ce13 feat(stablediffusion-ggml): make VAE tiling configurable (#11216)
GenerateImage hardcoded TilingParamsSetEnabled(vaep, false), so tiled VAE
decoding was unreachable from a model config even though all four upstream
setters were already bound in main.go.

Sampling runs in latent space, but the final VAE decode expands to full
resolution and needs one large compute buffer. At 1024x1024 that buffer
exceeds 8GB, which fails on two kinds of device: cards without the VRAM
for a full-frame decode, and drivers that cap a single allocation
regardless of how much memory is free. Mesa RADV reports a 4GiB
maxMemoryAllocationSize, so a Radeon 8060S with 74GiB of device-local
heap still cannot serve that decode:

    [INFO ] sampling completed, taking 251.82s
    [INFO ] decoding 1 latents
    ggml_vulkan: Requested buffer size exceeds device buffer size limit:
                 ErrorOutOfDeviceMemory
    [ERROR] vae: failed to allocate the compute buffer
    [ERROR] decode_first_stage failed for latent 1

Every sampling step completes and then the run is discarded at the last
stage, so the whole generation is wasted.

Add three options, parsed in Load and applied per generation:

    vae_tiling:true            enable tiled decoding (bare flag also works)
    vae_tile_size:512          tile size, or 512x384 for a rectangle
    vae_tile_overlap:0.25      overlap between tiles

Tiling stays off unless requested, so existing models are unaffected. Tile
size and overlap only reach the library when the operator set them, which
keeps upstream's defaults rather than pushing a zero, and an unparseable
value is treated as absent for the same reason.

Truthy spellings match what load_model already accepts for its own bool
options, and the bare-flag form matches diffusion_model, so no new
convention is introduced.

Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me>
2026-07-30 10:33:31 +02:00
localai-org-maint-bot
c10460d4de gallery: add Fara 1.5 27B GGUF variants (#11217)
Assisted-by: Codex:gpt-5 [web]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-30 10:23:22 +02:00
Ching
43b6ed2018 feat: add CAJAL gallery model (#9879)
Add the CAJAL GGUF gallery template and gallery index entry for local llama-cpp installs.

Assisted-by: Codex:gpt-5

Signed-off-by: Ching Kao <0980124jim@gmail.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-30 09:57:31 +02:00
ResearchForumOnline
7a7ebb5c2f add OpenZero Zero GGUF models to gallery (#11138)
Signed-off-by: ResearchForumOnline <116322650+ResearchForumOnline@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-30 09:47:38 +02:00
ghshhf
5b9aa02900 fix(ci): skip security scan on forks to avoid SARIF upload permission error (#10323)
The Security Scan workflow was failing on fork PRs because the workflow
does not have permission to upload SARIF files to the GitHub Security tab
when running from a fork.

This change adds '!github.repository.fork' checks to all steps
to prevent the workflow from running on fork repositories.

This fix should be applied to the main repository so that
all forks inherit the correct configuration.

Fixes #10322, #10318, #10320, #10321

Co-authored-by: ghshhf <ghshhf@users.noreply.github.com>
2026-07-30 09:42:48 +02:00
localai-org-maint-bot
5f055a407c gallery: add POCKET-35B GGUF variants (#11197)
Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-30 09:02:44 +02:00
Dimitris Karakasilis
d27c5e82ea Fix use case for video model (#11214)
Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me>
2026-07-30 08:57:09 +02:00
Tai An
efb43776ba fix(chatterbox): pin cublas12 torch/transformers and setuptools so the backend loads (fixes #11070) (#11074)
fix(chatterbox): pin cublas12 torch/transformers and setuptools so the backend loads

The cuda12-chatterbox gallery backend fails to load on a fresh install
because several deps in requirements-cublas12.txt are unpinned:

- torch/torchaudio: unlike requirements-cublas13.txt and
  requirements-cpu.txt, this file has no --extra-index-url, so pip pulls
  a wheel whose CUDA runtime (cu130) is newer than the host driver
  supports ("NVIDIA driver on your system is too old"). Add the cu124
  index and pin torch/torchaudio 2.6.0+cu124.
- transformers: resolves to 5.x, which dropped LlamaConfig.rope_theta
  that chatterbox-tts 0.3.1's T3 config still reads. Cap to <5.
- setuptools: 81+ dropped pkg_resources, which perth imports under a
  bare try/except and silently sets PerthImplicitWatermarker=None,
  making ChatterboxTTS.__init__ raise 'NoneType' object is not callable.
  Cap to <81 in requirements.txt.

Fixes #11070

Signed-off-by: Tai An <antai12232931@anaiguo.com>
Co-authored-by: Tai An <antai12232931@anaiguo.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-30 00:23:45 +02:00
mudler's LocalAI [bot]
d8a1e3c2e4 fix(realtime): echo response.metadata on response.created and response.done (#11198)
response.create accepts a metadata map and ResponseCreateParams has carried
the field all along, but triggerResponse never copied it onto the Response it
emits, so both terminals went out with metadata omitted.

That field is the only thing tying a terminal event back to the
response.create that asked for it. Our own doc comment on ResponseCreateEvent
says so — "the metadata field is a good way to disambiguate multiple
simultaneous Responses" — and it is what makes an out-of-band response
(conversation: "none") usable at all: a client running one alongside the
spoken conversation has no way to tell its own answer from the conversation's,
so it waits for a reply it already received and gave away.

Found from the client side: a headless text turn injected into a live session
was answered correctly in about a second, and the caller still blocked until
its own two-minute timeout because it could not recognise the answer.

Carry the map on liveResponse so all three terminals (in_progress, cancelled,
completed) report it, and leave it omitted when response.create sent none.

Assisted-by: Claude:claude-opus-5 gofmt

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-29 23:17:31 +02:00
localai-org-maint-bot
9bfd71387b feat(stores): add Valkey Search vector store backend (#11196)
* feat: add Valkey Search vector store backend

Add a new built-in Go gRPC store backend 'valkey-store' that implements the
four Stores RPCs (Set/Get/Delete/Find) against the Valkey Search module (FT.*)
using the pure-Go github.com/valkey-io/valkey-go client. It is selected via the
existing per-request 'backend' field on /stores, so there is no proto or HTTP
API change, and it mirrors the in-memory local-store while adding persistence
across restarts and opt-in HNSW.

Each vector is a Valkey HASH keyed by hex(little-endian float32); the index is
created lazily on first Set (FLAT+COSINE by default), cosine similarity is
derived as 1-distance, and namespaces get a collision-resistant token. Includes
unit tests (valkey-go mock) and env-gated integration tests against
valkey/valkey-bundle, plus build/matrix/gallery wiring and docs.

Assisted-by: Kiro:claude-opus-4.8 golangci-lint
Signed-off-by: Daria Korenieva <daric2612@gmail.com>

* Address review feedback: recover persisted index dimension, harden Find

- Load now recovers the persisted vector DIM from FT.INFO (not just index
  existence), so a post-restart Set/Find validates against the real DIM
  instead of silently re-learning a wrong one and dropping mismatched
  vectors from the index. This also restores Find's dimension check after
  a restart.
- StoresFind treats a dropped/missing index as an empty store (empty
  result, no error) and clears the stale indexCreated flag, matching
  local-store's empty-store behaviour.
- StoresSet reuses checkDims for its per-key length check so the four RPCs
  share one dimension-guard implementation.
- Add unit tests for FT.INFO dimension recovery, loadIndexState, and the
  dropped-index Find path.

Assisted-by: Kiro:claude-opus-4.8
Signed-off-by: Daria Korenieva <daric2612@gmail.com>

* Address review feedback: TLS ServerName/CA, Find nil-check, config fail-fast

Addresses external review comments on the valkey-store backend:

- StoresFind now rejects a nil/empty query Key before dereferencing it,
  so a malformed gRPC request can no longer panic the backend.
- TLS: derive ServerName (SNI) from the VALKEY_ADDR host so certificate
  verification works for IP-addressed endpoints, and add VALKEY_TLS_CA_CERT
  (custom CA bundle) and VALKEY_TLS_SKIP_VERIFY (testing-only) knobs.
- Config integer parsing now fails fast on a malformed value (e.g.
  VALKEY_HNSW_M=1x6) instead of silently defaulting, matching the
  fail-fast behaviour of the index-algo/distance-metric validation.
- Add VALKEY_DB (SELECT n) support for logical-DB isolation.
- Cap the human-readable part of a namespace token at 64 chars so a very
  long model name cannot produce an unbounded key prefix / index name
  (the appended short hash keeps distinct namespaces collision-free).
- Document the KNN-query injection-safety invariant (fields are constants)
  and why StoresGet uses a single aggregate DoMulti deadline for reads.
- Unit tests for the Find nil/empty-key guard, fail-fast HNSW parsing,
  and VALKEY_DB parsing/validation; docs + .env updated for the new vars.

Assisted-by: Kiro:claude-opus-4.8 golangci-lint
Signed-off-by: Daria Korenieva <daric2612@gmail.com>

* Address review feedback: configure valkey-store via model config

richiejp asked that the valkey-store backend take its configuration from
a model config rather than process-wide VALKEY_* environment variables,
so multiple stores can each have their own Valkey config within one
LocalAI process. This removes every env access from the backend and
routes config through the model-config seam every other backend uses.

- config.go: loadConfig(opts *pb.ModelOptions) now parses the model
  config `options:` list (key:value strings, split on the first ':')
  instead of os.Getenv. Option keys mirror the old VALKEY_* names without
  the prefix (addr, index_algo, distance_metric, ...). Defaults, fail-fast
  validation and the mandatory client name are unchanged.
- store.go: Load threads opts into loadConfig; TLS comments/errors renamed
  off the VALKEY_* names.
- core/backend/stores.go: StoreBackend and NewVectorStore take a
  *config.ModelConfigLoader, resolve the per-store ModelConfig by store
  name, and pass its Options (and Backend when unset) to the backend via
  WithLoadGRPCLoadModelOpts. No config -> default backend + built-in
  defaults, preserving the zero-config experience.
- Endpoints/routes/application: thread the config loader to StoreBackend.
- Unit + integration tests: configure via options; the integration test
  passes addr through the model-config path (VALKEY_ADDR is now only the
  test harness locating the server).
- docs + .env: document the model-config options, drop the env var table.

Assisted-by: Kiro:claude-opus-4.8
Signed-off-by: Daria Korenieva <daric2612@gmail.com>

* Remove valkey-store informational comment from .env The backend is configured via model config, not env vars — the comment was unnecessary noise in .env. The configuration is already documented in docs/content/features/stores.md.

Signed-off-by: Daria Korenieva <daric2612@gmail.com>

* feat(valkey-store): gate Load on NamespacePrefix to refuse autoload probing Mirror local-store's pattern: reject model names without store.NamespacePrefix so the model loader's greedy autoload probe cannot bind an arbitrary model name to the vector store backend (the #9287 failure mode). Also adds unit tests for the gate covering: prefixed namespace, prefix alone, unprefixed model name, empty model, and nil opts.

Signed-off-by: Daria Korenieva <daric2612@gmail.com>

* feat(valkey-store): add username_env/password_env credential indirection Add support for resolving Valkey credentials from environment variables named in the model config, mirroring cloud-proxy's api_key_env pattern. This keeps secrets out of model YAML files and lets distinct store configs each reference their own credentials. Options: username_env / password_env name the env var holding the value. The direct username / password options still work and take precedence when both are set (backward compatible). Includes 5 unit tests and updated stores.md documentation.

Signed-off-by: Daria Korenieva <daric2612@gmail.com>

* fix: correct rebase artifacts in backend-matrix.yml and Makefile Fix two issues introduced by the conflict-resolution script during the rebase onto master: 1. .github/backend-matrix.yml: valkey-store entries were merged INTO the cloud-proxy entries (duplicate keys in same YAML map items) instead of being separate list items. This broke cloud-proxy Linux builds and the cloud-proxy darwin entry lost its build-type/lang. Fixed by making them standalone entries and restoring cloud-proxy exactly as on master. 2. Makefile: duplicated .NOTPARALLEL and docker-build-backends lines. Collapsed to single lines that are master's current content plus the valkey-store additions. Also adds the three optional pickups from #10801: - /valkey-store in .gitignore (the built binary) - valkey-store row in docs/content/reference/compatibility-table.md - valkey-store line in backend/README.md

Signed-off-by: Daria Korenieva <daric2612@gmail.com>

---------

Signed-off-by: Daria Korenieva <daric2612@gmail.com>
Co-authored-by: Daria Korenieva <daric2612@gmail.com>
2026-07-29 20:12:29 +02:00
walcz-de
2f33d6dee0 docs(gpu): add ROCm 7.x and RDNA 3.5 / Strix Halo (gfx1151) to GPU acceleration guide (#9229)
* docs(gpu): add gfx1151 / ROCm 7.x and fix ROCm section

- Fix typo: "deditated" → "dedicated", "ROCm6" → "ROCm"
- Add ROCm 7.x to requirements (alongside ROCm 6.x)
- Add Ubuntu 24.04 to tested OS list
- Add AMD Strix Halo / gfx1151 section with kernel params,
  required env vars (HSA_OVERRIDE_GFX_VERSION, ROCBLAS_USE_HIPBLASLT),
  and Docker Compose example
- Add gfx1151 to the list of compiled GPU targets
- Add ROCm version column to verified devices table
- Add gfx1151 / Radeon 8060S (ROCm 7.11.0) as verified device

* fix(docs/gpu): correct gfx1151 section — env vars, image tag, safety warning

- Add all 4 required env vars (HSA_OVERRIDE_GFX_VERSION, ROCBLAS_USE_HIPBLASLT,
  HSA_XNACK=1, HSA_ENABLE_SDMA=0) with descriptions in a table
- Fix Docker Compose example to use the ROCm 7.x image tag (-gpu-hipblas-rocm7),
  not the ROCm 6.x image
- Add explicit warning: GGML_CUDA_ENABLE_UNIFIED_MEMORY must NOT be set
  (even =0 activates hipMallocManaged due to getenv != nullptr check)
- Add --force-recreate note (docker restart does not update container env)
- Add tested hardware note (Geekom A9 Mega / Ryzen AI MAX+ 395)

* docs(gpu): single ROCm image — drop -rocm7 tag suffix

Per maintainer feedback on PR #9229: there is only one ROCm/hipblas
main image, and it ships with ROCm 7.x by default — no separate
-rocm7 tag.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-07-29 20:09:52 +02:00
localai-org-maint-bot
ecdb32193d docs(proxy): cover long inference timeouts (#11065)
Document the reverse-proxy settings needed for long-running and multimodal requests, and distinguish edge-generated 504 responses from the optional LocalAI busy watchdog.

Assisted-by: Codex:gpt-5 [Codex]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-29 16:37:03 +02:00
Richard Palethorpe
9058a2bb46 feat: Add 3d generation UI/API and trellis2cpp backend (#10979)
* feat(3d): add Generate3D RPC, FLAG_3D capability, and /v1/3d/generations endpoint

Adds the plumbing for image-conditioned 3D asset generation (binary
glTF / GLB output), modeled on the video generation path:

- backend.proto: Generate3D RPC + Generate3DRequest (staged image src,
  glb dst, seed/step/cfg_scale/texture_steps, quality and background
  enums, params map for backend-specific extras)
- pkg/grpc: thread Generate3D through client, server, embed, base and
  the backend interfaces; connection-evicting and distributed-node
  wrappers (in-flight tracking + file staging) included
- core/config: FLAG_3D usecase (guessed only for the trellis2cpp
  backend), '3d' canonical usecase string mapped to the Generate3D
  method, and a '3d' output modality
- REST: POST /v1/3d/generations (+ unversioned alias) returning
  OpenAIResponse with a /generated-3d URL or b64_json; conditioning
  image accepted as URL, base64, or data URI; quality/background
  validated at the edge; .glb served as model/gltf-binary
- auth: '3d' route feature (default ON); /api/instructions entry

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* feat(trellis2cpp): add the trellis2.cpp image-to-3D backend

Wraps localai-org/trellis2cpp (C++/GGML port of Microsoft TRELLIS.2,
pbr-textures branch) as a Go+purego backend, following the
stablediffusion-ggml pattern:

- backend/go/trellis2cpp: purego bindings to the flat C ABI (v9,
  asserted at startup), eager pipeline load with model-set validation
  (refuses non-trellis GGUFs; degrades coarse/geometry-only/textured
  exactly like the upstream demo), Generate3D via t2_generate +
  t2_bake_glb writing a binary glTF to dst. Weight-free unit tests
  cover resolution/validation/param mapping — CI never downloads the
  multi-GB GGUF set or runs inference.
- CPU SIMD variants build into per-variant directories (the shared
  libggml sonames collide across variants, unlike sd-ggml's flat
  renamed-.so scheme); run.sh picks one via /proc/cpuinfo.
- CI wiring: backend-matrix entries (cpu, cuda12/13, vulkan
  amd64+arm64, l4t, l4t-cuda13, darwin metal), index.yaml meta +
  latest/master image entries, bump_deps tracking of the pbr-textures
  branch, changed-backends.js mapping, top-level Makefile targets.
- Importer: auto-detects trellis GGUF repos/URIs (registered before
  llama-cpp so the .gguf match isn't stolen) and expands any trellis
  URI to the full 10-file component set spanning the three LocalAI-io
  HF repos.
- Gallery: trellis2-4b (full PBR + 1024 cascade) and
  trellis2-4b-geometry (512 untextured) with verified sha256s.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* feat(ui): 3D generation page with native GLB viewer and IndexedDB history

Adds a Studio tab + /app/3d page for the new image-to-3D endpoint:

- GlbViewer ports the trellis2cpp demo's dependency-free WebGL2
  renderer (quaternion trackball, metallic-roughness PBR, ACES,
  hidden-line wireframe with a bounded index budget) and pairs it with
  a minimal GLB parser for the two forms t2_bake_glb emits — dense
  vertex-PBR (linear COLOR_0 + _METALLIC_ROUGHNESS, uploaded as
  normalized integers) and the opt-in UV-atlas textured form. Parsing
  happens before any GL so stats and errors render without WebGL2.
- use3DHistory stores past generations (params, input thumbnail, and
  the GLB blob itself) in IndexedDB with keep-newest-20 eviction —
  GLBs are multi-MB binaries localStorage can't hold — and the page
  offers a download button for the active GLB.
- Wiring: CAP_3D capability constant (FLAG_3D — the exact string
  /api/models/capabilities serves), threeDApi, router entries, Studio
  tab, vite dev proxy, en locale keys.
- e2e: render-smoke entry plus a focused spec that feeds a real
  one-triangle vertex-PBR GLB through the parser/viewer and exercises
  IndexedDB persistence, selection, deletion, and API errors.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* fix(3d): address API correctness and UX issues

Keep 3D generation on the LocalAI-specific /3d/generations route and ensure authentication and permissions cover it.

Propagate distributed transfer failures, publish a portable ARM64 backend image, honor importer overrides, and align discovery, upload validation, and touch controls.

Assisted-by: Codex:gpt-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* feat(3d): add previewable print remeshing

Add a single-detail CGAL Alpha Wrap workflow for existing Trellis GLBs, including PBR reprojection, API documentation, tracing, and an in-browser preview before download.

Allow the remesh route to enforce its 512 MiB upload cap independently of the smaller global default so generated high-resolution meshes can be processed.

Assisted-by: Codex:gpt-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* build(trellis2cpp): centralize remesh dependency pins

Assisted-by: Codex:GPT-5 [apply_patch] [exec_command]
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* fix(kokoros): implement Generate3D stub for new proto RPC

The Generate3D RPC added to backend.proto for the trellis2cpp backend
made tonic's generated Backend trait require generate3_d, breaking the
kokoros-grpc build. Return unimplemented like the other unsupported
modalities.

Assisted-by: Claude Code:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

---------

Signed-off-by: Richard Palethorpe <io@richiejp.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-29 16:15:04 +02:00
mudler's LocalAI [bot]
8089b2bf09 fix(ci): only rebuild the full backend matrix on breaking backend.proto edits (#11192)
backend/backend.proto is consumed by every language, so its SHARED_BUILD_INPUTS
rule could only ever be always/always: 417 Linux plus 56 Darwin builds. It fires
on ~1.3% of commits (10 of 767 over six months), which made it the single
largest CI cost driver in the repo.

On 2026-07-29 the queue reached 2178 jobs against 8 concurrent runners. Four
runs totalling 935 of those jobs were triggered by nothing but a proto edit. The
largest, 378 jobs on master, came from PR #11158, whose entire proto diff was
six lines adding `bool cache_prompt = 8;` to one message. No backend that does
not read that field behaves any differently for it.

Make the rule content-aware. changed-backends.js resolves backend.proto at the
base revision (the contents-API pattern already used for backend-matrix.yml) and
hands both texts to protoChangeIsAdditive(), which compares them structurally so
a comment reflow, reindent or field reorder does not read as a change. An
additive-only edit (new field with an unused number, new message, new enum
value, new RPC) suppresses the rule and rebuilds nothing; a removed, renumbered,
retyped or renamed field, a dropped RPC or a changed option still rebuilds
everything, as does an unresolvable base revision.

Every other matched rule is untouched, so a PR that edits the proto and
scripts/build/ is still a full rebuild, and the weekly full-matrix cron remains
the backstop for stale wheels.

Verified against all ten proto commits of the preceding six months: the nine
with a resolvable parent all classify as additive, and controls covering a
retyped-and-renumbered field, a deleted RPC, identical revisions and a
reindent-plus-comment-reflow all classify correctly.


Assisted-by: Claude:opus-5 [claude-code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-29 15:58:27 +02:00
localai-org-maint-bot
89ee62b2af gallery: add KAT-Coder V2.5 Dev GGUF variants (#11186)
* gallery: add KAT-Coder V2.5 Dev GGUF variants

Add Q4_K_M and Q8_0 builds of the newly released KAT-Coder-V2.5-Dev agentic coding model.

Assisted-by: Codex:gpt-5 [Hugging Face API]

* gallery: add KAT-Coder APEX variants

Assisted-by: Codex:gpt-5 [web]

* gallery: add KAT-Coder APEX checksums

Assisted-by: Codex:gpt-5

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-29 15:35:51 +02:00
Richard Palethorpe
49ef40a187 feat(classifier/VAD): support voice control on low power devices (#10804)
* feat(llama-cpp): route Score through the slot loop

Score previously bypassed the slot loop with a direct llama_decode: a
conflict guard aborted the whole process if scoring raced generation, the
config validator had to reject score alongside chat/completion/embeddings,
and every candidate re-decoded the full shared prompt.

Add SERVER_TASK_TYPE_SCORE to the (patched) upstream server so score tasks
are scheduled like any other slot work: generation and scoring serialize
naturally, the shared prompt is decoded once per call, and the slot's
prompt cache carries the conversation prefix across calls. Context
checkpoints at the score boundary and at the cache-divergence point keep
SWA/hybrid/recurrent models (e.g. LFM2.5) from re-prefilling the whole
prompt per candidate: warm-turn scoring on a 6-option set drops from ~8s
to ~0.5s on a desktop CPU.

The conflict guard and the validation split are removed; declaring score
with generation usecases on one config is now supported and shares the
slot cache.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* feat(realtime): classifier wire types and pipeline config

Wire types and YAML config for realtime classifier mode: sessions carry a
localai_classifier extension (options with canned replies/tool calls,
softmax threshold, normalization, history trimming, fallback modes, and a
deterministic wake-word address gate), mirrored by pipeline.classifier in
the model YAML and surfaced in the config-meta registry. The
localai.classifier.result server event reports the full score distribution
per turn.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* feat(realtime): classifier response flow

Classifier-mode responses: instead of autoregressive generation, each user
turn is prefill-scored against the option list (router.ScoreClassifier
prompt/candidate shapes over the Score primitive) and the winning option's
canned reply and tool call are emitted through the existing response
machinery. Below-threshold turns take the configured fallback (none /
canned reply / generate); empty transcripts and unaddressed turns (wake
word not mentioned) skip scoring entirely. The scoring probe defaults to
the latest user message only — small scorers echo canned replies from
prior turns back as the top option otherwise.

Built for hardware that can afford prompt processing but not decode: with
slot-based Score the option list stays KV-cached across turns, so a turn
costs roughly one forward pass over the new words.

session_update_error events now carry the validation cause instead of a
generic message.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* fix(realtime): bound the VAD tick's scan window and buffer retention

The VAD tick loop re-scanned the entire input buffer every 300ms and only
trimmed it on zero-segment ticks or commits. Audio that keeps producing
segments without a committing pause (steady noise a mic pipeline lets
through, music, continuous speech) grew the buffer toward the 100MB cap
with each tick rescanning all of it — O(n^2), measured at ~3.3ms of silero
per buffered second: past ~90s retained, ticks run back to back and pin
~4 cores until the stream stops.

Silero's recurrent state only carries a few hundred ms of context, so
rescanning old audio buys nothing. Clip the slice handed to the VAD to the
largest silence the commit test can need to measure (server_vad silence
window or the semantic eagerness fallback) plus a warm-up margin, and
rebase the returned segment times so every downstream consumer keeps
whole-buffer coordinates. An open turn whose clipped window is all silence
now commits (the silence outran the window) instead of being discarded as
no-speech. Independently, retain at most 90s of raw buffer, rebasing the
live-feed and EOU cursors on trim — this also bounds the previously
unbounded VAD-error path. Turn boundaries are otherwise unchanged: no
forced commits, no new coordinator states.

pipeline.turn_detection.vad_window_sec can widen the scan window; values
below the automatic floor are ignored. The tick body is extracted into
vadTick so specs can drive turn detection synchronously (same shape as
classifySoundWindow); the babble reproduction that pinned 4 cores now
plateaus under 10% of one core.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* fix(backend): let per-model threads override the global default

ModelOptions overrode a set per-model threads value with the app-level
--threads whenever the latter was non-zero — and WithThreads defaults it
to the physical core count, so it always was. The YAML threads: knob has
been dead config: a tiny VAD model could never opt down from the global
pool size.

SetDefaults already fills an unset per-model value from the app config,
which is the intended precedence; resolve threads through a helper that
honors it (explicit threads: 0 still means unset).

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* chore(gallery): single-thread the silero VAD

Silero is a ~2MB recurrent model with no exploitable graph parallelism:
measured per-call latency is identical at 1 and 10 ORT threads, while
every extra pool thread just spin-waits between the realtime loop's
frequent tiny inferences.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* docs(realtime): classifier mode, VAD scan window, threads precedence

Document the realtime classifier mode (options, threshold guidance,
wake-word address gate, empty-transcript handling), the VAD scan window
and 90s buffer retention (pipeline.turn_detection.vad_window_sec), the
per-model threads precedence, and the M3 classifier note in the realtime
state-machine design doc.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* perf(llama-cpp): score all candidates in one batched decode

One scoring call is now a single SERVER_TASK_TYPE_SCORE task: the slot
decodes the shared prefix (prompt + longest common candidate token
prefix) once, then forks one sequence per candidate off it
(metadata-only for the unified KV cache, copy-on-write for recurrent
state) and decodes every candidate's unique tail in one llama_decode.
Previously each candidate was its own task that restored the boundary
checkpoint and re-decoded its full tail sequentially, paying
per-candidate task and decode overhead.

The context reserves SERVER_SCORE_FORK_SEQS extra sequence ids (and
recurrent-state cells) beyond the parallel slots via the new
common_params::n_seq_score_forks. Forking requires the unified KV cache
(already this backend's default) since per-sequence streams would shrink
n_ctx_seq; an explicit kv_unified:false disables forking and Score calls
that need it fail cleanly. Candidates beyond the fork/output budget
decode in successive chunks.

Wire contract and scores are unchanged: per-token logprobs are stitched
from the shared region and the forked tails. Verified bitwise
deterministic call-to-call and independent of candidate order (no
cross-fork leakage via equal-length candidate swap); ranking matches the
per-candidate implementation on the drone battery (winner softmax
0.99996 vs 0.99997), and >16-candidate chunking, prefix-of-another and
empty candidates all pass.

Measured on a desktop CPU: warm /api/score calls 0.52s -> 0.23s; warm
realtime classifier turns 196-303ms. The 9-candidate drone turn decodes
~17 unique tail tokens in one batch instead of nine sequential ~220ms
checkpoint-restore tasks.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* fix(realtime): gate scoring capacity by model usecase

Reserve llama.cpp scoring slots only for models that explicitly declare the score usecase, while allowing score to coexist with chat and completion. Reject incompatible unified-KV settings and classifier activation on models without scoring capacity.

Propagate application defaults when resolving realtime and preload pipeline stages so unset thread counts are resolved consistently without overriding explicit model settings.

Assisted-by: Codex:gpt-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* fix(ci): honor APT mirrors in the prebuilt llama-cpp compile step

The builder-prebuilt path installs gcc-14 with apt directly and ignored
the APT_MIRROR/APT_PORTS_MIRROR build args the from-source path already
honors, so an ubuntu mirror outage broke every arm64 backend build. Pass
the args into the stage and run apt-mirror.sh (already in the build
context via COPY . /LocalAI) before the apt step.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* feat(realtime): classifier argument slots via constrained completion

Hybrid classify-then-complete: a classifier option's canned tool call can
declare typed argument slots (number | enum | string, with defaults and
prompt hints) referenced as "{{name}}" in the arguments template. When
the option wins, the slots are filled by a short grammar-constrained
completion that continues the exact scoring prompt — rendered by the same
cached ScoreClassifier, so the llama.cpp prompt cache is already warm —
with the chosen route JSON re-opened at the first slot field. A GBNF
grammar pins the field skeleton and frees only the values; temperature 0,
a couple dozen tokens at most (~300ms on a desktop CPU for two slots).

Slot declarations and hints ride the option descriptions in the shared
system prompt, informing scoring and the fill alike at no per-turn token
cost. The localai.classifier.result event carries the final arguments and
a fill_latency_ms. On inference failure the slots' defaults apply; a slot
without a default fails the response (or falls through with
fallback.mode: generate). Slot filling requires completion alongside
score in the scoring model's known_usecases.

Verified end-to-end on the Pi drone demo: "fly forward three meters" in
distance mode classifies forward and infers {"distance": 3, "units":
"meters"} in ~310ms, and the drone flies exactly 3 units.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* feat(realtime): splice filled slot values into classifier replies

A classifier option's spoken reply can now reference its tool's argument
slots ("Going forward {{distance}} {{units}}."): the values inferred by
the slot-fill completion — or the recovery defaults — are spliced into
the reply as plain text before it is emitted, so what the assistant says
confirms what it actually inferred. Placeholders without a value stay
literal, and options without slots are untouched.

FillToolArguments now returns the raw slot values alongside the spliced
arguments JSON to make the reply templating possible.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* fix(realtime): harden classifier slot completion

Reserve context for constrained slot filling, size completions from their encoded output, and encode enum grammar literals as valid JSON. Reject empty enum values and cover the failure modes with regression tests.

Assisted-by: Codex:gpt-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* feat(realtime): prewarm the classifier scoring prompt on registration

Swapping a session's classifier option list (a voice-switched command
mode, for instance) made the next turns pay a full re-prefill of the new
option-list prompt — measured 2.4s vs 0.3s warm on a desktop CPU, and
worse: on hybrid-memory models like LFM2.5, whose state cannot be
partially rewound (llama.cpp can only restore checkpoints), *every*
probe change re-prefilled from scratch whenever the last checkpoint
missed the probe boundary, so even same-list turns intermittently cost
full prefills.

Registering an option list (pipeline seed or session.update) now fires a
best-effort background prewarm: two throwaway scores with distinct
probes. The first prefills the new option-list prompt; the second,
diverging exactly where per-turn probe text starts, plants the backend's
rewind point (KV checkpoint) at the stable-prefix boundary that every
real turn reuses. The prewarm hides behind the canned mode-switch reply
— by the time it finishes speaking, the cache is warm. Idempotent per
option set, detached from the registering request's lifetime.

Measured on the drone demo (LFM2.5-1.2B, desktop CPU): first turn after
a mode switch 2374ms -> 340ms; intermittent same-list full prefills
(1.3-2.1s) all -> under 0.5s. For clients that swap lists frequently,
options: [parallel:2] on the scoring model additionally keeps one slot
per list via prefix-similarity routing (+26MB RSS, unified KV).

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* perf(llama-cpp): checkpoint scoring at the caller-declared stable prefix

Hybrid-memory models (LFM2.5 shortconv, Qwen3.5 deltanet — where new
small models are headed) cannot rewind their state, so any prompt-cache
reuse that needs a rewind falls back to a full re-prefill. For classifier
scoring that meant every probe change re-processed the whole option-list
prompt: the server's checkpoints were placed reactively (at wherever the
previous task happened to diverge), so a checkpoint past the next
divergence was erased rather than restored — measured as intermittent
2-10s turns on prompts with a 95%+ common prefix.

The classifier now computes the probe-invariant prompt prefix once (the
byte-wise common prefix of two synthetic probe renders) and declares its
length with every Score request; the server maps it to a token boundary
and forces a KV checkpoint exactly there on each score prefill. That
checkpoint sits at or before every future divergence under the same
option list, so it always survives and always restores — repeat scoring
costs probe+candidates regardless of how the probe changes.

Also:
- prewarm reruns on every option-list registration instead of memoizing
  per list: with boundary checkpoints a redundant rewarm costs two
  probe-sized decodes, while skipping one after a slot eviction (three
  lists sharing fewer slots evict in LRU cascades) silently moves a full
  re-prefill onto the user's next turn
- new llama.cpp backend option rs_seq:N exposes bounded recurrent-state
  rollback outside speculative decoding; measured impractical for
  deltanet-scale states (65GB for 64 snapshots on Qwen3.5-4B) but cheap
  insurance for small-state models
- docs: the multi-list recipe (parallel:N + sps:0.5 — the default slot
  similarity threshold funnels distinct lists onto one slot)

Measured on the drone demo (LFM2.5-1.2B scorer, desktop CPU), steady
state: every turn 285-421ms including mode switches, vs 2.4s post-switch
and intermittent 1.3-2.9s re-prefills before.

Assisted-by: Claude:claude-fable-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* fix(realtime): align classifier cache guidance

Document the single-score prewarm behavior and clean the vendored score patch formatting.

Assisted-by: Codex:gpt-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* fix(llama-cpp): guard score task for fork backends

TurboQuant and Bonsai reuse the primary gRPC server against llama.cpp forks that do not carry LocalAI's slot-based Score patches. Compile the Score integration only for the patched primary backend and return UNIMPLEMENTED from fork builds instead of referencing absent task types and common_params fields.

Assisted-by: Codex:gpt-5 [gh]
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* fix(dev): generate gRPC code before commit lint

The coverage phase regenerates ignored protobuf bindings, but lint runs first and can fail against missing or stale output. Generate the pinned bindings before lint so the gate always type-checks the current schema.

Assisted-by: Codex:gpt-5
Signed-off-by: Richard Palethorpe <io@richiejp.com>

---------

Signed-off-by: Richard Palethorpe <io@richiejp.com>
2026-07-29 12:50:22 +02:00
localai-org-maint-bot
bc21f832aa gallery: add Laguna S 2.1 GGUF variants (#11188)
Add the official Q4_K_M and Q8_0 builds plus the DFlash speculative-decoding pairing for llama.cpp.

Assisted-by: Codex:gpt-5 [Hugging Face API]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-29 10:27:09 +02:00
mudler's LocalAI [bot]
967becb365 chore: ⬆️ Update ikawrakow/ik_llama.cpp to b054a8b983827c01aec59d4dc273a27c492c51c4 (#11175)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-29 09:57:23 +02:00
mudler's LocalAI [bot]
0569bb30a2 chore: ⬆️ Update ggml-org/whisper.cpp to 97c56f1dc1d1100a9d859c865a20c82d22f823ed (#11182)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-29 09:57:12 +02:00
mudler's LocalAI [bot]
550545c03a chore: ⬆️ Update mudler/parakeet.cpp to e747acdaee69b916cef62263ae5f718bda9ff3f3 (#11181)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-29 09:57:01 +02:00
mudler's LocalAI [bot]
c5e5141010 fix(ci): unbreak the sglang and darwin nemo backend builds (#11168)
* fix(sglang): keep nvidia-modelopt on a stable release

Every cublas sglang image currently fails to build:

  Failed to build `nvidia-modelopt==0.46.0rc0`
  Call to `wheel_stub.buildapi.build_wheel` failed
  ModuleNotFoundError: No module named 'wheel_stub'

sglang[all] pulls nvidia-modelopt in through its `diffusion` extra with no
version bound of its own, and install.sh adds a GLOBAL --prerelease=allow so
that flash-attn-4, which only ships 4.0.0b* wheels, can resolve. Unbounded plus
prereleases-allowed picks 0.46.0rc0, whose build backend imports wheel_stub
without declaring it in build-system.requires. EXTRA_PIP_INSTALL_FLAGS also
starts with --no-build-isolation, so nothing installs wheel_stub and the build
dies. Latest stable is 0.45.0 and resolves cleanly.

Bounding this one package rather than dropping the global flag, because the
flag is load-bearing for flash-attn-4 and this is the narrower change with the
smaller blast radius. Raise the bound when 0.46.0 final ships.

This is invisible on master because the backend build is path-filtered: sglang
is only rebuilt when sglang changes. It surfaces on any PR touching a shared
build input such as backend/backend.proto, which rebuilds the whole matrix.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(nemo): build the darwin venv on Python 3.12

The darwin nemo image fails to build:

  ModuleNotFoundError: No module named 'maturin'

nemo_toolkit pulls in text2num, a Rust extension built with maturin, whose
macOS arm64 wheels start at cp311: 3.0.2 publishes cp311, cp312, cp313 and
cp314 and no cp310. libbackend.sh defaults PYTHON_VERSION to 3.10, so pip finds
no wheel, falls back to the sdist, and dies in the PEP 517 hook because
EXTRA_PIP_INSTALL_FLAGS carries --no-build-isolation and nothing installs the
build backend. Taking the prebuilt wheel avoids the source build entirely, so
the runner needs no Rust toolchain.

Darwin only, deliberately: the Linux profiles resolve a cp310 manylinux wheel
for the same package and have no reason to move. The override is set after
libbackend.sh is sourced and before installRequirements, the same shape
sglang's install.sh already uses for its l4t13 profile.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(nemo): pin the darwin portable-Python patch level too

The 3.12 bump alone traded one failure for another:

  curl: (56) The requested URL returned error: 404
  make[1]: *** [nemo-asr] Error 56

libbackend builds the portable-Python URL from
cpython-${PYTHON_VERSION}.${PYTHON_PATCH}+${PY_STANDALONE_TAG}, and
PYTHON_PATCH defaults to 18 because the default interpreter is 3.10.18. Setting
only PYTHON_VERSION asked for a 3.12.18 that was never released.

Patch 11, not the 12 that sglang/install.sh pairs with 3.12 for l4t13: at the
20250818 tag python-build-standalone published 3.12.12 for linux aarch64 but
not for aarch64-apple-darwin, where 3.12.11 is the newest. Both URLs were
checked against the release assets rather than assumed to match.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-29 09:56:50 +02:00
mudler's LocalAI [bot]
47c48e9409 chore: ⬆️ Update antirez/ds4 to 54b36ed9ba42da31b24f2d1a5feb075c2475dbb1 (#11178)
⬆️ Update antirez/ds4

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-29 08:45:31 +02:00
mudler's LocalAI [bot]
6aaf7db5f0 chore: ⬆️ Update mudler/depth-anything.cpp to 2028b47ac75a8659c6a9aa617baf09be193eb55f (#11179)
⬆️ Update mudler/depth-anything.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-29 08:45:18 +02:00
mudler's LocalAI [bot]
193d49001b chore: ⬆️ Update CrispStrobe/CrispASR to 754b67289cf1137e3ed722885705f94132fc614f (#11180)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
2026-07-29 08:44:45 +02:00
localai-org-maint-bot
8f9184fbb2 feat(cli): support systemd socket activation (#11169)
* feat(cli): support systemd socket activation

Serve the API from a single stream listener inherited through the systemd activation protocol while retaining the existing address bind path when no listener is provided. Validate activation metadata, preserve the public-bind safety check, and document an on-demand systemd setup.

Assisted-by: Codex:gpt-5

* fix(cli): satisfy listener cleanup lint

Make the best-effort close explicit so errcheck accepts the deferred systemd listener cleanup.

Assisted-by: Codex:gpt-5 [golangci-lint]

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-29 01:44:24 +00:00
mudler's LocalAI [bot]
84972cb745 chore(model-gallery): ⬆️ update checksum (#11176)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-29 00:11:10 +02:00
localai-org-maint-bot
034df6ceb1 fix(worker): report RAM alongside GPU memory (#11167)
* fix(worker): report RAM alongside GPU memory

Assisted-by: Codex:gpt-5

* feat(ui): show worker RAM on node views

Assisted-by: Codex:gpt-5

---------

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-28 23:55:35 +02:00
localai-org-maint-bot
996bdcecdc fix(mlx-vlm): install torch dependencies on Metal (#11164)
Some Transformers processors used by MLX-VLM, including Qwen vision models, import both PyTorch and Torchvision. Include them in the Metal backend environment so model loading does not fail with missing-library errors.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-28 19:45:07 +02:00
walcz-de
4b4faa4ac7 feat(cloud-proxy): optional Anthropic prompt-cache breakpoints in translate mode (#11158)
The Anthropic translate provider builds the upstream request from scratch and
never emitted cache_control, so prompt caching was impossible for OpenAI-format
clients routed through cloud-proxy — even though the entire system prompt + tools
prefix is re-sent on every agentic turn.

Add an opt-in cache_prompt flag (ProxyOptions.cache_prompt; model YAML
proxy.cache_prompt: true). On a translate+anthropic model, buildAnthropicRequest
injects cache_control:{type:ephemeral} on the stable prefix — the system block,
the last tool, and the last message block (at most 3 of Anthropic's 4 allowed
breakpoints). Anthropic then serves the repeated prefix at the cache-read rate
(0.1x input) on subsequent calls, cutting cost on multi-turn/agentic workloads.
No effect in passthrough mode, for non-Anthropic providers, or when unset.

System is widened to any so it can carry the block form required to attach
cache_control, while still marshalling as a bare string when caching is off.
Adds a unit test asserting exactly three breakpoints when on and none when off,
and documents the option in docs/content/operations/cloud-proxy.md.

Assisted-by: Claude:opus-4.8

Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>
2026-07-28 17:38:14 +00:00
mudler's LocalAI [bot]
0f7186f214 feat(ui): replace the stacked operations bar with a one-line strip and an Activity page (#11163)
* feat(ui): record finished gallery operations in a bounded history ring

The operations panel drops an operation the moment it succeeds, so a user
who steps away cannot tell whether an install finished, failed or was never
started. OpCache now keeps the last 50 terminal operations, recorded from
the point where an op leaves the cache.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(ui): pin the history ring's dedupe, outcome order and start stamp

Review of the history ring found four gaps. The dedupe guard and the
bounded seen set were unreachable through the exported API and so had no
coverage; an in-package spec file now drives opHistory directly. The
outcome switch claimed an ordering was load bearing that nothing pinned,
so an errored op that never reached Processed now has a spec.

Two behaviour fixes come with it. StartedAt was the zero time for ops
recovered from the store or replicated from a peer, since neither path
stamps a start time, which would have rendered as a two-millennia
duration; it now falls back to the finish time. Reusing a cache key with
a fresh job ID orphaned the previous stamp, so Set and SetBackend now
drop it.

The comment on the outcome switch described a state the code cannot be
in: CancelOperation sets Cancelled and Processed synchronously before the
handler removes the entry, so status.Cancelled already covers the cancel
endpoint. The !Processed clause stays for the dismiss endpoint firing on
an in-flight op, and the comments now say so.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): record operations that end on a peer replica

The NATS end event is the only signal a replica gets for an install another
replica ran. Record from applyEnd too, deduped by job ID so the originating
replica does not record its own broadcast twice.

Three start-stamp defects in the same path go with it. applyEnd now drops the
stamp unconditionally, since recordTerminal only cleans up on the path where it
found a cache key and an end event can overtake the local Set. applyStart drops
the stamp of the job whose cache key it replaces, which a peer-driven retry
previously stranded. And recordTerminal reads the stamp once instead of testing
Exists and then reading, so a concurrent record for the same job can no longer
delete the stamp between the two and let the zero time overwrite the
finish-time fallback, which the Activity page would render as a two-millennia
run.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): do not guess the outcome of a peer operation with no local status

A replica that restarts mid-operation hydrates its OpCache keys from
PostgreSQL, but gallery statuses are in-memory only and come back empty. The
end broadcast then landed on recordTerminal's nil-status branch, which reads a
missing status as queued-and-removed and filed a successful install as
cancelled. That reading is right locally and wrong on the peer path, where a
missing status means the outcome was never held here.

recordTerminal now takes the source of the terminal event and records nothing
when the peer path finds no status, restoring what the replica did before the
end event started recording. The local path is unchanged.

Also move the ApplyEndForTest seam to the conventional export_test.go, and stop
the dedupe spec from claiming to guard the ring's seen set: the local delete
removes the status keys, so the broadcast that follows returns before reaching
it. An in-package spec that calls recordTerminal twice does the pinning.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(api): add GET and DELETE /api/operations/history

Admin gated like the rest of the operations API. The live /api/operations
payload is unchanged so the one second poll stays small.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): expose operation history through OperationsContext

Fetched on demand and when the live list shrinks, never on the one second
poll interval.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): detect operation departure by identity and ignore committing ops in the ETA gate

Refetching history on a shrinking live count missed a completion that
coincided with a start, which is the common case during a batch install.
Track the live job IDs instead, so any departure triggers the refetch
regardless of how the count moved.

An operation that has finished downloading stays live at
currentBytes == totalBytes for the whole commit and install phase and can
never produce an estimate, so counting it in the all-or-nothing gate blanked
every other operation's time remaining for as long as it lasted. Only
operations still moving bytes get a vote.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): let only downloading operations gate the time remaining estimate

Verifying pins an operation flat below its total for the whole sha256 pass:
the AfterDownload hook reports completedBytes plus the finished file against
a total summed over every file, then hashes synchronously without emitting
progress. Files download sequentially, so a 15 shard model enters that
window 14 times, and a byte comparison cannot see it because the counter is
genuinely below the total throughout.

Gating on phase closes resolving, verifying, committing and persisting in
one predicate, so a quiet neighbour no longer blanks every other
operation's estimate for minutes at a time. The byte clauses stay: a
producer can report downloading with bytes already at the total.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): collapse the operations bar to a single line

Four concurrent installs used to take four rows above every page. The strip
now shows one operation, failure first, with a counter linking to Activity.
The close button hides the strip and no longer cancels an install.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): keep the operations strip from widening the page and from muting a failure

A long install error made the strip report a 1600px minimum width, which sized
main-content to fit and gave every page under it a horizontal scrollbar.
Inline-size containment plus shrinkable detail and bytes cells keep it inside
the viewport.

Hiding is no longer able to swallow the hidden job's own failure, a completed
removal or staging says so instead of claiming an install, a cancelling
operation renders as cancelling, and the live region no longer covers the
per-second percentage.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): shrink main-content instead of containing the strip, and expose progress

min-width on .main-content is what actually lets a long install error shrink,
and unlike inline-size containment it has no browser support floor and no
latent collapse if the strip ever lands in a shrink-to-fit context. It matches
what .app-layout-chat .main-content already does, and it clears pre-existing
horizontal overflow on narrow viewports as a side effect.

The progress track is now a labelled progressbar, so assistive tech can read
the value on demand rather than losing it to the aria-hidden that stopped the
live region re-announcing every poll.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): add the live operation card for the Activity page

Carries the detail the one-line strip has to drop: phase, bytes, the per-node
breakdown for cluster installs, and a labelled Cancel button. Cancelling is
destructive, so it gets a labelled button rather than a glyph.

A cancelling operation drops its progress bar and its time estimate, the same
call the strip makes: a percentage still climbing under "Cancelling" reads as
the cancel not having taken.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): give the operation card a verb, live node disclosure and its per-node detail

The card carried no verb, so an install, a removal and a staging op rendered as
spinner plus name plus kind tag and were indistinguishable. It now runs the same
verb and icon chain as the one-line strip, which is what stops the page that is
meant to carry more detail from carrying less.

The auto-expand default was evaluated once at mount. An operation is listed as
soon as it is admitted but its nodes are filled in only when the fan-out starts
reporting, so a card mounted at creation latched on the empty list and stayed
collapsed. The default is a live expression now, and state holds only an
explicit choice.

Also: an optional onRetry gates a Retry button, so the page can own the install
reconstruction without the card ever showing a control with nothing behind it;
the disclosure moved above the region it controls and gained aria-controls; the
toggle is gated at more than one node so the count is never "1 nodes"; an
unmapped node status is passed through instead of being relabelled "Queued";
error text is clamped with the full string in the title; and file_name plus the
per-node progress bar are rendered again, reviving three CSS rules that had gone
dead along with the detail they styled.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): add the Activity page

Live operations, unacknowledged failures and the record of what finished, at
/app/activity in the Operate console. Cancelling an install now lives here
behind a labelled button rather than on the strip, and a failed install can be
retried: the retry dismisses the failure first so it still reaches the record,
then reissues the model, backend or node-scoped backend install.

The sidebar Operate entry carries the operation count. The console rail is only
rendered on an Operate route and can be collapsed, so a badge there could
vanish while operations were still running.

Two follow-ups from review fold in here: a failed removal or staging job no
longer reports a failed install on either the card or the strip, and the card's
error text can shrink so one unbroken token cannot widen the card.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): dismiss operations by job, and stop the Activity page contradicting itself

Dismissing resolved the job by display id, but /api/operations strips the
"node:<nodeID>:" prefix before emitting, so a local install and a node-scoped
install of one backend arrive as two jobs sharing one id. Dismissing by id
retired whichever came first. That defeated the guarantee retry was built
around: with the wrong job dismissed, the reinstall overwrote the acted-on
failure's opcache entry in place, bypassing recordTerminal, while an unrelated
failure vanished from Needs attention. dismissFailedOp, the card's dismiss
control and the strip now all pass the jobID, which is what the endpoint takes.

A filter matching nothing rendered the "nothing has ever run" empty state while
the header counted the records the filter had hidden. The empty state is now
gated on the All chip and a narrowed view gets its own message plus a way back;
the header counts the instance rather than the chip, so selecting Backends no
longer reports "Nothing running" over running model installs.

Also: the summary drops a zero clause instead of rendering "0 needs attention"
on the happy path and pluralises both counts; a record duration is floored at
"< 1s" and rejected above a day, so a zero-value start stamp cannot render a
span of millennia and a zero span cannot render "installed in" with nothing
after it; a deletion cancelled mid-flight reports the cancellation rather than
claiming it was removed; and the retry variant comment names the fix instead of
calling the gap closed.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: document the Activity page and the operations history endpoints

Adds an Activity page under Operations covering the one-line operations
strip, the /app/activity sections and filters, per-operation cancel,
retry and dismiss, the in-memory 50-entry record, and the sidebar count.
Documents GET and DELETE /api/operations/history, and fills the gap in
the admin-only endpoint list, which also omitted the pre-existing
POST /api/operations/:jobID/dismiss.

Corrects the distributed-mode install-watching section: the per-node
breakdown now lives on the Activity page rather than on the strip, which
rolls a fan-out up into a single phrase.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: correct nine details in the Activity page documentation

The operations strip never renders a file name: its detail line is the
error, the node roll-up, the target node, the phase or the queued note.
Drops the stale clause in the distributed-mode section, where the
per-node bullet is now the only place a file name is described.

Scopes the phase vocabulary to artifact-backed gallery models, since a
plain GGUF install emits no phase. Corrects the per-node list: the
toggle exists for any fan-out of two or more workers and the four-node
threshold only governs whether it starts open, while the N nodes tag
needs more than one node. Notes that a cancelled operation can sit in
the live section reading Cancelling, that cluster staging never reaches
the record, and that Clear history appears only when the record has
something in it.

Names the operations response envelope, with a JSON example, so callers
do not index a bare array, and stops describing the icon-only dismiss
control as a labelled button.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: drop the unreachable Cancelling state and scope the byte claims

An operation can only report isCancelled while it is unprocessed, but
every writer of Cancelled sets Processed in the same breath, on the peer
path as much as the local one, and the cache evicts cancelled entries
before the handler sees them. The state cannot reach the page, so the
live section is described again as running or queued operations.

Byte counts come from the artifact bridge alone, the same producer as
the phase, so a plain GGUF install, a removal and a backend install
report none. Scopes both to artifact-backed gallery models and leaves
the verb, the name and the percentage as what every operation shows. A
worker backend install reports its bytes through fields the operations
payload does not carry, so the distributed section now describes the
percentage and the node roll-up, with per-file counts pointed at the
per-node detail.

Also: staging jobs carry no error, so they never reach Needs attention
and Retry never had a staging case to exclude; an install that involves
workers is no longer called node-scoped, which this page uses for
node-targeted installs; and the record timestamps carry nanoseconds.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: state only the verb and the name as unconditional on the strip

The percentage is as conditional as the bytes were: it renders only for
a running operation that has reported progress, so a queued operation, a
failed one and a removal never carry it. A removal in particular sits at
progress zero for its whole visible life, since the delete path reports
none and its completion is filtered out. Both the strip and the card
paragraphs now lead with what always shows and list the rest as
conditions.

The Cluster chip matches on a node list that finished operations do not
carry, so a fan-out install leaves the chip once it reaches the record.
Scoped that claim to the live sections.

Two more of the same shape, found by re-reading each clause alone: the
strip also appears for a failure, which is not running, and the
four-second hold only applies when nothing replaces the operation that
just finished.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): stop reporting a cancelled install as installed, and make queued real

Three defects that all trace to one root cause: `isCancelled: true` is
unreachable from /api/operations. Every writer of Cancelled=true also sets
Processed=true, the handler skips Processed && Cancelled, and OpCache.GetStatus
evicts a cancelled op before the handler iterates it.

Cancelling the last running operation put a green "Installed model X" on the
strip for four seconds: the completion hold was guarded by
`!previous.isCancelled`, which is dead. A cancellation deletes the operation
server side, so the strip sees exactly what it sees on a completion, and
nothing in the payload separates the two. The signal now comes from the side
that issued the cancel: the operations context remembers the job IDs it
cancelled (pruned after a minute) and the strip asks before it holds anything.
A cancelled operation goes as soon as it stops; the record already reports it
as cancelled.

isQueued was set only when the gallery status was missing, but markQueued
publishes a "queued" status at admission, so a queued op has a status for its
whole queued life and the state was unreachable outside a microsecond window.
Every operation waiting behind a running install rendered as "Installing model
X" with a spinner. The queued phase is now the signal, via an exported
PhaseQueued and a nil-safe OpStatus.IsQueued() next to the writer.

With those two fixed, the Cancelling state has no way to be entered: cancelling
is instantaneous from the API's point of view. Its branches, CSS, locale key
and the isCancelled field itself are removed rather than left for a future
reader to assume they work.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): keep a removal a removal, and say what an install is doing

OpStatus.Deletion was set once, at admission, and lost on the next status
write: UpdateStatus replaces the whole status and only carried Nodes
forward. Every later writer (the worker's first write, the progress
ticks, the failure path) leaves the field at its zero value, so the flag
survived only the queued window, and both surfaces test isQueued first.

The reachable consequence is that a failed removal reported itself as a
failed install, which is exactly the shape the Activity page offers Retry
for, and Retry installs: pressing it on a removal that failed
re-downloaded the model. A running delete also rendered as "Installing
model X" with a spinner, and a successful one as "Installed model X".

Carry Deletion forward the way Nodes already is. A job is a delete or an
install for its whole life; an unset flag means "no new information", not
"this is an install". Pinned by Go specs on both the service and
/api/operations: the existing Playwright specs were green only because
they stubbed a payload the server could not emit.

Also restore the operation's own status message on the Activity card.
Phases and byte counters exist only on the managed-artifact path, so a
legacy files: gallery model and every backend install rendered a sub-row
with nothing in it but the verb. The strip stays terse on purpose.

And give the strip's name a min-width floor: overflow: hidden zeroes its
automatic minimum, so a long error squeezed the name down to "mod…" and
the identity of the thing that broke was the first thing lost.

primaryOperation is made module-private: its comment claimed the Activity
page selected the same operation, but that page shows all of them,
partitioned into failed and running, and never imported it.

Assisted-by: Claude Code:Opus 5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(activity): read the operations record from PostgreSQL

The Activity page's record of finished installs and removals was a 50-entry
in-memory ring per frontend replica. In distributed mode that is the wrong
place for it: each replica keeps its own copy, a replica added by a scale-out
or a rolling deploy starts empty and never backfills, and "Clear history"
clears only the replica that served the request, so the record reappears on
the next poll routed elsewhere.

The data is already in gallery_operations. Read it from there.

GalleryStore gains ListTerminal and ClearTerminal, sharing a lifted
terminalStatuses set with CleanOld so there is one definition of "finished".
ListTerminal orders by updated_at, when the operation reached its terminal
status, because the record reports what finished and when.

OpCache.History and ClearHistory dispatch on whether a store is wired, so the
HTTP handlers and the OpRecord JSON shape are unchanged and the page needed no
change. A failed store read falls back to the local ring rather than blanking
the page, and ClearHistory empties the ring as well so a database blip cannot
resurrect a record the admin just cleared.

The name derivation in recordTerminal is lifted into operationDisplayName and
used by both paths, so the ring and the store cannot name the same operation
differently.

Also fixes a pre-existing bug the store path made visible: the backend channel
hardcoded op_type "backend_install" even for a removal, while the model channel
derives model_install/model_delete from op.Delete. Both channels carry the same
ManagementOp, whose Delete field the backend handler already branches on, so
the backend channel now derives backend_delete the same way.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(activity): keep a cancelled operation cancelled, and report a failed clear

Review follow-up on the store-backed Activity record.

A cancelled install was recorded as a failure. The cancel handler persists
"cancelled" synchronously, then the handler goroutine unwinds with the context
error and Start hands that to updateError unconditionally, which overwrote the
row with "failed: context canceled". The page rendered a cancelled install as a
red failure card offering Retry, with a raw context error as the reason.

Fixed in GalleryStore rather than in Start, because an operation finishes once
and the paths that retire one are not mutually exclusive: UpdateStatus now
refuses to rewrite a row that already reached a terminal status. That also pins
updated_at to when the operation really finished, which is the key the record is
ordered by, and Create's upsert now freezes the same columns so a worker
dequeuing an operation the admin cancelled while it was queued cannot reopen it
as pending.

ClearHistory returned nothing, so a failed delete logged a warning while the
handler still answered 200. The admin watched the record clear and come back on
the next fetch with nothing said about why. It now returns the error, the DELETE
handler answers 500, and the store is cleared before the local ring so a failure
leaves the fallback record intact rather than faking an empty one.

Hydrate is the only reader that decides from op_type whether an operation is a
removal, and it tested for "model_delete" exactly, so the backend_delete added
in the previous commit hydrated as an install: a replica restarting during a
backend removal rendered "Installing backend X". Both discriminations now go
through IsDeleteOpType/IsBackendOpType so a fifth op_type cannot silently read
as an install in whichever consumer was missed.

Also: the backend channel now persists Cancellable as !op.Delete, matching the
model channel; IsBackend falls back to the op_type prefix, since is_backend_op
is only written by UpsertCacheKey and the rows needing the name fallback were
reporting backend operations as models; and an unrecognized terminal status is
logged rather than quietly filed as a success, which is what the comment already
claimed.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(activity): keep a reaped operation correctable by its real outcome

The terminal-status freeze added in the previous commit was too wide. It froze
"failed" alongside "completed" and "cancelled", and the stale reaper writes
"failed" onto operations that are still going to run.

The gallery worker is a single goroutine consuming both channels serially, so
an operation queued behind a large download sits in "pending" with nothing
bumping updated_at, and ReapStaleOperations gives up on it after 30 minutes.
That used to be self-healing: the worker dequeued it, Create reset the row to
"pending", and the operation reported its real outcome. With the freeze the row
stayed "failed" forever while the install ran and succeeded underneath it: a red
failure card offering Retry for a model that is installed, omitted from
ListActive so no replica hydrates it, and no longer deduped cluster-wide by
FindDuplicate.

Freeze on ("completed", "cancelled") instead. That is all the cancelled-install
fix ever needed, and it leaves a failure correctable by what actually happened.
The set is separate from terminalStatuses, which ListTerminal, ClearTerminal and
CleanOld all still want in full, because the two mean different things: a
failure can be superseded by a real outcome, a completion or a cancellation is
the real outcome.

UpdateStatus now writes the error column unconditionally, so a corrected
outcome drops the previous attempt's reason rather than being recorded as
completed while still carrying "stale operation reaped" as its error.

Also adds the route-level spec for the 500 branch of DELETE
/api/operations/history, and trims a comment that credited the persisted
cancellable column with more than it survives long enough to do.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(activity): offer Cancel in the phase that can honour it

The cancellable flag was set at both ends of an operation's life and was wrong
at both, in opposite directions.

A queued operation is cancellable whatever it is. EnqueueModelOp and
EnqueueBackendOp select on the operation context, so cancelling one that is
still waiting releases the delivery goroutine and abandonQueued retires it: the
worker never sees it, nothing is downloaded, nothing is deleted. markQueued
nevertheless wrote Cancellable: !deletion, so a queued removal reported
cancellable: false and the UI hid the Cancel button in the one window where
pressing it both works and leaves no trace. A removal queued behind a large
install was stuck there until the install finished.

A running removal is not cancellable at all. DeleteModel and DeleteBackend take
no context, and modelHandler only checks the operation context after the call
returns, so a "cancelled" verdict would land after the model was already gone.
Both handlers nevertheless wrote Cancellable: true unconditionally at entry,
ahead of the op.Delete branch, offering a Cancel button the server cannot
honour.

So the queued phase is more cancellable than the running phase, which is the
reverse of the usual shape. markQueued now reports true unconditionally, and
the handler-entry writes report !op.Delete. Both sites carry a comment saying
why, because reading either one alone suggests the other is a bug.

GalleryStore.Create keeps !op.Delete: it runs at dequeue, so its value already
describes the running phase. Its comment now says so.

Specs cover queued removal, queued install, running removal and running install
through the handlers, plus the queued-removal case through /api/operations
where the flag is consumed, plus the behaviour the whole asymmetry rests on: a
removal cancelled while queued never reaches the worker and deletes nothing.
No existing spec asserted the old values.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Write] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(activity): clamp the installer message, and add a real-binary e2e spec

Running the page against a real local-ai showed the legacy installer message
wrapping to three lines and dominating the card: it embeds an absolute file
path, so it is both long and a single unbreakable token. One line, ellipsised,
full text in the title, matching what the error string already does.

The spec that found it runs with no route stubbing at all. Every other spec
here stubs /api/operations, which is how a payload the server cannot emit
(isDeletion true on a live operation) stayed green through a full review while
the UI rendered a removal as an install. It is skipped unless
LOCALAI_REAL_BINARY is set, so CI is unaffected.

Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-28 19:33:48 +02:00
localai-org-maint-bot
823fc25bb7 fix(kokoro): add CPU backend fallback (#11161)
Publish the existing Kokoro CPU profile for amd64 and arm64 and use it as the default gallery capability so Vulkan-only and CPU hosts can install the backend.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-28 18:02:39 +02:00
mudler's LocalAI [bot]
366db11c59 chore: ⬆️ Update mudler/parakeet.cpp to 3e1ddd8455ceb9bfae564f84db24ba068b00c56e (#11150)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-28 09:25:52 +02:00
mudler's LocalAI [bot]
54012002fd chore: ⬆️ Update ikawrakow/ik_llama.cpp to 5f063b7bbae8f9a34dfc5c704aa77939e76494a9 (#11153)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-28 09:25:39 +02:00
mudler's LocalAI [bot]
176000190e chore: ⬆️ Update ggml-org/llama.cpp to 1cbfd1988311775425d36c0ce066590f7d3049cf (#11155)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-28 09:25:25 +02:00
mudler's LocalAI [bot]
c4d0c060ed chore: ⬆️ Update CrispStrobe/CrispASR to 7bb8be77a8c1677e32bba58514bb2d42f29a7a48 (#11156)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-28 09:01:41 +02:00
mudler's LocalAI [bot]
cff31bbac0 chore: ⬆️ Update leejet/stable-diffusion.cpp to 22516991cbdf725e69b0b4a87e52ca16cce07c2d (#11157)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-28 08:23:55 +02:00
mudler's LocalAI [bot]
e218c7f56a chore(model-gallery): ⬆️ update checksum (#11152)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-27 23:36:07 +02:00
mudler's LocalAI [bot]
12f2e1b99c feat(swagger): update swagger (#11149)
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-27 23:35:42 +02:00
mudler's LocalAI [bot]
2ecf893c6c fix(sherpa-onnx): install cuDNN in the CUDA builder so the package can bundle it (#11145)
sherpa-onnx links onnxruntime's CUDA execution provider, and
libonnxruntime_providers_cuda.so carries cuDNN as a hard DT_NEEDED. The
onnxruntime GPU tarball ships no cuDNN of its own, and Dockerfile.golang
only installs libcudnn9 on the arm64 + CUDA 13 branch, so the amd64 CUDA
builders have none at all.

Since #10946 added the packaging guard, that combination is fatal rather
than silent: package-gpu-libs.sh reports 'cuDNN: venv=absent system=absent
-> bundle=detect', correctly detects the reference, finds nothing to copy
and refuses to emit the package. Both -gpu-nvidia-cuda-12-sherpa-onnx and
-gpu-nvidia-cuda-13-sherpa-onnx have failed to build since 2026-07-19, so
neither image has been published. Before the guard existed they shipped
without cuDNN and failed at load time instead.

Install the runtime package for this backend only. The auto-detection
bundles solely what a package references, so no other backend would grow,
but every Go CUDA builder would pay ~1.1 GB of layer and registry cache
for a library ggml never calls.


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-27 22:39:37 +02:00
mudler's LocalAI [bot]
05ff401de8 fix(test): stop the backend-trace specs racing the lossy trace channel (#11146)
RecordBackendTrace does a non-blocking send onto a 100-slot channel and
drops when it is full, so tracing never stalls inference. The payload
bounding specs pushed all 200 traces in one tight loop, which overruns
that channel on a loaded machine: entries are dropped for good and the
Eventually waiting for 200 can never be satisfied, no matter the timeout.
CI hit this on master at 0a8a7fbb, settling at 158/200.

Feed the traces in chunks of 50, draining after each, so the channel is
never overrun and the count stays exact. Reproduced with 60 busy loops on
a 20-core box at GOMAXPROCS=2: 0/12 runs passed before, 12/12 after.


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-27 22:38:23 +02:00
Tai An
a7fa678d83 fix(tts): forward the OpenAI speed field to the backend (#11097) (#11120)
* fix(tts): forward the OpenAI speed field to the backend (#11097)

/v1/audio/speech accepted the documented OpenAI `speed` field and then
dropped it: schema.TTSRequest had no Speed member, so the value never
reached proto.TTSRequest and the request returned 200 with an unchanged
playback rate.

Accept speed and normalise it into the existing per-request params map,
which core/backend forwards verbatim to the backend. An explicit
params["speed"] still wins, and a value outside the documented 0.25-4.0
range is now rejected with 400 instead of being silently ignored.

Signed-off-by: Anai-Guo <antai12232931@outlook.com>

* fix(tts): distinguish explicit speed=0 from an omitted field

Make TTSRequest.Speed a *float32 so an explicit `"speed": 0` (invalid,
below the documented 0.25 minimum) is rejected with 400 instead of being
treated as unset and silently defaulted. An omitted field stays nil and
leaves the backend default untouched.

Add a request-boundary regression that distinguishes an omitted speed from
an explicit zero, addressing review feedback.

Signed-off-by: Anai-Guo <antai12232931@outlook.com>

* docs: drop the speed field from the TTS docs

Per review: no backend consumes params.speed today, so documenting it
would be misleading. The API-level plumbing and validation stay.

Signed-off-by: Anai-Guo <antai12232931@outlook.com>

---------

Signed-off-by: Anai-Guo <antai12232931@outlook.com>
2026-07-27 19:03:32 +02:00
Anupam Mediratta
3698361510 fix: upgrade hono to 4.12.25 (CVE-2026-54290) (#11023)
* fix: CVE-2026-54290 security vulnerability

Automated dependency upgrade by OrbisAI Security

Signed-off-by: orbisai0security <mediratta@gmail.com>
Signed-off-by: Anupam Mediratta <mediratta@gmail.com>

* fix(deps): override hono transitive dep to eliminate CVE-2026-54290

Add package.json `overrides` field to force hono@4.12.25 across the
entire dependency graph, including the transitive copy pulled in by
@modelcontextprotocol/sdk. Previously bun.lock retained a scoped
`@modelcontextprotocol/sdk/hono` entry resolved to the vulnerable
hono@4.12.8; the override removes that entry so only the patched
version ships.

Assisted-by: Claude Code:claude-sonnet-4-6
Signed-off-by: Anupam Mediratta <mediratta@gmail.com>

---------

Signed-off-by: orbisai0security <mediratta@gmail.com>
Signed-off-by: Anupam Mediratta <mediratta@gmail.com>
2026-07-27 19:02:37 +02:00
mudler's LocalAI [bot]
856b0ea951 fix(ci): build the CUDA 13 image on Ubuntu 24.04 (#11143)
* fix(ci): build the CUDA 13 image on Ubuntu 24.04

The amd64 `-gpu-nvidia-cuda-13` image is the only runtime image still
built FROM ubuntu:22.04. The Ubuntu 24.04 migration (#7769) bumped its
`ubuntu-version` to 2404 but left `base-image` on jammy, so the image
ships glibc 2.35 while adding the noble CUDA apt repository, and every
backend it unpacks is built on noble.

Backends therefore cannot dlopen the libraries they bundle. The vLLM
backend dies at import time with:

  OSError: /lib/x86_64-linux-gnu/libc.so.6: version `GLIBC_2.38' not
  found (required by /backends/cuda13-vllm/lib/libnuma.so.1)

and torchcodec finds no usable libavutil because jammy ships ffmpeg 4.x
(libavutil.so.56) while torchcodec looks for .so.57 through .so.60.

Add a spec over the build matrices that fails when a base image and the
`ubuntu-version`/`ubuntu-codename` it is paired with disagree, or when
the runtime images are split across Ubuntu releases. Entries whose base
image does not name a release (JetPack) are left alone.

The `base-grpc-cuda-13-amd64` builder base stays on jammy: it only
compiles backends, and a lower glibc floor in a builder is safe.

Fixes #11059

Assisted-by: Claude:claude-opus-5 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ci): drop the build matrix invariant spec

Per review, the CI matrix guard does not belong in the tree. Only the
base image bump remains.

Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-27 18:49:02 +02:00
mudler's LocalAI [bot]
878a0d00a1 fix(distributed): reaper reaps live backends, ghost model stubs, in_flight leak, sidecar staging runaway (#11142)
* fix(distributed): stop the probe reaper from orphaning busy backends

The reconciler's liveness probe is a 1s gRPC HealthCheck, and a single
failed probe deleted the model's node_models row. A backend that is
merely busy cannot answer it: single-threaded Python backends (video and
avatar generation) block for minutes inside one request, so the reaper
was deleting registry rows for backends that were alive and mid-request.

The model then vanished from the nodes page while it was still
generating, and because the row was gone the in-flight decrement had
nothing to decrement ("DecrementInFlight: no matching row or already
zero"). Every subsequent request re-routed and re-staged the full model
from scratch.

Two guards:

  - Replicas with in-flight requests are excluded in SQL. A row that is
    actively serving is proof of life, and the running request is
    exactly what stops the backend from answering the probe.

  - Idle replicas must miss three CONSECUTIVE probes before removal, so
    a transient blip cannot orphan a live replica. A successful probe
    resets the streak.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(distributed): drop the local model stub when its last replica goes

In distributed mode every routed model leaves an in-process stub in the
frontend's ModelLoader, and DistributedModelStore.Range reports local
stubs UNION the registry rows. Every registry removal path deletes only
the DB row, so the stub outlived the replica and the model was reported
as loaded forever.

That is the "loaded on the home page, absent from every node" ghost:
/system reads the union and still sees the stub, while /api/nodes/models
reads the registry and correctly sees nothing. It never self-healed,
and both frontend replicas showed it independently.

The replica-removed chokepoint could not fix this as it stood, because
it held a SINGLE hook that the prefix cache already owned, and it was
registered only when the prefix cache was enabled. Registering a second
listener would have silently displaced the first.

  - Turn replicaRemovedHook into a list (AddReplicaRemovedHook), so
    independent subsystems can each register without displacing others.
  - Add NewLocalStubInvalidator, which drops the local stub once no
    healthy replica of the model remains anywhere in the cluster, and
    wire it unconditionally in startup.

The stub is kept while another node still serves the model: the
frontend is right to consider it loaded, and each request re-routes
through SmartRouter to pick a live replica anyway.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(distributed): stop staging checksum sidecars back to workers

The file transfer server writes a "<file>.sha256" sidecar next to every
file it accepts. The sender walked the model directory with no filter,
so it staged those sidecars too, and the receiver duly wrote a sidecar
for each sidecar. Every staging pass multiplied the tree:

  config.json -> config.json.sha256 -> config.json.sha256.sha256 -> ...

One LongCat snapshot had grown to 498 files, 466 of them chained, up to
29 levels deep, and the staged file count climbed on every pass. This
inflates each transfer and grows disk without bound on both ends.

Skip hash sidecars in stageDirectory, and mirror the skip in
countStageableFiles so the progress bar still reaches 100%. The check is
"a sidecar sitting next to a real file" rather than a blanket suffix
ban, so a model that genuinely ships a .sha256 payload with no
corresponding base file is still transferred.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(distributed): classify the liveness probe instead of gating on in_flight

The previous commit excluded replicas with in-flight requests from the probe
reaper. That was the wrong guard, and could invert the bug it fixed.

in_flight has no decrement guarantee: track() balances its increment with a
defer, but a frontend killed mid-request never runs it, and the load-time
reservation is released only when the first inference completes. Nothing
resets a leaked counter. Gating the reaper on it therefore meant a leaked
counter would shield a genuinely dead replica from ever being reaped.

Nor was patience alone a fix: three misses at the default interval is ~90s of
silence, while the generation that triggered this blocks for 15+ minutes.

The real conflation was in the probe itself. A gRPC HealthCheck against the
backend's serving port measures "is it idle enough to answer", not "does the
process exist", and probeLoadedModels discarded the error that tells them
apart. Because the gRPC client is lazy, the status code is decisive:

  - DeadlineExceeded: transport fine, nothing serviced the RPC. Busy.
  - Unavailable: nothing is listening. Gone.

ModelProber now returns a ProbeOutcome, and only ProbeUnreachable counts
toward the reap threshold. ProbeBusy clears the streak: it is evidence of
life. A blackholed network reads as busy too, deliberately, since whole-node
failure is the health monitor's job.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* feat(distributed): reconcile replicas against worker-reported processes

Probing a backend's own serving port cannot distinguish "busy" from "gone"
without inferring it from an error code. The worker can answer directly: it
spawned the process, holds the handle, and its reply is not blocked by
whatever that backend is doing.

Adds a models.running request-reply subject. The worker answers out of its
in-memory process table, reporting each live process as (modelID,
replicaIndex, address) — the supervisor's process keys are `modelID#replica`,
which is isomorphic to a NodeModel row, so the reconciler can diff the two
directly.

reconcileNodeProcesses runs before the port probe and reaps rows for models
the worker is not running. Models the worker vouches for get updated_at
bumped, which takes them out of the port prober's stale set entirely: that is
what keeps a backend deep in a long generation away from the probe in the
first place, rather than relying on classifying its silence after the fact.

A worker that does not answer is skipped, not assumed empty. A messaging
failure says nothing about the processes, and assuming the worst would delete
a node's rows on a transient NATS blip; the port probe stays as the fallback
for those nodes. Rows younger than probeStaleAfter are ignored so a freshly
created row is never judged against a process table that has not caught up.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

* fix(distributed): stop in_flight leaking and pin replicas against eviction

A leaked in_flight counter is not cosmetic. FindLRUModel,
FindGlobalLRUModelWithZeroInFlight and the router's eviction query all require
in_flight = 0, so a replica whose counter never came back is pinned and its
VRAM is unreclaimable for the lifetime of the process.

Two halves.

The source: routing reserves in_flight = 1 at load time so a freshly loaded
replica is not evicted out from under the request that caused the load. That
reservation was released ONLY by the first inference completing, so a route
torn down before any inference ran (client disconnect, handler error, failure
between load and the backend call) stranded it. newRouteResult now wires the
reservation to a sync.Once fired by whichever comes first, the first inference
or route teardown, and replaces three copies of the old wiring.

The backstop: a sweeper for counters leaked by paths that cannot run a defer
at all, such as a frontend killed mid-request.

Identifying a leak by elapsed time alone is unsafe. IncrementInFlight stamps
last_used at request START and nothing moves it while the request runs, so a
long generation is indistinguishable from a leak by age, and resetting there
would expose a serving model to eviction. The probe supplies the missing bit:
a backend that answers a health check promptly is not inside a request,
because that is precisely what a busy one cannot do. Requiring the row to also
be idle for 30 minutes covers backends that serve in parallel and can answer
while working, since those keep last_used fresh through each new increment.

Two existing tests asserted the old behaviour ("No decrement on Release").
That assertion was the leak, so both now pin the release instead.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-27 18:36:31 +02:00
mudler's LocalAI [bot]
e9c2754fc2 chore(model-gallery): propose variant groupings for review (#11139)
chore(model-gallery): propose variant groupings

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-27 09:48:23 +02:00
mudler's LocalAI [bot]
0a8a7fbbb4 chore(llama-cpp): bump llama.cpp and adapt to the load-mode refactor (#11140)
Bump LLAMA_VERSION to 0d47ea7427463093e69128bf2c2f9cd06b3ee5b3 (73 commits
touching common/, src/ and tools/server/). Two upstream changes break the
backend:

* ggml-org/llama.cpp#20834 folded common_params::use_mmap / use_mlock /
  use_direct_io into a single `load_mode` enum. LocalAI still exposes the three
  as independent settings (`mmap`, `mmlock`, and the `direct_io` option), so
  params_parse folds them once all three have been read, keeping the precedence
  the separate booleans had: direct I/O bypasses the page cache, mlock implies
  mmap, everything off is a plain buffered read. turboquant and bonsai compile
  this same grpc-server.cpp against forks that predate the refactor, so
  prepare.sh probes the checkout for LLAMA_LOAD_MODE_MMAP and generates
  llama_compat.h with LOCALAI_LEGACY_LOAD_MODE set accordingly. Probing beats a
  per-fork build flag here because the fork flavor targets disagree on whether
  they forward CMAKE_ARGS or EXTRA_CMAKE_ARGS, and it heals itself once a fork
  rebases past the refactor.

* The MiniMax M3 patch no longer applies. Upstream merged the model half of
  llama.cpp#24523 (LLM_ARCH_MINIMAX_M3, src/models/minimax-m3.cpp, the gguf-py
  constants and conversion/minimax.py) but not the chat half, so the patch is
  re-cut to carry only the common/chat.cpp template detection and PEG parser,
  rebased onto the new pin and onto the thinking_end_tag -> thinking_end_tags
  rename. Dropping it wholesale (as #11008 did, reverted in #11136) would have
  silently regressed MiniMax M3 tool calling and thinking.

Verified with a CPU docker build of the backend plus LoadModel and Predict
against a real GGUF over gRPC in all four load modes.


Assisted-by: Claude:claude-opus-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-27 07:02:29 +00:00
Ettore Di Giacinto
d9f3007876 Revert "chore: ⬆️ Update ggml-org/llama.cpp to d2a818231effb12b7b20b80b3b8c7756a9a33a04" (#11136)
Revert "chore: ⬆️ Update ggml-org/llama.cpp to `d2a818231effb12b7b20b…"

This reverts commit 6e69dbd617.
2026-07-27 01:16:04 +02:00
mudler's LocalAI [bot]
6e69dbd617 chore: ⬆️ Update ggml-org/llama.cpp to d2a818231effb12b7b20b80b3b8c7756a9a33a04 (#11008)
* ⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(llama-cpp): drop upstreamed MiniMax M3 patch

The pinned llama.cpp revision already contains MiniMax M3 support, so the downstream patch rejects during backend preparation on every platform.

Assisted-by: Codex:gpt-5

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-27 01:15:06 +02:00
mudler's LocalAI [bot]
19ffc33011 chore: ⬆️ Update leejet/stable-diffusion.cpp to 2d0385ba85af358f7115dda608a63eafd9de7ffd (#11132)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-27 01:00:05 +02:00
mudler's LocalAI [bot]
76ce4f59b6 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260726174827 (#11133)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 23:14:03 +02:00
dependabot[bot]
4f883c86a9 chore(deps): bump postcss from 8.5.15 to 8.5.23 in /core/http/react-ui in the npm_and_yarn group across 1 directory (#11106)
chore(deps): bump postcss

Bumps the npm_and_yarn group with 1 update in the /core/http/react-ui directory: [postcss](https://github.com/postcss/postcss).


Updates `postcss` from 8.5.15 to 8.5.23
- [Release notes](https://github.com/postcss/postcss/releases)
- [Changelog](https://github.com/postcss/postcss/blob/main/CHANGELOG.md)
- [Commits](https://github.com/postcss/postcss/compare/8.5.15...8.5.23)

---
updated-dependencies:
- dependency-name: postcss
  dependency-version: 8.5.23
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-26 23:13:42 +02:00
mudler's LocalAI [bot]
c56373d772 chore(model-gallery): ⬆️ update checksum (#11134)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 23:10:11 +02:00
Tai An
53006bb8e1 fix(realtime): accept legacy 'modalities' alias for output_modalities (fixes #11103) (#11104)
* fix(realtime): accept legacy 'modalities' alias for output_modalities

OpenAI's Realtime *beta* used the field name `modalities`; the GA field is
`output_modalities`. LocalAI only binds `output_modalities`, so a client
sending the still-common beta field `modalities: ["text"]` has it silently
dropped by encoding/json and the session falls back to audio: TTS runs and the
client receives large response.output_audio.* frames even though it asked for
text-only.

Accept `modalities` as an alias on both session.update (RealtimeSession) and
response.create (ResponseCreateParams). The GA `output_modalities` wins when
both are present, so GA clients are unaffected. Applied at the two existing
resolution points via a small modalitiesWithAlias helper.

Fixes #11103

Signed-off-by: Anai-Guo <antai12232931@anaiguo.com>

* test(realtime): add JSON-boundary regression for modalities alias

Decode representative session.update and response.create payloads that
carry only the legacy beta `modalities` key and assert the effective
output modality resolves to text (not audio), reproducing the exact
expressions used in updateSession and triggerResponseAtTurn. This guards
against a wrong JSON tag or a missed call site letting encoding/json drop
the alias silently.

Also document output_modalities (and the accepted legacy modalities
alias) for text-only sessions in the realtime feature docs.

Signed-off-by: Tai An <antai12232931@outlook.com>

---------

Signed-off-by: Anai-Guo <antai12232931@anaiguo.com>
Signed-off-by: Tai An <antai12232931@outlook.com>
Co-authored-by: Anai-Guo <antai12232931@anaiguo.com>
2026-07-26 23:09:30 +02:00
mudler's LocalAI [bot]
62e3b8304e chore: ⬆️ Update CrispStrobe/CrispASR to 306faee45fab641d54f9f941f075de1e9c0d3278 (#11131)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 23:05:50 +02:00
mudler's LocalAI [bot]
2476509321 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 0a4e10c7fb65d2dd5a4afb78339c7d373a8cdfaa (#11128)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 23:05:22 +02:00
dependabot[bot]
ac9352ef54 chore(deps): bump actions/setup-node from 4 to 7 (#11080)
Bumps [actions/setup-node](https://github.com/actions/setup-node) from 4 to 7.
- [Release notes](https://github.com/actions/setup-node/releases)
- [Commits](https://github.com/actions/setup-node/compare/v4...v7)

---
updated-dependencies:
- dependency-name: actions/setup-node
  dependency-version: '7'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-26 23:05:07 +02:00
mudler's LocalAI [bot]
4baa36ddd8 feat(backend): vllm-cpp - text-generation backend for vllm.cpp with llama.cpp-parity tool calling (#11100)
* feat(backend): add vllm-cpp text-generation backend (vllm.cpp)

Wrap https://github.com/mudler/vllm.cpp - the LocalAI-team from-scratch C++20
port of vLLM (paged KV cache, continuous batching, prefix caching, safetensors
+ GGUF loading, no Python at inference) - as a Go gRPC backend over its stable
C ABI (ABI v2) via purego.

Backend (backend/go/vllm-cpp):
- Load -> vllm_engine_load: accepts a .gguf file or a config.json model dir
  (anything else is refused, satisfying the greedy-probe rule); context_size
  maps to max_model_len, options block_size/num_blocks/max_num_seqs size the
  KV cache and scheduler admission.
- Predict -> vllm_complete (blocking); PredictStream -> vllm_complete_stream
  with the per-delta C callback bridged into the gRPC stream. The backend
  embeds base.Base (not SingleThread): concurrent requests batch continuously
  in the engine's shared AsyncLLM scheduler.
- PredictOptions.Grammar -> the ABI's structured_grammar (GBNF), giving
  grammar-constrained tool calling at parity with llama-cpp; the ABI also
  exposes JSON-schema/regex/choice constraints.
- Hand-mirrored POD structs with layout locked by unit tests
  (unsafe.Offsetof vs the C offsets) and a runtime vllm_abi_version gate.
- One portable library per platform (vllm.cpp uses per-file SIMD tiers with
  runtime dispatch), so no avx/avx2/avx512 variant builds.

Wiring:
- backend-matrix: CPU amd64+arm64 (per-arch + manifest merge), CUDA 12/13
  amd64 (120a;121a Blackwell fat binary), L4T arm64 (121a, GB10/DGX Spark -
  the runtime-proven GPU target), Vulkan amd64, and Darwin arm64 Metal.
- backend/index.yaml meta + 12 image entries (latest/development x cpu,
  cuda12, cuda13, l4t, vulkan, metal); bump_deps registration for the
  VLLM_CPP_VERSION pin; root Makefile registration; test-extra runs the unit
  specs (pure Go, no engine build).
- Importers: preference-only swaps - llama-cpp (GGUF) and vllm (safetensors)
  advertise vllm-cpp via AdditionalBackends and emit backend: vllm-cpp
  without tokenizer templating (the C ABI takes the FINAL prompt; templating
  and tool parsing stay LocalAI-side). No auto-detect importer.
- Docs: backends list, top-level README maintained-engines table,
  compatibility table.

Verified: 20/20 Ginkgo specs against the real pinned engine and
Qwen3.5-2B-UD-Q8_K_XL.gguf on CPU - blocking + streaming parity, greedy
determinism, stop words, GBNF-constrained generation, and 4 concurrent
streams; plus a dlopen/ABI-gate smoke of the built gRPC server binary.
Upstream ABI v2 + production structured-output wiring landed as
mudler/vllm.cpp@86013f3.

Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vllm-cpp): ride the autoparser code path - engine-side chat templating and tool engagement (ABI v3)

The backend now implements AIModelRich (PredictRich / PredictStreamRich) over
vllm.cpp's ABI v3 chat entry points, so chat and tool calling ride the SAME
code path as the llama.cpp autoparser: the ENGINE renders the model's chat
template, decides when a tool call engages, and parses it - LocalAI receives
pre-parsed ChatDelta / ToolCallDelta protos exactly as it does from llama-cpp.

- With use_tokenizer_template + structured Messages, PredictOptions lowers to
  ONE OpenAI chat request JSON (messages, tools, tool_choice, sampling,
  stream_options.include_usage) for vllm_chat / vllm_chat_stream. tool_choice
  auto lowers engine-side to a LAZY structural-tag decode constraint - free
  text until the model emits the tool trigger, then the call is
  grammar-constrained; required/named force a call. Tool output is parsed by
  the engine's streaming Hermes-style parser; each chat.completion.chunk maps
  onto ChatDeltas (content / reasoning_content / tool_calls) which the host
  already prefers over Go-side tag extraction. Without structured messages the
  plain path (LocalAI templating + optional GBNF grammar) applies unchanged.
- The engine resolves the chat template from the GGUF tokenizer.chat_template
  metadata (or tokenizer_config.json); templates beyond its minja subset -
  e.g. the full Qwen3.5 namespace()/macro template - degrade engine-side to a
  Hermes-aware fallback prompt (tools schemas + <tool_call> instruction) with
  a stderr witness, so structural-tag engagement keeps working.
- Importers now emit the same config shape as llama-cpp for vllm-cpp
  (use_tokenizer_template: true, no-grammar autoparser flow); only the
  llama-cpp-specific use_jinja option and the vllm-python parser options are
  dropped.
- Pin bumped to mudler/vllm.cpp@aaed7ec (ABI v3 + chat-prompt resolution).

Verified against the real engine and Qwen3.5-2B-UD-Q8_K_XL.gguf on CPU: full
suite green - blocking chat, streaming deltas concatenating byte-equal to the
blocking answer, a REQUIRED tool call returning schema-valid arguments JSON,
and an AUTO run where the engine itself engages get_weather and streams parsed
tool deltas; plus unit specs for the request lowering, chunk->ChatDelta
mapping, and the C struct mirrors (ABI gate now v3).

Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vllm-cpp): ABI v5 - engine-side parser selection for 30 tool dialects + reasoning

Bump the vllm.cpp pin to the autoparser-parity engine: 30 tool-call dialects
(every pure-text parser in the pinned vLLM registry, each ported 1:1 with its
upstream tests), 7 reasoning parsers, google/minja as the template renderer
(the full Qwen3.5 template now renders engine-side), per-family structural
tags (tool_choice required/named compiles the model's NATIVE syntax where
expressible), and template auto-detection for both parser axes.

Backend changes:
- cModelParams mirrors ABI v5 (tool_parser + reasoning_parser fields,
  layout-locked by the offset tests; ABI gate now v5).
- New model options tool_parser:<name> / reasoning_parser:<name> pass through
  to the engine; unset means template auto-detection (18-row tool marker
  table; [THINK]->mistral, <think>->think_auto for reasoning); "none"
  disables the reasoning split; unknown names fail the first chat call.
- Chat chunks parse the `reasoning` field (the pin renamed
  reasoning_content), flowing into ChatDelta.ReasoningContent which the host
  already prefers.

Live e2e against Qwen3.5-2B-UD-Q8_K_XL.gguf on CPU, full suite green: the
real chat template renders (no more fallback), reasoning auto-detection picks
think_auto so markerless answers stay pure content (the live run caught the
deepseek_r1 content-swallow upstream and drove the think_auto fix), required
tool_choice returns schema-valid arguments, auto tool_choice engages
engine-side and streams parsed deltas, and blocking/streaming stay
byte-identical. Turn latency also dropped (proper template EOS behavior).

Upstream program landed as mudler/vllm.cpp 86013f3..5fffe7e (ABI v2-v5,
minja, parser waves B1/B2/B4, reasoning seam, structural-tag registry,
think_auto).

Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(vllm-cpp): bump the engine pin to the ENG-wave close-out

mudler/vllm.cpp@df8909b: the six engine-backed vLLM tool-parser families
(qwen3-coder/xml/mimo, kimi_k2, glm45/47, minimax_m2, gemma4, seed_oss)
text-reimplemented from their wire formats and held to the upstream test
suites - 39 registered dialects; the pinned vLLM registry is now covered
except the three Rust/Harmony-backed families, descoped by decision. kimi_k2
also gains a full native structural-tag builder; four new template
auto-detection rows land with test-pinned ordering.

Full backend e2e re-run green against Qwen3.5-2B-UD-Q8_K_XL.gguf on CPU.

Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): add the vllm-cpp-development gallery meta

The gallery grew the twelve latest/development image entries but was missing
the separate vllm-cpp-development meta (own capabilities map targeting the
-development image names), which every backend ships so the development
gallery resolves per-platform. Validated: all capability targets in both
metas resolve to existing entries, and every image URI's tag suffix matches
a backend-matrix build.

Also full-stack verified in this change's context (single-node local-ai from
this branch, locally-built backend under --backends-path, Qwen3.5-2B GGUF):
/v1/chat/completions non-stream (clean content + usage), streaming (SSE
deltas), tool_choice auto engaging get_weather engine-side with schema-valid
arguments and finish_reason=tool_calls, and streamed tool-call deltas in the
standard name-first cadence.

Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): repair the CI backend builds - gcc-14 -Werror + fat-arch Triton

Two distinct failures took down all five vllm-cpp backend builds on the PR:

1. gcc-14 (ubuntu:24.04 CI images; the local toolchain is gcc-13) fails the
   engine build with -Werror=maybe-uninitialized in InputBatch::condense - a
   false positive through a staging std::optional's raw storage. Fixed
   upstream (mudler/vllm.cpp@61f3e85) by moving slot-to-slot directly;
   verified BOTH ways under dockerized g++-14.2 (unfixed reproduces CI's two
   diagnostics exactly, fixed compiles clean) with the engine's behavior
   suites green. Pin bumped to that sha.

2. The amd64 CUDA builds died at CMake configure: the vendored Triton-AOT
   cubin trees are per-arch and the engine refuses -DVLLM_CPP_TRITON=ON on a
   multi-arch (120a;121a) fat build unless pinned to one tree, which would be
   unsound for the other arch. Triton is now enabled only on the single-arch
   arm64/GB10 build (where the cubins matter); the fat amd64 binary uses the
   engine's non-AOT GDN path.

Backend e2e re-run green at the new pin (Qwen3.5-2B on CPU, full suite).

Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): cuda-12 images cannot compile compute_121a - target 120a only

The second CI round surfaced a CUDA-version constraint: the cuda-12 (12.8)
image's nvcc rejects 'compute_121a' (GB10 arch support landed with CUDA 13),
killing the amd64 cuda-12 build at nvcc. Gate the architecture list on
CUDA_MAJOR_VERSION (exported by Dockerfile.golang): cuda-12 builds consumer
Blackwell 120a only, cuda-13 keeps the 120a;121a fat binary, arm64/l4t
(cuda-13) keeps single-arch 121a with the Triton cubins. GB10 is arm64, so
the amd64 cuda-12 image never served it - no capability change.

Verified by Makefile dry-run variable dumps for all three combinations
(cuda12 -> 120a; cuda13 -> 120a;121a; cpu -> CUDA off).

Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): drop the cuda-12 variant - the engine needs the CUDA 13 toolchain

Third CI round, third layer: with the arch list already narrowed to 120a,
the cuda-12 (12.8) build still dies in ptxas compiling the sm_120a NVFP4 MMA
kernels ("Vector type too large, exceeds 128 bit limit") - the Blackwell fp4
path genuinely requires the CUDA 13 toolchain, and vllm.cpp supports
Blackwell-family GPUs only. Shipping a cuda-12 image without the fp4 kernels
would be a crippled build of an engine whose whole GPU story is fp4, so the
variant is dropped instead:

- backend-matrix: cuda-12 vllm-cpp entry removed (cuda-13 amd64, l4t arm64,
  cpu, vulkan, metal remain).
- gallery: cuda12 image entries removed; the nvidia capability now resolves
  to the cuda13 image in both metas; the nvidia-cuda-12 key is dropped so
  older-driver hosts fall back to the CPU image instead of an unrunnable one.
- backend Makefile: BUILD_TYPE=cublas under CUDA_MAJOR_VERSION=12 now fails
  fast with a clear message; cuda-13 keeps the 120a;121a fat binary and
  arm64/l4t keeps 121a with the Triton cubins.

Verified: Makefile branch dumps for all four combinations (cuda12 loud
error, cuda13 fat, arm64 121a+Triton, cpu off), YAML parses, matrix filter
tests green, gallery capability targets all resolve.

Assisted-by: Claude Code:claude-fable-5

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): forward multi-turn tool identity and reasoning to the engine

chatRequestJSON dropped Message.ToolCallId and Message.Name on role="tool"
replies and Message.ReasoningContent on assistant history, so a second
turn after tool execution reached the engine's chat template without the
fields that bind a tool result to the call it answers. Forward all three
(present-only, matching the OpenAI wire shape) and pin vllm.cpp to
6a0bd3e7, where ChatMessage parses/round-trips tool_calls, tool_call_id,
name and reasoning and the minja adapter exposes them to the template
context.

Adds the round-trip request-lowering spec (user -> assistant tool_call ->
tool reply -> lowered request) and re-ran the gated e2e suite against the
new engine pin with a real Qwen3.5 GGUF: chat, reasoning split, streaming
parity, required-tool and auto-tool cases all green.

Assisted-by: Claude Code:claude-fable-5 [Bash] [Edit] [Read]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): bump vllm.cpp for the darwin arm64 i8mm build fix

The darwin-metal CI job was the first build to compile the engine's arm
CPU-quant files on macOS and hit their Linux-only <asm/hwcap.h> /
<sys/auxv.h> includes. vllm.cpp 9e1c9025 detects i8mm per-OS (auxv on
Linux, sysctl on Apple Silicon) with kernels untouched. Gated e2e suite
re-run green against the new pin with a real Qwen3.5 GGUF.

Assisted-by: Claude Code:claude-fable-5 [Bash] [Read]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): darwin build - bound cmake parallelism when nproc is absent

The macOS runners have no nproc, so JOBS evaluated empty and
`cmake --build -j$(JOBS)` became bare `-j`: unlimited clang jobs on a
3-core/7GB Mac, which swap-thrashed until the 6h GHA timeout (the log
shows "nproc: Command not found" and 7+ concurrent clang processes being
reaped at the cutoff). Use the same portable fallback chain as the other
darwin backends: nproc, then sysctl hw.ncpu, then 4.

Assisted-by: Claude Code:claude-fable-5 [Bash] [Edit] [Read]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-26 23:04:48 +02:00
mudler's LocalAI [bot]
90355cd444 chore: ⬆️ Update mudler/magpie-tts.cpp to 3008ff73fc2d2da9e4d743b09350aa7023e8980c (#11126)
⬆️ Update mudler/magpie-tts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 00:37:10 +02:00
mudler's LocalAI [bot]
02bccadef9 chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to 35ebe5376b82a0a59d008586d55bbe623d449011 (#11127)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 00:36:58 +02:00
mudler's LocalAI [bot]
0b216b63f0 chore: ⬆️ Update CrispStrobe/CrispASR to b516d8402c994f8455701f38dfbe578907328db7 (#11124)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 00:03:49 +02:00
mudler's LocalAI [bot]
86c81c0e56 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260725151812 (#11123)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 00:02:17 +02:00
mudler's LocalAI [bot]
9091a1f2e4 chore: ⬆️ Update vllm-project/vllm cu130 wheel to 0.26.0 (#11125)
⬆️ Update vllm-project/vllm cu130 wheel

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 00:02:02 +02:00
mudler's LocalAI [bot]
6dadaea91c chore(model-gallery): ⬆️ update checksum (#11129)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 00:01:41 +02:00
mudler's LocalAI [bot]
decb606216 feat(swagger): update swagger (#11122)
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-26 00:01:19 +02:00
mudler's LocalAI [bot]
5d57c08c6e feat(distributed): cache staged-artifact hashes and publish the model load lifecycle (#11121)
feat(distributed): cache staged-artifact hashes and publish load lifecycle

Every load request re-hashed every staged artifact on the controller
(probeExisting and the upload path both re-read the full file), which for
a large multi-file model on NAS-backed storage is minutes of pure
re-reading per request even when nothing changed - observed as ~9 minutes
of "Upload skipped (file already exists with matching hash)" before every
avatar generation. Cache the local hash in the same .sha256 sidecar the
worker-side transfer server already maintains, invalidated whenever the
sidecar is older than the file.

The whole staging+loading phase was also invisible: the NodeModel row was
only written after LoadModel succeeded, so /api/nodes and the UI showed
nothing while a cold load spent 10+ minutes staging - indistinguishable
from nothing happening. Publish the lifecycle instead: "staging" as soon
as the node is chosen, "loading" when the checkpoint load starts, and the
existing "loaded" on success, with the row removed on any failure so a
dead load does not leave a phantom replica. The early row also reserves
the replica slot against concurrent schedulers. The nodes view already
renders non-loaded states on model chips; style "staging" like "loading".

Audited every state-filtered registry/router query: eviction, routing,
reconciler and idle-model queries all filter state='loaded' explicitly,
so the new transitional rows are visible to observability surfaces but
inert to scheduling decisions (except slot occupancy, intentionally).

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 16:05:44 +02:00
mudler's LocalAI [bot]
0d82efde2b fix(gallery): coalesce Hugging Face artifact progress (#11117)
* fix(gallery): coalesce artifact download progress

Buffer high-frequency downloading events and forward only the latest event on a periodic tick. Flush progress synchronously at phase boundaries and shutdown to preserve ordering and final state.

Assisted-by: Codex:gpt-5

* fix(gallery): wire progress coalescing into model installs

Route artifact progress through the 250 ms coalescer and flush it on every model operation exit. Keep the legacy download callback unchanged.

Assisted-by: Codex:gpt-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-25 08:39:07 +02:00
mudler's LocalAI [bot]
2d889e61a6 feat(backend): add magpie-tts-cpp text-to-speech backend (#11115)
* feat(backend): add magpie-tts-cpp text-to-speech backend

Add a Go + purego backend wrapping the magpie-tts.cpp ggml port of NVIDIA's
Magpie TTS Multilingual 357M (encoder + autoregressive decoder over NanoCodec
tokens), producing 22.05 kHz mono audio in 5 baked voices (Aria, Jason, John,
Leo, Sofia; case-insensitive names or indices 0-4) across 9+ languages from a
single self-contained GGUF. Mirrors qwen3-tts-cpp / moss-tts-cpp: dlopen the
static-ggml shared library, bind the flat magpie_tts_capi_* C-API via purego
(no local C shim needed, the upstream .so exports it directly), and serve the
gRPC TTS + TTSStream methods behind base.SingleThread (the C context is not
reentrant across synthesize calls).

The backend CMakeLists translates the Makefile's -DGGML_{CUDA,METAL,VULKAN,HIP}
flags into upstream's MAGPIE_GGML_* toggles (upstream FORCE-overwrites the ggml
cache entries from those), pinned to magpie-tts.cpp v0.1.1
(e3f3dd1ebe22b64e7405f93b519f2d1930712568), which statically links ggml into
libmagpie-tts.so (ldd shows only system libs).

Wires the full registration: backend-matrix.yml (CPU amd64/arm64, CUDA 12/13,
Intel SYCL f16/f32, Vulkan amd64/arm64, ROCm, NVIDIA L4T + L4T CUDA 13, and
Darwin metal), backend/index.yaml metas and image entries, the root Makefile
build targets, the changed-backends backend-filter path mapping, the bump_deps
auto-bump matrix, a test-extra per-backend smoke job, the /backends/known
pref-only importer entry, the backend capabilities map (TTS + TTSStream, no
voice cloning), and the README / compatibility-table docs rows.

Verified locally: unit + e2e Ginkgo suites pass against the real q8_0 GGUF
(22.05 kHz mono WAV, RMS > 0.01), a live gRPC LoadModel + TTS round-trip
returns valid non-silent audio, and the pre-commit gates (make lint,
make test-coverage-check) pass, run manually with LOCALAI_TEST_HTTP_PORT
overriding the locally-occupied 9090.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* gallery: add magpie-tts-cpp model entries (q8_0 + f16)

Add the Magpie TTS Multilingual 357M GGUFs from mudler/magpie-tts.cpp-gguf to
the model gallery: q8_0 (~624 MB, near-lossless, fastest decode, recommended)
with an f16 (~784 MB) variant, both served by the magpie-tts-cpp backend.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* magpie-tts-cpp: bump pin to rewritten upstream v0.1.1 SHA

Upstream history was rewritten to purge accidentally committed build
artifacts; v0.1.1 now resolves to 6f7696cf.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 08:38:48 +02:00
mudler's LocalAI [bot]
130daa0c55 chore: ⬆️ Update leejet/stable-diffusion.cpp to 87a01773be23b996e38217a6a574c2de08ac560f (#11111)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-25 01:20:23 +02:00
mudler's LocalAI [bot]
cda67dfb87 chore(model-gallery): ⬆️ update checksum (#11108)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-25 01:20:12 +02:00
mudler's LocalAI [bot]
a32a53fa8a chore: ⬆️ Update CrispStrobe/CrispASR to 2f26702117b4c7697d0fd421b6ee77cd7757ca0d (#11109)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-25 01:19:46 +02:00
mudler's LocalAI [bot]
05e16e0fa8 chore: remove local pre-commit gates (#11116)
Remove the versioned pre-commit hook and its installer while retaining CI coverage and conformance checks.

Assisted-by: Codex:gpt-5

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-25 01:19:32 +02:00
mudler's LocalAI [bot]
0d26de23ca chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260724093932 (#11105)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-25 00:05:55 +02:00
mudler's LocalAI [bot]
f92301f40c chore: ⬆️ Update ikawrakow/ik_llama.cpp to f359df4bc9a5e864029cbec4cb608e95f3500ce6 (#11107)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-24 23:52:06 +02:00
dependabot[bot]
7c95b25bd9 chore(deps): bump grpcio from 1.82.1 to 1.83.0 in /backend/python/transformers (#11085)
chore(deps): bump grpcio in /backend/python/transformers

Bumps [grpcio](https://github.com/grpc/grpc) from 1.82.1 to 1.83.0.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.82.1...v1.83.0)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.83.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-24 22:48:23 +02:00
dependabot[bot]
6ad5df4f7c chore(deps): bump sentence-transformers from 5.6.0 to 5.6.1 in /backend/python/transformers (#11084)
chore(deps): bump sentence-transformers in /backend/python/transformers

Bumps [sentence-transformers](https://github.com/huggingface/sentence-transformers) from 5.6.0 to 5.6.1.
- [Release notes](https://github.com/huggingface/sentence-transformers/releases)
- [Commits](https://github.com/huggingface/sentence-transformers/compare/v5.6.0...v5.6.1)

---
updated-dependencies:
- dependency-name: sentence-transformers
  dependency-version: 5.6.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-24 22:46:38 +02:00
dependabot[bot]
8aa5305790 chore(deps): bump grpcio from 1.82.1 to 1.83.0 in /backend/python/vllm (#11086)
Bumps [grpcio](https://github.com/grpc/grpc) from 1.82.1 to 1.83.0.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.82.1...v1.83.0)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.83.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-24 22:46:02 +02:00
dependabot[bot]
c786edec47 chore(deps): bump grpcio from 1.82.1 to 1.83.0 in /backend/python/coqui (#11083)
Bumps [grpcio](https://github.com/grpc/grpc) from 1.82.1 to 1.83.0.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.82.1...v1.83.0)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.83.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-24 22:33:32 +02:00
mudler's LocalAI [bot]
6124498e17 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260723125609 (#11087)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-24 22:31:28 +02:00
mudler's LocalAI [bot]
35d3a43053 chore: ⬆️ Update leejet/stable-diffusion.cpp to 5114672c482012d77d24bbd09eae86b53c48256b (#11088)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-24 22:31:11 +02:00
mudler's LocalAI [bot]
983d77ed27 chore: ⬆️ Update antirez/ds4 to 0a7ad776b9068348e6cb09df8cafa9cadd285298 (#11089)
⬆️ Update antirez/ds4

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-24 18:33:57 +02:00
mudler's LocalAI [bot]
07dc4439aa chore: ⬆️ Update CrispStrobe/CrispASR to cf0fdbbe38ad0aa107e3250f6ee5bdc755aced45 (#11090)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-24 16:12:51 +02:00
mudler's LocalAI [bot]
a66f904a3c chore: ⬆️ Update ikawrakow/ik_llama.cpp to 31018dc51135a8a3ded085fa7e198befff19ebf4 (#11091)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-24 15:05:53 +02:00
mudler's LocalAI [bot]
ce6c42d677 chore(model-gallery): ⬆️ update checksum (#11092)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-24 15:05:40 +02:00
mudler's LocalAI [bot]
90d93c71cd fix(downloader): hash the partial file before issuing the resume request (#11099)
The stall watchdog arms as soon as the response body exists, but the
downloader then re-hashed the entire existing .partial before reading a
single byte from the network. On slow models storage (a CIFS share
reading at ~117MB/s) hashing a multi-GB partial outlasts the 60s stall
window, so the watchdog aborted every healthy resume with 'download
stalled: no data received for 1m0s'. The partial never grew, so every
retry re-paid the same hash and failed identically, wedging the install
permanently (any partial over ~7GB on such storage).

Open the partial and hash it before the HTTP request instead: the
watchdog now only measures actual network idle time, and the origin no
longer sits on an idle connection while the hash runs.

Assisted-by: Claude:claude-fable-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-24 12:57:30 +02:00
Isabel Wu
977f663cb0 fix(trl): disable inline GRPO reward code by default (RCE, #11015) (#11068)
fix(trl): disable inline GRPO reward code by default (RCE)

POST /api/fine-tuning/jobs accepts reward_functions[].code, an inline Python
body, and compile_inline_reward() execs it against a restricted-builtins
allowlist (_SAFE_BUILTINS). That allowlist is not a security boundary:
().__class__.__bases__[0].__subclasses__() reaches os._wrap_close and thus
os.system, giving arbitrary code execution. The fine-tuning endpoint is
unauthenticated by default, so any caller could run code on the host.

Hardening the allowlist is a losing game against CPython introspection, so
inline reward code is now refused unless the operator explicitly opts in with
LOCALAI_TRL_ALLOW_INLINE_REWARD=true on the backend. Builtin reward functions
are unaffected. The gate lives in build_reward_functions(), the single point
all inline specs flow through. Docs updated to stop describing the allowlist as
a sandbox and to document the opt-in.

Fixes #11015

Signed-off-by: Isabel Wu <231155141+wuisabel-gif@users.noreply.github.com>
Co-authored-by: Isabel Wu <231155141+wuisabel-gif@users.noreply.github.com>
2026-07-23 23:10:34 +02:00
mudler's LocalAI [bot]
2fe10c3c4a fix(model-artifacts): persist companion artifacts so remote workers get the base_model option (#11075)
fix(model-artifacts): persist companion artifacts, not just the primary

A managed model can declare companion artifacts (LongCat-Video-Avatar-1.5
pulls its tokenizer, text encoder and VAE from the separate LongCat-Video
base repo via a target: companion artifact). preloadOne resolves the whole
set in memory, but the binding written back to disk carried only the
primary: persistArtifactBinding marshalled []Spec{result.Spec} and replaced
the entire artifacts: list with it, silently dropping every companion.

In a single process the loss is invisible because the in-memory config keeps
the companion. It bites on the next controller restart: the config reloads
from the mangled file with the primary alone, so withCompanionArtifactOptions
finds no resolved companion and synthesizes no base_model option. The remote
longcat-video backend then never receives base_model, falls back to
BASE_MODEL_ID and downloads the repo itself ("Downloading required files for
meituan-longcat/LongCat-Video"), failing the load with "base_model must point
to a LongCat-Video checkpoint".

This is why an explicit base_model:<path> added to the config options works
where the managed companion does not: an explicit option lives in options:,
which is never rewritten, while the managed companion lives in artifacts:,
which the binding overwrote.

Persist the full resolved set (primary + every companion), and widen
bindingNeedsPersistence to compare the whole artifact list so a companion
resolving for the first time still triggers a write. The single-node path is
unaffected: there the in-memory config already carried the companion, and the
staging/ModelPath resolution for a remote worker (nested per-model staged
root, #10949) is unchanged and already correct once the option is generated.

Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-23 16:50:50 +02:00
788 changed files with 70361 additions and 9103 deletions

View File

@@ -304,7 +304,9 @@ React pages that want to filter the ModelSelector by capability import this symb
### 4. `docs/content/` (user-facing documentation)
A new capability deserves its own page under `docs/content/features/`, plus cross-links from related features and an entry in `docs/content/whats-new.md`. See the pattern used by `face-recognition.md` / `object-detection.md`.
A new capability deserves its own page under `docs/content/features/`, plus cross-links from related features. See the pattern used by `face-recognition.md` / `object-detection.md`.
Announcing it is the release's job, not this page's: the capability gets covered in the release blog post under `website/content/blog/`. See [preparing-a-release.md](preparing-a-release.md). `docs/content/whats-new.md` is only a pointer at the blog and GitHub Releases, so there is nothing to add there.
## Path protection rules
@@ -334,7 +336,7 @@ When adding a new endpoint:
- [ ] Swagger block on the handler: `@Summary`, `@Tags`, `@Param`, `@Success`, `@Router`
- [ ] If new capability area (new swagger tag): entry in `instructionDefs` in `core/http/endpoints/localai/api_instructions.go` + test count bumped in `api_instructions_test.go`
- [ ] If new `FLAG_*` usecase flag: matching `CAP_*` symbol exported from `core/http/react-ui/src/utils/capabilities.js`
- [ ] `docs/content/features/<feature>.md` created; cross-links from related feature pages; entry in `docs/content/whats-new.md`
- [ ] `docs/content/features/<feature>.md` created; cross-links from related feature pages; capability covered in the release blog post (see [preparing-a-release.md](preparing-a-release.md))
**Quality**
- [ ] Error responses use `schema.ErrorResponse` format (or `echo.NewHTTPError` with a mapped gRPC status — see the `mapBackendError` helper in `core/http/endpoints/localai/images.go`)

View File

@@ -28,7 +28,6 @@ The core Go suites (`./pkg`, `./core`, plus the in-process integration suite `./
- **Build tags (`COVERAGE_TAGS`, passed via `GINKGO_TAGS`):** defaults to `debug auth`. The `auth` tag is required to compile the real (sqlite-backed) auth implementation and its ~150 `//go:build auth` tests — without it those files aren't built, the tests don't run, and the gate scores auth against a stub (~3.7% instead of ~38%). If you add new tag-gated tests, extend `COVERAGE_TAGS` or they won't count (and likely won't run in CI at all).
- `make test-coverage-check` — runs `test-coverage`, then `scripts/coverage-check.sh` fails the build if total coverage is **below** the committed baseline in `coverage-baseline.txt`. The Linux job in `.github/workflows/test.yml` runs this instead of `make test`.
- `make test-coverage-baseline` — regenerates and overwrites `coverage-baseline.txt` from the current run.
- `make install-hooks` — sets `core.hooksPath` to the versioned `.githooks/`, whose `pre-commit` runs checks scoped to what's staged: Go changes → `make lint` + `make test-coverage-check`; `core/http/react-ui/` changes → `make test-ui-coverage-check` (Playwright e2e + UI coverage gate). A commit touching neither is skipped; bypass with `git commit --no-verify`. The hook resolves golangci-lint's new-from base to `upstream/master``origin/master``master`, so it works from a fork clone where `origin/master` is stale (passed to `make lint` via `LINT_NEW_FROM`).
### React UI coverage
@@ -38,12 +37,11 @@ The React UI (`core/http/react-ui/`) has **no component/unit tests** — its onl
- **Browser:** the flake dev shell ships `chromium` and exports `PLAYWRIGHT_CHROMIUM_PATH`; `playwright.config.js` uses it via `launchOptions.executablePath`, and the Makefile skips `playwright install` when it's set. This avoids Playwright's downloaded browser, which can't resolve system libs (`libglib-2.0`, …) on NixOS. In CI (no `PLAYWRIGHT_CHROMIUM_PATH`) the Makefile falls back to `playwright install --with-deps chromium`.
- The app is a React SPA, so coverage accumulates across in-app navigation within a test; a full `page.goto`/reload resets it.
- `.nycrc.json` uses `all: true`, so **every `src/**` file is in the report**, including 0%-coverage ones — that's how you spot features with no test at all (sort the HTML report or `coverage-summary.json` by line% ascending).
- **UI coverage gate:** `make test-ui-coverage-check` runs the suite then `scripts/ui-coverage-check.sh`, failing if total line coverage drops more than `UI_COVERAGE_TOLERANCE` below `core/http/react-ui/coverage-baseline.txt`. `make test-ui-coverage-baseline` regenerates the baseline. Runs in CI (`tests-ui-e2e.yml`) and pre-commit on `core/http/react-ui/` changes.
- **UI coverage gate:** `make test-ui-coverage-check` runs the suite then `scripts/ui-coverage-check.sh`, failing if total line coverage drops more than `UI_COVERAGE_TOLERANCE` below `core/http/react-ui/coverage-baseline.txt`. `make test-ui-coverage-baseline` regenerates the baseline. Runs in CI (`tests-ui-e2e.yml`).
- **Why it has a tolerance (unlike the strict Go gate):** UI e2e coverage is *non-deterministic*. Specs that assert on state and end while async/lazy render work is still in flight collect those lines only when the render beats the coverage teardown — so the total drifts with machine speed/load (a fast local box reads higher than a slow CI runner), diffusely across many specs. The tolerance absorbs that drift, so set the baseline *below* the slow-CI floor, never to a fast-local `make test-ui-coverage-baseline` number, or CI flaps.
- **Raising coverage is cheap:** a *render-smoke* spec (navigate to a route, assert its header renders) mounts a lazy page and runs its full render + initial effects, capturing most of its lines in a few lines of test — see `e2e/page-render-smoke.spec.js`. Auth is disabled in the test server (`isAdmin=true`), so `RequireAdmin`/`RequireFeature` routes render without a mock. The most *deterministic* win is removing a race: make a spec `await` a rendered element before ending (see `e2e/agents.spec.js` → AgentCreate) so its lines count every run.
Rules (both gates):
- **Install the hooks:** `make install-hooks` once per clone so lint + coverage run pre-commit. Don't lean on CI for what the hook catches.
- **Don't work around the gate:** never `git commit --no-verify`, and never hand-lower a baseline or widen a tolerance to turn a red gate green. The ratchet only moves up.
- **Don't weaken the gate:** never hand-lower a baseline or widen a tolerance to turn a red gate green. The ratchet only moves up.
- If a change drops coverage, **add tests** (sort `coverage-summary.json` by line% ascending to find untested code) rather than editing the baseline. When coverage legitimately rises, commit the regenerated baseline (`make test-coverage-baseline` / `test-ui-coverage-baseline`).
- The Go gate is **strict — no tolerance**; `covermode=atomic` keeps it deterministic. The UI gate keeps a small tolerance only because its e2e coverage isn't.

View File

@@ -122,18 +122,89 @@ The per-backend prefix match only sees files under a backend's own directory, so
| Changed path | Rebuilds |
|---|---|
| `backend/backend.proto` | everything (all languages compile or copy it) |
| `backend/backend.proto` | nothing if the edit is additive-only, otherwise everything (see below) |
| `backend/Dockerfile.<x>` | the Linux entries whose `dockerfile:` names it |
| `backend/python/common/` | Python, Linux + Darwin |
| `scripts/build/package-gpu-libs.sh` | Python, Linux only |
| `scripts/build/package-gpu-libs.sh` | every Linux entry (Python, Go and C++ all run it) |
| `scripts/build/<lang>-darwin.sh` | the Darwin entries that build target routes to |
| `.github/workflows/backend_build[_darwin].yml` | everything on that OS |
| anything else under `scripts/build/` (except `*_test.sh`) | everything — conservative default for unclassified packaging inputs |
Deliberately excluded: `backend/index.yaml` (gallery metadata, never enters an image), `.github/backend-matrix.yml` (adding a backend would rebuild all of them), `backend/Dockerfile.base-grpc-builder` (owned by `base-images.yml`), and the root `Makefile` (touched in ~11% of commits, and its backend-relevant edits arrive alongside the backend directory anyway). `make test-ci-scripts` pins all of this.
#### `backend/backend.proto` is content-filtered, not path-filtered
Every language consumes the proto, so a path rule for it can only ever say "rebuild all 473 images". It changes in ~1.3% of commits, and that was enough to make it the single largest CI cost driver in the repo: on 2026-07-29 four runs totalling 935 queued jobs traced to nothing but a proto edit, one of which (#11158) was a six-line diff adding `bool cache_prompt = 8;`.
An additive proto edit cannot change how a backend that never references the new symbol behaves, so `filterMatrix()` suppresses the rule for one. `changed-backends.js` fetches `backend/backend.proto` at the base revision (same contents-API pattern as `.github/backend-matrix.yml`) and hands both texts to `protoChangeIsAdditive()`, which compares them structurally rather than textually:
- **Additive, rebuilds nothing**: a new field with an unused number, a new message, a new enum value, a new RPC. Comment, whitespace and ordering changes also land here.
- **Breaking, rebuilds everything**: a removed, renumbered, retyped or renamed field, a dropped RPC, a changed `option` or `package`. So does an unresolvable base revision, matching the run-all posture used for a truncated diff.
Checked against every proto commit in the preceding six months, all nine resolvable ones classify as additive. Note the tradeoff this accepts: generated stubs do change for an additive edit, so image bytes would differ on a rebuild even though behavior does not. That is the same standard already applied when the filter declines to rebuild on unrelated `pkg/` changes, and the weekly cron remains the backstop.
The Sunday 06:00 UTC cron on `backend.yml` exists specifically because path filtering can leave Python backends frozen on stale wheels. `DEPS_REFRESH` (below) only fires when the build actually runs, so an untouched Python backend would never re-resolve its unpinned deps. The weekly cron is the safety net.
## Content-blind PRs skip the workflows that cannot see them
`backend_pr.yml` and `test-extra.yml` filter themselves (matrix generation and a `detect-changes` job), so a gallery-only or docs-only PR costs them about one job each. The Go and image workflows had no filter of any kind, so a one-line `gallery/index.yaml` edit queued 20 jobs, and a docs-only PR queued the same.
This is worth more than it looks. Measured over the week to 2026-07-30, **97% of CI wall-clock is queueing, 3% is execution** (median queue ~5h against a 4-20min median job). Cutting job count is therefore the only lever that shortens feedback time; making individual jobs faster moves 3%.
The volume is real: 13 gallery-only PRs merged that week with 10 open at once, and 78 of the 137 PRs opened were bot-generated.
`paths-ignore` on the PR trigger of `image-pr.yml` (7 jobs), `build-test.yaml` (3), `lint.yml` (2) and `tests-e2e.yml` (1) drops 13 of those 20. The excluded set:
| Path | Why no image or Go build can see it |
|---|---|
| `gallery/**` | Model-gallery metadata, parsed at runtime, never copied into an image |
| `docs/**`, `examples/**`, `**/*.md` | Never enter an image or a binary. `lint.yml` already excluded these before gallery was added |
### `backend/{cpp,go,python}/**` on `image-pr.yml` and `build-test.yaml` only
Version-pin bumps dominate PR volume: 48 `update/*` PRs in the week to 2026-07-30, from 16 pins, each a two-line diff. Most edit nothing but one `backend/*/<name>/Makefile`.
Neither of those two workflows can observe such a change. `make build` is `go build ./cmd/local-ai`, GoReleaser builds the same plus `./cmd/launcher`, and the core image's final stage ships only `entrypoint.sh`, `healthcheck.sh` and that binary. The per-backend trees are copied into the builder but nothing in them reaches the output.
What still triggers a full run, because none of it lives under those prefixes:
- `backend/backend.proto` — feeds `protogen-go`, so it does change the binary.
- `go.mod` / `go.sum` — the `go mod tidy` before-hook.
- `backend/Dockerfile.*` and anything else directly under `backend/`.
Deliberately **not** applied to:
| Workflow | Why it must keep seeing `backend/**` |
|---|---|
| `test.yml` | `TEST_PATHS` explicitly includes `./backend/go/cloud-proxy/...`, `./backend/go/local-store/...` and `./backend/go/valkey-store/...` |
| `lint.yml` | `.golangci.yml` carries `backend/`-scoped rules, so golangci-lint covers that tree |
| `tests-e2e.yml` | The e2e suite drives real backends over gRPC |
| `backend_pr.yml` | This is the workflow whose entire job is to rebuild the changed backend |
What still runs, and why it has to:
| Workflow | Why it keeps running |
|---|---|
| `test.yml` (`tests`) | `core/gallery/variants_lint_test.go` reads the real `gallery/index.yaml` and asserts the index invariants (no duplicate entry names, no build claimed by two parents). This is the only schema-level check the gallery has. |
| `yaml-check.yml` (`Yamllint`) | Lints `gallery/` for syntax. |
| `backend_pr.yml`, `test-extra.yml` | Already self-filtering; they stop after the detect step. |
Two properties this relies on:
- `paths-ignore` skips a run only when **every** changed file matches, so a PR touching the gallery *and* Go code still runs everything. That is what makes the exclusion safe rather than a hole.
- `master` carries no branch protection and no rulesets, so a skipped workflow reports no status and nothing waits on it. If required status checks are ever introduced, these four entries must be excluded from the required set or PRs will hang on "Expected — Waiting for status to be reported".
### `image.yml` on master push is gated too, by a job rather than a path filter
The same reasoning applies to master pushes, and the volume is larger there: on 2026-07-30, **12 of the 23 queued `image.yml` runs** were commits like "add 1 new model to gallery" or a docs fix, each rebuilding all 18 container images.
`image.yml` now has a `changes` job that decides once whether the push can affect any image; the other 11 jobs carry `needs: changes` plus an `if:` on its output. Verified against the shipped `Dockerfile`: the final stage copies only `entrypoint.sh`, `healthcheck.sh` and the `local-ai` binary, there is no `go:embed` of `gallery/` or `docs/`, and the gallery is fetched at runtime from `github:mudler/LocalAI/gallery/index.yaml@master`. A gallery-only commit therefore produces byte-identical images, and the gallery change reaches users through GitHub immediately whether or not an image is rebuilt.
Two properties to preserve if you touch it:
- **It is a job gate, not `paths-ignore`.** `paths-ignore` on `push` also applies to tag pushes, and a tag created on an existing commit carries an empty commits list, which would silently skip the release image build. The gate short-circuits to "build" for `refs/tags/*`, and for any push whose base commit is missing, zero, or unresolvable.
- **The merge jobs must name the gate explicitly.** They use `if: ${{ !cancelled() && ... }}`, and `!cancelled()` is true when a dependency is *skipped*, so without the extra condition they would run and try to merge manifest lists for images that were never built.
## The `DEPS_REFRESH` cache-buster (Python backends)
Every Python backend goes through the shared `backend/Dockerfile.python`, which ends with:
@@ -169,15 +240,38 @@ RUN --mount=type=cache,target=/root/.ccache,id=<backend>-ccache-${TARGETARCH}-${
bash /usr/local/sbin/compile.sh
```
The compile script exports `CMAKE_C/CXX/CUDA_COMPILER_LAUNCHER=ccache` so CMake threads ccache through gcc/g++/nvcc. `cache-to: type=registry,mode=max` exports the cache mount data into the registry cache, so subsequent builds restore it.
The compile script exports `CMAKE_C/CXX/CUDA_COMPILER_LAUNCHER=ccache` so CMake threads ccache through gcc/g++/nvcc. Cache scope is per `(TARGETARCH, BUILD_TYPE)` so e.g. cublas-12 doesn't share with cublas-13 (their CUDA headers differ; cross-pollination would just be cache misses anyway).
On a `LLAMA_VERSION` bump, most translation units are byte-identical to the previous version's preprocessed source — ccache returns the previous `.o` and skips the real compile. Same for LocalAI source changes that don't actually touch llama.cpp's CMake inputs. Cache scope is per `(TARGETARCH, BUILD_TYPE)` so e.g. cublas-12 doesn't share with cublas-13 (their CUDA headers differ; cross-pollination would just be cache misses anyway).
### ⚠️ This ccache does nothing in CI today
This section previously claimed that `cache-to: type=registry,mode=max` "exports the cache mount data into the registry cache, so subsequent builds restore it". **That is not true.** BuildKit does not export the contents of a `--mount=type=cache` to a registry cache export. A cache mount lives in the builder's local state, and every CI job gets a fresh runner with a fresh builder, so `/root/.ccache` starts empty on every single build.
Measured on 2026-07-30 from the `ccache -s` output the compile script already prints (it runs `ccache -z` first, so the numbers are per-build):
| Job | Commit touched | Build time | ccache |
|---|---|---|---|
| 89766266951 (llama-cpp, cublas 13) | `backend/go/magpie-tts-cpp/Makefile` only | 6369s | **0 / 889 hits**, and 0 / 1778 |
| 89766267281 (llama-cpp, hipblas) | same commit | 8160s | **0 / 537 hits** |
| 90210828110 (llama-cpp, cublas 12.8) | `LLAMA_VERSION` bump | 5673s | **0 / 813 hits** |
The first two are the decisive control: commit `90355cd44` changed exactly one file, `backend/go/magpie-tts-cpp/Makefile`, nowhere near llama.cpp. The engine source was byte-identical to the previous build, which is precisely the case this section says ccache should serve, and the hit rate was still **0.00%**. A cache that was being restored but merely matching poorly would show partial hits; 0-of-N is the signature of an empty cache.
So the paragraph above about `LLAMA_VERSION` bumps reusing previous `.o` files describes an intended design that is not in effect. `Dockerfile.{llama-cpp,ik-llama-cpp,turboquant,bonsai,ds4,privacy-filter}` pay the ccache wrapper overhead and get nothing back. Multi-hour C++ rebuilds are recompiling identical translation units from scratch.
**Do not "fix" this by adding cache mounts to more Dockerfiles.** Wiring the same mount into `Dockerfile.golang` (215 of the 434 matrix entries) was measured locally at 18% faster on a rebuild after a source edit, with a 71.5% ccache hit rate — but only because the local test reused one builder across both builds. In CI it would be a no-op for exactly the reason above.
Making this actually work needs the cache to live outside the builder. The options, none of them free:
- **ccache `remote_storage`** (ccache ≥ 4.4, HTTP or Redis backend) or **sccache** with an S3/GCS/Redis backend. Genuinely works across runners; needs a cache service to point at. quay.io is a registry, not a blob store, so the existing infra does not cover it.
- **Round-trip the cache dir through `actions/cache` on the runner**: restore it, pass it in, and export it back out via a build stage output. No external infra, but clunky, and the repo already sits at GitHub's 10 GB cache ceiling while the llama-cpp ccache alone is capped at 5 GB.
Until one of those lands, treat C++ backend builds as always-cold and spend the effort on not running them instead (path filtering, see above).
## Composite actions
Two composite actions handle runner-side prep:
- **`.github/actions/free-disk-space/action.yml`** — wraps `jlumbroso/free-disk-space@main` plus an explicit apt purge of dotnet/android/ghc/mono/etc. Reclaims ~610 GB on `ubuntu-latest`. No-op on self-hosted runners. Used by `backend_build.yml`, `image_build.yml`, `test.yml`, `tests-aio.yml`, etc.
- **`.github/actions/free-disk-space/action.yml`** — wraps `jlumbroso/free-disk-space@main` plus an explicit apt purge of dotnet/android/ghc/mono/etc. Reclaims ~610 GB on `ubuntu-latest`. No-op on self-hosted runners. Used by `backend_build.yml`, `image_build.yml` and `base-images.yml` — the jobs that actually build images. Deliberately **not** used by `test.yml`, which runs no buildx step.
- **`.github/actions/setup-build-disk/action.yml`** — relocates Docker's data-root to `/mnt` on hosted X64 runners. GHA hosted `ubuntu-latest` ships ~75 GB of unused space at `/mnt`; combined with the free-disk-space cleanup this gives ~100 GB working space — enough for ROCm dev image + vLLM torch install + flash-attn intermediate layers. No-op on self-hosted and on non-X64 hosted runners. Used by `backend_build.yml`, `image_build.yml`, `base-images.yml`.
Both actions run before any docker buildx step.
@@ -218,10 +312,20 @@ Eviction is rarely needed in normal operation — `DEPS_REFRESH` handles weekly
## What the cache does **not** cover
- The `free-disk-space` and `setup-build-disk` composite actions run on every job — these reclaim runner-state, not Docker layers, so BuildKit caches don't apply.
- The `free-disk-space` and `setup-build-disk` composite actions run on every job — these reclaim runner-state, not Docker layers, so BuildKit caches don't apply. `test.yml` deliberately does **not** use `free-disk-space`: it runs no buildx step, and the multi-GB fixture downloads that once justified it left `make test` in the test-suite reorg.
- Intermediate artifacts of `Build (PR)` are not pushed anywhere — PRs only build for verification.
- Darwin builds (see below) — macOS runners have no Docker daemon, so the registry-backed BuildKit cache cannot apply.
### The Linux Go workflows set `cache: false` on purpose
`test.yml`, `lint.yml`, `tests-e2e.yml` and friends pass `cache: false` to `actions/setup-go@v5`, unlike the darwin jobs. This looks like an oversight and is not.
Measured over the week to 2026-07-30, the `Set up Go` step has a **median of 11 seconds** on these runners. There is essentially nothing to win: the module download is not where the time goes. The expensive steps are compilation and test execution (`Test (with coverage gate)` at ~18.6min, `Test Backend E2E` at ~14.5min), and Go's build cache would have to survive across runners to touch those.
Enabling it also has a real cost. GitHub caps Actions cache at **10 GB per repo and the repo already sits at that ceiling** (31 entries), so every `setup-go` entry written by a branch with a distinct `go.sum` (222-375 MB on Linux, up to 1.4 GB on macOS) evicts something else. See the darwin cache budget below.
Before re-enabling this, measure `Set up Go` again and confirm it has actually become slow. If room is needed in the 10 GB budget, the cheapest evictions are the `docker.io--tonistiigi--binfmt` entries (~30 MB each, trivially re-fetched).
## Darwin native caches
`backend_build_darwin.yml` runs natively on `macOS-14` GitHub-hosted runners — there is no Docker, no BuildKit, no cross-job registry cache. Instead, the reusable workflow uses `actions/cache@v4` for four native caches that mirror the spirit of the Linux cache (warm by default, weekly refresh for unpinned Python deps, PRs read-only).
@@ -255,6 +359,26 @@ GitHub Actions caches are limited to 10 GB per repo. Steady-state worst case: ~8
One residual self-hosted reference remains in `test-extra.yml` (`tests-vibevoice-cpp-grpc-transcription` uses `bigger-runner` for the 30s JFK-decode timeout headroom). That's a separate concern.
### Small always-on jobs routed to `arc-runner-set`
The hosted pool is shared across the whole *account*, not per repo, so a burst in one repo starves the others. On 2026-07-31 it went to **zero scheduled jobs for 35 consecutive minutes** with 39 jobs queued, while `arc-runner-set` completed 12 jobs without interruption over the same window. Actions was healthy globally at the time (other public repos were scheduling normally), so this is an account-level throttle, not an outage.
`gh-pages.yml` (`build` + `deploy`) is therefore routed to `arc-runner-set` when `github.repository == 'mudler/LocalAI'`. It needs no fork-safety clause because it only triggers on push-to-master and `workflow_dispatch`, so it never executes pull-request code. The repository guard keeps forks (which have no such runner label) from queueing forever. It fetches its own toolchains via `setup-go` / `actions-hugo` and uses no `sudo`/`apt`.
#### What the `arc-runner-set` image actually contains
Measured 2026-07-31 on run `30637392862` by a preflight step, not assumed:
| present | **absent** |
|---|---|
| `git`, `curl`, `unzip`, `tar`, `ldd`, `python3` | **`make`**, **`gcc`** |
That is why `lint.yml` is **not** on the self-hosted pool. Both of its jobs were routed there and both failed in one second: `golangci-lint` needs `make` (for `make protogen-go`, itself needing `curl`+`unzip` to fetch protoc, and for `make lint`), and `build-scripts` additionally needs a C toolchain because the packaging-script tests compile a throwaway binary and inspect it with `ldd`. Both jobs are back on `ubuntu-latest`.
The preflight steps were deliberately left in place. They cost about a second on the hosted pool and mean that whenever the runner image gains `make` + `gcc`, re-routing is one `runs-on:` line per job and any remaining gap reports itself by name rather than as an opaque mid-build failure.
Note for any future re-route: `lint.yml` also triggers on `pull_request`, and a fork PR runs untrusted contributor code. That must never reach a persistent self-hosted runner, so any re-route has to stay push-only, e.g. `${{ (github.event_name == 'push' && github.repository == 'mudler/LocalAI') && 'arc-runner-set' || 'ubuntu-latest' }}`.
## Touching the cache pipeline
When changing `image_build.yml`, `backend_build.yml`, any of the `backend/Dockerfile.*` files, `Dockerfile.base-grpc-builder`, `.docker/install-base-deps.sh`, `.docker/<backend>-compile.sh`, or `scripts/changed-backends.js`:

View File

@@ -70,3 +70,37 @@ The project documentation is located in `docs/content`. When adding new features
- **Configuration**: If you modify configuration options, update the relevant sections in `docs/content/`.
- **Examples**: providing concrete examples (like YAML configuration blocks) is highly encouraged to help users get started quickly.
- **Shortcodes**: Use `{{% notice note %}}`, `{{% notice tip %}}`, or `{{% notice warning %}}` for callout boxes. Do **not** use `{{% alert %}}` — that shortcode does not exist in this project's Hugo theme and will break the docs build.
## React UI styling
The React UI ships a design system in `core/http/react-ui/src/App.css`: design
tokens, form grids, data tables, stat cards, callouts, plus a small semantic
primitive layer (`.stack`, `.hstack`, `.text-note`, `.text-meta`, `.tone-*`,
`.icon-chip`). **Use it instead of `style={{ ... }}`.** Inline styles are a
spacing or colour decision made in one file, so no two pages end up sharing a
rhythm, which is the main reason the app reads as unfinished.
Inline styles are still correct for values that are genuinely computed at
runtime: `width: ${pct}%`, a data-driven `background`, a tooltip's coordinates.
Everything else belongs in a class.
A ratchet enforces this:
```sh
cd core/http/react-ui
npm run lint:inline-styles # fails if the count went UP
npm run lint:inline-styles:report # per-file counts, worst first
npm run lint:inline-styles:write # refresh the baseline after converting
```
The gate also fails on **duplicate `className` attributes on one element**. JSX
keeps the last and silently drops the first, so `<i className={icon}
className="text-xs" />` loses its icon while passing lint, the build and the e2e
suite. Converting a style to a class on an element that already has a
`className` is the usual way to introduce one; merge them into a single
attribute instead.
When converting a page, prefer naming the shapes it actually has
(`.p2p-diagram`, `.usage-tile`) over adding more utilities, and check whether an
existing block already covers it: the Nodes page reuses the P2P setup shapes,
and Model Editor reuses the Settings section rail.

View File

@@ -0,0 +1,26 @@
# Preparing a Release
A release is not finished when the tag is pushed. The GitHub release, the blog post and the demo clips ship together, because the changelog says what moved and the post and the clips are what make anyone care.
## What a release must include
1. **Labels on the merged PRs.** GitHub generates the raw notes from PR labels, so label first, generate second. Wrong labels mean a miscategorised changelog that has to be edited by hand.
2. **`RELEASE_NOTES_vX.Y.Z.md`** at the repository root, in the house style: what changed, why it matters, PR numbers so people can read the diffs.
3. **A blog post under `website/content/blog/`.** One post per release, front matter with `title`, `date`, `author`, `category: "Release"`, `tags`, `summary` and `extracss: ["blog.css"]`. Cover the two or three changes that alter what a user does day to day, not the whole changelog, and link the PR numbers. See `website/content/blog/what-landed-in-localai-4-8.md` for the shape.
4. **Demo clips for the notable features.** Anything visible (a new backend, a UI change, a new endpoint, a measured speedup) gets a short screen recording. Put the file in `website/static/media/`, reference it from the blog post, and reuse it on the marketing pages where it fits.
A release without a post and without clips is incomplete, in the same way a user-facing code change without a docs update is incomplete.
## Clip conventions
- MP4, H.264, no audio track unless the feature is about audio. Keep them short (10 to 30 seconds) and loopable.
- Record the real thing. A clip from the engine's own benchmark suite or a real session, never a mockup.
- Where the change is a speedup, record both sides on the same machine on the same input, so the comparison is honest.
- Name the file after the feature, not the release (`vllm-race.mp4`, not `v4-8-demo.mp4`), so it stays reusable once the release is old.
- The marketing site plays clips with `muted loop playsinline preload="none"` and a `data-lazy` attribute, which the site's IntersectionObserver uses to play and pause them on scroll. Follow that pattern for anything you add.
## Order of work
Label the PRs, generate and edit the release notes, cut the draft release, record the clips while the branch is still fresh in your head, then write the post against the notes and the clips. Publishing the release and merging the post should happen on the same day.
The `creating-localai-releases` skill drives steps 1 to 3 and captures the React UI screenshots that go into the notes.

View File

@@ -21,6 +21,17 @@ options:
- reasoning_parser:qwen3
```
## `Options[]` doubles as CLI-style engine flags
Beyond the parser names above, `Options[]` carries `--` prefixed engine flags (`--enable-prefix-caching`, `--kv-cache-dtype:fp8_e5m2`). `apply_options_to_engine_args` in `backend/python/common/vllm_utils.py` maps them onto `AsyncEngineArgs` fields, and it must run **before** `AsyncLLMEngine.from_engine_args()` - applying them afterwards is a silent no-op, which is exactly what issue #11130 was.
Things to keep straight when touching this:
- Precedence is typed proto fields → `options:``engine_args:`. `applyEngineArgDefaults` in `core/config/hooks_vllm.go` therefore skips seeding a production default whose key the user already set as an option, otherwise the later `engine_args:` pass would silently override them.
- Only `--` prefixed entries are engine flags; `tool_parser:`/`reasoning_parser:` and friends keep their meaning. Parser lookups accept both spellings via `normalize_option_key`.
- Unknown or uncoercible flags warn and are skipped, unlike `engine_args:` which is strict - `Options[]` is a shared bag and knows entries this mapping doesn't.
- Field types come from the annotation's *base* (`Literal["auto","float16"]` is not a float). The helper's tests are stdlib-only: `make test-python-helpers`.
Auto-defaults for known model families live in `core/config/parser_defaults.json` and are applied:
- at gallery import time by `core/gallery/importers/vllm.go`
- at model load time by the `vllm` / `vllm-omni` backend hook in `core/config/hooks_vllm.go`

View File

@@ -113,6 +113,54 @@ if [ "${BUILD_TYPE:-}" = "vulkan" ] && [ "${SKIP_DRIVERS:-false}" = "false" ]; t
rm -rf /var/lib/apt/lists/*
fi
# --- 2b. Intel graphics driver (BUILD_TYPE=sycl*) ---
# The Intel oneAPI base image brings the compilers and the oneAPI libraries, but
# not the driver that talks to the graphics card. The packaging step copies that
# driver into the backend, so that the backend works on a machine which has no
# Intel graphics packages of its own, for the same reason the Vulkan section
# above installs the Mesa drivers. Install it here so there is something to copy.
#
# Only the sycl builds are covered, because those are the ones whose packaging
# copies the driver. See package_intel_libs in scripts/build/package-gpu-libs.sh.
#
# The driver comes from Intel's own package repository, not from the Ubuntu
# archive. The archive has 23.43 from late 2023, which does not know any card
# released since, so a machine with a recent Intel GPU would end up carrying a
# driver that cannot drive it. Intel's repository has 25.18 for the same Ubuntu
# release.
#
# Anything that goes wrong here fails the build, on purpose. An unreachable
# repository is a passing problem that a retry fixes, whereas carrying a
# different driver than intended, or none, is a difference nobody would notice
# until a user reports an idle GPU.
if case "${BUILD_TYPE:-}" in sycl*) true;; *) false;; esac \
&& [ "${SKIP_DRIVERS:-false}" = "false" ]; then
# Ubuntu release name, which is what the repository is indexed by.
ubuntu_codename=$(. /etc/os-release && echo "${VERSION_CODENAME:-}")
if [ -z "$ubuntu_codename" ]; then
echo "ERROR: cannot tell which Ubuntu release this image is, so cannot pick the Intel driver repository" >&2
exit 1
fi
# The key is armored text, which apt reads directly from a .asc file, so
# there is no need for gnupg here. "unified" is the component Intel ships
# its current driver in.
mkdir -p /usr/share/keyrings
curl -fsSL https://repositories.intel.com/gpu/intel-graphics.key \
-o /usr/share/keyrings/intel-graphics.asc
echo "deb [arch=amd64 signed-by=/usr/share/keyrings/intel-graphics.asc] https://repositories.intel.com/gpu/ubuntu ${ubuntu_codename} unified" \
> /etc/apt/sources.list.d/intel-graphics.list
apt-get update
# The first package holds the driver OpenCL talks to, the second the driver
# Level Zero talks to. Between them they pull in the compiler and the memory
# manager that both need.
apt-get install -y --no-install-recommends \
intel-opencl-icd \
libze-intel-gpu1
apt-get clean
rm -rf /var/lib/apt/lists/*
fi
# --- 3. CUDA toolkit (BUILD_TYPE=cublas|l4t) ---
if { [ "${BUILD_TYPE:-}" = "cublas" ] || [ "${BUILD_TYPE:-}" = "l4t" ]; } && [ "${SKIP_DRIVERS:-false}" = "false" ]; then
apt-get update

View File

@@ -0,0 +1,32 @@
#!/usr/bin/env bash
set -euo pipefail
arch=${1:?target architecture is required}
build_type=${2-}
# SYCL compiles the whole tree with icpx -fsycl, and icpx never finishes
# ggml-cpu/arch/x86/repack.cpp at -march=sapphirerapids: the job sits on that one
# translation unit until GitHub kills it at 6h. gcc builds the same file in
# seconds, so only the SYCL images have to give up the CPU variant matrix.
#
# ROCm runs out of the same 6h budget for a different reason: volume, not a
# stall. hipcc compiles ggml's HIP kernels once per entry in AMDGPU_TARGETS,
# which is eleven architectures (gfx908 through gfx1201), and the CPU variant
# matrix lands on top of that. The job built in 2h27m before it was added and
# has been killed at exactly 6h00m on every run since, so no ROCm llama-cpp
# image has been published since 2026-08-01.
case "$build_type" in
sycl*|hipblas*)
echo llama-cpp-fallback
exit 0
;;
esac
# GPU arm64 base images do not consistently provide the gcc-14 toolchain needed
# to compile ggml's armv9.2 CPU variants. Keep their portable fallback until the
# builder images can supply that compiler.
if [ "$arch" = "arm64" ] && [ -n "$build_type" ]; then
echo llama-cpp-fallback
else
echo llama-cpp-cpu-all
fi

View File

@@ -18,27 +18,27 @@ if [[ -n "${CUDA_DOCKER_ARCH:-}" ]]; then
fi
cd /LocalAI/backend/cpp/llama-cpp
if [ -z "${BUILD_TYPE:-}" ]; then
# Pure CPU image (BUILD_TYPE empty): one build with ggml CPU_ALL_VARIANTS replaces the
# per-microarch binaries (x86: avx/avx2/avx512/fallback; arm64: armv8.x/armv9.x). ggml
# dlopens the best libggml-cpu-*.so at runtime by probing host CPU features.
BUILD_TARGET=$(/LocalAI/.docker/llama-cpp-build-target.sh "${TARGETARCH}" "${BUILD_TYPE:-}")
if [ "$BUILD_TARGET" = "llama-cpp-cpu-all" ]; then
# One build with ggml CPU_ALL_VARIANTS replaces the per-microarch binaries (x86:
# avx/avx2/avx512/fallback; arm64: armv8.x/armv9.x). BUILD_TYPE remains in the
# environment, so GPU builds retain their accelerator backend while ggml dlopens the
# best CPU library when work is offloaded to the host.
#
# arm64: the CPU_ALL_VARIANTS table includes armv9.2 SME variants whose -march=...+sme is
# rejected by the Ubuntu 24.04 default gcc-13. gcc-14 accepts it, so build the arm64
# variants with it (the host never *selects* SME unless it has it, but every variant must
# still compile).
if [ "${TARGETARCH}" = "arm64" ]; then
# The prebuilt base inherits default ports.ubuntu.com sources; honor the
# APT_*_MIRROR build args here like the from-source path does, so this
# apt step survives a mirror outage.
sh /LocalAI/.docker/apt-mirror.sh || true
apt-get update -qq && apt-get install -y -qq gcc-14 g++-14
export CC=gcc-14 CXX=g++-14
fi
make llama-cpp-cpu-all
else
# GPU build (cublas/hipblas/sycl/vulkan/...): the accelerator does the compute, so a
# single fallback CPU build is enough - no per-microarch CPU variants needed. (This also
# keeps the heavy GPU backend compile from also building the whole CPU variant matrix,
# and avoids the gcc-14 apt step on GPU base images such as nvidia l4t.)
make llama-cpp-fallback
fi
make "$BUILD_TARGET"
make llama-cpp-grpc
make llama-cpp-rpc-server

View File

@@ -0,0 +1,25 @@
#!/usr/bin/env bash
set -euo pipefail
arch=${1:?target architecture is required}
build_type=${2-}
# SYCL compiles the whole tree with icpx -fsycl, and icpx never finishes
# ggml-cpu/arch/x86/repack.cpp at -march=sapphirerapids: the job sits on that one
# translation unit until GitHub kills it at 6h. gcc builds the same file in
# seconds, so only the SYCL images have to give up the CPU variant matrix.
case "$build_type" in
sycl*)
echo turboquant-fallback
exit 0
;;
esac
# GPU arm64 base images do not consistently provide the gcc-14 toolchain needed
# to compile ggml's armv9.2 CPU variants. Keep their portable fallback until the
# builder images can supply that compiler.
if [ "$arch" = "arm64" ] && [ -n "$build_type" ]; then
echo turboquant-fallback
else
echo turboquant-cpu-all
fi

View File

@@ -19,20 +19,18 @@ fi
cd /LocalAI/backend/cpp/turboquant
if [ -z "${BUILD_TYPE:-}" ]; then
# Pure CPU image: one ggml CPU_ALL_VARIANTS build replaces the per-microarch binaries.
BUILD_TARGET=$(/LocalAI/.docker/turboquant-build-target.sh "${TARGETARCH}" "${BUILD_TYPE:-}")
if [ "$BUILD_TARGET" = "turboquant-cpu-all" ]; then
# BUILD_TYPE remains in the environment, so GPU builds retain their accelerator while
# ggml selects the best CPU library when model work is offloaded to the host.
# arm64: the armv9.2 SME variants need gcc-14 (gcc-13 rejects +sme).
if [ "${TARGETARCH}" = "arm64" ]; then
sh /LocalAI/.docker/apt-mirror.sh || true
apt-get update -qq && apt-get install -y -qq gcc-14 g++-14
export CC=gcc-14 CXX=g++-14
fi
make turboquant-cpu-all
else
# GPU build (cublas/hipblas/sycl/vulkan/...): single fallback CPU build, the accelerator
# does the compute. Keeps the GPU compile from also building the CPU variant matrix and
# avoids the gcc-14 apt step on GPU base images such as nvidia l4t.
make turboquant-fallback
fi
make "$BUILD_TARGET"
make turboquant-grpc
make turboquant-rpc-server

View File

@@ -40,6 +40,16 @@ backend/cpp/privacy-filter/build
backend/cpp/privacy-filter/grpc-server
backend/cpp/privacy-filter/package
# audio-cpp: same in-place pattern. The Makefile clones audio.cpp at the pinned
# AUDIO_CPP_VERSION and the `audio.cpp:` target is the directory itself, so a
# stale host checkout COPY'd in makes the build compile against whatever commit
# the host had. build/ is worse than stale: its CMakeCache.txt records the host
# source, prefix and compiler paths, and cmake refuses to reconfigure from it.
backend/cpp/audio-cpp/audio.cpp
backend/cpp/audio-cpp/build
backend/cpp/audio-cpp/grpc-server
backend/cpp/audio-cpp/package
# Rust backend build output (sources are tracked; target/ is generated)
backend/rust/*/target

View File

@@ -1,72 +0,0 @@
#!/usr/bin/env sh
#
# LocalAI pre-commit hook. Install it (once per clone) with:
#
# make install-hooks
#
# Runs only the checks relevant to what's staged:
# - Go files -> make lint + make test-coverage-check
# - core/http/react-ui -> make test-ui-coverage-check (Playwright e2e + gate)
# - realtime state machines / specs -> make test-realtime-conformance
# (respcoord/**, turncoord/**, or formal-verification/** -- a pure .fizz
# spec edit must still re-verify the design, detected separately from Go)
# A commit touching none of these is skipped entirely (other docs/YAML can't
# change lint findings, Go coverage, the UI, or the realtime conformance gate).
#
# To bypass for a single commit (e.g. a WIP checkpoint): git commit --no-verify
set -eu
repo_root="$(git rev-parse --show-toplevel)"
cd "$repo_root"
staged="$(git diff --cached --name-only --diff-filter=ACMRD)"
go_changed=0
ui_changed=0
rt_changed=0
if echo "$staged" | grep -qE '\.go$'; then go_changed=1; fi
if echo "$staged" | grep -qE '^core/http/react-ui/'; then ui_changed=1; fi
if echo "$staged" | grep -qE '^(core/http/endpoints/openai/(coordinator|respcoord|turncoord|conncoord|compactcoord|ttscoord)/|formal-verification/)'; then rt_changed=1; fi
if [ "$go_changed" -eq 0 ] && [ "$ui_changed" -eq 0 ] && [ "$rt_changed" -eq 0 ]; then
echo "pre-commit: no Go, React UI, or realtime-spec changes staged — skipping."
exit 0
fi
if [ "$go_changed" -eq 1 ]; then
# Resolve the ref golangci-lint's new-from-merge-base should compare
# against. .golangci.yml pins origin/master, which is correct in CI
# (origin == the canonical repo) but wrong from a fork clone, where
# origin/master lags behind and lint would report the whole upstream
# backlog. Prefer upstream/master, then origin/master, then master.
lint_base=""
for ref in upstream/master origin/master master; do
if git rev-parse --verify --quiet "${ref}^{commit}" >/dev/null 2>&1; then
lint_base="$ref"
break
fi
done
echo "pre-commit ▶ golangci-lint (make lint${lint_base:+, new-from $lint_base})"
make lint LINT_NEW_FROM="$lint_base"
echo "pre-commit ▶ coverage gate (make test-coverage-check) — builds and runs the"
echo " pkg/core suites plus tests/e2e; can take a few minutes."
make test-coverage-check
fi
if [ "$ui_changed" -eq 1 ]; then
echo "pre-commit ▶ React UI e2e + coverage gate (make test-ui-coverage-check) —"
echo " rebuilds the UI + ui-test-server, runs the Playwright specs, and"
echo " fails if line coverage regressed; can take a couple of minutes."
make test-ui-coverage-check
fi
if [ "$rt_changed" -eq 1 ]; then
echo "pre-commit ▶ realtime state-machine conformance (make test-realtime-conformance) —"
echo " Go transition/rapid tests under -race + FizzBee model check of the"
echo " authoritative specs. Fail-closed: needs FizzBee (make install-fizzbee)."
make test-realtime-conformance
fi
echo "pre-commit ✓ all relevant checks passed"

View File

@@ -66,6 +66,34 @@ include:
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-kokoro'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'true'
backend: "kokoro"
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-kokoro'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'true'
backend: "kokoro"
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
@@ -728,6 +756,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-12-trellis2cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "trellis2cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
@@ -871,6 +912,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-12-magpie-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
@@ -1675,6 +1729,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13-trellis2cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "trellis2cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1688,6 +1755,19 @@ include:
backend: "stablediffusion-ggml"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-cuda-13-arm64-trellis2cpp'
base-image: "ubuntu:24.04"
ubuntu-version: '2404'
runs-on: 'ubuntu-24.04-arm'
backend: "trellis2cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1935,6 +2015,32 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13-magpie-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13-vllm-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "vllm-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -2000,6 +2106,32 @@ include:
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/arm64'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-cuda-13-arm64-vllm-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "vllm-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-cuda-13-arm64-magpie-tts-cpp'
base-image: "ubuntu:24.04"
ubuntu-version: '2404'
runs-on: 'ubuntu-24.04-arm'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -2963,6 +3095,19 @@ include:
dockerfile: "./backend/Dockerfile.privacy-filter"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-vllm-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "vllm-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# Vulkan: base-grpc-vulkan-amd64 carries the SDK. arm64 vulkan is a one-line
# add once amd64 is proven in CI.
- build-type: 'vulkan'
@@ -2998,6 +3143,97 @@ include:
dockerfile: "./backend/Dockerfile.privacy-filter"
context: "./"
ubuntu-version: '2404'
# audio-cpp: 0xShug0/audio.cpp, a multi-family ggml audio engine (TTS, ASR,
# VAD, diarization, source separation, music generation).
#
# These entries deliberately carry NO builder-base-image, unlike the
# privacy-filter and llama-cpp blocks above. The prebuilt
# quay.io/go-skynet/ci-cache:base-grpc-* images ship a from-source gRPC whose
# protobuf is v26, and protobuf has depended on abseil since v22. audio.cpp
# links sentencepiece with SPM_PROTOBUF_PROVIDER=package (needed to stop
# sentencepiece's vendored protobuf 3.14 from colliding with the 3.21 the
# generated backend.pb.cc is built against, which broke every nested-message
# parse), so sentencepiece then sees real abseil's
# `absl::lts_20240116::internal` alongside its own vendored plain
# `absl::internal` and every `absl::internal::` reference becomes ambiguous.
# Verified, not theorised: building against base-grpc-amd64 fails at
# sentencepiece-static.dir/error.cc.o with "reference to 'internal' is
# ambiguous". Dockerfile.audio-cpp therefore installs Ubuntu Noble's apt
# gRPC/protobuf 3.21.12 itself and has a single `builder` stage, so the
# BUILDER_BASE_IMAGE / BUILDER_TARGET / SKIP_DRIVERS build-args are never
# consumed. Same reason CUDA needs its toolkit in base-image rather than in a
# builder image: this is the ds4 shape, not the llama-cpp one.
#
# No ROCm entry: upstream has no HIP configuration. No CUDA arm64 or L4T
# entry: upstream documents and validates CUDA on x86 only. Darwin/Metal is in
# the includeDarwin matrix below, built by scripts/build/audio-cpp-darwin.sh.
#
# No vulkan entry either, though Dockerfile.audio-cpp and the backend Makefile
# both handle BUILD_TYPE=vulkan for local builds. Every other vulkan backend
# gets its Mesa ICD drivers from .docker/install-base-deps.sh, which installs
# mesa-vulkan-drivers so package-gpu-libs.sh can bundle them; this Dockerfile
# calls neither, so the image would ship a Vulkan loader that finds no GPU. No
# CI job runs a vulkan image against real hardware, so it would pass green and
# fail in users' hands. The entry comes back once the ICD question is settled.
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-audio-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'true'
backend: "audio-cpp"
dockerfile: "./backend/Dockerfile.audio-cpp"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-audio-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'true'
backend: "audio-cpp"
dockerfile: "./backend/Dockerfile.audio-cpp"
context: "./"
ubuntu-version: '2404'
# cuda-major-version is forwarded into the build (Dockerfile.audio-cpp -> the
# backend Makefile) and picks the CMAKE_CUDA_ARCHITECTURES list, which upstream
# otherwise sets to `native` and no CI runner can enumerate. cuda-minor-version
# and the base-image tag encode the same toolkit and must move together;
# nothing checks that for you.
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-12-audio-cpp'
runs-on: 'ubuntu-latest'
base-image: "nvidia/cuda:12.8.1-devel-ubuntu24.04"
skip-drivers: 'true'
backend: "audio-cpp"
dockerfile: "./backend/Dockerfile.audio-cpp"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13-audio-cpp'
runs-on: 'ubuntu-latest'
base-image: "nvidia/cuda:13.0.0-devel-ubuntu24.04"
skip-drivers: 'true'
backend: "audio-cpp"
dockerfile: "./backend/Dockerfile.audio-cpp"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
@@ -3161,6 +3397,35 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# trellis2cpp
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-trellis2cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "trellis2cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-trellis2cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "trellis2cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# sam3-cpp
- build-type: ''
cuda-major-version: ""
@@ -3486,6 +3751,34 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-trellis2cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "trellis2cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-trellis2cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "trellis2cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "0"
@@ -3499,6 +3792,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2204'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-arm64-trellis2cpp'
base-image: "nvcr.io/nvidia/l4t-jetpack:r36.4.0"
runs-on: 'ubuntu-24.04-arm'
backend: "trellis2cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2204'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "0"
@@ -4583,6 +4889,20 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-magpie-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
@@ -4597,7 +4917,50 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# vllm-cpp
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-vllm-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "vllm-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-vllm-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "vllm-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# omnivoice-cpp
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-magpie-tts-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
@@ -4652,6 +5015,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f32'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-intel-sycl-f32-magpie-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "intel/oneapi-basekit:2025.3.0-0-devel-ubuntu24.04"
skip-drivers: 'false'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f32'
cuda-major-version: ""
cuda-minor-version: ""
@@ -4691,6 +5067,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f16'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-intel-sycl-f16-magpie-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "intel/oneapi-basekit:2025.3.0-0-devel-ubuntu24.04"
skip-drivers: 'false'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f16'
cuda-major-version: ""
cuda-minor-version: ""
@@ -4732,6 +5121,20 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-magpie-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
@@ -4774,6 +5177,20 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-magpie-tts-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
@@ -4814,6 +5231,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2204'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-arm64-magpie-tts-cpp'
base-image: "nvcr.io/nvidia/l4t-jetpack:r36.4.0"
runs-on: 'ubuntu-24.04-arm'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2204'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "0"
@@ -4853,6 +5283,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'hipblas'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-magpie-tts-cpp'
base-image: "rocm/dev-ubuntu-24.04:6.4.4"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "magpie-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'hipblas'
cuda-major-version: ""
cuda-minor-version: ""
@@ -5193,6 +5636,35 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# valkey-store
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-valkey-store'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "valkey-store"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-valkey-store'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "valkey-store"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# rfdetr
- build-type: ''
cuda-major-version: ""
@@ -5734,6 +6206,10 @@ includeDarwin:
tag-suffix: "-metal-darwin-arm64-stablediffusion-ggml"
build-type: "metal"
lang: "go"
- backend: "trellis2cpp"
tag-suffix: "-metal-darwin-arm64-trellis2cpp"
build-type: "metal"
lang: "go"
- backend: "whisper"
tag-suffix: "-metal-darwin-arm64-whisper"
build-type: "metal"
@@ -5774,6 +6250,14 @@ includeDarwin:
tag-suffix: "-metal-darwin-arm64-moss-tts-cpp"
build-type: "metal"
lang: "go"
- backend: "magpie-tts-cpp"
tag-suffix: "-metal-darwin-arm64-magpie-tts-cpp"
build-type: "metal"
lang: "go"
- backend: "vllm-cpp"
tag-suffix: "-metal-darwin-arm64-vllm-cpp"
build-type: "metal"
lang: "go"
- backend: "omnivoice-cpp"
tag-suffix: "-metal-darwin-arm64-omnivoice-cpp"
build-type: "metal"
@@ -5807,6 +6291,18 @@ includeDarwin:
- backend: "privacy-filter"
tag-suffix: "-metal-darwin-arm64-privacy-filter"
lang: "go"
# audio-cpp is the same shape: a C++/ggml backend built by a bespoke darwin
# script (make backends/audio-cpp-darwin), which reuses the backend's own
# package.sh so the Darwin package keeps the root-level layout the Linux image
# has (grpc-server, run.sh and assets/ in one directory, dylibs in lib/).
# No build-type: the backend Makefile turns ENGINE_ENABLE_METAL on from
# uname -s. lang=go drives runner/toolchain selection only - there is no
# backend/go/audio-cpp, which is why backend_build_darwin.yml and
# DARWIN_BESPOKE_BUILDERS in scripts/lib/backend-filter.mjs both route this
# backend away from the generic Go path.
- backend: "audio-cpp"
tag-suffix: "-metal-darwin-arm64-audio-cpp"
lang: "go"
# LocalVQE has no Metal path; on Apple Silicon it builds CPU-only (GGML_METAL
# OFF) but is still a native arm64 image. Uses the darwin/metal build profile.
- backend: "localvqe"
@@ -5903,6 +6399,10 @@ includeDarwin:
tag-suffix: "-metal-darwin-arm64-cloud-proxy"
build-type: "metal"
lang: "go"
- backend: "valkey-store"
tag-suffix: "-metal-darwin-arm64-valkey-store"
build-type: "metal"
lang: "go"
- backend: "llama-cpp-quantization"
tag-suffix: "-metal-darwin-arm64-llama-cpp-quantization"
build-type: "mps"

76
.github/ci/gen-redirects.sh vendored Executable file
View File

@@ -0,0 +1,76 @@
#!/usr/bin/env bash
#
# Generate client-side redirects for the documentation URLs that used to live at
# the site root.
#
# Until this site existed, the Hugo docs site WAS localai.io, so pages
# were published at /features/..., /getting-started/..., /faq/ and so on. The
# docs now build under /docs/, and GitHub Pages serves static files only: there
# is no server-side rewrite, no .htaccess, no _redirects. The only way to keep
# every published, bookmarked and search-indexed URL alive is to leave a real
# HTML file at the old address that sends the browser to the new one.
#
# Anything the main site already publishes wins: it owns /, /engines/,
# /blog/ and friends, so an existing file is never replaced.
#
# Usage: gen-redirects.sh <public-dir> [base-url]
# public-dir merged output directory (main site with docs/ inside it)
# base-url absolute or root-relative prefix the deployment is served from,
# trailing slash optional (default "/")
set -euo pipefail
PUBLIC_DIR=${1:?usage: gen-redirects.sh <public-dir> [base-url]}
BASE_URL=${2:-/}
# Normalise to exactly one trailing slash so concatenation below is predictable.
BASE_URL="${BASE_URL%/}/"
DOCS_DIR="${PUBLIC_DIR}/docs"
if [ ! -d "$DOCS_DIR" ]; then
echo "gen-redirects: no docs output at ${DOCS_DIR}" >&2
exit 1
fi
created=0
skipped=0
# Every .html file is a reachable old URL, not just directory indexes: the
# generated model gallery ships as a bare gallery.html and used to sit at the
# root too.
while IFS= read -r src; do
rel=${src#"$DOCS_DIR"/}
dst="${PUBLIC_DIR}/${rel}"
if [ -e "$dst" ]; then
skipped=$((skipped + 1))
continue
fi
# Link to the directory, not to its index.html, so the redirect target is the
# canonical URL the docs site itself advertises.
target="${BASE_URL}docs/${rel%index.html}"
mkdir -p "$(dirname "$dst")"
printf '%s' '<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>Moved</title>
<link rel="canonical" href="'"$target"'">
<meta name="robots" content="noindex">
<meta http-equiv="refresh" content="0; url='"$target"'">
</head>
<body>
<p>This page moved to <a href="'"$target"'">'"$target"'</a>.</p>
</body>
</html>
' > "$dst"
created=$((created + 1))
done <<EOF
$(find "$DOCS_DIR" -type f -name '*.html' | sort)
EOF
echo "gen-redirects: ${created} redirect(s) written, ${skipped} path(s) left to the main site"

64
.github/ci/refresh-site-counters.sh vendored Executable file
View File

@@ -0,0 +1,64 @@
#!/usr/bin/env bash
# Refreshes the counters shown on the landing page from the GitHub API.
#
# The numbers used to be typed into the templates by hand, which meant they
# only moved when somebody remembered, and a stale star count on the front
# page is worse than no star count. Everything the API can answer for lives
# in website/data/stats.yaml and is rewritten wholesale by this script.
#
# Anything the API cannot answer for (the Discord member count) is read back
# out of the existing file and carried through untouched.
set -euo pipefail
REPO="${REPO:-mudler/LocalAI}"
OUT="${OUT:-website/data/stats.yaml}"
# The contributors and releases endpoints are paginated and never report a
# total. Asking for one item per page makes the last page number equal to the
# item count, which the Link header hands over.
count_via_link_header() {
local path="$1" link last
link=$(gh api -i "${path}?per_page=1" 2>/dev/null | tr -d '\r' | grep -i '^link:' || true)
if [ -z "$link" ]; then
# No Link header means a single page, so count that page directly.
gh api "${path}?per_page=100" --jq 'length'
return
fi
last=$(sed -n 's/.*[?&]page=\([0-9]*\)>; rel="last".*/\1/p' <<<"$link")
[ -n "$last" ] || { gh api "${path}?per_page=100" --jq 'length'; return; }
printf '%s\n' "$last"
}
read -r stars forks < <(gh api "repos/${REPO}" --jq '"\(.stargazers_count) \(.forks_count)"')
contributors=$(count_via_link_header "repos/${REPO}/contributors")
releases=$(count_via_link_header "repos/${REPO}/releases")
# Not derivable from the GitHub API, so keep whatever is already on disk.
discord=$(sed -n 's/^discord: *\([0-9]*\).*/\1/p' "$OUT" 2>/dev/null | head -1)
discord="${discord:-0}"
for n in stars forks contributors releases; do
v="${!n}"
[[ "$v" =~ ^[0-9]+$ ]] && [ "$v" -gt 0 ] || {
echo "refusing to write: ${n} came back as '${v}'" >&2
exit 1
}
done
cat > "$OUT" <<YAML
# Counters shown on the landing page.
#
# The four GitHub fields are rewritten by .github/ci/refresh-site-counters.sh,
# which runs weekly from .github/workflows/refresh-site-counters.yml. Editing
# them by hand works but will be overwritten on the next run.
stars: ${stars}
forks: ${forks}
contributors: ${contributors}
releases: ${releases}
# The GitHub API cannot answer for this one, so it is maintained by hand and
# the refresh script carries it through untouched.
discord: ${discord}
YAML
echo "stars=${stars} forks=${forks} contributors=${contributors} releases=${releases} discord=${discord}"

View File

@@ -244,8 +244,19 @@ jobs:
make protogen-go
make backends/privacy-filter-darwin
# audio-cpp is a C++/ggml backend like ds4 and privacy-filter - a single
# grpc-server with otool dylib bundling, plus bundled VAD assets - so it
# gets its own bespoke darwin script rather than the generic
# build-darwin-go-backend path, which would look for a backend/go/audio-cpp
# that does not exist. Keep this set in sync with DARWIN_BESPOKE_BUILDERS
# in scripts/lib/backend-filter.mjs.
- name: Build audio-cpp backend (Darwin Metal)
if: inputs.backend == 'audio-cpp'
run: |
make backends/audio-cpp-darwin
- name: Build ${{ inputs.backend }}-darwin
if: inputs.backend != 'llama-cpp' && inputs.backend != 'ds4' && inputs.backend != 'privacy-filter'
if: inputs.backend != 'llama-cpp' && inputs.backend != 'ds4' && inputs.backend != 'privacy-filter' && inputs.backend != 'audio-cpp'
run: |
make protogen-go
BACKEND=${{ inputs.backend }} BUILD_TYPE=${{ inputs.build-type }} USE_PIP=${{ inputs.use-pip }} make build-darwin-${{ inputs.lang }}-backend

View File

@@ -5,6 +5,24 @@ on:
branches:
- master
pull_request:
# GoReleaser and the darwin launcher take no gallery, docs or markdown
# input, so a diff confined to these paths cannot change either binary.
# The darwin job matters here: macOS is the scarcest runner class.
#
# backend/{cpp,go,python}/**: GoReleaser builds ./cmd/local-ai and the
# launcher builds ./cmd/launcher; neither compiles a backend. The
# before-hooks still matter, but their input is backend/backend.proto
# (protogen-go) and go.mod/go.sum (go mod tidy), none of which live under
# these prefixes, so a change to any of those still triggers a full run.
# See .agents/ci-caching.md.
paths-ignore:
- 'gallery/**'
- 'docs/**'
- 'examples/**'
- '**/*.md'
- 'backend/cpp/**'
- 'backend/go/**'
- 'backend/python/**'
# Supersede an in-flight run when a PR gets a new push. Keyed on the PR number
# so every push to the same PR shares a group; on a master push the key falls
@@ -26,9 +44,15 @@ jobs:
uses: actions/setup-go@v5
with:
go-version: 1.25
# A PR builds only the host target. The three-platform cross-compile
# (linux/amd64, linux/arm64, darwin/arm64) is the bulk of this job's
# ~6.6min median and no PR consumes the resulting binaries. The
# before-hooks (protogen-go, react-ui, go mod tidy) run either way, so the
# "is the release build broken" signal is unchanged. master and tags still
# build everything.
- name: Run GoReleaser
run: |
make dev-dist
make ${{ github.event_name == 'pull_request' && 'dev-dist-single' || 'dev-dist' }}
launcher-build-darwin:
runs-on: macos-latest
steps:

View File

@@ -30,6 +30,10 @@ jobs:
variable: "DS4_VERSION"
branch: "main"
file: "backend/cpp/ds4/Makefile"
- repository: "0xShug0/audio.cpp"
variable: "AUDIO_CPP_VERSION"
branch: "main"
file: "backend/cpp/audio-cpp/Makefile"
- repository: "meituan-longcat/LongCat-Video"
variable: "LONGCAT_VIDEO_VERSION"
branch: "main"
@@ -50,6 +54,10 @@ jobs:
variable: "PARAKEET_VERSION"
branch: "master"
file: "backend/go/parakeet-cpp/Makefile"
- repository: "mudler/vllm.cpp"
variable: "VLLM_CPP_VERSION"
branch: "main"
file: "backend/go/vllm-cpp/Makefile"
- repository: "localai-org/moss-transcribe.cpp"
variable: "MOSS_VERSION"
branch: "master"
@@ -74,6 +82,10 @@ jobs:
variable: "STABLEDIFFUSION_GGML_VERSION"
branch: "master"
file: "backend/go/stablediffusion-ggml/Makefile"
- repository: "localai-org/trellis2cpp"
variable: "TRELLIS2CPP_VERSION"
branch: "pbr-textures"
file: "backend/go/trellis2cpp/Makefile"
- repository: "mudler/go-piper"
variable: "PIPER_VERSION"
branch: "master"
@@ -98,10 +110,14 @@ jobs:
variable: "LOCATEANYTHING_VERSION"
branch: "master"
file: "backend/go/locate-anything-cpp/Makefile"
- repository: "ServeurpersoCom/qwentts.cpp"
variable: "QWEN3TTS_CPP_VERSION"
branch: "master"
file: "backend/go/qwen3-tts-cpp/Makefile"
# qwentts.cpp is held, not tracked: upstream master hangs in synthesis
# (see the comment on QWEN3TTS_CPP_VERSION in the backend Makefile).
# Leaving it here would re-bump the pin back onto the hang every night.
# Restore this entry once the upstream fix lands.
# - repository: "ServeurpersoCom/qwentts.cpp"
# variable: "QWEN3TTS_CPP_VERSION"
# branch: "master"
# file: "backend/go/qwen3-tts-cpp/Makefile"
- repository: "ServeurpersoCom/omnivoice.cpp"
variable: "OMNIVOICE_VERSION"
branch: "master"
@@ -110,6 +126,10 @@ jobs:
variable: "VIBEVOICE_CPP_VERSION"
branch: "master"
file: "backend/go/vibevoice-cpp/Makefile"
- repository: "mudler/magpie-tts.cpp"
variable: "MAGPIETTS_CPP_VERSION"
branch: "main"
file: "backend/go/magpie-tts-cpp/Makefile"
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7

View File

@@ -1,4 +1,4 @@
name: Deploy docs to GitHub Pages
name: Deploy site to GitHub Pages
on:
push:
@@ -6,9 +6,11 @@ on:
- master
paths:
- 'docs/**'
- 'website/**'
- 'gallery/**'
- 'images/**'
- '.github/ci/modelslist.go'
- '.github/ci/gen-redirects.sh'
- '.github/workflows/gh-pages.yml'
workflow_dispatch:
@@ -23,7 +25,20 @@ concurrency:
jobs:
build:
runs-on: ubuntu-latest
# Self-hosted. This workflow is push-to-master + workflow_dispatch only, so
# it never executes pull-request code and a fork cannot reach the runner
# with untrusted changes. The repository guard keeps forks (whose own master
# pushes would otherwise queue forever against a label they do not have) on
# the hosted pool.
#
# Why: the GitHub-hosted pool is shared account-wide and has repeatedly
# starved (2026-07-31: 35 consecutive minutes at zero scheduled jobs, while
# arc-runner-set kept completing work throughout). Publishing the site is
# small, frequent, and must not sit behind a saturated hosted queue.
#
# Needs only git, tar and curl on the runner: setup-go and actions-hugo
# fetch their own toolchains, and no step uses sudo, apt, make or unzip.
runs-on: ${{ github.repository == 'mudler/LocalAI' && 'arc-runner-set' || 'ubuntu-latest' }}
env:
HUGO_VERSION: "0.146.3"
steps:
@@ -36,7 +51,16 @@ jobs:
- name: Setup Go
uses: actions/setup-go@v5
with:
go-version: '1.22'
# Track go.mod rather than a literal. Pinned at 1.22 this installed a
# toolchain older than the module's `go 1.26.0`, so the `go run` below
# downloaded the real one from proxy.golang.org on every run. That
# fetch is not always reachable from the runner and the deploy failed
# on five of eight consecutive master pushes with:
# go: download go1.26.0: ... connect: network is unreachable
# ##[error]Command failed: go env GOPATH
# Installing the version the module asks for removes the download
# instead of depending on it succeeding.
go-version-file: go.mod
cache: false
- name: Setup Hugo
@@ -49,25 +73,46 @@ jobs:
id: pages
uses: actions/configure-pages@v6
# The gallery page is generated from the model index and shipped as a
# static asset of the docs site, so it has to exist before Hugo runs.
- name: Generate gallery
run: go run ./.github/ci/modelslist.go ./gallery/index.yaml > docs/static/gallery.html
- name: Build site
# Two Hugo sites, one Pages artifact: the main site owns the root,
# the docs site is nested under /docs/.
- name: Build the main site
working-directory: website
run: hugo --minify --baseURL "${{ steps.pages.outputs.base_url }}/"
- name: Build documentation site
working-directory: docs
run: |
mkdir -p layouts/_default
hugo --minify --baseURL "${{ steps.pages.outputs.base_url }}/"
hugo --minify --baseURL "${{ steps.pages.outputs.base_url }}/docs/"
- name: Merge documentation into the main site
run: |
mkdir -p website/public/docs
cp -R docs/public/. website/public/docs/
# Keeps the pre-split URLs alive; see the script header.
- name: Generate legacy URL redirects
run: .github/ci/gen-redirects.sh website/public "${{ steps.pages.outputs.base_url }}/"
- name: Upload artifact
uses: actions/upload-pages-artifact@v5
with:
path: docs/public
path: website/public
deploy:
environment:
name: github-pages
url: ${{ steps.deployment.outputs.page_url }}
runs-on: ubuntu-latest
# Same routing as build: a hosted slot for a ~10s deploy is exactly the kind
# of job that should not block on a starved pool. deploy-pages authenticates
# with the job's OIDC token (id-token: write above), which self-hosted
# runners issue the same way hosted ones do.
runs-on: ${{ github.repository == 'mudler/LocalAI' && 'arc-runner-set' || 'ubuntu-latest' }}
needs: build
steps:
- name: Deploy to GitHub Pages

View File

@@ -3,7 +3,28 @@
on:
pull_request:
# None of these seven image builds can observe a diff confined to these
# paths. Gallery metadata is parsed at runtime and never copied into an
# image; docs and markdown never enter one at all. Gallery content is
# still checked by yaml-check.yml and by
# core/gallery/variants_lint_test.go under 'tests'.
#
# backend/{cpp,go,python}/**: this workflow builds the core image, whose
# only compiled output is `make build` -> `go build ./cmd/local-ai`. The
# per-backend trees are copied into the builder but nothing in them reaches
# the binary or the final stage. backend/backend.proto is deliberately not
# listed: it feeds protogen-go and so does change the binary, and it does
# not live under any of these prefixes, so it still triggers a full run.
# See .agents/ci-caching.md.
paths-ignore:
- 'gallery/**'
- 'docs/**'
- 'examples/**'
- '**/*.md'
- 'backend/cpp/**'
- 'backend/go/**'
- 'backend/python/**'
concurrency:
group: ci-${{ github.event.pull_request.number || github.sha }}-${{ github.repository }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
@@ -52,7 +73,7 @@
tag-latest: 'false'
tag-suffix: '-gpu-nvidia-cuda-13'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:22.04"
base-image: "ubuntu:24.04"
makeflags: "--jobs=3 --output-sync=target"
ubuntu-version: '2404'
- build-type: 'hipblas'

View File

@@ -13,8 +13,53 @@
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
jobs:
hipblas-jobs:
# Decide once whether this push can change any image. Gallery metadata is
# fetched at runtime and never baked into an image, and docs/markdown never
# enter one, so a push confined to those paths produces byte-identical
# images. On 2026-07-30, 12 of the 23 queued runs of this workflow were
# commits like "add 1 new model to gallery" or a docs fix, each rebuilding
# all 18 images.
#
# A job-level gate rather than `paths-ignore` on the trigger: paths-ignore
# would also apply to tag pushes, and a tag created on an existing commit
# carries an empty commits list, which would silently skip the release image
# build. Tags short-circuit to "build" below, as does a push whose base
# commit cannot be resolved -- the same run-everything posture the backend
# matrix filter takes for a truncated diff.
changes:
if: github.repository == 'mudler/LocalAI'
runs-on: ubuntu-latest
outputs:
build: ${{ steps.decide.outputs.build }}
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0
- id: decide
env:
BEFORE: ${{ github.event.before }}
AFTER: ${{ github.sha }}
run: |
set -euo pipefail
emit() { echo "$2"; echo "build=$1" >> "$GITHUB_OUTPUT"; exit 0; }
case "${GITHUB_REF}" in
refs/tags/*) emit true "tag push: building every image" ;;
esac
if [ -z "${BEFORE:-}" ] || [ "${BEFORE}" = "0000000000000000000000000000000000000000" ] \
|| ! git cat-file -e "${BEFORE}^{commit}" 2>/dev/null; then
emit true "no resolvable base commit: building every image"
fi
files="$(git diff --name-only "${BEFORE}" "${AFTER}")"
echo "changed files:"; echo "${files:-<none>}"
[ -z "${files}" ] && emit true "empty diff: building every image"
if echo "${files}" | grep -qvE '^(gallery/|docs/|examples/)|\.md$'; then
emit true "push touches image-visible content: building"
fi
emit false "only gallery/docs/markdown changed: images identical, skipping"
hipblas-jobs:
needs: changes
if: github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true'
uses: ./.github/workflows/image_build.yml
with:
tag-latest: ${{ matrix.tag-latest }}
@@ -47,7 +92,8 @@
ubuntu-codename: 'noble'
core-image-build:
if: github.repository == 'mudler/LocalAI'
needs: changes
if: github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true'
uses: ./.github/workflows/image_build.yml
with:
tag-latest: ${{ matrix.tag-latest }}
@@ -113,7 +159,7 @@
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:22.04"
base-image: "ubuntu:24.04"
skip-drivers: 'false'
makeflags: "--jobs=4 --output-sync=target"
ubuntu-version: '2404'
@@ -155,8 +201,8 @@
# merge whenever any matrix cell of the parent build fails or is
# cancelled. Same fix as backend.yml's merge jobs — we still want to
# publish the manifest list for tag-suffixes whose legs all succeeded.
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' }}
needs: core-image-build
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true' }}
needs: [changes, core-image-build]
uses: ./.github/workflows/image_merge.yml
with:
tag-latest: 'auto'
@@ -168,8 +214,8 @@
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
gpu-vulkan-image-merge:
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' }}
needs: core-image-build
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true' }}
needs: [changes, core-image-build]
uses: ./.github/workflows/image_merge.yml
with:
tag-latest: 'auto'
@@ -187,8 +233,8 @@
# Each merge job needs only its parent build matrix and is filtered by
# tag-suffix in image_merge.yml's artifact-download pattern.
gpu-nvidia-cuda-12-image-merge:
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' }}
needs: core-image-build
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true' }}
needs: [changes, core-image-build]
uses: ./.github/workflows/image_merge.yml
with:
tag-latest: 'auto'
@@ -200,8 +246,8 @@
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
gpu-nvidia-cuda-13-image-merge:
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' }}
needs: core-image-build
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true' }}
needs: [changes, core-image-build]
uses: ./.github/workflows/image_merge.yml
with:
tag-latest: 'auto'
@@ -213,8 +259,8 @@
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
gpu-intel-image-merge:
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' }}
needs: core-image-build
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true' }}
needs: [changes, core-image-build]
uses: ./.github/workflows/image_merge.yml
with:
tag-latest: 'auto'
@@ -226,8 +272,8 @@
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
gpu-hipblas-image-merge:
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' }}
needs: hipblas-jobs
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true' }}
needs: [changes, hipblas-jobs]
uses: ./.github/workflows/image_merge.yml
with:
tag-latest: 'auto'
@@ -239,8 +285,8 @@
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
nvidia-l4t-arm64-image-merge:
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' }}
needs: gh-runner
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true' }}
needs: [changes, gh-runner]
uses: ./.github/workflows/image_merge.yml
with:
tag-latest: 'auto'
@@ -252,8 +298,8 @@
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
nvidia-l4t-arm64-cuda-13-image-merge:
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' }}
needs: gh-runner
if: ${{ !cancelled() && github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true' }}
needs: [changes, gh-runner]
uses: ./.github/workflows/image_merge.yml
with:
tag-latest: 'auto'
@@ -265,7 +311,8 @@
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
gh-runner:
if: github.repository == 'mudler/LocalAI'
needs: changes
if: github.repository == 'mudler/LocalAI' && needs.changes.outputs.build == 'true'
uses: ./.github/workflows/image_build.yml
with:
tag-latest: ${{ matrix.tag-latest }}

View File

@@ -8,6 +8,9 @@ on:
- 'examples/**'
- 'README.md'
- '**/*.md'
# golangci-lint runs new-from-merge-base, so a diff with no touched Go
# lines can only ever be a no-op. See .agents/ci-caching.md.
- 'gallery/**'
push:
branches:
- master
@@ -18,8 +21,41 @@ concurrency:
jobs:
golangci-lint:
# Self-hosted for PUSH only, and only in the canonical repo.
#
# This workflow also runs on pull_request, which for a fork PR means
# executing untrusted contributor code. That must never land on a
# self-hosted runner, so anything that is not a push to mudler/LocalAI stays
# on the ephemeral hosted pool. Pushes to master are trusted code that has
# already been reviewed and merged.
#
# Why at all: the hosted pool is shared account-wide and starved for 35
# straight minutes on 2026-07-31 while arc-runner-set kept completing jobs.
# Lint is small and runs on every commit, so it is a good candidate to move
# off the contended pool.
# REVERTED to hosted: the arc-runner-set image has git, curl, unzip, tar,
# ldd and python3, but NOT make (nor gcc). Measured on run 30637392862,
# where the preflight below named both. Re-route here once the runner image
# ships a C toolchain and make; the preflight stays so the next attempt
# fails by name in one second instead of opaquely mid-build.
runs-on: ubuntu-latest
steps:
- name: Preflight - required host tools
# The hosted images ship these; a self-hosted container image may not.
# Check up front so a missing tool reports itself by name instead of
# surfacing as an opaque failure inside `make protogen-go` (which needs
# curl + unzip for protoc) or `make lint`.
run: |
missing=""
for t in git curl unzip make tar; do
command -v "$t" >/dev/null 2>&1 || missing="$missing $t"
done
echo "runner: ${RUNNER_NAME:-unknown} os: $(uname -sm)"
if [ -n "$missing" ]; then
echo "::error::missing required tools on this runner:$missing"
exit 1
fi
echo "all required tools present"
- uses: actions/checkout@v7
with:
# Full history so golangci-lint's new-from-merge-base can reach
@@ -52,8 +88,30 @@ jobs:
# container build (a missing transitive dep, a partial cuDNN family). Their
# shell tests need nothing but bash + gcc + ldd, so run them on every PR
# rather than waiting on a multi-GB cross-arch backend image build.
#
# Push-only self-hosted routing, same fork-safety reasoning as
# golangci-lint above.
# REVERTED to hosted: the arc-runner-set image has git, curl, unzip, tar,
# ldd and python3, but NOT make (nor gcc). Measured on run 30637392862,
# where the preflight below named both. Re-route here once the runner image
# ships a C toolchain and make; the preflight stays so the next attempt
# fails by name in one second instead of opaquely mid-build.
runs-on: ubuntu-latest
steps:
- name: Preflight - required host tools
# This job additionally needs a C toolchain: the packaging-script tests
# compile a throwaway binary and inspect it with ldd.
run: |
missing=""
for t in git make gcc ldd python3; do
command -v "$t" >/dev/null 2>&1 || missing="$missing $t"
done
echo "runner: ${RUNNER_NAME:-unknown} os: $(uname -sm)"
if [ -n "$missing" ]; then
echo "::error::missing required tools on this runner:$missing"
exit 1
fi
echo "all required tools present"
- uses: actions/checkout@v7
- name: run packaging script tests
run: make test-build-scripts
@@ -61,8 +119,14 @@ jobs:
# The backend matrix path filter fails silently: a miss emits an empty
# matrix, every job goes green, and the change reaches no image (#10946).
# Its tests need only node, so they ride along with this job.
- uses: actions/setup-node@v4
- uses: actions/setup-node@v7
with:
node-version: '20'
- name: run CI script tests
run: make test-ci-scripts
# The shared python backend helpers (Options[] parsing, engine-arg
# mapping, model reference resolution) are stdlib-only, so their tests
# ride along here instead of waiting on a multi-GB backend image build.
- name: run shared python backend helper tests
run: make test-python-helpers

View File

@@ -0,0 +1,44 @@
name: Refresh site counters
# The landing page shows a star count, a contributor count and a release
# count. They were typed in by hand, so they drifted the moment somebody
# forgot. This pulls the real numbers once a week and commits them only when
# they have actually moved, which in turn triggers the usual Pages deploy.
on:
schedule:
# Mondays, 06:17 UTC. Off the hour on purpose, since the scheduler queues
# everything that asks for :00 and drops what it cannot run.
- cron: '17 6 * * 1'
workflow_dispatch:
permissions:
contents: write
concurrency:
group: refresh-site-counters
cancel-in-progress: false
jobs:
refresh:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Read the counts off the GitHub API
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: ./.github/ci/refresh-site-counters.sh
- name: Commit only if something moved
run: |
if git diff --quiet -- website/data/stats.yaml; then
echo "counters unchanged, nothing to commit"
exit 0
fi
git diff --unified=0 -- website/data/stats.yaml
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git add website/data/stats.yaml
git commit -m "chore(website): refresh the counters"
git push

View File

@@ -28,9 +28,9 @@ jobs:
steps:
- name: Checkout Source
uses: actions/checkout@v7
if: ${{ github.actor != 'dependabot[bot]' }}
if: ${{ !github.repository.fork && github.actor != 'dependabot[bot]' }}
- name: Run Gosec Security Scanner
if: ${{ github.actor != 'dependabot[bot]' }}
if: ${{ !github.repository.fork && github.actor != 'dependabot[bot]' }}
uses: securego/gosec@v2.27.1
with:
# we let the report trigger content trigger a failure using the GitHub Security features.
@@ -39,7 +39,7 @@ jobs:
# noise, G104 unhandled errors) are inherent to that upstream code, not ours to rewrite.
args: '-no-fail -exclude-dir=backend/go/supertonic -fmt sarif -out results.sarif ./...'
- name: Upload SARIF file
if: ${{ github.actor != 'dependabot[bot]' }}
if: ${{ !github.repository.fork && github.actor != 'dependabot[bot]' }}
uses: github/codeql-action/upload-sarif@v4
with:
# Path to SARIF file relative to the root of the repository

View File

@@ -37,6 +37,8 @@ jobs:
sglang: ${{ steps.detect.outputs.sglang }}
acestep-cpp: ${{ steps.detect.outputs.acestep-cpp }}
qwen3-tts-cpp: ${{ steps.detect.outputs.qwen3-tts-cpp }}
magpie-tts-cpp: ${{ steps.detect.outputs.magpie-tts-cpp }}
trellis2cpp: ${{ steps.detect.outputs.trellis2cpp }}
rfdetr-cpp: ${{ steps.detect.outputs.rfdetr-cpp }}
locate-anything-cpp: ${{ steps.detect.outputs.locate-anything-cpp }}
vibevoice-cpp: ${{ steps.detect.outputs.vibevoice-cpp }}
@@ -866,6 +868,38 @@ jobs:
- name: Test qwen3-tts-cpp
run: |
make --jobs=5 --output-sync=target -C backend/go/qwen3-tts-cpp test
tests-magpie-tts-cpp:
needs: detect-changes
if: needs.detect-changes.outputs.magpie-tts-cpp == 'true' || needs.detect-changes.outputs.run-all == 'true'
runs-on: ubuntu-latest
steps:
- name: Clone
uses: actions/checkout@v7
with:
submodules: true
- name: Dependencies
run: |
sudo apt-get update
sudo apt-get install -y build-essential cmake curl libopenblas-dev ffmpeg
- name: Setup Go
uses: actions/setup-go@v5
- name: Display Go version
run: go version
- name: Proto Dependencies
run: |
# Install protoc
curl -L -s https://github.com/protocolbuffers/protobuf/releases/download/v26.1/protoc-26.1-linux-x86_64.zip -o protoc.zip && \
unzip -j -d /usr/local/bin protoc.zip bin/protoc && \
rm protoc.zip
go install google.golang.org/protobuf/cmd/protoc-gen-go@v1.34.2
go install google.golang.org/grpc/cmd/protoc-gen-go-grpc@1958fcbe2ca8bd93af633f11e97d44e567e945af
PATH="$PATH:$HOME/go/bin" make protogen-go
- name: Build magpie-tts-cpp
run: |
make --jobs=5 --output-sync=target -C backend/go/magpie-tts-cpp
- name: Test magpie-tts-cpp
run: |
make --jobs=5 --output-sync=target -C backend/go/magpie-tts-cpp test
# Per-backend smoke for rfdetr-cpp: builds the .so + Go binary and runs
# `make -C backend/go/rfdetr-cpp test`. test.sh fetches the small (~20 MB)
# rfdetr-nano-q8_0 GGUF from the published mudler/rfdetr-cpp-nano HF repo
@@ -902,6 +936,41 @@ jobs:
- name: Test rfdetr-cpp
run: |
make --jobs=5 --output-sync=target -C backend/go/rfdetr-cpp test
# Weight-free packaged-backend smoke for trellis2cpp. Starting run.sh loads
# libtrellis2 + ggml, resolves the complete C ABI (including remeshing), and
# answers gRPC Health without downloading or loading the multi-GB model set.
tests-trellis2cpp:
needs: detect-changes
if: needs.detect-changes.outputs.trellis2cpp == 'true' || needs.detect-changes.outputs.run-all == 'true'
runs-on: ubuntu-latest
timeout-minutes: 90
steps:
- name: Clone
uses: actions/checkout@v7
with:
submodules: true
- name: Dependencies
run: |
sudo apt-get update
sudo apt-get install -y build-essential cmake curl unzip
- name: Setup Go
uses: actions/setup-go@v5
- name: Display Go version
run: go version
- name: Proto Dependencies
run: |
curl -L -s https://github.com/protocolbuffers/protobuf/releases/download/v26.1/protoc-26.1-linux-x86_64.zip -o protoc.zip && \
unzip -j -d /usr/local/bin protoc.zip bin/protoc && \
rm protoc.zip
go install google.golang.org/protobuf/cmd/protoc-gen-go@v1.34.2
go install google.golang.org/grpc/cmd/protoc-gen-go-grpc@1958fcbe2ca8bd93af633f11e97d44e567e945af
PATH="$PATH:$HOME/go/bin" make protogen-go
- name: Build trellis2cpp
run: |
make --jobs=5 --output-sync=target -C backend/go/trellis2cpp
- name: Test trellis2cpp
run: |
make --jobs=5 --output-sync=target -C backend/go/trellis2cpp test
# Per-backend e2e for locate-anything-cpp: builds the .so + Go binary and
# runs `make -C backend/go/locate-anything-cpp test`. test.sh fetches the
# locate-anything-q8_0 GGUF (~6.3 GB, NVIDIA LocateAnything-3B) from the

View File

@@ -24,8 +24,13 @@ jobs:
uses: actions/checkout@v7
with:
submodules: true
- name: Free disk space
uses: ./.github/actions/free-disk-space
# No free-disk-space step here on purpose. That action exists to make room
# for docker buildx layers, and this job runs no buildx step. It was also
# sized for a `make test` that downloaded multi-GB GGUF/whisper fixtures
# and built llama-cpp/whisper/stablediffusion-ggml; after the test-suite
# reorg it does neither (see the Makefile test target). It cost ~3min of
# every run, and its tool-cache:true wipe also forced setup-go and
# setup-node to re-download toolchains that ship preinstalled.
- name: Setup Go ${{ matrix.go-version }}
uses: actions/setup-go@v5
with:

View File

@@ -3,6 +3,14 @@ name: 'E2E Backend Tests'
on:
pull_request:
# The e2e suite drives backends over gRPC directly and reads none of these
# paths, so a diff confined to them cannot move it.
# See .agents/ci-caching.md.
paths-ignore:
- 'gallery/**'
- 'docs/**'
- 'examples/**'
- '**/*.md'
push:
branches:
- master

11
.gitignore vendored
View File

@@ -30,6 +30,7 @@ LocalAI
# Go backend packages whose main lives under backend/go/.
/cloud-proxy
/local-store
/valkey-store
# prevent above rules from omitting the helm chart
!charts/*
# prevent above rules from omitting the api/localai folder
@@ -61,6 +62,11 @@ prepare
/ggml-metal.metal
docs/static/gallery.html
# Hugo build output and lock files (docs/ and website/)
docs/public/
website/public/
.hugo_build.lock
# Protobuf generated files
*.pb.go
*pb2.py
@@ -118,3 +124,8 @@ formal-verification/out/
# package directory itself and untrack the source.
/apexentries
/.github/ci/apexentries/apexentries
# Runtime state written by `local-ai run` when it is started from the repo
# root, which is what a contributor testing a build does. Nothing under here is
# source: it is the instance's own models, outputs, traces and identity.
/data/

48
ADOPTERS.md Normal file
View File

@@ -0,0 +1,48 @@
# Adopters
Organisations running LocalAI, listed by the people who run it.
If your organisation uses LocalAI and you are happy to say so publicly, open a
pull request adding a row to the table below. That pull request is how we know
we have permission to list you, which is why we do not add anybody ourselves.
You do not need to be a large company, and you do not need to disclose anything
sensitive. A sentence on what you use it for is more useful to other readers
than a logo.
## How to add yourself
1. Add a row to the table, in alphabetical order.
2. Use your organisation's usual name and a link to your site.
3. Say briefly what you use LocalAI for, and whether it is in production.
4. Open the pull request from an account that makes it plausible you speak for
the organisation, or say in the description who you are. We may ask.
To be removed, open a pull request deleting your row, or email
[info@localai.io](mailto:info@localai.io). We will not ask why.
## Who is using LocalAI
<!-- Keep alphabetical. Columns: Organisation | What for | Status -->
| Organisation | What they use it for | Status |
|---|---|---|
| _Your organisation here_ | | |
## What this list is not
This is not a list of everyone who has ever starred the repository, and it is
not a list of the employers of people who have contributed a patch. Both of
those are easy to scrape and neither means what a logo wall implies.
The website shows two separate things, both of which are checkable without
anybody's permission:
- **Engineers from these companies have contributed code.** Evidence is the
commit history plus the employer on that person's public GitHub profile. It
is a claim about a person, not about their employer.
- **These projects integrate LocalAI.** Evidence is a reference to LocalAI in
that project's own repository or documentation.
Those two lists live in [`website/data/ecosystem.yaml`](website/data/ecosystem.yaml).
This file is the third, stronger thing: organisations that chose to say so.

View File

@@ -32,16 +32,18 @@ LocalAI follows the Linux kernel project's [guidelines for AI coding assistants]
| [.agents/adding-gallery-models.md](.agents/adding-gallery-models.md) | Adding GGUF models from HuggingFace to the model gallery |
| [.agents/localai-assistant-mcp.md](.agents/localai-assistant-mcp.md) | LocalAI Assistant chat modality — adding admin tools to the in-process MCP server, editing skill prompts, keeping REST + MCP + skills in sync |
| [.agents/backend-signing.md](.agents/backend-signing.md) | Backend OCI image signing (keyless cosign + sigstore-go) — producer-side CI setup, consumer-side gallery `verification:` block, strict mode (`LOCALAI_REQUIRE_BACKEND_INTEGRITY`), revocation via `not_before` |
| [.agents/preparing-a-release.md](.agents/preparing-a-release.md) | Cutting a release: PR labels, `RELEASE_NOTES_vX.Y.Z.md`, the blog post under `website/content/blog/`, and the demo clips under `website/static/media/` |
## Quick Reference
- **Git hooks & coverage gates**: Run `make install-hooks` once per clone so the pre-commit lint + coverage gates run. **Never bypass them with `git commit --no-verify`, and never lower a coverage baseline or widen a gate's tolerance to turn a red gate green** — the coverage ratchet only moves up. If a change drops coverage, add tests to raise it (e.g. render-smoke specs). See [.agents/building-and-testing.md](.agents/building-and-testing.md).
- **Coverage gates**: Never lower a coverage baseline or widen a gate's tolerance to turn a red gate green — the coverage ratchet only moves up. If a change drops coverage, add tests to raise it (e.g. render-smoke specs). See [.agents/building-and-testing.md](.agents/building-and-testing.md).
- **Logging**: Use `github.com/mudler/xlog` (same API as slog)
- **Go style**: Prefer `any` over `interface{}`
- **Comments**: Explain *why*, not *what*
- **Docs (docs-with-code rule)**: When you change user-facing behavior (API endpoints, CLI flags, config keys, or features), update the corresponding page under `docs/content/` in the SAME change, not as a follow-up. A user-facing change without a matching docs update is incomplete. See also the documentation conventions in [.agents/coding-style.md](.agents/coding-style.md).
- **New API endpoints**: LocalAI advertises its capability surface in several independent places — swagger `@Tags`, `/api/instructions` registry, auth `RouteFeatureRegistry`, React UI `capabilities.js`, docs. Read [.agents/api-endpoints-and-auth.md](.agents/api-endpoints-and-auth.md) and follow its checklist — missing any surface means clients, admins, and the UI won't know the endpoint exists.
- **Admin endpoints → MCP tool**: every admin endpoint that an admin would manage conversationally (install/list/edit/toggle/upgrade) MUST also be exposed as an MCP tool in `pkg/mcp/localaitools/`. The LocalAI Assistant chat modality and the standalone `local-ai mcp-server` consume that package; drift between REST and MCP is a real risk. Read [.agents/localai-assistant-mcp.md](.agents/localai-assistant-mcp.md) — the `TestToolHTTPRouteMappingComplete` test fails until you wire the new tool and update the route map.
- **Releases ship with a post and clips**: a release is not done at the tag. It needs labelled PRs, `RELEASE_NOTES_vX.Y.Z.md`, a blog post under `website/content/blog/`, and a short demo clip in `website/static/media/` for each notable feature. See [.agents/preparing-a-release.md](.agents/preparing-a-release.md).
- **Build**: Inspect `Makefile` and `.github/workflows/` — ask the user before running long builds
- **Backend OS coverage**: a new backend must target every OS it can build for, not just Linux. `.github/backend-matrix.yml` has two matrices — `include:` (Linux) and `includeDarwin:` (macOS / Apple Silicon). Most C/C++/GGML and many Python backends build on Darwin too — wire the `includeDarwin` entry + `backend/index.yaml` `metal:` entries, or say in the PR why an OS is unsupported. See the darwin checklist in [.agents/adding-backends.md](.agents/adding-backends.md).
- **Gallery variant ranking**: a gallery entry can declare `variants` (alternative builds of the same weights), and LocalAI ranks the ones a host can run by engine preference first, size second. A new backend that should be preferred on some hardware must be listed in `engineNamePreferenceRules` in `pkg/system/capabilities.go`; the sibling `backendBuildTagPreferenceRules` speaks build tags rather than engine names, and using the wrong table matches nothing without erroring. See [.agents/adding-backends.md](.agents/adding-backends.md).

View File

@@ -198,7 +198,6 @@ For AI-assisted development, see [`AGENTS.md`](AGENTS.md) (or the equivalent [`C
- Prefer modern Go idioms — for example, use `any` instead of `interface{}`.
- Use [`golangci-lint`](https://golangci-lint.run) to catch common issues before submitting a PR.
- Run `make install-hooks` once per clone to enable the pre-commit hook: Go changes run `make lint` + the coverage gate (`make test-coverage-check`); `core/http/react-ui/` changes run the Playwright e2e suite (`make test-ui`). Bypass a single commit with `git commit --no-verify`.
- Use [`github.com/mudler/xlog`](https://github.com/mudler/xlog) for logging (same API as `slog`). Do not use `fmt.Println` or the standard `log` package for operational logging.
- Use tab indentation for Go files (as defined in `.editorconfig`).
@@ -268,7 +267,7 @@ make test-e2e
### React UI tests and coverage
The React UI (`core/http/react-ui/`) is covered by Playwright e2e specs, gated by a **monotonic line-coverage ratchet** (`make test-ui-coverage-check`, run in CI and pre-commit). The metric is non-deterministic — a fast local box reads higher than a slow CI runner for the same code — so a small tolerance is unavoidable.
The React UI (`core/http/react-ui/`) is covered by Playwright e2e specs, gated by a **monotonic line-coverage ratchet** (`make test-ui-coverage-check`, run in CI). The metric is non-deterministic — a fast local box reads higher than a slow CI runner for the same code — so a small tolerance is unavoidable.
**If your change lowers UI coverage, raise it back by adding specs — do not widen the tolerance or hand-lower the baseline.** A *render-smoke* spec (navigate to a page, assert its header is visible) cheaply covers an entire lazy page. See `core/http/react-ui/e2e/page-render-smoke.spec.js` and the full policy in [.agents/building-and-testing.md](.agents/building-and-testing.md#react-ui-coverage).

102
Makefile
View File

@@ -1,5 +1,5 @@
# Disable parallel execution for backend builds
.NOTPARALLEL: backends/diffusers backends/llama-cpp backends/turboquant backends/bonsai backends/outetts backends/piper backends/stablediffusion-ggml backends/whisper backends/crispasr backends/parakeet-cpp backends/moss-transcribe-cpp backends/faster-whisper backends/silero-vad backends/local-store backends/cloud-proxy backends/huggingface backends/rfdetr backends/rfdetr-cpp backends/insightface backends/speaker-recognition backends/kitten-tts backends/kokoro backends/chatterbox backends/llama-cpp-darwin backends/neutts build-darwin-python-backend build-darwin-go-backend backends/mlx backends/diffuser-darwin backends/mlx-vlm backends/mlx-audio backends/mlx-distributed backends/stablediffusion-ggml-darwin backends/vllm backends/vllm-omni backends/longcat-video backends/sglang backends/moonshine backends/pocket-tts backends/qwen-tts backends/faster-qwen3-tts backends/qwen-asr backends/nemo backends/voxcpm backends/whisperx backends/ace-step backends/acestep-cpp backends/fish-speech backends/voxtral backends/opus backends/trl backends/llama-cpp-quantization backends/kokoros backends/sam3-cpp backends/qwen3-tts-cpp backends/moss-tts-cpp backends/omnivoice-cpp backends/vibevoice-cpp backends/localvqe backends/tinygrad backends/sherpa-onnx backends/ds4 backends/ds4-darwin backends/liquid-audio backends/supertonic backends/depth-anything-cpp backends/privacy-filter backends/privacy-filter-darwin
.NOTPARALLEL: backends/diffusers backends/llama-cpp backends/turboquant backends/bonsai backends/outetts backends/piper backends/stablediffusion-ggml backends/trellis2cpp backends/trellis2cpp-darwin backends/whisper backends/crispasr backends/parakeet-cpp backends/moss-transcribe-cpp backends/faster-whisper backends/silero-vad backends/local-store backends/valkey-store backends/cloud-proxy backends/huggingface backends/rfdetr backends/rfdetr-cpp backends/insightface backends/speaker-recognition backends/kitten-tts backends/kokoro backends/chatterbox backends/llama-cpp-darwin backends/neutts build-darwin-python-backend build-darwin-go-backend backends/mlx backends/diffuser-darwin backends/mlx-vlm backends/mlx-audio backends/mlx-distributed backends/stablediffusion-ggml-darwin backends/vllm backends/vllm-omni backends/longcat-video backends/sglang backends/moonshine backends/pocket-tts backends/qwen-tts backends/faster-qwen3-tts backends/qwen-asr backends/nemo backends/voxcpm backends/whisperx backends/ace-step backends/acestep-cpp backends/fish-speech backends/voxtral backends/opus backends/trl backends/llama-cpp-quantization backends/kokoros backends/sam3-cpp backends/qwen3-tts-cpp backends/moss-tts-cpp backends/magpie-tts-cpp backends/vllm-cpp backends/omnivoice-cpp backends/vibevoice-cpp backends/localvqe backends/tinygrad backends/sherpa-onnx backends/ds4 backends/ds4-darwin backends/liquid-audio backends/supertonic backends/depth-anything-cpp backends/privacy-filter backends/privacy-filter-darwin backends/audio-cpp backends/audio-cpp-darwin
GOCMD=go
GOTEST=$(GOCMD) test
@@ -69,7 +69,7 @@ else
GORELEASER=$(shell which goreleaser)
endif
TEST_PATHS?=./api/... ./pkg/... ./core/... ./backend/go/cloud-proxy/... ./backend/go/local-store/...
TEST_PATHS?=./api/... ./pkg/... ./core/... ./backend/go/cloud-proxy/... ./backend/go/local-store/... ./backend/go/valkey-store/...
## Coverage output and the committed baseline that CI compares against.
## The gate is strict: total coverage must never decrease (no tolerance).
@@ -103,7 +103,7 @@ COVERAGE_E2E_LABELS?=!real-models
COVERAGE_EXCLUDE_RE?=grpc/proto/.*[.]pb[.]go
.PHONY: all test test-coverage test-coverage-baseline test-coverage-check test-backend-cpp test-build-scripts test-ui test-ui-coverage-baseline test-ui-coverage-check install-hooks build vendor lint lint-all
.PHONY: all test test-coverage test-coverage-baseline test-coverage-check test-backend-cpp test-build-scripts test-ui test-ui-coverage-baseline test-ui-coverage-check build vendor lint lint-all
all: help
@@ -172,6 +172,15 @@ build-dev: ## Run LocalAI in dev mode with live reload
dev-dist:
$(GORELEASER) build --snapshot --clean
## PR-time variant of dev-dist: builds only the host platform instead of all
## three release targets (linux/amd64, linux/arm64, darwin/arm64). The point of
## running goreleaser on a PR is to catch a broken config or a broken
## before-hook (protogen-go, react-ui, go mod tidy), and --single-target still
## exercises every one of those. Nothing consumes a PR's cross-compiled
## binaries. master pushes and tags still run the full dev-dist/dist.
dev-dist-single:
$(GORELEASER) build --snapshot --clean --single-target
dist:
$(GORELEASER) build --clean
@@ -222,6 +231,14 @@ test-build-scripts:
test-ci-scripts:
@set -e; for t in scripts/lib/*_test.mjs; do echo "== $$t"; node --test "$$t"; done
## Runs the unit tests for the shared python backend helpers. These modules are
## pure stdlib on purpose so they run without any backend venv; the list is
## explicit because their siblings (model_identity_test) import grpc and the
## generated protobufs, which only exist inside a built backend.
PYTHON_HELPER_TESTS?=python_utils_test vllm_utils_test model_utils_test mlx_utils_test parent_watch_test
test-python-helpers:
cd backend/python/common && python3 -m unittest $(PYTHON_HELPER_TESTS)
## Runs the core suite ($(TEST_PATHS)) with statement-coverage instrumentation
## and writes a merged profile to $(COVERAGE_PROFILE). Deliberately omits
## --fail-fast so a single failure doesn't truncate the coverage number, and
@@ -269,8 +286,7 @@ LINT_EXCLUDE_DIRS_RE=/(backend/go/(piper|silero-vad|llm)|cmd/launcher)(/|$$)
## Set LINT_NEW_FROM to a git ref to override .golangci.yml's
## new-from-merge-base (origin/master). Useful from a fork clone where
## origin/master is stale relative to the canonical repo — the pre-commit
## hook passes the resolved upstream ref here so local lint matches CI.
## origin/master is stale relative to the canonical repo.
LINT_NEW_FROM?=
lint:
@command -v golangci-lint >/dev/null 2>&1 || { \
@@ -289,17 +305,6 @@ lint-all:
}
golangci-lint run --new=false --new-from-merge-base= --new-from-rev= $$(go list -e -f '{{.Dir}}' ./... | grep -vE '$(LINT_EXCLUDE_DIRS_RE)')
########################################################
## Git hooks
########################################################
## Points git at the versioned .githooks/ directory so the pre-commit hook
## (lint + coverage gate) runs locally. Run once per clone. Undo with:
## `git config --unset core.hooksPath`. Skip a single commit with
## `git commit --no-verify`.
install-hooks:
git config core.hooksPath .githooks
@echo 'Installed git hooks: core.hooksPath -> .githooks (pre-commit runs lint + test-coverage-check on Go changes)'
########################################################
## E2E AIO tests (uses standard image with pre-configured models)
########################################################
@@ -398,6 +403,15 @@ test-stores: backends/local-store
BACKENDS_PATH=$(abspath ./)/backends \
$(GOCMD) run github.com/onsi/ginkgo/v2/ginkgo --flake-attempts $(TEST_FLAKES) -v -r tests/integration
## Valkey-backed vector-store integration. Requires a running Valkey Search
## server (valkey/valkey-bundle:9.1.0) reachable at $$VALKEY_ADDR — the suite
## skips itself when VALKEY_ADDR is unset. Builds the backend on demand and
## points the model loader at it via BACKENDS_PATH. Label-filtered to the
## valkey specs so it does not also run the in-memory local-store suite.
test-valkey-store: backends/valkey-store
BACKENDS_PATH=$(abspath ./)/backends \
$(GOCMD) run github.com/onsi/ginkgo/v2/ginkgo --flake-attempts $(TEST_FLAKES) --label-filter='valkey' -v -r tests/integration
test-opus:
@echo 'Running opus backend tests'
$(MAKE) -C backend/go/opus libopusshim.so
@@ -606,6 +620,8 @@ prepare-test-extra: protogen-python
$(MAKE) -C backend/rust/kokoros kokoros-grpc
$(MAKE) -C backend/go/rfdetr-cpp
$(MAKE) -C backend/go/locate-anything-cpp
$(MAKE) -C backend/go/trellis2cpp
$(MAKE) -C backend/go/valkey-store
test-extra: prepare-test-extra
$(MAKE) -C backend/python/transformers test
@@ -637,6 +653,9 @@ test-extra: prepare-test-extra
$(MAKE) -C backend/go/locate-anything-cpp test
$(MAKE) -C backend/go/depth-anything-cpp test
$(MAKE) -C backend/go/supertonic test
$(MAKE) -C backend/go/vllm-cpp test
$(MAKE) -C backend/go/trellis2cpp test
$(MAKE) -C backend/go/valkey-store test
##
## End-to-end gRPC tests that exercise a built backend container image.
@@ -1199,6 +1218,10 @@ backends/privacy-filter-darwin: build
bash ./scripts/build/privacy-filter-darwin.sh
./local-ai backends install "ocifile://$(abspath ./backend-images/privacy-filter.tar)"
backends/audio-cpp-darwin: build
bash ./scripts/build/audio-cpp-darwin.sh
./local-ai backends install "ocifile://$(abspath ./backend-images/audio-cpp.tar)"
build-darwin-python-backend: build
bash ./scripts/build/python-darwin.sh
@@ -1229,6 +1252,10 @@ backends/stablediffusion-ggml-darwin:
BACKEND=stablediffusion-ggml BUILD_TYPE=metal $(MAKE) build-darwin-go-backend
./local-ai backends install "ocifile://$(abspath ./backend-images/stablediffusion-ggml.tar)"
backends/trellis2cpp-darwin:
BACKEND=trellis2cpp BUILD_TYPE=metal $(MAKE) build-darwin-go-backend
./local-ai backends install "ocifile://$(abspath ./backend-images/trellis2cpp.tar)"
backend-images:
mkdir -p backend-images
@@ -1252,14 +1279,21 @@ BACKEND_DS4 = ds4|ds4|.|false|false
# openai-privacy-filter PII/NER token classifier) — the TokenClassify RPC for
# the PII redactor tier, on stock ggml with no llama.cpp carry-patches.
BACKEND_PRIVACY_FILTER = privacy-filter|privacy-filter|.|false|false
# audio-cpp wraps 0xShug0/audio.cpp, a multi-family ggml audio inference engine
# (TTS, ASR, VAD, diarization, source separation, music generation). Builds
# against apt gRPC/protobuf rather than a prebuilt base-grpc image; the reason
# is on the audio-cpp block in .github/backend-matrix.yml.
BACKEND_AUDIO_CPP = audio-cpp|audio-cpp|.|false|false
# Golang backends
BACKEND_PIPER = piper|golang|.|false|true
BACKEND_LOCAL_STORE = local-store|golang|.|false|true
BACKEND_VALKEY_STORE = valkey-store|golang|.|false|true
BACKEND_CLOUD_PROXY = cloud-proxy|golang|.|false|true
BACKEND_HUGGINGFACE = huggingface|golang|.|false|true
BACKEND_SILERO_VAD = silero-vad|golang|.|false|true
BACKEND_STABLEDIFFUSION_GGML = stablediffusion-ggml|golang|.|--progress=plain|true
BACKEND_TRELLIS2CPP = trellis2cpp|golang|.|--progress=plain|true
BACKEND_WHISPER = whisper|golang|.|false|true
BACKEND_CRISPASR = crispasr|golang|.|false|true
BACKEND_PARAKEET_CPP = parakeet-cpp|golang|.|false|true
@@ -1269,6 +1303,8 @@ BACKEND_VOXTRAL = voxtral|golang|.|false|true
BACKEND_ACESTEP_CPP = acestep-cpp|golang|.|false|true
BACKEND_QWEN3_TTS_CPP = qwen3-tts-cpp|golang|.|false|true
BACKEND_MOSS_TTS_CPP = moss-tts-cpp|golang|.|false|true
BACKEND_MAGPIE_TTS_CPP = magpie-tts-cpp|golang|.|false|true
BACKEND_VLLM_CPP = vllm-cpp|golang|.|false|true
BACKEND_OMNIVOICE_CPP = omnivoice-cpp|golang|.|false|true
BACKEND_VIBEVOICE_CPP = vibevoice-cpp|golang|.|false|true
BACKEND_LOCALVQE = localvqe|golang|.|false|true
@@ -1351,12 +1387,15 @@ $(eval $(call generate-docker-build-target,$(BACKEND_TURBOQUANT)))
$(eval $(call generate-docker-build-target,$(BACKEND_BONSAI)))
$(eval $(call generate-docker-build-target,$(BACKEND_DS4)))
$(eval $(call generate-docker-build-target,$(BACKEND_PRIVACY_FILTER)))
$(eval $(call generate-docker-build-target,$(BACKEND_AUDIO_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_PIPER)))
$(eval $(call generate-docker-build-target,$(BACKEND_LOCAL_STORE)))
$(eval $(call generate-docker-build-target,$(BACKEND_VALKEY_STORE)))
$(eval $(call generate-docker-build-target,$(BACKEND_CLOUD_PROXY)))
$(eval $(call generate-docker-build-target,$(BACKEND_HUGGINGFACE)))
$(eval $(call generate-docker-build-target,$(BACKEND_SILERO_VAD)))
$(eval $(call generate-docker-build-target,$(BACKEND_STABLEDIFFUSION_GGML)))
$(eval $(call generate-docker-build-target,$(BACKEND_TRELLIS2CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_WHISPER)))
$(eval $(call generate-docker-build-target,$(BACKEND_CRISPASR)))
$(eval $(call generate-docker-build-target,$(BACKEND_PARAKEET_CPP)))
@@ -1396,6 +1435,8 @@ $(eval $(call generate-docker-build-target,$(BACKEND_ACE_STEP)))
$(eval $(call generate-docker-build-target,$(BACKEND_ACESTEP_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_QWEN3_TTS_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_MOSS_TTS_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_MAGPIE_TTS_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_VLLM_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_OMNIVOICE_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_VIBEVOICE_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_LOCALVQE)))
@@ -1415,7 +1456,7 @@ $(eval $(call generate-docker-build-target,$(BACKEND_SUPERTONIC)))
docker-save-%: backend-images
docker save local-ai-backend:$* -o backend-images/$*.tar
docker-build-backends: docker-build-llama-cpp docker-build-ik-llama-cpp docker-build-turboquant docker-build-bonsai docker-build-ds4 docker-build-rerankers docker-build-vllm docker-build-vllm-omni docker-build-longcat-video docker-build-sglang docker-build-transformers docker-build-outetts docker-build-diffusers docker-build-kokoro docker-build-faster-whisper docker-build-crispasr docker-build-coqui docker-build-chatterbox docker-build-vibevoice docker-build-liquid-audio docker-build-moonshine docker-build-pocket-tts docker-build-qwen-tts docker-build-fish-speech docker-build-faster-qwen3-tts docker-build-qwen-asr docker-build-nemo docker-build-voxcpm docker-build-whisperx docker-build-ace-step docker-build-acestep-cpp docker-build-voxtral docker-build-mlx-distributed docker-build-trl docker-build-llama-cpp-quantization docker-build-tinygrad docker-build-kokoros docker-build-sam3-cpp docker-build-rfdetr-cpp docker-build-qwen3-tts-cpp docker-build-moss-tts-cpp docker-build-omnivoice-cpp docker-build-vibevoice-cpp docker-build-localvqe docker-build-insightface docker-build-speaker-recognition docker-build-sherpa-onnx docker-build-cloud-proxy docker-build-supertonic docker-build-depth-anything-cpp docker-build-moss-transcribe-cpp docker-build-privacy-filter
docker-build-backends: docker-build-llama-cpp docker-build-ik-llama-cpp docker-build-turboquant docker-build-bonsai docker-build-ds4 docker-build-rerankers docker-build-vllm docker-build-vllm-omni docker-build-longcat-video docker-build-sglang docker-build-transformers docker-build-outetts docker-build-diffusers docker-build-kokoro docker-build-faster-whisper docker-build-crispasr docker-build-coqui docker-build-chatterbox docker-build-vibevoice docker-build-liquid-audio docker-build-moonshine docker-build-pocket-tts docker-build-qwen-tts docker-build-fish-speech docker-build-faster-qwen3-tts docker-build-qwen-asr docker-build-nemo docker-build-voxcpm docker-build-whisperx docker-build-ace-step docker-build-acestep-cpp docker-build-voxtral docker-build-mlx-distributed docker-build-trl docker-build-llama-cpp-quantization docker-build-tinygrad docker-build-kokoros docker-build-sam3-cpp docker-build-rfdetr-cpp docker-build-qwen3-tts-cpp docker-build-moss-tts-cpp docker-build-magpie-tts-cpp docker-build-vllm-cpp docker-build-omnivoice-cpp docker-build-vibevoice-cpp docker-build-localvqe docker-build-insightface docker-build-speaker-recognition docker-build-sherpa-onnx docker-build-cloud-proxy docker-build-supertonic docker-build-depth-anything-cpp docker-build-moss-transcribe-cpp docker-build-privacy-filter docker-build-trellis2cpp docker-build-valkey-store docker-build-audio-cpp
########################################################
### Mock Backend for E2E Tests
@@ -1450,7 +1491,7 @@ test-ui-e2e: build-ui-test-server
UI_TEST_WORKERS ?=
PLAYWRIGHT_WORKERS_FLAG = $(if $(UI_TEST_WORKERS),--workers=$(UI_TEST_WORKERS),)
## Fast Playwright e2e run used by the pre-commit hook on React UI changes.
## Fast Playwright e2e run for local React UI validation.
## Force-rebuilds the (non-instrumented) dist so the suite tests the working
## tree — not a stale dist the `react-ui` skip-guard would leave — re-embeds
## it into ui-test-server, and runs the specs. Uses the nix-provided browser
@@ -1507,7 +1548,12 @@ swagger:
gen-assets:
$(GOCMD) run core/dependencies_manager/manager.go webui_static.yaml core/http/static/assets
## Documentation
## Documentation and website
# The published site is two Hugo sites: website/ owns the root, docs/ is nested
# under /docs/. Serve them separately while editing; use `make site` to get the
# merged tree (including the legacy URL redirects) that GitHub Pages deploys.
SITE_BASE_URL?=http://localhost:8000
docs/layouts/_default:
mkdir -p docs/layouts/_default
@@ -1519,12 +1565,30 @@ docs/public: docs/layouts/_default docs/static/gallery.html
docs-clean:
rm -rf docs/public
rm -rf website/public
rm -rf docs/static/gallery.html
.PHONY: docs
docs: docs/static/gallery.html
cd docs && hugo serve
.PHONY: website
website:
cd website && hugo serve
.PHONY: site
site: docs/static/gallery.html
rm -rf website/public docs/public
cd website && hugo --minify --baseURL "$(SITE_BASE_URL)/"
cd docs && hugo --minify --baseURL "$(SITE_BASE_URL)/docs/"
mkdir -p website/public/docs
cp -R docs/public/. website/public/docs/
./.github/ci/gen-redirects.sh website/public "$(SITE_BASE_URL)/"
.PHONY: site-serve
site-serve: site
cd website/public && python3 -m http.server 8000
########################################################
## Platform-specific builds
########################################################

View File

@@ -161,7 +161,7 @@ local-ai run https://gist.githubusercontent.com/.../phi-2.yaml
local-ai run oci://localai/phi-2:latest
```
To test a running LocalAI server from the terminal, open an interactive chat session from another shell. Inside the prompt, `/models` lists installed models and `/model <name>` switches between them.
To work with a running LocalAI server from the terminal, start the built-in agent from another shell. It answers questions, reads your files and runs commands on your machine, asking you to approve anything that changes state. Inside a session, `/models` lists installed models and `/model <name>` switches between them. See the [Terminal agent](https://localai.io/docs/features/terminal-agent/) docs.
```bash
# Terminal 1
@@ -195,7 +195,7 @@ For more details, see the [Getting Started guide](https://localai.io/basics/gett
- **August 2025**: MLX, MLX-VLM, Diffusers, llama.cpp now supported on Apple Silicon
- **July 2025**: All backends migrated outside the main binary — [lightweight, modular architecture](https://github.com/mudler/LocalAI/releases/tag/v3.2.0)
For older news and full release notes, see [GitHub Releases](https://github.com/mudler/LocalAI/releases) and the [News page](https://localai.io/basics/news/).
For older news and full release notes, see [GitHub Releases](https://github.com/mudler/LocalAI/releases) and the [blog](https://localai.io/blog/).
## Features
@@ -231,18 +231,21 @@ Most backends wrap a best-in-class upstream engine. A handful of them are native
| Backend | What it does |
|---------|-------------|
| [vllm.cpp](https://github.com/mudler/vllm.cpp) | From-scratch C++20 port of vLLM for text generation: paged KV cache, continuous batching, prefix caching, safetensors + GGUF loading, engine-enforced structured output, on CPU, CUDA, Metal and Vulkan |
| [parakeet.cpp](https://github.com/mudler/parakeet.cpp) | C++/GGML port of NVIDIA NeMo Parakeet ASR (tdt/ctc/rnnt/hybrid), with cache-aware streaming transcription |
| [moss-transcribe.cpp](https://github.com/localai-org/moss-transcribe.cpp) | C++/GGML port of OpenMOSS MOSS-Transcribe-Diarize: joint long-form transcription, speaker diarization and timestamping in a single pass |
| [moss-tts.cpp](https://github.com/mudler/moss-tts.cpp) | C++/GGML port of the OpenMOSS MOSS-TTS family: text-to-speech (MOSS-TTS-Local v1.5, 48 kHz stereo) with reference-audio voice cloning, through the MOSS-Audio-Tokenizer neural codec |
| [magpie-tts.cpp](https://github.com/mudler/magpie-tts.cpp) | C++/GGML port of NVIDIA's Magpie TTS Multilingual 357M: 22.05 kHz mono text-to-speech in 5 voices and 9+ languages, with the NanoCodec neural codec and tokenizer/G2P embedded in a single GGUF |
| [ced.cpp](https://github.com/localai-org/ced.cpp) | C++/GGML port of the CED audio-tagging models: sound-event classification (527-class AudioSet) over REST and the realtime API for live recognition |
| [voice-detect.cpp](https://github.com/localai-org/voice-detect.cpp) | Speaker recognition and voice analysis (ECAPA-TDNN, WeSpeaker, ERes2Net, CAM++, wav2vec2 age/gender/emotion), replacing the Python speaker-recognition backend |
| [voxtral-tts.c](https://github.com/mudler/voxtral-tts.c) | Voxtral Realtime 4B speech-to-text in pure C |
| [voxtral-tts.c](https://github.com/mudler/voxtral-tts.c) | Mistral Voxtral-4B-TTS text-to-speech in pure C: 20 preset voices across 9 languages, 24 kHz WAV output, no dependencies beyond libc |
| [vibevoice.cpp](https://github.com/mudler/vibevoice.cpp) | Native port of Microsoft VibeVoice for TTS (voice cloning) and long-form ASR with speaker diarization |
| [rf-detr.cpp](https://github.com/localai-org/rf-detr.cpp) | Native RF-DETR object detection and instance segmentation |
| [locate-anything.cpp](https://github.com/mudler/locate-anything.cpp) | Open-vocabulary object detection and visual grounding (LocateAnything-3B) |
| [depth-anything.cpp](https://github.com/mudler/depth-anything.cpp) | Depth Anything 3 monocular metric depth + camera pose estimation |
| [face-detect.cpp](https://github.com/mudler/face-detect.cpp) | Face detection, recognition, demographics and anti-spoofing (SCRFD/ArcFace, YuNet/SFace), replacing the Python insightface backend |
| [free-splatter.cpp](https://github.com/localai-org/free-splatter.cpp) | Pose-free 3D reconstruction (FreeSplatter): turns a handful of plain photos into 3D Gaussians, no camera poses or GPU required |
| [trellis2.cpp](https://github.com/localai-org/trellis2cpp) | C++/GGML port of Microsoft TRELLIS.2: single-image to textured 3D mesh (GLB with PBR materials) |
| [privacy-filter.cpp](https://github.com/localai-org/privacy-filter.cpp) | Standalone GGML PII/NER token-classification engine powering LocalAI's PII redaction tier |
| [LocalVQE](https://github.com/localai-org/LocalVQE) | Joint acoustic echo cancellation, noise suppression, and dereverberation |
| [local-store](https://github.com/mudler/LocalAI) | Local-first vector database for embeddings (shipped in-tree) |
@@ -257,7 +260,7 @@ We also maintain [apex-quant](https://github.com/localai-org/apex-quant), a per-
- [Kubernetes installation](https://localai.io/basics/getting_started/#run-localai-in-kubernetes)
- [Integrations & community projects](https://localai.io/docs/integrations/)
- [Installation video walkthrough](https://www.youtube.com/watch?v=cMVNnlqwfw4)
- [Media & blog posts](https://localai.io/basics/news/#media-blogs-social)
- [Blog: release write-ups, benchmarks and engineering notes](https://localai.io/blog/)
- [Examples](https://github.com/mudler/LocalAI-examples) — including the [realtime voice assistant demo](https://github.com/localai-org/localai-realtime-demo) (Go client for the Realtime API with tool calling)
## Team

View File

@@ -0,0 +1,120 @@
ARG BASE_IMAGE=ubuntu:24.04
ARG APT_MIRROR=""
ARG APT_PORTS_MIRROR=""
# audio-cpp: 0xShug0/audio.cpp, a ggml audio inference framework covering TTS,
# ASR, VAD, diarization, source separation and music generation, wrapped as a
# LocalAI gRPC backend.
#
# BASE_IMAGE is ubuntu:24.04 for cpu and vulkan builds, or
# nvidia/cuda:<ver>-devel-ubuntu24.04 for cublas builds; both ship apt and
# Ubuntu Noble packages, and the CUDA base additionally provides
# /usr/local/cuda. BUILD_TYPE selects the engine backend in the Makefile:
# "" = portable CPU with all ggml CPU variants, "cublas" ->
# -DENGINE_ENABLE_CUDA=ON, "vulkan" -> -DENGINE_ENABLE_VULKAN=ON. Darwin
# (Metal) builds bypass this Dockerfile entirely.
#
# Upstream needs GCC 13 or newer, which ubuntu:24.04 and the CUDA 12/13
# devel-ubuntu24.04 images all provide.
#
# THIS BACKEND CANNOT USE .docker/install-base-deps.sh OR THE PREBUILT
# quay.io/go-skynet/ci-cache:base-grpc-* IMAGES, AND THAT IS NOT A STYLE CHOICE.
#
# Both supply gRPC v1.65 built from source at /opt/grpc, which downstream
# Dockerfiles copy to /usr/local. That gRPC vendors protobuf v26, and protobuf
# has depended on abseil since v22: google/protobuf/message_lite.h includes
# absl/strings/cord.h. audio.cpp links sentencepiece, and our CMakeLists sets
# SPM_PROTOBUF_PROVIDER=package so sentencepiece uses the same protobuf the
# generated backend.pb.cc was built against (the alternative broke every
# nested-message parse; the full account is in backend/cpp/audio-cpp/CMakeLists.txt).
# That makes sentencepiece's init.h include the external message_lite.h while it
# still includes its own vendored mini-abseil from third_party/absl. The vendored
# copy declares `namespace absl { namespace internal { ... } }` and real abseil
# declares `namespace absl { inline namespace lts_20240116 { namespace internal
# { ... } } }`, so every `absl::internal::` reference becomes ambiguous and the
# compile dies in absl/base/casts.h. Verified, not theorised: building this image
# against the base-grpc-amd64 prebuilt fails at
# sentencepiece-static/error.cc.o with "reference to 'internal' is ambiguous".
#
# Ubuntu Noble's apt protobuf is 3.21.12, which predates the abseil dependency,
# so message_lite.h pulls in no abseil and the vendored copy is the only one in
# scope. That is also the exact protobuf/gRPC pair every unit and end-to-end run
# of this backend has been verified against. Keep it: a from-source gRPC here
# does not buy a faster build, it buys a broken one.
#
# The install-base-deps path is additionally unsafe because it drops protoc 27.1
# into /usr/local/bin, which shadows apt's protoc on PATH and would generate
# protobuf-27 sources to be compiled against 3.21 headers.
FROM ${BASE_IMAGE} AS builder
ARG BUILD_TYPE
ARG TARGETARCH
ARG TARGETVARIANT
ARG APT_MIRROR
ARG APT_PORTS_MIRROR
# Selects the CUDA architecture list in backend/cpp/audio-cpp/Makefile. It has
# to be forwarded: upstream compiles engine_runtime for `native` when
# CMAKE_CUDA_ARCHITECTURES is unset, and no CI runner has a GPU to enumerate.
# The value is the same cuda-major-version the matrix entry declares.
ARG CUDA_MAJOR_VERSION
ENV BUILD_TYPE=${BUILD_TYPE} \
CUDA_MAJOR_VERSION=${CUDA_MAJOR_VERSION} \
APT_MIRROR=${APT_MIRROR} \
APT_PORTS_MIRROR=${APT_PORTS_MIRROR} \
DEBIAN_FRONTEND=noninteractive \
PATH=/usr/local/cuda/bin:${PATH}
WORKDIR /build
# gRPC/protobuf from apt, deliberately; see the block above. libgrpc++-dev ships
# a CMake config so find_package(gRPC CONFIG) resolves, and libprotobuf-dev
# lands in the layout CMake's FindProtobuf module expects, which matters because
# sentencepiece runs a bare find_package(Protobuf REQUIRED) with no CONFIG
# fallback of its own.
#
# BUILD_TYPE=vulkan additionally needs the loader headers and glslc; both are in
# Noble. The CUDA toolkit for BUILD_TYPE=cublas comes from BASE_IMAGE.
RUN --mount=type=bind,source=.docker/apt-mirror.sh,target=/usr/local/sbin/apt-mirror \
sh /usr/local/sbin/apt-mirror && \
apt-get update && \
apt-get install -y --no-install-recommends \
git cmake build-essential pkg-config ca-certificates \
libgrpc++-dev libprotobuf-dev protobuf-compiler protobuf-compiler-grpc && \
if [ "${BUILD_TYPE}" = "vulkan" ]; then \
apt-get install -y --no-install-recommends libvulkan-dev glslc; \
fi && \
if [ "${TARGETARCH}" = "arm64" ]; then \
apt-get install -y --no-install-recommends gcc-14 g++-14; \
fi && \
apt-get clean && \
rm -rf /var/lib/apt/lists/*
COPY . /LocalAI
# gcc-14 on arm64, for the same reason llama-cpp does it in
# .docker/llama-cpp-compile.sh: ggml's CPU_ALL_VARIANTS table includes armv9.2
# variants built with -march=...+sme, and Noble's default gcc-13 rejects that
# feature modifier outright ("invalid feature modifier 'sme'"). Every variant in
# the table has to COMPILE even though a host only ever dlopens the one its own
# CPU supports, so one unbuildable variant fails the whole image.
#
# ON EVERY arm64 BUILD_TYPE, which is where this differs from llama-cpp's script
# and why that difference is spelled out rather than assumed. llama-cpp only
# needs gcc-14 for its pure-CPU image because its GPU builds run
# llama-cpp-fallback, which has no variant table at all. This backend's Makefile
# sets ENGINE_ENABLE_CPU_ALL_VARIANTS for every non-Darwin build, GPU included,
# so an arm64 GPU image would hit the identical compile error. Gating this on an
# empty BUILD_TYPE would leave that trap armed for the first arm64 GPU entry
# added to the matrix, which today has none.
RUN --mount=type=cache,target=/root/.ccache,id=audio-cpp-ccache-${TARGETARCH}-${BUILD_TYPE},sharing=locked \
if [ "${TARGETARCH}" = "arm64" ]; then \
export CC=gcc-14 CXX=g++-14; \
fi && \
make -C /LocalAI/backend/cpp/audio-cpp BUILD_TYPE=${BUILD_TYPE} \
CUDA_MAJOR_VERSION=${CUDA_MAJOR_VERSION} NATIVE=false grpc-server package
# The package directory is the whole image: run.sh, grpc-server, the dlopened
# ggml CPU variants, the bundled loader and its library closure, and the
# bundled silero_vad / marblenet_vad assets. Nothing else exists at run time.
FROM scratch
COPY --from=builder /LocalAI/backend/cpp/audio-cpp/package/. ./

View File

@@ -221,10 +221,75 @@ RUN if [ "${BACKEND}" = "crispasr" ]; then \
apt-get clean && rm -rf /var/lib/apt/lists/*; \
fi
COPY . /LocalAI
# sherpa-onnx links onnxruntime's CUDA execution provider, and
# libonnxruntime_providers_cuda.so has cuDNN as a hard DT_NEEDED. The
# onnxruntime GPU tarball does not ship cuDNN itself, so without this the
# builder has none (the arm64 + CUDA 13 branch above is the only other place
# that installs it) and package-gpu-libs.sh correctly refuses to produce a
# package that references cuDNN with no cuDNN available to it.
#
# Installed per-backend rather than for every cublas build: the auto-detection
# in package-gpu-libs.sh bundles only what a package actually references, so
# the ggml backends would not grow either way, but they would all pay ~1.1 GB
# of builder layer and registry cache for a library they never call.
#
# Runtime package only, no -dev: sherpa-onnx consumes onnxruntime's prebuilt
# CUDA provider and never compiles against cuDNN headers. libcudnn9-cuda-N
# carries the dispatcher plus all seven dlopen()ed sublibraries, which is what
# complete_cudnn_family needs to assemble a whole bundle.
RUN <<EOT bash
if [ "${BACKEND}" = "sherpa-onnx" ] && [ "${BUILD_TYPE}" = "cublas" ] && [ "${SKIP_DRIVERS}" = "false" ]; then
apt-get update && \
apt-get install -y --no-install-recommends \
libcudnn9-cuda-${CUDA_MAJOR_VERSION} && \
ldconfig && \
apt-get clean && \
rm -rf /var/lib/apt/lists/*
fi
EOT
RUN git config --global --add safe.directory /LocalAI
# Prebuild the native engine from a layer that depends on this backend's own
# directory and nothing else.
#
# The expensive part of a C++ backend build is the engine: each of these
# Makefiles clones an upstream repo at a pinned SHA and compiles it once per
# SIMD variant (depth-anything-cpp builds four: avx, avx2, avx512, fallback),
# and those variant targets depend only on the clone. They cannot observe a
# change anywhere else in the LocalAI tree. Building them below `COPY . /LocalAI`
# threw that away: any Go-side edit invalidated the layer and recompiled C++ that
# had not changed. Measured on 2026-07-30, that is a 100+ minute rebuild for the
# larger engines.
#
# Copying only this backend's directory first keeps the compile in a layer that
# survives any change elsewhere in the tree, so `cache-from: type=registry`
# restores it. That covers the expensive cases directly: a shared-build-input or
# backend.proto change, the weekly full-matrix cron and a tag push all rebuild
# every backend while touching none of their directories. This is the mechanism
# behind base-grpc-* applied one level down; unlike a --mount=type=cache it is a
# real layer, which is what actually survives to the registry.
#
# The whole directory rather than just the Makefile: the CMake targets also need
# CMakeLists.txt, and the file list differs per backend. The cost is that editing
# this backend's Go sources also invalidates the engine layer.
#
# Backends whose Makefile has no `engine` target are unaffected: the guard skips
# the prebuild and their engine still compiles in the `build` step below.
COPY backend/go/${BACKEND}/ /LocalAI/backend/go/${BACKEND}/
RUN cd /LocalAI/backend/go/${BACKEND} && \
if make -n engine >/dev/null 2>&1; then \
echo "==> prebuilding engine for ${BACKEND} (cacheable layer)" && \
make engine; \
else \
echo "==> ${BACKEND} has no engine target; it builds with the backend"; \
fi
COPY . /LocalAI
# The engine variants built above survive this COPY (they are build outputs, not
# tracked files) and are newer than the pinned clone, so make treats them as up
# to date and goes straight to the Go binary.
RUN cd /LocalAI && make protogen-go && make -C /LocalAI/backend/go/${BACKEND} build
FROM scratch

View File

@@ -111,6 +111,10 @@ RUN make -BC /LocalAI/backend/cpp/llama-cpp package
# ============================================================================
FROM ${BUILDER_BASE_IMAGE} AS builder-prebuilt
ARG APT_MIRROR
ENV APT_MIRROR=${APT_MIRROR}
ARG APT_PORTS_MIRROR
ENV APT_PORTS_MIRROR=${APT_PORTS_MIRROR}
ARG BUILD_TYPE
ENV BUILD_TYPE=${BUILD_TYPE}
ARG CUDA_DOCKER_ARCH

View File

@@ -56,6 +56,7 @@ The backend system provides language-specific Dockerfiles that handle the build
- **stablediffusion-ggml**: Stable Diffusion in Go with GGML Cpp backend
- **piper**: Text-to-speech synthesis Golang with C bindings using rhaspy/piper
- **local-store**: Vector storage backend
- **valkey-store**: Durable vector storage backend backed by Valkey Search (FT.*)
#### C++ Backends (`cpp/`)
- **llama-cpp**: Llama.cpp integration

View File

@@ -15,7 +15,9 @@ service Backend {
rpc PredictStream(PredictOptions) returns (stream Reply) {}
rpc Embedding(PredictOptions) returns (EmbeddingResult) {}
rpc GenerateImage(GenerateImageRequest) returns (Result) {}
rpc UpscaleImage(UpscaleImageRequest) returns (Result) {}
rpc GenerateVideo(GenerateVideoRequest) returns (Result) {}
rpc Generate3D(Generate3DRequest) returns (Result) {}
rpc AudioTranscription(TranscriptRequest) returns (TranscriptResult) {}
rpc AudioTranscriptionStream(TranscriptRequest) returns (stream TranscriptStreamResponse) {}
// AudioTranscriptionLive is the bidirectional live-microphone ASR RPC. The
@@ -34,6 +36,7 @@ service Backend {
rpc TTSStream(TTSRequest) returns (stream Reply) {}
rpc SoundGeneration(SoundGenerationRequest) returns (Result) {}
rpc TokenizeString(PredictOptions) returns (TokenizationResponse) {}
rpc Detokenize(DetokenizeRequest) returns (DetokenizeResponse) {}
rpc Status(HealthMessage) returns (StatusResponse) {}
rpc Detect(DetectOptions) returns (DetectResponse) {}
// SoundDetection runs an audio-tagging / sound-event-classification model
@@ -181,6 +184,13 @@ message ScoreRequest {
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 5;
// Byte length of the prompt prefix that stays identical across
// repeated scoring calls (e.g. a classifier's option-list system
// prompt — everything before the per-turn probe text). Backends that
// snapshot state (hybrid/recurrent models cannot rewind otherwise)
// use it to place a reuse point exactly at the boundary, so the next
// call re-processes only the tokens after it. 0 means unknown.
int32 stable_prefix_len = 6;
}
// CandidateScore is one row in the ScoreResponse, matching by index
@@ -493,6 +503,11 @@ message ModelOptions {
// Proxy carries the cloud-proxy backend's per-model configuration.
// Empty for non-proxy backends.
ProxyOptions Proxy = 74;
// EnableScore reserves backend resources for the Score RPC. It is derived
// from the model's explicit `known_usecases: [score]` declaration so models
// that never score retain their ordinary serving footprint.
bool EnableScore = 75;
}
// ProxyOptions configures the cloud-proxy backend. UpstreamURL and
@@ -508,6 +523,12 @@ message ProxyOptions {
string api_key_file = 5;
string upstream_model = 6;
int32 request_timeout_seconds = 7;
// cache_prompt enables automatic Anthropic prompt-cache breakpoints
// (cache_control: ephemeral) on the stable prefix — system, tools, and
// the last message block — when translating to the Anthropic provider.
// Cuts input cost on repeated/agentic calls (cache read = 0.1x). Only
// meaningful for mode=translate + provider=anthropic; ignored otherwise.
bool cache_prompt = 8;
}
message Result {
@@ -617,6 +638,12 @@ message GenerateImageRequest {
string ModelIdentity = 13;
}
message UpscaleImageRequest {
string src = 1; // input image path
string dst = 2; // output image path
int32 scale = 3; // upscale factor (e.g. 2 or 4)
}
message GenerateVideoRequest {
string prompt = 1;
string negative_prompt = 2; // Negative prompt for video generation
@@ -640,6 +667,20 @@ message GenerateVideoRequest {
string ModelIdentity = 15;
}
message Generate3DRequest {
string src = 1; // Path to the staged conditioning image (3D generation is image-conditioned)
string dst = 2; // Output path for the generated binary glTF (.glb) asset
int32 seed = 3; // <=0 lets the backend pick a random seed
int32 step = 4; // Flow sampling steps; <=0 uses the backend default
float cfg_scale = 5; // Classifier-free guidance scale; <=0 uses the backend default
int32 texture_steps = 6; // Texture flow sampling steps; <=0 uses the backend default
string quality = 7; // Mesh pipeline: ""|"auto"|"coarse"|"512"|"1024"
string background = 8; // Conditioning-image background handling: ""|"auto"|"keep"|"black"|"white"
// Backend-specific per-request generation parameters. Values are strings
// and are validated/coerced by the selected backend.
map<string, string> params = 9;
}
message TTSRequest {
string text = 1;
string model = 2;
@@ -763,6 +804,14 @@ message TokenizationResponse {
repeated int32 tokens = 2;
}
message DetokenizeRequest {
repeated int32 tokens = 1;
}
message DetokenizeResponse {
string content = 1;
}
message MemoryUsageData {
uint64 total = 1;
map<string, uint64> breakdown = 2;
@@ -1089,11 +1138,28 @@ message AudioTransformRequest {
string ModelIdentity = 5;
}
// One named output of a transform that produces several from a single run.
// Source separation is the case that needs it: htdemucs yields drums, bass,
// other and vocals from one pass over the input.
message AudioTransformStem {
string name = 1; // the model's own stem id, e.g. "vocals"
string dst = 2; // path of the file written for that stem
}
message AudioTransformResult {
string dst = 1;
int32 sample_rate = 2;
int32 samples = 3;
bool reference_provided = 4;
// Every named output the run produced, in the model's own order, including
// the one copied into dst. Empty for a transform with a single output.
//
// It exists because dst carries one file while separation produces several,
// and running the model once per stem would cost four full separations of
// the same audio. The backend runs once, writes each stem beside dst, and
// names them here; without this field the other stems are on disk but no
// caller can find them, which is the same as not having produced them.
repeated AudioTransformStem stems = 5;
}
// Bidirectional streaming audio transform. The first message MUST carry a

8
backend/cpp/audio-cpp/.gitignore vendored Normal file
View File

@@ -0,0 +1,8 @@
audio.cpp/
build/
package/
grpc-server
backend.pb.cc
backend.pb.h
backend.grpc.pb.cc
backend.grpc.pb.h

View File

@@ -0,0 +1,331 @@
cmake_minimum_required(VERSION 3.20)
project(audio-cpp-grpc-server LANGUAGES C CXX)
set(CMAKE_CXX_STANDARD 17)
set(CMAKE_CXX_STANDARD_REQUIRED ON)
set(TARGET grpc-server)
set(AUDIO_CPP_DIR "${CMAKE_CURRENT_SOURCE_DIR}/audio.cpp"
CACHE PATH "Path to the pinned audio.cpp checkout")
option(AUDIO_CPP_GRPC_BUILD_TESTS "Build engine-linked ctest binaries" OFF)
if(NOT EXISTS "${AUDIO_CPP_DIR}/CMakeLists.txt")
message(FATAL_ERROR
"AUDIO_CPP_DIR does not contain an audio.cpp checkout: ${AUDIO_CPP_DIR}. "
"Run 'make audio.cpp' first.")
endif()
if(APPLE)
# Homebrew installs protobuf/grpc under a non-default prefix.
if(CMAKE_HOST_SYSTEM_PROCESSOR MATCHES "arm64")
set(HOMEBREW_DEFAULT_PREFIX "/opt/homebrew")
else()
set(HOMEBREW_DEFAULT_PREFIX "/usr/local")
endif()
link_directories("${HOMEBREW_DEFAULT_PREFIX}/lib")
include_directories("${HOMEBREW_DEFAULT_PREFIX}/include")
endif()
find_package(Threads REQUIRED)
find_package(Protobuf CONFIG QUIET)
if(NOT Protobuf_FOUND)
find_package(Protobuf REQUIRED)
endif()
find_package(gRPC CONFIG QUIET)
if(NOT gRPC_FOUND)
# Reached only on distros whose grpc++ packaging ships no CMake config.
# Ubuntu's libgrpc-dev does ship one, so this is dead code on LocalAI's own
# build distro. Kept for the distros that do not.
find_library(GRPCPP_LIB grpc++ REQUIRED)
find_library(GRPCPP_REFLECTION_LIB grpc++_reflection REQUIRED)
add_library(gRPC::grpc++ INTERFACE IMPORTED)
set_target_properties(gRPC::grpc++ PROPERTIES
INTERFACE_LINK_LIBRARIES "${GRPCPP_LIB}")
add_library(gRPC::grpc++_reflection INTERFACE IMPORTED)
set_target_properties(gRPC::grpc++_reflection PROPERTIES
INTERFACE_LINK_LIBRARIES "${GRPCPP_REFLECTION_LIB}")
endif()
find_program(_PROTOC NAMES protoc REQUIRED)
find_program(_GRPC_CPP_PLUGIN NAMES grpc_cpp_plugin REQUIRED)
get_filename_component(HW_PROTO "${CMAKE_CURRENT_SOURCE_DIR}/../../backend.proto" ABSOLUTE)
get_filename_component(HW_PROTO_PATH "${HW_PROTO}" PATH)
set(HW_PROTO_SRCS "${CMAKE_CURRENT_BINARY_DIR}/backend.pb.cc")
set(HW_PROTO_HDRS "${CMAKE_CURRENT_BINARY_DIR}/backend.pb.h")
set(HW_GRPC_SRCS "${CMAKE_CURRENT_BINARY_DIR}/backend.grpc.pb.cc")
set(HW_GRPC_HDRS "${CMAKE_CURRENT_BINARY_DIR}/backend.grpc.pb.h")
add_custom_command(
OUTPUT "${HW_PROTO_SRCS}" "${HW_PROTO_HDRS}" "${HW_GRPC_SRCS}" "${HW_GRPC_HDRS}"
COMMAND ${_PROTOC}
ARGS --grpc_out "${CMAKE_CURRENT_BINARY_DIR}"
--cpp_out "${CMAKE_CURRENT_BINARY_DIR}"
-I "${HW_PROTO_PATH}"
--plugin=protoc-gen-grpc="${_GRPC_CPP_PLUGIN}"
"${HW_PROTO}"
DEPENDS "${HW_PROTO}")
add_library(hw_grpc_proto STATIC
${HW_GRPC_SRCS} ${HW_GRPC_HDRS}
${HW_PROTO_SRCS} ${HW_PROTO_HDRS})
target_include_directories(hw_grpc_proto PUBLIC ${CMAKE_CURRENT_BINARY_DIR})
# Required on macOS: without these the Homebrew protobuf/grpc include dirs never
# reach this target and google/protobuf/runtime_version.h is not found.
target_link_libraries(hw_grpc_proto PUBLIC protobuf::libprotobuf gRPC::grpc++)
# TWO PROTOBUF RUNTIMES IN ONE BINARY, AND THE ONE THAT WON WAS THE WRONG ONE.
#
# engine_runtime links sentencepiece, whose default SPM_PROTOBUF_PROVIDER
# ("internal") builds the protobuf-lite 3.14.0 sources vendored under
# external/sentencepiece/third_party/protobuf-lite. Our generated backend.pb.cc
# is compiled against the toolchain's protobuf 3.21.12 headers and links
# libprotobuf.so 3.21.12. Both used to end up in the executable: 476
# google::protobuf:: symbols from that archive, 278 of them also defined by
# libprotobuf.so.
#
# The binding is decided at STATIC LINK time. Once ld pulls a sentencepiece
# member in for sentencepiece's own code, that member's protobuf definitions are
# in the executable and references from libhw_grpc_proto.a bind to them. Do NOT
# reach for -Wl,--exclude-libs: it flips those symbols to LOCAL in .dynsym and
# the breakage is unchanged, because no visibility flag revisits a static
# binding already made.
#
# What broke, measured rather than assumed:
# google::protobuf::internal::ParseContext::ParseMessage(MessageLite*, const char*)
# is what every generated _InternalParse calls for a SUBMESSAGE field and for
# nothing else. Bound to the 3.14 definition it fails, so a flat message parsed
# and every nested one did not: a TranscriptResult carrying segments serialized
# to correct bytes that the same process could not read back, and
# TranscriptLiveRequest, a oneof of submessages, could not have been parsed at
# all. 3.21 generated code was also running 3.14 arena, ArenaStringPtr and
# ExtensionSet code, which is an ABI mismatch rather than a missing feature, so
# "not observed to bite yet" was never a reason to leave it.
#
# "package" makes sentencepiece use the protobuf found above, which is the one
# the generated code was built against. It must be set before add_subdirectory,
# since that is when sentencepiece's own cache entry is created.
#
# WHAT THAT BUYS IS ONE PROTOBUF RUNTIME, not an executable free of protobuf
# symbols, and the difference matters to whoever checks this next. Measured with
# nm -C --defined-only on the linked grpc-server, 2515 google::protobuf::
# symbols are still DEFINED in it, and that is what should be there: they are
# generated code, sentencepiece::ModelProto's own _InternalParse and
# CheckTypeAndMergeFrom among them, which name protobuf types in their
# signatures and are compiled into every user of a .proto. Expecting zero would
# send a reader looking for a regression that is not one.
#
# The claim that decides whether the ABI mismatch above is gone is the RUNTIME
# one, and it holds: google::protobuf::internal::ParseContext::ParseMessage is
# UNDEFINED in the executable, so every generated _InternalParse resolves it to
# libprotobuf.so at load instead of to a vendored 3.14 copy. No vendored
# protobuf-lite archive is pulled in at all, and citrinet_asr, which parses a
# SentencePiece ModelProto at load, tokenizes correctly as a result.
set(SPM_PROTOBUF_PROVIDER "package" CACHE STRING
"Make sentencepiece use the found protobuf, not its vendored 3.14 copy" FORCE)
# Upstream's global add_compile_options(-Wall -Wextra -Wpedantic -pedantic-errors)
# is a directory property of the subdirectory and does not reach our targets.
#
# EXCLUDE_FROM_ALL is load-bearing, do not drop it: upstream's default target set
# includes its CLI, server, converter and test binaries, none of which we ship.
# Without it every build would compile all of them. The targets we do name in
# target_link_libraries below are still built on demand, so nothing is lost.
add_subdirectory("${AUDIO_CPP_DIR}" "${CMAKE_CURRENT_BINARY_DIR}/audio-cpp" EXCLUDE_FROM_ALL)
add_executable(${TARGET}
grpc-server.cpp
model_options.cpp
capability_routing.cpp
family_gate.cpp
loaded_model.cpp
audio_io.cpp
audio_units.cpp
transcript_assembly.cpp
result_map.cpp
stem_selection.cpp
generation_request.cpp
stream_delta.cpp
wav_header.cpp
inference_lane.cpp
live_watchdog.cpp
)
# Two files carry a switch over an enum with no `default:` label, deliberately,
# so that -Wswitch reports an enumerator nobody handled. -Wswitch is only a
# warning by default, and a warning in a 600-file build log is a warning nobody
# reads, so it is promoted to an error on exactly these two translation units.
# Not project-wide: upstream's own sources are not held to this, and they are
# where the churn is.
#
# loaded_model.cpp mirrors engine::runtime::VoiceTaskKind onto its own Task enum.
# Its static_asserts catch an insertion or a reorder, but an enumerator APPENDED
# after the last one shifts no value, so no assertion can see it. What does see
# it is from_engine_task's switch over the engine enum. This is the difference
# between a build failure and a backend that silently runs the wrong task.
#
# capability_routing.cpp's unsupported_surface() switches UnsupportedRpc onto the
# row of unsupported_surfaces() that explains it. Left as a warning, a sixth
# enumerator added without a row BUILDS AND SHIPS, and its trailing
# `return surfaces[0];` then answers the new RPC with AudioEncode's codec reason:
# a confident, specific and false statement about audio.cpp, on the wire, on the
# one code path whose entire job is to be truthful about what this backend
# cannot do. Verified rather than assumed: adding a sixth enumerator and building
# the shipping target produced exit 0, a binary, and one warning. A compile-time
# check is the better trade than the runtime fallback it replaced only if it is
# fatal, so here it is fatal.
if(NOT MSVC)
set_source_files_properties(loaded_model.cpp capability_routing.cpp
PROPERTIES COMPILE_OPTIONS "-Werror=switch")
endif()
target_include_directories(${TARGET} PRIVATE
"${AUDIO_CPP_DIR}/include"
"${CMAKE_CURRENT_SOURCE_DIR}")
# The shipping binary is held to the same bar as the tests below. Upstream's own
# add_compile_options is a property of its directory and never reached this
# target, so until now "the build was clean" meant only that nothing was being
# checked.
if(NOT MSVC)
target_compile_options(${TARGET} PRIVATE -Wall -Wextra -Wpedantic)
endif()
target_link_libraries(${TARGET} PRIVATE
hw_grpc_proto
engine_runtime
ggml
gRPC::grpc++
gRPC::grpc++_reflection
protobuf::libprotobuf
Threads::Threads)
# ENGINE_ENABLE_CPU_ALL_VARIANTS builds ggml backends as shared objects that sit
# next to the binary in the package, so the binary must search its own directory.
# BUILD_WITH_INSTALL_RPATH keeps the build-tree binary at exactly "$ORIGIN".
# Upstream sets CMAKE_BUILD_WITH_INSTALL_RPATH in its own directory scope, which
# does not reach ours, so without this CMake also appends its build-tree library
# directory. That absolute build-host path would survive into the copied binary
# and let a package.sh that forgot to bundle libggml*.so still pass on the build
# machine while failing everywhere else.
set_target_properties(${TARGET} PROPERTIES
BUILD_RPATH "$ORIGIN"
INSTALL_RPATH "$ORIGIN"
BUILD_WITH_INSTALL_RPATH TRUE)
if(AUDIO_CPP_GRPC_BUILD_TESTS)
enable_testing()
# These are the units whose tests CANNOT run under
# backend/cpp/run-unit-tests.sh, because that script compiles each
# *_test.cpp standalone with no protobuf and no audio.cpp include path.
# They are named *_ctest.cpp so the script's glob does not pick them up and
# fail every backend's suite; everything that can be stdlib-only still is,
# and still lives in a *_test.cpp beside its unit.
add_executable(result_map_ctest
result_map_ctest.cpp
result_map.cpp
transcript_assembly.cpp
audio_units.cpp)
target_include_directories(result_map_ctest PRIVATE
"${AUDIO_CPP_DIR}/include"
"${CMAKE_CURRENT_SOURCE_DIR}"
# session.h reaches ggml.h through core/backend.h. Every other target
# here inherits that directory from the ggml target it links; this one
# links no ggml, so it has to name it.
"${AUDIO_CPP_DIR}/external/ggml/include")
# No engine_runtime: result_map touches only the plain structs in
# engine/framework/runtime/session.h, so the header is all it needs.
target_link_libraries(result_map_ctest PRIVATE
hw_grpc_proto
protobuf::libprotobuf
Threads::Threads)
target_compile_options(result_map_ctest PRIVATE -Wall -Wextra -Wpedantic)
add_test(NAME result_map COMMAND result_map_ctest)
# Same shape as result_map_ctest: generation_request touches only the plain
# structs in engine/framework/runtime/session.h plus the generated protobuf
# messages, so the headers are all it needs and no engine_runtime is linked.
add_executable(generation_request_ctest
generation_request_ctest.cpp
generation_request.cpp)
target_include_directories(generation_request_ctest PRIVATE
"${AUDIO_CPP_DIR}/include"
"${CMAKE_CURRENT_SOURCE_DIR}"
# session.h reaches ggml.h through core/backend.h, and this target links
# no ggml, so it has to name the include directory itself.
"${AUDIO_CPP_DIR}/external/ggml/include")
target_link_libraries(generation_request_ctest PRIVATE
hw_grpc_proto
protobuf::libprotobuf
Threads::Threads)
target_compile_options(generation_request_ctest PRIVATE -Wall -Wextra -Wpedantic)
add_test(NAME generation_request COMMAND generation_request_ctest)
add_executable(audio_io_ctest
audio_io_ctest.cpp
audio_io.cpp)
target_include_directories(audio_io_ctest PRIVATE
"${AUDIO_CPP_DIR}/include"
"${CMAKE_CURRENT_SOURCE_DIR}")
target_link_libraries(audio_io_ctest PRIVATE
engine_runtime
ggml
Threads::Threads)
target_compile_options(audio_io_ctest PRIVATE -Wall -Wextra -Wpedantic)
set_target_properties(audio_io_ctest PROPERTIES
BUILD_RPATH "$ORIGIN"
INSTALL_RPATH "$ORIGIN"
BUILD_WITH_INSTALL_RPATH TRUE)
add_test(NAME audio_io COMMAND audio_io_ctest)
# The streaming drivers live in loaded_model.cpp, which links the engine, so
# this cannot be a standalone *_test.cpp. It builds no model and reads no
# file: LoadedModel::Session is a plain struct holding a pointer to an
# engine interface, so the drivers are exercised against fake sessions.
add_executable(streaming_driver_ctest
streaming_driver_ctest.cpp
loaded_model.cpp
capability_routing.cpp
family_gate.cpp
model_options.cpp
inference_lane.cpp)
target_include_directories(streaming_driver_ctest PRIVATE
"${AUDIO_CPP_DIR}/include"
"${CMAKE_CURRENT_SOURCE_DIR}")
target_link_libraries(streaming_driver_ctest PRIVATE
engine_runtime
ggml
Threads::Threads)
target_compile_options(streaming_driver_ctest PRIVATE -Wall -Wextra -Wpedantic)
# No "$ORIGIN" rpath override here, unlike the shipping target and unlike
# audio_io_ctest. loaded_model.cpp reaches make_default_registry, so this
# binary genuinely links libggml, and CMake's own build-tree rpath is what
# finds it: the ggml shared objects land in ${CMAKE_CURRENT_BINARY_DIR}/bin
# while the test binary sits one directory up. A build-host absolute path in
# a test binary is harmless, since package.sh ships only grpc-server, and
# forcing "$ORIGIN" here means ctest cannot start the binary at all.
add_test(NAME streaming_driver COMMAND streaming_driver_ctest)
# Asserts that the upstream ABSENCES capability_routing.cpp's refusal
# messages rest on are still absences, by querying make_default_registry()
# rather than by re-reading upstream. This is what makes an AUDIO_CPP_VERSION
# bump that adds a codec task kind, an spk family or a streaming converter
# fail the build instead of leaving a false statement on the wire.
#
# It links engine_runtime purely to run that query, which is why it lives
# here rather than with the standalone *_test.cpp files, and it needs no
# "$ORIGIN" rpath override for the same reason streaming_driver_ctest does
# not: see the note above.
add_executable(upstream_absence_ctest upstream_absence_ctest.cpp)
target_include_directories(upstream_absence_ctest PRIVATE
"${AUDIO_CPP_DIR}/include"
"${CMAKE_CURRENT_SOURCE_DIR}")
target_link_libraries(upstream_absence_ctest PRIVATE
engine_runtime
ggml
Threads::Threads)
target_compile_options(upstream_absence_ctest PRIVATE -Wall -Wextra -Wpedantic)
add_test(NAME upstream_absence COMMAND upstream_absence_ctest)
endif()

View File

@@ -0,0 +1,172 @@
# audio.cpp backend Makefile.
#
# Upstream pin lives below in the AUDIO_CPP_VERSION variable, so
# .github/bump_deps.sh can find and update it, matching the llama-cpp / ds4
# convention. That script seds every line matching the variable name followed by
# an assignment, so this comment deliberately spells the name on its own: a
# comment repeating the full assignment token gets rewritten and mangled by the
# first auto-bump (backend/cpp/ds4/Makefile shows the damage). The clone
# recipe is a make target (not a prepare.sh) so 'make purge && make' is a clean
# rebuild and so the bump bot can see the pin.
AUDIO_CPP_VERSION?=4e3aea2fd99aeaa5924e71c51eb2793846045332
AUDIO_CPP_REPO?=https://github.com/0xShug0/audio.cpp
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))
BUILD_DIR := build
BUILD_TYPE ?=
NATIVE ?= false
JOBS ?= $(shell nproc 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 4)
UNAME_S := $(shell uname -s)
# AUDIOCPP_DEPLOYMENT_BUILD compiles the model_specs/*.json catalog into
# engine_runtime, so the shipped package needs no model_specs directory and a
# safetensors model tree still resolves its family spec.
CMAKE_ARGS ?= -DCMAKE_BUILD_TYPE=Release -DAUDIOCPP_DEPLOYMENT_BUILD=ON
# CMAKE_CUDA_ARCHITECTURES must be set explicitly for a cublas build, and this
# is not a tuning knob: upstream's CMakeLists sets CUDA_ARCHITECTURES to
# `native` on the engine_runtime target whenever the root-scope variable is
# unset (audio.cpp/CMakeLists.txt, the `if (CMAKE_CUDA_ARCHITECTURES)` branch
# next to the istft/torch_random .cu sources), and docs/build/linux.md says so
# outright: "Leave CMAKE_CUDA_ARCHITECTURES unset to build for the GPUs present
# at build time (native)". No CI runner has a GPU, so `native` has nothing to
# enumerate. ggml's own default (external/ggml/src/ggml-cuda/CMakeLists.txt)
# does not rescue this: it list(APPEND)s in the ggml subdirectory scope, which
# never reaches the root scope where the engine_runtime property is decided.
#
# The values below are ggml's list for the matching toolkit, copied rather than
# invented, so the two targets compile for exactly the same set:
# - CUDA 13 drops the Maxwell/Pascal/Volta virtual archs (50/61/70).
# - 121a-real needs CUDA >= 12.9, so the CUDA 12 list (built against 12.8)
# stops at 120a-real.
# - `a`-suffixed archs are used rather than ggml's rejected 120f-virtual: the
# `f` suffix needs CMake >= 3.31.8, and Ubuntu Noble ships 3.28.3. The
# 3.28 validator (Modules/Internal/CMakeCUDAArchitecturesValidate.cmake)
# accepts `[0-9]+a?(-real|-virtual)?`.
#
# Setting it here also pins ggml's copy, since its default is guarded by
# `if (NOT DEFINED CMAKE_CUDA_ARCHITECTURES)`. CUDA_MAJOR_VERSION is the CI
# build-arg, forwarded by Dockerfile.audio-cpp.
#
# An EMPTY major maps to `native`, NOT to the CUDA 12 list. Only CI declares a
# major; a developer running `BUILD_TYPE=cublas make` locally declares none, and
# the CUDA 12 list contains 120a-real, which needs nvcc >= 12.8. Falling through
# to it turned every local build on a CUDA 12.0-12.7 host into a compile error,
# where upstream's documented behaviour ("Leave CMAKE_CUDA_ARCHITECTURES unset
# to build for the GPUs present at build time") worked. `native` restores that.
# It does require a GPU to enumerate, so the escape hatch for a GPU-less local
# cross-build is to set CUDA_ARCHITECTURES on the command line, which the ?=
# assignments below leave untouched.
CUDA_MAJOR_VERSION ?=
ifeq ($(CUDA_MAJOR_VERSION),13)
CUDA_ARCHITECTURES ?= 75-virtual;80-virtual;86-real;89-real;120a-real;121a-real
else ifeq ($(CUDA_MAJOR_VERSION),12)
CUDA_ARCHITECTURES ?= 50-virtual;61-virtual;70-virtual;75-virtual;80-virtual;86-real;89-real;120a-real
else ifeq ($(CUDA_MAJOR_VERSION),)
CUDA_ARCHITECTURES ?= native
else ifeq ($(BUILD_TYPE),cublas)
# Gated on cublas because the variable means nothing to any other build, so a
# stray CUDA_MAJOR_VERSION in the environment must not break `make clean` or
# a CPU build. It does still error for `BUILD_TYPE=cublas make clean`, which
# is the right trade: that invocation is asking about a CUDA build tree.
$(error CUDA_MAJOR_VERSION=$(CUDA_MAJOR_VERSION) has no architecture list here (12 and 13 do). Leave it empty for a native build, or pass CUDA_ARCHITECTURES explicitly.)
endif
ifeq ($(BUILD_TYPE),cublas)
CMAKE_ARGS += -DENGINE_ENABLE_CUDA=ON "-DCMAKE_CUDA_ARCHITECTURES=$(CUDA_ARCHITECTURES)"
else ifeq ($(BUILD_TYPE),vulkan)
CMAKE_ARGS += -DENGINE_ENABLE_VULKAN=ON
else ifeq ($(UNAME_S),Darwin)
# Metal. ggml embeds the shader library by default (GGML_METAL_EMBED_LIBRARY
# defaults to GGML_METAL), so the package needs no .metallib beside the
# binary. Darwin builds go through scripts/build/audio-cpp-darwin.sh.
CMAKE_ARGS += -DENGINE_ENABLE_METAL=ON
# AppleClang ships no OpenMP runtime and Homebrew's libomp is keg-only, so
# neither libomp.dylib nor omp.h is symlinked into /opt/homebrew and CMake's
# FindOpenMP cannot find them on its own (the workflow's `brew link libomp`
# is a no-op for a keg-only formula, and its failure is swallowed).
# audio.cpp calls find_package(OpenMP REQUIRED COMPONENTS CXX) whenever
# ENGINE_ENABLE_OPENMP is ON, so with no hint the macOS build dies at
# configure time before compiling anything. OpenMP_ROOT is honoured by the
# find_library/find_path calls inside FindOpenMP under CMP0074, which is NEW
# here because audio.cpp requires CMake 3.20.
#
# If the keg is absent, turn OpenMP off rather than fail: the tree's only
# <omp.h> include is guarded by #ifdef _OPENMP and a #pragma omp without
# -fopenmp is simply ignored, so an OpenMP-less build is CORRECT. It is not
# cheap, though: 108 `#pragma omp` directives across ~30 files (roformer,
# demucs, chatterbox, moss, supertonic, seed_vc, framework/audio/dsp) are
# compiled out, and clang says nothing about an ignored omp pragma unless
# -Wsource-uses-openmp is on. A green package that is quietly single-threaded
# in every host DSP loop gets blamed on Metal, not on packaging, so the
# fallback announces itself.
ifeq ($(origin LIBOMP_PREFIX),undefined)
LIBOMP_PREFIX := $(shell brew --prefix libomp 2>/dev/null)
endif
# Nested ifneq rather than $(and ...): $(and) needs GNU make 3.81, and while
# that is what Apple ships, an older make expands it to empty and would take
# the OpenMP-OFF branch with no way to tell that from a genuinely missing
# keg. Two plain conditionals cannot fail that way.
LIBOMP_USABLE :=
ifneq ($(wildcard $(LIBOMP_PREFIX)/lib/libomp.dylib),)
ifneq ($(wildcard $(LIBOMP_PREFIX)/include/omp.h),)
LIBOMP_USABLE := yes
endif
endif
ifeq ($(LIBOMP_USABLE),yes)
CMAKE_ARGS += "-DOpenMP_ROOT=$(LIBOMP_PREFIX)"
else
$(warning audio-cpp: libomp not found at '$(LIBOMP_PREFIX)'; building without OpenMP (single-threaded host DSP). Install it with `brew install libomp`, or set LIBOMP_PREFIX.)
CMAKE_ARGS += -DENGINE_ENABLE_OPENMP=OFF
endif
else
# Portable Linux CPU. Upstream wires this to GGML_BACKEND_DL +
# GGML_CPU_ALL_VARIANTS + $ORIGIN rpath, so one build serves every CPU
# tier instead of an AVX-tier image fan-out.
CMAKE_ARGS += -DENGINE_ENABLE_CPU_ALL_VARIANTS=ON
endif
ifneq ($(NATIVE),true)
CMAKE_ARGS += -DENGINE_ENABLE_NATIVE_CPU=OFF
endif
.PHONY: all grpc-server package test test-engine clean purge
all: grpc-server
# Clone the upstream source at the pinned commit. The directory is the target
# so make only re-clones when it is missing. After bumping AUDIO_CPP_VERSION,
# run 'make purge && make' to refetch.
audio.cpp:
mkdir -p audio.cpp
cd audio.cpp && \
git init -q && \
git remote add origin $(AUDIO_CPP_REPO) && \
git fetch --depth 1 origin $(AUDIO_CPP_VERSION) && \
git checkout FETCH_HEAD
grpc-server: audio.cpp
mkdir -p $(BUILD_DIR)
cd $(BUILD_DIR) && cmake $(CMAKE_ARGS) $(CURRENT_MAKEFILE_DIR) && \
cmake --build . --config Release -j $(JOBS)
cp $(BUILD_DIR)/grpc-server grpc-server
package: grpc-server
bash package.sh
test:
@echo "audio-cpp: standalone unit tests run from the repo root via 'make test-backend-cpp'"
# Engine-linked tests. Needs the upstream checkout and a full engine build.
test-engine: audio.cpp
mkdir -p $(BUILD_DIR)
cd $(BUILD_DIR) && cmake $(CMAKE_ARGS) -DAUDIO_CPP_GRPC_BUILD_TESTS=ON $(CURRENT_MAKEFILE_DIR) && \
cmake --build . --config Release -j $(JOBS) && ctest --output-on-failure --no-tests=error
clean:
rm -rf $(BUILD_DIR) grpc-server package
purge: clean
rm -rf audio.cpp

View File

@@ -0,0 +1,124 @@
#include "audio_io.h"
#include "loaded_model.h"
#include "engine/framework/audio/conversion.h"
#include "engine/framework/audio/wav_reader.h"
#include "engine/framework/audio/wav_writer.h"
#include <filesystem>
#include <string>
#include <utility>
namespace audiocpp_backend {
engine::runtime::AudioBuffer read_audio_file(const std::string &path,
int target_sample_rate) {
if (path.empty()) {
throw ConfigError("audio-cpp: no input audio path was supplied");
}
std::error_code ec;
const bool present = std::filesystem::exists(std::filesystem::path(path), ec);
if (ec) {
// exists() returning false with ec set does NOT mean the file is
// absent, it means the question could not be answered: most often a
// parent directory is not searchable. Reporting that as "does not
// exist" sends the operator after the file when the fault is the
// permissions on the directory above it.
throw ConfigError("audio-cpp: cannot stat input audio " + path + ": " +
ec.message());
}
if (!present) {
throw ConfigError("audio-cpp: input audio does not exist: " + path);
}
engine::audio::WavData wav;
try {
wav = engine::audio::read_wav_f32(std::filesystem::path(path));
} catch (const std::exception &err) {
throw ConfigError("audio-cpp: cannot read " + path +
" as WAV: " + err.what());
}
if (wav.sample_rate <= 0) {
throw ConfigError("audio-cpp: " + path +
" declares a non-positive sample rate; every "
"timestamp derived from it would be zero");
}
// AudioBuffer's own default is 1, and a reader that reports 0 channels
// still gave us an interleaving of one. Normalised before the conversion
// below rather than after, because mixdown_interleaved_to_mono_average
// throws on a non-positive channel count.
if (wav.channels <= 0) {
wav.channels = 1;
}
engine::runtime::AudioBuffer buffer;
if (target_sample_rate <= 0) {
buffer.sample_rate = wav.sample_rate;
buffer.channels = wav.channels;
buffer.samples = std::move(wav.samples);
return buffer;
}
buffer.sample_rate = target_sample_rate;
buffer.channels = 1;
try {
// A no-op copy when the rates already match, so the common 16 kHz
// upload pays only the mono mixdown it would have paid inside the
// family anyway.
buffer.samples =
engine::audio::convert_wav_to_mono_linear_resampled(wav, target_sample_rate);
} catch (const std::exception &err) {
// ConfigError, so this is INVALID_ARGUMENT rather than INTERNAL. What
// reaches here is a malformed input: a sample count that is not a whole
// number of frames is the realistic one, and it is the uploader's file
// that is truncated, not this backend that is broken.
throw ConfigError("audio-cpp: cannot resample " + path + " from " +
std::to_string(wav.sample_rate) + " Hz to " +
std::to_string(target_sample_rate) +
" Hz: " + err.what());
}
return buffer;
}
void write_audio_file(const std::string &path,
const engine::runtime::AudioBuffer &audio) {
if (path.empty()) {
throw ConfigError("audio-cpp: no output path was supplied");
}
const std::filesystem::path destination(path);
if (destination.has_parent_path()) {
// Best effort: a failure here shows up as a write failure below, with a
// message naming the file the caller actually asked for.
std::error_code ec;
std::filesystem::create_directories(destination.parent_path(), ec);
}
try {
engine::audio::write_pcm16_wav(destination, audio.sample_rate,
audio.channels > 0 ? audio.channels : 1,
audio.samples);
} catch (const std::exception &err) {
// NOT a ConfigError, and the distinction is not cosmetic. The
// destination is chosen by LocalAI rather than by the caller: it is a
// unique name inside GeneratedContentDir. A failure to write it is a
// full disk, a permission fault on the server's own directory, or a bad
// mount, none of which the caller can fix or is to blame for. As a
// ConfigError this surfaced as INVALID_ARGUMENT, which tells a client
// its request was wrong and not to retry; a plain runtime_error maps to
// INTERNAL, which is both true and retryable. The empty path above
// stays INVALID_ARGUMENT, because that one really is a malformed
// request.
throw std::runtime_error("audio-cpp: cannot write " + path + ": " +
err.what());
}
}
engine::runtime::AudioBuffer buffer_from_mono(std::vector<float> samples,
int sample_rate) {
engine::runtime::AudioBuffer buffer;
buffer.sample_rate = sample_rate;
buffer.channels = 1;
buffer.samples = std::move(samples);
return buffer;
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,61 @@
#pragma once
// Thin wrappers over the framework's public audio IO. Engine-linked, so this
// unit is built and tested through the CMake target rather than by
// backend/cpp/run-unit-tests.sh. The pure part of the arithmetic these
// wrappers feed lives in audio_units, which is stdlib-only and does have a
// standalone test.
#include "engine/framework/runtime/session.h"
#include <string>
#include <vector>
namespace audiocpp_backend {
// Reads a WAV file. Throws ConfigError when the file is missing, is not
// readable as WAV, or declares a non-positive sample rate: all three are
// user-fixable input problems rather than backend faults.
//
// A declared sample rate of zero is refused rather than passed on, because
// every downstream conversion in audio_units answers 0 for a non-positive rate.
// Accepting it would turn a corrupt header into a response full of zero
// timestamps, which reads as a real answer.
//
// `target_sample_rate` is the rate the CALLER needs, in Hz:
//
// 0 (or negative) keep the file's own rate and channel count.
// positive downmix to mono and resample to that rate. Resampling is
// skipped when the file already declares it, so passing the
// rate a route needs costs nothing on the common input.
//
// It is a parameter, and not a constant inside this function, because the
// routes that read audio do not agree on an answer. Speech routes want 16 kHz
// mono; source separation does not, and folding a 44.1 kHz stereo input to
// 16 kHz mono for demucs or roformer would destroy the very thing they separate
// (both refuse a rate other than their own outright). Making the caller name
// the rate keeps that decision where the route is known.
//
// Downmixing along with the resample is not an extra liberty: every family a
// positive rate is used for (silero_vad, sortformer_diar and every ASR family)
// begins by calling the same mixdown_interleaved_to_mono_average on whatever it
// is given. Doing it once here produces the identical samples and halves the
// buffer that is then moved through the request.
engine::runtime::AudioBuffer read_audio_file(const std::string &path,
int target_sample_rate);
// Writes 16-bit PCM WAV, creating parent directories.
//
// Throws ConfigError, i.e. INVALID_ARGUMENT, ONLY for an empty path, which is a
// malformed request. Every other failure throws a plain runtime_error, i.e.
// INTERNAL: the destination is LocalAI's own generated-content directory and
// not anything the caller named, so a full disk or a permission fault there is
// a server fault and is worth retrying, which is the opposite of what
// INVALID_ARGUMENT tells a client.
void write_audio_file(const std::string &path,
const engine::runtime::AudioBuffer &audio);
engine::runtime::AudioBuffer buffer_from_mono(std::vector<float> samples,
int sample_rate);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,217 @@
// Tests for audio_io's reading contract, and in particular for the resampling
// that keeps a 44.1 or 48 kHz upload from reaching a family that only accepts
// 16 kHz.
//
// NAMED _ctest AND NOT _test ON PURPOSE: see the note at the top of
// result_map_ctest.cpp. This file links the audio.cpp engine, so it is built
// and run by ctest, not by backend/cpp/run-unit-tests.sh.
//
// make -C backend/cpp/audio-cpp test-engine
#include "audio_io.h"
#include "loaded_model.h"
#include <algorithm>
#include <cmath>
#include <cstdio>
#include <filesystem>
#include <string>
#include <vector>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
using namespace audiocpp_backend;
// A one-second tone, interleaved across `channels`. Real audio rather than
// silence so a resample that dropped its input would be visible as a flat
// buffer, not just as a different length.
static engine::runtime::AudioBuffer tone(int sample_rate, int channels,
float seconds) {
engine::runtime::AudioBuffer buffer;
buffer.sample_rate = sample_rate;
buffer.channels = channels;
const auto frames =
static_cast<size_t>(static_cast<double>(sample_rate) * seconds);
buffer.samples.reserve(frames * static_cast<size_t>(channels));
for (size_t frame = 0; frame < frames; ++frame) {
const float value = 0.5f * std::sin(2.0f * 3.14159265f * 220.0f *
static_cast<float>(frame) /
static_cast<float>(sample_rate));
for (int channel = 0; channel < channels; ++channel) {
buffer.samples.push_back(value);
}
}
return buffer;
}
static float peak(const std::vector<float> &samples) {
float highest = 0.0f;
for (const float sample : samples) {
highest = std::max(highest, std::abs(sample));
}
return highest;
}
static std::filesystem::path scratch_dir() {
const auto dir = std::filesystem::temp_directory_path() / "audiocpp-io-ctest";
std::filesystem::create_directories(dir);
return dir;
}
// The I2 fixture. Before the resample this returned a 44.1 kHz buffer, which
// silero_vad and sortformer_diar both reject with a plain runtime_error, which
// the server maps to INTERNAL. A 44.1 kHz WAV is an ordinary upload.
static void test_441k_stereo_is_read_as_16k_mono() {
const auto path = scratch_dir() / "input-44100-stereo.wav";
write_audio_file(path.string(), tone(44100, 2, 1.0f));
const auto audio = read_audio_file(path.string(), 16000);
check(audio.sample_rate == 16000, "44.1 kHz input is resampled to 16 kHz");
check(audio.channels == 1, "stereo input is downmixed to mono");
// Linear resampling lands within a sample or two of the exact ratio.
const auto frames = static_cast<long long>(audio.samples.size());
check(frames > 15990 && frames < 16010,
"one second in stays one second out");
check(peak(audio.samples) > 0.2f,
"the resampled buffer still carries the signal");
}
static void test_48k_is_read_as_16k() {
const auto path = scratch_dir() / "input-48000-mono.wav";
write_audio_file(path.string(), tone(48000, 1, 0.5f));
const auto audio = read_audio_file(path.string(), 16000);
check(audio.sample_rate == 16000, "48 kHz input is resampled to 16 kHz");
const auto frames = static_cast<long long>(audio.samples.size());
check(frames > 7990 && frames < 8010, "half a second in, half a second out");
}
// The common case: the upload is already 16 kHz mono, and nothing is resampled.
static void test_16k_mono_passes_through_unchanged() {
const auto path = scratch_dir() / "input-16000-mono.wav";
const auto source = tone(16000, 1, 1.0f);
write_audio_file(path.string(), source);
const auto audio = read_audio_file(path.string(), 16000);
check(audio.sample_rate == 16000, "16 kHz stays 16 kHz");
check(audio.channels == 1, "mono stays mono");
check(audio.samples.size() == source.samples.size(),
"a matching rate resamples nothing");
}
// Rate 0 means "give me the file as it is", which is what a source separation
// route needs: demucs and roformer refuse anything but their own 44.1 kHz and
// work on stereo, so the reader must not force them to 16 kHz mono.
static void test_zero_target_keeps_the_native_format() {
const auto path = scratch_dir() / "input-native.wav";
write_audio_file(path.string(), tone(44100, 2, 0.25f));
const auto audio = read_audio_file(path.string(), 0);
check(audio.sample_rate == 44100, "a zero target keeps the file's rate");
check(audio.channels == 2, "a zero target keeps the file's channels");
}
static void test_missing_file_is_a_config_error() {
bool threw_config_error = false;
try {
read_audio_file((scratch_dir() / "does-not-exist.wav").string(), 16000);
} catch (const ConfigError &) {
threw_config_error = true;
} catch (const std::exception &) {
// Any other type maps to INTERNAL, which is what this asserts against.
}
check(threw_config_error, "a missing input file is INVALID_ARGUMENT, not INTERNAL");
}
static void test_unreadable_file_is_a_config_error() {
const auto path = scratch_dir() / "not-a-wav.wav";
{
FILE *file = fopen(path.string().c_str(), "wb");
if (file != nullptr) {
fputs("this is not a RIFF header", file);
fclose(file);
}
}
bool threw_config_error = false;
try {
read_audio_file(path.string(), 16000);
} catch (const ConfigError &) {
threw_config_error = true;
} catch (const std::exception &) {
}
check(threw_config_error, "a non-WAV input is INVALID_ARGUMENT, not INTERNAL");
}
// The write side of the same distinction. The destination is LocalAI's own
// generated-content directory, not a caller-supplied path, so a failure to
// write it is a server fault: INTERNAL, which a client may retry, and not
// INVALID_ARGUMENT, which tells it the request itself was wrong.
static void test_write_failure_is_not_a_config_error() {
// A regular file where a directory has to be. ENOTDIR defeats root as well
// as an ordinary user, unlike a chmod, which CI running as root would walk
// straight through.
const auto blocker = scratch_dir() / "blocking-file";
{
FILE *file = fopen(blocker.string().c_str(), "wb");
if (file != nullptr) {
fputs("not a directory", file);
fclose(file);
}
}
const auto path = blocker / "nested" / "out.wav";
bool threw_config_error = false;
bool threw_something = false;
try {
write_audio_file(path.string(), tone(16000, 1, 0.05f));
} catch (const ConfigError &) {
threw_config_error = true;
threw_something = true;
} catch (const std::exception &) {
threw_something = true;
}
check(threw_something, "an unwritable destination is reported at all");
check(!threw_config_error,
"a failed write is INTERNAL, not INVALID_ARGUMENT: the caller did not "
"choose the destination and cannot fix it");
check(!std::filesystem::exists(path), "and nothing was written");
}
static void test_empty_output_path_is_a_config_error() {
// The one write failure that IS the caller's: no path at all.
bool threw_config_error = false;
try {
write_audio_file("", tone(16000, 1, 0.05f));
} catch (const ConfigError &) {
threw_config_error = true;
} catch (const std::exception &) {
}
check(threw_config_error, "an empty output path stays INVALID_ARGUMENT");
}
int main() {
test_441k_stereo_is_read_as_16k_mono();
test_48k_is_read_as_16k();
test_16k_mono_passes_through_unchanged();
test_zero_target_keeps_the_native_format();
test_missing_file_is_a_config_error();
test_unreadable_file_is_a_config_error();
test_write_failure_is_not_a_config_error();
test_empty_output_path_is_a_config_error();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all audio_io checks passed\n");
return 0;
}

View File

@@ -0,0 +1,139 @@
#include "audio_units.h"
#include <algorithm>
#include <cmath>
#include <limits>
namespace audiocpp_backend {
std::int64_t interleaved_frame_count(std::size_t sample_count, int channels) {
const std::size_t lanes = channels > 0 ? static_cast<std::size_t>(channels)
: static_cast<std::size_t>(1);
// Truncating division is deliberate: a trailing partial frame is not a
// position every channel reached, so counting it would overstate the length.
return static_cast<std::int64_t>(sample_count / lanes);
}
std::int64_t samples_to_nanoseconds(std::int64_t samples, int sample_rate) {
if (sample_rate <= 0) {
return 0;
}
// Split into whole seconds plus a remainder so the intermediate product
// cannot overflow on long recordings, and so rates like 44100 stay exact.
// The remainder division truncates deliberately: that matches Go's
// time.Duration conventions and keeps successive sample indices monotonic.
const std::int64_t rate = static_cast<std::int64_t>(sample_rate);
const std::int64_t whole_seconds = samples / rate;
const std::int64_t remainder = samples % rate;
return whole_seconds * 1000000000LL + (remainder * 1000000000LL) / rate;
}
float samples_to_seconds(std::int64_t samples, int sample_rate) {
if (sample_rate <= 0) {
return 0.0f;
}
return static_cast<float>(static_cast<double>(samples) /
static_cast<double>(sample_rate));
}
std::int64_t seconds_to_samples(double seconds, int sample_rate) {
// !(seconds > 0.0) rather than seconds <= 0.0: every comparison against NaN
// is false, so the <= form lets NaN reach the cast below, which is undefined
// behaviour and lands on INT64_MIN in practice. This is the one entry point
// fed by untrusted-shaped input (a float-seconds timestamp off the wire, or
// a boundary from a model that diverged), and a hugely negative sample index
// used later as an offset or a length is a wild pointer rather than merely a
// wrong timestamp.
if (sample_rate <= 0 || !(seconds > 0.0)) {
return 0;
}
const double scaled = seconds * static_cast<double>(sample_rate);
// Bound before the cast for the same reason: converting a double at or above
// 2^63 (infinity included) is undefined behaviour, so saturate instead.
const double limit =
static_cast<double>(std::numeric_limits<std::int64_t>::max());
if (scaled >= limit) {
return std::numeric_limits<std::int64_t>::max();
}
// Round rather than truncate: these functions exist to cross the float
// seconds boundary the VAD and diarize messages use, so a value that came
// from samples_to_seconds converts back to the sample it started as.
// Truncation lost one sample about half the time, starting at n=1.
//
// That round trip is exact only below roughly 2^23 samples. Past that the
// float samples_to_seconds returns can no longer resolve adjacent indices
// and the trip fails whatever the rounding. Both the first failing INDEX
// and the duration it stands for depend on the rate, so they are listed per
// rate rather than folded into one range; measured:
//
// 16 kHz 16384001 samples 17.1 min
// 44.1 kHz 11289602 samples 4.3 min
// 48 kHz 12288002 samples 4.3 min
// 96 kHz 12288002 samples 2.1 min
//
// The shortest recording this bites is therefore a couple of minutes of
// 96 kHz audio. It is a property of the float seconds API itself, not of
// the rounding here, and it is why nothing should use these to carry a
// sample-accurate position in a long recording.
return static_cast<std::int64_t>(std::llround(scaled));
}
std::vector<float> s16le_to_f32(const std::string &bytes) {
std::vector<float> samples;
const size_t count = bytes.size() / 2;
samples.reserve(count);
for (size_t i = 0; i < count; ++i) {
const auto low = static_cast<unsigned char>(bytes[i * 2]);
const auto high = static_cast<unsigned char>(bytes[i * 2 + 1]);
const auto raw = static_cast<std::int16_t>(
static_cast<std::uint16_t>(low) |
(static_cast<std::uint16_t>(high) << 8));
// 32768 on decode against 32767 on encode is deliberate, not a typo.
// 32768 is what keeps INT16_MIN at exactly -1.0 and every other code
// inside the [-1, 1] range this header promises; dividing by 32767
// would decode INT16_MIN to -1.00003. See f32_to_s16le for the other
// half of the pair. The cost is that a round trip shrinks a sample by
// 32767/32768, well under one LSB.
samples.push_back(static_cast<float>(raw) / 32768.0f);
}
return samples;
}
std::string f32_to_s16le(const std::vector<float> &samples) {
std::string bytes;
bytes.reserve(samples.size() * 2);
for (const float sample : samples) {
// NaN maps to silence. A NaN sample rendered as a full-scale click is
// worse audio than a dropped one, and this unit converts audio that may
// have originated off the wire.
//
// This guard also removes what used to be a spelling hazard in the
// clamp below. std::min and std::max return their first argument when
// the comparison is false, and every comparison against NaN is false,
// so before this branch existed the choice of spelling silently decided
// whether a NaN reached std::lround, whose result is unspecified for
// NaN. These three leaked it, the last being the idiomatic C++17 way to
// write a clamp and so the likeliest future edit:
// std::min(std::max(sample, -1.0f), 1.0f)
// std::max(std::min(sample, 1.0f), -1.0f)
// std::clamp(sample, -1.0f, 1.0f)
// The order is no longer load-bearing now that the guard runs first,
// but the history is why the guard is here, so do not drop it.
if (std::isnan(sample)) {
bytes.push_back(0);
bytes.push_back(0);
continue;
}
const float clamped = std::max(-1.0f, std::min(1.0f, sample));
// 32767 rather than 32768 so +1.0 saturates at INT16_MAX instead of
// overflowing to INT16_MIN. See s16le_to_f32 for why decode differs.
const auto value =
static_cast<std::int16_t>(std::lround(clamped * 32767.0f));
const auto raw = static_cast<std::uint16_t>(value);
bytes.push_back(static_cast<char>(raw & 0xFF));
bytes.push_back(static_cast<char>((raw >> 8) & 0xFF));
}
return bytes;
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,50 @@
#pragma once
// Time and sample-format conversion between audio.cpp's runtime types (sample
// indices, float PCM) and LocalAI's proto types. Standard library only.
//
// LocalAI uses three different time units:
// TranscriptSegment / TranscriptWord start,end : int64 nanoseconds
// VADSegment start,end : float seconds
// DiarizeSegment start,end : float seconds
#include <cstddef>
#include <cstdint>
#include <string>
#include <vector>
namespace audiocpp_backend {
// Frames in an interleaved buffer of `sample_count` floats laid out across
// `channels` channels. A frame is one per-channel position, which is the unit
// every duration and every span boundary in this backend is expressed in, so a
// stereo buffer must not report twice its real length: feeding sample_count
// straight to samples_to_seconds makes a 3 second stereo clip come back as 6.
//
// A non-positive channel count is treated as mono, matching
// engine::runtime::AudioBuffer's own default of 1 and keeping a reader that
// reports 0 channels from dividing by zero.
std::int64_t interleaved_frame_count(std::size_t sample_count, int channels);
// Returns 0 when sample_rate is not positive rather than dividing by zero.
// Uses integer arithmetic so 44.1 kHz does not lose precision.
std::int64_t samples_to_nanoseconds(std::int64_t samples, int sample_rate);
float samples_to_seconds(std::int64_t samples, int sample_rate);
// Rounds to nearest. Negative seconds and NaN both yield 0, and a value too
// large to convert saturates at INT64_MAX rather than overflowing. Round trips
// with samples_to_seconds only below roughly 2^23 samples, past which the float
// seconds can no longer resolve adjacent sample indices.
std::int64_t seconds_to_samples(double seconds, int sample_rate);
// Decodes little-endian signed 16-bit PCM. A trailing odd byte is dropped.
std::vector<float> s16le_to_f32(const std::string &bytes);
// Encodes to little-endian signed 16-bit PCM, clamping to [-1, 1] first so an
// overshooting sample saturates instead of wrapping to the opposite sign.
// A NaN sample encodes to 0, on the grounds that silence beats a full-scale
// click.
std::string f32_to_s16le(const std::vector<float> &samples);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,226 @@
// Unit tests for audio_units. Standard library only. The harness compiles this
// as a single translation unit, so the implementation is included directly.
#include "audio_units.cpp"
#include <cfenv>
#include <cmath>
#include <cstdio>
#include <limits>
#include <string>
#include <vector>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
static bool close_to(float a, float b, float tol) { return std::fabs(a - b) <= tol; }
using namespace audiocpp_backend;
static void test_nanoseconds() {
// LocalAI TranscriptSegment/TranscriptWord times are nanoseconds
// (Go reads them as time.Duration).
check(samples_to_nanoseconds(16000, 16000) == 1000000000LL, "1s at 16k is 1e9 ns");
check(samples_to_nanoseconds(8000, 16000) == 500000000LL, "0.5s at 16k");
check(samples_to_nanoseconds(0, 16000) == 0, "zero samples is zero ns");
check(samples_to_nanoseconds(1000, 0) == 0, "zero sample rate yields zero, not UB");
// 44.1 kHz must not lose precision to float arithmetic.
check(samples_to_nanoseconds(44100, 44100) == 1000000000LL, "1s at 44.1k");
check(samples_to_nanoseconds(22050, 44100) == 500000000LL, "0.5s at 44.1k");
// The cases above all land on values a float happens to hold exactly, so
// they do not actually rule float arithmetic out. These do:
// a fraction that does not divide evenly, and a duration whose magnitude
// exceeds a float's 24-bit mantissa at nanosecond resolution.
check(samples_to_nanoseconds(44099, 44100) == 999977324LL,
"44.1k fraction is exact, not rounded through a float");
check(samples_to_nanoseconds(44100LL * 3600, 44100) == 3600000000000LL,
"one hour at 44.1k is exact to the nanosecond");
// A naive samples * 1e9 would overflow int64 here; the split into whole
// seconds plus a remainder is what keeps this correct.
check(samples_to_nanoseconds(44100LL * 360000, 44100) == 360000000000000LL,
"100 hours at 44.1k does not overflow");
// Double arithmetic is close enough to pass everything above, but still
// truncates this one a nanosecond short. Integer division does not.
check(samples_to_nanoseconds(4004, 8000) == 500500000LL,
"0.5005s at 8k is exact to the nanosecond");
// Truncation, not rounding: this matches Go's time.Duration conventions and
// keeps successive sample indices monotonic. The exact value here is
// 22675.7...; rounding to nearest would give 22676.
check(samples_to_nanoseconds(1, 44100) == 22675LL,
"a sub-nanosecond fraction truncates rather than rounding up");
}
static void test_seconds() {
check(close_to(samples_to_seconds(24000, 24000), 1.0f, 1e-6f), "1s at 24k");
check(close_to(samples_to_seconds(12000, 24000), 0.5f, 1e-6f), "0.5s at 24k");
check(close_to(samples_to_seconds(100, 0), 0.0f, 1e-6f), "zero sample rate is 0s");
check(seconds_to_samples(1.0, 16000) == 16000, "1s to samples at 16k");
check(seconds_to_samples(0.5, 16000) == 8000, "0.5s to samples at 16k");
check(seconds_to_samples(1.0, 0) == 0, "zero sample rate yields zero samples");
check(seconds_to_samples(-1.0, 16000) == 0, "negative seconds clamps to zero");
// seconds_to_samples is the one entry point fed by untrusted-shaped input:
// a float-seconds timestamp off the wire, or a VAD boundary from a model
// that diverged. A hugely negative sample index used later as an offset or
// a length is a wild pointer, not merely a wrong timestamp.
const double nan_seconds = std::numeric_limits<double>::quiet_NaN();
const double inf_seconds = std::numeric_limits<double>::infinity();
const std::int64_t max_samples = std::numeric_limits<std::int64_t>::max();
check(seconds_to_samples(nan_seconds, 16000) == 0, "NaN seconds yields zero");
check(seconds_to_samples(inf_seconds, 16000) == max_samples,
"infinite seconds saturates instead of overflowing");
check(seconds_to_samples(1e30, 16000) == max_samples,
"out of range seconds saturates instead of overflowing");
check(seconds_to_samples(-inf_seconds, 16000) == 0,
"negative infinity clamps to zero");
// Crossing the float-seconds boundary and back is the expected round trip
// for the VAD and diarize messages, so it must not lose a sample.
// Truncation loses one about half the time, starting at n=1.
check(seconds_to_samples(samples_to_seconds(1, 44100), 44100) == 1,
"one sample survives the seconds round trip at 44.1k");
check(seconds_to_samples(samples_to_seconds(1, 16000), 16000) == 1,
"one sample survives the seconds round trip at 16k");
check(seconds_to_samples(samples_to_seconds(4001, 8000), 8000) == 4001,
"4001 samples survive the seconds round trip at 8k");
}
static void test_s16le_round_trip() {
const std::vector<float> original = {0.0f, 0.5f, -0.5f, 1.0f, -1.0f};
const std::string encoded = f32_to_s16le(original);
check(encoded.size() == original.size() * 2, "two bytes per sample");
const std::vector<float> decoded = s16le_to_f32(encoded);
check(decoded.size() == original.size(), "round trip keeps the sample count");
for (size_t i = 0; i < original.size(); ++i) {
// 16-bit quantisation: one LSB is ~3.05e-5. Guard the index so a short
// result reports a named failure instead of aborting the whole suite.
check(i < decoded.size() && close_to(decoded[i], original[i], 1e-4f),
"round trip preserves sample " + std::to_string(i));
}
}
static void test_s16le_endianness() {
// 0.5 encodes to 16384 = 0x4000, little endian is 0x00 0x40.
const std::string encoded = f32_to_s16le({0.5f});
check(encoded.size() == 2, "one sample is two bytes");
check(static_cast<unsigned char>(encoded[0]) == 0x00, "low byte first");
check(static_cast<unsigned char>(encoded[1]) == 0x40, "high byte second");
}
static void test_s16le_clamping() {
// Values outside [-1, 1] must clamp, not wrap around to the opposite sign.
const std::string encoded = f32_to_s16le({2.0f, -2.0f});
const std::vector<float> decoded = s16le_to_f32(encoded);
check(decoded.size() == 2, "two samples survive clamping");
check(decoded.size() > 0 && decoded[0] > 0.99f,
"positive overshoot clamps to full scale");
check(decoded.size() > 1 && decoded[1] < -0.99f,
"negative overshoot clamps to full scale");
}
static void test_s16le_decode_range() {
// INT16_MIN is the one value that pins the decode scale. Dividing by 32767
// instead of 32768 would decode it to -1.00003, outside the [-1, 1] range
// the header promises, and every other test would still pass.
const std::vector<float> decoded = s16le_to_f32(std::string("\x00\x80", 2));
check(decoded.size() == 1, "INT16_MIN decodes to one sample");
check(decoded.size() == 1 && decoded[0] == -1.0f,
"INT16_MIN decodes to exactly -1.0, not past full scale");
}
static void test_s16le_nan_input() {
// A NaN sample must not reach std::lround, whose result is unspecified for
// NaN. Asserting a range is not enough to pin this: the three outcomes the
// plausible clamp spellings produce (full scale, negative full scale, zero)
// are all finite and all inside [-1, 1], so a range check passes for every
// one of them. Only an exact value distinguishes them.
// NaN maps to silence, not to full scale: a NaN sample rendered as a
// full-scale click is worse audio than a dropped one, and this unit
// converts audio that may have originated off the wire.
//
// volatile so the NaN cannot be constant-folded, which would let the
// compiler evaluate the conversion at compile time and raise no
// floating-point exception at run time for the check below to observe.
volatile float nan_source = std::numeric_limits<float>::quiet_NaN();
const std::vector<float> input = {nan_source};
std::feclearexcept(FE_ALL_EXCEPT);
const std::string encoded = f32_to_s16le(input);
const bool raised_invalid = std::fetestexcept(FE_INVALID) != 0;
const std::vector<float> decoded = s16le_to_f32(encoded);
check(decoded.size() == 1, "a NaN sample still encodes to one sample");
check(decoded.size() == 1 && decoded[0] == 0.0f,
"a NaN sample encodes to exactly zero, not to a full-scale click");
// Independent of the value: a quiet NaN raises invalid-operation as soon as
// it reaches any ordered comparison, which is what std::min and std::max
// use, so this fails unless the NaN is diverted before the clamp runs at
// all. That is what stops the explicit guard from being dropped in favour
// of a clamp spelling that happens to yield zero.
check(!raised_invalid,
"encoding a NaN sample raises no invalid-operation exception");
}
static void test_s16le_odd_length() {
// A truncated frame must drop the dangling byte rather than read past it.
const std::string odd(5, '\0');
check(s16le_to_f32(odd).size() == 2, "odd byte count drops the trailing byte");
check(s16le_to_f32(std::string()).empty(), "empty input yields no samples");
}
static void test_interleaved_frame_count() {
// Mono is a pass-through, which is the only case the VAD path exercises.
check(interleaved_frame_count(16000, 1) == 16000, "mono frames equal samples");
// The case that matters: a stereo buffer holds two floats per position, so a
// one second 16 kHz stereo clip is 32000 floats and still one second. Handing
// the raw float count to samples_to_seconds reports two seconds instead.
check(interleaved_frame_count(32000, 2) == 16000,
"stereo frames are half the samples");
check(samples_to_seconds(interleaved_frame_count(32000, 2), 16000) == 1.0f,
"a one second stereo clip measures one second, not two");
check(interleaved_frame_count(48000, 3) == 16000,
"three channels divide by three");
// engine::runtime::AudioBuffer defaults channels to 1, but a reader is free
// to report 0, and dividing by that is undefined rather than merely wrong.
check(interleaved_frame_count(1000, 0) == 1000,
"zero channels is treated as mono");
check(interleaved_frame_count(1000, -2) == 1000,
"a negative channel count is treated as mono");
// A dangling partial frame is not a position every channel reached.
check(interleaved_frame_count(3, 2) == 1,
"a trailing partial frame is not counted");
check(interleaved_frame_count(0, 2) == 0, "an empty buffer has no frames");
// Past 2^32 floats, so a size_t narrowed to 32 bits on the way in, or a
// signed 32-bit intermediate, shows up here rather than in a multi-hour
// recording nobody tests with.
check(interleaved_frame_count(static_cast<std::size_t>(9000000000ULL), 2) ==
4500000000LL,
"a buffer beyond 2^32 floats counts frames without truncating");
}
int main() {
test_interleaved_frame_count();
test_nanoseconds();
test_seconds();
test_s16le_round_trip();
test_s16le_endianness();
test_s16le_clamping();
test_s16le_decode_range();
test_s16le_nan_input();
test_s16le_odd_length();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all audio_units checks passed\n");
return 0;
}

View File

@@ -0,0 +1,411 @@
#include "capability_routing.h"
#include <algorithm>
namespace audiocpp_backend {
namespace {
struct NamedTask {
Task task;
const char *name;
};
// Short names are exactly the strings audio.cpp prints and parses in
// framework/runtime/session.cpp, so a name pinned here survives conversion at
// the engine boundary and a name copied out of audio.cpp is accepted here. All
// thirteen have an upstream name; only "spk" is absent from the --task table in
// docs/usage.md.
const NamedTask kTaskNames[] = {
{Task::Vad, "vad"},
{Task::Asr, "asr"},
{Task::Diarization, "diar"},
{Task::SourceSeparation, "sep"},
{Task::AudioGeneration, "gen"},
{Task::Tts, "tts"},
{Task::VoiceCloning, "clon"},
{Task::VoiceConversion, "vc"},
{Task::SpeechToSpeech, "s2s"},
{Task::Alignment, "align"},
{Task::VoiceDesign, "vdes"},
{Task::SpeakerRecognition, "spk"},
{Task::Svc, "svc"},
};
// Accepted on input but never emitted. "spkrec" was this backend's own earlier
// name for the kind; upstream only ever knew "spk".
const NamedTask kTaskAliases[] = {
{Task::SpeakerRecognition, "spkrec"},
};
// First match wins, which is safe because a Capabilities value holds at most one
// entry per task: it mirrors upstream runtime::TaskCapability (model.h), which
// pairs one kind with a modes vector, and no loader's supported_tasks list
// repeats a kind.
bool family_supports(const Capabilities &caps, Task task, Mode mode) {
for (const auto &capability : caps.tasks) {
if (capability.task != task) {
continue;
}
return std::find(capability.modes.begin(), capability.modes.end(), mode) !=
capability.modes.end();
}
return false;
}
// Mode preference per RPC. Only AudioTranscriptionStream has a fallback: a
// server-streaming transcription can be satisfied by an offline run that emits
// one delta then the final result. Live transcription cannot, because it is
// bidirectional and must consume audio incrementally.
std::vector<Mode> mode_candidates(Rpc rpc) {
switch (rpc) {
case Rpc::TtsStream:
case Rpc::AudioTranscriptionLive:
return {Mode::Streaming};
case Rpc::AudioTranscriptionStream:
return {Mode::Streaming, Mode::Offline};
default:
return {Mode::Offline};
}
}
std::vector<Task> task_candidates(Rpc rpc, const RequestShape &shape) {
switch (rpc) {
case Rpc::Tts:
case Rpc::TtsStream:
// A supplied speaker clip is the strongest signal: the caller named the
// voice they want. Free-form instructions come next. Both fall back to
// plain Tts so a family without the specialised task still answers.
if (shape.has_voice_reference) {
return {Task::VoiceCloning, Task::Tts, Task::VoiceDesign};
}
if (shape.has_instructions) {
return {Task::VoiceDesign, Task::Tts, Task::VoiceCloning};
}
return {Task::Tts, Task::VoiceCloning, Task::VoiceDesign};
case Rpc::AudioTranscription:
case Rpc::AudioTranscriptionStream:
case Rpc::AudioTranscriptionLive:
// Asr first: `prompt` is also whisper-style decoding context, so its
// presence must not hijack a real ASR family into forced alignment.
if (shape.has_prompt_text) {
return {Task::Asr, Task::Alignment};
}
return {Task::Asr};
case Rpc::Vad:
return {Task::Vad};
case Rpc::Diarize:
return {Task::Diarization};
case Rpc::SoundGeneration:
return {Task::AudioGeneration};
case Rpc::AudioTransform:
// Svc is listed for completeness but is unreachable by auto-routing, by
// design: the only families advertising it (seed_vc, vevo2) also
// advertise VoiceConversion, which always wins, and no request signal
// means "this input is singing". Singing voice conversion therefore
// requires an explicit task:svc pin.
return {Task::SourceSeparation, Task::VoiceConversion, Task::Svc,
Task::SpeechToSpeech};
}
return {};
}
// The reasons behind unsupported_surfaces(), spelled once because AudioEncode
// and AudioDecode share theirs. Each is phrased in terms of what upstream does
// and does not have, so a reader can check it against the pinned checkout
// rather than take it on trust. Every one of them was checked against
// audio.cpp e800d435d130dc776baf6f3e6129bb62b1495c89, and one of the four
// claims this backend was planned against did not survive that check: see
// kTransformStreamReason.
//
// A latent upstream inconsistency worth knowing about but deliberately NOT put
// on the wire, because it would mislead: model_spec/schema.cpp's task-string
// whitelist does accept "codec" (and "dialogue"), while
// model_spec/metadata.cpp's parse_task_kind has no branch for either and
// throws "unknown model spec task". So a spec declaring "codec" validates and
// then fails to load. That is a hole in upstream's own validation, not a codec
// task this backend could reach.
const char *const kCodecReason =
"audio.cpp's VoiceTaskKind has no codec entry, so no family can be asked to "
"turn PCM into codec frames or back; miocodec carries a Codec tag in "
"upstream's README but its loader advertises only vc and s2s";
// NOT "streaming exists for tts and asr only", which is what this backend was
// planned to say and is false: silero_vad advertises vad with RunMode::Streaming
// (src/models/silero_vad/session.cpp). The claim that actually holds is the
// narrower one below, about the four tasks AudioTransform routes to.
// The trailing clause is not padding. The premise is an absence, and an absence
// does not on its own make the RPC impossible: an offline sep family could be
// buffered and emitted as a stream, which is what several LocalAI backends do.
// Stopping at "nothing advertises streaming" would imply an impossibility the
// evidence does not support. What is true, and what the caller needs, is that
// this backend declines to dress an offline call up as a streaming one.
const char *const kTransformStreamReason =
"no audio.cpp family advertises streaming for any task AudioTransform routes "
"to (sep, vc, svc, s2s); upstream advertises RunMode::Streaming for tts, asr "
"and vad only, and no conversion or separation family even implements its "
"IStreamingVoiceTaskSession interface, so a streaming transform here would be "
"a buffered offline call in disguise, which this backend does not pretend to "
"offer";
// "clip-to-clip processing against a target voice", NOT "voice conversion". The
// latter is true of miocodec and FALSE of vevo2, whose s2s route is `editing`
// and only `editing`: src/models/vevo2/session.cpp's default_route_for_task maps
// SpeechToSpeech to Editing and route_matches_task accepts nothing else, and
// docs/models/vevo2.md defines that route as "Edit source speech into new target
// text while using the target voice", requiring --target-text. It rewrites what
// was said. vevo2's actual voice conversion is its separate vc task, which is
// why upstream's README tags the family "TTS, Music, VC, Edit". The conclusion
// is unaffected: neither family converses.
const char *const kAudioToAudioReason =
"LocalAI's contract here is OpenAI-Realtime shaped, an audio conversation "
"emitting audio, transcript and tool-call deltas from a system prompt and a "
"tool list; audio.cpp's s2s is offline clip-to-clip processing against a "
"target voice, declared only by miocodec (voice conversion) and vevo2 "
"(speech editing), with no conversation, system prompt or tool loop";
const char *const kVoiceEmbedReason =
"no audio.cpp family advertises the spk (SpeakerRecognition) task, so "
"nothing in the engine can produce a speaker embedding; the task kind "
"itself exists upstream, and TitaNet and ECAPA-TDNN exist as internal "
"conditioning encoders, but neither is registered as a loadable family";
// The tasks an RPC is ever willing to route to, independent of request shape.
//
// DERIVED from task_candidates rather than restated, so a task added to an
// RPC's candidate list cannot become inadmissible as a pin by omission. Setting
// every shape flag yields each RPC's widest list: the per-flag branches only
// reorder the same three tasks for Tts, and only ADD Alignment for
// transcription, so the union is what comes back.
std::vector<Task> admissible_tasks(Rpc rpc) {
RequestShape widest;
widest.has_voice_reference = true;
widest.has_instructions = true;
widest.has_prompt_text = true;
return task_candidates(rpc, widest);
}
std::string join_task_names(const std::vector<Task> &tasks) {
std::string out;
for (const Task task : tasks) {
if (!out.empty()) {
out += ", ";
}
out += task_name(task);
}
if (out.empty()) {
out = "nothing";
}
return out;
}
std::string join_attempts(const std::vector<Task> &tasks,
const std::vector<Mode> &modes) {
std::string out;
for (const Task task : tasks) {
for (const Mode mode : modes) {
if (!out.empty()) {
out += ", ";
}
out += task_name(task);
out += "/";
out += mode_name(mode);
}
}
return out;
}
} // namespace
const char *task_name(Task task) {
for (const auto &entry : kTaskNames) {
if (entry.task == task) {
return entry.name;
}
}
return "unknown";
}
const char *mode_name(Mode mode) {
return mode == Mode::Streaming ? "streaming" : "offline";
}
const char *rpc_name(Rpc rpc) {
switch (rpc) {
case Rpc::Tts:
return "TTS";
case Rpc::TtsStream:
return "TTSStream";
case Rpc::AudioTranscription:
return "AudioTranscription";
case Rpc::AudioTranscriptionStream:
return "AudioTranscriptionStream";
case Rpc::AudioTranscriptionLive:
return "AudioTranscriptionLive";
case Rpc::Vad:
return "VAD";
case Rpc::Diarize:
return "Diarize";
case Rpc::SoundGeneration:
return "SoundGeneration";
case Rpc::AudioTransform:
return "AudioTransform";
}
return "unknown";
}
bool parse_task_name(const std::string &value, Task &out) {
for (const auto &entry : kTaskNames) {
if (value == entry.name) {
out = entry.task;
return true;
}
}
for (const auto &entry : kTaskAliases) {
if (value == entry.name) {
out = entry.task;
return true;
}
}
return false;
}
std::string describe_capabilities(const Capabilities &caps) {
std::string out;
for (const auto &capability : caps.tasks) {
for (const Mode mode : capability.modes) {
if (!out.empty()) {
out += ", ";
}
out += task_name(capability.task);
out += "/";
out += mode_name(mode);
}
}
if (out.empty()) {
out = "nothing";
}
return out;
}
const std::vector<UnsupportedSurface> &unsupported_surfaces() {
// Ordered as UnsupportedRpc declares them. unsupported_surface() names each
// index in a switch rather than casting the enum, so the order is checked at
// compile time rather than trusted.
static const std::vector<UnsupportedSurface> kSurfaces = {
{"AudioEncode", kCodecReason},
{"AudioDecode", kCodecReason},
{"AudioTransformStream", kTransformStreamReason},
{"AudioToAudioStream", kAudioToAudioReason},
{"VoiceEmbed", kVoiceEmbedReason},
};
return kSurfaces;
}
// A switch with NO default label, deliberately. -Wswitch is on under -Wall, so a
// sixth UnsupportedRpc added without a case here is a BUILD diagnostic, which is
// the only place this class of mistake can be caught for free: a positional
// static_cast<size_t>(rpc) would compile fine and read past the end of the table
// at run time, on the one code path whose entire job is to be diagnosable. The
// table stays a table because the tests iterate it.
//
// The trailing return is unreachable through the enum and exists only for a
// caller that hands over a value outside it, which is already undefined
// behaviour by the time it arrives.
const UnsupportedSurface &unsupported_surface(UnsupportedRpc rpc) {
const std::vector<UnsupportedSurface> &surfaces = unsupported_surfaces();
switch (rpc) {
case UnsupportedRpc::AudioEncode:
return surfaces[0];
case UnsupportedRpc::AudioDecode:
return surfaces[1];
case UnsupportedRpc::AudioTransformStream:
return surfaces[2];
case UnsupportedRpc::AudioToAudioStream:
return surfaces[3];
case UnsupportedRpc::VoiceEmbed:
return surfaces[4];
}
return surfaces[0];
}
std::string unsupported_surface_message(const Capabilities &caps, const char *rpc,
const char *reason) {
return std::string("audio-cpp: the ") + rpc +
" RPC is not available through this backend because " + reason +
". Loaded family '" + caps.family +
"' supports: " + describe_capabilities(caps);
}
std::string unsupported_surface_message(const char *rpc, const char *reason) {
return std::string("audio-cpp: the ") + rpc +
" RPC is not available through this backend because " + reason +
". No model is loaded, so there is no family to list; loading one "
"would not change this answer";
}
Route resolve_route(Rpc rpc, const RequestShape &shape,
const Capabilities &caps) {
Route route;
std::vector<Task> tasks;
if (!shape.pinned_task.empty()) {
Task pinned = Task::Tts;
if (!parse_task_name(shape.pinned_task, pinned)) {
route.error = "audio-cpp: unknown task option '" + shape.pinned_task +
"'. Known tasks: gen, tts, clon, vc, svc, s2s, asr, "
"align, vad, diar, sep, vdes, spk";
return route;
}
// A pin is honoured exactly, but ONLY on an RPC that could have routed
// to it anyway. It used to replace the candidate list wholesale for
// every RPC, and because the model's `task:` option is copied into the
// shape by all nine handlers, one pin bled across all nine surfaces and
// produced wrong 200s rather than errors: nemotron with task:asr made
// Vad return 200 with zero segments after a full ASR decode, so 14
// seconds of speech was reported as silence, and silero_vad with
// task:vad made AudioTranscription return 200 with empty text and four
// segments whose spans were VAD segments, which the srt/vtt/lrc writers
// then rendered as a well formed subtitle file of four timed EMPTY
// cues. Refusing is what the docs already promise: "if the family
// cannot serve it, the request is refused rather than rerouted".
//
// Every legitimate pin survives, because a pin only ever names the task
// its own RPC already routes to: svc is in AudioTransform's candidates,
// tts/clon/vdes in TTS's, asr in transcription's, vad and diar in
// theirs.
const std::vector<Task> admissible = admissible_tasks(rpc);
if (std::find(admissible.begin(), admissible.end(), pinned) ==
admissible.end()) {
route.error = std::string("audio-cpp: this model pins task '") +
task_name(pinned) + "', which the " + rpc_name(rpc) +
" RPC never routes to (it routes to " +
join_task_names(admissible) +
"). Remove the task option to reach this RPC, or call "
"the RPC the pinned task serves";
return route;
}
tasks = {pinned};
} else {
tasks = task_candidates(rpc, shape);
}
const std::vector<Mode> modes = mode_candidates(rpc);
// Task-major: prefer the right task in a fallback mode over the wrong task
// in the preferred mode.
for (const Task task : tasks) {
for (const Mode mode : modes) {
if (family_supports(caps, task, mode)) {
route.ok = true;
route.task = task;
route.mode = mode;
return route;
}
}
}
route.error = std::string("audio-cpp: family '") + caps.family +
"' cannot serve the " + rpc_name(rpc) + " RPC (tried " +
join_attempts(tasks, modes) + "); it supports: " +
describe_capabilities(caps);
return route;
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,131 @@
#pragma once
// Decides which audio.cpp (task, mode) pair serves a given LocalAI RPC, or
// produces the capability error when none can. Standard library only, so this
// unit is tested without an audio.cpp checkout; loaded_model.cpp converts
// to and from engine::runtime types at the boundary.
#include <string>
#include <vector>
namespace audiocpp_backend {
// Mirrors engine::runtime::VoiceTaskKind, same members and same order.
enum class Task {
Vad,
Asr,
Diarization,
SourceSeparation,
AudioGeneration,
Tts,
VoiceCloning,
VoiceConversion,
SpeechToSpeech,
Alignment,
VoiceDesign,
SpeakerRecognition,
Svc,
};
// Mirrors engine::runtime::RunMode.
enum class Mode { Offline, Streaming };
struct TaskCapability {
Task task = Task::Vad;
std::vector<Mode> modes;
};
struct Capabilities {
std::string family;
std::vector<TaskCapability> tasks;
};
// The LocalAI RPCs this backend serves. The ones it cannot serve at all are in
// UnsupportedRpc below rather than here: they never reach routing, because no
// family could satisfy them.
enum class Rpc {
Tts,
TtsStream,
AudioTranscription,
AudioTranscriptionStream,
AudioTranscriptionLive,
Vad,
Diarize,
SoundGeneration,
AudioTransform,
};
struct RequestShape {
// A speaker reference clip was supplied (TTSRequest.voice resolved to audio).
bool has_voice_reference = false;
// TTSRequest.instructions is set.
bool has_instructions = false;
// TranscriptRequest.prompt is set.
bool has_prompt_text = false;
// The model's `task:` option, empty when unset. Overrides routing.
std::string pinned_task;
};
struct Route {
bool ok = false;
Task task = Task::Tts;
Mode mode = Mode::Offline;
// Set when ok is false. Suitable verbatim as an UNIMPLEMENTED message.
std::string error;
};
Route resolve_route(Rpc rpc, const RequestShape &shape, const Capabilities &caps);
// Canonical audio.cpp short names: gen, tts, clon, vc, svc, s2s, asr, align,
// vad, diar, sep, vdes, spk. parse_task_name additionally accepts "spkrec" as
// a legacy alias; task_name only ever emits "spk".
const char *task_name(Task task);
const char *mode_name(Mode mode);
const char *rpc_name(Rpc rpc);
bool parse_task_name(const std::string &value, Task &out);
// "asr/offline, asr/streaming", for error messages.
std::string describe_capabilities(const Capabilities &caps);
// The RPCs in LocalAI's backend contract that audio.cpp has no counterpart for,
// as opposed to the ones in Rpc above, which a particular family may or may not
// be able to serve. Nothing routes to these: the refusal is a property of the
// engine, not of the loaded model, so loading a different family cannot change
// it.
//
// This is deliberately NOT deferred work. Each entry names the upstream
// limitation that keeps it out of Rpc, and each becomes an ordinary routing
// entry the day upstream lifts that limitation.
enum class UnsupportedRpc {
AudioEncode,
AudioDecode,
AudioTransformStream,
AudioToAudioStream,
VoiceEmbed,
};
struct UnsupportedSurface {
// The RPC's name as backend.proto spells it, for the message.
const char *rpc;
// Why audio.cpp cannot serve it, in terms of what upstream does and does
// not have. Stated so a caller can tell "not built yet" from "not possible".
const char *reason;
};
// The table behind the five refusals. Exposed whole so a test can assert every
// entry rather than the one somebody remembered to cover, and so the reasons
// are data in one place instead of string literals hand-copied into handlers.
const std::vector<UnsupportedSurface> &unsupported_surfaces();
const UnsupportedSurface &unsupported_surface(UnsupportedRpc rpc);
// Message for an RPC this backend cannot serve at all, as opposed to one this
// particular family cannot serve. `reason` states the upstream limitation.
std::string unsupported_surface_message(const Capabilities &caps, const char *rpc,
const char *reason);
// Same, for the no-model-loaded case. Says so explicitly rather than naming an
// empty family, and says that loading one would not help, because the caller's
// obvious next move otherwise is to load a model and try again.
std::string unsupported_surface_message(const char *rpc, const char *reason);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,600 @@
// Unit tests for capability_routing. Standard library only. The harness
// compiles this as a single translation unit, so the implementation is
// included directly rather than linked.
#include "capability_routing.cpp"
#include <cstdio>
#include <string>
#include <vector>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
using namespace audiocpp_backend;
// Mirrors what supertonic advertises: TTS offline and streaming.
static Capabilities supertonic() {
return Capabilities{"supertonic",
{{Task::Tts, {Mode::Offline, Mode::Streaming}}}};
}
// Mirrors chatterbox: TTS, cloning and voice conversion, offline only.
static Capabilities chatterbox() {
return Capabilities{"chatterbox",
{{Task::Tts, {Mode::Offline}},
{Task::VoiceCloning, {Mode::Offline}},
{Task::VoiceConversion, {Mode::Offline}}}};
}
// Mirrors nemotron_asr: ASR offline and streaming.
static Capabilities nemotron() {
return Capabilities{"nemotron_asr",
{{Task::Asr, {Mode::Offline, Mode::Streaming}}}};
}
// Mirrors qwen3_asr: ASR offline only.
static Capabilities qwen3_asr() {
return Capabilities{"qwen3_asr", {{Task::Asr, {Mode::Offline}}}};
}
// Mirrors qwen3_forced_aligner: alignment only.
static Capabilities aligner() {
return Capabilities{"qwen3_forced_aligner",
{{Task::Alignment, {Mode::Offline}}}};
}
// Mirrors htdemucs: separation only.
static Capabilities htdemucs() {
return Capabilities{"htdemucs", {{Task::SourceSeparation, {Mode::Offline}}}};
}
static void test_plain_tts() {
const auto r = resolve_route(Rpc::Tts, RequestShape{}, chatterbox());
check(r.ok, "plain TTS routes");
check(r.task == Task::Tts, "plain TTS picks Tts, not VoiceCloning");
check(r.mode == Mode::Offline, "TTS runs offline");
}
static void test_tts_with_voice_reference_prefers_cloning() {
RequestShape shape;
shape.has_voice_reference = true;
const auto r = resolve_route(Rpc::Tts, shape, chatterbox());
check(r.ok, "TTS with a voice reference routes");
check(r.task == Task::VoiceCloning, "voice reference prefers VoiceCloning");
}
// supertonic has no VoiceCloning: a voice reference must fall back to Tts
// rather than failing the request.
static void test_tts_voice_reference_falls_back_to_tts() {
RequestShape shape;
shape.has_voice_reference = true;
const auto r = resolve_route(Rpc::Tts, shape, supertonic());
check(r.ok, "voice reference on a clone-less family still routes");
check(r.task == Task::Tts, "falls back to Tts");
}
static void test_tts_instructions_prefer_voice_design() {
RequestShape shape;
shape.has_instructions = true;
Capabilities caps{"qwen3_tts",
{{Task::Tts, {Mode::Offline}},
{Task::VoiceDesign, {Mode::Offline}}}};
const auto r = resolve_route(Rpc::Tts, shape, caps);
check(r.ok, "TTS with instructions routes");
check(r.task == Task::VoiceDesign, "instructions prefer VoiceDesign");
}
// A voice reference is a stronger signal than free-form instructions: cloning
// a specific voice is what the user asked for.
static void test_voice_reference_beats_instructions() {
RequestShape shape;
shape.has_voice_reference = true;
shape.has_instructions = true;
Capabilities caps{"omnivoice",
{{Task::Tts, {Mode::Offline}},
{Task::VoiceCloning, {Mode::Offline}},
{Task::VoiceDesign, {Mode::Offline}}}};
const auto r = resolve_route(Rpc::Tts, shape, caps);
check(r.ok, "both signals present routes");
check(r.task == Task::VoiceCloning, "voice reference outranks instructions");
}
static void test_tts_stream_requires_streaming() {
const auto ok = resolve_route(Rpc::TtsStream, RequestShape{}, supertonic());
check(ok.ok, "streaming TTS routes on supertonic");
check(ok.mode == Mode::Streaming, "TTSStream runs in streaming mode");
const auto bad = resolve_route(Rpc::TtsStream, RequestShape{}, chatterbox());
check(!bad.ok, "streaming TTS is refused on an offline-only family");
check(bad.error.find("chatterbox") != std::string::npos,
"error names the family");
check(bad.error.find("tts/offline") != std::string::npos,
"error lists what the family does support");
check(bad.error.find("TTSStream") != std::string::npos,
"error names the RPC that was refused");
check(bad.error.find("tts/streaming") != std::string::npos,
"error lists the (task, mode) pairs that were tried");
}
static void test_transcription_stream_falls_back_to_offline() {
const auto streaming =
resolve_route(Rpc::AudioTranscriptionStream, RequestShape{}, nemotron());
check(streaming.ok && streaming.mode == Mode::Streaming,
"streaming ASR uses streaming mode when offered");
const auto offline =
resolve_route(Rpc::AudioTranscriptionStream, RequestShape{}, qwen3_asr());
check(offline.ok, "streaming ASR falls back on an offline-only family");
check(offline.mode == Mode::Offline, "fallback mode is offline");
check(offline.task == Task::Asr, "fallback task is still Asr");
}
// Task preference dominates mode preference: it is better to run the right
// task in a fallback mode than the wrong task in the preferred mode. This is
// the only RPC where the two orderings can disagree, because it is the only
// one with more than one acceptable mode.
static void test_task_preference_beats_mode_preference() {
RequestShape shape;
shape.has_prompt_text = true;
Capabilities mixed{"mixed_asr_aligner",
{{Task::Asr, {Mode::Offline}},
{Task::Alignment, {Mode::Streaming}}}};
const auto r = resolve_route(Rpc::AudioTranscriptionStream, shape, mixed);
check(r.ok, "mixed family routes");
check(r.task == Task::Asr,
"the preferred task wins even in its fallback mode");
check(r.mode == Mode::Offline,
"the fallback mode is accepted to keep the preferred task");
}
// Live transcription is bidirectional and cannot be faked from an offline run.
static void test_live_transcription_has_no_offline_fallback() {
const auto r =
resolve_route(Rpc::AudioTranscriptionLive, RequestShape{}, qwen3_asr());
check(!r.ok, "live transcription is refused on an offline-only family");
}
static void test_alignment_needs_prompt_text() {
const auto without =
resolve_route(Rpc::AudioTranscription, RequestShape{}, aligner());
check(!without.ok, "aligner without a transcript is refused");
RequestShape shape;
shape.has_prompt_text = true;
const auto with = resolve_route(Rpc::AudioTranscription, shape, aligner());
check(with.ok, "aligner with a transcript routes");
check(with.task == Task::Alignment, "routes to Alignment");
}
// A real ASR family must not be hijacked to Alignment just because the caller
// passed a prompt: `prompt` is also whisper-style decoding context.
static void test_prompt_does_not_hijack_asr() {
RequestShape shape;
shape.has_prompt_text = true;
const auto r = resolve_route(Rpc::AudioTranscription, shape, nemotron());
check(r.ok, "ASR with a prompt routes");
// nemotron advertises Asr alone, so the assertion has to be made against a
// family that advertises both: otherwise "Asr is preferred" only restates
// that Asr is the only option, and reversing the preference order passes.
Capabilities both{"asr_with_aligner",
{{Task::Asr, {Mode::Offline}},
{Task::Alignment, {Mode::Offline}}}};
const auto pref = resolve_route(Rpc::AudioTranscription, shape, both);
check(pref.ok, "a family offering both routes");
check(pref.task == Task::Asr, "Asr is preferred over Alignment");
}
static void test_audio_transform_prefers_separation() {
const auto sep =
resolve_route(Rpc::AudioTransform, RequestShape{}, htdemucs());
check(sep.ok && sep.task == Task::SourceSeparation, "separation routes");
// htdemucs advertises separation alone, so the check above cannot fail on
// ordering. This family advertises both, which is what pins the preference.
Capabilities sep_and_vc{"sep_and_vc",
{{Task::SourceSeparation, {Mode::Offline}},
{Task::VoiceConversion, {Mode::Offline}}}};
const auto pref =
resolve_route(Rpc::AudioTransform, RequestShape{}, sep_and_vc);
check(pref.ok && pref.task == Task::SourceSeparation,
"separation is preferred over voice conversion");
Capabilities miocodec{"miocodec",
{{Task::VoiceConversion, {Mode::Offline}},
{Task::SpeechToSpeech, {Mode::Offline}}}};
const auto vc = resolve_route(Rpc::AudioTransform, RequestShape{}, miocodec);
check(vc.ok && vc.task == Task::VoiceConversion,
"voice conversion is preferred over speech-to-speech");
}
static void test_pinned_task_overrides_routing() {
RequestShape shape;
shape.pinned_task = "s2s";
Capabilities miocodec{"miocodec",
{{Task::VoiceConversion, {Mode::Offline}},
{Task::SpeechToSpeech, {Mode::Offline}}}};
const auto r = resolve_route(Rpc::AudioTransform, shape, miocodec);
check(r.ok && r.task == Task::SpeechToSpeech, "pinned task wins");
RequestShape bad;
bad.pinned_task = "not-a-task";
const auto e = resolve_route(Rpc::AudioTransform, bad, miocodec);
check(!e.ok, "an unknown pinned task is an error");
check(e.error.find("not-a-task") != std::string::npos,
"error names the bad task");
// A pinned task the family does not offer must fail, not silently reroute.
RequestShape unsupported;
unsupported.pinned_task = "sep";
const auto u = resolve_route(Rpc::AudioTransform, unsupported, miocodec);
check(!u.ok, "a pinned but unsupported task is refused");
}
// A pin lives on the MODEL, and every one of the nine handlers copies it into
// the shape, so a pin set for one RPC arrives at all of them. It used to
// replace the candidate list wholesale, which turned the other eight into wrong
// 200s rather than errors: nemotron pinned to asr made Vad answer with zero
// segments after a full ASR decode, and silero_vad pinned to vad made
// AudioTranscription answer with empty text and four segments whose spans were
// VAD segments, which the srt/vtt/lrc writers rendered as timed EMPTY cues.
static void test_pin_must_be_admissible_for_the_rpc() {
Capabilities nemotron_asr{"nemotron_asr",
{{Task::Asr, {Mode::Offline, Mode::Streaming}},
{Task::Vad, {Mode::Offline}}}};
// The pin is legitimate on the RPC it was meant for.
RequestShape asr_pin;
asr_pin.pinned_task = "asr";
const auto transcription =
resolve_route(Rpc::AudioTranscription, asr_pin, nemotron_asr);
check(transcription.ok && transcription.task == Task::Asr,
"an admissible pin is still honoured exactly");
// ...and refused on one that never routes to it, EVEN THOUGH the family
// advertises the pinned task. That is the whole point: family support is
// not the question, RPC admissibility is.
const auto vad = resolve_route(Rpc::Vad, asr_pin, nemotron_asr);
check(!vad.ok, "an inadmissible pin is refused rather than served");
check(vad.error.find("asr") != std::string::npos,
"the refusal names the pinned task");
check(vad.error.find(rpc_name(Rpc::Vad)) != std::string::npos,
"the refusal names the RPC that cannot serve it");
Capabilities silero{"silero_vad", {{Task::Vad, {Mode::Offline}}}};
RequestShape vad_pin;
vad_pin.pinned_task = "vad";
const auto vad_ok = resolve_route(Rpc::Vad, vad_pin, silero);
check(vad_ok.ok && vad_ok.task == Task::Vad, "vad is admissible on Vad");
const auto transcribe_vad =
resolve_route(Rpc::AudioTranscription, vad_pin, silero);
check(!transcribe_vad.ok,
"a vad pin cannot make a transcription request return empty cues");
// Every pin a shipped configuration could sensibly set stays reachable on
// the RPC that serves it. This is the list the fix was checked against.
struct AdmissibleCase {
Rpc rpc;
const char *task;
};
const AdmissibleCase kAdmissible[] = {
{Rpc::AudioTransform, "svc"}, {Rpc::AudioTransform, "sep"},
{Rpc::AudioTransform, "vc"}, {Rpc::AudioTransform, "s2s"},
{Rpc::Tts, "tts"}, {Rpc::Tts, "clon"},
{Rpc::Tts, "vdes"}, {Rpc::TtsStream, "tts"},
{Rpc::AudioTranscription, "asr"},
{Rpc::AudioTranscription, "align"},
{Rpc::AudioTranscriptionStream, "asr"},
{Rpc::AudioTranscriptionLive, "asr"},
{Rpc::Vad, "vad"}, {Rpc::Diarize, "diar"},
{Rpc::SoundGeneration, "gen"},
};
for (const auto &entry : kAdmissible) {
Task task = Task::Tts;
check(parse_task_name(entry.task, task),
std::string("known task name: ") + entry.task);
// A family that advertises the pinned task offline and nothing else, so
// the ONLY thing that can refuse the route is the admissibility check.
Capabilities only{"probe",
{{task, {Mode::Offline, Mode::Streaming}}}};
RequestShape pin;
pin.pinned_task = entry.task;
const auto route = resolve_route(entry.rpc, pin, only);
check(route.ok && route.task == task,
std::string("pin '") + entry.task + "' stays admissible on " +
rpc_name(entry.rpc));
}
// And the pins that must NOT cross over, one per RPC pair that was
// observed producing a wrong 200.
const AdmissibleCase kInadmissible[] = {
{Rpc::Vad, "asr"}, {Rpc::Diarize, "asr"},
{Rpc::AudioTranscription, "vad"}, {Rpc::AudioTranscription, "diar"},
{Rpc::Tts, "asr"}, {Rpc::Vad, "tts"},
{Rpc::SoundGeneration, "tts"}, {Rpc::AudioTransform, "asr"},
};
for (const auto &entry : kInadmissible) {
Task task = Task::Tts;
check(parse_task_name(entry.task, task),
std::string("known task name: ") + entry.task);
Capabilities only{"probe",
{{task, {Mode::Offline, Mode::Streaming}}}};
RequestShape pin;
pin.pinned_task = entry.task;
const auto route = resolve_route(entry.rpc, pin, only);
check(!route.ok,
std::string("pin '") + entry.task + "' is refused on " +
rpc_name(entry.rpc));
}
}
static void test_vad_and_diarize() {
Capabilities silero{"silero_vad", {{Task::Vad, {Mode::Offline, Mode::Streaming}}}};
const auto v = resolve_route(Rpc::Vad, RequestShape{}, silero);
check(v.ok && v.task == Task::Vad && v.mode == Mode::Offline, "VAD routes offline");
const auto d = resolve_route(Rpc::Diarize, RequestShape{}, silero);
check(!d.ok, "diarization is refused on a VAD-only family");
Capabilities sortformer{"sortformer_diar", {{Task::Diarization, {Mode::Offline}}}};
const auto ok = resolve_route(Rpc::Diarize, RequestShape{}, sortformer);
check(ok.ok && ok.task == Task::Diarization, "diarization routes");
}
static void test_sound_generation() {
Capabilities stable{"stable_audio", {{Task::AudioGeneration, {Mode::Offline}}}};
const auto r = resolve_route(Rpc::SoundGeneration, RequestShape{}, stable);
check(r.ok && r.task == Task::AudioGeneration, "sound generation routes");
}
static void test_names_round_trip() {
const Task all[] = {Task::Vad, Task::Asr, Task::Diarization,
Task::SourceSeparation, Task::AudioGeneration, Task::Tts,
Task::VoiceCloning, Task::VoiceConversion,
Task::SpeechToSpeech, Task::Alignment, Task::VoiceDesign,
Task::SpeakerRecognition, Task::Svc};
for (const Task t : all) {
Task parsed = Task::Vad;
const bool ok = parse_task_name(task_name(t), parsed);
check(ok && parsed == t,
std::string("task name round-trips: ") + task_name(t));
}
check(std::string(mode_name(Mode::Offline)) == "offline", "offline name");
check(std::string(mode_name(Mode::Streaming)) == "streaming", "streaming name");
// The emitted name must be the one audio.cpp itself prints and parses
// (framework/runtime/session.cpp), because `task:` is user-facing: a name
// copied out of audio.cpp has to be accepted here, and a name pinned here
// has to survive conversion at the engine boundary.
check(std::string(task_name(Task::SpeakerRecognition)) == "spk",
"speaker recognition emits upstream's name 'spk'");
Task pinned = Task::Vad;
check(parse_task_name("spk", pinned) && pinned == Task::SpeakerRecognition,
"'spk' parses to SpeakerRecognition");
// Accepted as a legacy alias so configs written against the earlier name
// keep working, but never emitted.
Task alias = Task::Vad;
check(parse_task_name("spkrec", alias) && alias == Task::SpeakerRecognition,
"'spkrec' is still accepted as an alias");
}
static void test_describe_capabilities() {
const std::string described = describe_capabilities(nemotron());
check(described.find("asr/offline") != std::string::npos,
"description lists asr/offline");
check(described.find("asr/streaming") != std::string::npos,
"description lists asr/streaming");
}
static void test_empty_capabilities() {
const auto r = resolve_route(Rpc::Tts, RequestShape{}, Capabilities{"mystery", {}});
check(!r.ok, "a family advertising nothing is refused");
check(r.error.find("mystery") != std::string::npos, "error names the family");
}
// The table is indexed by UnsupportedRpc's underlying value, so a reordering of
// either list silently pairs an RPC with another's reason. Nothing else would
// catch that: both sides still compile and every message still reads plausibly.
static void test_unsupported_surface_table_matches_the_enum() {
check(unsupported_surfaces().size() == 5,
"all five unsupported surfaces are tabulated");
check(std::string(unsupported_surface(UnsupportedRpc::AudioEncode).rpc) ==
"AudioEncode",
"UnsupportedRpc::AudioEncode indexes AudioEncode");
check(std::string(unsupported_surface(UnsupportedRpc::AudioDecode).rpc) ==
"AudioDecode",
"UnsupportedRpc::AudioDecode indexes AudioDecode");
check(std::string(
unsupported_surface(UnsupportedRpc::AudioTransformStream).rpc) ==
"AudioTransformStream",
"UnsupportedRpc::AudioTransformStream indexes AudioTransformStream");
check(std::string(
unsupported_surface(UnsupportedRpc::AudioToAudioStream).rpc) ==
"AudioToAudioStream",
"UnsupportedRpc::AudioToAudioStream indexes AudioToAudioStream");
check(std::string(unsupported_surface(UnsupportedRpc::VoiceEmbed).rpc) ==
"VoiceEmbed",
"UnsupportedRpc::VoiceEmbed indexes VoiceEmbed");
// Two entries may share a reason (the codec pair does), but two entries
// naming the same RPC would mean one of the five is unreachable.
for (size_t i = 0; i < unsupported_surfaces().size(); ++i) {
for (size_t j = i + 1; j < unsupported_surfaces().size(); ++j) {
check(std::string(unsupported_surfaces()[i].rpc) !=
unsupported_surfaces()[j].rpc,
std::string("no duplicate RPC name at ") + std::to_string(i) +
"/" + std::to_string(j));
}
}
}
// There is no out-of-range test for unsupported_surface(). It switches over the
// enumerators with no default label, so a sixth UnsupportedRpc without a case is
// a -Wswitch diagnostic at build time and cannot reach a run-time check at all.
// Every entry, not just the one somebody remembered to cover. A refusal that
// drops the family, the RPC or the reason is a refusal the caller cannot act
// on, which is the entire point of this surface existing.
static void test_every_unsupported_surface_message_is_diagnosable() {
for (const auto &surface : unsupported_surfaces()) {
const std::string label = std::string(" [") + surface.rpc + "]";
const std::string loaded =
unsupported_surface_message(nemotron(), surface.rpc, surface.reason);
check(loaded.find(surface.rpc) != std::string::npos,
"message names the RPC" + label);
check(std::string(surface.reason).size() > 20 &&
loaded.find(surface.reason) != std::string::npos,
"message gives a substantive upstream reason" + label);
check(loaded.find("nemotron_asr") != std::string::npos,
"message names the loaded family" + label);
check(loaded.find("asr/offline") != std::string::npos &&
loaded.find("asr/streaming") != std::string::npos,
"message lists what the family does support" + label);
// The no-model form keeps the two facts that do not depend on a model
// and drops only the one that does, so the caller still learns why.
const std::string unloaded =
unsupported_surface_message(surface.rpc, surface.reason);
check(unloaded.find(surface.rpc) != std::string::npos,
"no-model message names the RPC" + label);
check(unloaded.find(surface.reason) != std::string::npos,
"no-model message gives the upstream reason" + label);
check(unloaded.find("nemotron_asr") == std::string::npos,
"no-model message names no family" + label);
// Without this the caller's obvious next move is to load a model and
// retry, which cannot work: the refusal is a property of the engine.
check(unloaded.find("would not change this answer") != std::string::npos,
"no-model message says loading a model would not help" + label);
}
}
// The reasons are the load-bearing half of this feature and each was checked
// against the pinned upstream checkout. Pinning the distinguishing phrase here
// means a later edit that guts one into a generic "not supported" fails rather
// than passes quietly.
static void test_unsupported_reasons_name_the_upstream_limitation() {
const auto &encode = unsupported_surface(UnsupportedRpc::AudioEncode);
const auto &decode = unsupported_surface(UnsupportedRpc::AudioDecode);
check(std::string(encode.reason).find("VoiceTaskKind") != std::string::npos &&
std::string(encode.reason).find("codec") != std::string::npos,
"the AudioEncode reason names the missing VoiceTaskKind entry");
// miocodec is the family a reader will reach for first, because upstream's
// README tags it Codec. Naming it and its actual advertised tasks is what
// stops the next person re-deriving the same dead end.
check(std::string(encode.reason).find("miocodec") != std::string::npos,
"the AudioEncode reason disposes of miocodec's README Codec tag");
check(std::string(encode.reason) == decode.reason,
"AudioEncode and AudioDecode refuse for the same reason");
const auto &transform =
unsupported_surface(UnsupportedRpc::AudioTransformStream);
// The reason must be scoped to the tasks AudioTransform routes to. The
// broader claim, "upstream streams tts and asr only", is FALSE: silero_vad
// advertises vad with RunMode::Streaming. A refusal resting on a false
// premise is worse than a bare UNIMPLEMENTED, because it will be believed.
check(std::string(transform.reason).find("sep, vc, svc, s2s") !=
std::string::npos,
"the AudioTransformStream reason is scoped to the routed tasks");
check(std::string(transform.reason).find("tts, asr and vad") !=
std::string::npos,
"the AudioTransformStream reason counts vad among the streaming tasks");
check(std::string(transform.reason).find("tts and asr only") ==
std::string::npos,
"the AudioTransformStream reason does not repeat the refuted claim");
// An absence is not an impossibility. A sep family could be buffered and
// emitted as a stream, so the reason has to say this backend declines to
// rather than cannot, or it overreaches on a true premise.
check(std::string(transform.reason).find("buffered offline call in disguise") !=
std::string::npos,
"the AudioTransformStream reason does not overclaim impossibility");
const auto &s2s = unsupported_surface(UnsupportedRpc::AudioToAudioStream);
check(std::string(s2s.reason).find("Realtime") != std::string::npos &&
std::string(s2s.reason).find("clip-to-clip") != std::string::npos,
"the AudioToAudioStream reason contrasts the two contracts");
// Naming both s2s families, and what each of them actually does, makes the
// claim checkable. It must NOT say s2s is voice conversion full stop: that
// is true of miocodec and false of vevo2, whose s2s route is `editing` and
// rewrites the spoken content against a target voice.
check(std::string(s2s.reason).find("miocodec (voice conversion)") !=
std::string::npos &&
std::string(s2s.reason).find("vevo2 (speech editing)") !=
std::string::npos,
"the AudioToAudioStream reason names each s2s family's actual task");
check(std::string(s2s.reason).find("s2s is offline voice conversion") ==
std::string::npos,
"the AudioToAudioStream reason does not miscast vevo2 as conversion");
const auto &embed = unsupported_surface(UnsupportedRpc::VoiceEmbed);
check(std::string(embed.reason).find("spk") != std::string::npos,
"the VoiceEmbed reason names the task no family advertises");
// spk IS a VoiceTaskKind upstream; what is missing is any family that
// advertises it. Saying the kind does not exist would be false, and would
// send a reader looking in the wrong place.
check(std::string(embed.reason).find("no audio.cpp family") !=
std::string::npos,
"the VoiceEmbed reason blames the families, not the enum");
// The speaker encoders DO exist upstream, as conditioning modules inside
// TTS and VC families. Not saying so invites "but audio.cpp ships TitaNet".
check(std::string(embed.reason).find("TitaNet") != std::string::npos,
"the VoiceEmbed reason disposes of the internal speaker encoders");
Task parsed = Task::Vad;
check(parse_task_name("spk", parsed) && parsed == Task::SpeakerRecognition,
"spk is a real task kind, so the reason must not claim otherwise");
}
// A family advertising nothing still gets a message that reads, rather than one
// trailing off after "supports: ".
static void test_unsupported_surface_message_with_empty_capabilities() {
const auto &embed = unsupported_surface(UnsupportedRpc::VoiceEmbed);
const std::string message = unsupported_surface_message(
Capabilities{"mystery", {}}, embed.rpc, embed.reason);
check(message.find("mystery") != std::string::npos,
"empty-capability message still names the family");
check(message.find("supports: nothing") != std::string::npos,
"empty-capability message says the family supports nothing");
}
int main() {
test_plain_tts();
test_tts_with_voice_reference_prefers_cloning();
test_tts_voice_reference_falls_back_to_tts();
test_tts_instructions_prefer_voice_design();
test_voice_reference_beats_instructions();
test_tts_stream_requires_streaming();
test_transcription_stream_falls_back_to_offline();
test_task_preference_beats_mode_preference();
test_live_transcription_has_no_offline_fallback();
test_alignment_needs_prompt_text();
test_prompt_does_not_hijack_asr();
test_audio_transform_prefers_separation();
test_pinned_task_overrides_routing();
test_pin_must_be_admissible_for_the_rpc();
test_vad_and_diarize();
test_sound_generation();
test_names_round_trip();
test_describe_capabilities();
test_empty_capabilities();
test_unsupported_surface_table_matches_the_enum();
test_every_unsupported_surface_message_is_diagnosable();
test_unsupported_reasons_name_the_upstream_limitation();
test_unsupported_surface_message_with_empty_capabilities();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all capability_routing checks passed\n");
return 0;
}

View File

@@ -0,0 +1,163 @@
#include "family_gate.h"
#include <cctype>
#include <cstddef>
namespace audiocpp_backend {
namespace {
// PRECONDITION: `suffix` must already be lowercase. Both sides are folded, so
// this reads as symmetric, but only `value` can carry case in practice and a
// caller passing ".GGUF" would still work today for that reason alone. Do not
// rely on it: the fold on the suffix side is the only thing standing between
// this and a helper that answers false for every input, and it is not covered
// by any test, because with a lowercase suffix no input can distinguish it.
bool ends_with_ci(const std::string &value, const std::string &suffix) {
if (value.size() <= suffix.size()) {
return false; // a bare ".gguf" is an extension, not a model file
}
const size_t offset = value.size() - suffix.size();
for (size_t i = 0; i < suffix.size(); ++i) {
const auto lhs = static_cast<unsigned char>(value[offset + i]);
const auto rhs = static_cast<unsigned char>(suffix[i]);
if (std::tolower(lhs) != std::tolower(rhs)) {
return false;
}
}
return true;
}
} // namespace
bool path_looks_like_gguf(const std::string &path) {
return ends_with_ci(path, ".gguf");
}
FamilyDecision decide_family(bool path_is_gguf, const std::string &embedded_family,
const std::string &configured_family) {
FamilyDecision decision;
if (!configured_family.empty()) {
decision.ok = true;
decision.family = configured_family;
return decision;
}
if (path_is_gguf) {
if (!embedded_family.empty()) {
decision.ok = true;
decision.family = embedded_family;
return decision;
}
decision.error =
"audio-cpp: this GGUF carries no 'audiocpp.model_spec.family' "
"metadata key, so it is not an audio.cpp model. Convert it with "
"audiocpp_gguf, or name the family explicitly with the model option "
"'family:<name>'";
return decision;
}
decision.error =
"audio-cpp: a model path that is not a standalone audio.cpp GGUF needs "
"an explicit 'family:<name>' model option, because the audio.cpp family "
"cannot be inferred from a safetensors or package directory";
return decision;
}
namespace {
// Families that ABORT THE PROCESS on a weight dtype they cannot handle, and the
// dtypes they can. See the header for why this is a list of crashes rather than
// a list of preferences.
//
// supertonic: upstream's docs/gguf.md:90 records its 16-bit GGUF column as
// "---", i.e. NOT TESTED, and its q8_0 as "No (unsupported weight dtype)". Only
// the `orig` package is marked Pass, and its 698 weight tensors are f32 while
// its 72 index and shape constants are i64. The f16 abort is a LOCAL
// OBSERVATION rather than an upstream claim, and it is attributed rather than
// assumed: it is identical through the unary TTS RPC and through TTSStream, so
// it is the packaging and not the streaming path. q8_0 was never run here and is
// refused on upstream's "unsupported weight dtype" alone, which is the weaker of
// the two claims. See the header for why keeping them apart matters.
//
// TO REMOVE AN ENTRY: bump AUDIO_CPP_VERSION past a fix, load a package in the
// refused dtype, and synthesise. If audio comes out, delete the entry. No test
// can do that for you, which is exactly why it is written here: the test beside
// this file pins WHAT the table says, not whether upstream has moved on. Do not
// widen an entry without running that, because what it prevents is a process
// death rather than a wrong answer.
struct DtypeAllowList {
// NULL TERMINATED, and the terminator occupies one of these slots: both
// loops below stop at the first nullptr and have no other bound, so an entry
// that named three dtypes would leave them reading past the end of the
// array. That is undefined behaviour rather than a wrong answer, and it is
// one keystroke away from any edit that widens an entry, so the terminator
// is asserted at compile time below rather than trusted.
static constexpr std::size_t kSlots = 3;
const char *family;
const char *allowed[kSlots];
};
constexpr DtypeAllowList kDtypeAllowLists[] = {
{"supertonic", {"f32", "i64", nullptr}},
};
constexpr bool allow_lists_are_terminated() {
for (const auto &entry : kDtypeAllowLists) {
if (entry.allowed[DtypeAllowList::kSlots - 1] != nullptr) {
return false;
}
}
return true;
}
static_assert(allow_lists_are_terminated(),
"every DtypeAllowList must leave its last slot null: the lookups "
"below stop at the first nullptr and would otherwise read past "
"the end of the array");
const DtypeAllowList *find_allow_list(const std::string &family) {
for (const auto &entry : kDtypeAllowLists) {
if (family == entry.family) {
return &entry;
}
}
return nullptr;
}
} // namespace
bool family_has_weight_dtype_allow_list(const std::string &family) {
return find_allow_list(family) != nullptr;
}
bool weight_dtype_is_supported(const std::string &family, const std::string &dtype) {
const DtypeAllowList *list = find_allow_list(family);
if (list == nullptr) {
return true;
}
for (const char *const *name = list->allowed; *name != nullptr; ++name) {
if (dtype == *name) {
return true;
}
}
return false;
}
std::string supported_weight_dtypes(const std::string &family) {
const DtypeAllowList *list = find_allow_list(family);
if (list == nullptr) {
return {};
}
std::string out;
for (const char *const *name = list->allowed; *name != nullptr; ++name) {
if (!out.empty()) {
out += ", ";
}
out += *name;
}
return out;
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,86 @@
#pragma once
// Decides which audio.cpp family a model path belongs to, and refuses paths
// this backend must not claim. Standard library only.
//
// This is the guard against issue #9287. A model config with no explicit
// backend makes LocalAI probe every installed backend and bind to the first
// Load that succeeds, so accepting an arbitrary GGUF here would capture
// unrelated LLMs. audio.cpp GGUFs carry an audiocpp.model_spec.family metadata
// key; llama.cpp GGUFs do not.
#include <string>
namespace audiocpp_backend {
// True when the path ends in ".gguf", case insensitively, and has a stem.
bool path_looks_like_gguf(const std::string &path);
struct FamilyDecision {
bool ok = false;
std::string family;
// Set when ok is false. Suitable verbatim as an INVALID_ARGUMENT message.
std::string error;
};
// Precedence:
// 1. an explicit `family:` option, so a user can override wrong metadata;
// 2. for a GGUF, the family embedded in audiocpp.model_spec.family;
// 3. otherwise refuse.
// A directory path never consults embedded metadata: there is no single GGUF
// to read it from.
FamilyDecision decide_family(bool path_is_gguf, const std::string &embedded_family,
const std::string &configured_family);
// True when `family` can run weights stored as `dtype`, where dtype is the
// string a TensorMetadata carries ("f32", "f16", "q8_0", "i64", ...).
//
// This is a LIST OF FAMILIES THAT CRASH THE PROCESS, not a list of families that
// perform badly. It exists because the failure is not an exception: loading the
// supertonic f16 GGUF package reaches ggml_concat with one f16 operand and one
// f32 one, GGML_ASSERT(a->type == b->type) fails (external/ggml/src/ggml.c:2595)
// and ggml_abort takes the backend down with SIGABRT on the FIRST request.
// Nothing upstream of the load can catch that, so an operator sees a model that
// loaded successfully and a backend that dies on every request with no status
// and no message.
//
// EVIDENCE, per dtype, because the two are not equally attested:
// - f16 was OBSERVED to abort here, identically through the unary TTS RPC and
// through TTSStream, so it is the packaging and not the streaming path.
// Upstream's docs/gguf.md:90 has supertonic's 16-bit column as "---", which
// its own legend (:53) defines as not tested, so upstream neither confirms
// nor contradicts it.
// - q8_0 was NOT run here. Upstream records it as "No (unsupported weight
// dtype)" in the same row, which is a weaker claim than the f16 abort: it
// says the format is unusable, not that it takes the process down.
// Both are refused, because the allow list is what the family CAN run (f32 for
// weights, i64 for the shape and index constants) rather than a list of the
// dtypes that fail, and a format upstream calls unusable has no business being
// loaded either way.
//
// A family with no entry is unrestricted, which is every family but one.
//
// Split out of loaded_model.cpp, where the caller lives, so that the policy is
// stdlib-only and can be held by a test: the caller needs a real GGUF on disk
// and an engine, and neither is available to a unit test. What the test pins is
// that the table says what it is meant to say, so widening it is a deliberate
// act rather than a typo. It CANNOT pin the removal criterion, which is
// "upstream fixed it": no test can know that without downloading the package and
// synthesising, so that step stays a documented manual one at the table itself.
bool weight_dtype_is_supported(const std::string &family, const std::string &dtype);
// True when `family` has an entry in the table at all, which is the question a
// caller deciding whether to OPEN THE FILE has to ask. Distinct from
// "supported_weight_dtypes(family) is empty": that string is also empty for an
// entry with an empty allow list, and such an entry means "this family can run
// nothing", which weight_dtype_is_supported already answers by refusing every
// dtype. Deciding from the string would skip the check on precisely the entry
// that most needs it.
bool family_has_weight_dtype_allow_list(const std::string &family);
// The dtypes `family` is restricted to, as "f32, i64", or empty when it is not
// restricted at all. For the refusal message, so the operator is told what to
// look for rather than only what is wrong.
std::string supported_weight_dtypes(const std::string &family);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,179 @@
// Unit tests for family_gate. Standard library only. The harness compiles this
// as a single translation unit, so the implementation is included directly.
//
// This unit is the guard against issue #9287: when a model config has no
// explicit backend, LocalAI probes every installed backend and binds to the
// first Load that succeeds. Accepting an arbitrary GGUF here would capture
// unrelated LLMs.
#include "family_gate.cpp"
#include <cstdio>
#include <string>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
static void check_eq(const std::string &got, const std::string &want,
const std::string &name) {
check(got == want, name + " (got \"" + got + "\" want \"" + want + "\")");
}
using namespace audiocpp_backend;
static void test_gguf_suffix_detection() {
check(path_looks_like_gguf("/models/chatterbox-q8_0.gguf"), "plain .gguf");
check(path_looks_like_gguf("/models/CHATTERBOX.GGUF"), "uppercase .GGUF");
check(path_looks_like_gguf("/models/x.GgUf"), "mixed case .GgUf");
check(path_looks_like_gguf("a.gguf"), "a one character stem is still a stem");
check(!path_looks_like_gguf("/models/chatterbox"), "extensionless directory");
check(!path_looks_like_gguf("/models/model.safetensors"), "safetensors");
check(!path_looks_like_gguf("/models/gguf"), "a name that is merely 'gguf'");
check(!path_looks_like_gguf("/models/GGUF"), "an uppercase name that is merely 'GGUF'");
check(!path_looks_like_gguf(""), "empty path");
check(!path_looks_like_gguf(".gguf"), "a bare extension is not a model file");
// The suffix has to be at the end. A prefix or infix match would let
// ".gguf.tmp" download artefacts and ".ggufx" siblings through.
check(!path_looks_like_gguf("/models/model.gguf.tmp"), ".gguf in the middle");
check(!path_looks_like_gguf("/models/model.ggufx"), "a longer extension");
// Every character of the suffix has to match, including the last one.
check(!path_looks_like_gguf("/models/model.ggug"), "a near miss in the final character");
check(!path_looks_like_gguf("/models/model_gguf"), "a near miss in the first character");
check(!path_looks_like_gguf("/models/.gguf-notes"), "a leading .gguf");
}
static void test_explicit_family_always_wins() {
// Explicit configuration beats metadata, so a user can force a family when
// upstream metadata is wrong or absent.
const auto gguf = decide_family(true, "chatterbox", "omnivoice");
check(gguf.ok && gguf.family == "omnivoice", "explicit family overrides GGUF metadata");
check(gguf.error.empty(), "an accepted decision carries no error text");
const auto dir = decide_family(false, "", "qwen3_tts");
check(dir.ok && dir.family == "qwen3_tts", "explicit family satisfies a directory path");
// A GGUF with no embedded spec is still loadable when the user names the
// family: the option is an override, not a tie-break that needs metadata to
// break against.
const auto bare = decide_family(true, "", "supertonic");
check(bare.ok && bare.family == "supertonic",
"explicit family rescues a GGUF that carries no spec");
}
static void test_gguf_metadata_supplies_the_family() {
const auto d = decide_family(true, "nemotron_asr", "");
check(d.ok, "an audio.cpp GGUF loads with no family option");
check(d.family == "nemotron_asr", "family comes from the embedded spec");
check(d.error.empty(), "an accepted GGUF carries no error text");
}
// THE GATE. A llama.cpp GGUF has no audiocpp.model_spec.family key.
static void test_foreign_gguf_is_refused() {
const auto d = decide_family(true, "", "");
check(!d.ok, "a GGUF with no audio.cpp spec is refused");
check(d.family.empty(), "no family is guessed");
check(!d.error.empty(), "a refusal always says why");
check(d.error.find("audiocpp.model_spec.family") != std::string::npos,
"error names the missing metadata key so the cause is diagnosable");
check(d.error.find("family:") != std::string::npos,
"error names the option that would override it");
}
static void test_directory_without_family_is_refused() {
const auto d = decide_family(false, "", "");
check(!d.ok, "a non-GGUF path with no family option is refused");
check(d.family.empty(), "a refused directory guesses no family");
check(d.error.find("family:") != std::string::npos,
"error names the required option");
}
// A directory path never consults embedded metadata, because there is no single
// GGUF to read it from.
static void test_directory_ignores_embedded_family() {
const auto d = decide_family(false, "chatterbox", "");
check(!d.ok, "a directory is refused even when an embedded family is supplied");
check(d.family.empty(), "a refused directory does not adopt the embedded family");
// If the GGUF branch ever leaked into the directory branch this message
// would start blaming a metadata key that a directory has no place to carry.
check(d.error.find("audiocpp.model_spec.family") == std::string::npos,
"a directory refusal does not blame GGUF metadata it could not have");
}
// Pins the weight-dtype allow list. Not a style preference: an entry here is a
// family that ABORTS THE PROCESS on the first request when handed the wrong
// dtype, so the model loads and then every request kills the backend with no
// status and no message.
//
// What this test can and cannot do, stated so the next reader does not expect
// more of it: it pins WHAT THE TABLE SAYS, so widening an entry is a deliberate
// act rather than a typo, and it pins that an unlisted family is unrestricted.
// It CANNOT pin the removal criterion, which is "upstream fixed it": knowing
// that needs the package downloaded and a synthesis run, so it stays a manual
// step documented at the table in family_gate.cpp.
static void test_weight_dtype_allow_list() {
// The entry that exists, and the exact reason it exists.
check(!weight_dtype_is_supported("supertonic", "f16"),
"supertonic refuses f16, the package that aborts the process");
check(!weight_dtype_is_supported("supertonic", "q8_0"),
"supertonic refuses q8_0, which upstream records as unsupported");
check(!weight_dtype_is_supported("supertonic", "bf16"),
"supertonic refuses bf16, which is untested rather than known good");
check(weight_dtype_is_supported("supertonic", "f32"),
"supertonic accepts f32, which is what the orig package stores");
check(weight_dtype_is_supported("supertonic", "i64"),
"supertonic accepts i64: the orig package carries 72 such tensors and "
"refusing them would refuse the artifact that works");
// Every other family is unrestricted, and must stay that way: this guard is
// for process death, not for quality.
check(weight_dtype_is_supported("nemotron_asr", "q8_0"),
"an unlisted family is not restricted");
check(weight_dtype_is_supported("citrinet_asr", "f16"),
"an unlisted family is not restricted by another family's entry");
check(weight_dtype_is_supported("", "anything"),
"an empty family name is not restricted");
// The message the operator reads has to name the remedy, so the refusal is
// actionable rather than only correct.
check_eq(supported_weight_dtypes("supertonic"), "f32, i64",
"the refusal can name what to look for");
check_eq(supported_weight_dtypes("nemotron_asr"), "",
"an unlisted family reports no restriction");
// What the caller actually decides on, and it is a DIFFERENT question from
// "is the description empty": an entry with an empty allow list would
// describe itself as "" while refusing every dtype, so a caller that skipped
// the file read on the empty string would skip the check on the one entry
// that refuses everything.
check(family_has_weight_dtype_allow_list("supertonic"),
"a listed family has an allow list");
check(!family_has_weight_dtype_allow_list("nemotron_asr"),
"an unlisted family has none, which is what lets the caller skip "
"opening the file at all");
check(!family_has_weight_dtype_allow_list(""),
"an empty family name has no allow list");
}
int main() {
test_gguf_suffix_detection();
test_explicit_family_always_wins();
test_gguf_metadata_supplies_the_family();
test_foreign_gguf_is_refused();
test_directory_without_family_is_refused();
test_directory_ignores_embedded_family();
test_weight_dtype_allow_list();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all family_gate checks passed\n");
return 0;
}

View File

@@ -0,0 +1,280 @@
#include "generation_request.h"
#include <filesystem>
#include <string>
#include <system_error>
#include <utility>
namespace audiocpp_backend {
namespace {
const char *bool_option(bool value) { return value ? "true" : "false"; }
} // namespace
bool voice_is_reference_file(const std::string &voice) {
if (voice.empty()) {
return false;
}
std::error_code ec;
return std::filesystem::is_regular_file(std::filesystem::path(voice), ec);
}
RequestShape build_tts_shape(const backend::TTSRequest &request) {
RequestShape shape;
shape.has_voice_reference = voice_is_reference_file(request.voice());
// !empty() as well as has_instructions(), and it must match the guard in
// build_tts_request: a request whose instructions are an empty string
// carries no style condition, so telling routing to prefer VoiceDesign for
// it would route to a task with nothing to design from.
shape.has_instructions =
request.has_instructions() && !request.instructions().empty();
return shape;
}
engine::runtime::TaskRequest
build_tts_request(const backend::TTSRequest &request,
std::optional<engine::runtime::AudioBuffer> reference_audio) {
engine::runtime::TaskRequest task;
// The Transcript, not an option, is where every TTS family reads its
// language: chatterbox normalises request.text_input->language into its
// voice-clone config, qwen3_tts reads it as out.language, ace_step turns it
// into vocal_language. Upstream's own HTTP server does the same and sets no
// language option at all (app/server/runtime.cpp build_speech_request).
engine::runtime::Transcript transcript;
transcript.text = request.text();
if (request.has_language()) {
transcript.language = request.language();
}
task.text_input = std::move(transcript);
engine::runtime::VoiceCondition condition;
bool condition_used = false;
if (reference_audio.has_value()) {
// A clip: VoiceReference::audio, at the file's own rate and channel
// count. See kVoiceReferenceSampleRate in grpc-server.cpp for why it is
// not folded first.
engine::runtime::VoiceReference reference;
reference.audio = std::move(*reference_audio);
condition.speaker = std::move(reference);
condition_used = true;
} else if (!request.voice().empty()) {
// A named preset. cached_voice_id is the channel that actually lands:
// supertonic (options.voice), pocket_tts (voice_config.preset_name),
// voxcpm2, vibevoice, fish_audio (its saved-reference lookup) and
// qwen3_tts CustomVoice all read request.voice->speaker->cached_voice_id,
// and upstream's own server puts a non-preset `voice` body field in
// exactly this slot (app/server/runtime.cpp build_speech_request).
engine::runtime::VoiceReference reference;
reference.cached_voice_id = request.voice();
condition.speaker = std::move(reference);
condition_used = true;
// Forward-tolerant alias only. A bare "voice" REQUEST OPTION is read by
// no family in the pinned upstream: grepping find_option for it returns
// nothing. It is sent so a family adopting the name later works with no
// change here, not because it does anything today.
task.options["voice"] = request.voice();
}
if (request.has_instructions() && !request.instructions().empty()) {
// "instruct" is the key upstream itself maps the OpenAI `instructions`
// body field onto (app/server/runtime.cpp: request.options["instruct"]
// = value->as_string()), and it is read: qwen3_tts VoiceDesign and
// CustomVoice both take find_option(options, {"instruct"}) first, and
// omnivoice reads it in resolve_instruct.
task.options["instruct"] = request.instructions();
// "caption" is irodori_tts's name for the same thing, read in its
// make_request and documented as the voice-design caption for the 600M
// VoiceDesign model (docs/tts.md). Without this, that family's voice
// design cannot be driven from this RPC at all.
task.options["caption"] = request.instructions();
// The proto field's own name, forwarded for the same forward-tolerant
// reason as "voice" above and with the same honest accounting: NO family
// in the pinned upstream reads a request option called "instructions".
task.options["instructions"] = request.instructions();
engine::runtime::StyleCondition style;
// "instruct", not "instructions". This tag IS read, and only under that
// spelling: omnivoice and qwen3_tts both fall back to
// request.voice->style->tags.find("instruct") when the option is absent.
// Spelling it "instructions" here would have made the whole
// StyleCondition dead weight.
style.tags["instruct"] = request.instructions();
// !empty(), matching the option emission below, and load-bearing rather
// than tidiness. core/backend/tts.go's newTTSRequest sets
// `Language: &language` UNCONDITIONALLY, so has_language() is true on
// every request LocalAI sends and carries "" whenever the caller named
// no language. An engaged-but-empty style language is WORSE than an
// absent one: supertonic reads text_input->language behind its own
// !empty() guard and then OVERRIDES it from style->language with no
// guard at all (supertonic/session.cpp), so "" would replace its "en"
// default and tokenizer_text.cpp would throw
// "invalid Supertonic language: " on every request that set
// instructions and no language.
if (request.has_language() && !request.language().empty()) {
style.language = request.language();
}
condition.style = std::move(style);
condition_used = true;
}
if (condition_used) {
task.voice = std::move(condition);
}
if (request.has_language() && !request.language().empty()) {
// Forward-tolerant alias, exactly as in build_transcription_request. The
// families that read a "language" request option are the ASR ones
// (nemotron_asr, hviske_asr, vibevoice_asr, higgs_audio_stt), none of
// which this RPC can route to; pocket_tts reads one but from its
// ModelLoadRequest at load time, not from here. The Transcript above is
// what actually carries the language to a TTS family.
task.options["language"] = request.language();
}
// LAST, so an explicit params entry wins over anything derived above. That
// matters for "caption": a caller who sets params[caption] has named the
// exact string they want, and it must not be overwritten by `instructions`.
for (const auto &param : request.params()) {
task.options[param.first] = param.second;
}
return task;
}
engine::runtime::TaskRequest
build_sound_generation_request(const backend::SoundGenerationRequest &request,
std::optional<engine::runtime::AudioBuffer> source_audio) {
engine::runtime::TaskRequest task;
engine::runtime::Transcript transcript;
transcript.text = request.text();
if (request.has_language()) {
transcript.language = request.language();
}
task.text_input = std::move(transcript);
// src is the input clip for the editing routes. ace_step's repaint, cover
// and edit routes need it; stable_audio uses it as init_audio or
// inpaint_audio; heartmula refuses it outright.
if (source_audio.has_value()) {
task.audio_input = std::move(*source_audio);
}
// WHAT LANDS AND WHAT DOES NOT. Three families advertise AudioGeneration in
// the pinned upstream: ace_step, heartmula and stable_audio. Every key below
// was grepped against find_option/parse_*_option in src/ and include/ rather
// than assumed, because a key nobody reads is not a feature and shipping one
// while implying it works is the mistake this comment exists to prevent.
//
// Unknown REQUEST options cannot turn a valid request into an error:
// families look theirs up by name and ignore the rest, and the unknown-key
// refusals upstream does have are on SESSION options, which arrive at load
// time. So a forward-tolerant alias is free; it is just not a feature.
if (request.has_duration()) {
// duration_seconds is the key that works, and it works everywhere:
// ace_step (request_parser.cpp), heartmula (session.cpp, which also
// refuses a non-positive value) and stable_audio (request.cpp) all read
// it. This is the SoundGeneration analogue of Task 9's return_timestamps.
task.options["duration_seconds"] = std::to_string(request.duration());
// The proto field's own name. Read by exactly one family, omnivoice, and
// omnivoice advertises Tts rather than AudioGeneration, so this RPC can
// never route to it: DEAD here, kept only as a forward-tolerant alias.
task.options["duration"] = std::to_string(request.duration());
}
if (request.has_temperature()) {
// Read by heartmula. ace_step's sampling temperature is a different,
// narrower knob it calls lm_temperature (it drives the caption/thinking
// LM, not the audio diffusion), so this is deliberately NOT mapped onto
// it; a caller who wants it sets it through the request options that
// reach ace_step by name. stable_audio has no temperature at all.
task.options["temperature"] = std::to_string(request.temperature());
}
if (request.has_sample()) {
// do_sample is read widely upstream, but only by TTS and ASR families
// (chatterbox, index_tts2, miotts, moss, qwen3_tts, vibevoice,
// hviske_asr, voxtral_realtime). NO AudioGeneration family reads it, so
// it is dead on this route.
task.options["do_sample"] = bool_option(request.sample());
}
if (request.has_src_divisor()) {
// Read by nobody, anywhere in the pinned upstream. Forwarded because the
// proto documents it as part of this request and a family adopting it
// then works unchanged.
task.options["src_divisor"] = std::to_string(request.src_divisor());
}
if (request.has_think()) {
// "thinking" is the key ace_step actually reads (request_parser.cpp),
// which is why it is sent alongside the proto's own "think". "think" on
// its own is read by nobody.
task.options["thinking"] = bool_option(request.think());
task.options["think"] = bool_option(request.think());
}
if (request.has_caption()) {
// Read only by irodori_tts, which advertises Tts/VoiceCloning/VoiceDesign
// and not AudioGeneration, so it is unreachable from this RPC: DEAD here.
task.options["caption"] = request.caption();
}
if (request.has_lyrics()) {
// Read by ace_step and heartmula.
task.options["lyrics"] = request.lyrics();
}
if (request.has_bpm()) {
// Read by ace_step.
task.options["bpm"] = std::to_string(request.bpm());
}
if (request.has_keyscale()) {
// Read by ace_step.
task.options["keyscale"] = request.keyscale();
}
if (request.has_timesignature()) {
// Read by ace_step.
task.options["timesignature"] = request.timesignature();
}
if (request.has_instrumental()) {
// Read by nobody: "instrumental" appears in the pinned upstream only as
// a roformer STEM NAME, never as a request option. A forward-tolerant
// alias and nothing more.
task.options["instrumental"] = bool_option(request.instrumental());
}
if (request.has_language() && !request.language().empty()) {
// Alias again: the Transcript above is what ace_step reads as
// vocal_language. No AudioGeneration family reads a "language" option.
task.options["language"] = request.language();
}
return task;
}
bool apply_transform_text_input(engine::runtime::TaskRequest &task) {
// Canonical first, alias second, and an empty value falls through to the
// next candidate rather than ending the search: a caller who sent
// target_text="" and text="the real one" meant the second one.
static const char *const kTextKeys[] = {"target_text", "text"};
std::string text;
for (const char *key : kTextKeys) {
const auto found = task.options.find(key);
if (found != task.options.end() && !found->second.empty()) {
text = found->second;
break;
}
}
if (text.empty()) {
return false;
}
engine::runtime::Transcript transcript;
transcript.text = std::move(text);
// Inside the has-text branch on purpose. See the header: a language on its
// own conditions nothing and must not manufacture a text_input.
const auto language = task.options.find("language");
if (language != task.options.end()) {
transcript.language = language->second;
}
task.text_input = std::move(transcript);
return true;
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,108 @@
#pragma once
// Builds the engine::runtime::TaskRequest for the two audio-PRODUCING offline
// RPCs, TTS and SoundGeneration, and answers the one filesystem question TTS
// routing depends on.
//
// It is a unit of its own rather than a pair of statics in grpc-server.cpp so
// that it can be tested: grpc-server.cpp has a main() and cannot be linked into
// a test binary, and everything here is a pure function of its arguments once
// the file read has been lifted out (which is why the reference clip arrives as
// an already-read buffer rather than a path). TTSStream reuses build_tts_request
// unchanged.
//
// Only the plain structs in engine/framework/runtime/session.h are touched, so
// this compiles against the header without linking engine_runtime, the same way
// result_map does.
#include "backend.pb.h"
#include "capability_routing.h"
#include "engine/framework/runtime/session.h"
#include <optional>
#include <string>
namespace audiocpp_backend {
// True when TTSRequest.voice names an existing regular file, in which case it
// is a speaker reference clip and routing prefers VoiceCloning; false when it is
// a named preset (or empty).
//
// The overload is LocalAI's, not this backend's: `voice` is the OpenAI speech
// field and different LocalAI backends have always read it both ways. Deciding
// it from the filesystem needs no new option and matches how somebody actually
// configures a cloning family, which is by pointing at a clip.
//
// A DIRECTORY is deliberately not a reference: is_regular_file, not exists. A
// directory named as a voice cannot be read as a WAV, and treating it as a
// reference would turn a preset typo into "cannot read /x as WAV" instead of
// letting it travel as the preset name it looks like.
//
// The error_code overload is used so an unreadable parent directory answers
// false rather than throwing. That is the right answer here: the name is then
// passed on as a preset, and if it really was meant to be a clip the family
// refuses a request it cannot serve, which is a better message than a
// filesystem exception thrown while classifying a string.
bool voice_is_reference_file(const std::string &voice);
// Everything routing needs to know about a TTSRequest, in one place, so that TTS
// and TTSStream cannot describe the same request differently.
//
// `pinned_task` is deliberately NOT filled here: it comes off the LoadedModel,
// not off the request, and this unit links no engine. The caller must still
// write `shape.pinned_task = model->pinned_task();` or the model's `task:`
// option is dead. That is the one field a new handler can forget, so it is the
// one field left visible at the call site rather than hidden behind this
// helper.
RequestShape build_tts_shape(const backend::TTSRequest &request);
// `reference_audio` is the already-read speaker clip, present exactly when
// voice_is_reference_file(request.voice()) was true. Passing it in rather than a
// path keeps this function pure and lets the caller do the read where the
// ordering rules (capability refusal first, then the lane) are enforced.
//
// It is taken BY VALUE and moved in: a reference clip is seconds of audio and
// the caller has no use for it afterwards.
engine::runtime::TaskRequest
build_tts_request(const backend::TTSRequest &request,
std::optional<engine::runtime::AudioBuffer> reference_audio);
// `source_audio` is SoundGenerationRequest.src already read, present exactly
// when the field was set and non-empty. Same reasoning as above.
engine::runtime::TaskRequest
build_sound_generation_request(const backend::SoundGenerationRequest &request,
std::optional<engine::runtime::AudioBuffer> source_audio);
// Lifts a text-conditioned transform route's text out of the request params
// into TaskRequest.text_input, and reports whether it set one.
//
// WHY THIS EXISTS. AudioTransform is an audio-in / audio-out RPC and its proto
// message has no text field, but not every task it routes to is audio-only.
// vevo2's speech-to-speech and prosody routes read their text from
// request.text_input (src/models/vevo2/session.cpp fills refs.target_text from
// exactly there and nowhere else) and refuse the run without one: "Vevo2
// text/prosody route requires text_input or target_text". The params map is the
// only channel AudioTransform has that reaches the engine, so the text travels
// through it and is unpacked here. Without this, s2s is not merely awkward to
// reach through this RPC, it is unreachable.
//
// CALL IT AFTER the params have been copied into task.options, and note that it
// does NOT erase the keys it reads. vevo2's loader advertises "target_text" in
// its own documented request-option table, so a family that looks there keeps
// finding it; the copy in text_input is what the session actually reads today.
//
// "target_text" is canonical and "text" is its alias, the same order vevo2's
// option table declares them in. A request setting both gets target_text, so
// the canonical spelling wins rather than whichever the map happened to store
// first. An empty value is not a text: it means the caller sent the key with
// nothing in it, and a family asked to vocalise "" should say so itself rather
// than be handed an empty Transcript that looks deliberate.
//
// "language" rides along when a text was found, and only then. On its own it
// conditions nothing, and setting text_input for it alone would turn a plain
// separation request that happened to carry a language hint into a text-routed
// one.
bool apply_transform_text_input(engine::runtime::TaskRequest &task);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,575 @@
// Tests for the TTS and SoundGeneration request builders, and for the
// filesystem rule that decides whether TTSRequest.voice is a speaker reference
// clip or a named preset.
//
// NAMED _ctest AND NOT _test ON PURPOSE: see the note at the top of
// result_map_ctest.cpp. This file needs the generated protobuf messages and the
// audio.cpp include path, neither of which backend/cpp/run-unit-tests.sh
// provides, so it is built and run by ctest:
//
// make -C backend/cpp/audio-cpp test-engine
//
// The assertions on OPTION KEYS are the point of this file, not decoration.
// Every one of them names a key that was grepped against the pinned upstream:
// "instruct" is read, "instructions" is not; "duration_seconds" is read,
// "duration" is not. A rename that looks harmless is exactly the change that
// silently stops a family honouring the request, so the spellings are pinned
// here rather than left to a comment.
#include "generation_request.h"
#include <cstdio>
#include <filesystem>
#include <fstream>
#include <string>
#include <unordered_map>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
using namespace audiocpp_backend;
static bool has_key(const std::unordered_map<std::string, std::string> &options,
const std::string &key) {
return options.find(key) != options.end();
}
static std::string option_or(
const std::unordered_map<std::string, std::string> &options,
const std::string &key, const std::string &fallback) {
const auto it = options.find(key);
return it == options.end() ? fallback : it->second;
}
static engine::runtime::AudioBuffer clip(int sample_rate, int channels) {
engine::runtime::AudioBuffer buffer;
buffer.sample_rate = sample_rate;
buffer.channels = channels;
// Distinguishable content, so a builder that swapped one buffer for another
// (or default-constructed one) is visible rather than merely a size change.
buffer.samples = {0.25f, -0.5f, 0.75f, -1.0f};
return buffer;
}
static void test_voice_is_reference_file() {
const auto dir = std::filesystem::temp_directory_path() /
"audiocpp-generation-request-ctest";
std::filesystem::remove_all(dir);
std::filesystem::create_directories(dir);
const auto file = dir / "reference.wav";
{
std::ofstream out(file, std::ios::binary);
out << "not really a wav, but a regular file";
}
const auto subdir = dir / "a-directory";
std::filesystem::create_directories(subdir);
check(!voice_is_reference_file(""), "voice_is_reference_file: empty");
check(!voice_is_reference_file("alloy"),
"voice_is_reference_file: bare preset name");
check(!voice_is_reference_file((dir / "absent.wav").string()),
"voice_is_reference_file: missing path");
check(voice_is_reference_file(file.string()),
"voice_is_reference_file: existing regular file");
// A directory is NOT a reference. exists() would say yes here and the read
// would then fail with "cannot read <dir> as WAV", which sends the operator
// after a file problem instead of a preset typo.
check(!voice_is_reference_file(subdir.string()),
"voice_is_reference_file: directory is not a reference");
std::filesystem::remove_all(dir);
}
// build_tts_shape is what TTS and TTSStream both hand to routing, so a wrong
// answer here silently changes which task a request runs as, with a 200 and no
// diagnostic. Every field is asserted in both directions.
static void test_tts_shape() {
const auto dir = std::filesystem::temp_directory_path() /
"audiocpp-generation-request-ctest-shape";
std::filesystem::remove_all(dir);
std::filesystem::create_directories(dir);
const auto file = dir / "reference.wav";
{
std::ofstream out(file, std::ios::binary);
out << "a regular file";
}
{
backend::TTSRequest request;
request.set_text("hello");
const auto shape = build_tts_shape(request);
check(!shape.has_voice_reference && !shape.has_instructions,
"shape: bare request has neither signal");
// Never filled here: it comes off the LoadedModel, and leaving it empty
// is what makes the caller's assignment visible at the call site.
check(shape.pinned_task.empty(), "shape: pinned_task is left to the caller");
}
{
backend::TTSRequest request;
request.set_voice(file.string());
const auto shape = build_tts_shape(request);
check(shape.has_voice_reference,
"shape: an existing file is a voice reference");
check(!shape.has_instructions, "shape: a clip is not an instruction");
}
{
backend::TTSRequest request;
request.set_voice("alloy");
const auto shape = build_tts_shape(request);
check(!shape.has_voice_reference,
"shape: a preset name is not a voice reference");
}
{
backend::TTSRequest request;
request.set_voice(dir.string());
const auto shape = build_tts_shape(request);
check(!shape.has_voice_reference,
"shape: a directory is not a voice reference");
}
{
backend::TTSRequest request;
request.set_instructions("a calm older man");
const auto shape = build_tts_shape(request);
check(shape.has_instructions, "shape: instructions are seen");
check(!shape.has_voice_reference,
"shape: instructions do not imply a reference");
}
{
// The guard that has to match build_tts_request's. An empty
// instructions string builds no style condition, so telling routing to
// prefer VoiceDesign for it would route to a task with nothing to
// design from.
backend::TTSRequest request;
request.set_instructions("");
const auto shape = build_tts_shape(request);
check(!shape.has_instructions,
"shape: an empty instructions string is not an instruction");
}
{
backend::TTSRequest request;
request.set_voice(file.string());
request.set_instructions("a calm older man");
const auto shape = build_tts_shape(request);
check(shape.has_voice_reference && shape.has_instructions,
"shape: both signals are reported when both are set");
}
std::filesystem::remove_all(dir);
}
static void test_tts_plain() {
backend::TTSRequest request;
request.set_text("hello there");
const auto task = build_tts_request(request, std::nullopt);
check(task.text_input.has_value() && task.text_input->text == "hello there",
"tts: text reaches the transcript");
// No voice and no instructions means NO voice condition at all. A builder
// that always emitted one would make every family think a speaker was
// named, and chatterbox in particular refuses a prepare whose voice
// condition carries neither audio nor anything else it can use.
check(!task.voice.has_value(), "tts: no voice condition when nothing is set");
check(task.options.empty(), "tts: no options when nothing is set");
check(!task.audio_input.has_value(), "tts: no audio input");
}
static void test_tts_named_preset() {
backend::TTSRequest request;
request.set_text("hello");
request.set_voice("alloy");
const auto task = build_tts_request(request, std::nullopt);
check(task.voice.has_value() && task.voice->speaker.has_value(),
"tts preset: speaker condition present");
check(task.voice->speaker->cached_voice_id.has_value() &&
*task.voice->speaker->cached_voice_id == "alloy",
"tts preset: lands in cached_voice_id");
// The clip slot must stay empty, or a cloning family would try to prepare
// conditionals from a default-constructed buffer.
check(!task.voice->speaker->audio.has_value(),
"tts preset: no reference audio");
check(!task.voice->style.has_value(), "tts preset: no style condition");
check(option_or(task.options, "voice", "") == "alloy",
"tts preset: forwarded as the voice option too");
}
static void test_tts_reference_clip() {
backend::TTSRequest request;
request.set_text("hello");
request.set_voice("/tmp/reference.wav");
const auto task = build_tts_request(request, clip(44100, 2));
check(task.voice.has_value() && task.voice->speaker.has_value(),
"tts clip: speaker condition present");
check(task.voice->speaker->audio.has_value(),
"tts clip: reference audio present");
// Rate and channels survive untouched. This is the assertion that fails if
// anybody decides to fold the clip to 16 kHz mono on the way in.
check(task.voice->speaker->audio->sample_rate == 44100 &&
task.voice->speaker->audio->channels == 2 &&
task.voice->speaker->audio->samples.size() == 4,
"tts clip: rate, channels and samples pass through unchanged");
// A clip is NOT also a cached voice id, and the path must not travel as a
// preset name: a family reading cached_voice_id would then look up a voice
// called "/tmp/reference.wav".
check(!task.voice->speaker->cached_voice_id.has_value(),
"tts clip: no cached_voice_id");
check(!has_key(task.options, "voice"), "tts clip: no voice option");
}
static void test_tts_instructions() {
backend::TTSRequest request;
request.set_text("hello");
request.set_instructions("a calm older man, speaking slowly");
const auto task = build_tts_request(request, std::nullopt);
check(option_or(task.options, "instruct", "") ==
"a calm older man, speaking slowly",
"tts instructions: instruct option is the one qwen3_tts reads");
check(option_or(task.options, "caption", "") ==
"a calm older man, speaking slowly",
"tts instructions: caption option is the one irodori_tts reads");
check(option_or(task.options, "instructions", "") ==
"a calm older man, speaking slowly",
"tts instructions: proto field name forwarded as an alias");
check(task.voice.has_value() && task.voice->style.has_value(),
"tts instructions: style condition present");
// "instruct", not "instructions". omnivoice and qwen3_tts both look this tag
// up by that exact name and by no other.
check(option_or(task.voice->style->tags, "instruct", "") ==
"a calm older man, speaking slowly",
"tts instructions: style tag is spelled instruct");
check(!has_key(task.voice->style->tags, "instructions"),
"tts instructions: style tag is NOT spelled instructions");
// Instructions alone must not invent a speaker: has_voice_reference is what
// routing keys VoiceCloning off, and a speaker here would make a
// clone-capable family expect a clip it never received.
check(!task.voice->speaker.has_value(),
"tts instructions: no speaker without a voice");
}
static void test_tts_empty_instructions_are_not_instructions() {
backend::TTSRequest request;
request.set_text("hello");
request.set_instructions("");
const auto task = build_tts_request(request, std::nullopt);
// has_instructions() is true here, because the field was set. An empty
// string is still no instruction, and forwarding it would set an empty
// instruct option that qwen3_tts would prefer over its style tag fallback.
check(!has_key(task.options, "instruct"),
"tts: an empty instructions string sets no instruct option");
check(!task.voice.has_value(),
"tts: an empty instructions string sets no voice condition");
}
// THE EXACT SHAPE LocalAI PUTS ON THE WIRE. core/backend/tts.go's newTTSRequest
// sets Language: &language UNCONDITIONALLY, so has_language() is true on every
// request that ever reaches this backend, carrying an empty string whenever the
// caller named no language.
//
// An empty StyleCondition::language is not a harmless default. supertonic reads
// text_input->language behind a !empty() guard and then OVERRIDES it from
// style->language whenever that optional is engaged, with no guard at all
// (supertonic/session.cpp generation_options_from_request), so an empty style
// language replaces its "en" default (session.h) with "" and
// tokenizer_text.cpp's preprocess throws "invalid Supertonic language: ".
// Every /v1/audio/speech request carrying instructions and no language would be
// an INTERNAL. A plain request never sees it, because the style condition only
// exists when instructions are non-empty.
static void test_tts_empty_language_is_not_a_language() {
backend::TTSRequest request;
request.set_text("hello");
request.set_instructions("a calm older man");
request.set_language("");
const auto task = build_tts_request(request, std::nullopt);
check(task.voice.has_value() && task.voice->style.has_value(),
"tts empty language: the style condition still exists");
check(!task.voice->style->language.has_value(),
"tts empty language: style language is left unset, not set to empty");
// The option emission has always guarded on !empty(); this pins the two to
// the same rule so they cannot drift apart again.
check(!has_key(task.options, "language"),
"tts empty language: no language option");
check(task.text_input.has_value() && task.text_input->language.empty(),
"tts empty language: transcript language stays empty");
}
// The one shape routing treats specially: a clip outranks instructions, and the
// VoiceCondition then has to carry BOTH, because the family that wins is chosen
// on the clip but may still read the style tag.
static void test_tts_clip_and_instructions() {
backend::TTSRequest request;
request.set_text("hello");
request.set_voice("/tmp/reference.wav");
request.set_instructions("bright and fast");
request.set_language("en");
const auto task = build_tts_request(request, clip(22050, 1));
check(task.voice.has_value(), "tts clip+instructions: voice condition present");
check(task.voice->speaker.has_value() &&
task.voice->speaker->audio.has_value() &&
task.voice->speaker->audio->sample_rate == 22050,
"tts clip+instructions: speaker carries the clip");
check(!task.voice->speaker->cached_voice_id.has_value(),
"tts clip+instructions: the clip path is not also a preset id");
check(task.voice->style.has_value() &&
option_or(task.voice->style->tags, "instruct", "") == "bright and fast",
"tts clip+instructions: style carries the instruct tag");
check(task.voice->style->language.has_value() &&
*task.voice->style->language == "en",
"tts clip+instructions: a real language does reach the style condition");
check(option_or(task.options, "instruct", "") == "bright and fast",
"tts clip+instructions: instruct option still emitted");
check(!has_key(task.options, "voice"),
"tts clip+instructions: still no voice option for a clip");
}
static void test_tts_language_and_params() {
backend::TTSRequest request;
request.set_text("ciao");
request.set_language("it");
request.set_instructions("warm");
(*request.mutable_params())["exaggeration"] = "0.7";
// An explicit param must win over the value derived from instructions.
(*request.mutable_params())["caption"] = "explicitly chosen caption";
const auto task = build_tts_request(request, std::nullopt);
check(task.text_input.has_value() && task.text_input->language == "it",
"tts language: reaches the transcript");
check(option_or(task.options, "language", "") == "it",
"tts language: forwarded as an option alias");
check(task.voice.has_value() && task.voice->style.has_value() &&
task.voice->style->language.has_value() &&
*task.voice->style->language == "it",
"tts language: reaches the style condition");
check(option_or(task.options, "exaggeration", "") == "0.7",
"tts params: passed through verbatim");
check(option_or(task.options, "caption", "") == "explicitly chosen caption",
"tts params: an explicit param overrides the derived caption");
}
static void test_sound_generation_minimal() {
backend::SoundGenerationRequest request;
request.set_text("a distant thunderstorm");
const auto task = build_sound_generation_request(request, std::nullopt);
check(task.text_input.has_value() &&
task.text_input->text == "a distant thunderstorm",
"sound: text reaches the transcript");
// Unset optionals must emit NOTHING. Emitting a zero for an unset duration
// would make heartmula refuse the request ("duration_seconds must be
// positive") on a request that never mentioned a duration.
check(task.options.empty(), "sound: unset optionals emit no options");
check(!task.audio_input.has_value(), "sound: no audio input without src");
check(!task.voice.has_value(), "sound: no voice condition");
}
static void test_sound_generation_full() {
backend::SoundGenerationRequest request;
request.set_text("a slow blues in E");
request.set_duration(30.0f);
request.set_temperature(0.8f);
request.set_sample(false);
request.set_src_divisor(4);
request.set_think(true);
request.set_caption("smoky bar recording");
request.set_lyrics("first line\nsecond line");
request.set_bpm(72);
request.set_keyscale("E minor");
request.set_language("en");
request.set_timesignature("4/4");
request.set_instrumental(true);
const auto task = build_sound_generation_request(request, clip(48000, 2));
// duration_seconds is the key every AudioGeneration family actually reads;
// "duration" rides along as an alias. Both spellings are pinned so a
// "cleanup" that keeps only the proto's own name is a test failure and not
// a silent loss of the duration.
check(option_or(task.options, "duration_seconds", "").rfind("30.", 0) == 0,
"sound: duration lands as duration_seconds");
check(option_or(task.options, "duration", "").rfind("30.", 0) == 0,
"sound: duration also forwarded under its own name");
check(option_or(task.options, "temperature", "").rfind("0.8", 0) == 0,
"sound: temperature forwarded");
// Set to FALSE, so this also proves the key is written whenever the field is
// present rather than only when the value is truthy.
check(option_or(task.options, "do_sample", "") == "false",
"sound: sample=false is forwarded as do_sample=false");
check(option_or(task.options, "src_divisor", "") == "4",
"sound: src_divisor forwarded");
check(option_or(task.options, "thinking", "") == "true",
"sound: think lands as thinking, the key ace_step reads");
check(option_or(task.options, "think", "") == "true",
"sound: think also forwarded under its own name");
check(option_or(task.options, "caption", "") == "smoky bar recording",
"sound: caption forwarded");
check(option_or(task.options, "lyrics", "") == "first line\nsecond line",
"sound: lyrics forwarded");
check(option_or(task.options, "bpm", "") == "72", "sound: bpm forwarded");
check(option_or(task.options, "keyscale", "") == "E minor",
"sound: keyscale forwarded");
check(option_or(task.options, "timesignature", "") == "4/4",
"sound: timesignature forwarded");
check(option_or(task.options, "instrumental", "") == "true",
"sound: instrumental forwarded");
check(option_or(task.options, "language", "") == "en",
"sound: language forwarded as an option alias");
check(task.text_input->language == "en",
"sound: language reaches the transcript, which is what ace_step reads");
check(task.audio_input.has_value() &&
task.audio_input->sample_rate == 48000 &&
task.audio_input->channels == 2 &&
task.audio_input->samples.size() == 4,
"sound: src passes through at its own rate and channel count");
}
// ---------------------------------------------------------------------------
// apply_transform_text_input
//
// AudioTransform has no text field on the wire, so a text-conditioned route
// (vevo2's speech-to-speech) can only be reached if the text travels as a
// param and is unpacked into text_input. Every assertion below pins a spelling
// or a precedence that a family actually depends on, not a shape that merely
// looks tidy.
static void test_transform_text_absent() {
engine::runtime::TaskRequest task;
task.options["stem"] = "vocals";
check(!apply_transform_text_input(task),
"transform text: reports false when no text key is present");
check(!task.text_input.has_value(),
"transform text: a request with no text keeps text_input unset");
}
static void test_transform_text_canonical_key() {
engine::runtime::TaskRequest task;
task.options["target_text"] = "sing this line";
check(apply_transform_text_input(task), "transform text: target_text reports true");
check(task.text_input.has_value() && task.text_input->text == "sing this line",
"transform text: target_text becomes text_input.text");
check(has_key(task.options, "target_text"),
"transform text: target_text survives in options for families that read it there");
}
static void test_transform_text_alias_key() {
engine::runtime::TaskRequest task;
task.options["text"] = "say this instead";
check(apply_transform_text_input(task), "transform text: text alias reports true");
check(task.text_input.has_value() && task.text_input->text == "say this instead",
"transform text: the text alias becomes text_input.text");
}
static void test_transform_text_canonical_wins() {
engine::runtime::TaskRequest task;
task.options["target_text"] = "canonical";
task.options["text"] = "alias";
check(apply_transform_text_input(task), "transform text: both keys reports true");
check(task.text_input.has_value() && task.text_input->text == "canonical",
"transform text: target_text wins over text, not whichever hashed first");
}
static void test_transform_text_empty_is_not_a_text() {
engine::runtime::TaskRequest task;
task.options["target_text"] = "";
check(!apply_transform_text_input(task),
"transform text: an empty target_text reports false");
check(!task.text_input.has_value(),
"transform text: an empty target_text leaves text_input unset");
}
static void test_transform_text_empty_canonical_falls_through_to_alias() {
engine::runtime::TaskRequest task;
task.options["target_text"] = "";
task.options["text"] = "the real one";
check(apply_transform_text_input(task),
"transform text: an empty canonical key does not mask a usable alias");
check(task.text_input.has_value() && task.text_input->text == "the real one",
"transform text: the alias is used when the canonical key is empty");
}
static void test_transform_text_language_rides_along() {
engine::runtime::TaskRequest task;
task.options["target_text"] = "vocalise me";
task.options["language"] = "ja";
check(apply_transform_text_input(task), "transform text: text plus language reports true");
check(task.text_input.has_value() && task.text_input->language == "ja",
"transform text: language lands on the Transcript alongside the text");
check(has_key(task.options, "language"),
"transform text: language survives in options too");
}
static void test_transform_language_alone_is_not_a_text() {
engine::runtime::TaskRequest task;
task.options["language"] = "ja";
check(!apply_transform_text_input(task),
"transform text: a language with no text reports false");
check(!task.text_input.has_value(),
"transform text: a language alone must not route a separation request through text");
}
static void test_transform_text_preserves_other_inputs() {
engine::runtime::TaskRequest task;
engine::runtime::AudioBuffer audio;
audio.sample_rate = 44100;
audio.channels = 2;
audio.samples = {0.1f, 0.2f, 0.3f, 0.4f};
task.audio_input = audio;
task.options["target_text"] = "keep the audio";
check(apply_transform_text_input(task), "transform text: with audio present reports true");
check(task.audio_input.has_value() && task.audio_input->samples.size() == 4 &&
task.audio_input->sample_rate == 44100,
"transform text: the source audio is untouched");
}
int main() {
test_voice_is_reference_file();
test_tts_shape();
test_tts_plain();
test_tts_named_preset();
test_tts_reference_clip();
test_tts_instructions();
test_tts_empty_instructions_are_not_instructions();
test_tts_empty_language_is_not_a_language();
test_tts_clip_and_instructions();
test_tts_language_and_params();
test_sound_generation_minimal();
test_sound_generation_full();
test_transform_text_absent();
test_transform_text_canonical_key();
test_transform_text_alias_key();
test_transform_text_canonical_wins();
test_transform_text_empty_is_not_a_text();
test_transform_text_empty_canonical_falls_through_to_alias();
test_transform_text_language_rides_along();
test_transform_language_alone_is_not_a_text();
test_transform_text_preserves_other_inputs();
if (failures != 0) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all checks passed\n");
return 0;
}

View File

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,114 @@
#include "inference_lane.h"
#include <algorithm>
#include <chrono>
namespace audiocpp_backend {
std::int64_t monotonic_millis() {
return std::chrono::duration_cast<std::chrono::milliseconds>(
std::chrono::steady_clock::now().time_since_epoch())
.count();
}
int resolve_wait_budget_ms(int policy_ceiling_ms, int request_hint_ms) {
// Both sides normalise to 0 for "not specified", which is also the value
// that means unbounded on the way out, so the unspecified cases fall out of
// the arithmetic instead of needing their own branches.
const int ceiling = policy_ceiling_ms > 0 ? policy_ceiling_ms : 0;
const int hint = request_hint_ms > 0 ? request_hint_ms : 0;
if (ceiling == 0) {
return hint;
}
if (hint == 0) {
return ceiling;
}
// The hint can tighten the ceiling but never loosen it.
return std::min(ceiling, hint);
}
bool run_exceeds_budget(bool lane_occupied, std::int64_t run_started_ms,
std::int64_t now_ms, int budget_ms) {
if (!lane_occupied || budget_ms <= 0) {
return false;
}
return now_ms - run_started_ms > static_cast<std::int64_t>(budget_ms);
}
namespace {
std::string busy_prefix(const std::string &model_label) {
return "inference lane for model '" + model_label + "' is busy: ";
}
// States the measurement and nothing else. A short-budget caller meeting a
// legitimately long run lands here too, so this text must not declare the run
// broken; the numbers let a reader decide that for themselves.
std::string overrun_message(const std::string &model_label,
std::int64_t run_age_ms, int budget_ms) {
return busy_prefix(model_label) + "the in-flight run has been running for " +
std::to_string(run_age_ms) + " ms, longer than this request's " +
std::to_string(budget_ms) + " ms wait budget";
}
std::string wait_exhausted_message(const std::string &model_label,
int budget_ms) {
return busy_prefix(model_label) + "timed out after " +
std::to_string(budget_ms) +
" ms waiting for the in-flight run to finish";
}
} // namespace
void InferenceLane::occupy(int budget_ms) {
std::unique_lock<std::mutex> lock(state_mutex_);
// Checked once, on arrival: if the run already in the lane has outlived what
// this caller brought, no amount of waiting can help it, and queueing here
// is exactly how a wedged run swallows every handler thread.
const std::int64_t arrived_ms = monotonic_millis();
if (run_exceeds_budget(occupied_, run_started_ms_, arrived_ms, budget_ms)) {
throw LaneUnavailable(overrun_message(
model_label_, arrived_ms - run_started_ms_, budget_ms));
}
if (budget_ms > 0) {
if (!vacated_.wait_for(lock, std::chrono::milliseconds(budget_ms),
[this] { return !occupied_; })) {
throw LaneUnavailable(
wait_exhausted_message(model_label_, budget_ms));
}
} else {
vacated_.wait(lock, [this] { return !occupied_; });
}
// Only now, holding both the mutex and the lane. Stamping any earlier would
// restart the age of the run for every caller behind this one and make a
// genuinely wedged holder look freshly started forever.
occupied_ = true;
run_started_ms_ = monotonic_millis();
}
void InferenceLane::vacate() {
{
std::lock_guard<std::mutex> lock(state_mutex_);
occupied_ = false;
}
// notify_all, not notify_one: a waiter that times out concurrently with a
// notification can consume it, and losing the only wakeup would park the
// remaining waiters for the rest of the run's lifetime. Waking all of them
// still admits exactly one, since the rest re-test occupancy under the mutex
// and go back to waiting, and ordering among waiters is not a requirement.
vacated_.notify_all();
}
LaneEntry::LaneEntry(InferenceLane &lane, int budget_ms) : lane_(lane) {
// If this throws, the object never existed, so ~LaneEntry does not run and
// cannot hand back a lane this caller never held.
lane_.occupy(budget_ms);
}
LaneEntry::~LaneEntry() { lane_.vacate(); }
} // namespace audiocpp_backend

View File

@@ -0,0 +1,138 @@
#pragma once
// Serializes inference against the single audio.cpp model a backend process
// owns. Standard library only.
//
// Why a plain mutex is not enough: an audio.cpp session is not reentrant, so
// concurrent gRPC handlers have to take turns. But once a GPU call stops making
// progress there is nothing a host thread can do to take it back, and an
// unbounded queue behind such a run would absorb the handler threads one by one
// until nothing is left to answer with. A caller therefore needs to be able to
// walk away, and needs to be able to tell "the lane is busy with normal work and
// I ran out of patience" apart from "the run in the lane has already outlived
// the patience I brought".
#include <condition_variable>
#include <cstdint>
#include <mutex>
#include <stdexcept>
#include <string>
namespace audiocpp_backend {
// Reading off a monotonic clock, in milliseconds. Monotonic on purpose: a wall
// clock adjustment must never make an in-flight run look younger or older than
// it is, because that reading decides whether callers give up.
std::int64_t monotonic_millis();
// Collapses the per-model configured ceiling and the optional per-request hint
// into the wait budget a caller actually gets. Returns 0 for "wait
// indefinitely".
//
// Either input may be non-positive, which means "not specified":
// - an unspecified hint yields the ceiling,
// - an unspecified ceiling means no policy limit, so the hint stands,
// - unspecified on both sides is unbounded.
// A specified hint may only tighten the ceiling. A client asking for a longer
// wait than the model's policy allows does not get it, because that would let a
// request weaken an operator's choice.
int resolve_wait_budget_ms(int policy_ceiling_ms, int request_hint_ms);
// True when a caller carrying budget_ms should give up on arrival rather than
// queue up. Deliberately a pure function of the lane's observable state so the
// decision can be tested without threads or sleeping.
//
// Only occupancy makes a start timestamp meaningful: a lane nobody holds is
// never overrunning, whatever timestamp the last holder left behind. An
// unbounded caller (non-positive budget) has no budget to exceed. And the
// comparison is strict, so a run whose age exactly equals the budget still has
// its last millisecond.
bool run_exceeds_budget(bool lane_occupied, std::int64_t run_started_ms,
std::int64_t now_ms, int budget_ms);
// Thrown when a caller cannot take the lane, in either of the two situations
// resolve_wait_budget_ms allows for. The message distinguishes them; callers
// that need to report a status code can treat them alike.
class LaneUnavailable : public std::runtime_error {
public:
explicit LaneUnavailable(const std::string &reason)
: std::runtime_error(reason) {}
};
class LaneEntry;
// One lane per loaded model. Shared by every handler thread; not copyable.
class InferenceLane {
public:
explicit InferenceLane(std::string model_label)
: model_label_(std::move(model_label)) {}
InferenceLane(const InferenceLane &) = delete;
InferenceLane &operator=(const InferenceLane &) = delete;
const std::string &model_label() const { return model_label_; }
private:
// Occupancy is only reachable through LaneEntry, so there is no way to take
// the lane without also having something that gives it back.
friend class LaneEntry;
void occupy(int budget_ms);
void vacate();
const std::string model_label_;
std::mutex state_mutex_;
std::condition_variable vacated_;
bool occupied_ = false;
// Only meaningful while occupied_ is true.
std::int64_t run_started_ms_ = 0;
};
// Scoped occupancy of a lane. Construct it where the inference happens and it
// is given back on every exit from that scope, including an exception and
// including a caller that returns from the middle of a long stream. Throws
// LaneUnavailable if the lane could not be taken, in which case there is no
// object and nothing to release.
//
// Not reentrant, and it does not detect reentrancy: a second entry constructed
// while the calling thread already holds the same lane waits for a lane only
// that thread can release. With a positive budget that surfaces as
// LaneUnavailable, but in unbounded mode the thread parks with no diagnostic at
// all. Keep entries one per call: a handler that holds one across a stream must
// not let a helper it calls construct another.
class LaneEntry {
public:
// budget_ms <= 0 waits indefinitely. Pass the output of
// resolve_wait_budget_ms.
LaneEntry(InferenceLane &lane, int budget_ms);
~LaneEntry();
LaneEntry(const LaneEntry &) = delete;
LaneEntry &operator=(const LaneEntry &) = delete;
// Deliberately immovable rather than carefully movable. A moved-from entry
// would have to stop releasing the lane while the lane still records it as
// occupied, and that hazard is not worth the convenience: the lane can only
// be recovered by whoever took it.
//
// To hold a lane for longer than one scope, construct the entry in place
// instead of moving one in. Two shapes work:
// std::optional<LaneEntry> held; // member or local
// held.emplace(lane, budget_ms); // takes the lane, held.reset() gives it back
// auto held = std::make_unique<LaneEntry>(lane, budget_ms); // also returnable
// Both outlive the acquiring scope and still release exactly once, when they
// are reset or destroyed. Prefer the optional for a member whose lifetime is
// the handler's; use the unique_ptr when the entry has to be returned, since
// an optional of an immovable type is itself immovable and cannot be. A
// factory may instead write `return LaneEntry(lane, budget_ms);`, which C++17
// guarantees to elide, whereas `LaneEntry entry(...); return entry;` does not
// compile, because that form is a move.
LaneEntry(LaneEntry &&) = delete;
LaneEntry &operator=(LaneEntry &&) = delete;
private:
InferenceLane &lane_;
};
} // namespace audiocpp_backend

View File

@@ -0,0 +1,618 @@
// Unit tests for inference_lane. Standard library only. The harness compiles
// this file as a single translation unit, so the implementation is included
// directly rather than linked.
//
// Two kinds of test live here:
//
// * The pure ones (budget negotiation, the overrun predicate) run with no
// threads and no sleeping. They carry the arithmetic, so they are the tests
// that must be exhaustive.
// * The threaded ones exercise the lane itself. Every one of them is bounded:
// contenders use a generous wait budget instead of the unbounded mode
// wherever the point of the test does not require unbounded, and a watchdog
// in main() puts a ceiling on the whole file. A broken implementation must
// go red, not hang, because a hung job costs CI far more than a red one.
//
// Wall-clock margins are called out individually. The rule applied throughout:
// a margin is only allowed if a slow or loaded machine pushes the measurement
// deeper into the passing region.
#include "inference_lane.cpp"
#include <atomic>
#include <chrono>
#include <condition_variable>
#include <cstdio>
#include <cstdlib>
#include <mutex>
#include <string>
#include <thread>
#include <vector>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
using audiocpp_backend::InferenceLane;
using audiocpp_backend::LaneEntry;
using audiocpp_backend::LaneUnavailable;
using audiocpp_backend::resolve_wait_budget_ms;
using audiocpp_backend::run_exceeds_budget;
// A one-shot, level-triggered signal with a bounded wait. Preferred over sleeps
// for "the other thread got there" so the tests do not encode a guess about
// scheduling.
class Signal {
public:
void raise() {
{
std::lock_guard<std::mutex> lock(mutex_);
raised_ = true;
}
cv_.notify_all();
}
bool await(int timeout_ms) {
std::unique_lock<std::mutex> lock(mutex_);
return cv_.wait_for(lock, std::chrono::milliseconds(timeout_ms),
[this] { return raised_; });
}
private:
std::mutex mutex_;
std::condition_variable cv_;
bool raised_ = false;
};
static std::int64_t elapsed_ms_since(
const std::chrono::steady_clock::time_point &start) {
return std::chrono::duration_cast<std::chrono::milliseconds>(
std::chrono::steady_clock::now() - start)
.count();
}
static void nap(int ms) {
std::this_thread::sleep_for(std::chrono::milliseconds(ms));
}
static bool mentions(const std::string &haystack, const std::string &needle) {
return haystack.find(needle) != std::string::npos;
}
// Wording that separates the two failure modes. Kept here so a message reword
// that erases the distinction breaks these tests loudly.
static const char *const kTimedOutPhrase = "timed out after";
static const char *const kStillRunningPhrase = "has been running for";
// Attempts an entry and reports what happened, so the threaded tests can assert
// on the message rather than only on the exception type.
struct EntryOutcome {
bool acquired = false;
std::string message;
std::int64_t took_ms = 0;
};
static EntryOutcome try_entry(InferenceLane &lane, int budget_ms) {
EntryOutcome out;
const auto start = std::chrono::steady_clock::now();
try {
LaneEntry entry(lane, budget_ms);
out.acquired = true;
} catch (const LaneUnavailable &refused) {
out.message = refused.what();
}
out.took_ms = elapsed_ms_since(start);
return out;
}
// ---------------------------------------------------------------------------
// B9: budget negotiation. Pure, no threads.
// ---------------------------------------------------------------------------
static void test_budget_negotiation() {
check(resolve_wait_budget_ms(0, 0) == 0,
"B9 no ceiling and no request hint is unbounded");
check(resolve_wait_budget_ms(-1, -1) == 0,
"B9 negative ceiling and negative hint is unbounded");
check(resolve_wait_budget_ms(5000, 0) == 5000,
"B9 absent hint yields the ceiling");
check(resolve_wait_budget_ms(5000, -250) == 5000,
"B9 negative hint yields the ceiling");
check(resolve_wait_budget_ms(5000, 1200) == 1200,
"B9 a shorter request hint is granted");
check(resolve_wait_budget_ms(5000, 1) == 1,
"B9 a much shorter request hint is granted");
check(resolve_wait_budget_ms(5000, 9000) == 5000,
"B9 a longer request hint cannot weaken the ceiling");
check(resolve_wait_budget_ms(5000, 5001) == 5000,
"B9 a hint one ms over the ceiling is clamped");
check(resolve_wait_budget_ms(5000, 5000) == 5000,
"B9 a hint equal to the ceiling is the ceiling");
check(resolve_wait_budget_ms(0, 1200) == 1200,
"B9 without a policy limit the request hint applies");
check(resolve_wait_budget_ms(-5, 1200) == 1200,
"B9 a negative ceiling is no policy limit");
check(resolve_wait_budget_ms(1, 0) == 1,
"B9 a one ms ceiling survives negotiation");
}
// ---------------------------------------------------------------------------
// B5 and B6: the overrun predicate. Pure, no threads.
// ---------------------------------------------------------------------------
static void test_overrun_predicate() {
// B5: strictly longer.
check(run_exceeds_budget(true, 1000, 1100, 100) == false,
"B5 elapsed exactly equal to the budget is not an overrun");
check(run_exceeds_budget(true, 1000, 1101, 100) == true,
"B5 one ms past the budget is an overrun");
check(run_exceeds_budget(true, 1000, 1099, 100) == false,
"B5 one ms short of the budget is not an overrun");
check(run_exceeds_budget(true, 0, 1, 1) == false,
"B5 a one ms budget at one ms elapsed is not an overrun");
check(run_exceeds_budget(true, 0, 2, 1) == true,
"B5 a one ms budget at two ms elapsed is an overrun");
// B6: an unoccupied lane is never stuck, whatever the leftover timestamp
// says. This is the pure half of B6; the wiring half is threaded below.
check(run_exceeds_budget(false, 0, 10000000, 1) == false,
"B6 an idle lane with an ancient start timestamp is not an overrun");
check(run_exceeds_budget(false, 5, 5, 5) == false,
"B6 an idle lane is not an overrun at any elapsed value");
// An unbounded caller has no budget to exceed, so it never fails fast.
check(run_exceeds_budget(true, 0, 10000000, 0) == false,
"B2 an unbounded caller never sees an overrun");
check(run_exceeds_budget(true, 0, 10000000, -1) == false,
"B2 a negative budget never sees an overrun");
}
// ---------------------------------------------------------------------------
// B1: mutual exclusion under real contention, and a release admits a waiter.
// ---------------------------------------------------------------------------
static void test_mutual_exclusion() {
InferenceLane lane("exclusion-model");
constexpr int kContenders = 4; // "at least three simultaneous contenders"
constexpr int kHoldMs = 15;
// Generous on purpose: the point of this test is exclusion, not timeouts.
// A larger budget only makes a healthy run more likely to pass, while still
// bounding a broken one at roughly five seconds instead of forever.
constexpr int kBudgetMs = 5000;
std::atomic<int> in_flight{0};
std::atomic<int> peak_in_flight{0};
std::atomic<int> completed{0};
std::atomic<int> refused{0};
Signal go;
std::vector<std::thread> contenders;
for (int i = 0; i < kContenders; i++) {
contenders.emplace_back([&] {
go.await(5000);
try {
LaneEntry entry(lane, kBudgetMs);
const int now_inside = in_flight.fetch_add(1) + 1;
int seen = peak_in_flight.load();
while (now_inside > seen &&
!peak_in_flight.compare_exchange_weak(seen, now_inside)) {
// retry with the refreshed value
}
nap(kHoldMs);
in_flight.fetch_sub(1);
completed.fetch_add(1);
} catch (const LaneUnavailable &) {
refused.fetch_add(1);
}
});
}
go.raise();
for (auto &t : contenders) {
t.join();
}
check(refused.load() == 0, "B1 no contender was refused within its budget");
check(completed.load() == kContenders,
"B1 every contender eventually got the lane");
check(peak_in_flight.load() == 1,
"B1 never more than one holder inside the lane at once");
}
// ---------------------------------------------------------------------------
// B2: unbounded mode waits out a run longer than any bound would allow.
// ---------------------------------------------------------------------------
static void test_unbounded_waits_out_the_holder() {
InferenceLane lane("patient-model");
constexpr int kHoldMs = 300;
Signal held;
Signal patient_done;
std::int64_t patient_wait_ms = -1;
bool patient_acquired = false;
std::thread holder([&] {
LaneEntry entry(lane, 0);
held.raise();
nap(kHoldMs);
});
check(held.await(5000), "B2 holder took the lane");
std::thread patient([&] {
const auto start = std::chrono::steady_clock::now();
try {
LaneEntry entry(lane, 0);
patient_acquired = true;
} catch (const LaneUnavailable &) {
patient_acquired = false;
}
patient_wait_ms = elapsed_ms_since(start);
patient_done.raise();
});
// Bound on the unbounded mode: a lane that never wakes its waiters goes red
// here instead of hanging in join(). The named failure is printed before the
// join, so even a hard hang leaves a diagnosis behind for the watchdog.
check(patient_done.await(10000), "B2 unbounded caller returned at all");
patient.join();
holder.join();
check(patient_acquired, "B2 unbounded caller acquired instead of failing");
// Margin: the holder holds for 300 ms, so the waiter must block for about
// that long. Asserting only half of it means a loaded machine, which makes
// the wait longer, drifts further into passing.
check(patient_wait_ms >= kHoldMs / 2,
"B2 unbounded caller actually waited for the in-flight run");
}
// ---------------------------------------------------------------------------
// B3: a bounded caller that cannot get in gives up with the timeout wording.
// ---------------------------------------------------------------------------
static void test_bounded_wait_times_out() {
InferenceLane lane("impatient-model");
constexpr int kBudgetMs = 120;
Signal held;
Signal release;
std::thread holder([&] {
LaneEntry entry(lane, 0);
held.raise();
release.await(10000);
});
check(held.await(5000), "B3 holder took the lane");
// The waiter arrives immediately, so the holder's elapsed time is far below
// the budget and the fail-fast path must not trigger here.
//
// Load-sensitive margin, and the tightest one in this file: what makes this
// the timeout path rather than the fail-fast path is the holder's age
// staying under 120 ms at the arrival check. All that sits between the
// holder's stamp and this call is one signal handover, microseconds against
// a 120 ms allowance, but unlike the other margins here load pushes this one
// toward failing rather than away from it. If it ever does flip, the symptom
// is the wording assertions below going red, not a hang, and the fix is a
// larger budget rather than a weaker assertion.
const EntryOutcome outcome = try_entry(lane, kBudgetMs);
release.raise();
holder.join();
check(!outcome.acquired, "B3 bounded caller did not acquire a held lane");
check(mentions(outcome.message, kTimedOutPhrase),
"B3 failure names the exhausted wait, not an overrunning run");
check(!mentions(outcome.message, kStillRunningPhrase),
"B3 failure is not worded as an overrun");
check(mentions(outcome.message, "impatient-model"),
"B3 failure names the model");
check(mentions(outcome.message, "120"),
"B3 failure reports the budget it waited out");
// Margin: wait_for cannot return before its deadline, so the true value is
// at least 120 ms and load only raises it. Asserting 100 leaves room for
// clock granularity while still catching an implementation that returns
// early without waiting.
check(outcome.took_ms >= 100, "B3 bounded caller waited out its budget");
}
// ---------------------------------------------------------------------------
// B4: a caller whose budget is already exceeded fails at once.
// ---------------------------------------------------------------------------
static void test_fail_fast_against_a_long_run() {
InferenceLane lane("wedged-model");
constexpr int kBudgetMs = 200;
constexpr int kRunAgeMs = 400;
Signal held;
Signal release;
std::thread holder([&] {
LaneEntry entry(lane, 0);
held.raise();
release.await(10000);
});
check(held.await(5000), "B4 holder took the lane");
// Margin: the arriving caller needs the holder's elapsed time to exceed
// 200 ms. Sleeping 400 ms means a loaded machine oversleeps and pushes the
// elapsed time further past the budget, never below it.
nap(kRunAgeMs);
const EntryOutcome outcome = try_entry(lane, kBudgetMs);
release.raise();
holder.join();
check(!outcome.acquired, "B4 caller did not acquire a long-running lane");
check(mentions(outcome.message, kStillRunningPhrase),
"B4 failure states the measured age of the in-flight run");
check(!mentions(outcome.message, kTimedOutPhrase),
"B4 failure is not worded as an exhausted wait");
check(mentions(outcome.message, "wedged-model"),
"B4 failure names the model");
// Margin: a fail-fast return takes microseconds. 150 ms of headroom under a
// 200 ms budget separates "returned at once" from "waited out the budget"
// by a wide enough gap that scheduler noise cannot close it. The message
// assertions above are the load-independent proof; this one pins the timing.
check(outcome.took_ms < 150, "B4 caller failed without waiting out its budget");
}
// ---------------------------------------------------------------------------
// B6 wiring: an idle lane never looks stuck, however old the last run is.
// ---------------------------------------------------------------------------
static void test_idle_lane_is_never_stuck() {
InferenceLane lane("idle-model");
{
LaneEntry entry(lane, 0);
}
// Ages the leftover start timestamp well past the tiny budget used below.
// A longer sleep only makes a stale-timestamp bug more visible, so load
// helps this test rather than hurting it.
nap(80);
const EntryOutcome first = try_entry(lane, 20);
check(first.acquired, "B6 tiny budget still acquires an idle lane");
nap(80);
const EntryOutcome second = try_entry(lane, 1);
check(second.acquired, "B6 a one ms budget still acquires an idle lane");
}
// ---------------------------------------------------------------------------
// B7: every ownership exit clears the busy state, including an exception
// thrown from inside the guarded region.
// ---------------------------------------------------------------------------
static void test_release_on_exception() {
InferenceLane lane("throwing-model");
struct GuardedRegionFailure {};
bool propagated = false;
try {
LaneEntry entry(lane, 0);
throw GuardedRegionFailure{};
} catch (const GuardedRegionFailure &) {
propagated = true;
}
check(propagated, "B7 the guarded region's own exception propagated");
// If the throw had leaked the busy state, this tiny budget would fail.
const EntryOutcome after_throw = try_entry(lane, 20);
check(after_throw.acquired, "B7 lane is free after an exception unwound it");
// Same check for a holder that unwinds on another thread, which is the shape
// a gRPC handler failing mid-inference actually has.
Signal thrown;
std::thread unlucky([&] {
try {
LaneEntry entry(lane, 0);
throw GuardedRegionFailure{};
} catch (const GuardedRegionFailure &) {
thrown.raise();
}
});
check(thrown.await(5000), "B7 worker thread unwound its guarded region");
unlucky.join();
const EntryOutcome after_worker = try_entry(lane, 20);
check(after_worker.acquired,
"B7 lane is free after a worker thread unwound it");
}
// ---------------------------------------------------------------------------
// B8: a waiter must not publish itself as the holder. If it did, its arrival
// would restart the elapsed-time measurement and hide the real holder.
//
// Timeline, with the holder taking the lane at t0 and never letting go:
//
// t0 holder acquires, elapsed measurement starts here and only here
// t0+100 waiter arrives with a 400 ms budget, blocks, and times out
// t0+500 late caller arrives with a 450 ms budget
//
// A correct lane measures 500 ms of holding at the late arrival, which is more
// than 450, so the late caller fails fast. An implementation that let the
// waiter stamp itself as holder measures only the 400 ms since the waiter
// arrived, which is under 450, so the late caller would queue behind a stuck
// run instead. The two paths are told apart by their wording.
// ---------------------------------------------------------------------------
static void test_waiter_does_not_become_the_holder() {
InferenceLane lane("stamp-model");
constexpr int kWaiterArrivesAfterMs = 100;
constexpr int kWaiterBudgetMs = 400;
constexpr int kLateBudgetMs = 450;
Signal held;
Signal release;
std::thread holder([&] {
LaneEntry entry(lane, 0);
held.raise();
release.await(10000);
});
check(held.await(5000), "B8 holder took the lane");
// Margin: the waiter must not fail fast on arrival, which needs the
// holder's elapsed time to stay under 400 ms. Arriving at 100 ms leaves
// 300 ms of slack, so oversleeping under load does not flip the path.
nap(kWaiterArrivesAfterMs);
EntryOutcome waiter_outcome;
std::thread waiter([&] { waiter_outcome = try_entry(lane, kWaiterBudgetMs); });
waiter.join();
check(!waiter_outcome.acquired, "B8 mid-queue waiter did not acquire");
check(mentions(waiter_outcome.message, kTimedOutPhrase),
"B8 mid-queue waiter waited out its budget and timed out");
// Margin: the holder has now been in the lane for at least 500 ms against a
// 450 ms budget. Load lengthens both sleeps, so the measured age only grows
// and the fail-fast path only becomes more certain.
const EntryOutcome late = try_entry(lane, kLateBudgetMs);
release.raise();
holder.join();
check(!late.acquired, "B8 late caller did not acquire");
check(mentions(late.message, kStillRunningPhrase),
"B8 elapsed time is still measured from the real holder's acquisition");
check(late.took_ms < 200,
"B8 late caller failed fast rather than queueing behind the holder");
}
// ---------------------------------------------------------------------------
// B8, other direction: the age of a run is measured from the moment its holder
// acquired, not from the moment that holder arrived. A caller that queued for a
// while and then got in is starting a fresh run, and its time in the queue must
// not be billed to it: if it were, every handover would hand the new holder a
// head start towards looking overrun, and short-budget callers would be turned
// away from a run that has barely begun.
//
// t0 first holder acquires and holds for 300 ms
// t0 second caller arrives and queues
// t0+300 second caller acquires, so its own run age is ~0 here
// t0+300 a third caller arrives with a 300 ms budget
//
// Correct: the third caller sees a run that just started, so it queues and then
// times out. Billing the queue time to the second caller would show a 300 ms old
// run instead, and the third caller would be turned away as an overrun.
// ---------------------------------------------------------------------------
static void test_run_age_starts_at_acquisition() {
InferenceLane lane("handover-model");
constexpr int kFirstHoldMs = 300;
constexpr int kThirdBudgetMs = 200;
Signal first_held;
Signal handed_over;
Signal release_second;
std::thread first([&] {
LaneEntry entry(lane, 0);
first_held.raise();
nap(kFirstHoldMs);
});
// Ordering matters only for the diagnosis, not for the assertion: waiting
// for the first holder guarantees the second caller really does queue, which
// is what gives it queue time to be wrongly billed for.
check(first_held.await(5000), "B8 first holder took the lane");
std::thread second([&] {
LaneEntry entry(lane, 0);
handed_over.raise();
release_second.await(10000);
});
check(handed_over.await(10000), "B8 queued caller was handed the lane");
// Margin: the new holder's run is a few ms old against a 200 ms budget, so
// this caller must queue. Load can only add a few ms of handover latency,
// well inside that slack, while it lengthens the queue time that the buggy
// version would bill, making the bug more visible rather than less.
const EntryOutcome third = try_entry(lane, kThirdBudgetMs);
release_second.raise();
second.join();
first.join();
check(!third.acquired, "B8 lane was still held by the queued caller");
check(mentions(third.message, kTimedOutPhrase),
"B8 a fresh holder's run age excludes the time it spent queueing");
}
// ---------------------------------------------------------------------------
// B10: both failure modes carry a usable message, and the fail-fast one reports
// a measurement rather than diagnosing a cause.
// ---------------------------------------------------------------------------
static void test_failure_messages_are_diagnosable() {
InferenceLane lane("diagnosable-model");
Signal held;
Signal release;
std::thread holder([&] {
LaneEntry entry(lane, 0);
held.raise();
release.await(10000);
});
check(held.await(5000), "B10 holder took the lane");
// The two budgets have to straddle the run's age, or both callers take the
// same path and the comparisons below are between two fail-fast messages
// that differ only in the budget they print.
//
// Margin: 30 ms is far under the ~120 ms age, and load only ages the run
// further, so `fast` fails fast. 400 ms is far over it, with the same
// slack and the same safe direction as the B3 test, so `slow` queues and
// then times out.
nap(120);
const EntryOutcome fast = try_entry(lane, 30);
const EntryOutcome slow = try_entry(lane, 400);
release.raise();
holder.join();
check(mentions(fast.message, kStillRunningPhrase) &&
mentions(slow.message, kTimedOutPhrase),
"B10 one caller took the fail-fast path and the other timed out");
check(fast.message != slow.message,
"B10 the two failure modes do not share one message");
check(!fast.message.empty() && !slow.message.empty(),
"B10 both failures carry text");
check(mentions(fast.message, "diagnosable-model") &&
mentions(slow.message, "diagnosable-model"),
"B10 both failures name the model");
// The fail-fast wording must not accuse the run of being stuck: a short
// budget meeting a legitimately long run reaches this path too.
for (const char *verdict : {"stuck", "wedged", "hung", "deadlock"}) {
check(!mentions(fast.message, verdict),
std::string("B10 fail-fast message avoids diagnosing '") +
verdict + "'");
}
}
int main() {
// Last resort only. Every threaded test above is individually bounded, so
// this should never fire; it exists so that an implementation which parks a
// thread forever still ends the job instead of occupying a CI runner.
std::thread watchdog([] {
std::this_thread::sleep_for(std::chrono::seconds(30));
fprintf(stderr, "FAIL: watchdog fired, an inference_lane test hung\n");
fflush(stderr);
std::_Exit(1);
});
watchdog.detach();
test_budget_negotiation();
test_overrun_predicate();
test_mutual_exclusion();
test_unbounded_waits_out_the_holder();
test_bounded_wait_times_out();
test_fail_fast_against_a_long_run();
test_idle_lane_is_never_stuck();
test_release_on_exception();
test_waiter_does_not_become_the_holder();
test_run_age_starts_at_acquisition();
test_failure_messages_are_diagnosable();
if (failures == 0) {
fprintf(stderr, "\nAll inference_lane tests passed.\n");
return 0;
}
fprintf(stderr, "\n%d inference_lane test(s) failed.\n", failures);
return 1;
}

View File

@@ -0,0 +1,75 @@
#include "live_watchdog.h"
#include <utility>
namespace audiocpp_backend {
IdleWatchdog::IdleWatchdog(std::chrono::milliseconds window,
std::function<void()> on_idle)
: window_(window), on_idle_(std::move(on_idle)),
last_(std::chrono::steady_clock::now()) {
if (window_.count() <= 0) {
// Disabled: no thread at all, rather than a thread with an infinite
// deadline. A thread that exists is a thread that has to be joined on
// every exit path, and there is nothing for this one to do.
return;
}
thread_ = std::thread([this] { run(); });
}
IdleWatchdog::~IdleWatchdog() { disarm(); }
void IdleWatchdog::touch() {
std::lock_guard<std::mutex> lock(mu_);
last_ = std::chrono::steady_clock::now();
// Deliberately does NOT notify. The waiter recomputes its deadline from
// last_ every time it wakes, so a touch that lands mid-window is picked up
// when the old deadline expires, and a touch is the hot path: it runs once
// per frame on the wire.
}
void IdleWatchdog::disarm() {
{
std::lock_guard<std::mutex> lock(mu_);
stop_ = true;
}
cv_.notify_all();
if (thread_.joinable()) {
thread_.join();
}
}
bool IdleWatchdog::fired() const {
std::lock_guard<std::mutex> lock(mu_);
return fired_;
}
void IdleWatchdog::run() {
std::unique_lock<std::mutex> lock(mu_);
while (!stop_) {
const auto deadline = last_ + window_;
if (cv_.wait_until(lock, deadline, [this] { return stop_; })) {
return; // disarmed
}
// The deadline passed, but last_ may have moved while this thread was
// waiting, and a condition variable may also wake spuriously. Re-read
// it: without this check a touch that landed mid-window would still be
// followed by a cancellation, i.e. a live client cut off mid-sentence.
if (std::chrono::steady_clock::now() < last_ + window_) {
continue;
}
fired_ = true;
auto callback = on_idle_;
lock.unlock();
if (callback) {
callback();
}
return; // one shot
}
}
bool live_frame_carries_audio(bool has_audio, bool pcm_empty) {
return has_audio && !pcm_empty;
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,99 @@
#pragma once
// A one-shot idle timer for a bidirectional stream. Standard library only, so
// it is tested without an audio.cpp checkout or a gRPC server.
//
// WHY IT EXISTS. AudioTranscriptionLive holds the model's inference lane for the
// whole stream, because the streaming session is stateful and a concurrent run
// would interleave two callers' audio. Every other RPC in this backend holds the
// lane across COMPUTE, or across a write to a slow reader, and both of those
// terminate on their own. A live stream instead blocks in a client-driven read,
// and a peer that goes silent WITHOUT closing the stream never terminates
// anything: the lane stays taken and every other RPC against that model queues
// behind a client that has stopped speaking. A websocket death does cancel the
// RPC and free it, but "the peer's TCP connection eventually dies" is not a
// bound anyone can state, so this supplies one.
//
// HOW IT ENDS THE STREAM, and the part that is not obvious: gRPC's synchronous
// ServerReaderWriter::Read has no timeout and cannot be given one. The only way
// to unblock it from another thread is ServerContext::TryCancel, which is what
// the callback is for. That means the client sees CANCELLED rather than whatever
// status the handler goes on to return: the returned status is for the server's
// own record. Releasing the lane is the point.
//
// ONE SHOT on purpose. Once the callback has run the stream is being torn down,
// so there is nothing left to watch, and a repeating timer would call TryCancel
// on a context the handler may already have returned from.
#include <chrono>
#include <condition_variable>
#include <functional>
#include <mutex>
#include <thread>
namespace audiocpp_backend {
class IdleWatchdog {
public:
// A window that is not positive DISABLES the watchdog entirely: no thread is
// started and fired() never becomes true. That is the operator's escape
// hatch for a client that legitimately holds a stream open through long
// pauses, and it is why the option carrying it documents 0 as "no limit"
// rather than as "expire immediately".
//
// `on_idle` runs on the watchdog's own thread with no lock held. It must be
// safe to call while the watched thread is blocked in a read, which is the
// only reason this class exists; ServerContext::TryCancel is documented as
// exactly that.
IdleWatchdog(std::chrono::milliseconds window, std::function<void()> on_idle);
// Joins the thread, so the callback can safely capture anything that
// outlives this object's scope and nothing else has to be reasoned about.
~IdleWatchdog();
IdleWatchdog(const IdleWatchdog &) = delete;
IdleWatchdog &operator=(const IdleWatchdog &) = delete;
// Restarts the window. Call it whenever the peer proves it is still there.
void touch();
// Stops watching and joins. Idempotent, and REQUIRED before any long
// non-read work the window must not cover: the caller's own decode is not
// the peer going quiet, and cancelling in the middle of it would throw away
// a transcript the client is waiting for.
void disarm();
// True once the window elapsed and the callback ran. Stays true after
// disarm, so the caller can tell "the peer closed" from "we cancelled it".
bool fired() const;
private:
void run();
const std::chrono::milliseconds window_;
std::function<void()> on_idle_;
mutable std::mutex mu_;
std::condition_variable cv_;
std::chrono::steady_clock::time_point last_;
bool stop_ = false;
bool fired_ = false;
std::thread thread_;
};
// Whether one message read off a live stream is a frame the decoder can
// actually consume, which is the ONLY thing that counts as the peer proving it
// is still there.
//
// Split out of the read loop so the distinction is testable, and because
// getting it wrong is silent. The loop used to touch the watchdog on ANY
// message, before it filtered on has_audio and on an empty pcm field, so a peer
// writing unset-oneof or zero-length frames faster than the window held the
// lane forever: no audio was ever fed, no work was ever done, and the timer
// that exists to break exactly that grip was reset by the frames doing it.
// There is one lane per model and one model per process, so that is a single
// client denying the whole backend. The thrown message already said "no audio
// frame arrived"; this is the code agreeing with it.
bool live_frame_carries_audio(bool has_audio, bool pcm_empty);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,225 @@
// Unit tests for the live stream idle watchdog. Standard library only; the
// harness compiles this as a single translation unit, so the implementation is
// included directly.
//
// These are TIMING tests, which is unavoidable: what is under test is a
// deadline. Every window here is short and every assertion waits several
// multiples of it, so a loaded machine slows the test down rather than
// flipping its answer. The one thing never asserted is how SOON something
// happens, only that it eventually does or never does.
#include "live_watchdog.cpp"
#include <atomic>
#include <cstdio>
#include <string>
#include <thread>
using namespace std::chrono_literals;
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
// A peer that goes quiet without closing. This is the whole point: the lane it
// holds has to come back.
static void test_it_fires_when_nothing_touches_it() {
std::atomic<int> calls{0};
audiocpp_backend::IdleWatchdog watchdog(100ms, [&calls] { ++calls; });
std::this_thread::sleep_for(600ms);
check(watchdog.fired(), "a window that elapses untouched fires");
check(calls.load() == 1, "the callback runs exactly once, not once per window");
}
// A peer that is still speaking must never be cut off. Touches land at a third
// of the window, for six windows' worth of wall clock.
static void test_touching_defers_it_indefinitely() {
std::atomic<int> calls{0};
audiocpp_backend::IdleWatchdog watchdog(300ms, [&calls] { ++calls; });
for (int i = 0; i < 20; ++i) {
std::this_thread::sleep_for(100ms);
watchdog.touch();
}
check(!watchdog.fired(),
"a stream touched inside every window is never cancelled");
check(calls.load() == 0, "no callback runs while the peer is still there");
}
// Disarm is what the handler calls when the read side closes, before a decode
// that can take longer than the window. Firing after that would throw away the
// transcript the client is waiting for.
static void test_disarm_stops_it_before_the_window() {
std::atomic<int> calls{0};
audiocpp_backend::IdleWatchdog watchdog(200ms, [&calls] { ++calls; });
std::this_thread::sleep_for(20ms);
watchdog.disarm();
std::this_thread::sleep_for(600ms);
check(!watchdog.fired(), "a disarmed watchdog does not fire");
check(calls.load() == 0, "a disarmed watchdog runs no callback");
}
static void test_disarm_is_idempotent() {
audiocpp_backend::IdleWatchdog watchdog(50ms, [] {});
watchdog.disarm();
watchdog.disarm();
watchdog.disarm();
check(true, "disarming three times joins once and does not abort");
}
// The operator's escape hatch, for a client that legitimately holds a stream
// open through long pauses. Not "expire immediately", which is what a naive
// reading of a zero timeout would give.
static void test_a_non_positive_window_disables_it() {
std::atomic<int> calls{0};
{
audiocpp_backend::IdleWatchdog watchdog(0ms, [&calls] { ++calls; });
std::this_thread::sleep_for(300ms);
check(!watchdog.fired(), "a zero window never fires");
}
{
audiocpp_backend::IdleWatchdog watchdog(-5ms, [&calls] { ++calls; });
std::this_thread::sleep_for(300ms);
check(!watchdog.fired(), "a negative window never fires");
}
check(calls.load() == 0, "a disabled watchdog runs no callback");
}
// fired() has to survive the disarm, because the handler reads it AFTER the
// read loop ends to tell "the peer closed" from "we cancelled the peer", and
// those two get different statuses.
static void test_fired_survives_a_later_disarm() {
audiocpp_backend::IdleWatchdog watchdog(80ms, [] {});
std::this_thread::sleep_for(500ms);
watchdog.disarm();
check(watchdog.fired(), "a watchdog that fired still says so after disarm");
}
// The destructor joins, so a callback capturing the handler's frame cannot run
// after that frame is gone. Without the join this is a use after free that only
// shows up under load.
//
// The window is LONGER than the scope on purpose. An earlier version of this
// test slept past the window inside the scope, so the callback had already run
// by the time the object was destroyed and a destructor that DETACHED the thread
// instead of joining it passed unnoticed. Mutation testing is what found that;
// the shape below kills it, because a detached thread wakes after the object is
// gone and calls a callback that must never run.
static void test_the_destructor_joins() {
std::atomic<int> calls{0};
std::atomic<bool> alive{true};
{
audiocpp_backend::IdleWatchdog watchdog(200ms, [&calls, &alive] {
check(alive.load(),
"the callback never runs after the watched scope ended");
++calls;
});
std::this_thread::sleep_for(20ms);
}
alive.store(false);
std::this_thread::sleep_for(600ms);
check(calls.load() == 0,
"destruction stops the timer rather than leaving it running against a "
"dead frame");
}
// The other half of that pair: a callback that DOES fire inside the scope runs
// exactly once, so the test above is not passing merely because nothing ever
// fires.
static void test_a_firing_watchdog_still_joins_cleanly() {
std::atomic<int> calls{0};
{
audiocpp_backend::IdleWatchdog watchdog(50ms, [&calls] { ++calls; });
std::this_thread::sleep_for(400ms);
}
check(calls.load() == 1, "the callback ran once, inside the scope");
}
// The predicate the live read loop filters on.
static void test_only_a_frame_with_audio_counts() {
using audiocpp_backend::live_frame_carries_audio;
check(live_frame_carries_audio(true, false),
"a frame with a non-empty pcm field carries audio");
check(!live_frame_carries_audio(true, true),
"an empty pcm field does not");
check(!live_frame_carries_audio(false, false),
"an unset audio oneof does not, whatever the pcm field looks like");
check(!live_frame_carries_audio(false, true),
"and neither does an unset oneof with an empty pcm field");
}
// The defect this closes, expressed as behaviour rather than as a call order:
// a peer writing frames the decoder cannot consume, faster than the window,
// used to hold the model's only inference lane forever, because the read loop
// touched the watchdog before it filtered them out. One lane per model and one
// model per process, so that is a single client denying the whole backend,
// which is exactly what the watchdog exists to prevent.
static void test_empty_frames_do_not_hold_the_lane() {
std::atomic<int> cancels{0};
std::atomic<bool> stop{false};
audiocpp_backend::IdleWatchdog watchdog(80ms, [&cancels] { ++cancels; });
// The read loop with the real filter in it: frames arrive continuously,
// none of them carries audio, and only a frame that does may touch.
std::thread peer([&] {
while (!stop.load()) {
if (audiocpp_backend::live_frame_carries_audio(false, true)) {
watchdog.touch();
}
std::this_thread::sleep_for(5ms);
}
});
std::this_thread::sleep_for(600ms);
stop.store(true);
peer.join();
watchdog.disarm();
check(cancels.load() == 1,
"a flood of frames with no audio in them still releases the lane");
// The mirror image, so this cannot pass merely because the watchdog always
// fires: a peer that keeps sending audio is left alone, exactly as before.
std::atomic<int> live_cancels{0};
std::atomic<bool> live_stop{false};
audiocpp_backend::IdleWatchdog live(80ms, [&live_cancels] { ++live_cancels; });
std::thread speaker([&] {
while (!live_stop.load()) {
if (audiocpp_backend::live_frame_carries_audio(true, false)) {
live.touch();
}
std::this_thread::sleep_for(5ms);
}
});
std::this_thread::sleep_for(600ms);
live_stop.store(true);
speaker.join();
live.disarm();
check(live_cancels.load() == 0,
"a peer that keeps sending audio is never cancelled");
}
int main() {
test_it_fires_when_nothing_touches_it();
test_touching_defers_it_indefinitely();
test_disarm_stops_it_before_the_window();
test_disarm_is_idempotent();
test_a_non_positive_window_disables_it();
test_fired_survives_a_later_disarm();
test_the_destructor_joins();
test_a_firing_watchdog_still_joins_cleanly();
test_only_a_frame_with_audio_counts();
test_empty_frames_do_not_hold_the_lane();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all live_watchdog checks passed\n");
return 0;
}

View File

@@ -0,0 +1,802 @@
#include "loaded_model.h"
#include "family_gate.h"
#include "engine/framework/assets/tensor_source.h"
#include <algorithm>
#include <cstddef>
#include <cstdint>
#include <filesystem>
#include <utility>
#if defined(__APPLE__)
#include <mach-o/dyld.h>
#include <cstdint>
#include <vector>
#elif defined(__linux__)
#include <unistd.h>
#endif
namespace audiocpp_backend {
// --------------------------------------------------------------------------
// Enum coupling
//
// audiocpp_backend::Task mirrors engine::runtime::VoiceTaskKind positionally so
// capability_routing can stay stdlib-only and testable without an audio.cpp
// checkout. Nothing about that mirroring is enforced by the type system, and a
// drift is silent in the worst possible way: every unit still compiles, every
// test still passes, and the backend runs a different task than the one the
// caller asked for.
//
// Two mechanisms pin it, and both are needed because they catch different edits:
//
// 1. The assertions below pin every enumerator's value on both sides. An
// insertion or a reorder anywhere before the last member shifts the values
// after it and fails the build here.
// 2. An enumerator APPENDED after the last one shifts nothing, so no value
// assertion can see it. What sees it is the switch in from_engine_task,
// which covers the engine enum with no `default:` label. CMakeLists.txt
// compiles this file with -Werror=switch so that omission is an error
// rather than a warning nobody reads.
//
// Neither mechanism catches a pure RENAME of an upstream enumerator, but that
// does not need catching: the switch stops naming an enumerator that exists and
// the build fails on its own.
// --------------------------------------------------------------------------
namespace {
constexpr int kEngine(engine::runtime::VoiceTaskKind kind) {
return static_cast<int>(kind);
}
constexpr int kMirror(Task task) { return static_cast<int>(task); }
} // namespace
static_assert(kEngine(engine::runtime::VoiceTaskKind::Vad) == 0, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::Asr) == 1, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::Diarization) == 2, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::SourceSeparation) == 3, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::AudioGeneration) == 4, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::Tts) == 5, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::VoiceCloning) == 6, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::VoiceConversion) == 7, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::SpeechToSpeech) == 8, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::Alignment) == 9, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::VoiceDesign) == 10, "VoiceTaskKind drifted");
static_assert(kEngine(engine::runtime::VoiceTaskKind::SpeakerRecognition) == 11, "VoiceTaskKind drifted");
// The last member. Pinning it pins the member count too, as long as the
// enumerators stay contiguous and unassigned, which upstream's declaration is.
static_assert(kEngine(engine::runtime::VoiceTaskKind::Svc) == 12,
"engine::runtime::VoiceTaskKind gained, lost or reordered a member. "
"audiocpp_backend::Task mirrors it positionally: update capability_routing.h, "
"to_engine_task and from_engine_task together, then move this pin.");
static_assert(kMirror(Task::Vad) == 0, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::Asr) == 1, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::Diarization) == 2, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::SourceSeparation) == 3, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::AudioGeneration) == 4, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::Tts) == 5, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::VoiceCloning) == 6, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::VoiceConversion) == 7, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::SpeechToSpeech) == 8, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::Alignment) == 9, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::VoiceDesign) == 10, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::SpeakerRecognition) == 11, "Task drifted from VoiceTaskKind");
static_assert(kMirror(Task::Svc) == 12, "Task drifted from VoiceTaskKind");
static_assert(static_cast<int>(engine::runtime::RunMode::Offline) == 0, "RunMode drifted");
static_assert(static_cast<int>(engine::runtime::RunMode::Streaming) == 1,
"engine::runtime::RunMode gained, lost or reordered a member. "
"audiocpp_backend::Mode mirrors it positionally.");
static_assert(static_cast<int>(Mode::Offline) == 0, "Mode drifted from RunMode");
static_assert(static_cast<int>(Mode::Streaming) == 1, "Mode drifted from RunMode");
namespace {
engine::core::BackendType parse_backend_type(const std::string &value) {
if (value == "cuda") {
return engine::core::BackendType::Cuda;
}
if (value == "vulkan") {
return engine::core::BackendType::Vulkan;
}
if (value == "metal") {
return engine::core::BackendType::Metal;
}
if (value == "best") {
return engine::core::BackendType::BestAvailable;
}
if (value == "cpu" || value.empty()) {
return engine::core::BackendType::Cpu;
}
throw ConfigError("audio-cpp: unknown backend option '" + value +
"'. Known backends: cpu, cuda, vulkan, metal, best");
}
std::filesystem::path executable_directory() {
#if defined(__APPLE__)
std::uint32_t size = 0;
_NSGetExecutablePath(nullptr, &size);
std::vector<char> buffer(size + 1, '\0');
if (_NSGetExecutablePath(buffer.data(), &size) != 0) {
return std::filesystem::current_path();
}
return std::filesystem::path(buffer.data()).parent_path();
#elif defined(__linux__)
std::error_code ec;
const auto self = std::filesystem::read_symlink("/proc/self/exe", ec);
if (ec) {
return std::filesystem::current_path();
}
return self.parent_path();
#else
return std::filesystem::current_path();
#endif
}
// Runs the load gate. Throws ConfigError rather than returning a decision,
// because the only caller is a delegating constructor whose member initializer
// list has nowhere to put a failure.
std::string require_family(const std::string &resolved_path,
const ModelOptions &options) {
const std::filesystem::path path(resolved_path);
std::error_code ec;
if (!std::filesystem::exists(path, ec)) {
throw ConfigError("audio-cpp: model path does not exist: " + resolved_path);
}
const bool is_gguf = path_looks_like_gguf(resolved_path);
// Only a GGUF is asked for embedded metadata. A directory has no single
// file to read it from, and probing one would make the gate's refusal
// depend on which file happened to be inside.
const std::string embedded =
is_gguf ? read_gguf_family(resolved_path) : std::string();
const FamilyDecision decision =
decide_family(is_gguf, embedded, options.family);
if (!decision.ok) {
throw ConfigError(decision.error);
}
return decision.family;
}
// Refuses a GGUF whose weights are stored in a dtype the family cannot survive.
//
// The POLICY lives in family_gate's weight_dtype_is_supported, which is
// stdlib-only and therefore testable; this is only the part that needs a file
// and an engine to read one. See the table there for why an entry exists and
// what has to be run before deleting it.
//
// Only GGUF paths are inspected. A directory of safetensors carries its dtypes
// per file and has not been tested against this failure, so it is passed
// through rather than guessed at.
void require_supported_weight_dtypes(const std::string &family,
const std::string &resolved_path) {
// Asked as "is there an entry", not as "is the description non-empty": an
// entry with an empty allow list describes a family that can run nothing,
// and reading the description would skip the check on exactly that entry
// while weight_dtype_is_supported refused every dtype. No such entry exists
// today; the two questions are different ones and only one of them is this
// guard's.
if (!family_has_weight_dtype_allow_list(family) ||
!path_looks_like_gguf(resolved_path)) {
return;
}
std::string offending_dtype;
std::string offending_tensor;
try {
const auto source =
engine::assets::open_tensor_source(std::filesystem::path(resolved_path));
if (source == nullptr) {
return;
}
for (const auto &tensor : source->tensors()) {
if (!weight_dtype_is_supported(family, tensor.dtype)) {
offending_dtype = tensor.dtype;
offending_tensor = tensor.name;
break;
}
}
} catch (const std::exception &) {
// Unreadable as a tensor source. Not this guard's problem to report:
// the registry load below produces a message naming the real fault, and
// refusing here would turn every unusual packaging into this error.
return;
}
if (offending_dtype.empty()) {
return;
}
throw ConfigError(
"audio-cpp: family '" + family + "' cannot run weights stored as '" +
offending_dtype + "' (tensor '" + offending_tensor + "' in " +
resolved_path +
"); it aborts the backend process on the first request rather than "
"failing the request. Use the 'orig' GGUF package, whose weights are " +
supported_weight_dtypes(family) + ".");
}
} // namespace
engine::runtime::VoiceTaskKind to_engine_task(Task task) {
using K = engine::runtime::VoiceTaskKind;
// An explicit switch, never a cast: a cast would keep compiling through
// exactly the drift the assertions above exist to catch.
switch (task) {
case Task::Vad: return K::Vad;
case Task::Asr: return K::Asr;
case Task::Diarization: return K::Diarization;
case Task::SourceSeparation: return K::SourceSeparation;
case Task::AudioGeneration: return K::AudioGeneration;
case Task::Tts: return K::Tts;
case Task::VoiceCloning: return K::VoiceCloning;
case Task::VoiceConversion: return K::VoiceConversion;
case Task::SpeechToSpeech: return K::SpeechToSpeech;
case Task::Alignment: return K::Alignment;
case Task::VoiceDesign: return K::VoiceDesign;
case Task::SpeakerRecognition: return K::SpeakerRecognition;
case Task::Svc: return K::Svc;
}
// Unreachable for any valid enumerator. No `default:` label, so -Wswitch
// still reports a member this switch stops covering.
return K::Vad;
}
Task from_engine_task(engine::runtime::VoiceTaskKind kind) {
using K = engine::runtime::VoiceTaskKind;
switch (kind) {
case K::Vad: return Task::Vad;
case K::Asr: return Task::Asr;
case K::Diarization: return Task::Diarization;
case K::SourceSeparation: return Task::SourceSeparation;
case K::AudioGeneration: return Task::AudioGeneration;
case K::Tts: return Task::Tts;
case K::VoiceCloning: return Task::VoiceCloning;
case K::VoiceConversion: return Task::VoiceConversion;
case K::SpeechToSpeech: return Task::SpeechToSpeech;
case K::Alignment: return Task::Alignment;
case K::VoiceDesign: return Task::VoiceDesign;
case K::SpeakerRecognition: return Task::SpeakerRecognition;
case K::Svc: return Task::Svc;
}
return Task::Vad;
}
engine::runtime::RunMode to_engine_mode(Mode mode) {
using M = engine::runtime::RunMode;
switch (mode) {
case Mode::Offline: return M::Offline;
case Mode::Streaming: return M::Streaming;
}
return M::Offline;
}
Mode from_engine_mode(engine::runtime::RunMode mode) {
using M = engine::runtime::RunMode;
switch (mode) {
case M::Offline: return Mode::Offline;
case M::Streaming: return Mode::Streaming;
}
return Mode::Offline;
}
Capabilities to_capabilities(const std::string &family,
const engine::runtime::CapabilitySet &set) {
Capabilities caps;
caps.family = family;
caps.tasks.reserve(set.supported_tasks.size());
for (const auto &supported : set.supported_tasks) {
TaskCapability capability;
capability.task = from_engine_task(supported.task);
capability.modes.reserve(supported.modes.size());
for (const auto mode : supported.modes) {
capability.modes.push_back(from_engine_mode(mode));
}
caps.tasks.push_back(std::move(capability));
}
return caps;
}
std::string read_gguf_family(const std::string &path) {
try {
const auto spec =
engine::assets::read_gguf_embedded_model_spec(std::filesystem::path(path));
if (spec.has_value()) {
return spec->family;
}
} catch (...) {
// A file that is not a readable GGUF simply has no family. The load
// gate turns that into a clear refusal; a throw here would surface as
// an opaque internal error during backend probing.
}
return {};
}
std::string resolve_model_path(const std::string &model_path_dir,
const std::string &model_file,
const std::string &model_name) {
const std::string candidate = !model_file.empty() ? model_file : model_name;
// The bundled: form is looked for in BOTH fields, and in ModelOptions.Model
// FIRST, because that is the only field it survives in. LocalAI fills
// ModelFile by joining ModelPath onto the configured model string
// (pkg/model/loader.go, LoadModelWithFile), so a model YAML saying
// `model: bundled:silero_vad` arrives here as ModelFile
// "/models/bundled:silero_vad" and Model "bundled:silero_vad". Testing
// `candidate` alone therefore made the zero-download VAD path reachable only
// from a hand-written LoadModel call that left ModelFile empty, and every
// model YAML using it failed with "model path does not exist".
const std::string bundled_prefix = "bundled:";
for (const std::string *field : {&model_name, &model_file}) {
if (field->rfind(bundled_prefix, 0) == 0) {
const std::string name = field->substr(bundled_prefix.size());
return (executable_directory() / "assets" / name).string();
}
}
std::filesystem::path path(candidate);
if (path.is_absolute() || model_path_dir.empty()) {
return path.string();
}
return (std::filesystem::path(model_path_dir) / path).string();
}
LoadedModel::LoadedModel(const std::string &resolved_path,
const ModelOptions &options,
std::string model_identity)
: LoadedModel(resolved_path, options, require_family(resolved_path, options),
std::move(model_identity)) {}
LoadedModel::LoadedModel(const std::string &resolved_path,
const ModelOptions &options, std::string family,
std::string model_identity)
: lane_(family), registry_(engine::runtime::make_default_registry()),
identity_(std::move(model_identity)) {
if (!registry_.supports_family(family)) {
throw ConfigError("audio-cpp: unknown audio.cpp family '" + family + "'");
}
// Before the load, for the same reason parse_backend_type runs before it:
// a refusal a metadata read can produce should not cost a full model load.
// More importantly it must precede the FIRST REQUEST, since that is where
// an unsupported dtype aborts the process rather than failing.
require_supported_weight_dtypes(family, resolved_path);
// Session options are built BEFORE the load, because parse_backend_type
// rejects an unknown backend name. Validating after the load would make
// `backend:cudaa` cost a full model load, on a fault a string comparison
// could have caught.
session_options_.backend.type = parse_backend_type(options.backend);
session_options_.backend.device = options.device;
if (options.threads > 0) {
session_options_.backend.threads = options.threads;
}
for (const auto &entry : options.session_options) {
session_options_.options[entry.first] = entry.second;
}
pinned_task_ = options.task;
wait_budget_ceiling_ms_ = options.busy_timeout_ms;
live_idle_timeout_ms_ = options.live_idle_timeout_ms;
engine::runtime::ModelLoadRequest request;
request.model_path = std::filesystem::path(resolved_path);
request.family_hint = family;
if (!options.model_spec_override.empty()) {
request.model_spec_override =
std::filesystem::path(options.model_spec_override);
}
for (const auto &entry : options.load_options) {
request.options[entry.first] = entry.second;
}
try {
model_ = registry_.load(request);
} catch (const std::exception &err) {
throw ConfigError("audio-cpp: failed to load family '" + family +
"' from " + resolved_path + ": " + err.what());
}
if (model_ == nullptr) {
throw ConfigError("audio-cpp: the registry returned no model for " +
resolved_path);
}
const auto &metadata = model_->metadata();
const auto &engine_caps = model_->capabilities();
variant_ = metadata.variant;
description_ = metadata.description;
languages_ = engine_caps.languages;
supports_timestamps_ = engine_caps.supports_timestamps;
capabilities_ = to_capabilities(family, engine_caps);
}
Route LoadedModel::check_can_serve(Rpc rpc, const RequestShape &shape) const {
const Route route = resolve_route(rpc, shape, capabilities_);
if (!route.ok) {
throw CapabilityError(route.error);
}
// Returned so a handler can act on the task before running it. It is the
// same route session_for will resolve, since both read the immutable
// capabilities_ from the same shape.
return route;
}
LoadedModel::Session LoadedModel::session_for(Rpc rpc, const RequestShape &shape,
LaneEntry &lane) {
// Proof of holding only. Nothing here reads it, and nothing should: its
// whole job is to make a caller that has not taken the lane fail to
// compile. Non-const so it cannot bind to an inline acquire(), whose
// temporary would be released at the end of this call.
(void)lane;
const Route route = resolve_route(rpc, shape, capabilities_);
if (!route.ok) {
throw CapabilityError(route.error);
}
const SessionKey key{static_cast<int>(route.task), static_cast<int>(route.mode)};
auto found = sessions_.find(key);
const bool cache_hit = found != sessions_.end();
if (!cache_hit) {
engine::runtime::TaskSpec spec;
spec.task = to_engine_task(route.task);
spec.mode = to_engine_mode(route.mode);
std::unique_ptr<engine::runtime::IVoiceTaskSession> created;
try {
created = model_->create_task_session(spec, session_options_);
} catch (const std::exception &err) {
// NOT a CapabilityError. The family said it supports this pair, and
// a throw from here is overwhelmingly an environment fault: a ggml
// backend .so that package.sh did not ship, an out of memory, a CUDA
// device that is not there. UNIMPLEMENTED would tell LocalAI and
// every client "this model cannot do this, never retry", and send an
// operator hunting a capability bug instead of a packaging one. A
// plain runtime_error maps to INTERNAL, which is what a fixable
// deployment fault should look like.
throw std::runtime_error(
std::string("audio-cpp: family '") + capabilities_.family +
"' advertises " + task_name(route.task) + "/" +
mode_name(route.mode) + " but refused to create the session: " +
err.what());
}
if (created == nullptr) {
// A null return with no throw is the family declining, which is a
// genuine capability answer and stays UNIMPLEMENTED.
throw CapabilityError(std::string("audio-cpp: family '") +
capabilities_.family +
"' returned no session for " +
task_name(route.task) + "/" +
mode_name(route.mode));
}
found = sessions_.emplace(key, std::move(created)).first;
}
Session session;
session.task = route.task;
session.mode = route.mode;
engine::runtime::IVoiceTaskSession *raw = found->second.get();
if (route.mode == Mode::Streaming) {
session.streaming =
dynamic_cast<engine::runtime::IStreamingVoiceTaskSession *>(raw);
if (session.streaming == nullptr) {
throw CapabilityError(std::string("audio-cpp: family '") +
capabilities_.family +
"' advertises " + task_name(route.task) +
"/streaming but its session is not streaming");
}
// Deliberately NOT reset here, though a cached streaming session does
// carry state across chunks. reset() is not callable at this point:
// silero_vad's implementation throws "session prepare() must be called
// before Silero VAD reset()", so resetting on a cache hit would turn an
// ordinary second fetch into a hard error, which is worse than the leak
// it would prevent.
//
// The state is instead cleared by the sequence every streaming caller
// owes anyway. IStreamingVoiceTaskSession::start_stream's base
// implementation IS a call to reset(), so a caller that runs
// prepare(...) then start_stream(...) at the top of each stream gets a
// clean session for free. See the STATE CONTRACT in loaded_model.h.
} else {
session.offline =
dynamic_cast<engine::runtime::IOfflineVoiceTaskSession *>(raw);
if (session.offline == nullptr) {
throw CapabilityError(std::string("audio-cpp: family '") +
capabilities_.family +
"' advertises " + task_name(route.task) +
"/offline but its session is not offline");
}
}
return session;
}
LaneEntry LoadedModel::acquire(int requested_timeout_ms) {
// Constructed straight into the return value. C++17 requires that, which is
// what lets an immovable type be returned at all; a named local here would
// not compile.
return LaneEntry(lane_,
resolve_wait_budget_ms(wait_budget_ceiling_ms_,
requested_timeout_ms));
}
std::unique_ptr<LaneEntry> LoadedModel::acquire_owned(int requested_timeout_ms) {
return std::make_unique<LaneEntry>(
lane_,
resolve_wait_budget_ms(wait_budget_ceiling_ms_, requested_timeout_ms));
}
engine::runtime::TaskResult run_offline(const LoadedModel::Session &session,
const engine::runtime::TaskRequest &request,
LaneEntry &lane) {
// Proof of holding only, as in session_for.
(void)lane;
if (session.offline == nullptr) {
throw CapabilityError("audio-cpp: no offline session for this request");
}
session.offline->prepare(engine::runtime::build_preparation_request(request));
return session.offline->run(request);
}
namespace {
engine::runtime::IStreamingVoiceTaskSession &
require_streaming(const LoadedModel::Session &session) {
if (session.streaming == nullptr) {
throw CapabilityError("audio-cpp: no streaming session for this request");
}
return *session.streaming;
}
// Clears the stream event sink on every exit from the driver, including the
// exception path. The session is CACHED and outlives the call that installed
// the sink, so a std::function left behind holding references into that call's
// frame is called with dangling captures by whoever streams next.
class ScopedStreamSink {
public:
ScopedStreamSink(engine::runtime::IStreamingVoiceTaskSession &session,
engine::runtime::StreamEventCallback sink)
: session_(session) {
session_.set_stream_event_sink(std::move(sink));
}
~ScopedStreamSink() { session_.set_stream_event_sink(nullptr); }
ScopedStreamSink(const ScopedStreamSink &) = delete;
ScopedStreamSink &operator=(const ScopedStreamSink &) = delete;
private:
engine::runtime::IStreamingVoiceTaskSession &session_;
};
// Frames per chunk to feed a streaming session, from its own policy.
//
// FRAMES, not floats. preferred_audio_chunk_samples is a per-channel count
// everywhere upstream sets it (nemotron_asr uses its frontend sample rate,
// i.e. one second), and vibevoice_asr refuses a chunk whose float count is not
// divisible by its channel count, so slicing on floats would both mis-size the
// window and hand a family a half frame.
std::int64_t chunk_frames_for(const engine::runtime::StreamingPolicy &policy,
int sample_rate) {
if (policy.preferred_audio_chunk_samples > 0) {
return policy.preferred_audio_chunk_samples;
}
// higgs_audio_stt states its window in seconds (4.0) and leaves the sample
// count at zero, so this branch is real rather than defensive.
if (policy.preferred_audio_chunk_seconds > 0.0 && sample_rate > 0) {
const auto frames = static_cast<std::int64_t>(
policy.preferred_audio_chunk_seconds * static_cast<double>(sample_rate));
if (frames > 0) {
return frames;
}
}
// The interface's own default, from IStreamingVoiceTaskSession::streaming_policy.
return 512;
}
} // namespace
void begin_stream(const LoadedModel::Session &session,
const engine::runtime::TaskRequest &request, LaneEntry &lane) {
// Proof of holding only, as in session_for.
(void)lane;
auto &streaming = require_streaming(session);
// Order is load-bearing: start_stream's reset() is illegal before prepare().
streaming.prepare(engine::runtime::build_preparation_request(request));
streaming.start_stream(request);
}
engine::runtime::TaskResult run_streaming_pull(
const LoadedModel::Session &session,
const engine::runtime::TaskRequest &request,
const std::function<void(const engine::runtime::StreamEvent &)> &on_event,
LaneEntry &lane) {
auto &streaming = require_streaming(session);
begin_stream(session, request, lane);
while (const auto event = streaming.next_stream_event()) {
if (on_event) {
on_event(*event);
}
// No pinned family sets is_final on a pulled event, so this is not what
// ends the loop today; the nullopt above is. Honoured anyway, because a
// family that does set it is saying the stream is over and pulling once
// more would be asking a finished session for another chunk.
if (event->is_final) {
break;
}
}
return streaming.finish_stream();
}
engine::runtime::TaskResult run_streaming_audio(
const LoadedModel::Session &session,
const engine::runtime::TaskRequest &request,
const engine::runtime::AudioBuffer &audio,
const std::function<void(const engine::runtime::StreamEvent &)> &on_event,
LaneEntry &lane) {
auto &streaming = require_streaming(session);
const int channels = audio.channels > 0 ? audio.channels : 1;
// REFUSED, not rounded away, and checked before anything is touched so a
// refusal leaves no half-started stream on a cached session.
//
// An interleaved buffer whose float count is not a whole number of frames
// is a truncated input, and the integer division below would silently drop
// the tail floats: they would never be fed, never reach the transcript, and
// nothing would say so. Upstream refuses the same condition rather than
// tolerating it, in two places: vibevoice_asr's audio_frame_count throws
// "VibeVoice-ASR audio samples must be divisible by channel count"
// (session.cpp:70-76), and its process_audio_chunk throws the same with
// "streamed" in the text about the chunks this driver hands it
// (session.cpp:742-747).
//
// ConfigError, i.e. INVALID_ARGUMENT, because the buffer came from the
// caller's file. read_audio_file's positive-rate path always answers mono
// and so cannot reach this, but its native-rate path passes the reader's
// sample count through unchanged, and a driver does not get to assume which
// path its caller took.
if (audio.samples.size() % static_cast<std::size_t>(channels) != 0) {
throw ConfigError(
"audio-cpp: streaming input is not a whole number of frames: " +
std::to_string(audio.samples.size()) + " samples across " +
std::to_string(channels) + " channels");
}
// Installed BEFORE the stream begins, so a family that reports during
// start_stream is not silently dropped, and destroyed after finish_stream,
// because nemotron_asr emits every one of its partials from inside
// finalize().
ScopedStreamSink sink(streaming,
[&on_event](const engine::runtime::StreamEvent &event) {
if (on_event) {
on_event(event);
}
});
begin_stream(session, request, lane);
const auto total_frames =
static_cast<std::int64_t>(audio.samples.size() / static_cast<size_t>(channels));
const std::int64_t chunk_frames =
chunk_frames_for(streaming.streaming_policy(), audio.sample_rate);
for (std::int64_t offset = 0; offset < total_frames; offset += chunk_frames) {
const std::int64_t end = std::min(offset + chunk_frames, total_frames);
engine::runtime::AudioChunk chunk;
chunk.sample_rate = audio.sample_rate;
chunk.channels = channels;
// A FRAME index, which is what every span in a returned event is
// expressed in. vibevoice_asr adds the chunk's own frame count to it to
// offset the spans it reports, so a float index here would place every
// span of a stereo stream at twice its real time.
chunk.start_sample = offset;
chunk.samples.assign(
audio.samples.begin() + static_cast<std::ptrdiff_t>(offset * channels),
audio.samples.begin() + static_cast<std::ptrdiff_t>(end * channels));
const auto event = streaming.process_audio_chunk(chunk);
if (on_event) {
on_event(event);
}
}
return streaming.finish_stream();
}
engine::runtime::TaskResult run_streaming_live(
const LoadedModel::Session &session,
const engine::runtime::TaskRequest &request,
const std::function<bool(std::vector<float> &)> &next_frames,
const std::function<void(const engine::runtime::StreamEvent &)> &on_event,
LaneEntry &lane) {
auto &streaming = require_streaming(session);
// The contract is the only thing that says what rate and layout the frames
// about to arrive are in, and prepare() needs it: see the header.
if (!request.audio_input.has_value()) {
throw ConfigError(
"audio-cpp: a live streaming request carries no audio contract");
}
const int sample_rate = request.audio_input->sample_rate;
const int channels =
request.audio_input->channels > 0 ? request.audio_input->channels : 1;
// Installed BEFORE the stream begins and cleared on every exit, including
// the exception path, for the reasons spelled out in run_streaming_audio.
ScopedStreamSink sink(streaming,
[&on_event](const engine::runtime::StreamEvent &event) {
if (on_event) {
on_event(event);
}
});
begin_stream(session, request, lane);
const std::int64_t chunk_frames =
chunk_frames_for(streaming.streaming_policy(), sample_rate);
// chunk_frames_for never returns a non-positive count, so this is never
// zero and the accumulation loop below always terminates.
const std::size_t chunk_floats = static_cast<std::size_t>(chunk_frames) *
static_cast<std::size_t>(channels);
std::int64_t fed_frames = 0;
const auto feed = [&](std::vector<float> samples) {
engine::runtime::AudioChunk chunk;
chunk.sample_rate = sample_rate;
chunk.channels = channels;
// A FRAME index, counted across the whole stream: vibevoice_asr offsets
// every span it reports by it, so restarting it per chunk would put
// every word at the top of the recording.
chunk.start_sample = fed_frames;
chunk.samples = std::move(samples);
fed_frames +=
static_cast<std::int64_t>(chunk.samples.size()) / channels;
const auto event = streaming.process_audio_chunk(chunk);
if (on_event) {
on_event(event);
}
};
std::vector<float> pending;
std::vector<float> incoming;
while (true) {
incoming.clear();
if (!next_frames(incoming)) {
break;
}
pending.insert(pending.end(), incoming.begin(), incoming.end());
while (pending.size() >= chunk_floats) {
std::vector<float> window(pending.begin(),
pending.begin() +
static_cast<std::ptrdiff_t>(chunk_floats));
pending.erase(pending.begin(),
pending.begin() +
static_cast<std::ptrdiff_t>(chunk_floats));
feed(std::move(window));
}
}
if (!pending.empty()) {
// The tail is whatever did not fill a window. Refused rather than
// truncated when it is not a whole number of frames, exactly as in
// run_streaming_audio: the division above would drop the stray floats
// from the transcript with no diagnostic. Unreachable for a mono live
// stream, which is every live stream today.
if (pending.size() % static_cast<std::size_t>(channels) != 0) {
throw ConfigError(
"audio-cpp: live stream ended mid-frame: " +
std::to_string(pending.size()) + " trailing samples across " +
std::to_string(channels) + " channels");
}
feed(std::move(pending));
}
if (fed_frames == 0) {
// Nothing was spoken. See the header: finalizing an empty stream is not
// legal for every family, and an empty transcript is the truthful
// answer rather than an engine-internal INTERNAL.
return engine::runtime::TaskResult{};
}
return streaming.finish_stream();
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,378 @@
#pragma once
// Owns one audio.cpp model for the life of the process, plus a lazily created
// session per (task, mode) so the same model answers both TTS and TTSStream.
// This is the only unit that converts between the stdlib-only mirror types and
// engine::runtime types.
#include "capability_routing.h"
#include "inference_lane.h"
#include "model_options.h"
#include "engine/framework/runtime/model.h"
#include "engine/framework/runtime/registry.h"
#include "engine/framework/runtime/session.h"
#include <functional>
#include <map>
#include <memory>
#include <stdexcept>
#include <string>
#include <utility>
#include <vector>
namespace audiocpp_backend {
// User-fixable configuration problem. grpc-server maps this to INVALID_ARGUMENT.
class ConfigError : public std::runtime_error {
public:
using std::runtime_error::runtime_error;
};
// The family cannot serve the requested RPC. Maps to UNIMPLEMENTED.
class CapabilityError : public std::runtime_error {
public:
using std::runtime_error::runtime_error;
};
engine::runtime::VoiceTaskKind to_engine_task(Task task);
engine::runtime::RunMode to_engine_mode(Mode mode);
Task from_engine_task(engine::runtime::VoiceTaskKind kind);
Mode from_engine_mode(engine::runtime::RunMode mode);
Capabilities to_capabilities(const std::string &family,
const engine::runtime::CapabilitySet &set);
// Reads audiocpp.model_spec.family from a GGUF. Returns an empty string when
// the file is not a GGUF, carries no audio.cpp spec, or cannot be read. Never
// throws: an unreadable file is the load gate's problem, not a crash.
std::string read_gguf_family(const std::string &path);
// Builds the absolute model path from LocalAI's (ModelPath, ModelFile, Model)
// triple. Either the Model or the ModelFile field may carry the form
// "bundled:<name>", which resolves to <executable dir>/assets/<name>, where
// package.sh puts upstream's bundled silero_vad and marblenet_vad assets. BOTH
// are checked because LocalAI fills ModelFile by joining ModelPath onto the
// configured model string, so a model YAML using the form has it intact only in
// Model.
std::string resolve_model_path(const std::string &model_path_dir,
const std::string &model_file,
const std::string &model_name);
class LoadedModel {
public:
struct Session {
Task task = Task::Tts;
Mode mode = Mode::Offline;
// Exactly one of these is non-null, matching the resolved mode.
engine::runtime::IOfflineVoiceTaskSession *offline = nullptr;
engine::runtime::IStreamingVoiceTaskSession *streaming = nullptr;
};
// Throws ConfigError when the path does not exist, the family cannot be
// determined, or the registry rejects the family.
//
// `model_identity` is ModelOptions.Model verbatim: the UNTRANSLATED
// controller-side name. It is a constructor argument rather than a setter
// so identity and model are inseparable. llama-cpp keeps its equivalent in
// a separate global from the model, which leaves a window where a handler
// can read one without the other; here a handler that holds the model
// through snapshot() necessarily holds the identity it was loaded with.
LoadedModel(const std::string &resolved_path, const ModelOptions &options,
std::string model_identity);
LoadedModel(const LoadedModel &) = delete;
LoadedModel &operator=(const LoadedModel &) = delete;
const std::string &family() const noexcept { return capabilities_.family; }
// Empty when the controller predates ModelOptions.ModelIdentity, which the
// identity check reads as "skip". See check_model_identity in grpc-server.
const std::string &identity() const noexcept { return identity_; }
const std::string &variant() const noexcept { return variant_; }
const std::string &description() const noexcept { return description_; }
const std::vector<std::string> &languages() const noexcept { return languages_; }
const Capabilities &capabilities() const noexcept { return capabilities_; }
bool supports_timestamps() const noexcept { return supports_timestamps_; }
const engine::runtime::SessionOptions &session_options() const noexcept {
return session_options_;
}
// The model's `task:` option, empty when unset. Every handler must copy it
// into RequestShape::pinned_task before calling session_for: routing is
// otherwise derived from the RPC alone, and this is the option's only route
// from the load to the request that honours it.
const std::string &pinned_task() const noexcept { return pinned_task_; }
// The `live_idle_timeout_ms` option: how long AudioTranscriptionLive waits
// for the next audio frame before cancelling the stream to give this
// model's lane back. 0 means no limit. See the option in model_options.h
// for why it exists and how the default was chosen.
int live_idle_timeout_ms() const noexcept { return live_idle_timeout_ms_; }
// Throws the same CapabilityError session_for would throw when this family
// cannot serve the RPC, and RETURNS THE RESOLVED ROUTE otherwise.
//
// The route is returned rather than computed and dropped because a handler
// often has to know which task it is about to run BEFORE running it.
// AudioTransform refuses params[stem] on any route but source separation,
// and reading that off the route costs microseconds where reading it off
// the result costs a whole inference first. A caller with no such need
// ignores the value, which is what the three transcription-shaped handlers
// do.
//
// It exists so a refusal does not have to buy a place in the queue first.
// resolve_route is a pure function of capabilities_, which is fixed at
// construction and never written again, so unlike the session cache it
// needs no lane and no lock: a model that cannot transcribe can say so
// while another request is halfway through a thirty second run. Without
// this the refusal waits for that run to finish only to be told no.
//
// It does NOT replace the routing inside session_for, and must not be made
// to: session_for still needs the route to key the session cache. The two
// calls agree because both read the same immutable capabilities. What this
// one adds is only the ordering, so call it before acquire().
//
// Const and lane-free on purpose. If a future edit makes routing depend on
// mutable state, this must grow the lane parameter its siblings carry.
Route check_can_serve(Rpc rpc, const RequestShape &shape) const;
// Routes the RPC and returns the cached session, creating it on first use.
// Throws CapabilityError when this family cannot serve the RPC, and a plain
// runtime_error when it can but the session could not be built, which is an
// environment fault rather than a capability answer.
//
// The `lane` parameter is a PROOF OF HOLDING and is otherwise unused: it
// exists so the rule below is a compile error rather than prose. The
// session cache is an unsynchronised std::map and the sessions themselves
// are not reentrant, so this must only be called with the lane held; the
// lane admits one caller at a time, which is exactly the constraint the
// sessions impose. Pass the LaneEntry from acquire().
//
// NON-CONST reference on purpose, and do not "tidy" it to const. A const
// reference binds to a temporary, which makes this compile:
//
// auto session = model->session_for(rpc, shape, model->acquire(0));
// auto result = run_offline(session, task, model->acquire(0));
//
// and each temporary dies at the end of its own full-expression, so the
// lane is released between the two calls. That is precisely the split this
// parameter exists to prevent, and it is the form a future caller is most
// likely to reach for because it reads as tidy. Requiring an lvalue forces
// a named entry whose scope spans both calls.
//
// What it proves is bounded, so do not over-trust it: it proves A lane was
// taken, not THIS model's lane. A caller determined to defeat it can
// construct an entry on an unrelated InferenceLane and pass that. It
// therefore catches the two mistakes that actually happen, forgetting the
// lane entirely and taking it after routing, and does not catch lane
// identity.
//
// STATE CONTRACT, and it is the CALLER'S to honour. Sessions are cached per
// (task, mode), so a streaming session is normally the same warm object the
// previous stream used, carrying that stream's state. session_for hands it
// back as it is.
//
// Every streaming caller must therefore begin a stream through
// begin_stream() below, which is prepare() then start_stream() in that
// order and is the ONLY implementation of that sequence. start_stream's
// base implementation is a call to reset(), which is what clears the
// previous stream, and reset() is only legal after prepare(): silero_vad
// throws "session prepare() must be called before Silero VAD reset()"
// otherwise. That ordering constraint is also why session_for cannot do
// this for you. Skipping it does not raise an error, it silently continues
// the previous stream.
//
// Offline sessions need no such care: their interface has no reset and
// run() takes a whole request.
Session session_for(Rpc rpc, const RequestShape &shape, LaneEntry &lane);
// Takes the inference lane, or throws LaneUnavailable. Serializes runs
// against this model. `requested_timeout_ms` is a per-request wait hint
// where a value <= 0 means "use the model's configured ceiling"; a hint may
// only tighten that ceiling, never loosen it.
//
// LaneEntry is deliberately immovable, so bind the result to a named local
// in the scope the inference happens in:
//
// LaneEntry entry = model.acquire(request_hint_ms);
//
// which C++17 initializes in place. A handler that has to keep the lane
// beyond one scope, for instance in a member that outlives the call that
// took it, wants acquire_owned instead.
LaneEntry acquire(int requested_timeout_ms);
// Same lane, heap-allocated so it can be stored or handed on. Prefer
// acquire: this one adds a null state that the scoped form does not have.
std::unique_ptr<LaneEntry> acquire_owned(int requested_timeout_ms);
private:
// Keyed by the enum values so the map needs no custom comparator.
using SessionKey = std::pair<int, int>;
// The public constructor runs the load gate, then delegates here. The
// detour exists because lane_ has to be built from the family in the member
// initializer list, and the family is only known after the gate has run.
// Four parameters rather than three so it cannot be confused with the
// public constructor, whose third argument is also a std::string.
LoadedModel(const std::string &resolved_path, const ModelOptions &options,
std::string family, std::string model_identity);
// MEMBER ORDER IS LOAD-BEARING BELOW THIS LINE. Members are destroyed in
// reverse declaration order.
//
// lane_ is first so it is destroyed last: nothing that runs during teardown
// can then find a lane that has already gone.
InferenceLane lane_;
// registry_ before model_: the registry owns the loader that produced the
// model, and the model may hold loader-owned state.
engine::runtime::ModelRegistry registry_;
// model_ before sessions_, so sessions_ is destroyed FIRST and the model
// second. A session is created from the model and must not outlive it. Do
// not reorder these two.
std::unique_ptr<engine::runtime::ILoadedVoiceModel> model_;
std::map<SessionKey, std::unique_ptr<engine::runtime::IVoiceTaskSession>> sessions_;
engine::runtime::SessionOptions session_options_;
Capabilities capabilities_;
std::string variant_;
std::string description_;
std::vector<std::string> languages_;
std::string pinned_task_;
std::string identity_;
bool supports_timestamps_ = false;
int wait_budget_ceiling_ms_ = 0;
int live_idle_timeout_ms_ = 0;
};
// Prepares and runs an offline session. prepare() is called for every run
// rather than once per session, because SessionPreparationRequest is derived
// from the request itself (audio contract, text, voice condition) and not from
// the model: a second request with a different sample rate or length would
// otherwise run against the first request's contract.
//
// `lane` is a PROOF OF HOLDING, unused at runtime, for the same reason
// session_for takes one: the session is not reentrant and prepare() mutates it,
// so running without the lane is a data race. Making it a parameter turns that
// into a compile error instead of a comment. Non-const for the same reason as
// session_for's: a const reference would bind to `model.acquire(0)` written
// inline, and that temporary dies at the end of this call, releasing the lane
// before the caller's next one.
//
// Throws CapabilityError when the session is not an offline one. That should be
// unreachable through session_for, which already refuses a non-offline session
// for an offline route, and is checked anyway because the alternative is a null
// dereference.
engine::runtime::TaskResult run_offline(const LoadedModel::Session &session,
const engine::runtime::TaskRequest &request,
LaneEntry &lane);
// THE ONE IMPLEMENTATION of the streaming state obligation described in
// session_for's STATE CONTRACT: prepare(), then start_stream(), in that order.
//
// It is a function rather than a comment because the obligation is invisible
// when it is broken. Streaming sessions are CACHED per (task, mode), so the
// object a second stream gets is the warm one the first stream left behind,
// still holding its audio, its tokens and its started flag. What clears it is
// start_stream, whose base implementation IS a reset() and whose seven family
// overrides (nemotron_asr, vibevoice_asr, higgs_audio_stt, voxtral_realtime,
// supertonic, omnivoice, voxcpm2) every one call reset() as their first
// statement, verified in the pinned checkout. Nothing in the type system pins
// that. A future override that dropped the reset would break every call site
// at once with no compile error and no exception, only a second transcript
// that begins with the first one's audio, so the fewer call sites there are to
// break, the better: this is the only one.
//
// prepare() must come first and cannot be folded into session_for, because
// reset() is illegal before prepare() (silero_vad throws "session prepare()
// must be called before Silero VAD reset()"), and because the preparation
// request is derived from the REQUEST, not the model: build_preparation_request
// reads the audio contract, the text and the voice condition off it, so a
// second stream with a different sample rate or length would otherwise run
// against the first stream's contract.
//
// `lane` is a PROOF OF HOLDING, unused at runtime, exactly as in run_offline.
//
// Throws CapabilityError when the session is not a streaming one.
void begin_stream(const LoadedModel::Session &session,
const engine::runtime::TaskRequest &request, LaneEntry &lane);
// Drives a streaming session that takes NO incremental input, which is the TTS
// shape (StreamingInputKind::None, StreamingOutputKind::PullEvents): begin the
// stream, pull events until the session says there are no more, then finish.
//
// NO STREAM EVENT SINK IS INSTALLED HERE, and that is deliberate rather than an
// omission. voxcpm2's start_stream runs the whole synthesis and pushes every
// chunk to the sink, then its next_stream_event replays those same chunks out
// of the stored result, so a sink on this path would put every chunk of audio
// on the wire twice. supertonic and omnivoice ignore set_stream_event_sink
// outright. The pull loop is therefore the single delivery channel.
//
// The returned TaskResult is the session's own merged whole for all three
// families, NOT a tail the pull loop missed. A caller that already emitted the
// pulled events must not also emit its audio; see the TTSStream handler.
engine::runtime::TaskResult run_streaming_pull(
const LoadedModel::Session &session,
const engine::runtime::TaskRequest &request,
const std::function<void(const engine::runtime::StreamEvent &)> &on_event,
LaneEntry &lane);
// Drives a streaming session that CONSUMES audio chunks, which is the ASR shape
// (StreamingInputKind::AudioChunks): begin the stream, feed the buffer in
// policy-sized chunks, then finalize.
//
// A STREAM EVENT SINK IS INSTALLED HERE, and it is not optional: nemotron_asr
// reports its partial text ONLY through the sink, and only from inside
// finalize(), because its decode does not start until the audio is complete.
// Without the sink that family streams a transcript with no partials at all.
// The sink is cleared again before returning, including on the exception path:
// the session is cached and outlives this call, so a sink left holding a
// reference to the caller's frame is a use after free waiting for the next
// stream.
//
// Both delivery channels are consumed, the sink and the value process_audio_chunk
// returns, because the families do not agree on which they use, and
// voxtral_realtime uses BOTH for the same event. The duplicate that produces is
// absorbed by TranscriptDeltaTracker in stream_delta.h rather than here.
engine::runtime::TaskResult run_streaming_audio(
const LoadedModel::Session &session,
const engine::runtime::TaskRequest &request,
const engine::runtime::AudioBuffer &audio,
const std::function<void(const engine::runtime::StreamEvent &)> &on_event,
LaneEntry &lane);
// Drives the same ASR shape as run_streaming_audio when the audio DOES NOT
// EXIST YET, which is the live-microphone case: instead of slicing a buffer it
// pulls frames from the caller until the input side closes.
//
// `next_frames` fills `out` with interleaved float PCM and returns true, or
// returns false when there is no more input. It is expected to BLOCK, since the
// only real implementation is a gRPC stream Read, and it may throw: a request
// the handler has to refuse mid-stream unwinds through here, and the sink is
// cleared on that path like every other.
//
// The audio contract comes from `request.audio_input`, which for a live stream
// is an EMPTY buffer carrying only the sample rate and channel count. It is not
// optional: nemotron_asr's streaming prepare() throws "Nemotron ASR streaming
// prepare() requires an audio contract" without one, and there is no buffer to
// derive it from here.
//
// Frames are BUFFERED to the family's own preferred window rather than fed in
// whatever sizes the wire delivered them in, because that window is a family's
// statement about what it can decode (nemotron_asr asks for one second, higgs
// for four), and a 512-sample gRPC frame is a property of the client's audio
// callback rather than of the model. The tail shorter than a window is fed at
// the end.
//
// A stream that carried NO AUDIO returns an empty TaskResult and never calls
// finish_stream. Finalizing an empty stream is not universally legal:
// nemotron_asr throws "Nemotron ASR finalize requires streamed audio", so a
// client that opens a session and closes it without speaking would receive an
// INTERNAL naming an engine internal instead of an empty transcript, which is
// the truthful answer to "transcribe nothing".
engine::runtime::TaskResult run_streaming_live(
const LoadedModel::Session &session,
const engine::runtime::TaskRequest &request,
const std::function<bool(std::vector<float> &)> &next_frames,
const std::function<void(const engine::runtime::StreamEvent &)> &on_event,
LaneEntry &lane);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,153 @@
#include "model_options.h"
#include <cctype>
#include <cerrno>
#include <climits>
#include <cstdlib>
namespace audiocpp_backend {
namespace {
std::string trim(const std::string &value) {
size_t begin = 0;
while (begin < value.size() &&
std::isspace(static_cast<unsigned char>(value[begin])) != 0) {
++begin;
}
size_t end = value.size();
while (end > begin &&
std::isspace(static_cast<unsigned char>(value[end - 1])) != 0) {
--end;
}
return value.substr(begin, end - begin);
}
// Parses a non-negative integer. Returns false on anything else, including
// empty strings, signs, trailing garbage, and values too large for int.
//
// strtol rather than atoi: atoi is undefined behaviour once the digits exceed
// long, and in practice it hands back a wrapped value. That would let
// "device:2147483648" through as -2147483648 and send a negative index to the
// ggml backend selector, from a function whose error text promises the caller a
// non-negative integer.
bool parse_non_negative_int(const std::string &value, int &out) {
if (value.empty()) {
return false;
}
for (const char ch : value) {
if (std::isdigit(static_cast<unsigned char>(ch)) == 0) {
return false;
}
}
errno = 0;
char *end = nullptr;
const long parsed = std::strtol(value.c_str(), &end, 10);
if (errno == ERANGE || end == nullptr || *end != '\0') {
return false;
}
if (parsed < 0 || parsed > INT_MAX) {
return false;
}
out = static_cast<int>(parsed);
return true;
}
bool starts_with(const std::string &value, const std::string &prefix) {
return value.size() >= prefix.size() &&
value.compare(0, prefix.size(), prefix) == 0;
}
} // namespace
ParsedOptions parse_model_options(const std::vector<std::string> &entries) {
ParsedOptions parsed;
for (const auto &raw : entries) {
const std::string entry = trim(raw);
if (entry.empty()) {
continue;
}
// Split on the FIRST colon: values are often paths that contain more.
const size_t sep = entry.find(':');
if (sep == std::string::npos) {
parsed.error = "audio-cpp: option '" + entry +
"' is not in key:value form";
return parsed;
}
const std::string key = trim(entry.substr(0, sep));
const std::string value = trim(entry.substr(sep + 1));
if (starts_with(key, "load.")) {
const std::string inner = key.substr(5);
if (inner.empty()) {
parsed.error = "audio-cpp: option '" + entry +
"' has an empty load option name";
return parsed;
}
parsed.options.load_options[inner] = value;
continue;
}
if (starts_with(key, "session.")) {
const std::string inner = key.substr(8);
if (inner.empty()) {
parsed.error = "audio-cpp: option '" + entry +
"' has an empty session option name";
return parsed;
}
parsed.options.session_options[inner] = value;
continue;
}
if (key == "family") {
parsed.options.family = value;
} else if (key == "task") {
parsed.options.task = value;
} else if (key == "backend") {
parsed.options.backend = value;
} else if (key == "model_spec_override") {
parsed.options.model_spec_override = value;
} else if (key == "device") {
if (!parse_non_negative_int(value, parsed.options.device)) {
parsed.error = "audio-cpp: option 'device' needs a non-negative "
"integer, got '" + value + "'";
return parsed;
}
parsed.options.device_set = true;
} else if (key == "threads") {
if (!parse_non_negative_int(value, parsed.options.threads)) {
parsed.error = "audio-cpp: option 'threads' needs a non-negative "
"integer, got '" + value + "'";
return parsed;
}
} else if (key == "busy_timeout_ms") {
if (!parse_non_negative_int(value, parsed.options.busy_timeout_ms)) {
parsed.error = "audio-cpp: option 'busy_timeout_ms' needs a "
"non-negative integer, got '" + value + "'";
return parsed;
}
} else if (key == "live_idle_timeout_ms") {
if (!parse_non_negative_int(value,
parsed.options.live_idle_timeout_ms)) {
parsed.error = "audio-cpp: option 'live_idle_timeout_ms' needs a "
"non-negative integer, got '" + value + "'";
return parsed;
}
} else {
// Quotes the whole entry, not just the key: an entry like ":value"
// has an empty key and would otherwise leave nothing to grep for.
parsed.error = "audio-cpp: unknown option key '" + entry +
"'. Known keys: family, task, backend, device, "
"threads, model_spec_override, busy_timeout_ms, "
"live_idle_timeout_ms, load.<key>, session.<key>";
return parsed;
}
}
return parsed;
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,66 @@
#pragma once
// Parses the model YAML's `options:` list (ModelOptions.Options in
// backend.proto) into a struct. Standard library only: this unit is compiled
// and tested by backend/cpp/run-unit-tests.sh without an audio.cpp checkout.
#include <map>
#include <string>
#include <vector>
namespace audiocpp_backend {
struct ModelOptions {
// audio.cpp model family. Empty means "derive from the GGUF's embedded
// audiocpp.model_spec.family key"; a non-GGUF path with an empty family is
// rejected at load time, not here.
std::string family;
// Pins the audio.cpp task, overriding RPC-based routing. Empty means route.
std::string task;
// ggml backend: cpu, cuda, vulkan, metal, best.
std::string backend = "cpu";
int device = 0;
// True once a `device:` entry has been seen. 0 is both the default and a
// legitimate device index, so the value alone cannot tell an explicit
// `device:0` from an unset option, and a caller merging in its own fallback
// would silently override the explicit choice.
bool device_set = false;
// 0 means "let the runtime decide".
int threads = 0;
std::string model_spec_override;
// 0 disables the run guard's fail-fast, restoring an unbounded wait.
int busy_timeout_ms = 0;
// How long AudioTranscriptionLive waits for the next audio frame before it
// cancels the stream and gives the model's lane back. 0 means NO LIMIT.
//
// It exists because that RPC holds the lane for the whole stream, so a peer
// that stops sending WITHOUT closing blocks every other request against this
// model for as long as its socket stays up. No other RPC can do that: they
// hold the lane across compute, which ends on its own.
//
// 30 seconds, and the number is picked from what the only in-tree client
// does. core/http/endpoints/openai/realtime.go drives a 300 ms ticker and
// feeds every tick that produced new audio while a turn is open, so 30 s of
// silence is a hundred ticks that delivered nothing: the peer is gone, or
// its socket is wedged. It is also comfortably longer than any pause a
// speaker takes mid-utterance, which is the case that must never be cut off,
// and backend.proto allows one stream to span many utterances, so a client
// that pauses for longer than this between them should raise it rather than
// discover it. Lowering it below a few seconds risks cancelling a live
// speaker; 0 turns the limit off for a client that legitimately idles.
int live_idle_timeout_ms = 30000;
// `load.<key>:<value>` entries, prefix stripped.
std::map<std::string, std::string> load_options;
// `session.<key>:<value>` entries, prefix stripped.
std::map<std::string, std::string> session_options;
};
struct ParsedOptions {
ModelOptions options;
// Non-empty means the caller must fail the load with INVALID_ARGUMENT.
std::string error;
};
ParsedOptions parse_model_options(const std::vector<std::string> &entries);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,165 @@
// Unit tests for model_options. Standard library only, so
// backend/cpp/run-unit-tests.sh picks this up with no engine checkout.
//
// The harness compiles this file as a single translation unit with no other
// sources, so the implementation is included directly rather than linked.
//
// Build and run standalone:
// g++ -std=c++17 -I. model_options_test.cpp -o t && ./t
#include "model_options.cpp"
#include <cstdio>
#include <map>
#include <string>
#include <vector>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
// Returns the mapped value, or an empty string when the key is absent. map::at
// would throw on a miss and abort the whole binary, so one prefix off-by-one
// would hide every check that follows it instead of failing a single one.
static std::string lookup(const std::map<std::string, std::string> &values,
const std::string &key) {
const auto found = values.find(key);
return found == values.end() ? std::string() : found->second;
}
using audiocpp_backend::parse_model_options;
static void test_defaults() {
auto r = parse_model_options({});
check(r.error.empty(), "empty option list is not an error");
check(r.options.family.empty(), "family defaults to empty");
check(r.options.task.empty(), "task defaults to empty");
check(r.options.backend == "cpu", "backend defaults to cpu");
check(r.options.device == 0, "device defaults to 0");
check(r.options.threads == 0, "threads defaults to 0");
check(r.options.busy_timeout_ms == 0, "busy_timeout_ms defaults to 0");
// NOT zero, unlike every other numeric option here. A live stream holds the
// model's lane while it waits on the client, so the default has to bound
// that wait; 0 is the explicit "no limit" the operator opts into.
check(r.options.live_idle_timeout_ms == 30000,
"live_idle_timeout_ms defaults to 30000");
check(r.options.load_options.empty(), "load_options defaults empty");
check(r.options.session_options.empty(), "session_options defaults empty");
}
static void test_scalar_options() {
auto r = parse_model_options({
"family:qwen3_tts",
"task:tts",
"backend:cuda",
"device:1",
"threads:8",
"busy_timeout_ms:30000",
"live_idle_timeout_ms:5000",
});
check(r.error.empty(), "scalar options parse without error");
check(r.options.family == "qwen3_tts", "family parsed");
check(r.options.task == "tts", "task parsed");
check(r.options.backend == "cuda", "backend parsed");
check(r.options.device == 1, "device parsed");
check(r.options.threads == 8, "threads parsed");
check(r.options.busy_timeout_ms == 30000, "busy_timeout_ms parsed");
check(r.options.live_idle_timeout_ms == 5000, "live_idle_timeout_ms parsed");
check(parse_model_options({"live_idle_timeout_ms:0"}).options.live_idle_timeout_ms == 0,
"an explicit 0 turns the live idle limit off rather than reverting to "
"the default");
}
// Values containing colons must survive: split on the FIRST colon only.
static void test_value_containing_colon() {
auto r = parse_model_options({"model_spec_override:/models/a:b/spec.json"});
check(r.error.empty(), "colon-bearing value is not an error");
check(r.options.model_spec_override == "/models/a:b/spec.json",
"value keeps every colon after the first separator");
}
static void test_namespaced_options() {
auto r = parse_model_options({
"load.weight_type:q8_0",
"session.miocodec.weight_type:f16",
"session.graph_capacity:tiered",
});
check(r.error.empty(), "namespaced options parse without error");
check(r.options.load_options.size() == 1, "one load option");
check(lookup(r.options.load_options, "weight_type") == "q8_0", "load prefix stripped");
check(r.options.session_options.size() == 2, "two session options");
check(lookup(r.options.session_options, "miocodec.weight_type") == "f16",
"session prefix stripped, inner dots kept");
check(lookup(r.options.session_options, "graph_capacity") == "tiered",
"second session option parsed");
}
static void test_errors() {
check(!parse_model_options({"family"}).error.empty(),
"entry without a colon is rejected");
check(!parse_model_options({"nonsense:1"}).error.empty(),
"unknown key is rejected");
check(!parse_model_options({"device:abc"}).error.empty(),
"non-numeric device is rejected");
check(!parse_model_options({"threads:-1"}).error.empty(),
"negative threads is rejected");
check(!parse_model_options({"load.:x"}).error.empty(),
"empty load key is rejected");
check(!parse_model_options({"session.:x"}).error.empty(),
"empty session key is rejected");
check(!parse_model_options({"busy_timeout_ms:abc"}).error.empty(),
"non-numeric busy_timeout_ms is rejected");
check(!parse_model_options({"live_idle_timeout_ms:abc"}).error.empty(),
"non-numeric live_idle_timeout_ms is rejected");
check(!parse_model_options({"live_idle_timeout_ms:-1"}).error.empty(),
"negative live_idle_timeout_ms is rejected");
check(!parse_model_options({"device:-1"}).error.empty(),
"negative device is rejected");
check(!parse_model_options({"threads:x"}).error.empty(),
"non-numeric threads is rejected");
// Values too large for int must be rejected, not silently wrapped into a
// negative device index that then reaches the ggml backend selector.
check(!parse_model_options({"device:2147483648"}).error.empty(),
"device above INT_MAX is rejected");
check(!parse_model_options({"threads:99999999999999"}).error.empty(),
"threads above INT_MAX is rejected");
// The error text must name the offending entry so a user can fix their YAML.
const auto r = parse_model_options({"nonsense:1"});
check(r.error.find("nonsense") != std::string::npos,
"error names the offending key");
// An empty key still has to give the user something to grep for.
const auto empty_key = parse_model_options({":value"});
check(empty_key.error.find(":value") != std::string::npos,
"unknown-key error names the entry even when the key is empty");
}
static void test_blank_entries_ignored() {
auto r = parse_model_options({"", " ", "family:supertonic"});
check(r.error.empty(), "blank entries are skipped, not rejected");
check(r.options.family == "supertonic", "real entry still parsed");
}
int main() {
test_defaults();
test_scalar_options();
test_value_containing_colon();
test_namespaced_options();
test_errors();
test_blank_entries_ignored();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all model_options checks passed\n");
return 0;
}

228
backend/cpp/audio-cpp/package.sh Executable file
View File

@@ -0,0 +1,228 @@
#!/bin/bash
# Assemble backend/cpp/audio-cpp/package, which becomes the whole content of the
# FROM scratch backend image. Nothing outside this directory exists at run time.
set -euo pipefail
CURDIR=$(dirname "$(realpath "$0")")
REPO_ROOT="${CURDIR}/../../.."
PACKAGE_DIR="$CURDIR/package"
BUILD_DIR="$CURDIR/build"
rm -rf "$PACKAGE_DIR"
mkdir -p "$PACKAGE_DIR/lib" "$PACKAGE_DIR/assets"
cp -avf "$CURDIR/grpc-server" "$PACKAGE_DIR/"
cp -fv "$CURDIR/run.sh" "$PACKAGE_DIR/"
# ENGINE_ENABLE_CPU_ALL_VARIANTS builds the ggml backends as shared objects that
# are dlopened at run time, so ldd cannot see them and the dependency walk below
# would leave the image with no CPU backend at all. They also cannot go in lib/:
# ggml DISCOVERS them by listing dirname(/proc/self/exe) and the current
# directory, so being on a library path is not enough, they have to be in a
# directory ggml scans. run.sh execs the bundled loader from the package root
# exactly so that directory is this one. cmake writes them to build/bin, not
# next to build/grpc-server, which is why this reads from bin/.
#
# -a keeps the libggml.so -> libggml.so.0 -> libggml.so.0.12.0 symlink chain,
# so the SONAME the binary asks for still names a file here.
for pattern in '*.so*' '*.dylib*'; do
if compgen -G "$BUILD_DIR/bin/$pattern" > /dev/null; then
# shellcheck disable=SC2086
cp -avf "$BUILD_DIR/bin/"$pattern "$PACKAGE_DIR/"
fi
done
# Upstream ships silero_vad and marblenet_vad as small runtime assets.
# resolve_model_path() expands "bundled:<name>" to
# dirname(/proc/self/exe)/assets/<name>, so copying them here is what makes VAD
# work with nothing downloaded.
for asset in silero_vad marblenet_vad; do
src="$CURDIR/audio.cpp/assets/framework/models/$asset"
if [ -d "$src" ]; then
cp -rfv "$src" "$PACKAGE_DIR/assets/"
else
echo "package.sh: bundled asset missing: $src" >&2
echo "package.sh: run 'make audio.cpp' before packaging" >&2
exit 1
fi
done
# Everything below this point is Linux-only: a bundled ELF loader, an ldd walk
# and an ld.so --list validation. The macOS equivalent is the otool -L closure in
# scripts/build/audio-cpp-darwin.sh, which picks up from the exit 0 below.
#
# WARNING FOR ANYONE REWORKING THAT SCRIPT. The obvious move is to copy
# scripts/build/privacy-filter-darwin.sh, and that script assembles its own
# package under build/darwin and never calls package.sh at all. Adapted as-is it
# will silently omit assets/, and the bundled: model path form then resolves to
# nothing, which takes the only zero-download verification path in this backend
# with it. That is why audio-cpp-darwin.sh copies THIS directory instead of
# rebuilding one. The Darwin package needs the same root-level layout as the
# Linux one: grpc-server, run.sh, the ggml dylibs and assets/ in ONE directory,
# with lib/ for the rest. run.sh's Darwin branch execs grpc-server directly, so
# _NSGetExecutablePath already names the package root; nothing else is needed
# beyond putting the files there.
UNAME_S=$(uname -s)
if [ "$UNAME_S" = "Darwin" ]; then
echo "package.sh: Darwin dylib bundling is deferred to scripts/build/audio-cpp-darwin.sh"
ls -lah "$PACKAGE_DIR/" "$PACKAGE_DIR/assets/"
exit 0
fi
# The loader goes in the package ROOT, not in lib/. run.sh explains why at
# length; the short version is that exec'ing it makes dirname(/proc/self/exe)
# the directory it sits in, and both the ggml backend scan and the bundled:
# asset lookup need that to be the package root.
if [ -f "/lib64/ld-linux-x86-64.so.2" ]; then
cp -arfLv /lib64/ld-linux-x86-64.so.2 "$PACKAGE_DIR/ld.so"
elif [ -f "/lib/ld-linux-aarch64.so.1" ]; then
cp -arfLv /lib/ld-linux-aarch64.so.1 "$PACKAGE_DIR/ld.so"
else
echo "package.sh: unknown architecture" >&2
exit 1
fi
# THE LAYOUT ASSERTION. Everything else in this script checks that the package
# can LINK. This checks that it can RESOLVE, which is a different property and
# the one with no other guard on it.
#
# The loader, assets/ and the dlopened ggml objects only agree while they share
# one directory, because run.sh execs the loader and all three are reached
# through dirname(/proc/self/exe). A tidy-up that moves the loader into lib/,
# following the llama-cpp layout, produces a package that builds, ships, and
# then fails at run time with "model path does not exist: <pkg>/lib/assets/..."
# or "Failed to initialize CPU backend". Fail the build instead.
#
# This sits immediately after the loader copy rather than at the end of the
# script on purpose: everything below dereferences $PACKAGE_DIR/ld.so, so a
# misplaced loader would otherwise surface as "No such file or directory" from
# the validation gate and never reach an assertion that could explain it.
if [ ! -f "$PACKAGE_DIR/ld.so" ]; then
echo "package.sh: the bundled loader must be at the package root, not in lib/." >&2
echo "package.sh: run.sh execs it, so its directory is dirname(/proc/self/exe)," >&2
echo "package.sh: which is where resolve_model_path looks for assets/ and where" >&2
echo "package.sh: ggml looks for the CPU variants." >&2
exit 1
fi
if [ ! -d "$PACKAGE_DIR/assets" ]; then
echo "package.sh: assets/ must sit beside the loader at the package root." >&2
exit 1
fi
# Only assert the ggml half when this build produced CPU variants at all: a
# cublas or vulkan build links ggml statically and ships none.
if compgen -G "$BUILD_DIR/bin/libggml-cpu-*.so" > /dev/null && \
! compgen -G "$PACKAGE_DIR/libggml-cpu-*.so" > /dev/null; then
echo "package.sh: the build produced libggml-cpu-*.so but none reached the" >&2
echo "package.sh: package root, so ggml's scan of dirname(/proc/self/exe)" >&2
echo "package.sh: will find no CPU backend." >&2
exit 1
fi
# Libraries the host GPU driver stack owns. package_gpu_libs deliberately ships
# the CUDA/Vulkan runtime but not the driver, because the driver has to match
# the kernel module on whatever host runs the image. Copying the build host's
# copy in would pin it to the build host instead.
#
# One regex, used by both the copy loop and the validation gate below. They have
# to agree: exempting a library from the copy but not from the gate makes the
# gate reject the very absence the copy loop just created.
DRIVER_LIB_RE='^(libcuda\.so|libnvidia-)'
# awk applies string-escape processing to a -v assignment before compiling the
# regex, so a lone backslash is eaten and awk warns about it. Double them here
# rather than keeping a second hand-written copy of the pattern, which is the
# drift this single-source-of-truth exists to prevent.
DRIVER_LIB_RE_AWK=${DRIVER_LIB_RE//\\/\\\\}
is_driver_lib() {
[[ "$(basename "$1")" =~ $DRIVER_LIB_RE ]]
}
# Bundle the full dependency closure. grpc-server links the distro gRPC,
# protobuf and absl stack; copying only the C/C++ runtime leaves the scratch
# image unable to start. The walk runs over the PACKAGED binary, not the one in
# $CURDIR, because its RUNPATH is $ORIGIN: only from inside the package does
# libggml.so.0 resolve to the copy shipped above rather than to nothing.
# The dlopened ggml objects are walked too, since a dependency of theirs that
# grpc-server does not itself link would otherwise be missed.
{
ldd "$PACKAGE_DIR/grpc-server"
for so in "$PACKAGE_DIR"/*.so*; do
[ -f "$so" ] || continue
ldd "$so"
done
} | awk '$2 == "=>" && $3 ~ /^\// { print $3 }' | sort -u | \
while read -r so; do
# Skip what is already inside the package: the ggml objects resolve through
# $ORIGIN and re-copying them into lib/ would ship two copies of each.
case "$so" in "$PACKAGE_DIR"/*) continue ;; esac
if is_driver_lib "$so"; then
echo "package.sh: leaving driver-owned library to the host: $so"
continue
fi
cp -arfLv "$so" "$PACKAGE_DIR/lib/"
done
GPU_LIB_SCRIPT="${REPO_ROOT}/scripts/build/package-gpu-libs.sh"
if [ -f "$GPU_LIB_SCRIPT" ]; then
echo "Packaging GPU libraries for BUILD_TYPE=${BUILD_TYPE:-cpu}..."
# shellcheck source=/dev/null
source "$GPU_LIB_SCRIPT" "$PACKAGE_DIR/lib"
package_gpu_libs
fi
# Resolve every dependency through the same loader and library path the
# from-scratch image uses. Two distinct failures are rejected, because the
# loader can still fall back to the host's default directories: a dependency it
# could not resolve at all, and one it resolved to a file OUTSIDE the package,
# which would validate here and be absent in the image.
#
# The driver libraries are exempt from BOTH rejections, and that exemption is
# load-bearing on GPU builds rather than tidiness. With BUILD_TYPE=cublas ggml
# is static (no CPU_ALL_VARIANTS), and ggml/CMakeLists.txt defaults
# GGML_CUDA_NO_VMM=OFF, so ggml-cuda links CUDA::cuda_driver and grpc-server
# itself carries DT_NEEDED libcuda.so.1. The copy loop above deliberately leaves
# that to the host, so inside the CUDA builder it resolves either to a host path
# or to nothing. Without this exemption every cublas build would fail here and
# CI would produce no image at all.
#
# LD_TRACE_LOADED_OBJECTS + LD_LIBRARY_PATH, NOT `ld.so --library-path --list`,
# and the difference is not cosmetic. Measured on a stub object built to
# DT_NEEDED an absent libcuda.so.1: `--list` refuses to trace at all, printing
# "libdrivertest.so: error while loading shared libraries: libcuda.so.1: cannot
# open shared object file" and exiting 127, so no per-library line is ever
# produced and no exemption below could apply. The env form prints
# "libcuda.so.1 => not found" and exits 0, which is what makes both the
# unresolved rule and its driver exemption reachable. It is also closer to what
# run.sh actually does, since run.sh exports LD_LIBRARY_PATH rather than passing
# --library-path.
validation_failed=0
validate_object() {
local object="$1"
LD_TRACE_LOADED_OBJECTS=1 LD_LIBRARY_PATH="$PACKAGE_DIR/lib:$PACKAGE_DIR" \
"$PACKAGE_DIR/ld.so" "$object" | awk -v pkg="$PACKAGE_DIR/" -v obj="$object" \
-v driver_re="$DRIVER_LIB_RE_AWK" '
function base(p, n, parts) { n = split(p, parts, "/"); return parts[n] }
$2 == "=>" && $3 == "not" {
if ($1 ~ driver_re) next
print "package.sh: unresolved dependency of " obj ": " $1 > "/dev/stderr"
bad = 1
}
$2 == "=>" && $3 ~ /^\// && index($3, pkg) != 1 {
if (base($3) ~ driver_re) next
print "package.sh: dependency of " obj " resolved outside the package: " $0 > "/dev/stderr"
bad = 1
}
END { exit bad }
'
}
validate_object "$PACKAGE_DIR/grpc-server" || validation_failed=1
for so in "$PACKAGE_DIR"/*.so*; do
[ -f "$so" ] || continue
validate_object "$so" || validation_failed=1
done
if [ "$validation_failed" -ne 0 ]; then
exit 1
fi
echo "audio-cpp package contents:"
ls -lah "$PACKAGE_DIR/" "$PACKAGE_DIR/lib/" "$PACKAGE_DIR/assets/"

View File

@@ -0,0 +1,85 @@
#include "result_map.h"
#include "transcript_assembly.h"
#include <string>
#include <vector>
namespace audiocpp_backend {
void fill_transcript_result(const engine::runtime::TaskResult &result,
int sample_rate, float duration_seconds,
backend::TranscriptResult *out) {
// No null guard on `out`, deliberately. gRPC always hands a handler a
// response message, so a null here would be a programming error in a
// caller, and a guard that returned quietly would answer the client with an
// untouched, empty transcript and an OK status. That is the same
// indistinguishable-from-silence failure the rest of this unit exists to
// prevent; crashing on the developer's machine is the cheaper outcome.
std::vector<Span> speech_segments;
speech_segments.reserve(result.speech_segments.size());
for (const auto &segment : result.speech_segments) {
speech_segments.push_back(
Span{segment.span.start_sample, segment.span.end_sample});
}
std::vector<SpeakerSpan> speaker_turns;
speaker_turns.reserve(result.speaker_turns.size());
for (const auto &turn : result.speaker_turns) {
speaker_turns.push_back(
SpeakerSpan{Span{turn.span.start_sample, turn.span.end_sample},
turn.speaker_id});
}
std::vector<WordSpan> words;
words.reserve(result.word_timestamps.size());
for (const auto &word : result.word_timestamps) {
words.push_back(
WordSpan{Span{word.span.start_sample, word.span.end_sample},
word.word});
}
// The ONLY read of transcript text in this function, and the only one there
// may ever be. See THE RULE in the header.
const std::string text =
result.text_output.has_value() ? result.text_output->text : std::string();
const AssembledTranscript assembled = assemble_transcript(
text, speech_segments, speaker_turns, words, sample_rate);
out->set_text(assembled.text);
// language has no source inside transcript_assembly, which is span-shaped
// only, so it is read straight off the engine result here. Left untouched
// when the family reported no text output at all: an empty string would be
// indistinguishable from a family that genuinely detected no language, and
// the field is documented as optional.
if (result.text_output.has_value()) {
out->set_language(result.text_output->language);
}
out->set_duration(duration_seconds);
// Cleared rather than appended to. A caller that fills the same message
// twice (a stream's final_result being rebuilt, say) would otherwise emit
// every segment twice, and the second call's ids would restart at 0 and
// collide with the first call's.
out->clear_segments();
for (const auto &segment : assembled.segments) {
auto *out_segment = out->add_segments();
out_segment->set_id(segment.id);
// NANOSECONDS. TranscriptSegment and TranscriptWord are the only
// messages in backend.proto that use them; VADSegment and DiarizeSegment
// are float seconds. assemble_transcript has already converted.
out_segment->set_start(segment.start_ns);
out_segment->set_end(segment.end_ns);
out_segment->set_text(segment.text);
out_segment->set_speaker(segment.speaker);
for (const auto &word : segment.words) {
auto *out_word = out_segment->add_words();
out_word->set_start(word.start_ns);
out_word->set_end(word.end_ns);
out_word->set_text(word.text);
}
}
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,37 @@
#pragma once
// Converts engine::runtime results into LocalAI proto messages. All of the
// non-trivial shaping lives in transcript_assembly, which is stdlib-only and
// unit tested; this unit is the thin engine-typed boundary around it.
#include "backend.pb.h"
#include "engine/framework/runtime/session.h"
namespace audiocpp_backend {
// Fills text, language, duration, segments and per-segment words.
//
// THE RULE: the top-level text is TaskResult.text_output verbatim. It is never
// derived from segments or words. audio.cpp carries transcript text in
// text_output and nowhere else: speech_segments, speaker_turns and
// word_timestamps carry spans and labels and no text at all. Deriving the
// transcript from them therefore returns an EMPTY text for every producer that
// reports segments without word timing, which real VibeVoice diarized ASR does.
// An earlier attempt at this backend shipped exactly that bug. assemble_transcript
// enforces the rule and is heavily tested; this unit's job is not to re-derive
// it but to not undo it at the proto boundary.
//
// `sample_rate` is the rate the result's spans are expressed in, which is the
// rate of the AudioBuffer that was handed to the session, NOT the rate of the
// file the caller uploaded. Those differ whenever read_audio_file resampled,
// which is why the handler passes the buffer's rate rather than the file's.
//
// Segments are replaced, not appended to, so a message filled twice does not
// accumulate. `out` must be non-null and is not checked; see the note at the
// top of the implementation for why that is not an oversight.
void fill_transcript_result(const engine::runtime::TaskResult &result,
int sample_rate, float duration_seconds,
backend::TranscriptResult *out);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,247 @@
// Tests for result_map, the engine-to-proto boundary.
//
// NAMED _ctest AND NOT _test ON PURPOSE. backend/cpp/run-unit-tests.sh globs
// every *_test.cpp under backend/cpp/ and compiles it as a single standalone
// translation unit with no include path beyond its own directory. This file
// needs backend.pb.h and the audio.cpp framework headers, so it is built and
// run by ctest instead:
//
// make -C backend/cpp/audio-cpp test-engine
//
// Renaming it to *_test.cpp would break the standalone suite for every backend.
//
// What is worth testing here is exactly one thing, and it is not the field
// copying: THE RULE. TaskResult carries transcript text in text_output and
// nowhere else, so the proto's text must be that string verbatim. An earlier
// attempt at this backend derived it from the segments, which returns an empty
// transcript for every producer that reports segments without word timing.
// transcript_assembly already enforces the rule and is tested on its own; these
// checks are here so that a future edit cannot undo it at the boundary.
#include "result_map.h"
#include <cstdio>
#include <string>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
using namespace audiocpp_backend;
namespace rt = engine::runtime;
static const int kRate = 16000;
static rt::SpeechSegment speech(std::int64_t start, std::int64_t end) {
rt::SpeechSegment segment;
segment.span.start_sample = start;
segment.span.end_sample = end;
return segment;
}
static rt::SpeakerTurn turn(std::int64_t start, std::int64_t end,
const std::string &speaker) {
rt::SpeakerTurn out;
out.span.start_sample = start;
out.span.end_sample = end;
out.speaker_id = speaker;
return out;
}
static rt::WordTimestamp word(std::int64_t start, std::int64_t end,
const std::string &text) {
rt::WordTimestamp out;
out.span.start_sample = start;
out.span.end_sample = end;
out.word = text;
return out;
}
// THE REGRESSION. A diarized ASR result: real text, real speaker turns, and no
// word timing at all. This is the vibevoice_asr shape, and it is the one that
// came back empty before.
static void test_text_survives_segments_without_words() {
rt::TaskResult result;
rt::Transcript transcript;
transcript.text = "hello there general kenobi";
transcript.language = "en";
result.text_output = transcript;
result.speaker_turns.push_back(turn(0, 16000, "speaker_0"));
result.speaker_turns.push_back(turn(16000, 32000, "speaker_1"));
backend::TranscriptResult out;
fill_transcript_result(result, kRate, 2.0f, &out);
check(out.text() == "hello there general kenobi",
"diarized result keeps text_output verbatim");
check(out.language() == "en", "language comes from text_output");
check(out.segments_size() == 2, "both speaker turns become segments");
if (out.segments_size() == 2) {
check(out.segments(0).speaker() == "speaker_0",
"first segment keeps its own speaker label");
check(out.segments(1).speaker() == "speaker_1",
"second segment keeps its own speaker label");
check(out.segments(1).start() == 1000000000LL,
"segment start is nanoseconds, not samples");
check(out.segments(1).end() == 2000000000LL,
"segment end is nanoseconds, not samples");
}
}
// The same rule seen from the other side: text present, spans present, and the
// per-segment text empty because there is nothing truthful to split. A boundary
// that derived the top-level text from these segments would produce "".
static void test_speech_segments_do_not_supply_the_text() {
rt::TaskResult result;
rt::Transcript transcript;
transcript.text = "one two three";
result.text_output = transcript;
result.speech_segments.push_back(speech(0, 8000));
result.speech_segments.push_back(speech(8000, 16000));
backend::TranscriptResult out;
fill_transcript_result(result, kRate, 1.0f, &out);
check(out.text() == "one two three",
"speech segments without words do not empty the transcript");
check(out.segments_size() == 2, "both speech segments are emitted");
if (out.segments_size() == 2) {
check(out.segments(0).text().empty() && out.segments(1).text().empty(),
"per-segment text stays empty when there is no word timing");
}
}
static void test_words_reach_the_proto_in_nanoseconds() {
rt::TaskResult result;
rt::Transcript transcript;
transcript.text = "hi there";
result.text_output = transcript;
result.word_timestamps.push_back(word(0, 8000, "hi"));
result.word_timestamps.push_back(word(8000, 16000, "there"));
backend::TranscriptResult out;
fill_transcript_result(result, kRate, 1.0f, &out);
check(out.text() == "hi there", "word-timed result keeps text_output");
check(out.segments_size() == 1, "words with no spans yield one covering segment");
if (out.segments_size() == 1) {
const auto &segment = out.segments(0);
check(segment.words_size() == 2, "both words are emitted");
if (segment.words_size() == 2) {
check(segment.words(0).text() == "hi", "first word text");
check(segment.words(0).start() == 0, "first word start");
check(segment.words(0).end() == 500000000LL,
"first word end is 0.5 s in nanoseconds");
check(segment.words(1).start() == 500000000LL, "second word start");
check(segment.words(1).end() == 1000000000LL, "second word end");
}
}
}
// The buffer's rate, not the file's, is what the spans mean. Passing 8000 for
// the same spans has to halve every timestamp, which is what makes resampling
// the input at read time load-bearing rather than cosmetic.
static void test_sample_rate_scales_the_timestamps() {
rt::TaskResult result;
rt::Transcript transcript;
transcript.text = "x";
result.text_output = transcript;
result.speech_segments.push_back(speech(0, 8000));
backend::TranscriptResult out;
fill_transcript_result(result, 8000, 1.0f, &out);
check(out.segments_size() == 1, "one segment at 8 kHz");
if (out.segments_size() == 1) {
check(out.segments(0).end() == 1000000000LL,
"8000 samples at 8 kHz is one second");
}
}
static void test_duration_is_carried_through() {
rt::TaskResult result;
rt::Transcript transcript;
transcript.text = "x";
result.text_output = transcript;
backend::TranscriptResult out;
fill_transcript_result(result, kRate, 14.07f, &out);
check(out.duration() > 14.06f && out.duration() < 14.08f,
"duration is set from the argument");
}
// No text output at all. A VAD-shaped result reaching this boundary must not
// invent a transcript, and must not overwrite a language the caller had already
// decided on.
static void test_missing_text_output_leaves_language_alone() {
rt::TaskResult result;
result.speech_segments.push_back(speech(0, 16000));
backend::TranscriptResult out;
out.set_language("it");
fill_transcript_result(result, kRate, 1.0f, &out);
check(out.text().empty(), "no text_output means no text");
check(out.language() == "it",
"a result with no text_output does not clear the language");
check(out.segments_size() == 1, "spans are still emitted");
}
// Filling the same message twice must replace, not accumulate: the second
// call's ids restart at 0 and would collide with the first call's.
static void test_refilling_replaces_the_segments() {
rt::TaskResult first;
rt::Transcript transcript;
transcript.text = "first";
first.text_output = transcript;
first.speech_segments.push_back(speech(0, 16000));
first.speech_segments.push_back(speech(16000, 32000));
backend::TranscriptResult out;
fill_transcript_result(first, kRate, 2.0f, &out);
rt::TaskResult second;
rt::Transcript replacement;
replacement.text = "second";
second.text_output = replacement;
second.speech_segments.push_back(speech(0, 16000));
fill_transcript_result(second, kRate, 1.0f, &out);
check(out.text() == "second", "the second fill replaces the text");
check(out.segments_size() == 1,
"the second fill replaces the segments instead of appending");
}
static void test_empty_result_is_empty() {
rt::TaskResult result;
backend::TranscriptResult out;
fill_transcript_result(result, kRate, 0.0f, &out);
check(out.text().empty(), "empty result has no text");
check(out.segments_size() == 0, "empty result has no segments");
}
int main() {
test_text_survives_segments_without_words();
test_speech_segments_do_not_supply_the_text();
test_words_reach_the_proto_in_nanoseconds();
test_sample_rate_scales_the_timestamps();
test_duration_is_carried_through();
test_missing_text_output_leaves_language_alone();
test_refilling_replaces_the_segments();
test_empty_result_is_empty();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all result_map checks passed\n");
return 0;
}

58
backend/cpp/audio-cpp/run.sh Executable file
View File

@@ -0,0 +1,58 @@
#!/bin/bash
# Entry point for the audio-cpp backend image and for BACKEND_BINARY mode.
#
# The image's final stage is FROM scratch, so the package root is / and there is
# no system loader, no system libc and no fallback library path. Everything the
# process opens has to be inside the package, and it has to be findable by the
# two mechanisms that actually do the finding: the dynamic linker, and
# audio.cpp's own directory scans.
set -e
CURDIR=$(dirname "$(realpath "$0")")
if [ "$(uname -s)" = "Darwin" ]; then
export DYLD_LIBRARY_PATH="$CURDIR/lib:$CURDIR:$DYLD_LIBRARY_PATH"
exec "$CURDIR/grpc-server" "$@"
fi
# $CURDIR is on the path as well as $CURDIR/lib: the ggml shared objects the
# CPU-all-variants build produces sit in the package root, next to the binary,
# not in lib/. See the comment below for why they cannot live in lib/.
export LD_LIBRARY_PATH="$CURDIR/lib:$CURDIR:$LD_LIBRARY_PATH"
# THE BUNDLED LOADER IS AT THE PACKAGE ROOT, NOT AT lib/ld.so. DO NOT MOVE IT.
#
# Exec'ing the loader is what pins the bundled glibc to the matching ld.so, and
# every other C++ backend here does it. The cost is that /proc/self/exe then
# names the LOADER rather than grpc-server, and this backend has two consumers
# of /proc/self/exe that both have to land on the package root:
#
# - ggml's backend registry DISCOVERS the per-microarch libggml-cpu-*.so by
# scanning dirname(/proc/self/exe) and the current directory. Those files
# are dlopened, never linked, so no library path and no RUNPATH reaches
# them: they have to be in a directory ggml scans.
# - resolve_model_path() turns "bundled:<name>" into
# dirname(/proc/self/exe)/assets/<name>, which is how the bundled
# silero_vad and marblenet_vad models resolve with nothing downloaded.
#
# backend/cpp/llama-cpp/package.sh answers the first of these by keeping
# lib/ld.so and moving the ggml objects INTO lib/. That does not generalise
# here, because it would also drag assets/ into lib/ to keep the second
# consumer working. Putting the loader in the package root instead makes
# dirname(/proc/self/exe) the package root, so the binary, the ggml objects and
# assets/ all sit in the one directory that all three mechanisms agree on.
#
# The ggml half has a second chance that the bundled: half does not: LocalAI
# sets the backend process cwd to the directory holding run.sh
# (pkg/model/process.go), so ggml's fs::current_path() fallback would find the
# objects in normal operation whatever the loader's placement. That fallback is
# worth little here. It holds only for the launcher that sets that cwd, it is
# gone the moment anyone runs the binary by hand or through a wrapper that
# chdirs, and resolve_model_path has no equivalent, which would leave the only
# zero-download path in this backend resting on it. Rooting the loader is the
# one layout where all three mechanisms agree without depending on the cwd.
if [ -f "$CURDIR/ld.so" ]; then
exec "$CURDIR/ld.so" "$CURDIR/grpc-server" "$@"
fi
exec "$CURDIR/grpc-server" "$@"

View File

@@ -0,0 +1,129 @@
#include "stem_selection.h"
#include <cstddef>
#include <filesystem>
#include <string>
namespace audiocpp_backend {
namespace {
// The stem every separation family this backend can reach names its lead vocal
// track, and the one a caller who names no stem almost always wants: it is what
// the OpenAI-shaped "isolate the voice" request means. htdemucs (drums, bass,
// other, vocals) and mel_band_roformer (vocals, instrumental) both have it, and
// in htdemucs's case it is NOT the first output, which is the whole reason this
// preference is written down rather than left as "take index 0".
const char *const kPreferredStem = "vocals";
std::string join_names(const std::vector<std::string> &names) {
std::string out;
for (const auto &name : names) {
if (!out.empty()) {
out += ", ";
}
out += name;
}
return out;
}
// Whether a model-supplied stem name can be used as one component of a file
// name. Deliberately a whitelist of refusals rather than a sanitiser: silently
// rewriting "vo/cals" to "cals" would make the file the caller receives
// disagree with the name they would have to ask for.
bool name_is_writable(const std::string &name) {
if (name.empty() || name == "." || name == "..") {
return false;
}
// Control bytes, NUL above all. GGUF strings are length prefixed and demucs
// reads its source names out of JSON, which can encode one, so a
// std::string holding an embedded NUL survives all the way here. Two such
// names differing only AFTER the NUL are distinct std::strings, so the
// duplicate check below waves them through, and then path::c_str()
// truncates both at the NUL and they open the same file: precisely the
// silent overwrite the duplicate check exists to prevent, with the ".wav"
// stripped off as well. The rest of the range goes with it, since a newline
// or an escape sequence in a file name is a terminal and log injection
// nuisance with no legitimate use.
for (const char byte : name) {
const auto value = static_cast<unsigned char>(byte);
if (value < 0x20 || value == 0x7f) {
return false;
}
}
// Both separators, not just the host's. A GGUF is a downloaded file and its
// strings are not this host's to trust, so a name written on Windows must
// not become a directory traversal wherever the check happens to run.
return name.find('/') == std::string::npos &&
name.find('\\') == std::string::npos;
}
} // namespace
StemChoice select_named_output(const std::vector<std::string> &names,
const std::string &requested) {
StemChoice choice;
if (names.empty()) {
// Not an error here. The caller distinguishes "this family produces one
// unnamed output" from "this family produced nothing", and only it can
// tell them apart.
return choice;
}
// Every name is checked, not merely the selected one, because every stem is
// written. A bad name in the fourth output would otherwise be discovered
// only after three files had already been created.
for (std::size_t i = 0; i < names.size(); ++i) {
if (!name_is_writable(names[i])) {
choice.error = "audio-cpp: this model names an output stem '" +
names[i] +
"' that cannot be used as a file name; stems: " +
join_names(names);
return choice;
}
for (std::size_t seen = 0; seen < i; ++seen) {
if (names[seen] == names[i]) {
choice.error =
"audio-cpp: this model produces two output stems both named '" +
names[i] + "'; one would silently overwrite the other";
return choice;
}
}
}
if (!requested.empty()) {
for (std::size_t i = 0; i < names.size(); ++i) {
if (names[i] == requested) {
choice.index = static_cast<int>(i);
return choice;
}
}
choice.error = "audio-cpp: no stem named '" + requested +
"' in this model's output; available stems: " +
join_names(names);
return choice;
}
for (std::size_t i = 0; i < names.size(); ++i) {
if (names[i] == kPreferredStem) {
choice.index = static_cast<int>(i);
return choice;
}
}
choice.index = 0;
return choice;
}
std::string sibling_stem_path(const std::string &dst, const std::string &name) {
const std::filesystem::path path(dst);
// The dst extension is reused rather than forced to ".wav" so the siblings
// look like the file the caller named. write_audio_file writes WAV bytes
// whatever the extension says, for dst as much as for the siblings, so this
// keeps the set consistent instead of making the siblings honest about a
// format dst is already lying about.
const std::string extension =
path.has_extension() ? path.extension().string() : std::string(".wav");
return (path.parent_path() / (path.stem().string() + "." + name + extension))
.string();
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,63 @@
#pragma once
// Decides which of a separation model's named stems the AudioTransform response
// carries in `dst`, and names the sibling files every other stem is written to.
//
// It exists because AudioTransformResult carries ONE dst while htdemucs and
// mel_band_roformer produce several named outputs from a single run. Running
// once per stem would cost four full inferences for a four stem model, so the
// handler runs once, writes every stem beside dst, and puts the selected one in
// dst itself.
//
// Standard library only, so backend/cpp/run-unit-tests.sh compiles and runs its
// test without an audio.cpp checkout. grpc-server.cpp flattens the engine's
// NamedAudioBuffer list into the plain name vector taken here; nothing in this
// unit knows about engine::runtime.
#include <string>
#include <vector>
namespace audiocpp_backend {
struct StemChoice {
// Index into the `names` vector. -1 means nothing was chosen, which happens
// for an empty list (the family produced a single unnamed output) and for
// every refusal.
//
// THE CONTRACT THE CALLER INDEXES ON: when `error` is empty and `names` was
// not, this is always a valid index into `names`. It is never -1 in that
// case, so a caller that checks `error` first can index without a further
// guard, and a caller that does not check `error` first would index with a
// negative value. Check the error.
int index = -1;
// Non-empty when the request must be refused, and suitable verbatim as an
// INVALID_ARGUMENT message. Two things land here: the caller named a stem
// this model does not produce, and the model named stems that cannot both
// be written (an unusable file name, or two stems sharing one).
std::string error;
};
// Picks the stem that goes to dst. Preference order: an explicit `requested`,
// then "vocals", then the first output.
//
// An explicit but unknown `requested` is an ERROR rather than a fallback. A
// caller who asks for "drums" and silently receives "vocals" gets a 200 and a
// wrong file, which is the failure mode nobody can see; the message therefore
// lists the stem names this model really has.
//
// The names are also validated, because they come from the MODEL (htdemucs
// reads them from the GGUF's config.sources) and each one becomes a component
// of a file path this backend writes. A name carrying a path separator would
// write outside the caller's output directory, and two stems sharing a name
// would silently overwrite each other. Both are refused before anything is
// written, which is also why selection has to happen before the first write
// rather than after the loop: a refused request must leave no files behind.
StemChoice select_named_output(const std::vector<std::string> &names,
const std::string &requested);
// "/generated/transform-1.wav" + "drums" -> "/generated/transform-1.drums.wav".
// A dst with no extension gets ".wav", since that is what write_audio_file
// produces whatever the caller called the file.
std::string sibling_stem_path(const std::string &dst, const std::string &name);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,250 @@
// Unit tests for stem_selection. Standard library only. The harness
// (backend/cpp/run-unit-tests.sh) compiles this as a single translation unit,
// so the implementation is included directly.
//
// What is actually at stake here: AudioTransformResult carries one dst, a
// separation model produces several stems, and the caller cannot see which one
// they got. Every check below is about a wrong file arriving with a 200.
#include "stem_selection.cpp"
#include <cstdio>
#include <string>
#include <vector>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
static void check_equal(const std::string &got, const std::string &want,
const std::string &name) {
check(got == want, name + " (got \"" + got + "\", want \"" + want + "\")");
}
using namespace audiocpp_backend;
// htdemucs's real source order, taken from its GGUF config.sources. vocals is
// LAST, which is why "first output" is not the default.
static const std::vector<std::string> kDemucs = {"drums", "bass", "other",
"vocals"};
// mel_band_roformer's, where vocals is first.
static const std::vector<std::string> kRoformer = {"vocals", "instrumental"};
static void test_default_selection() {
const auto demucs = select_named_output(kDemucs, "");
check(demucs.error.empty(), "an unrequested selection is not an error");
check(demucs.index == 3, "no stem asked for picks vocals, not the first output");
const auto roformer = select_named_output(kRoformer, "");
check(roformer.index == 0, "vocals is picked when it is already first");
// No vocals anywhere: the first output is the documented fallback.
const auto novocals = select_named_output({"accompaniment", "drums"}, "");
check(novocals.index == 0 && novocals.error.empty(),
"a family with no vocals stem falls back to the first output");
// Substring matches must not count: "vocals_2" is a different stem.
const auto near = select_named_output({"vocals_2", "backing"}, "");
check(near.index == 0 && near.error.empty(),
"'vocals_2' is not 'vocals', so the fallback and not the preference applies");
const auto near_second = select_named_output({"backing", "vocals_2"}, "");
check(near_second.index == 0,
"a near miss on the preferred name does not pull it to the front");
}
static void test_explicit_selection() {
for (int i = 0; i < 4; ++i) {
const auto choice = select_named_output(kDemucs, kDemucs[static_cast<size_t>(i)]);
check(choice.index == i && choice.error.empty(),
"explicit '" + kDemucs[static_cast<size_t>(i)] + "' selects its own index");
}
// Including the one the default would have chosen anyway: asking for it
// must not be treated as "no request".
const auto vocals = select_named_output(kDemucs, "vocals");
check(vocals.index == 3, "explicitly asking for vocals still selects vocals");
}
static void test_unknown_stem_is_refused() {
const auto choice = select_named_output(kDemucs, "kazoo");
check(choice.index == -1, "an unknown stem selects nothing");
check(!choice.error.empty(), "an unknown stem is refused rather than substituted");
check(choice.error.find("kazoo") != std::string::npos,
"the refusal names the stem that was asked for");
// The real names, so the caller can fix the request without guessing.
for (const auto &name : kDemucs) {
check(choice.error.find(name) != std::string::npos,
"the refusal lists the real stem '" + name + "'");
}
check(choice.error.find("drums, bass, other, vocals") != std::string::npos,
"the refusal lists the stems in the model's own order");
// Case matters: the engine's ids are exact, so a wrong case is a wrong name
// rather than a near miss to be forgiven.
const auto wrong_case = select_named_output(kDemucs, "Vocals");
check(wrong_case.index == -1 && !wrong_case.error.empty(),
"stem names are matched case sensitively");
}
static void test_no_named_outputs() {
const auto choice = select_named_output({}, "");
check(choice.index == -1, "an empty output list selects nothing");
check(choice.error.empty(),
"an empty output list is not an error here: the caller decides");
const auto requested = select_named_output({}, "vocals");
check(requested.index == -1 && requested.error.empty(),
"an empty output list stays the caller's decision even when a stem was asked for");
}
static void test_unwritable_names_are_refused() {
// Model-supplied names become file path components. A separator would write
// outside the caller's output directory.
const std::vector<std::string> traversal = {"vocals", "../../etc/passwd"};
const auto escaped = select_named_output(traversal, "vocals");
check(escaped.index == -1 && !escaped.error.empty(),
"a stem name containing a path separator is refused");
check(escaped.error.find("../../etc/passwd") != std::string::npos,
"the refusal names the offending stem");
check(!select_named_output({"vo\\cals", "drums"}, "").error.empty(),
"a backslash separator is refused too");
check(!select_named_output({"drums", ""}, "").error.empty(),
"an empty stem name is refused");
check(!select_named_output({"drums", "."}, "").error.empty(),
"a stem named '.' is refused");
check(!select_named_output({"drums", ".."}, "").error.empty(),
"a stem named '..' is refused");
// The check covers EVERY name, not only the selected one: all of them are
// written, so a bad fourth name must not be found after three files exist.
const auto late = select_named_output({"vocals", "drums", "bass", "a/b"}, "vocals");
check(late.index == -1 && !late.error.empty(),
"an unwritable name after the selected one still refuses the whole request");
// Control bytes, and the NUL case is why the whole range is refused. These
// two names are DIFFERENT std::strings, so the duplicate check does not
// fire, yet both truncate to "vocals" at path::c_str() and would open one
// file: the silent overwrite the duplicate check exists to prevent, with
// the ".wav" stripped off into the bargain.
const std::string nul_a("vocals\0drums", 12);
const std::string nul_b("vocals\0bass", 11);
check(nul_a != nul_b, "the two NUL names really are distinct std::strings");
check(std::string(nul_a.c_str()) == "vocals" &&
std::string(nul_b.c_str()) == "vocals",
"and both truncate to the same C string, which is the hazard");
const auto nul_pair = select_named_output({nul_a, nul_b}, "");
check(nul_pair.index == -1 && !nul_pair.error.empty(),
"two stem names differing only after an embedded NUL are refused");
check(!select_named_output({"drums", std::string("vo\0cals", 7)}, "").error.empty(),
"a single embedded NUL is refused on its own");
check(!select_named_output({"drums", "voc\nals"}, "").error.empty(),
"a newline in a stem name is refused");
check(!select_named_output({"drums", "voc\tals"}, "").error.empty(),
"a tab in a stem name is refused");
check(!select_named_output({"drums", "voc\033[31mals"}, "").error.empty(),
"an escape sequence in a stem name is refused");
check(!select_named_output({"drums", "voc\177als"}, "").error.empty(),
"DEL in a stem name is refused");
// The boundary below the refused range is the space, which is an ordinary
// file name character and must stay usable, or this check would be
// refusing real stem names.
const auto spaced = select_named_output({"lead vocals", "drums"}, "lead vocals");
check(spaced.index == 0 && spaced.error.empty(),
"a space is not a control character and stays usable");
// And every byte above DEL: a UTF-8 stem name is ordinary, and signed char
// would make those bytes compare as negative.
const auto utf8 = select_named_output({"vocals", "b\xc3\xa4sse"}, "b\xc3\xa4sse");
check(utf8.index == 1 && utf8.error.empty(),
"a UTF-8 stem name is not mistaken for a control character");
// A leading dot is not a traversal and must stay usable.
const auto dotted = select_named_output({".vocals", "drums"}, ".vocals");
check(dotted.index == 0 && dotted.error.empty(),
"a leading dot in a stem name is allowed");
}
static void test_duplicate_names_are_refused() {
const auto choice = select_named_output({"vocals", "drums", "vocals"}, "vocals");
check(choice.index == -1 && !choice.error.empty(),
"two stems sharing a name are refused: one file would overwrite the other");
check(choice.error.find("vocals") != std::string::npos,
"the duplicate refusal names the repeated stem");
}
static void test_sibling_paths() {
check_equal(sibling_stem_path("/generated/transform-1.wav", "drums"),
"/generated/transform-1.drums.wav", "sibling beside an absolute dst");
check_equal(sibling_stem_path("sep.wav", "vocals"), "sep.vocals.wav",
"sibling of a bare file name has no directory");
check_equal(sibling_stem_path("/out/sep", "vocals"), "/out/sep.vocals.wav",
"an extensionless dst gets .wav");
check_equal(sibling_stem_path("/out/take.2.wav", "bass"), "/out/take.2.bass.wav",
"only the final extension is treated as the extension");
check_equal(sibling_stem_path("/out/sep.WAV", "bass"), "/out/sep.bass.WAV",
"the caller's extension spelling is preserved");
check_equal(sibling_stem_path("/a b/c d.wav", "other"), "/a b/c d.other.wav",
"spaces in the destination survive");
// The property that matters: no stem can ever be written over dst itself,
// or the "dst holds the selected stem" contract would depend on write order.
const std::string dst = "/out/sep.wav";
for (const auto &name : kDemucs) {
check(sibling_stem_path(dst, name) != dst,
"the sibling for '" + name + "' is not dst itself");
}
// Distinct stems must land in distinct files.
check(sibling_stem_path(dst, "drums") != sibling_stem_path(dst, "bass"),
"two stems get two different sibling paths");
}
// The contract grpc-server.cpp indexes on: an accepted choice over a non-empty
// name list is always in range, so the handler needs no bounds guard of its own.
// A -1 reaching the subscript would become a colossal size_t.
static void test_accepted_index_is_always_in_range() {
const std::vector<std::vector<std::string>> lists = {
kDemucs, kRoformer, {"solo"}, {"accompaniment", "drums"}, {"a", "b", "c"}};
const std::vector<std::string> requests = {"", "vocals", "drums", "solo", "c",
"kazoo", "..", "a/b"};
for (const auto &names : lists) {
for (const auto &requested : requests) {
const auto choice = select_named_output(names, requested);
if (!choice.error.empty()) {
check(choice.index == -1,
"a refusal never carries an index (request '" + requested + "')");
continue;
}
check(choice.index >= 0 &&
choice.index < static_cast<int>(names.size()),
"an accepted choice is in range (request '" + requested + "')");
// And the selected name is the one that was asked for, when one was.
if (!requested.empty()) {
check(names[static_cast<size_t>(choice.index)] == requested,
"an accepted explicit request selects that exact name");
}
}
}
}
int main() {
test_accepted_index_is_always_in_range();
test_default_selection();
test_explicit_selection();
test_unknown_stem_is_refused();
test_no_named_outputs();
test_unwritable_names_are_refused();
test_duplicate_names_are_refused();
test_sibling_paths();
if (failures != 0) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all stem_selection checks passed\n");
return 0;
}

View File

@@ -0,0 +1,204 @@
#include "stream_delta.h"
namespace audiocpp_backend {
namespace {
// True when `text` begins with `prefix`. An empty prefix matches everything,
// which is what makes the first fragment take the cumulative branch and the
// incremental branch alike: they agree there.
bool starts_with(const std::string &text, const std::string &prefix) {
return text.size() >= prefix.size() &&
text.compare(0, prefix.size(), prefix) == 0;
}
// Number of leading bytes of `text` that cannot BEGIN a character, i.e. orphan
// continuation bytes with no lead byte in front of them.
//
// They are unrecoverable rather than early: the byte that would have led them
// has already gone past, and nothing can be prepended to a fragment after the
// fact. Holding them would stall the stream for good, and emitting them puts
// invalid UTF-8 on the wire, so the caller DROPS them. Losing a byte keeps the
// stream alive; emitting one ends it, and takes the final_result still to come
// with it.
std::size_t utf8_orphan_prefix_length(const std::string &text) {
std::size_t index = 0;
while (index < text.size() &&
(static_cast<unsigned char>(text[index]) & 0xC0) == 0x80) {
++index;
}
return index;
}
// Length of the longest prefix of `text` that does NOT end inside a multi-byte
// UTF-8 sequence, i.e. the most that can go on the wire without splitting a
// character in half.
//
// Only the TRAILING sequence is examined; the leading side is
// utf8_orphan_prefix_length's job. Bytes in the MIDDLE are neither one's
// business: a family that emitted a malformed sequence inside its own text
// cannot be repaired here without deleting part of that transcript, and the
// stream is lost anyway, because final_result.text carries the same bytes
// through the same proto3 string field.
//
// Anything that can never complete is reported as complete, so it goes out
// rather than being held forever: a lead byte the encoding does not define, and
// a run of five or more continuation bytes, are both passed through. A tracker
// that stalled on undecodable input would turn one bad byte into a permanently
// silent stream, which is worse than the bad byte.
std::size_t utf8_complete_prefix_length(const std::string &text) {
std::size_t index = text.size();
std::size_t continuations = 0;
while (index > 0 && continuations < 4) {
const auto byte = static_cast<unsigned char>(text[index - 1]);
if ((byte & 0xC0) == 0x80) {
--index;
++continuations;
continue;
}
std::size_t needed = 1;
if ((byte & 0x80) == 0x00) {
needed = 1;
} else if ((byte & 0xE0) == 0xC0) {
needed = 2;
} else if ((byte & 0xF0) == 0xE0) {
needed = 3;
} else if ((byte & 0xF8) == 0xF0) {
needed = 4;
} else {
// Not a lead byte this encoding defines, so nothing is waiting on
// it and it must not be held.
needed = 1;
}
if (continuations + 1 >= needed) {
return text.size();
}
// The trailing sequence is short by at least one byte: cut before its
// lead byte and keep the rest for the next fragment.
return index - 1;
}
return text.size();
}
} // namespace
std::string TranscriptDeltaTracker::release(const std::string &fragment) {
// Held-back bytes go in FRONT of whatever arrived next, or the character
// they begin is reassembled in the wrong order.
std::string candidate = pending_ + fragment;
// A fragment must not BEGIN mid-character either. pending_ always starts on
// a lead byte, so this only bites when nothing was held and the caller's
// rules dropped the lead byte somewhere upstream; it is the backstop that
// makes "no delta this class returns is ever invalid UTF-8" true of the
// FRONT as well as the back, independently of those rules being right.
candidate.erase(0, utf8_orphan_prefix_length(candidate));
const std::size_t cut = utf8_complete_prefix_length(candidate);
pending_ = candidate.substr(cut);
std::string emitted = candidate.substr(0, cut);
assembled_ += emitted;
return emitted;
}
std::string TranscriptDeltaTracker::observe(const std::string &partial_text) {
if (partial_text.empty()) {
return {};
}
// The comparisons run against everything KNOWN, delivered plus held back,
// rather than against the delivered text alone. Comparing against the
// delivered text would treat the held-back byte as new on the very next
// report and emit it twice.
const std::string known = assembled_ + pending_;
// An EXACT repeat, and nothing looser. This absorbs the duplicate delivery
// voxtral_realtime produces, which is the ONLY thing rule 2 was ever needed
// for: process_available_stream_chunks hands each event it produces to the
// sink from inside its loop and RETURNS the last of the batch
// (session.cpp:385-386), so that last event arrives twice with byte-equal
// text both times. A duplicate IS an exact repeat, so equality covers it.
//
// It used to discard any report the known text merely STARTED WITH, and that
// cost far more than it bought. Two separate defects came out of it, and
// both were found by randomized traces rather than by reading:
//
// 1. A short INCREMENTAL fragment that happens to be a byte prefix of the
// transcript so far was read as a repeat and dropped, losing text with
// a 200 and no diagnostic. Both incremental families emit fragments
// that small routinely: nemotron_asr cuts at a byte offset
// (decoder.cpp:550) and vibevoice_asr at a common prefix
// (session.cpp:89-105). 9.50% of randomized pure-ASCII traces and
// 29.12% of French ones ended with a corrupted transcript.
// 2. When such a fragment was the LEAD BYTE of a multi-byte character, its
// continuation bytes then arrived alone and began the next delta, which
// is invalid UTF-8, which the Go runtime refuses to unmarshal, which
// ends the stream and the final_result with it.
//
// What the narrowing gives up is the shrinking-hypothesis case: a cumulative
// report SHORTER than what is known is now read as an incremental fragment
// and duplicates those bytes at the client. No pinned family produces one.
// voxtral is the only cumulative reporter, and its hypothesis is
// tokenizer_.decode(streaming_token_ids_) over a vector that is only ever
// push_back'ed (session.cpp:436) and cleared by reset() (session.cpp:257),
// so within a stream it can only grow. Measured: narrowing this changed not
// one byte of 30,000 randomized cumulative traces.
//
// KEPT DELIBERATELY THOUGH NO TEST CAN SEE IT. Once narrowed to equality
// this rule became redundant with rule 3 below: an equal partial has an
// empty suffix, so rule 3 would call release("") and emit nothing either
// way. Deleting it is therefore an equivalent mutation, and the mutation
// harness reports it as a survivor, which is the honest result and not a
// gap in the tests. It stays for two reasons: it states the duplicate
// absorption where the citation for it lives, and it is independent of rule
// 3's condition. A future tightening of rule 3 to, say, require a STRICTLY
// longer partial would otherwise send every duplicate down the incremental
// branch and put the whole transcript on the wire a second time.
if (partial_text == known) {
return {};
}
if (starts_with(partial_text, known)) {
// Cumulative: the report is the whole transcript so far.
return release(partial_text.substr(known.size()));
}
// Incremental: the fragment is new text to append.
//
// A family that REWRITES its hypothesis lands here too, and the client's
// view is then wrong in a way nothing downstream can fix. nemotron's
// decoder has such a branch (decoder.cpp:552-554): when the new text is not
// an extension of what it already emitted, it emits the whole new text. So
// "the cat sat" followed by "the cat sap" leaves the client holding
// "the cat satthe cat sap", and reconcile then correctly refuses to append
// to a contradicted assembly, which leaves concat(deltas) != final with no
// signal on the wire. This is NOT repaired here, and the reason is that a
// delta stream has no retraction: emitting only the differing suffix would
// read as "sap" appended to "the cat sat", which is a different wrong
// answer, and emitting a correction would need a wire field that does not
// exist. final_result carries the authoritative text either way. It did not
// fire in a 331 delta run, because an RNN-T decode is monotonic in practice.
return release(partial_text);
}
std::string TranscriptDeltaTracker::reconcile(const std::string &final_text) {
if (final_text.empty() || final_text == assembled_) {
// Nothing further is owed. Held-back bytes are dropped rather than
// flushed: they are not in the authoritative text, so sending them
// would contradict it.
pending_.clear();
return {};
}
if (!starts_with(final_text, assembled_)) {
// Contradicted. Nothing sent can be taken back, so nothing more is
// sent; final_result carries the authoritative text.
pending_.clear();
return {};
}
// Compared against the DELIVERED text, so the fragment below already
// contains whatever was held back. pending_ is therefore cleared rather
// than prepended, or those bytes would go out twice.
//
// The fragment ends on a character boundary whenever final_text is
// well-formed, which is the normal case and the reason a held-back sequence
// is always flushed here. It is still cut, so a family handing back a final
// text that is itself truncated mid-character cannot put a partial sequence
// on the wire through this path either.
pending_.clear();
return release(final_text.substr(assembled_.size()));
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,133 @@
#pragma once
// Turns whatever a streaming session calls a "partial transcript" into the
// incremental deltas AudioTranscriptionStream is contracted to send. Standard
// library only, so it is tested without an audio.cpp checkout.
//
// THIS UNIT EXISTS BECAUSE THE FAMILIES DISAGREE, and the disagreement is
// invisible at the interface: StreamEvent::partial_text is a Transcript either
// way. Read out of the pinned upstream, one family at a time:
//
// nemotron_asr INCREMENTAL. decoder.cpp emits
// current_text.substr(emitted_text.size()) per non-blank
// token, and only through the stream event SINK, during
// finalize(). process_audio_chunk returns empty events.
// vibevoice_asr INCREMENTAL. process_audio_chunk returns
// text.substr(common_prefix_size(...)); the sink is
// deliberately swapped out around its internal run_single,
// so the fragment arrives once, on the return value.
// higgs_audio_stt INCREMENTAL. Same shape as vibevoice_asr.
// voxtral_realtime CUMULATIVE. partial_text is
// tokenizer_.decode(streaming_token_ids_), the whole
// hypothesis so far. process_available_stream_chunks hands
// every event it produces to the sink from INSIDE its loop
// (session.cpp:385-386) and RETURNS only the last of the
// batch, so the last event of each batch arrives twice and
// the others arrive once.
//
// Applying either convention to the other family corrupts the transcript: read
// a cumulative report as a delta and the client sees the transcript repeated on
// every event; read an incremental fragment as cumulative and the suffix
// arithmetic eats the front of it. So the tracker decides per fragment, from
// what it has already delivered, and the one rule it enforces is that TEXT THE
// CLIENT HAS ALREADY BEEN SENT IS NEVER SENT AGAIN.
//
// The cumulative reading is provably safe for voxtral, which is the family it
// matters for: its decode is a pure concatenation of per-token byte strings
// (tokenizer_text.cpp:171-183), so decode(ids[0..n]) is an unconditional BYTE
// PREFIX of decode(ids[0..n+1]) and one of its reports can never be mistaken
// for an incremental fragment.
//
// UTF-8 IS THE OTHER HALF OF THAT SAME FACT. Because that decode concatenates
// raw token BYTES, a multi-byte character is split across token boundaries, and
// the difference between two consecutive cumulative reports is then a lone
// continuation byte. TranscriptStreamResponse.delta is a proto3 `string`, whose
// wire format REQUIRES valid UTF-8: the C++ runtime serializes an invalid one
// with at most a warning, but the Go runtime refuses to unmarshal it, and the
// client loses every remaining delta AND the final_result. So no fragment this
// class returns ever BEGINS OR ENDS inside a character: an incomplete trailing
// sequence is held back and merged into the next fragment, and a leading orphan
// continuation byte, which nothing can ever complete, is dropped.
//
// It is NOT a voxtral-only concern, which is what the first attempt at this
// assumed. The incremental families split characters by the same arithmetic:
// nemotron_asr's decoder cuts at a BYTE offset (decoder.cpp:550) and
// vibevoice_asr's common_prefix_size compares BYTES (session.cpp:80-86). Nor is
// European text the worst case: a Japanese transcript, whose every character is
// three bytes, carried at least one invalid delta in 33.74% of traces until
// rule 2 below learned to leave an incomplete fragment alone.
#include <string>
namespace audiocpp_backend {
class TranscriptDeltaTracker {
public:
// Takes one StreamEvent::partial_text and returns the fragment to put on
// the wire, empty when there is nothing new.
//
// The rules, in order, all of them against everything KNOWN (delivered
// plus held back), never against the delivered text alone:
// 1. An empty partial says nothing.
// 2. A partial IDENTICAL to the known text is a repeat: nothing is
// emitted. Identical, not merely a prefix of it. That absorbs voxtral's
// repeat of the last event in each batch, which is the only duplicate
// any pinned family produces and which carries byte-equal text both
// times. Discarding a mere PREFIX used to swallow an incremental
// fragment that coincided with the start of the transcript, corrupting
// the text silently and, when that fragment was a character's lead
// byte, killing the stream outright; see the note at the rule in the
// implementation.
// 3. A partial that EXTENDS the known text is a cumulative report: only
// its new suffix is emitted.
// 4. Anything else is an incremental fragment: it is emitted whole and
// appended.
//
// Rule 3 is the one judgement call, since a fragment that happens to begin
// with the entire transcript so far is indistinguishable from a cumulative
// report. It is read as cumulative because every cumulative family produces
// that shape on EVERY event, while an incremental family produces it only
// when one fragment repeats everything before it, which no tokenizer output
// does in practice.
//
// What comes back is the fragment MINUS any incomplete trailing UTF-8
// sequence, which is carried into the next call, and minus any leading
// orphan continuation byte, which is dropped. So an empty return can also
// mean "the only new bytes were half a character", and the caller needs no
// knowledge of that: writing nothing is exactly right.
std::string observe(const std::string &partial_text);
// Reconciles against TaskResult::text_output, which is authoritative, and
// returns the fragment that makes appending every delta equal it.
//
// This is what makes the OFFLINE FALLBACK a single line rather than its own
// branch: with no partials observed, the assembly is empty and the whole
// final text comes back as one delta.
//
// It is also what FLUSHES a held-back UTF-8 sequence, and it can always do
// so: the final text is complete, so the fragment from the last delivered
// byte to its end ends on a character boundary.
//
// A final text that CONTRADICTS what was already sent returns empty. A
// fragment on the wire cannot be retracted, so the alternative would be to
// send the transcript a second time and let the client hold it twice.
// final_result carries the authoritative text either way.
std::string reconcile(const std::string &final_text);
// Everything the client has been sent, concatenated. Held-back bytes are
// deliberately NOT included: this is what the client holds, not what the
// tracker knows.
const std::string &assembled() const noexcept { return assembled_; }
private:
// Appends the emittable prefix of `fragment` to assembled_ and returns it,
// keeping any incomplete trailing UTF-8 sequence in pending_.
std::string release(const std::string &fragment);
std::string assembled_;
// An incomplete trailing UTF-8 sequence, computed but not sent. Always a
// proper prefix of one character, so at most three bytes.
std::string pending_;
};
} // namespace audiocpp_backend

View File

@@ -0,0 +1,486 @@
// Unit tests for stream_delta. Standard library only. The harness compiles this
// as a single translation unit, so the implementation is included directly.
//
// The traces below are transcribed from the pinned upstream sessions rather
// than invented, because the whole reason this unit exists is that the four
// streaming ASR families do NOT agree on what partial_text means.
#include "stream_delta.cpp"
#include <cstdio>
#include <string>
#include <vector>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
static void check_eq(const std::string &got, const std::string &want,
const std::string &name) {
check(got == want, name + " (got \"" + got + "\" want \"" + want + "\")");
}
using audiocpp_backend::TranscriptDeltaTracker;
// Feeds a whole trace and returns what a client appending every emitted
// fragment would end up holding, which is the only property that matters.
static std::string client_view(TranscriptDeltaTracker &tracker,
const std::vector<std::string> &partials,
std::vector<std::string> *emitted = nullptr) {
std::string view;
for (const auto &partial : partials) {
const std::string fragment = tracker.observe(partial);
if (emitted != nullptr && !fragment.empty()) {
emitted->push_back(fragment);
}
view += fragment;
}
return view;
}
// nemotron_asr: decoder.cpp emits current_text.substr(emitted_text.size()) on
// every non-blank token, i.e. INCREMENTAL fragments, through the stream event
// sink during finalize().
static void test_incremental_family() {
TranscriptDeltaTracker tracker;
std::vector<std::string> emitted;
const std::string view =
client_view(tracker, {"Local", " AI", " now", " speaks."}, &emitted);
check_eq(view, "Local AI now speaks.", "incremental deltas concatenate");
check(emitted.size() == 4, "incremental family emits one fragment per partial");
check_eq(tracker.assembled(), "Local AI now speaks.",
"incremental family assembles the whole transcript");
check_eq(tracker.reconcile("Local AI now speaks."), "",
"a final result the deltas already cover adds nothing");
}
// voxtral_realtime: process_one_stream_chunk sets partial_text to
// tokenizer_.decode(streaming_token_ids_), the WHOLE accumulated hypothesis,
// and process_available_stream_chunks then hands the SAME event to both the
// sink and the caller, so every partial arrives twice.
static void test_cumulative_family_with_duplicate_delivery() {
TranscriptDeltaTracker tracker;
std::vector<std::string> emitted;
const std::string view = client_view(
tracker, {"Local", "Local", "Local AI", "Local AI", "Local AI now",
"Local AI now"},
&emitted);
check_eq(view, "Local AI now", "cumulative partials are not repeated to the client");
check(emitted.size() == 3,
"the duplicate delivery of each cumulative event emits nothing twice");
check_eq(emitted.empty() ? "" : emitted[0], "Local", "first cumulative fragment");
check_eq(emitted.size() < 2 ? "" : emitted[1], " AI", "second cumulative fragment");
check_eq(emitted.size() < 3 ? "" : emitted[2], " now", "third cumulative fragment");
}
// The rule that separates the two: a report IDENTICAL to everything known is a
// repeat and is never sent again.
//
// Only an identical one. Rule 2 used to discard any report the known text merely
// STARTED WITH, and that cost more than it bought: see
// test_a_short_fragment_is_not_mistaken_for_a_repeat.
static void test_an_exact_repeat_is_never_resent() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe("hello world"), "hello world", "first fragment");
check_eq(tracker.observe("hello world"), "", "an identical repeat emits nothing");
check_eq(tracker.observe("hello world"), "", "and a third delivery emits nothing");
check_eq(tracker.assembled(), "hello world", "the assembly is unchanged by repeats");
}
// THE TRADE, pinned so it is a decision rather than a surprise. A CUMULATIVE
// report that SHRINKS is no longer absorbed: it is read as an incremental
// fragment and duplicates a few bytes at the client.
//
// No pinned family produces one. voxtral_realtime is the only cumulative
// reporter, and its hypothesis is tokenizer_.decode(streaming_token_ids_) over a
// vector that is only ever push_back'ed (session.cpp:436) and cleared by reset()
// (session.cpp:257), so within a stream it grows and never shrinks. The
// duplicate delivery rule 2 really exists for is an EXACT repeat, which the test
// above still covers.
static void test_a_shrinking_hypothesis_is_read_as_incremental() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe("hello world"), "hello world", "first fragment");
check_eq(tracker.observe("hello"), "hello",
"a shortened hypothesis is now read as an incremental fragment");
check_eq(tracker.assembled(), "hello worldhello",
"which duplicates those bytes at the client: the accepted cost");
}
static void test_empty_partials_are_ignored() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe(""), "", "an empty partial emits nothing");
check_eq(tracker.observe("a"), "a", "a real partial after an empty one still emits");
check_eq(tracker.observe(""), "", "a later empty partial emits nothing");
check_eq(tracker.assembled(), "a", "empty partials do not disturb the assembly");
}
// The offline fallback: a family with no streaming ASR runs once, so nothing is
// ever observed and the reconciliation IS the single delta the RPC promises.
static void test_offline_fallback_is_one_delta() {
TranscriptDeltaTracker tracker;
check_eq(tracker.reconcile("the whole transcript"), "the whole transcript",
"with no partials the final text is emitted whole");
check_eq(tracker.assembled(), "the whole transcript",
"the reconciliation is recorded as delivered");
check_eq(tracker.reconcile("the whole transcript"), "",
"reconciling twice does not duplicate");
}
// A streaming family whose partials stopped short of the final text: the tail
// is emitted so that appending every delta still equals final_result.text.
static void test_reconcile_emits_the_tail() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe("Local AI"), "Local AI", "partial arrives");
check_eq(tracker.reconcile("Local AI now speaks."), " now speaks.",
"the untold tail of the final text is emitted");
check_eq(tracker.assembled(), "Local AI now speaks.", "tail is recorded");
}
// Divergence. nemotron's decoder has a rewrite branch: when the new hypothesis
// is NOT an extension of what it already emitted, it emits the whole new text.
// Nothing can retract a fragment already written to the wire, so the tracker
// must not try: it emits nothing further and leaves final_result authoritative.
static void test_divergent_final_text_is_not_appended() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe("the cat"), "the cat", "first hypothesis");
check_eq(tracker.reconcile("the dog"), "",
"a final text that contradicts the deltas is not appended to them");
check_eq(tracker.assembled(), "the cat",
"a contradicted assembly is left as it was actually sent");
}
static void test_empty_final_text() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe("something"), "something", "partial arrives");
check_eq(tracker.reconcile(""), "", "an empty final text emits nothing");
check_eq(tracker.assembled(), "something", "an empty final text changes nothing");
}
// A whitespace-only fragment is real text: the space between two words is
// exactly what an incremental family delivers on its own.
static void test_whitespace_fragments_survive() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe("one"), "one", "word");
check_eq(tracker.observe(" "), " ", "a bare separator is emitted");
check_eq(tracker.observe("two"), "two", "next word");
check_eq(tracker.assembled(), "one two", "separator is kept in the assembly");
}
// The ambiguity this unit cannot resolve, pinned so that a future reader sees
// the choice rather than rediscovering it: a fragment that EXTENDS everything
// delivered so far is read as a cumulative report, because that is what every
// cumulative family produces on every event, while an incremental family
// producing one is the rare coincidence of a fragment repeating the whole
// transcript so far.
static void test_prefix_extension_is_read_as_cumulative() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe("I"), "I", "first fragment");
check_eq(tracker.observe("I'm"), "'m",
"a fragment extending the assembly is treated as a cumulative report");
check_eq(tracker.assembled(), "I'm", "cumulative reading assembles once");
}
// --------------------------------------------------------------------------
// UTF-8 boundaries
//
// TranscriptStreamResponse.delta is a proto3 `string`, and the wire format
// REQUIRES a string field to be valid UTF-8. The C++ runtime serializes an
// invalid one with at most a warning; the Go runtime refuses to unmarshal it,
// so the client loses every delta AND the final_result still to come.
//
// Not hypothetical. voxtral_realtime reports the whole hypothesis as
// tokenizer_.decode(streaming_token_ids_), and that decode is a pure
// concatenation of raw token BYTES (tokenizer_text.cpp:171-183), so a
// multi-byte character is split across token boundaries and the cumulative
// difference between two consecutive reports is a lone continuation byte.
// --------------------------------------------------------------------------
// True when `text` is well-formed UTF-8. Written out here rather than reused
// from the implementation on purpose: a test that shares the implementation's
// idea of a boundary cannot catch the implementation's idea being wrong.
static bool is_valid_utf8(const std::string &text) {
size_t i = 0;
while (i < text.size()) {
const auto lead = static_cast<unsigned char>(text[i]);
size_t length = 0;
if ((lead & 0x80) == 0x00) {
length = 1;
} else if ((lead & 0xE0) == 0xC0) {
length = 2;
} else if ((lead & 0xF0) == 0xE0) {
length = 3;
} else if ((lead & 0xF8) == 0xF0) {
length = 4;
} else {
return false;
}
if (i + length > text.size()) {
return false;
}
for (size_t k = 1; k < length; ++k) {
if ((static_cast<unsigned char>(text[i + k]) & 0xC0) != 0x80) {
return false;
}
}
i += length;
}
return true;
}
// Written as byte escapes so the test does not depend on the encoding of this
// source file.
static const std::string kEAcute = "\xC3\xA9"; // 2 bytes
static const std::string kEuro = "\xE2\x82\xAC"; // 3 bytes
static const std::string kEmoji = "\xF0\x9F\x8E\xA7"; // 4 bytes
// A cumulative family advancing its hypothesis one BYTE at a time, which is
// what voxtral_realtime does across a multi-byte character.
static void test_cumulative_split_multibyte_character() {
TranscriptDeltaTracker tracker;
std::vector<std::string> emitted;
const std::string full = "5" + kEuro;
std::vector<std::string> partials;
for (size_t n = 1; n <= full.size(); ++n) {
partials.push_back(full.substr(0, n));
}
const std::string view = client_view(tracker, partials, &emitted);
check_eq(view, full, "a byte-at-a-time cumulative report still assembles");
for (size_t i = 0; i < emitted.size(); ++i) {
check(is_valid_utf8(emitted[i]),
"cumulative fragment " + std::to_string(i) + " is valid UTF-8");
}
check_eq(tracker.reconcile(full), "", "the final text adds nothing");
}
// An incremental family splitting a character across two fragments.
static void test_incremental_split_multibyte_character() {
TranscriptDeltaTracker tracker;
std::vector<std::string> emitted;
const std::string view = client_view(
tracker, {"caf" + kEAcute.substr(0, 1), kEAcute.substr(1), " au lait"},
&emitted);
check_eq(view, "caf" + kEAcute + " au lait",
"an incremental split character still assembles");
for (size_t i = 0; i < emitted.size(); ++i) {
check(is_valid_utf8(emitted[i]),
"incremental fragment " + std::to_string(i) + " is valid UTF-8");
}
check(emitted.size() == 3, "one fragment out per partial, none swallowed");
if (emitted.size() == 3) {
check_eq(emitted[0], "caf", "the lead byte of the character is held back");
check_eq(emitted[1], kEAcute,
"the held byte is merged into the next fragment, not sent alone");
check_eq(emitted[2], " au lait", "the rest follows unchanged");
}
}
// A 4 byte character split three ways, so the held-back buffer has to survive
// more than one round.
static void test_four_byte_character_split_three_ways() {
TranscriptDeltaTracker tracker;
std::vector<std::string> emitted;
const std::string view =
client_view(tracker,
{"listen " + kEmoji.substr(0, 1), kEmoji.substr(1, 2),
kEmoji.substr(3), " now"},
&emitted);
check_eq(view, "listen " + kEmoji + " now",
"a 4 byte character survives three splits");
for (size_t i = 0; i < emitted.size(); ++i) {
check(is_valid_utf8(emitted[i]),
"4 byte fragment " + std::to_string(i) + " is valid UTF-8");
}
}
// The held-back bytes must reach the client. reconcile can always flush them,
// because the final text is complete by construction.
static void test_reconcile_flushes_a_held_back_sequence() {
TranscriptDeltaTracker tracker;
const std::string first = tracker.observe("done" + kEuro.substr(0, 2));
check_eq(first, "done", "the incomplete trailing sequence is held back");
check(is_valid_utf8(first), "what was emitted is valid UTF-8");
const std::string tail = tracker.reconcile("done" + kEuro);
check_eq(tail, kEuro, "reconcile flushes the completed character");
check(is_valid_utf8(tail), "the flushed tail is valid UTF-8");
check_eq(tracker.assembled(), "done" + kEuro, "the client holds the whole text");
}
// Held-back bytes are not lost track of: the next cumulative report emits the
// whole character rather than only the bytes that just arrived.
static void test_held_bytes_join_the_next_fragment() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe("a" + kEuro.substr(0, 1)), "a",
"only the complete prefix goes out");
check_eq(tracker.observe("a" + kEuro), kEuro,
"the next cumulative report emits the whole character at once");
check_eq(tracker.assembled(), "a" + kEuro, "assembly is correct");
}
// A whole multi-byte character arriving at once must NOT be held back: holding
// a complete sequence would stall every stream by one character.
static void test_a_complete_character_is_not_held() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe("x" + kEuro), "x" + kEuro,
"a fragment ending on a boundary is emitted immediately");
check_eq(tracker.observe("x" + kEuro + kEmoji), kEmoji,
"and so is the next one");
}
// Bytes that can never complete must not be held forever: a stray continuation
// byte or an invalid lead is passed through rather than stalling the stream.
// Repairing a family's malformed output is not something a delta tracker can do
// without altering the transcript.
static void test_undecodable_bytes_are_not_held_forever() {
TranscriptDeltaTracker tracker;
check_eq(tracker.observe(std::string("ok\x80")), std::string("ok\x80"),
"a stray continuation byte is passed through, not held");
check_eq(tracker.observe(std::string("ok\x80") + "next"), "next",
"the stream continues");
// A lead byte the encoding does not define (0xF8 and above). Nothing can
// ever complete it, so holding it would stall the stream for good.
TranscriptDeltaTracker invalid_lead;
check_eq(invalid_lead.observe(std::string("ok\xFE")), std::string("ok\xFE"),
"an undefined lead byte is passed through, not held");
check_eq(invalid_lead.observe(std::string("ok\xFE") + "more"), "more",
"the stream continues past an undefined lead byte");
// The same byte followed by continuation bytes, which is the shape that
// looks most like a real sequence waiting to be completed.
TranscriptDeltaTracker invalid_run;
check_eq(invalid_run.observe(std::string("\xFE\x80\x80")), std::string("\xFE\x80\x80"),
"an undefined lead with continuations is passed through");
// Five continuation bytes with no lead in sight. They are DROPPED, not
// held: nothing can ever precede them, so holding would stall the stream
// for good, and emitting them would put invalid UTF-8 on the wire. See
// test_a_fragment_never_begins_mid_character.
TranscriptDeltaTracker orphans;
check_eq(orphans.observe(std::string("\x80\x80\x80\x80\x80")), "",
"a run of orphan continuation bytes is dropped, not emitted");
check_eq(orphans.observe("after"), "after",
"and the stream continues past them");
}
// THE SECOND HALF OF THE SAME BUG, and the one that survived fix round 1.
//
// Rule 2 discards a fragment the known text already starts with. When that
// fragment is the LEAD BYTE of a NEW character it looks exactly like a repeat of
// an earlier character beginning with the same byte, so it was discarded and
// never held. Its continuation bytes then arrived on their own and began the
// next delta, which is invalid UTF-8 at the FRONT, and utf8_complete_prefix_length
// only ever inspected the TRAILING sequence.
//
// Reachable from shipping families, not synthetic: nemotron_asr's
// decoder.cpp:550 cuts at a BYTE offset (current_text.substr(emitted_text.size()))
// and vibevoice_asr's common_prefix_size (session.cpp:80-86) compares BYTES, so
// both split characters mid-sequence. The trace below is exactly how they split
// "ssee" spelled with the German sharp s, an e-acute, a euro sign and an o-grave,
// three of which begin with the same 0xC3 lead byte.
static void test_a_repeated_lead_byte_is_not_swallowed() {
TranscriptDeltaTracker tracker;
std::vector<std::string> emitted;
const std::string sharp_s = "\xC3\x9F"; // U+00DF
const std::string e_acute = "\xC3\xA9"; // U+00E9
const std::string euro = "\xE2\x82\xAC"; // U+20AC
const std::string o_grave = "\xC3\xB2"; // U+00F2
const std::string full = sharp_s + e_acute + euro + o_grave;
const std::string view = client_view(tracker,
{sharp_s.substr(0, 1), sharp_s.substr(1),
e_acute.substr(0, 1),
e_acute.substr(1) + euro,
o_grave.substr(0, 1), o_grave.substr(1)},
&emitted);
for (size_t i = 0; i < emitted.size(); ++i) {
check(is_valid_utf8(emitted[i]),
"repeated-lead fragment " + std::to_string(i) + " is valid UTF-8");
}
check_eq(view, full, "no character is lost to a repeated lead byte");
check_eq(tracker.reconcile(full), "", "the final text adds nothing");
}
// The same shape one layer down, as a backstop: a fragment that BEGINS with
// orphan continuation bytes must never go on the wire, whatever produced it.
// Dropping bytes keeps the stream alive; emitting them ends it, and takes the
// final_result that was still to come with it.
static void test_a_fragment_never_begins_mid_character() {
TranscriptDeltaTracker tracker;
const std::string first = tracker.observe(std::string("\xA9") + "rest");
check(is_valid_utf8(first), "a leading orphan continuation byte is not emitted");
check_eq(first, "rest", "the rest of the fragment still goes out");
TranscriptDeltaTracker all_orphans;
check_eq(all_orphans.observe(std::string("\x82\xAC")), "",
"a fragment that is nothing but orphans emits nothing");
check_eq(all_orphans.assembled(), "",
"and nothing is recorded as delivered");
check_eq(all_orphans.observe("after"), "after", "the stream continues");
}
// PURE ASCII, no multi-byte character anywhere, and the transcript still comes
// out wrong: a short incremental fragment that happens to be a byte prefix of
// everything known was read as an already-delivered repeat and discarded.
//
// This is the shape both incremental families produce. nemotron_asr emits
// current_text.substr(emitted_text.size()) per non-blank token (decoder.cpp:550)
// and vibevoice_asr emits text.substr(common_prefix_size(...)) (session.cpp:89-105),
// so a one-character fragment is ordinary output, and any of the transcript's
// own leading characters will eventually arrive as one.
//
// Measured over 5,000 randomized traces per transcript before this was fixed:
// 9.50% of pure-ASCII traces and 29.12% of French ones ended with the client
// holding something other than final_result.text, with a 200 and no diagnostic.
static void test_a_short_fragment_is_not_mistaken_for_a_repeat() {
TranscriptDeltaTracker tracker;
std::vector<std::string> emitted;
const std::string full = "pure ascii transcript";
const std::string view = client_view(
tracker, {"pure ", "ascii ", "trans", "c", "ri", "p", "t"}, &emitted);
check_eq(view, full, "the whole transcript reaches the client");
check_eq(tracker.assembled(), full, "and the assembly agrees with it");
check_eq(tracker.reconcile(full), "",
"the final text adds nothing, because nothing was lost");
check(emitted.size() == 7, "every fragment produced exactly one delta");
}
int main() {
test_incremental_family();
test_cumulative_family_with_duplicate_delivery();
test_an_exact_repeat_is_never_resent();
test_a_shrinking_hypothesis_is_read_as_incremental();
test_a_short_fragment_is_not_mistaken_for_a_repeat();
test_empty_partials_are_ignored();
test_offline_fallback_is_one_delta();
test_reconcile_emits_the_tail();
test_divergent_final_text_is_not_appended();
test_empty_final_text();
test_whitespace_fragments_survive();
test_prefix_extension_is_read_as_cumulative();
test_cumulative_split_multibyte_character();
test_incremental_split_multibyte_character();
test_four_byte_character_split_three_ways();
test_reconcile_flushes_a_held_back_sequence();
test_held_bytes_join_the_next_fragment();
test_a_complete_character_is_not_held();
test_undecodable_bytes_are_not_held_forever();
test_a_repeated_lead_byte_is_not_swallowed();
test_a_fragment_never_begins_mid_character();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all stream_delta checks passed\n");
return 0;
}

View File

@@ -0,0 +1,976 @@
// Tests for the streaming drivers in loaded_model: begin_stream,
// run_streaming_pull, run_streaming_audio and run_streaming_live, plus
// resolve_model_path, which lives in the same engine-linked unit.
//
// Engine-linked, so this runs through ctest rather than
// backend/cpp/run-unit-tests.sh. It builds no model and loads no file: a
// LoadedModel::Session is a plain struct holding a pointer to an engine
// interface, so a fake session exercises the drivers directly, which is the
// only way to assert the STATE OBLIGATION (prepare, then start_stream, on every
// stream) without a GPU and a gigabyte of weights.
#include "inference_lane.h"
#include "loaded_model.h"
#include "engine/framework/runtime/session.h"
#include <cstddef>
#include <cstdio>
#include <functional>
#include <memory>
#include <optional>
#include <stdexcept>
#include <string>
#include <vector>
namespace rt = engine::runtime;
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
static void check_eq(const std::string &got, const std::string &want,
const std::string &name) {
check(got == want, name + " (got \"" + got + "\" want \"" + want + "\")");
}
static std::string join(const std::vector<std::string> &parts) {
std::string out;
for (const auto &part : parts) {
if (!out.empty()) {
out += "|";
}
out += part;
}
return out;
}
// --------------------------------------------------------------------------
// Fakes
// --------------------------------------------------------------------------
// Consumes audio chunks, like every streaming ASR family.
//
// It deliberately does NOT override start_stream, so every test that drives it
// also pins the claim the drivers rely on: IStreamingVoiceTaskSession's BASE
// start_stream is a call to reset(). If upstream ever changes that base, the
// replay test below fails rather than the backend silently continuing the
// previous stream.
class FakeAudioSession : public rt::IStreamingVoiceTaskSession {
public:
std::vector<std::string> calls;
std::vector<rt::AudioChunk> chunks;
rt::StreamingPolicy policy;
// Set to make process_audio_chunk throw on the nth call (1-based).
int throw_on_chunk = 0;
// Cumulative partial text, which is voxtral_realtime's convention.
bool report_partials = true;
std::string family() const override { return "fake_audio"; }
rt::VoiceTaskKind task_kind() const override { return rt::VoiceTaskKind::Asr; }
rt::RunMode run_mode() const override { return rt::RunMode::Streaming; }
void prepare(const rt::SessionPreparationRequest &request) override {
calls.push_back("prepare");
prepared_ = true;
prepared_rate_ = request.audio.has_value() ? request.audio->sample_rate : 0;
}
rt::StreamingPolicy streaming_policy() const override { return policy; }
void set_stream_event_sink(rt::StreamEventCallback sink) override {
calls.push_back(sink ? "sink+" : "sink-");
sink_ = std::move(sink);
}
void reset() override {
if (!prepared_) {
// Exactly what silero_vad does, and the reason prepare() has to come
// first rather than being folded into session_for.
throw std::runtime_error("fake: prepare() must be called before reset()");
}
calls.push_back("reset");
seen_frames_ = 0;
seen_chunks_ = 0;
text_.clear();
}
rt::StreamEvent process_audio_chunk(const rt::AudioChunk &chunk) override {
calls.push_back("chunk");
chunks.push_back(chunk);
++seen_chunks_;
if (throw_on_chunk == seen_chunks_) {
throw std::runtime_error("fake: chunk failure");
}
const int channels = chunk.channels > 0 ? chunk.channels : 1;
seen_frames_ += static_cast<std::int64_t>(chunk.samples.size()) / channels;
text_ += "w" + std::to_string(seen_chunks_);
rt::StreamEvent event;
if (report_partials) {
event.partial_text = rt::Transcript{text_, "en"};
}
return event;
}
rt::TaskResult finalize() override {
calls.push_back("finalize");
rt::TaskResult result;
result.text_output = rt::Transcript{
text_ + "/frames=" + std::to_string(seen_frames_), "en"};
return result;
}
// Emits through the SINK the way nemotron_asr does, from inside the final
// step rather than from process_audio_chunk.
void emit_through_sink(const std::string &fragment) {
if (!sink_) {
return;
}
rt::StreamEvent event;
event.partial_text = rt::Transcript{fragment, "en"};
sink_(event);
}
bool sink_installed() const { return static_cast<bool>(sink_); }
int prepared_rate() const { return prepared_rate_; }
private:
rt::StreamEventCallback sink_;
bool prepared_ = false;
int prepared_rate_ = 0;
std::int64_t seen_frames_ = 0;
int seen_chunks_ = 0;
std::string text_;
};
// nemotron_asr's shape: partials arrive only through the sink, and only from
// inside the finalize step.
class SinkOnlyAudioSession : public FakeAudioSession {
public:
SinkOnlyAudioSession() { report_partials = false; }
rt::TaskResult finalize() override {
emit_through_sink("late ");
emit_through_sink("partial");
return FakeAudioSession::finalize();
}
};
// Pulls events, like every streaming TTS family. Overrides start_stream the way
// the seven real families do, calling reset() first.
class FakePullSession : public rt::IStreamingVoiceTaskSession {
public:
std::vector<std::string> calls;
std::size_t event_count = 3;
bool final_on_second = false;
std::string family() const override { return "fake_pull"; }
rt::VoiceTaskKind task_kind() const override { return rt::VoiceTaskKind::Tts; }
rt::RunMode run_mode() const override { return rt::RunMode::Streaming; }
void prepare(const rt::SessionPreparationRequest &) override {
calls.push_back("prepare");
prepared_ = true;
}
rt::StreamingPolicy streaming_policy() const override {
rt::StreamingPolicy policy;
policy.input = rt::StreamingInputKind::None;
policy.output = rt::StreamingOutputKind::PullEvents;
return policy;
}
void start_stream(const rt::TaskRequest &request) override {
calls.push_back("start_stream");
(void)request;
reset();
}
void set_stream_event_sink(rt::StreamEventCallback sink) override {
calls.push_back(sink ? "sink+" : "sink-");
sink_ = std::move(sink);
}
void reset() override {
if (!prepared_) {
throw std::runtime_error("fake: prepare() must be called before reset()");
}
calls.push_back("reset");
emitted_ = 0;
}
std::optional<rt::StreamEvent> next_stream_event() override {
if (emitted_ >= event_count) {
return std::nullopt;
}
rt::StreamEvent event;
rt::AudioBuffer audio;
audio.sample_rate = 24000;
audio.channels = 1;
audio.samples.assign(4, 0.25F);
// named_audio_outputs, NOT audio_output: this is where supertonic,
// omnivoice and voxcpm2 all put their streamed chunks.
event.named_audio_outputs.push_back(
{"chunk_" + std::to_string(emitted_), std::move(audio), {}});
++emitted_;
if (final_on_second && emitted_ == 2) {
event.is_final = true;
}
calls.push_back("pull");
return event;
}
rt::StreamEvent process_audio_chunk(const rt::AudioChunk &) override {
throw std::runtime_error("fake_pull consumes no audio");
}
rt::TaskResult finalize() override {
calls.push_back("finalize");
rt::TaskResult result;
rt::AudioBuffer merged;
merged.sample_rate = 24000;
merged.channels = 1;
merged.samples.assign(4 * emitted_, 0.25F);
result.audio_output = std::move(merged);
return result;
}
bool sink_installed() const { return static_cast<bool>(sink_); }
private:
rt::StreamEventCallback sink_;
bool prepared_ = false;
std::size_t emitted_ = 0;
};
// --------------------------------------------------------------------------
// Helpers
// --------------------------------------------------------------------------
static audiocpp_backend::LoadedModel::Session
streaming_session(rt::IStreamingVoiceTaskSession &fake,
audiocpp_backend::Task task) {
audiocpp_backend::LoadedModel::Session session;
session.task = task;
session.mode = audiocpp_backend::Mode::Streaming;
session.streaming = &fake;
return session;
}
static rt::TaskRequest audio_request(int sample_rate, int channels,
std::int64_t frames) {
rt::TaskRequest request;
rt::AudioBuffer audio;
audio.sample_rate = sample_rate;
audio.channels = channels;
audio.samples.assign(static_cast<std::size_t>(frames * channels), 0.5F);
request.audio_input = std::move(audio);
return request;
}
// --------------------------------------------------------------------------
// Tests
// --------------------------------------------------------------------------
static void test_begin_stream_prepares_then_starts() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakePullSession fake;
const auto session = streaming_session(fake, audiocpp_backend::Task::Tts);
rt::TaskRequest request;
audiocpp_backend::begin_stream(session, request, entry);
check_eq(join(fake.calls), "prepare|start_stream|reset",
"begin_stream prepares before it starts, and start_stream resets");
}
// The base implementation of start_stream IS a reset(). FakeAudioSession does
// not override start_stream, so this is that guarantee, read out of the pinned
// header rather than assumed.
static void test_base_start_stream_resets() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
audiocpp_backend::begin_stream(session, audio_request(16000, 1, 10), entry);
check_eq(join(fake.calls), "prepare|reset",
"the interface's own start_stream resets the session");
}
static void test_begin_stream_refuses_a_non_streaming_session() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
audiocpp_backend::LoadedModel::Session session;
session.mode = audiocpp_backend::Mode::Offline;
bool threw_capability = false;
try {
rt::TaskRequest request;
audiocpp_backend::begin_stream(session, request, entry);
} catch (const audiocpp_backend::CapabilityError &) {
threw_capability = true;
} catch (const std::exception &) {
}
check(threw_capability,
"begin_stream on an offline session throws CapabilityError, not a null deref");
}
// THE ONE THIS TASK IS ABOUT. A streaming session is cached, so the second
// stream gets the object the first one left behind. Two identical runs against
// the SAME session must produce identical output.
static void test_a_refetched_session_replays_identically() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 512;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
const auto request = audio_request(16000, 1, 1536);
std::vector<std::string> first_fragments;
const auto first = audiocpp_backend::run_streaming_audio(
session, request, *request.audio_input,
[&](const rt::StreamEvent &event) {
if (event.partial_text.has_value()) {
first_fragments.push_back(event.partial_text->text);
}
},
entry);
std::vector<std::string> second_fragments;
const auto second = audiocpp_backend::run_streaming_audio(
session, request, *request.audio_input,
[&](const rt::StreamEvent &event) {
if (event.partial_text.has_value()) {
second_fragments.push_back(event.partial_text->text);
}
},
entry);
check_eq(join(second_fragments), join(first_fragments),
"a re-fetched streaming session replays the same partials");
check_eq(second.text_output.has_value() ? second.text_output->text : "",
first.text_output.has_value() ? first.text_output->text : "",
"a re-fetched streaming session replays the same final text");
check_eq(first.text_output.has_value() ? first.text_output->text : "",
"w1w2w3/frames=1536",
"the first run saw exactly the audio it was given");
// Not a tautology: without the reset the second run reports six words and
// 3072 frames, and both checks above fail.
check_eq(join(first_fragments), "w1|w1w2|w1w2w3", "cumulative partials");
}
static void test_run_streaming_audio_installs_and_clears_the_sink() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
SinkOnlyAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 1024;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
const auto request = audio_request(16000, 1, 1024);
std::vector<std::string> fragments;
const auto result = audiocpp_backend::run_streaming_audio(
session, request, *request.audio_input,
[&](const rt::StreamEvent &event) {
if (event.partial_text.has_value()) {
fragments.push_back(event.partial_text->text);
}
},
entry);
check_eq(join(fragments), "late |partial",
"a family that reports only through the sink is not silent");
check(!fake.sink_installed(),
"the sink is cleared before returning, so the cached session holds no "
"reference to the caller's frame");
check_eq(join(fake.calls), "sink+|prepare|reset|chunk|finalize|sink-",
"the sink is installed before the stream begins and cleared after it ends");
check(result.text_output.has_value(), "the final result still comes back");
}
static void test_the_sink_is_cleared_when_the_stream_throws() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 512;
fake.throw_on_chunk = 1;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
const auto request = audio_request(16000, 1, 1024);
bool threw = false;
try {
audiocpp_backend::run_streaming_audio(
session, request, *request.audio_input,
[](const rt::StreamEvent &) {}, entry);
} catch (const std::exception &) {
threw = true;
}
check(threw, "a failing chunk propagates");
check(!fake.sink_installed(),
"the sink is cleared on the exception path too");
}
static void test_chunking_honours_the_policy_sample_count() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 16000;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
// 2.5 chunks, so the last one is short.
const auto request = audio_request(16000, 1, 40000);
audiocpp_backend::run_streaming_audio(session, request, *request.audio_input,
[](const rt::StreamEvent &) {}, entry);
check(fake.chunks.size() == 3, "40000 frames at 16000 per chunk is three chunks");
if (fake.chunks.size() == 3) {
check(fake.chunks[0].samples.size() == 16000, "first chunk is full");
check(fake.chunks[1].samples.size() == 16000, "second chunk is full");
check(fake.chunks[2].samples.size() == 8000, "last chunk is the remainder");
check(fake.chunks[0].start_sample == 0, "first chunk starts at zero");
check(fake.chunks[1].start_sample == 16000, "second chunk start index");
check(fake.chunks[2].start_sample == 32000, "third chunk start index");
check(fake.chunks[0].sample_rate == 16000, "chunk carries the buffer's rate");
check(fake.chunks[0].channels == 1, "chunk carries the buffer's channel count");
}
}
// A buffer whose float count is not a whole number of frames is REFUSED rather
// than truncated. The integer division would otherwise drop the tail floats
// from the fed audio, and therefore from the transcript, with no diagnostic.
static void test_a_partial_trailing_frame_is_refused() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 100;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
// 501 floats across 2 channels: 250 whole frames and one stray float.
rt::TaskRequest request;
rt::AudioBuffer audio;
audio.sample_rate = 48000;
audio.channels = 2;
audio.samples.assign(501, 0.5F);
request.audio_input = std::move(audio);
bool threw_config = false;
try {
audiocpp_backend::run_streaming_audio(session, request, *request.audio_input,
[](const rt::StreamEvent &) {}, entry);
} catch (const audiocpp_backend::ConfigError &) {
threw_config = true;
} catch (const std::exception &) {
}
check(threw_config,
"a buffer that is not a whole number of frames is refused with ConfigError");
check(fake.calls.empty(),
"the refusal precedes every call into the session, so no half-started "
"stream is left on the cached one");
}
// higgs_audio_stt states its window in seconds and leaves the sample count at
// zero, so this branch is a real family's path rather than a defensive one.
static void test_chunking_falls_back_to_the_policy_seconds() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 0;
fake.policy.preferred_audio_chunk_seconds = 4.0;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
const auto request = audio_request(16000, 1, 96000); // 6 s
audiocpp_backend::run_streaming_audio(session, request, *request.audio_input,
[](const rt::StreamEvent &) {}, entry);
check(fake.chunks.size() == 2, "6 s at a 4 s window is two chunks");
if (fake.chunks.size() == 2) {
check(fake.chunks[0].samples.size() == 64000, "first window is 4 s");
check(fake.chunks[1].samples.size() == 32000, "second window is the 2 s remainder");
}
}
static void test_chunking_falls_back_to_the_interface_default() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 0;
fake.policy.preferred_audio_chunk_seconds = 0.0;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
const auto request = audio_request(16000, 1, 1024);
audiocpp_backend::run_streaming_audio(session, request, *request.audio_input,
[](const rt::StreamEvent &) {}, entry);
check(fake.chunks.size() == 2, "a policy naming no window uses the interface's 512");
if (!fake.chunks.empty()) {
check(fake.chunks[0].samples.size() == 512, "default window is 512 frames");
}
}
// A zero sample rate must not turn a seconds-only policy into a zero-length
// chunk, which would loop forever.
static void test_a_seconds_policy_with_no_rate_falls_through() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 0;
fake.policy.preferred_audio_chunk_seconds = 4.0;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
const auto request = audio_request(0, 1, 1024);
audiocpp_backend::run_streaming_audio(session, request, *request.audio_input,
[](const rt::StreamEvent &) {}, entry);
check(fake.chunks.size() == 2, "a rateless buffer still chunks at the default 512");
}
// FRAMES, not floats. vibevoice_asr refuses a chunk whose sample count is not
// divisible by its channel count, and offsets every span it reports by the
// chunk's start_sample.
static void test_stereo_chunks_are_frame_aligned() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 300;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
const auto request = audio_request(48000, 2, 750);
audiocpp_backend::run_streaming_audio(session, request, *request.audio_input,
[](const rt::StreamEvent &) {}, entry);
check(fake.chunks.size() == 3, "750 frames at 300 frames per chunk is three chunks");
for (const auto &chunk : fake.chunks) {
check(chunk.samples.size() % 2 == 0, "every stereo chunk is a whole number of frames");
}
if (fake.chunks.size() == 3) {
check(fake.chunks[0].samples.size() == 600, "300 stereo frames is 600 floats");
check(fake.chunks[1].start_sample == 300,
"start_sample counts frames, not floats");
check(fake.chunks[2].samples.size() == 300, "the remainder is 150 frames");
}
}
static void test_pull_drains_every_event_and_installs_no_sink() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakePullSession fake;
fake.event_count = 3;
const auto session = streaming_session(fake, audiocpp_backend::Task::Tts);
std::vector<std::string> ids;
rt::TaskRequest request;
const auto result = audiocpp_backend::run_streaming_pull(
session, request,
[&](const rt::StreamEvent &event) {
for (const auto &named : event.named_audio_outputs) {
ids.push_back(named.id);
}
},
entry);
check_eq(join(ids), "chunk_0|chunk_1|chunk_2", "every pulled event reaches the caller");
check(!fake.sink_installed(),
"no stream event sink is installed on the pull path, so voxcpm2 cannot "
"deliver every chunk twice");
check_eq(join(fake.calls), "prepare|start_stream|reset|pull|pull|pull|finalize",
"prepare, start, drain, finish");
check(result.audio_output.has_value(), "the merged result comes back");
check(result.audio_output.has_value() && result.audio_output->samples.size() == 12,
"the merged result is the whole synthesis, not a tail");
}
static void test_pull_stops_on_a_final_event() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakePullSession fake;
fake.event_count = 5;
fake.final_on_second = true;
const auto session = streaming_session(fake, audiocpp_backend::Task::Tts);
int events = 0;
rt::TaskRequest request;
audiocpp_backend::run_streaming_pull(
session, request, [&](const rt::StreamEvent &) { ++events; }, entry);
check(events == 2, "an event marked final ends the pull loop");
}
static void test_pull_refuses_a_non_streaming_session() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
audiocpp_backend::LoadedModel::Session session;
bool threw_capability = false;
try {
rt::TaskRequest request;
audiocpp_backend::run_streaming_pull(
session, request, [](const rt::StreamEvent &) {}, entry);
} catch (const audiocpp_backend::CapabilityError &) {
threw_capability = true;
} catch (const std::exception &) {
}
check(threw_capability, "run_streaming_pull refuses a session with no streaming half");
}
// prepare() runs on EVERY stream, not once per session: the preparation request
// is derived from the request (audio contract, text, voice), so a second stream
// at a different rate would otherwise run against the first one's contract.
static void test_prepare_tracks_the_request_not_the_session() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 4096;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
const auto first = audio_request(16000, 1, 4096);
audiocpp_backend::run_streaming_audio(session, first, *first.audio_input,
[](const rt::StreamEvent &) {}, entry);
check(fake.prepared_rate() == 16000, "the first stream prepares at its own rate");
const auto second = audio_request(44100, 1, 4096);
audiocpp_backend::run_streaming_audio(session, second, *second.audio_input,
[](const rt::StreamEvent &) {}, entry);
check(fake.prepared_rate() == 44100,
"the second stream prepares at ITS rate, not the first one's");
}
// --------------------------------------------------------------------------
// run_streaming_live
// --------------------------------------------------------------------------
// Hands the driver a fixed list of wire frames, the way a client's audio
// callback would, and then closes.
static std::function<bool(std::vector<float> &)>
frames_from(const std::vector<std::size_t> &sizes) {
auto index = std::make_shared<std::size_t>(0);
auto list = std::make_shared<std::vector<std::size_t>>(sizes);
return [index, list](std::vector<float> &out) {
if (*index >= list->size()) {
return false;
}
out.assign((*list)[*index], 0.5F);
++*index;
return true;
};
}
static rt::TaskRequest live_request(int sample_rate, int channels) {
rt::TaskRequest request;
rt::AudioBuffer contract;
contract.sample_rate = sample_rate;
contract.channels = channels;
request.audio_input = std::move(contract); // no samples: none exist yet
return request;
}
// The wire's frame size is a property of the client's audio callback. The
// family's window is a statement about what it can decode. The driver feeds the
// second, not the first.
static void test_live_buffers_wire_frames_into_policy_windows() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 1600;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
// Ten 512-sample frames: 5120 samples, i.e. three full 1600 windows and a
// 320 sample tail.
const auto result = audiocpp_backend::run_streaming_live(
session, live_request(16000, 1),
frames_from(std::vector<std::size_t>(10, 512)),
[](const rt::StreamEvent &) {}, entry);
check(fake.chunks.size() == 4,
"5120 wire samples at a 1600 frame window is three windows and a tail");
if (fake.chunks.size() == 4) {
check(fake.chunks[0].samples.size() == 1600, "first window is full");
check(fake.chunks[2].samples.size() == 1600, "third window is full");
check(fake.chunks[3].samples.size() == 320, "the tail is what was left");
check(fake.chunks[0].start_sample == 0, "the first window starts at zero");
check(fake.chunks[1].start_sample == 1600, "start_sample counts frames");
check(fake.chunks[3].start_sample == 4800, "the tail is offset by all of it");
check(fake.chunks[0].sample_rate == 16000, "the chunk carries the session rate");
}
check(result.text_output.has_value() &&
result.text_output->text == "w1w2w3w4/frames=5120",
"every wire sample reaches the family exactly once");
}
// nemotron_asr's shape: no partials from process_audio_chunk, every one of them
// through the sink from inside finalize.
static void test_live_installs_and_clears_the_sink() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
SinkOnlyAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 512;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
std::vector<std::string> fragments;
audiocpp_backend::run_streaming_live(
session, live_request(16000, 1), frames_from({512}),
[&](const rt::StreamEvent &event) {
if (event.partial_text.has_value()) {
fragments.push_back(event.partial_text->text);
}
},
entry);
check_eq(join(fragments), "late |partial",
"a family that reports only through the sink is not silent live either");
check(!fake.sink_installed(),
"the sink is cleared before returning, so the cached session holds no "
"reference to this call's frame");
check_eq(join(fake.calls), "sink+|prepare|reset|chunk|finalize|sink-",
"sink installed before the stream begins, cleared after it ends");
}
// A client that opens a session and closes it without speaking. finalize is NOT
// called: nemotron_asr throws "finalize requires streamed audio", and an empty
// transcript is the truthful answer to transcribing nothing.
static void test_live_with_no_audio_never_finalizes() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 512;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
const auto result = audiocpp_backend::run_streaming_live(
session, live_request(16000, 1), frames_from({}),
[](const rt::StreamEvent &) {}, entry);
check_eq(join(fake.calls), "sink+|prepare|reset|sink-",
"an empty live stream begins and ends without a chunk or a finalize");
check(!result.text_output.has_value(),
"an empty live stream reports no transcript rather than an error");
}
// A tail shorter than a window is still fed. Without this the last fragment of
// speech never reaches the model, and nothing says so.
static void test_live_feeds_a_short_tail() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 16000;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
audiocpp_backend::run_streaming_live(session, live_request(16000, 1),
frames_from({100, 200}),
[](const rt::StreamEvent &) {}, entry);
check(fake.chunks.size() == 1,
"300 samples against a 16000 frame window is one short chunk, not none");
if (!fake.chunks.empty()) {
check(fake.chunks[0].samples.size() == 300, "the tail carries everything fed");
}
}
// A live request carries no samples, so the CONTRACT is the only thing that says
// what rate the frames are in, and prepare() needs it: nemotron_asr's streaming
// prepare throws without one.
static void test_live_prepares_at_the_contract_rate() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 512;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
audiocpp_backend::run_streaming_live(session, live_request(16000, 1),
frames_from({512}),
[](const rt::StreamEvent &) {}, entry);
check(fake.prepared_rate() == 16000,
"the empty contract buffer still carries the rate into prepare()");
bool threw_config = false;
try {
rt::TaskRequest bare; // no audio_input at all
audiocpp_backend::run_streaming_live(session, bare, frames_from({512}),
[](const rt::StreamEvent &) {}, entry);
} catch (const audiocpp_backend::ConfigError &) {
threw_config = true;
} catch (const std::exception &) {
}
check(threw_config, "a live request with no audio contract is refused");
}
// The pull function is the gRPC read, and a request the handler has to refuse
// mid-stream unwinds through the driver. The sink must not survive it.
static void test_live_clears_the_sink_when_the_puller_throws() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 512;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
bool threw = false;
try {
audiocpp_backend::run_streaming_live(
session, live_request(16000, 1),
[](std::vector<float> &) -> bool {
throw std::runtime_error("fake: the client vanished");
},
[](const rt::StreamEvent &) {}, entry);
} catch (const std::exception &) {
threw = true;
}
check(threw, "a failing pull propagates");
check(!fake.sink_installed(), "the sink is cleared on the pull's exception path");
}
// Two live streams over the SAME cached session must not run into each other.
static void test_live_replays_identically_on_a_refetched_session() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
FakeAudioSession fake;
fake.policy.preferred_audio_chunk_samples = 512;
const auto session = streaming_session(fake, audiocpp_backend::Task::Asr);
const auto first = audiocpp_backend::run_streaming_live(
session, live_request(16000, 1), frames_from({512, 512}),
[](const rt::StreamEvent &) {}, entry);
const auto second = audiocpp_backend::run_streaming_live(
session, live_request(16000, 1), frames_from({512, 512}),
[](const rt::StreamEvent &) {}, entry);
check_eq(second.text_output.has_value() ? second.text_output->text : "",
first.text_output.has_value() ? first.text_output->text : "",
"a re-fetched live session replays the same transcript");
// Not a tautology: without the reset the second run reports w1..w4 and 2048
// frames.
check_eq(first.text_output.has_value() ? first.text_output->text : "",
"w1w2/frames=1024", "the first live run saw exactly what was fed");
}
static void test_live_refuses_a_non_streaming_session() {
audiocpp_backend::InferenceLane lane("test");
audiocpp_backend::LaneEntry entry(lane, 0);
audiocpp_backend::LoadedModel::Session session;
session.mode = audiocpp_backend::Mode::Offline;
bool threw_capability = false;
try {
audiocpp_backend::run_streaming_live(session, live_request(16000, 1),
frames_from({512}),
[](const rt::StreamEvent &) {}, entry);
} catch (const audiocpp_backend::CapabilityError &) {
threw_capability = true;
} catch (const std::exception &) {
}
check(threw_capability,
"run_streaming_live refuses a session with no streaming half");
}
// --------------------------------------------------------------------------
// resolve_model_path
// --------------------------------------------------------------------------
//
// It lives in loaded_model.cpp and is a pure (dir, file, name) -> string, so it
// is tested here rather than in a standalone unit: that file cannot compile
// without the engine headers.
//
// It is tested at all because it shipped a bug no test could have caught. THE
// SHAPES BELOW ARE THE PRODUCTION SHAPES, not convenient ones, and that
// distinction is the entire point. Task 15 verified the bundled: form with a
// hand-written LoadModel that left ModelFile empty, which is the one shape the
// server never produces: pkg/model/loader.go's LoadModelWithFile always fills
// ModelFile with filepath.Join(ModelPath, model), and core/backend/options.go
// only overrides it for a managed artifact. The first case below is therefore
// the regression test; the other three are what it must not have broken.
// std::string::ends_with is C++20 and this target is C++17.
static bool ends_with(const std::string &value, const std::string &suffix) {
return value.size() >= suffix.size() &&
value.compare(value.size() - suffix.size(), suffix.size(), suffix) == 0;
}
// THE REGRESSION CASE. What a model YAML saying `model: bundled:silero_vad`
// actually arrives as: Model intact, ModelFile joined onto the models directory.
static void test_bundled_in_model_survives_a_joined_model_file() {
const std::string resolved = audiocpp_backend::resolve_model_path(
"/models", "/models/bundled:silero_vad", "bundled:silero_vad");
check(ends_with(resolved, "/assets/silero_vad"),
"bundled: in Model resolves under the package assets dir (got \"" +
resolved + "\")");
// Checked separately from the suffix because this is the failure that
// shipped: the joined ModelFile came back verbatim and the load died on
// "model path does not exist: /models/bundled:silero_vad".
check(resolved.find("/models/") == std::string::npos,
"bundled: in Model is not resolved against the models directory (got \"" +
resolved + "\")");
}
// Task 15's shape: the form in ModelFile with Model empty. It worked before the
// fix and must keep working.
static void test_bundled_in_model_file_still_resolves() {
const std::string resolved =
audiocpp_backend::resolve_model_path("", "bundled:marblenet_vad", "");
check(ends_with(resolved, "/assets/marblenet_vad"),
"bundled: in ModelFile still resolves under the package assets dir (got \"" +
resolved + "\")");
}
// The ordinary case, and the one the bundled: lookup must not capture: a real
// artifact path in ModelFile with a plain name in Model.
static void test_a_plain_name_resolves_to_the_model_file() {
const std::string resolved = audiocpp_backend::resolve_model_path(
"/models", "/models/chatterbox-q8_0.gguf", "chatterbox-q8_0.gguf");
check_eq(resolved, "/models/chatterbox-q8_0.gguf",
"a plain name resolves to the absolute ModelFile");
}
// A relative ModelFile is still joined onto ModelPath. The fix does not touch
// this branch, which is why it is pinned: the bundled: lookup now runs before it
// and has to fall through for every non-bundled input.
static void test_a_relative_model_file_joins_the_model_path() {
const std::string resolved = audiocpp_backend::resolve_model_path(
"/models", "sub/nemotron-asr-q8_0.gguf", "nemotron-asr");
check_eq(resolved, "/models/sub/nemotron-asr-q8_0.gguf",
"a relative ModelFile joins the models directory");
}
int main() {
test_bundled_in_model_survives_a_joined_model_file();
test_bundled_in_model_file_still_resolves();
test_a_plain_name_resolves_to_the_model_file();
test_a_relative_model_file_joins_the_model_path();
test_begin_stream_prepares_then_starts();
test_base_start_stream_resets();
test_begin_stream_refuses_a_non_streaming_session();
test_a_refetched_session_replays_identically();
test_run_streaming_audio_installs_and_clears_the_sink();
test_the_sink_is_cleared_when_the_stream_throws();
test_chunking_honours_the_policy_sample_count();
test_a_partial_trailing_frame_is_refused();
test_chunking_falls_back_to_the_policy_seconds();
test_chunking_falls_back_to_the_interface_default();
test_a_seconds_policy_with_no_rate_falls_through();
test_stereo_chunks_are_frame_aligned();
test_pull_drains_every_event_and_installs_no_sink();
test_pull_stops_on_a_final_event();
test_pull_refuses_a_non_streaming_session();
test_prepare_tracks_the_request_not_the_session();
test_live_buffers_wire_frames_into_policy_windows();
test_live_installs_and_clears_the_sink();
test_live_with_no_audio_never_finalizes();
test_live_feeds_a_short_tail();
test_live_prepares_at_the_contract_rate();
test_live_clears_the_sink_when_the_puller_throws();
test_live_replays_identically_on_a_refetched_session();
test_live_refuses_a_non_streaming_session();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all streaming driver checks passed\n");
return 0;
}

View File

@@ -0,0 +1,215 @@
#include "transcript_assembly.h"
#include "audio_units.h"
#include <algorithm>
#include <cstdlib>
#include <limits>
namespace audiocpp_backend {
namespace {
std::int64_t midpoint(const Span &span) {
return span.start_sample + (span.end_sample - span.start_sample) / 2;
}
bool contains(const Span &span, std::int64_t sample) {
return sample >= span.start_sample && sample < span.end_sample;
}
std::int64_t overlap(const Span &a, const Span &b) {
const std::int64_t begin = std::max(a.start_sample, b.start_sample);
const std::int64_t end = std::min(a.end_sample, b.end_sample);
return end > begin ? end - begin : 0;
}
// Joins a segment's words into that segment's text.
//
// THE SEPARATOR IS NOT ALWAYS A SPACE, and getting it wrong is visible to every
// caller rather than cosmetic: core/http/endpoints/openai/transcription.go
// routes response_format text, srt, vtt and lrc through
// schema.TranscriptionResponse, which builds the entire body out of
// Segments[].Text and never reads the top-level text. For those four formats
// the segment text IS the response.
//
// Two producer conventions have to be told apart:
//
// whole words "Some", "call", "me" -> join with a space
// subword pieces "So", "me", " call" -> concatenate
//
// The second is SentencePiece, where a word boundary is carried as a LEADING
// SPACE on the piece; nemotron_asr emits one entry per token in exactly that
// form. Space-joining those produced "So me call me na ture ,", which is
// what response_format=text returned while the correct sentence sat unread in
// the top-level field. Concatenating them reproduces text_output exactly.
//
// The convention is read off the words themselves, because nothing else in the
// result declares it. One leading space anywhere is enough to decide: a
// whole-word producer has no reason to emit one, and a subword producer emits
// one at every word boundary, so the two populations do not overlap. A producer
// that mixed both conventions inside one segment could not be served correctly
// by any single separator; this picks concatenation for it.
//
// This does NOT touch the top-level text, which stays text_output verbatim. The
// rule that forbids deriving the transcript from the segments is about the
// direction segments -> text. Segment text has no source other than its words
// and is necessarily derived.
std::string join_words(const std::vector<OutWord> &words) {
const bool subword_pieces =
std::any_of(words.begin(), words.end(), [](const OutWord &word) {
return !word.text.empty() && word.text.front() == ' ';
});
std::string out;
for (const auto &word : words) {
if (word.text.empty()) {
continue;
}
if (!subword_pieces && !out.empty()) {
out += " ";
}
out += word.text;
}
return out;
}
std::string speaker_for(const Span &segment,
const std::vector<SpeakerSpan> &turns) {
std::string best;
std::int64_t best_overlap = 0;
for (const auto &turn : turns) {
const std::int64_t shared = overlap(segment, turn.span);
if (shared > best_overlap) {
best_overlap = shared;
best = turn.speaker;
}
}
return best;
}
// The chosen segmentation. labels is empty unless the spans were sourced from
// the speaker turns themselves, in which case it is parallel to spans and holds
// the label each span arrived with.
struct SegmentSource {
std::vector<Span> spans;
std::vector<std::string> labels;
};
// Chooses the segment spans, per the documented precedence.
SegmentSource choose_segment_spans(const std::string &text_output,
const std::vector<Span> &speech_segments,
const std::vector<SpeakerSpan> &speaker_turns,
const std::vector<WordSpan> &words) {
if (!speech_segments.empty()) {
return {speech_segments, {}};
}
if (!speaker_turns.empty()) {
// The labels are carried out rather than re-derived by overlap later. A
// turn wholly contained in another speaker's turn overlaps its own span
// completely, which is the largest overlap possible, so it can only tie
// with the containing turn and would then lose that tie on order.
// sortformer_diar binarizes each speaker's track independently and
// sorts the result by start sample, so the container always comes
// first, and the interjecting speaker would be silently relabelled to
// the speaker it interrupted.
SegmentSource source;
source.spans.reserve(speaker_turns.size());
source.labels.reserve(speaker_turns.size());
for (const auto &turn : speaker_turns) {
source.spans.push_back(turn.span);
source.labels.push_back(turn.speaker);
}
return source;
}
if (!words.empty()) {
Span covering = words.front().span;
for (const auto &word : words) {
covering.start_sample =
std::min(covering.start_sample, word.span.start_sample);
covering.end_sample = std::max(covering.end_sample, word.span.end_sample);
}
return {{covering}, {}};
}
if (!text_output.empty()) {
// A zero span rather than a fabricated duration: the model reported no
// timing, and inventing one would be a lie the caller cannot detect.
return {{Span{0, 0}}, {}};
}
return {};
}
// Returns the index of the segment a word belongs to, or the nearest segment
// when the word falls outside all of them.
size_t segment_index_for_word(const std::vector<Span> &spans, const Span &word) {
const std::int64_t centre = midpoint(word);
for (size_t i = 0; i < spans.size(); ++i) {
if (contains(spans[i], centre)) {
return i;
}
}
size_t nearest = 0;
std::int64_t best_distance = std::numeric_limits<std::int64_t>::max();
for (size_t i = 0; i < spans.size(); ++i) {
const std::int64_t distance = std::llabs(midpoint(spans[i]) - centre);
if (distance < best_distance) {
best_distance = distance;
nearest = i;
}
}
return nearest;
}
} // namespace
AssembledTranscript assemble_transcript(const std::string &text_output,
const std::vector<Span> &speech_segments,
const std::vector<SpeakerSpan> &speaker_turns,
const std::vector<WordSpan> &words,
int sample_rate) {
AssembledTranscript assembled;
// THE RULE. Never derived from spans.
assembled.text = text_output;
const SegmentSource source =
choose_segment_spans(text_output, speech_segments, speaker_turns, words);
const std::vector<Span> &spans = source.spans;
if (spans.empty()) {
return assembled;
}
assembled.segments.resize(spans.size());
for (size_t i = 0; i < spans.size(); ++i) {
OutSegment &segment = assembled.segments[i];
segment.id = static_cast<int>(i);
segment.start_ns = samples_to_nanoseconds(spans[i].start_sample, sample_rate);
segment.end_ns = samples_to_nanoseconds(spans[i].end_sample, sample_rate);
// A segment that came from a speaker turn already knows its speaker.
// Only the other three sources have to look one up by overlap.
segment.speaker = source.labels.empty()
? speaker_for(spans[i], speaker_turns)
: source.labels[i];
}
for (const auto &word : words) {
const size_t index = segment_index_for_word(spans, word.span);
OutWord out;
out.start_ns = samples_to_nanoseconds(word.span.start_sample, sample_rate);
out.end_ns = samples_to_nanoseconds(word.span.end_sample, sample_rate);
out.text = word.word;
assembled.segments[index].words.push_back(out);
}
for (auto &segment : assembled.segments) {
segment.text = join_words(segment.words);
}
// A single segment with no word timing carries the whole transcript. With
// several segments there is no defensible way to split the text, so their
// per-segment text stays empty and only the top-level text is authoritative.
if (assembled.segments.size() == 1 && assembled.segments[0].words.empty()) {
assembled.segments[0].text = text_output;
}
return assembled;
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,87 @@
#pragma once
// Builds LocalAI's TranscriptResult shape from audio.cpp's TaskResult spans.
// Standard library only; result_map.cpp converts the engine types into these
// PODs at the boundary.
//
// THE RULE: the top-level transcript text is text_output verbatim, always.
// audio.cpp carries transcript text in exactly one place, TaskResult.text_output.
// speech_segments, speaker_turns and word_timestamps carry spans and labels but
// no text. Deriving the top-level text by concatenating per-segment text
// therefore yields an empty transcript for every producer that reports segments
// without word timestamps, which includes VibeVoice diarized ASR.
#include <cstdint>
#include <string>
#include <vector>
namespace audiocpp_backend {
// Sample-index span, mirroring engine::runtime::TimeSpan.
struct Span {
std::int64_t start_sample = 0;
std::int64_t end_sample = 0;
};
struct WordSpan {
Span span;
std::string word;
};
struct SpeakerSpan {
Span span;
std::string speaker;
};
struct OutWord {
std::int64_t start_ns = 0;
std::int64_t end_ns = 0;
std::string text;
};
struct OutSegment {
int id = 0;
std::int64_t start_ns = 0;
std::int64_t end_ns = 0;
std::string text;
std::string speaker;
std::vector<OutWord> words;
};
struct AssembledTranscript {
std::string text;
std::vector<OutSegment> segments;
};
// Segment source, first non-empty wins:
// 1. speech_segments
// 2. speaker_turns
// 3. one segment spanning all words, when words are present
// 4. one zero-span segment carrying the full text, when text is present
// 5. no segments
//
// Words attach to the segment whose range contains their midpoint; a word
// outside every segment attaches to the nearest one by midpoint distance so it
// is never silently dropped. A lone segment with no words carries the full text.
//
// A segment's text is its words joined, and the separator depends on the
// producer's convention: whole words ("Some", "call") are joined with a space,
// while SentencePiece-style subword pieces, which carry the word boundary as a
// LEADING SPACE (" call"), are concatenated. One leading space anywhere in the
// segment selects concatenation. This matters beyond tidiness: response_format
// text, srt, vtt and lrc build their entire body out of the segment text and
// never read the top-level text.
//
// A segment's speaker is the speaker turn with the greatest overlap, except
// when the segments came from the speaker turns themselves (source 2), where
// each segment keeps its own turn's label. Re-deriving it there loses a turn
// nested inside another speaker's turn: the nested turn overlaps its own span
// completely, so it can only tie with the containing turn, which is listed
// first and wins the tie.
AssembledTranscript assemble_transcript(const std::string &text_output,
const std::vector<Span> &speech_segments,
const std::vector<SpeakerSpan> &speaker_turns,
const std::vector<WordSpan> &words,
int sample_rate);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,532 @@
// Unit tests for transcript_assembly. Standard library only. The harness
// compiles this as a single translation unit, so both implementations are
// included directly rather than linked.
//
// Every fixture below either mirrors a producer shape actually observed from
// audio.cpp families, and names the families it was checked against, or says in
// its own comment that it is defensive. Do not replace an observed shape with an
// invented one and do not quietly promote a defensive fixture to an observed
// one: an invented shape is what let the earlier attempt ship an empty
// transcript.
#include "audio_units.cpp"
#include "transcript_assembly.cpp"
#include <cstddef>
#include <cstdio>
#include <string>
#include <vector>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
using namespace audiocpp_backend;
// Indexed access that reports a named failure instead of running off the end.
// std::vector::operator[] past the end is undefined behaviour, so a regression
// that drops a segment would crash the process here and take every later check
// with it. Returning a default element keeps the rest of the suite reporting.
static const OutSegment &segment_at(const AssembledTranscript &out, size_t index,
const std::string &name) {
static const OutSegment missing;
if (index >= out.segments.size()) {
failures++;
fprintf(stderr, "FAIL: %s (segment %zu is missing)\n", name.c_str(), index);
return missing;
}
return out.segments[index];
}
static const OutWord &word_at(const OutSegment &segment, size_t index,
const std::string &name) {
static const OutWord missing;
if (index >= segment.words.size()) {
failures++;
fprintf(stderr, "FAIL: %s (word %zu is missing)\n", name.c_str(), index);
return missing;
}
return segment.words[index];
}
static const int kRate = 16000;
// Shape A: word timestamps only. Emitted by nemotron_asr, qwen3_asr and
// qwen3_forced_aligner, all of which set text_output plus word_timestamps and
// leave speech_segments empty.
static void test_words_only() {
const std::vector<WordSpan> words = {
{{0, 8000}, "hello"},
{{8000, 16000}, "world"},
};
const auto out = assemble_transcript("hello world", {}, {}, words, kRate);
check(out.text == "hello world", "text is text_output verbatim");
check(out.segments.size() == 1, "words with no segments yield one segment");
const OutSegment &first = segment_at(out, 0, "words only segment");
check(first.start_ns == 0, "segment starts at the first word");
check(first.end_ns == 1000000000LL, "segment ends at the last word");
check(first.words.size() == 2, "both words attached");
check(word_at(first, 0, "first word").text == "hello", "first word text");
check(word_at(first, 1, "second word").start_ns == 500000000LL,
"second word start in ns");
check(first.text == "hello world", "segment text joins its words");
check(first.id == 0, "ids are zero based");
}
// Shape A-whole: the whole-word convention, stated explicitly rather than left
// implicit in the shape A tests. qwen3_forced_aligner emits one entry per WORD
// (processor.cpp parses per-word timestamp tokens), so its pieces carry no
// leading space and must be joined with one.
static void test_whole_words_are_space_joined() {
const std::vector<WordSpan> words = {
{{0, 8000}, "Some"},
{{8000, 16000}, "call"},
{{16000, 24000}, "me"},
};
const auto out = assemble_transcript("Some call me", {}, {}, words, kRate);
check(segment_at(out, 0, "whole words").text == "Some call me",
"whole words are joined with a single space");
}
// Shape A-subword: the SentencePiece convention, where the word boundary is a
// LEADING SPACE on the piece. These are the first eleven word_timestamps
// nemotron_asr actually returned for audio.cpp/assets/resources/sample_16k.wav
// with the q8_0 GGUF, copied verbatim rather than invented, including the lone
// " " piece at index 3.
//
// Space-joining these produced "So me call me na ture , other s call",
// which is not a cosmetic problem: response_format text, srt, vtt and lrc build
// their entire body from the segment text and never read the top-level text, so
// that string WAS the transcription response for those formats.
static void test_subword_pieces_are_concatenated() {
const std::vector<WordSpan> words = {
{{15360, 16640}, "So"}, {{15360, 16640}, "me"},
{{23040, 24320}, " call"}, {{28160, 29440}, " "},
{{28160, 29440}, "me"}, {{30720, 32000}, " na"},
{{33280, 34560}, "ture"}, {{35840, 37120}, ","},
{{38400, 39680}, " other"}, {{40960, 42240}, "s"},
{{43520, 44800}, " call"},
};
const auto out = assemble_transcript(
"Some call me nature, others call me mother nature.", {}, {}, words, kRate);
check(out.text == "Some call me nature, others call me mother nature.",
"the top-level text is still text_output verbatim");
check(segment_at(out, 0, "subword pieces").text ==
"Some call me nature, others call",
"subword pieces are concatenated, reproducing text_output");
}
// One leading space anywhere decides for the whole segment. A subword producer
// emits a boundary space at every word start, so its first piece, which is
// sentence-initial, does not have one; keying off the first piece alone would
// therefore pick the wrong convention on every segment.
static void test_a_single_leading_space_selects_concatenation() {
const std::vector<WordSpan> words = {
{{0, 8000}, "al"},
{{8000, 16000}, "pha"},
{{16000, 24000}, " beta"},
};
const auto out = assemble_transcript("alpha beta", {}, {}, words, kRate);
check(segment_at(out, 0, "mixed").text == "alpha beta",
"a leading space on a later piece selects concatenation");
}
// Shape A': the same producer, but text_output is punctuated and cased while
// the word timestamps are not. qwen3_asr rebuilds text_output from its word
// list only when timestamps are requested, so the two genuinely differ; this
// pins the sole segment's text to its words rather than to the top-level text.
static void test_words_only_with_punctuated_text_output() {
const std::vector<WordSpan> words = {
{{0, 8000}, "hello"},
{{8000, 16000}, "world"},
};
const auto out = assemble_transcript("Hello, world!", {}, {}, words, kRate);
check(out.text == "Hello, world!", "punctuated text_output is untouched");
check(out.segments.size() == 1, "one segment");
check(segment_at(out, 0, "punctuated segment").text == "hello world",
"a segment with words takes its text from the words, not text_output");
}
// Shape A'': a merged word list whose last word is not the one that ends
// latest. audio.cpp concatenates per-chunk word lists in chunk order
// (append_chunk_word_timestamps in framework/audio/chunking.cpp). It drops a
// word whose global start falls before the chunk's keep span, but it never
// clips a word's end to that boundary, so the last word kept from one chunk can
// outlast the first word kept from the next. The covering span must therefore
// be the extent of every word, not the span from the first to the last.
static void test_covering_span_spans_every_word() {
const std::vector<WordSpan> words = {
{{0, 4000}, "a"},
// Kept from the earlier chunk, ending past the chunk boundary.
{{4000, 10000}, "b"},
// First word of the next chunk, shorter, so it ends earlier.
{{8000, 9000}, "c"},
};
const auto out = assemble_transcript("a b c", {}, {}, words, kRate);
check(out.segments.size() == 1, "one covering segment");
check(segment_at(out, 0, "covering segment").end_ns == 625000000LL,
"the covering span reaches the latest word end, not the last word's");
}
// Defensive, not observed: no pinned family emits a word with no text.
// nemotron_asr's build_token_timestamps (models/nemotron_asr/decoder.cpp:97)
// skips a token that decodes to an empty chunk before it ever becomes a
// WordTimestamp. join_words guards against one anyway, and an unexercised guard
// is a guard the next reader deletes as dead weight.
static void test_empty_word_contributes_no_separator() {
const std::vector<WordSpan> words = {
{{0, 4000}, "alpha"},
{{4000, 8000}, ""},
{{8000, 12000}, "beta"},
};
const auto out = assemble_transcript("alpha beta", {}, {}, words, kRate);
check(out.segments.size() == 1, "one segment");
check(segment_at(out, 0, "sole segment").text == "alpha beta",
"an empty word adds no separator to the segment text");
check(segment_at(out, 0, "sole segment").words.size() == 3,
"the empty word still reports its span");
}
// Shape B: speech segments, no words. Emitted by ASR families that report
// utterance boundaries without word-level timing.
static void test_segments_without_words() {
const std::vector<Span> segments = {{0, 16000}, {16000, 32000}};
const auto out = assemble_transcript("one two three", segments, {}, {}, kRate);
// The regression: this must NOT be empty.
check(out.text == "one two three", "multi-segment text is not empty");
check(out.segments.size() == 2, "both segments survive");
check(segment_at(out, 0, "first segment").end_ns == 1000000000LL,
"first segment ends at 1s");
check(segment_at(out, 1, "second segment").start_ns == 1000000000LL,
"second segment starts at 1s");
check(segment_at(out, 0, "first segment").text.empty(),
"per-segment text stays empty when there are no words to split by");
check(segment_at(out, 1, "second segment").id == 1, "ids increment");
}
// Shape C: speech segments plus speaker turns, no words. This is the real
// VibeVoice diarized ASR shape that broke the earlier attempt.
static void test_segments_with_speaker_turns_no_words() {
const std::vector<Span> segments = {{0, 16000}, {16000, 32000}};
const std::vector<SpeakerSpan> turns = {
{{0, 16000}, "SPEAKER_00"},
{{16000, 32000}, "SPEAKER_01"},
};
const auto out = assemble_transcript("hi there", segments, turns, {}, kRate);
check(out.text == "hi there", "diarized multi-segment text is not empty");
check(out.segments.size() == 2, "two segments");
check(segment_at(out, 0, "first diarized segment").speaker == "SPEAKER_00",
"first speaker assigned");
check(segment_at(out, 1, "second diarized segment").speaker == "SPEAKER_01",
"second speaker assigned");
}
// Defensive, not observed: speech segments and speaker turns that disagree.
// vibevoice_asr builds each SpeakerTurn with turn.span = speech_segment.span in
// one loop (models/vibevoice_asr/session.cpp:965) and shifts and clips both
// lists identically when merging chunks, so in practice the two lists are 1:1
// with identical spans. That is exactly why the shape C fixture above cannot
// show which list is the segment source: swapping the precedence there produces
// byte-identical output. This fixture pins the precedence, and it is the shape
// any future family that segments and diarizes separately would produce.
static void test_speech_segments_outrank_speaker_turns() {
const std::vector<Span> segments = {{0, 32000}};
const std::vector<SpeakerSpan> turns = {
{{0, 16000}, "SPEAKER_00"},
{{16000, 32000}, "SPEAKER_01"},
};
const auto out = assemble_transcript("hi there", segments, turns, {}, kRate);
check(out.segments.size() == 1,
"speech segments decide the segmentation, not speaker turns");
check(segment_at(out, 0, "single utterance").end_ns == 2000000000LL,
"the utterance keeps its own span");
}
// Defensive, not observed: no pinned family emits speaker turns and word
// timestamps together. It pins rule 2 against rule 3, which nothing else does:
// a diarized result is segmented by who spoke, and words only fill the turns in.
static void test_speaker_turns_outrank_words() {
const std::vector<SpeakerSpan> turns = {
{{0, 16000}, "SPEAKER_00"},
{{16000, 32000}, "SPEAKER_01"},
};
const std::vector<WordSpan> words = {
{{0, 8000}, "hi"},
{{16000, 24000}, "there"},
};
const auto out = assemble_transcript("hi there", {}, turns, words, kRate);
check(out.segments.size() == 2, "the two turns segment the result");
check(segment_at(out, 0, "turn 0").text == "hi", "first turn takes its word");
check(segment_at(out, 1, "turn 1").text == "there", "second turn takes its word");
}
// Shape D: text only. Emitted by ASR families that report no timing at all,
// such as hviske_asr and citrinet_asr.
static void test_text_only() {
const auto out = assemble_transcript("just text", {}, {}, {}, kRate);
check(out.text == "just text", "text survives");
check(out.segments.size() == 1, "a single synthetic segment is emitted");
const OutSegment &only = segment_at(out, 0, "synthetic segment");
check(only.start_ns == 0 && only.end_ns == 0,
"synthetic segment has zero span, not a fabricated duration");
check(only.text == "just text", "the sole segment carries the full text");
}
// Shape E: speaker turns only, no speech segments and no text. This is
// sortformer_diar, reached through the Diarize RPC.
static void test_speaker_turns_only() {
const std::vector<SpeakerSpan> turns = {
{{0, 24000}, "0"},
{{24000, 48000}, "1"},
};
const auto out = assemble_transcript("", {}, turns, {}, kRate);
check(out.text.empty(), "no text is reported when the model produced none");
check(out.segments.size() == 2, "turns become segments");
check(segment_at(out, 0, "turn 0").speaker == "0",
"speaker label preserved verbatim");
check(segment_at(out, 1, "turn 1").start_ns == 1500000000LL,
"second turn starts at 1.5s");
}
// Shape E', the same producer with one speaker talking over another.
// decode_sortformer_speaker_turns (models/sortformer_diar/postprocess.cpp)
// binarizes each speaker's probability track independently, which is the whole
// point of sortformer, then sorts the turns by start sample. So a turn can be
// wholly contained in another speaker's turn, and the containing turn always
// comes first. A segment sourced from a speaker turn must keep that turn's own
// label: re-deriving it by overlap can only ever tie with the containing turn,
// which then wins on order and silently erases the interjecting speaker.
static void test_nested_speaker_turn_keeps_its_own_label() {
const std::vector<SpeakerSpan> turns = {
{{0, 100000}, "speaker_0"},
{{10000, 20000}, "speaker_1"},
};
const auto out = assemble_transcript("", {}, turns, {}, kRate);
check(out.segments.size() == 2, "both turns become segments");
check(segment_at(out, 0, "containing turn").speaker == "speaker_0",
"the containing turn keeps its label");
check(segment_at(out, 1, "nested turn").speaker == "speaker_1",
"a turn nested inside another is not relabelled to the container");
}
// Shape F: nothing at all. A model that ran but produced no output must not
// crash or fabricate a segment.
static void test_empty() {
const auto out = assemble_transcript("", {}, {}, {}, kRate);
check(out.text.empty(), "empty stays empty");
check(out.segments.empty(), "no segments are invented");
}
// Shape H: speech segments with no text and no words at all. This is the VAD
// path, silero_vad and marblenet_vad, which fill speech_segments and never
// touch text_output. It reaches the lone-segment rule with nothing to carry.
static void test_vad_segments_without_text() {
const std::vector<Span> segments = {{0, 16000}, {24000, 32000}};
const auto out = assemble_transcript("", segments, {}, {}, kRate);
check(out.text.empty(), "VAD reports no text");
check(out.segments.size() == 2, "both speech regions survive");
check(segment_at(out, 1, "second speech region").start_ns == 1500000000LL,
"second region starts at 1.5s");
check(segment_at(out, 0, "first speech region").text.empty(),
"a VAD segment carries no text");
const std::vector<Span> one = {{0, 16000}};
const auto single = assemble_transcript("", one, {}, {}, kRate);
check(single.segments.size() == 1, "a single speech region survives");
check(segment_at(single, 0, "lone speech region").text.empty(),
"a lone VAD segment does not fabricate text");
}
// Shape G: segments and words together. Words are assigned by midpoint so a
// word straddling a boundary lands in exactly one segment.
static void test_words_distributed_into_segments() {
const std::vector<Span> segments = {{0, 16000}, {16000, 32000}};
const std::vector<WordSpan> words = {
{{0, 4000}, "alpha"},
{{4000, 8000}, "beta"},
// Straddles the boundary; midpoint 16000 falls in the second segment.
{{12000, 20000}, "gamma"},
{{20000, 28000}, "delta"},
};
const auto out = assemble_transcript("alpha beta gamma delta", segments, {},
words, kRate);
check(out.text == "alpha beta gamma delta", "top level text unchanged");
check(out.segments.size() == 2, "two segments");
check(segment_at(out, 0, "first segment").words.size() == 2,
"first segment takes two words");
check(segment_at(out, 1, "second segment").words.size() == 2,
"second segment takes two words");
check(segment_at(out, 0, "first segment").text == "alpha beta",
"first segment text");
check(segment_at(out, 1, "second segment").text == "gamma delta",
"boundary-straddling word lands by midpoint");
}
// The midpoint rule is not the same as either endpoint rule. "early" starts in
// the first segment but ends in the second, and "late" the other way round;
// each must land where its midpoint says, which no start-only or end-only rule
// reproduces.
static void test_words_assigned_by_midpoint_not_endpoint() {
const std::vector<Span> segments = {{0, 16000}, {16000, 32000}};
const std::vector<WordSpan> words = {
// Midpoint 12000 -> first segment, although it ends in the second.
{{4000, 20000}, "early"},
// Midpoint 20000 -> second segment, although it starts in the first.
{{12000, 28000}, "late"},
};
const auto out = assemble_transcript("early late", segments, {}, words, kRate);
check(segment_at(out, 0, "first segment").text == "early",
"a word ending past the boundary stays where its midpoint is");
check(segment_at(out, 1, "second segment").text == "late",
"a word starting before the boundary follows its midpoint");
}
// A word outside every segment must still be reachable rather than dropped
// silently, so it attaches to the nearest segment by midpoint distance.
static void test_word_outside_all_segments() {
const std::vector<Span> segments = {{0, 16000}};
const std::vector<WordSpan> words = {
{{0, 8000}, "inside"},
{{40000, 48000}, "outside"},
};
const auto out = assemble_transcript("inside outside", segments, {}, words,
kRate);
check(out.segments.size() == 1, "one segment");
check(segment_at(out, 0, "sole segment").words.size() == 2,
"the stray word is not dropped");
}
// The fallback picks the nearest segment, which is not the same as picking the
// first. With one segment the two are indistinguishable, so this uses three and
// puts the stray word past the last one.
static void test_stray_word_goes_to_the_nearest_segment() {
const std::vector<Span> segments = {{0, 8000}, {8000, 16000}, {16000, 24000}};
const std::vector<WordSpan> words = {
// Midpoint 44000, nearest the third segment.
{{40000, 48000}, "trailing"},
};
const auto out = assemble_transcript("trailing", segments, {}, words, kRate);
check(segment_at(out, 0, "first segment").words.empty(),
"the stray word does not fall back to the first segment");
check(segment_at(out, 2, "third segment").text == "trailing",
"the stray word attaches to the nearest segment");
}
// "Nearest" is measured from the segment's midpoint, and it is neither "the
// first segment" nor "the last". A leading stray word is the case a
// trailing-only fixture cannot reach: forced-aligner words scored against VAD
// segments produce one, and with only trailing coverage it would land at the
// end of the transcript with the suite green. Here the leading word's nearest
// midpoint is the first segment while its nearest start is the second, and the
// trailing word's nearest midpoint is the third while its nearest end is the
// second, so no endpoint rule reproduces this assignment either.
static void test_stray_word_distance_is_measured_from_the_midpoint() {
const std::vector<Span> segments = {{0, 2000}, {8000, 200000}, {300000, 302000}};
const std::vector<WordSpan> words = {
{{4000, 6000}, "lead"},
{{249000, 251000}, "trail"},
};
const auto out = assemble_transcript("lead trail", segments, {}, words, kRate);
check(out.segments.size() == 3, "three segments");
check(segment_at(out, 0, "first segment").text == "lead",
"the leading stray word goes to the nearest segment by midpoint");
check(segment_at(out, 2, "third segment").text == "trail",
"the trailing stray word goes to the nearest segment by midpoint");
check(segment_at(out, 1, "middle segment").words.empty(),
"the long middle segment claims neither stray word");
}
// Speaker assignment uses greatest overlap, not first match, so a turn that
// barely touches a segment does not win over one that covers it.
static void test_speaker_assigned_by_greatest_overlap() {
const std::vector<Span> segments = {{8000, 24000}};
const std::vector<SpeakerSpan> turns = {
{{0, 9000}, "brief"}, // overlaps 1000 samples
{{9000, 24000}, "main"} // overlaps 15000 samples
};
const auto out = assemble_transcript("x", segments, turns, {}, kRate);
check(out.segments.size() == 1, "one segment");
check(segment_at(out, 0, "sole segment").speaker == "main",
"greatest overlap wins");
}
// A segment no turn touches gets no speaker rather than the label of whichever
// turn happened to be listed first.
static void test_segment_without_any_overlapping_turn_has_no_speaker() {
const std::vector<Span> segments = {{0, 8000}, {40000, 48000}};
const std::vector<SpeakerSpan> turns = {{{0, 8000}, "SPEAKER_00"}};
const auto out = assemble_transcript("x", segments, turns, {}, kRate);
check(segment_at(out, 0, "overlapped segment").speaker == "SPEAKER_00",
"the overlapped segment is labelled");
check(segment_at(out, 1, "unlabelled segment").speaker.empty(),
"a segment no turn overlaps is left unlabelled");
}
static void test_zero_sample_rate_is_safe() {
const std::vector<Span> segments = {{0, 16000}};
const auto out = assemble_transcript("x", segments, {}, {}, 0);
check(out.segments.size() == 1, "a zero sample rate still yields the segment");
const OutSegment &only = segment_at(out, 0, "sole segment");
check(only.start_ns == 0 && only.end_ns == 0,
"unknown sample rate yields zero timings rather than garbage");
}
int main() {
test_words_only();
test_whole_words_are_space_joined();
test_subword_pieces_are_concatenated();
test_a_single_leading_space_selects_concatenation();
test_words_only_with_punctuated_text_output();
test_covering_span_spans_every_word();
test_empty_word_contributes_no_separator();
test_segments_without_words();
test_segments_with_speaker_turns_no_words();
test_speech_segments_outrank_speaker_turns();
test_speaker_turns_outrank_words();
test_text_only();
test_speaker_turns_only();
test_nested_speaker_turn_keeps_its_own_label();
test_empty();
test_vad_segments_without_text();
test_words_distributed_into_segments();
test_words_assigned_by_midpoint_not_endpoint();
test_word_outside_all_segments();
test_stray_word_goes_to_the_nearest_segment();
test_stray_word_distance_is_measured_from_the_midpoint();
test_speaker_assigned_by_greatest_overlap();
test_segment_without_any_overlapping_turn_has_no_speaker();
test_zero_sample_rate_is_safe();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all transcript_assembly checks passed\n");
return 0;
}

View File

@@ -0,0 +1,283 @@
// Asserts the ABSENCES that capability_routing.cpp's five refusal reasons rest
// on, against the engine itself rather than against somebody's reading of it.
//
// Those reasons are prose making checkable claims about a pinned third-party
// checkout: "VoiceTaskKind has no codec entry", "no family advertises spk",
// "miocodec advertises only vc and s2s". Prose rots silently across an
// AUDIO_CPP_VERSION bump, and it rots in the worst possible place, since a
// refusal that states a false fact is worse than a bare UNIMPLEMENTED: it will
// be believed. One of the four claims this backend was planned against was
// already false when it was written ("streaming exists for tts and asr only" is
// contradicted by silero_vad). This test is what turns the next such change
// from a silent lie on the wire into a build failure.
//
// Pinned at audio.cpp e800d435d130dc776baf6f3e6129bb62b1495c89. What follows
// was true of that commit; a bump is exactly when it needs to be re-run.
//
// It links engine_runtime and queries make_default_registry(), touching only
// include/engine/framework/**, like every other unit in this backend. It loads
// no model and reads no file: advertise_loaders() is the path-free catalog
// upstream publishes for --list-loaders.
//
// FOUR CAVEATS, so nobody reads more into a green run than it earns:
//
// 1. It cannot assert the AudioToAudioStream reason. That one contrasts
// LocalAI's OpenAI-Realtime contract (conversation, system prompt, tool
// loop) with what audio.cpp's s2s families actually do, and semantics are
// not a queryable property. What IS asserted is the enumerable half: that
// s2s is advertised by exactly miocodec and vevo2, which is the clause the
// message names by hand.
// 2. It queries the LOADER catalog, while grpc-server.cpp reads the LOADED
// model's own capabilities(). The two agree today, cross-checked on the
// wire: this test asserts miocodec advertises {vc/offline, s2s/offline},
// and a live LoadModel of miocodec-q8_0.gguf reports exactly
// "vc/offline, s2s/offline". A family whose loaded capabilities diverged
// from its advertisement would slip past, but nothing loads without a
// model file and a ctest cannot depend on one.
// 3. Capabilities need not come from a loader at all.
// src/framework/model_spec/metadata.cpp's advertised_capabilities() builds
// a CapabilitySet from a spec's "capabilities"/"tasks"/"modes" keys, which
// is a second route by which a bump could falsify claim 3. Today no
// shipped model_specs/*.json carries a top-level "tasks" key and no loader
// calls that function, so the route is dead. It is covered anyway to the
// extent that a loader adopting it would surface through advertise_loaders
// like any other capability, which is why every assertion below queries
// advertised capabilities rather than loader source.
// 4. An absence test passes trivially when the query is broken, so
// test_catalog_is_populated below is a POSITIVE control and is not
// optional. It proves the registry is non-empty and that the query does
// find a capability it should, before any absence is believed.
#include "engine/framework/runtime/registry.h"
#include "engine/framework/runtime/session.h"
#include <algorithm>
#include <cstdio>
#include <set>
#include <stdexcept>
#include <string>
#include <vector>
namespace {
int failures = 0;
void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
using engine::runtime::LoaderAdvertisement;
using engine::runtime::RunMode;
using engine::runtime::VoiceTaskKind;
// Built once. make_default_registry constructs every loader, which is the
// expensive part, and none of the assertions mutates it.
const std::vector<LoaderAdvertisement> &catalog() {
static const std::vector<LoaderAdvertisement> kCatalog =
engine::runtime::make_default_registry().advertise_loaders();
return kCatalog;
}
bool advertises(const LoaderAdvertisement &loader, VoiceTaskKind task,
RunMode mode) {
for (const auto &capability : loader.capabilities.supported_tasks) {
if (capability.task != task) {
continue;
}
return std::find(capability.modes.begin(), capability.modes.end(), mode) !=
capability.modes.end();
}
return false;
}
bool advertises_any_mode(const LoaderAdvertisement &loader, VoiceTaskKind task) {
return advertises(loader, task, RunMode::Offline) ||
advertises(loader, task, RunMode::Streaming);
}
std::string describe(const LoaderAdvertisement &loader) {
std::string out;
for (const auto &capability : loader.capabilities.supported_tasks) {
for (const RunMode mode : capability.modes) {
if (!out.empty()) {
out += ", ";
}
out += engine::runtime::to_string(capability.task);
out += "/";
out += engine::runtime::to_string(mode);
}
}
return out.empty() ? "nothing" : out;
}
// THE POSITIVE CONTROL. Every other test here asserts an absence, and an
// absence is what a broken query returns for everything. If make_default_registry
// ever returns an empty registry, or advertise_loaders stops populating modes,
// this is the test that fails instead of the suite going quietly green while
// asserting nothing.
void test_catalog_is_populated() {
check(catalog().size() >= 20,
"the default registry advertises a plausible number of families (" +
std::to_string(catalog().size()) + ")");
bool found_streaming_asr = false;
for (const auto &loader : catalog()) {
if (advertises(loader, VoiceTaskKind::Asr, RunMode::Streaming)) {
found_streaming_asr = true;
break;
}
}
check(found_streaming_asr,
"the query finds a capability that IS advertised (asr/streaming)");
}
// The AudioEncode and AudioDecode premise. Not "no family does codec" but the
// stronger "the task kind does not exist", which is what makes those two RPCs
// unroutable rather than merely unserved.
//
// Note what is NOT asserted: model_spec/schema.cpp's task whitelist DOES accept
// the string "codec", so a spec declaring it validates and then fails here. That
// is a hole in upstream's own validation, and it is deliberately kept off the
// wire; the refusal rests on this parser, which is the thing routing would have
// to go through.
void test_no_codec_task_kind() {
bool threw = false;
try {
(void)engine::runtime::parse_voice_task_kind("codec");
} catch (const std::exception &) {
threw = true;
}
check(threw, "no codec task kind: parse_voice_task_kind(\"codec\") throws");
// The control for the line above: a real name must NOT throw, or the test
// would pass against a parser that rejected everything.
bool tts_threw = false;
try {
(void)engine::runtime::parse_voice_task_kind("tts");
} catch (const std::exception &) {
tts_threw = true;
}
check(!tts_threw, "parse_voice_task_kind accepts a real name (\"tts\")");
}
// The VoiceEmbed premise. SpeakerRecognition IS in the enum, and TitaNet and
// ECAPA-TDNN exist as internal conditioning encoders; what is missing is any
// registered family advertising the task, which is what the message says.
void test_no_family_advertises_speaker_recognition() {
std::string offenders;
for (const auto &loader : catalog()) {
if (advertises_any_mode(loader, VoiceTaskKind::SpeakerRecognition)) {
if (!offenders.empty()) {
offenders += ", ";
}
offenders += loader.family;
}
}
check(offenders.empty(),
"no family advertises spk" +
(offenders.empty() ? std::string() : " (found: " + offenders + ")"));
}
// The AudioTransformStream premise, and the one that has to be scoped exactly.
// The broad claim "upstream streams tts and asr only" is FALSE: silero_vad
// advertises vad/streaming. The claim that holds is that none of the four tasks
// AudioTransform routes to is advertised streaming by anybody.
//
// NEGATIVE CONTROL: add VoiceTaskKind::Tts to kTransformTasks and this test must
// fail, naming the tts families. A run where that edit stays green means the
// query is broken and every absence above is worthless.
void test_no_streaming_for_the_transform_tasks() {
const VoiceTaskKind kTransformTasks[] = {
VoiceTaskKind::SourceSeparation,
VoiceTaskKind::VoiceConversion,
VoiceTaskKind::Svc,
VoiceTaskKind::SpeechToSpeech,
};
std::string offenders;
for (const VoiceTaskKind task : kTransformTasks) {
for (const auto &loader : catalog()) {
if (advertises(loader, task, RunMode::Streaming)) {
if (!offenders.empty()) {
offenders += ", ";
}
offenders += loader.family;
offenders += "/";
offenders += engine::runtime::to_string(task);
}
}
}
check(offenders.empty(),
"no streaming for sep/vc/svc/s2s" +
(offenders.empty() ? std::string() : " (found: " + offenders + ")"));
}
// The clause the AudioEncode message names by hand: miocodec carries a Codec tag
// in upstream's README, and its loader advertises only vc and s2s. Asserted
// exactly, not as a subset, so a bump that ADDS a codec capability to miocodec
// fails here rather than leaving the message stale.
void test_miocodec_advertises_exactly_vc_and_s2s() {
const LoaderAdvertisement *miocodec = nullptr;
for (const auto &loader : catalog()) {
if (loader.family == "miocodec") {
miocodec = &loader;
break;
}
}
if (miocodec == nullptr) {
check(false, "miocodec is a registered family");
return;
}
check(describe(*miocodec) == "vc/offline, s2s/offline",
"miocodec advertises exactly vc/offline, s2s/offline (got: " +
describe(*miocodec) + ")");
}
// The "declared only by" clause in the AudioToAudioStream message. Exact set,
// for the same reason as above: a third s2s family would make the message stale
// without making it obviously wrong.
void test_speech_to_speech_is_exactly_miocodec_and_vevo2() {
std::set<std::string> families;
for (const auto &loader : catalog()) {
if (advertises_any_mode(loader, VoiceTaskKind::SpeechToSpeech)) {
families.insert(loader.family);
}
}
const std::set<std::string> expected = {"miocodec", "vevo2"};
std::string found;
for (const auto &family : families) {
if (!found.empty()) {
found += ", ";
}
found += family;
}
check(families == expected,
"s2s is advertised by exactly miocodec and vevo2 (got: " +
(found.empty() ? "nothing" : found) + ")");
}
} // namespace
int main() {
test_catalog_is_populated();
test_no_codec_task_kind();
test_no_family_advertises_speaker_recognition();
test_no_streaming_for_the_transform_tasks();
test_miocodec_advertises_exactly_vc_and_s2s();
test_speech_to_speech_is_exactly_miocodec_and_vevo2();
if (failures) {
fprintf(stderr,
"%d upstream absence check(s) failed. A refusal message in "
"capability_routing.cpp now states something that is not true "
"of the pinned audio.cpp; fix the message, not this test.\n",
failures);
return 1;
}
fprintf(stderr, "all upstream absence checks passed\n");
return 0;
}

View File

@@ -0,0 +1,63 @@
#include "wav_header.h"
#include <cstdint>
#include <limits>
namespace audiocpp_backend {
namespace {
void append_u32(std::string &out, std::uint32_t value) {
out.push_back(static_cast<char>(value & 0xFF));
out.push_back(static_cast<char>((value >> 8) & 0xFF));
out.push_back(static_cast<char>((value >> 16) & 0xFF));
out.push_back(static_cast<char>((value >> 24) & 0xFF));
}
void append_u16(std::string &out, std::uint16_t value) {
out.push_back(static_cast<char>(value & 0xFF));
out.push_back(static_cast<char>((value >> 8) & 0xFF));
}
// The unknown-length sentinel, in both the RIFF and the data chunk size.
constexpr std::uint32_t kStreamingSize = 0xFFFFFFFFu;
constexpr std::uint16_t kBitsPerSample = 16;
constexpr std::uint32_t kPcmFmtChunkSize = 16;
constexpr std::uint16_t kFormatTagPcm = 1;
constexpr int kMaxChannels = 65535;
} // namespace
std::string streaming_wav_header(int sample_rate, int channels) {
int clamped_channels = channels > 0 ? channels : 1;
if (clamped_channels > kMaxChannels) {
clamped_channels = kMaxChannels;
}
const auto channel_count = static_cast<std::uint16_t>(clamped_channels);
const auto rate = static_cast<std::uint32_t>(sample_rate > 0 ? sample_rate : 0);
const auto block_align =
static_cast<std::uint16_t>(channel_count * (kBitsPerSample / 8));
// uint32 arithmetic on purpose: 384 kHz by 8 channels is 6.1 MB/s, which
// does not fit the uint16 block align it is derived from.
const std::uint32_t byte_rate = rate * static_cast<std::uint32_t>(block_align);
static_assert(std::numeric_limits<std::uint32_t>::max() >= 0xFFFFFFFFu,
"the streaming sentinel must be representable");
std::string header;
header.reserve(44);
header += "RIFF";
append_u32(header, kStreamingSize); // unknown total length
header += "WAVE";
header += "fmt ";
append_u32(header, kPcmFmtChunkSize);
append_u16(header, kFormatTagPcm);
append_u16(header, channel_count);
append_u32(header, rate);
append_u32(header, byte_rate);
append_u16(header, block_align);
append_u16(header, kBitsPerSample);
header += "data";
append_u32(header, kStreamingSize); // unknown payload length
return header;
}
} // namespace audiocpp_backend

View File

@@ -0,0 +1,39 @@
#pragma once
// Builds the 44 byte canonical WAV header that precedes a streamed PCM body.
// Standard library only.
//
// TTSStream chunks travel in Reply.audio and the FIRST chunk must be this
// header, or an HTTP client has no format to decode the PCM with. Because the
// total length is unknown while the model is still generating, both size fields
// carry 0xFFFFFFFF; that is the convention backend/go/vibevoice-cpp established
// (govibevoicecpp.go, TTSStream) and the one core/backend/tts.go writes when it
// synthesises a header itself.
//
// WHY THIS BACKEND SENDS THE HEADER RATHER THAN LETTING GO DO IT.
// core/backend/tts.go's ModelTTSStream will emit a header of its own, but only
// when the FIRST Reply carries a non-empty `message` field holding a JSON blob
// with a sample_rate. This backend sends audio and never sets `message`, so
// that branch never fires and there is exactly one header on the wire: this
// one. Do not start setting Reply.message on this RPC without deleting the
// header below, or every stream gains a second header 44 bytes into the PCM.
//
// The header this produces is byte-identical to the one pkg/audio.WAVHeader
// serialises, which is what core/http/endpoints/openai/realtime_model.go
// assumes when it reads the sample rate out of byte offset 24 of the first
// callback.
#include <string>
namespace audiocpp_backend {
// 16-bit PCM, little endian, interleaved.
//
// `channels` is clamped to at least 1 and at most 65535, so a garbage channel
// count can never write a zero block align, which is what a reader divides the
// data size by. `sample_rate` is clamped to at least 0 rather than wrapped:
// a zero rate is visibly wrong to whoever reads the header, whereas the
// 4294967295 an unsigned conversion of -1 would write looks like a real field.
std::string streaming_wav_header(int sample_rate, int channels);
} // namespace audiocpp_backend

View File

@@ -0,0 +1,163 @@
// Unit tests for wav_header. Standard library only. The harness compiles this
// as a single translation unit, so the implementation is included directly.
#include "wav_header.cpp"
#include <cstdint>
#include <cstdio>
#include <string>
static int failures = 0;
static void check(bool ok, const std::string &name) {
if (!ok) {
failures++;
fprintf(stderr, "FAIL: %s\n", name.c_str());
} else {
fprintf(stderr, "ok: %s\n", name.c_str());
}
}
using audiocpp_backend::streaming_wav_header;
static std::uint32_t read_u32(const std::string &data, size_t offset) {
return static_cast<std::uint32_t>(static_cast<unsigned char>(data[offset])) |
(static_cast<std::uint32_t>(static_cast<unsigned char>(data[offset + 1])) << 8) |
(static_cast<std::uint32_t>(static_cast<unsigned char>(data[offset + 2])) << 16) |
(static_cast<std::uint32_t>(static_cast<unsigned char>(data[offset + 3])) << 24);
}
static std::uint16_t read_u16(const std::string &data, size_t offset) {
return static_cast<std::uint16_t>(
static_cast<std::uint16_t>(static_cast<unsigned char>(data[offset])) |
(static_cast<std::uint16_t>(static_cast<unsigned char>(data[offset + 1])) << 8));
}
static std::string to_hex(const std::string &data) {
static const char *digits = "0123456789abcdef";
std::string out;
out.reserve(data.size() * 2);
for (const char byte : data) {
const auto value = static_cast<unsigned char>(byte);
out.push_back(digits[value >> 4]);
out.push_back(digits[value & 0x0F]);
}
return out;
}
static void test_header_layout() {
const std::string header = streaming_wav_header(24000, 1);
check(header.size() == 44, "canonical 44 byte header");
check(header.compare(0, 4, "RIFF") == 0, "RIFF magic");
check(header.compare(8, 4, "WAVE") == 0, "WAVE magic");
check(header.compare(12, 4, "fmt ") == 0, "fmt chunk id");
check(header.compare(36, 4, "data") == 0, "data chunk id");
check(read_u32(header, 16) == 16, "PCM fmt chunk is 16 bytes");
check(read_u16(header, 20) == 1, "format tag 1 is PCM");
check(read_u16(header, 22) == 1, "mono channel count");
check(read_u32(header, 24) == 24000, "sample rate");
// byte rate = rate * channels * bytes per sample
check(read_u32(header, 28) == 24000 * 1 * 2, "byte rate");
check(read_u16(header, 32) == 2, "block align for mono 16 bit");
check(read_u16(header, 34) == 16, "16 bits per sample");
}
// The field-wise checks above can all pass while the fields sit in the wrong
// ORDER, since several of them hold the same value. This pins the whole 44 byte
// string against a literal transcribed from the layout in pkg/audio/audio.go,
// which is the struct binary.Write serializes for every Go LocalAI backend.
// Independent of the implementation: it was written out by hand rather than
// captured from a run.
static void test_exact_bytes_match_the_go_layout() {
const std::string expected =
"52494646" // "RIFF"
"ffffffff" // chunk size: streaming sentinel
"57415645" // "WAVE"
"666d7420" // "fmt "
"10000000" // subchunk1 size 16
"0100" // audio format 1 (PCM)
"0100" // channels 1
"c05d0000" // sample rate 24000
"80bb0000" // byte rate 48000
"0200" // block align 2
"1000" // bits per sample 16
"64617461" // "data"
"ffffffff"; // subchunk2 size: streaming sentinel
check(to_hex(streaming_wav_header(24000, 1)) == expected,
"byte for byte match with the canonical mono 24 kHz header");
}
// The whole point: a streaming header cannot know the final length, so both
// size fields are the sentinel. A client that sees a real size stops early.
static void test_streaming_sentinels() {
const std::string header = streaming_wav_header(16000, 1);
check(read_u32(header, 4) == 0xFFFFFFFFu, "RIFF chunk size is the sentinel");
check(read_u32(header, 40) == 0xFFFFFFFFu, "data chunk size is the sentinel");
}
static void test_stereo() {
const std::string header = streaming_wav_header(44100, 2);
check(read_u16(header, 22) == 2, "stereo channel count");
check(read_u32(header, 28) == 44100 * 2 * 2, "stereo byte rate");
check(read_u16(header, 32) == 4, "block align for stereo 16 bit");
check(header.size() == 44, "stereo header is still 44 bytes");
}
static void test_degenerate_inputs() {
// A zero or negative channel count must not produce a header that divides
// by zero downstream; clamp to mono.
const std::string zero_channels = streaming_wav_header(16000, 0);
check(read_u16(zero_channels, 22) == 1, "zero channels clamps to mono");
check(read_u16(zero_channels, 32) == 2, "zero channels still block aligns as mono");
check(read_u32(zero_channels, 28) == 32000, "zero channels byte rate is the mono one");
const std::string negative_channels = streaming_wav_header(16000, -3);
check(read_u16(negative_channels, 22) == 1, "negative channels clamps to mono");
// A channel count past the field's range must saturate rather than wrap:
// 65536 truncated to uint16 is 0, and a zero channel count writes a zero
// block align, which is what a reader divides the data size by.
const std::string too_many = streaming_wav_header(16000, 65536);
check(read_u16(too_many, 22) == 65535, "an out of range channel count saturates");
check(read_u16(too_many, 32) != 0, "an out of range channel count never writes a zero block align");
// A non-positive rate is written as zero rather than wrapping through the
// unsigned conversion: 0 is visibly wrong to whoever reads the header,
// 4294967295 looks like a plausible field nobody checks.
const std::string zero_rate = streaming_wav_header(0, 1);
check(read_u32(zero_rate, 24) == 0, "zero sample rate stays zero");
check(read_u32(zero_rate, 28) == 0, "zero sample rate yields a zero byte rate");
const std::string negative_rate = streaming_wav_header(-48000, 1);
check(read_u32(negative_rate, 24) == 0, "negative sample rate is clamped to zero");
check(read_u32(negative_rate, 28) == 0, "negative sample rate yields a zero byte rate");
// The sentinels are unconditional. A degenerate rate must not turn the
// stream into one a client thinks it can measure.
check(read_u32(zero_rate, 4) == 0xFFFFFFFFu, "degenerate input keeps the RIFF sentinel");
check(read_u32(zero_rate, 40) == 0xFFFFFFFFu, "degenerate input keeps the data sentinel");
}
// Large but legal: 384 kHz 8 channel would overflow a 16 bit byte rate and
// must not overflow the 32 bit one either.
static void test_large_but_legal() {
const std::string header = streaming_wav_header(384000, 8);
check(read_u32(header, 28) == 384000u * 8u * 2u, "high rate multichannel byte rate");
check(read_u16(header, 32) == 16, "high channel count block align");
}
int main() {
test_header_layout();
test_exact_bytes_match_the_go_layout();
test_streaming_sentinels();
test_stereo();
test_degenerate_inputs();
test_large_but_legal();
if (failures) {
fprintf(stderr, "%d check(s) failed\n", failures);
return 1;
}
fprintf(stderr, "all wav_header checks passed\n");
return 0;
}

View File

@@ -1,7 +1,7 @@
# Pinned to the HEAD of the `prism` branch on https://github.com/PrismML-Eng/llama.cpp.
# Auto-bumped nightly by .github/workflows/bump_deps.yaml.
BONSAI_VERSION?=7529fdaaf99ffdc5ca71ace9c7409a56b27ad92f
BONSAI_VERSION?=9ca265a57f85f2117942490f421f64a226dd9847
LLAMA_REPO?=https://github.com/PrismML-Eng/llama.cpp
CMAKE_ARGS?=
@@ -41,6 +41,7 @@ define bonsai-build
# and are applied by apply-patches.sh below.
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build/patches
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build purge
bash $(LLAMA_CPP_DIR)/disable-score-task.sh $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build/grpc-server.cpp
$(info $(GREEN)I bonsai build info:$(1)$(RESET))
LLAMA_REPO=$(LLAMA_REPO) LLAMA_VERSION=$(BONSAI_VERSION) \
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build llama.cpp
@@ -77,6 +78,7 @@ bonsai-cpu-all:
# and are applied by apply-patches.sh below.
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/patches
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build purge
bash $(LLAMA_CPP_DIR)/disable-score-task.sh $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/grpc-server.cpp
$(info $(GREEN)I bonsai build info:cpu-all-variants$(RESET))
LLAMA_REPO=$(LLAMA_REPO) LLAMA_VERSION=$(BONSAI_VERSION) \
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build llama.cpp

View File

@@ -24,34 +24,7 @@ if [ -d "$CURDIR/ggml-shared-libs" ]; then
fi
# Detect architecture and copy appropriate libraries
if [ -f "/lib64/ld-linux-x86-64.so.2" ]; then
# x86_64 architecture
echo "Detected x86_64 architecture, copying x86_64 libraries..."
cp -arfLv /lib64/ld-linux-x86-64.so.2 $CURDIR/package/lib/ld.so
cp -arfLv /lib/x86_64-linux-gnu/libc.so.6 $CURDIR/package/lib/libc.so.6
cp -arfLv /lib/x86_64-linux-gnu/libgcc_s.so.1 $CURDIR/package/lib/libgcc_s.so.1
cp -arfLv /lib/x86_64-linux-gnu/libstdc++.so.6 $CURDIR/package/lib/libstdc++.so.6
cp -arfLv /lib/x86_64-linux-gnu/libm.so.6 $CURDIR/package/lib/libm.so.6
cp -arfLv /lib/x86_64-linux-gnu/libgomp.so.1 $CURDIR/package/lib/libgomp.so.1
cp -arfLv /lib/x86_64-linux-gnu/libdl.so.2 $CURDIR/package/lib/libdl.so.2
cp -arfLv /lib/x86_64-linux-gnu/librt.so.1 $CURDIR/package/lib/librt.so.1
cp -arfLv /lib/x86_64-linux-gnu/libpthread.so.0 $CURDIR/package/lib/libpthread.so.0
elif [ -f "/lib/ld-linux-aarch64.so.1" ]; then
# ARM64 architecture
echo "Detected ARM64 architecture, copying ARM64 libraries..."
cp -arfLv /lib/ld-linux-aarch64.so.1 $CURDIR/package/lib/ld.so
cp -arfLv /lib/aarch64-linux-gnu/libc.so.6 $CURDIR/package/lib/libc.so.6
cp -arfLv /lib/aarch64-linux-gnu/libgcc_s.so.1 $CURDIR/package/lib/libgcc_s.so.1
cp -arfLv /lib/aarch64-linux-gnu/libstdc++.so.6 $CURDIR/package/lib/libstdc++.so.6
cp -arfLv /lib/aarch64-linux-gnu/libm.so.6 $CURDIR/package/lib/libm.so.6
cp -arfLv /lib/aarch64-linux-gnu/libgomp.so.1 $CURDIR/package/lib/libgomp.so.1
cp -arfLv /lib/aarch64-linux-gnu/libdl.so.2 $CURDIR/package/lib/libdl.so.2
cp -arfLv /lib/aarch64-linux-gnu/librt.so.1 $CURDIR/package/lib/librt.so.1
cp -arfLv /lib/aarch64-linux-gnu/libpthread.so.0 $CURDIR/package/lib/libpthread.so.0
else
echo "Error: Could not detect architecture"
exit 1
fi
source "$CURDIR/../../../scripts/build/package-system-libs.sh" "$CURDIR/package/lib" ""
# Package GPU libraries based on BUILD_TYPE
GPU_LIB_SCRIPT="${REPO_ROOT}/scripts/build/package-gpu-libs.sh"

View File

@@ -40,6 +40,27 @@ else
if [ -d "$CURDIR/lib/hipblaslt/library" ]; then
export HIPBLASLT_TENSILE_LIBPATH="$CURDIR"/lib/hipblaslt/library
fi
# Backends built for Intel GPUs carry a copy of the Intel graphics driver,
# and libze_loader is only there in those builds. Level Zero looks for a
# driver on its own, so point it at the copy that came with this backend: it
# was built against the same C library, while the machine's own driver may
# not have been, and loading that one can crash on start.
#
# Anything the user set is left alone, so a machine with a graphics card
# newer than the driver carried here can still be told to use its own.
# Nothing is said about OpenCL: no OpenCL driver is carried, so anything we
# set there would leave OpenCL worse off than the machine's own setup.
if [ -e "$CURDIR/lib/libze_loader.so.1" ]; then
if [ -e "$CURDIR/lib/libze_intel_gpu.so.1" ] && [ -z "${ZE_ENABLE_ALT_DRIVERS:-}" ]; then
export ZE_ENABLE_ALT_DRIVERS="$CURDIR"/lib/libze_intel_gpu.so.1
fi
# Ask the driver how much graphics memory is free. Without this, the
# backend reads zero on an integrated graphics chip, because such a chip
# shares the system memory instead of having its own.
if [ -z "${ZES_ENABLE_SYSMAN:-}" ]; then
export ZES_ENABLE_SYSMAN=1
fi
fi
fi
# If there is a lib/ld.so, use it

View File

@@ -69,7 +69,15 @@ target_include_directories(hw_grpc_proto PUBLIC ${CMAKE_CURRENT_BINARY_DIR})
set(DS4_OBJS "${DS4_DIR}/ds4.o")
if(DS4_GPU STREQUAL "cuda")
list(APPEND DS4_OBJS "${DS4_DIR}/ds4_cuda.o")
list(APPEND DS4_OBJS
"${DS4_DIR}/ds4_cuda.o"
"${DS4_DIR}/cuda/mmq/ds4_ggml_stubs.o"
"${DS4_DIR}/cuda/mmq/ds4_mmq.o"
"${DS4_DIR}/cuda/mmq/ds4_mmq_d2r.o"
"${DS4_DIR}/cuda/mmq/quantize.o"
"${DS4_DIR}/cuda/mmq/mmid.o"
"${DS4_DIR}/cuda/mmq/mmvq.o"
"${DS4_DIR}/cuda/mmq/ds4_repack.o")
elseif(DS4_GPU STREQUAL "metal")
list(APPEND DS4_OBJS "${DS4_DIR}/ds4_metal.o")
elseif(DS4_GPU STREQUAL "cpu")

View File

@@ -1,10 +1,10 @@
# ds4 backend Makefile.
#
# Upstream pin lives below as DS4_VERSION?=efdadd41e20134af4f3381e1ed90e96fe4faef6f
# Upstream pin lives below as DS4_VERSION?=b7e9f0091139999b6c070a57590c447c5741da5c
# (.github/bump_deps.sh) can find and update it - matches the
# llama-cpp / ik-llama-cpp / turboquant convention.
DS4_VERSION?=efdadd41e20134af4f3381e1ed90e96fe4faef6f
DS4_VERSION?=b7e9f0091139999b6c070a57590c447c5741da5c
DS4_REPO?=https://github.com/antirez/ds4
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))
@@ -23,7 +23,9 @@ CMAKE_ARGS ?= -DCMAKE_BUILD_TYPE=Release
# are shared by every GPU mode, so append them unconditionally below.
ifeq ($(BUILD_TYPE),cublas)
CMAKE_ARGS += -DDS4_GPU=cuda
DS4_OBJ_TARGET := ds4.o ds4_cuda.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
DS4_OBJ_TARGET := ds4.o ds4_cuda.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o \
cuda/mmq/ds4_ggml_stubs.o cuda/mmq/ds4_mmq.o cuda/mmq/ds4_mmq_d2r.o \
cuda/mmq/quantize.o cuda/mmq/mmid.o cuda/mmq/mmvq.o cuda/mmq/ds4_repack.o
else ifeq ($(UNAME_S),Darwin)
CMAKE_ARGS += -DDS4_GPU=metal
DS4_OBJ_TARGET := ds4.o ds4_metal.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
@@ -55,7 +57,7 @@ ds4:
# the right per-platform compile flags (Objective-C/Metal on Darwin, nvcc on Linux+CUDA).
ds4/ds4.o: ds4
ifeq ($(BUILD_TYPE),cublas)
+$(MAKE) -C ds4 ds4.o ds4_cuda.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
+$(MAKE) -C ds4 $(DS4_OBJ_TARGET)
else ifeq ($(UNAME_S),Darwin)
+$(MAKE) -C ds4 ds4.o ds4_metal.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
else

View File

@@ -17,13 +17,7 @@ if [ "$UNAME_S" = "Darwin" ]; then
exit 0
fi
if [ -f "/lib64/ld-linux-x86-64.so.2" ]; then
cp -arfLv /lib64/ld-linux-x86-64.so.2 "$PACKAGE_DIR/lib/ld.so"
elif [ -f "/lib/ld-linux-aarch64.so.1" ]; then
cp -arfLv /lib/ld-linux-aarch64.so.1 "$PACKAGE_DIR/lib/ld.so"
else
echo "package.sh: unknown architecture" >&2; exit 1
fi
source "$CURDIR/../../../scripts/build/package-system-libs.sh" "$CURDIR/package/lib" ""
# Bundle the complete dependency closure for both executables. In particular,
# grpc-server links the distro gRPC/protobuf/absl stack; copying only the core

View File

@@ -1,5 +1,5 @@
IK_LLAMA_VERSION?=e5357286c0d433cd4384e82ed7e2b6d655f57087
IK_LLAMA_VERSION?=60389410a1ff01f9d37dcc6261db33b3183bdea2
LLAMA_REPO?=https://github.com/ikawrakow/ik_llama.cpp
CMAKE_ARGS?=

View File

@@ -15,34 +15,7 @@ cp -avrf $CURDIR/ik-llama-cpp-* $CURDIR/package/
cp -rfv $CURDIR/run.sh $CURDIR/package/
# Detect architecture and copy appropriate libraries
if [ -f "/lib64/ld-linux-x86-64.so.2" ]; then
# x86_64 architecture
echo "Detected x86_64 architecture, copying x86_64 libraries..."
cp -arfLv /lib64/ld-linux-x86-64.so.2 $CURDIR/package/lib/ld.so
cp -arfLv /lib/x86_64-linux-gnu/libc.so.6 $CURDIR/package/lib/libc.so.6
cp -arfLv /lib/x86_64-linux-gnu/libgcc_s.so.1 $CURDIR/package/lib/libgcc_s.so.1
cp -arfLv /lib/x86_64-linux-gnu/libstdc++.so.6 $CURDIR/package/lib/libstdc++.so.6
cp -arfLv /lib/x86_64-linux-gnu/libm.so.6 $CURDIR/package/lib/libm.so.6
cp -arfLv /lib/x86_64-linux-gnu/libgomp.so.1 $CURDIR/package/lib/libgomp.so.1
cp -arfLv /lib/x86_64-linux-gnu/libdl.so.2 $CURDIR/package/lib/libdl.so.2
cp -arfLv /lib/x86_64-linux-gnu/librt.so.1 $CURDIR/package/lib/librt.so.1
cp -arfLv /lib/x86_64-linux-gnu/libpthread.so.0 $CURDIR/package/lib/libpthread.so.0
elif [ -f "/lib/ld-linux-aarch64.so.1" ]; then
# ARM64 architecture
echo "Detected ARM64 architecture, copying ARM64 libraries..."
cp -arfLv /lib/ld-linux-aarch64.so.1 $CURDIR/package/lib/ld.so
cp -arfLv /lib/aarch64-linux-gnu/libc.so.6 $CURDIR/package/lib/libc.so.6
cp -arfLv /lib/aarch64-linux-gnu/libgcc_s.so.1 $CURDIR/package/lib/libgcc_s.so.1
cp -arfLv /lib/aarch64-linux-gnu/libstdc++.so.6 $CURDIR/package/lib/libstdc++.so.6
cp -arfLv /lib/aarch64-linux-gnu/libm.so.6 $CURDIR/package/lib/libm.so.6
cp -arfLv /lib/aarch64-linux-gnu/libgomp.so.1 $CURDIR/package/lib/libgomp.so.1
cp -arfLv /lib/aarch64-linux-gnu/libdl.so.2 $CURDIR/package/lib/libdl.so.2
cp -arfLv /lib/aarch64-linux-gnu/librt.so.1 $CURDIR/package/lib/librt.so.1
cp -arfLv /lib/aarch64-linux-gnu/libpthread.so.0 $CURDIR/package/lib/libpthread.so.0
else
echo "Error: Could not detect architecture"
exit 1
fi
source "$CURDIR/../../../scripts/build/package-system-libs.sh" "$CURDIR/package/lib" ""
# Package GPU libraries based on BUILD_TYPE
# The GPU library packaging script will detect BUILD_TYPE and copy appropriate GPU libraries

View File

@@ -110,4 +110,9 @@ if(LLAMA_GRPC_BUILD_TESTS)
target_link_libraries(parent_watch_test PRIVATE Threads::Threads)
target_compile_features(parent_watch_test PRIVATE cxx_std_17)
add_test(NAME parent_watch_test COMMAND parent_watch_test)
add_executable(passthrough_options_test passthrough_options_test.cpp passthrough_options.h)
target_include_directories(passthrough_options_test PRIVATE ${CMAKE_CURRENT_SOURCE_DIR})
target_compile_features(passthrough_options_test PRIVATE cxx_std_17)
add_test(NAME passthrough_options_test COMMAND passthrough_options_test)
endif()

View File

@@ -1,5 +1,5 @@
LLAMA_VERSION?=571d0d540df04f25298d0e159e520d9fc62ed121
LLAMA_VERSION?=221f0f6356efe2260023208365705ec5d5a7c8f5
LLAMA_REPO?=https://github.com/ggerganov/llama.cpp
CMAKE_ARGS?=

View File

@@ -0,0 +1,43 @@
#!/bin/bash
# Mark a copied gRPC server as targeting a llama.cpp fork that does not carry
# LocalAI's slot-based Score patches. The RPC remains present in the shared
# protobuf service, but responds with UNIMPLEMENTED instead of referencing
# server task types and common_params fields absent from those forks.
set -euo pipefail
if [[ $# -ne 1 ]]; then
echo "usage: $0 <grpc-server.cpp>" >&2
exit 2
fi
SRC=$1
if [[ ! -f "$SRC" ]]; then
echo "grpc-server.cpp not found at $SRC" >&2
exit 2
fi
if grep -q '^#define LOCALAI_LLAMA_CPP_NO_SCORE_TASK' "$SRC"; then
echo "==> $SRC already disables the LocalAI score task, skipping"
exit 0
fi
awk '
!done && /^#include/ {
print "#define LOCALAI_LLAMA_CPP_NO_SCORE_TASK 1"
print "// ^ injected by disable-score-task.sh for an unpatched llama.cpp fork"
print ""
done = 1
}
{ print }
END {
if (!done) {
print "disable-score-task.sh: no #include anchor found" > "/dev/stderr"
exit 1
}
}
' "$SRC" > "$SRC.tmp"
mv "$SRC.tmp" "$SRC"
echo "==> LocalAI score task disabled in $SRC"

View File

@@ -52,7 +52,9 @@
#include "common.h"
#include "arg.h"
#include "chat-auto-parser.h"
#include "llama_compat.h" // fork-skew switches, generated by prepare.sh
#include "message_content.h"
#include "passthrough_options.h"
#include <getopt.h>
#include <grpcpp/ext/proto_server_reflection_plugin.h>
#include <grpcpp/grpcpp.h>
@@ -151,40 +153,6 @@ static std::string base64_encode_bytes(const unsigned char* data, size_t len) {
bool loaded_model; // TODO: add a mutex for this, but happens only once loading the model
// Score bypasses the slot loop (see the comment on Score below) so it
// must not run concurrently with any slot-loop RPC. These counters
// are a defence-in-depth tripwire — ModelConfig.Validate already
// rejects llama-cpp configs that mix score with chat/completion/
// embeddings, so a healthy deployment never trips them. seq_cst is
// load-bearing for the increment-then-check pattern below.
static std::atomic<int> slot_loop_inflight{0};
static std::atomic<int> score_inflight{0};
// Increment-then-check, not check-then-increment: two simultaneous
// racers both observe the other's increment and both abort cleanly.
// Reversed, both could see zero and proceed.
struct conflict_guard {
std::atomic<int>& self;
conflict_guard(const char* rpc, std::atomic<int>& self_, std::atomic<int>& other, const char* other_name)
: self(self_) {
self.fetch_add(1, std::memory_order_seq_cst);
int o = other.load(std::memory_order_seq_cst);
if (o > 0) {
fprintf(stderr,
"FATAL: %s called with %s=%d. The llama-cpp backend cannot "
"service Score and slot-loop RPCs concurrently — Score "
"bypasses the slot loop and races the llama_context. Bind "
"Score-using features to a model dedicated to scoring "
"(known_usecases: [score] with no chat/completion/embeddings).\n",
rpc, other_name, o);
std::abort();
}
}
~conflict_guard() {
self.fetch_sub(1, std::memory_order_seq_cst);
}
};
static std::function<void(int)> shutdown_handler;
static std::atomic_flag is_terminating = ATOMIC_FLAG_INIT;
@@ -612,6 +580,15 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
// Raw upstream llama-server flags collected from any option entry that
// starts with '-'. Applied once after the loop via common_params_parse.
std::vector<std::string> extra_argv;
bool passthrough_main_gpu_layers = false;
bool passthrough_draft_gpu_layers = false;
// O_DIRECT intent from the `direct_io` option. Upstream folded
// use_mmap/use_mlock/use_direct_io into a single common_params::load_mode
// enum (ggml-org/llama.cpp#20834), so the three independent LocalAI settings
// can only be reduced to one value once all of them have been read, held
// here until the mmap/mlock fields arrive further down.
bool want_direct_io = false;
auto add_device_options = [&](const std::string & devices) {
const std::regex regex{ R"([,]+)" };
@@ -724,6 +701,22 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
// If conversion fails, keep default value (0)
}
}
#ifndef LOCALAI_LLAMA_CPP_NO_SCORE_TASK
} else if (!strcmp(optname, "n_rs_seq") || !strcmp(optname, "rs_seq")) {
// Recurrent-state rollback snapshots per sequence. Hybrid models
// (deltanet/conv layers) cannot rewind their state, so without
// snapshots any prompt-cache reuse that needs a rewind — e.g. a
// score task whose probe changed under a stable option-list
// prefix — falls back to a full re-prefill. Costs recurrent-state
// memory x (1 + N) per sequence; unsupported archs clamp to 0.
if (optval != NULL) {
try {
params.n_rs_seq = std::stoi(optval_str);
} catch (const std::exception& e) {
// If conversion fails, keep default value (0)
}
}
#endif
} else if (!strcmp(optname, "slot_prompt_similarity") || !strcmp(optname, "sps")) {
if (optval != NULL) {
try {
@@ -868,9 +861,9 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
// --- O_DIRECT model loading (upstream --direct-io) ---
} else if (!strcmp(optname, "direct_io") || !strcmp(optname, "use_direct_io")) {
if (optval_str == "true" || optval_str == "1" || optval_str == "yes" || optval_str == "on" || optval_str == "enabled") {
params.use_direct_io = true;
want_direct_io = true;
} else if (optval_str == "false" || optval_str == "0" || optval_str == "no" || optval_str == "off" || optval_str == "disabled") {
params.use_direct_io = false;
want_direct_io = false;
}
// --- embedding normalization (upstream --embd-normalize) ---
@@ -1196,6 +1189,17 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
flag.c_str());
} else {
extra_argv.push_back(flag);
passthrough_main_gpu_layers =
passthrough_main_gpu_layers ||
flag == "-ngl" ||
flag == "--gpu-layers" ||
flag == "--n-gpu-layers";
passthrough_draft_gpu_layers =
passthrough_draft_gpu_layers ||
flag == "--spec-draft-ngl" ||
flag == "-ngld" ||
flag == "--gpu-layers-draft" ||
flag == "--n-gpu-layers-draft";
// Preserve the whole value after the first ':' so embedded
// colons (e.g. host:port) survive strtok's truncation of optval.
auto colon = opt.find(':');
@@ -1278,8 +1282,28 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
lora_info.ptr = nullptr;
params.lora_adapters.push_back(std::move(lora_info));
}
params.use_mlock = request->mlock();
params.use_mmap = request->mmap();
// LocalAI keeps mmap, mlock and direct-I/O as three independent settings,
// while upstream now carries a single load mode. Fold them with the
// precedence the separate booleans used to give: direct I/O bypasses the
// page cache entirely, mlock implies mmap, and everything off is a plain
// buffered read. Forks that branched before ggml-org/llama.cpp#20834 still
// expose the booleans; prepare.sh probes the checkout and sets
// LOCALAI_LEGACY_LOAD_MODE in the generated llama_compat.h accordingly.
#if LOCALAI_LEGACY_LOAD_MODE
params.use_mlock = request->mlock();
params.use_mmap = request->mmap();
params.use_direct_io = want_direct_io;
#else
if (want_direct_io) {
params.load_mode = LLAMA_LOAD_MODE_DIRECT_IO;
} else if (request->mlock()) {
params.load_mode = LLAMA_LOAD_MODE_MLOCK;
} else if (request->mmap()) {
params.load_mode = LLAMA_LOAD_MODE_MMAP;
} else {
params.load_mode = LLAMA_LOAD_MODE_NONE;
}
#endif
if (request->flashattention() == "on" || request->flashattention() == "enabled") {
params.flash_attn_type = LLAMA_FLASH_ATTN_TYPE_ENABLED;
@@ -1339,6 +1363,14 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
// (n_parallel -> -1, use_color). Snapshot n_parallel so an unrelated
// passthrough flag can't silently clobber LocalAI's resolved value.
const int saved_n_parallel = params.n_parallel;
// Newer upstream parsers assert that these fields still contain their
// negative initialization sentinels. LocalAI resolves them from the
// model request before applying passthrough options, so stage the
// sentinels and restore the values unless a raw flag overrides them.
const auto saved_gpu_layers =
llama_grpc::prepare_passthrough_gpu_layers(
params.n_gpu_layers,
params.speculative.draft.n_gpu_layers);
std::vector<char *> argv;
std::string prog = "llama-server";
@@ -1362,8 +1394,25 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
if (params.n_parallel == -1) {
params.n_parallel = saved_n_parallel;
}
llama_grpc::restore_passthrough_gpu_layers(
params.n_gpu_layers,
params.speculative.draft.n_gpu_layers,
saved_gpu_layers,
passthrough_main_gpu_layers,
passthrough_draft_gpu_layers);
}
#ifndef LOCALAI_LLAMA_CPP_NO_SCORE_TASK
// Score-task suffix forking: reserve seq ids (and recurrent-state cells)
// beyond the slots so one scoring call decodes all candidate tails in a
// single batch (SERVER_TASK_TYPE_SCORE, patches/). Requires the unified
// KV cache — with per-sequence streams the extra ids would shrink every
// sequence's context to n_ctx / n_seq_max. Decided after both option
// passes so an explicit kv_unified:false wins and disables forking.
params.score_enabled = request->enablescore();
params.n_seq_score_forks = params.score_enabled && params.kv_unified ? SERVER_SCORE_FORK_SEQS : 0;
#endif
// Terminate/pad the override vectors only after BOTH the named-option loop
// and the generic passthrough (common_params_parse above) have pushed their
// real entries, so back() is the null sentinel the model loader asserts on.
@@ -1450,6 +1499,16 @@ public:
common_params params;
params_parse(ctx_server, request, params);
#ifndef LOCALAI_LLAMA_CPP_NO_SCORE_TASK
if (params.score_enabled && !params.kv_unified) {
const std::string error_msg =
"Score requires the unified KV cache; remove kv_unified:false or remove score from known_usecases";
result->set_message(error_msg);
result->set_success(false);
return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT, error_msg);
}
#endif
common_init();
// Ensure debug logs are enabled after common_init() sets up logging
common_log_set_verbosity_thold(params.verbosity);
@@ -1652,7 +1711,6 @@ public:
if (params_base.model.path.empty()) {
return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, "Model not loaded");
}
conflict_guard guard("PredictStream", slot_loop_inflight, score_inflight, "score_inflight");
json data = parse_options(true, request, params_base, ctx_server.get_llama_context());
@@ -2221,7 +2279,6 @@ public:
if (params_base.model.path.empty()) {
return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, "Model not loaded");
}
conflict_guard guard("Predict", slot_loop_inflight, score_inflight, "score_inflight");
json data = parse_options(true, request, params_base, ctx_server.get_llama_context());
data["stream"] = false;
@@ -2755,7 +2812,6 @@ public:
if (params_base.model.path.empty()) {
return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, "Model not loaded");
}
conflict_guard guard("Embedding", slot_loop_inflight, score_inflight, "score_inflight");
json body = parse_options(false, request, params_base, ctx_server.get_llama_context());
body["stream"] = false;
@@ -2865,7 +2921,6 @@ public:
return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT, "\"documents\" must be a non-empty string array");
}
conflict_guard guard("Rerank", slot_loop_inflight, score_inflight, "score_inflight");
// Create and queue the task
auto rd = ctx_server.get_response_reader();
@@ -2942,37 +2997,16 @@ public:
// Score returns the model's joint log-probability of each candidate
// continuation given a shared prompt.
//
// WHY bypass the slot/task queue: upstream server_context exposes
// get_llama_context as "main thread only" and the slot loop's
// update_slots() owns the context whenever a task is in flight.
// No public synchronization primitive is available — so Score is
// unsafe to call concurrently with active generation through this
// backend. In practice routing-classifier calls happen before the
// request is routed to a generation backend, so the model used
// for Score is typically idle. Concurrent Score calls are
// serialised by a local mutex; KV-cache state is isolated behind
// a dedicated sequence ID cleared between candidates.
//
// A patch to server-context.cpp that adds SERVER_TASK_TYPE_SCORE
// and routes scoring through the slot loop would be the correct
// long-term fix; tracked as a follow-up.
//
// Perf TODO (measured: ~450 ms warm for 3 candidates on Arch-
// Router-1.5B Q4_K_M + Intel SYCL): the current loop re-decodes
// `prompt + candidate` from scratch for every candidate, throwing
// away the prompt's KV cache between iterations. A smarter
// version would:
// 1. Decode just the prompt once into score_seq_id.
// 2. Snapshot/cp that sequence (llama_memory_seq_cp) into a
// per-candidate sequence id.
// 3. For each candidate, decode only its tokens onto the copy
// (continuing from the saved prompt state), read logits.
// 4. llama_memory_seq_rm the copy.
// Estimated speedup: 3-candidate calls 450 ms -> ~150-200 ms,
// 6-candidate calls 630 ms -> ~220 ms. Single source-file change,
// no proto / Go-side changes needed. Worth doing once routing is
// wired into the middleware and Score is on the hot path of every
// chat request.
// Scoring runs as a single SERVER_TASK_TYPE_SCORE task through the
// slot loop (added by patches/ on top of upstream server-context), so
// it is safe to interleave with generation on the same process and it
// reuses any KV prefix the slot already holds across turns. The task
// decodes the shared prefix (prompt + longest common candidate token
// prefix) once on the slot's sequence; every candidate's unique tail
// then rides its own forked sequence and all tails are decoded
// together in one batch, so a warm scoring call costs roughly one
// forward pass over the new prompt tokens plus one batched pass over
// the candidate tails.
grpc::Status Score(ServerContext* context, const backend::ScoreRequest* request, backend::ScoreResponse* response) override {
auto auth = checkAuth(context);
if (!auth.ok()) return auth;
@@ -2981,40 +3015,21 @@ public:
if (params_base.model.path.empty()) {
return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, "Model not loaded");
}
#ifdef LOCALAI_LLAMA_CPP_NO_SCORE_TASK
(void) request;
(void) response;
return grpc::Status(grpc::StatusCode::UNIMPLEMENTED,
"Score is unavailable in this llama.cpp fork backend");
#else
if (!params_base.score_enabled) {
return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION,
"Score was not enabled when the model was loaded; add score to known_usecases");
}
if (request->candidates_size() == 0) {
return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT, "candidates must be non-empty");
}
// Tripwire against the slot loop. Acquired before score_mutex
// so it fires even when this Score is queued behind another.
conflict_guard guard("Score", score_inflight, slot_loop_inflight, "slot_loop_inflight");
// Serialise concurrent Score calls. The slot loop is still
// free to race with us — see the class comment above.
static std::mutex score_mutex;
std::lock_guard<std::mutex> score_lock(score_mutex);
llama_context * lctx = ctx_server.get_llama_context();
if (lctx == nullptr) {
return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, "llama context unavailable (sleeping?)");
}
const llama_vocab * vocab = ctx_server.impl->vocab;
const int32_t n_vocab = llama_vocab_n_tokens(vocab);
const int32_t n_ctx = llama_n_ctx(lctx);
llama_memory_t mem = llama_get_memory(lctx);
// The KV-cache is sized to seq_to_stream.size() at load
// (typically equal to n_slots, often 1). Sequence IDs must
// be in [0, n_seq_max), so we can't pick a high-value
// "private" ID — we have to share with the slot. We clear
// the cache before AND after each candidate to keep
// scoring isolated from whatever state the slot held, and
// the static mutex above guarantees no other Score call is
// racing in the meantime. The slot loop is still free to
// race (see comment on this method) — Score must not run
// concurrently with generation through this backend.
const llama_seq_id score_seq_id = 0;
llama_memory_seq_rm(mem, score_seq_id, -1, -1);
// Tokenize the shared prompt once with add_special=true so
// BOS is prepended when the model requires it. parse_special
@@ -3023,6 +3038,15 @@ public:
std::vector<llama_token> prompt_tokens = common_tokenize(vocab, prompt, /*add_special=*/true, /*parse_special=*/true);
const int32_t prompt_len = (int32_t) prompt_tokens.size();
// Per candidate: full prompt+candidate token list and the
// divergence point, kept for piece rendering and empty-candidate
// handling after the task comes back.
std::vector<std::vector<llama_token>> cand_tokens(request->candidates_size());
std::vector<int32_t> cand_divergence(request->candidates_size(), 0);
// candidates that actually have tokens to score
std::vector<int32_t> included;
for (int ci = 0; ci < request->candidates_size(); ci++) {
const std::string & candidate_text = request->candidates(ci);
@@ -3039,9 +3063,135 @@ public:
break;
}
}
divergence = std::min<int32_t>(divergence, (int32_t) full_tokens.size());
const int32_t cand_len = (int32_t) full_tokens.size() - divergence;
if (cand_len > 0 && divergence < 1) {
// Need at least one prior token (typically BOS) to
// predict the first candidate token's logit. Tokeniser
// models without BOS + an empty prompt fall in here.
return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT,
"Score: prompt produced no leading tokens; need at least one (e.g. BOS) to predict candidate");
}
if (cand_len > SERVER_SCORE_MAX_CAND_TOKENS) {
// The context reserves logits outputs for at most this many
// candidate tokens per slot (server_n_outputs_max).
return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT,
"Score: candidate " + std::to_string(ci) + " is " + std::to_string(cand_len) +
" tokens; the maximum is " + std::to_string(SERVER_SCORE_MAX_CAND_TOKENS));
}
cand_divergence[ci] = divergence;
cand_tokens[ci] = std::move(full_tokens);
if (cand_len > 0) {
included.push_back(ci);
}
}
auto rd = ctx_server.get_response_reader();
bool posted_task = false;
// Shared prefix bounds, needed again when stitching the results:
// n_shared is the longest common token prefix of the scored
// candidates, n_score_prompt the earliest divergence from the
// bare prompt (scored logprobs start there).
int32_t n_shared = 0;
int32_t n_score_prompt = 0;
if (!included.empty()) {
const auto & first = cand_tokens[included[0]];
// the common prefix of a set is the shortest common prefix
// against any fixed member
n_shared = (int32_t) first.size();
for (int32_t ci : included) {
const auto & ft = cand_tokens[ci];
const int32_t lim = std::min<int32_t>(n_shared, (int32_t) ft.size());
int32_t match = 0;
while (match < lim && ft[match] == first[match]) {
match++;
}
n_shared = match;
}
// below its divergence every candidate equals the prompt
// tokens, so n_score_prompt <= n_shared always holds
n_score_prompt = cand_divergence[included[0]];
for (int32_t ci : included) {
n_score_prompt = std::min(n_score_prompt, cand_divergence[ci]);
}
// Map the caller's stable-prefix byte length onto a token
// index: the last prompt token that ends at or before the
// boundary. A checkpoint forced there survives every future
// probe under the same option list, which is what keeps
// repeat scoring cheap on models that cannot rewind state.
int32_t n_stable_prompt = 0;
if (request->stable_prefix_len() > 0) {
size_t consumed = 0;
for (int32_t ti = 0; ti < n_score_prompt; ti++) {
const size_t piece_len = common_token_to_piece(vocab, prompt_tokens[ti]).size();
// BOS and other zero-length specials consume no prompt bytes
if (consumed + piece_len > (size_t) request->stable_prefix_len()) {
break;
}
consumed += piece_len;
n_stable_prompt = ti + 1;
}
}
server_task task(SERVER_TASK_TYPE_SCORE);
task.id = rd.queue_tasks.get_new_id();
task.index = 0;
task.tokens = server_tokens(llama_tokens(first.begin(), first.begin() + n_shared), false);
task.n_score_prompt = n_score_prompt;
task.n_stable_prompt = n_stable_prompt;
task.score_suffixes.reserve(included.size());
for (int32_t ci : included) {
task.score_suffixes.emplace_back(cand_tokens[ci].begin() + n_shared, cand_tokens[ci].end());
}
std::vector<server_task> tasks;
tasks.push_back(std::move(task));
rd.post_tasks(std::move(tasks));
posted_task = true;
}
// Wait for the shared-prefix and per-candidate logprob vectors.
// Context overflow and decode failures surface here as task errors.
std::vector<float> shared_logprobs;
std::vector<std::vector<float>> cand_logprobs;
if (posted_task) {
auto all_results = rd.wait_for_all([&context]() { return context->IsCancelled(); });
if (all_results.is_terminated) {
return grpc::Status(grpc::StatusCode::CANCELLED, "Request cancelled by client");
}
if (all_results.error) {
return grpc::Status(grpc::StatusCode::INTERNAL,
all_results.error->to_json().value("message", "Error in receiving score results"));
}
if (all_results.results.size() != 1) {
return grpc::Status(grpc::StatusCode::INTERNAL, "expected a single score result");
}
auto * score_res = dynamic_cast<server_task_result_score*>(all_results.results[0].get());
if (score_res == nullptr) {
return grpc::Status(grpc::StatusCode::INTERNAL, "unexpected result type for score task");
}
shared_logprobs = std::move(score_res->shared_logprobs);
cand_logprobs = std::move(score_res->cand_logprobs);
if (cand_logprobs.size() != included.size()) {
return grpc::Status(grpc::StatusCode::INTERNAL, "score result candidate count mismatch");
}
}
size_t inc = 0; // index into included / cand_logprobs
for (int ci = 0; ci < request->candidates_size(); ci++) {
const int32_t divergence = cand_divergence[ci];
const int32_t cand_len = (int32_t) cand_tokens[ci].size() - divergence;
backend::CandidateScore * cs = response->add_candidates();
cs->set_num_tokens(cand_len);
cs->set_num_tokens(cand_len > 0 ? cand_len : 0);
if (cand_len <= 0) {
cs->set_log_prob(0.0);
if (request->length_normalize()) {
@@ -3049,101 +3199,57 @@ public:
}
continue;
}
if (divergence < 1) {
// Need at least one prior token (typically BOS) to
// predict the first candidate token's logit. Tokeniser
// models without BOS + an empty prompt fall in here.
return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT,
"Score: prompt produced no leading tokens; need at least one (e.g. BOS) to predict candidate");
// Stitch the candidate's scored logprobs back together: the
// stretch inside the shared prefix (identical for every
// candidate) followed by its forked suffix. Suffix entries
// before the candidate's own divergence are prompt tokens
// decoded only as context — not scored.
std::vector<float> lp;
lp.reserve(cand_len);
for (int32_t t = divergence; t < n_shared; t++) {
const int32_t idx = t - n_score_prompt;
if (idx < 0 || idx >= (int32_t) shared_logprobs.size()) {
return grpc::Status(grpc::StatusCode::INTERNAL,
"Score: shared logprob index out of range for candidate " + std::to_string(ci));
}
lp.push_back(shared_logprobs[idx]);
}
if ((int32_t) full_tokens.size() > n_ctx) {
return grpc::Status(grpc::StatusCode::OUT_OF_RANGE,
"Score: prompt+candidate exceeds context size (got " +
std::to_string(full_tokens.size()) + ", n_ctx=" + std::to_string(n_ctx) + ")");
const auto & sfx_lp = cand_logprobs[inc++];
for (int32_t j = std::max(0, divergence - n_shared); j < (int32_t) sfx_lp.size(); j++) {
lp.push_back(sfx_lp[j]);
}
// Build a batch covering the entire prompt+candidate. We
// need logits at (divergence-1) onward — those are the
// predictions for each candidate token.
llama_batch batch = llama_batch_init((int32_t) full_tokens.size(), 0, 1);
for (int32_t i = 0; i < (int32_t) full_tokens.size(); i++) {
batch.token[i] = full_tokens[i];
batch.pos[i] = i;
batch.n_seq_id[i] = 1;
batch.seq_id[i][0] = score_seq_id;
// logits[i] is "do we want the prediction *for the
// next token*, computed from this position?"
// We want predictions for candidate tokens at
// positions divergence .. full_tokens.size()-1, which
// come from logits at positions (divergence-1) ..
// (full_tokens.size()-2).
bool need_logit = (i >= divergence - 1) && (i < (int32_t) full_tokens.size() - 1);
batch.logits[i] = need_logit ? 1 : 0;
}
batch.n_tokens = (int32_t) full_tokens.size();
// Decode the batch. If decode fails (e.g. KV slot
// exhaustion), surface as INTERNAL — the caller will
// typically fall back to a sampling-based classifier.
int decode_err = llama_decode(lctx, batch);
if (decode_err != 0) {
llama_batch_free(batch);
llama_memory_seq_rm(mem, score_seq_id, -1, -1);
if ((int32_t) lp.size() != cand_len) {
return grpc::Status(grpc::StatusCode::INTERNAL,
"llama_decode failed during Score: " + std::to_string(decode_err));
"Score: result for candidate " + std::to_string(ci) + " is missing token logprobs");
}
// Sum log-probabilities of the actual candidate tokens.
double total_log_prob = 0.0;
for (int32_t k = 0; k < cand_len; k++) {
// The k-th candidate token sits at full_tokens index
// (divergence + k). Its predicting logit is at batch
// position (divergence + k - 1).
int32_t logit_pos = divergence + k - 1;
const float * logits = llama_get_logits_ith(lctx, logit_pos);
if (logits == nullptr) {
llama_batch_free(batch);
llama_memory_seq_rm(mem, score_seq_id, -1, -1);
const float token_log_prob = lp[k];
if (std::isnan(token_log_prob)) {
return grpc::Status(grpc::StatusCode::INTERNAL,
"llama_get_logits_ith returned null at position " + std::to_string(logit_pos));
"Score: incomplete result for candidate " + std::to_string(ci) +
" at token " + std::to_string(k));
}
llama_token target_token = full_tokens[divergence + k];
// Compute log_softmax(logits)[target_token] with the
// max-subtraction stability trick.
float max_logit = logits[0];
for (int32_t v = 1; v < n_vocab; v++) {
if (logits[v] > max_logit) max_logit = logits[v];
}
double sum_exp = 0.0;
for (int32_t v = 0; v < n_vocab; v++) {
sum_exp += std::exp((double)(logits[v] - max_logit));
}
double token_log_prob = (double)(logits[target_token] - max_logit) - std::log(sum_exp);
total_log_prob += token_log_prob;
total_log_prob += (double) token_log_prob;
if (request->include_token_logprobs()) {
backend::TokenLogProb * tlp = cs->add_tokens();
std::string piece = common_token_to_piece(lctx, target_token);
tlp->set_token(piece);
tlp->set_token(common_token_to_piece(vocab, cand_tokens[ci][divergence + k]));
tlp->set_log_prob(token_log_prob);
}
}
cs->set_log_prob(total_log_prob);
if (request->length_normalize() && cand_len > 0) {
if (request->length_normalize()) {
cs->set_length_normalized_log_prob(total_log_prob / (double) cand_len);
}
llama_batch_free(batch);
// Drop this candidate's KV-cache contribution so the next
// candidate starts from a clean state. Without this, the
// next decode would conflict at positions 0..N-1 for our
// sequence ID.
llama_memory_seq_rm(mem, score_seq_id, -1, -1);
}
return grpc::Status::OK;
#endif
}
grpc::Status TokenizeString(ServerContext* context, const backend::PredictOptions* request, backend::TokenizationResponse* response) override {
@@ -3154,7 +3260,6 @@ public:
if (params_base.model.path.empty()) {
return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, "Model not loaded");
}
conflict_guard guard("TokenizeString", slot_loop_inflight, score_inflight, "score_inflight");
json body = parse_options(false, request, params_base, ctx_server.get_llama_context());
body["stream"] = false;
@@ -3174,9 +3279,23 @@ public:
return grpc::Status::OK;
}
grpc::Status Detokenize(ServerContext* context, const backend::DetokenizeRequest* request, backend::DetokenizeResponse* response) override {
auto auth = checkAuth(context);
if (!auth.ok()) return auth;
if (params_base.model.path.empty()) {
return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, "Model not loaded");
}
std::string content;
for (const auto token : request->tokens()) {
content.append(common_token_to_piece(ctx_server.get_llama_context(), token));
}
response->set_content(content);
return grpc::Status::OK;
}
grpc::Status GetMetrics(ServerContext* /*context*/, const backend::MetricsRequest* /*request*/, backend::MetricsResponse* response) override {
conflict_guard guard("GetMetrics", slot_loop_inflight, score_inflight, "score_inflight");
// request slots data using task queue
auto rd = ctx_server.get_response_reader();

Some files were not shown because too many files have changed in this diff Show More