mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-18 18:59:29 -04:00
144baaa809c4ab14282ee4bfeb3235bac0d28fc8
7438
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
144baaa809 |
chore: ⬆️ Update ggml-org/whisper.cpp to 306c88f4d1286aec1bf96e544632897886af5501 (#11353)
⬆️ Update ggml-org/whisper.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
86c2e9a273 |
chore: ⬆️ Update leejet/stable-diffusion.cpp to ea7f0c87cfe4c673263b4c201c596c7f1cbe2528 (#11354)
⬆️ Update leejet/stable-diffusion.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
89995d7535 |
chore: ⬆️ Update 0xShug0/audio.cpp to 238ab6a9e321c17de8e120559f57efeedaeb1345 (#11355)
⬆️ Update 0xShug0/audio.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
1f4ec3bdf8 |
chore: ⬆️ Update CrispStrobe/CrispASR to ec730908a418b6032f9e69ded6186d3f042a7747 (#11356)
⬆️ Update CrispStrobe/CrispASR Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
1466aaa9f7 |
chore: ⬆️ Update antirez/ds4 to 6747e7718dd08f00b680d0c16231f2d59ec3747e (#11357)
⬆️ Update antirez/ds4 Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
e6712844ee |
chore: ⬆️ Update ikawrakow/ik_llama.cpp to 6b55d2c7504f482e7c8ec6cbf22a19f3778c522b (#11358)
⬆️ Update ikawrakow/ik_llama.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
b1d964ef7b |
chore(model-gallery): ⬆️ update checksum (#11359)
⬆️ Checksum updates in gallery/index.yaml Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
0d342c61d8 |
docs(backends): correct the vllm-cpp description in the gallery (#11363)
This is the text users read in the backends list and the gallery, and it was the last place still describing vllm.cpp as "a from-scratch C++20 port of vLLM created and maintained by the LocalAI team" with no indication of maturity. Three corrections, matching the v4.8 release notes and blog post: - It leads with ALPHA. These are alpha development builds and llama-cpp stays the recommendation for production, which is the single most useful thing to know before clicking install. - It is maintained by the LocalAI team but developed in its own repository and usable without LocalAI. vLLM is named for what it actually is, the reference implementation that output is checked against and benchmarked against, rather than just the thing that was ported. - It records the featureset that has grown past vLLM: GGUF loading, speculative decoding and KV offload, alongside the architecture and hardware coverage that were already listed. Also notes that the project is expected to be renamed, with the new name still to be decided, so anyone who installs it now is not surprised later. vllm-cpp-development inherits all of this through the YAML anchor, so both entries are covered by the one edit. Verified the file still parses and that both entries carry the new text. Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
4fec33966a |
docs(blog): final figures for the 4.8 post, and the MLX provider (#11362)
* docs(blog): final figures for the 4.8 post, and the MLX provider The cycle closed at 374 PRs over twenty-one days, not the 321 over eighteen the post was written against. Corrects the summary, the opening line, the contributor count and the gallery total, and moves the date to the day the release is cut. Adds the MLX GEMM provider (#11137), which merged after the post was written and is the one number an Apple Silicon reader wants: 1.54x to 2.19x on an M4 with time to first token roughly halving, both arms toggled on one binary. The +/-10% caveat travels with the table rather than being left in the PR. Two lines edited against the no-ai-slop skill while I was in the file, the same pass #11324 ran over the engines post: - The opener balanced two clauses across a colon and closed on "without lying to you", which is the built-to-be-quoted shape readers picked out of the HN thread. It is a flat statement now. - "This is a new modality rather than a new backend under an existing one" is a binary contrast that says nothing the next clause does not. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash] * docs(blog): call vllm.cpp alpha, and finish the no-ai-slop pass vllm.cpp is not a released backend and the post read like it was. The old wording buried the caveat in a block quote at the end of the section and still said "first release of a young engine". It now says plainly, before the caveat can be skipped, that these are alpha development builds, that shipping them in 4.8 is about letting people try the thing rather than recommending it, and that llama-cpp stays the default. Also completes the no-ai-slop pass I had only half run. Counting the lines built to be quoted, headings and section endings included, the post is in reasonable shape: long flat informational stretches, tables followed by a plain finding, headings that are labels rather than epigram-verdicts. Three patterns survived, each one an item in eval.md: - "and inverts that:" set the usual shape against ours across a colon. The sentence works without the frame. - "Two things were conflated there: a signal, which needs one line, and the detail, which needs somewhere to put it" is a role-assignment pair. Says what happens instead. - "The maturity statement from the release notes is worth repeating in full" is throat-clearing in front of a quote, and the quote is gone. Left the rest alone. Minimum effective edit, not a rewrite. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash] * docs(blog): present vllm.cpp as a community project, with its own numbers The post described vllm.cpp as "a from-scratch port of vLLM, written and maintained by the LocalAI team". Two things wrong with that. It is a community project, and it has stopped being only a port: it loads GGUF, runs on CPU, Metal and Vulkan, ships speculative decoding and KV offload, and its benchmark page measures against llama.cpp, MLX-LM and DwarfStar as well as vLLM, because those are the engines it competes with on that hardware. vLLM's role is now stated for what it is, the reference implementation. Correctness is checked against it and the scoreboard is kept against it. Also flags that the name will probably change, since it is drifting far enough that vllm.cpp will eventually mislead. Adds real numbers from the project's own docs/BENCHMARKS.md rather than adjectives: 1.045x vLLM at concurrency 1 on Qwen3.6-27B NVFP4 with token-for-token identical output, 1.010x and 1.013x at c16 and c32 on the 35B MoE and behind below that, prefill 1.18x over llama.cpp on CPU aarch64, 97.6% of MLX-LM warm total on an M4. Upstream's own caution travels with them: it treats c2 through c32 as ties because its noise band is 0.5% and those margins are 0.7% to 1.7%. Every figure was checked against ~/_git/vllm.cpp/docs/BENCHMARKS.md rather than restated from memory. The heading is marked alpha to match the section body. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash] * docs(blog): say who maintains vllm.cpp, and add the DeepSeek Flash result Two corrections to the previous commit. "A community project" says nothing and was not quite true either. The LocalAI team maintains vllm.cpp. Community-first is the intent, not a description, so it now says that and says what backs it: its own repository, its own docs, benchmark record and issue tracker, and it runs without LocalAI anywhere in the picture. Adds the DeepSeek-V4-Flash result, which makes the divergence point better than any of the prose around it. That model does not run on vLLM on a single GB10: every vLLM-loadable checkpoint is 156 GB or more against a 119 GiB unified pool, and the only quant that fits is an extreme-low-bit GGUF that vLLM cannot load. vllm.cpp reads GGUF and runs it at 16.28 tok/s against ds4's 16.33, a parity result. Also notes MTP speculative decoding, token-identical to vLLM's and about 4% faster at concurrency 1. Both figures checked against ~/_git/vllm.cpp/docs/BENCHMARKS.md. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash] * docs(blog): lead the DeepSeek result with what we run, not with what vLLM cannot The previous version opened on "that model does not run on vLLM on a single GB10 at all". Wrong emphasis twice over: it makes a strong negative claim about another project the headline, and it buries the actual result, which is that vllm.cpp runs DeepSeek-V4-Flash at roughly 2-bit (IQ2_XXS mixed, about 80 GB) on a single DGX Spark and decodes at 16.28 tok/s against DwarfStar's 16.33. The size constraint is still there, stated as the reason the quant is what it is rather than as a point about vLLM: at 300B+ total parameters even a 4-bit checkpoint is 156 GB or more, so a 2-bit GGUF is what fits the Spark's 119 GiB unified pool. The table row now names the quant and the box (IQ2_XXS, one DGX Spark) instead of just "GGUF, GB10", since that is the part a reader with a Spark wants. Figures unchanged and still from ~/_git/vllm.cpp/docs/BENCHMARKS.md. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash] * docs(blog): say the new name is undecided "The name will probably change at some point" invited the obvious question. It now says the rename is expected and the name is still to be decided, which is the actual state. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
8f52437c81 |
fix(gallery): describe Genesis Hermes model accurately (#11342)
Replace copied HauhauCS base-model text with metadata for the actual Genesis Hermes V6 artifact and link its upstream base model. Assisted-by: Codex:gpt-5 [Hugging Face] Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
cd516452dd |
fix(rocm): stop building the ggml CPU variant matrix for hipblas llama.cpp (#11346)
No -gpu-rocm-hipblas-llama-cpp image has been published since 2026-08-01.
Every build since has been killed by GitHub at exactly its 6h job limit:
job 91830652349 cancelled 6h00m (2026-08-04)
job 91763226161 cancelled 6h00m (2026-08-03)
job 91466626154 cancelled 6h00m (2026-08-02 full matrix)
The registry shows the damage: master-gpu-rocm-hipblas-llama-cpp last
built 2026-08-01 05:53, latest-gpu-rocm-hipblas-llama-cpp 2026-07-15,
against master-cpu-llama-cpp which is current.
Same cause as #11321, different mechanism. Since #11255 every x86 GPU
image also builds ggml's CPU_ALL_VARIANTS matrix. SYCL died because icpx
stalls on one translation unit; ROCm dies on volume. hipcc compiles the
HIP kernels once per entry in AMDGPU_TARGETS, and that list is eleven
architectures (gfx908, gfx90a, gfx942, gfx950, gfx1030, gfx1100, gfx1101,
gfx1102, gfx1151, gfx1200, gfx1201). The CPU matrix lands on top of that.
The numbers are unambiguous. The same job took 2h27m in the 2026-07-26
full matrix, before #11255. #11255 merged 2026-08-01 07:26, an hour and a
half after the last image was published, and it has been 6h00m ever since.
The tail of the last run shows it 61% through ggml-hip at the 83 minute
mark, still building HIP template instances.
Route hipblas to the portable fallback, exactly as #11321 did for SYCL and
for the same practical reason: it is what these images shipped before
#11255, and run.sh already prefers *-cpu-all when present and falls back
otherwise. Expected to restore the 2h27m build with room to spare.
Not fixed here: the CPU variant matrix is genuinely wanted on ROCm for
partial offload. Getting it needs the build to fit in 6h, which means
trimming AMDGPU_TARGETS or splitting the job per architecture. Both are
larger changes than unbreaking the image, and neither should ride along
with a build that is currently not shipping at all.
Verified: make test-build-scripts passes, including the extended
llama-cpp-build-target_test.sh. bonsai is unaffected (own compile script,
ROCm builds in 1h52m) and turboquant has no hipblas variant.
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
3f0db2a9c2 |
feat(vllm-cpp): enable and vendor the MLX GEMM provider on darwin/metal (#11137)
* feat(vllm-cpp): enable and vendor the MLX GEMM provider on darwin/metal
The darwin vllm-cpp image built the Metal backend with vllm.cpp's native MSL
GEMM only. vllm.cpp also ships an optional MLX provider for the dense GEMM,
kept OFF upstream because it costs a ~19 MB libmlx.dylib plus a ~105 MB
mlx.metallib, on the stated position that it must earn that cost by
measurement.
Measured on an Apple M4 (16 GiB, macOS 26.5.2) it does. One binary, arms
toggled with VT_OP_PROVIDER_DISABLE=mlx so there is no build-difference
confound, Qwen3-1.7B-bf16 p=512 g=128, 2 reps, arm order alternated per rep:
B=1 5.79 vs 3.08 agg tok/s (1.88x) TTFT 3.32 s vs 7.68 s
B=8 25.70 vs 13.69 (1.88x) TTFT 13.95 s vs 34.38 s
B=16 38.65 vs 17.69 (2.19x) TTFT 18.33 s vs 54.48 s
Peak RSS is unchanged (6.65 to 7.50 GB in both arms) and the output is
bit-identical: vllm.cpp's three-way parity test measures mlx-vs-msl NMSE of 0
on all six shapes, and mlx-vs-cpu equal to msl-vs-cpu, against a 5e-4 bar. MLX
serves the dense GEMM alone; paged attention stays vllm.cpp's own kernel
because MLX has no paged-KV primitive. Full disposition, including the
INDICATIVE status and the isolation actually achieved, is in vllm.cpp
docs/BENCHMARKS.md "MLX GEMM provider A/B on Apple M4".
Build: MLX comes from the pinned prebuilt pip wheel (MLX_VERSION, default
0.29.3) into a venv under the backend dir. Building MLX from source needs
`xcrun metal`, i.e. a full Xcode the macOS runners do not have, while the wheel
ships include/, lib/libmlx.dylib and the compiled metallib ready to link. The
install is a stamp FILE rather than a phony target, because a phony
prerequisite is always newer than libvllm and would re-link it every
invocation. VLLM_CPP_MLX=off restores the previous Metal build.
Packaging vendors libmlx.dylib, mlx.metallib and MLX's MIT license into
package/lib/. Three things this had to get right, each verified on the M4
before it was written rather than after:
1. libvllm.dylib links @rpath/libmlx.dylib and its build-time LC_RPATH points
inside the build venv, a path no user has. Every build rpath is deleted
and replaced with @loader_path/lib.
2. MLX loads its metallib from beside its OWN dylib, so both files must land
in the same directory or every Metal op fails with "Failed to load the
default metallib".
3. install_name_tool invalidates the code signature and macOS refuses to load
an arm64 image with a stale one, so the patched library is re-signed
ad-hoc.
Verified end to end on the M4 by building through this Makefile and running the
packaged artifact: `DYLD_PRINT_LIBRARIES` resolves libmlx from package/lib/,
`codesign -v` passes, no build-venv path survives in the load commands, and a
real generation runs with the provider selected (op=65 selected=mlx) and zero
metallib failures. A missing rpath now fails the build instead of the user's
first inference.
Cost: the darwin vllm-cpp image grows by about 124 MB.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
* fix(vllm-cpp): default the MLX GEMM provider OFF on darwin
This branch opened with VLLM_CPP_MLX=on, justified by an A/B that measured the
MLX provider at 1.88x to 2.19x against the native MSL GEMM. That measurement was
correct when taken and is now stale: vllm.cpp's own Metal kernels have improved
several-fold since, through mma prefill attention, a vectorised decode V
accumulation, vectorised attention staging, a fused qk-norm-RoPE preamble and a
simdgroup-per-row softmax. The native path MLX was compared against no longer
exists.
Re-measured on the same Apple M4, in the same binary, with the arms toggled by
VT_OP_PROVIDER_DISABLE=mlx, on Qwen3-1.7B-bf16 warm at p=512 g=128:
MLX provider ON prefill TTFT 1370 ms warm throughput 11.98 tok/s
MLX provider OFF prefill TTFT 1400 ms warm throughput 22.06 tok/s
Shipping the previous default would have halved Apple Silicon throughput.
MLX's steel GEMM is still about 20% faster than ours in isolation, but the
provider pays a per-op mx::eval synchronisation plus an output memcpy, because it
cannot write into our buffer. Across prefill's roughly 112 GEMMs that overhead
leaves a 2% gain; on decode, where the same synchronisation is paid once per
matmul per token, it costs 46%. The option is kept for prefill-dominated
workloads, where the margin is small but real.
The README section is rewritten rather than patched: it previously presented the
stale table as the reason for the default, so leaving it in place would have made
the new default look arbitrary.
Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(vllm-cpp): bump vllm.cpp and default MLX ON, gated to prefill
Bumps VLLM_CPP_VERSION from 9e1c9025 to eec09bed and turns VLLM_CPP_MLX back on.
These two must move together, which is why they are one commit.
Upstream now shape-gates the MLX provider to prefill: it declines m < 2, which is
exactly the decode GEMV. MLX's steel GEMM wins prefill, 524.5 ms of TTFT against
602 for the native path, but loses decode badly because the provider pays an
mx::eval synchronisation and an output memcpy on every call while decode makes
about 112 calls per token. Ungated it does both; gated it does only the good half.
Measured on an Apple M4 with Qwen3-1.7B-bf16 warm at p=512 g=128:
MLX gated to prefill (pin >= 89c46aeb) TTFT 524.5 ms 24.40 tok/s, 99.1% of MLX-LM
MLX ungated (older pins) TTFT 537 ms 12.7 tok/s
MLX off TTFT 602 ms 23.9 tok/s
This branch briefly defaulted the provider off, which was the correct call for an
ungated provider at the old pin. The gate is what makes on correct again, so the
pin and the flag are coupled: rolling VLLM_CPP_VERSION back before 89c46aeb while
leaving MLX on would select the middle row and roughly halve throughput. Both the
Makefile comment and the README state that dependency explicitly.
The bump also brings six Metal kernels landed upstream since the old pin — mma
prefill attention, a vectorised decode V accumulation, vectorised attention
staging, a fused qk-norm-RoPE preamble, a simdgroup-per-row softmax and a
simdgroup-per-head preamble — which take the non-MLX Metal path from 89.4% to
96.4% of MLX-LM on their own.
One caveat, recorded in the README: MLX's GEMM is not bit-identical to the native
kernel, so an MLX build produces a different greedy sequence than a non-MLX build.
That is a property of the provider rather than of the gate and predates this
packaging.
Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs(vllm-cpp): correct the MLX-gated figure to 97.6%, from 99.1%
The previous commit quoted 99.1% of MLX-LM for the prefill-gated MLX build. That
figure divided by a two-run MLX-LM baseline, 27.135 and 27.744 generation tok/s
averaged to 27.44. Re-measured interleaved with ours over four ABBA blocks,
MLX-LM's decode is 27.848 with a 0.34% spread across six runs, so the 27.135 was
an outlier and averaging it in overstated us by roughly 1.5 points.
Corrected: the gated configuration is 24.37 tok/s, or 97.6% of MLX-LM, and the
MLX-off build is 23.9 tok/s or 95.9%. Prefill TTFT is unchanged at 524.5 ms
against MLX-LM's 532.6, so we remain about 1.5% faster there.
Nothing else changes. MLX still wins prefill and loses decode, the shape gate is
still the right disposition, and the pin and the flag are still coupled. The gate
is worth about 1.7 points over the MLX-off build rather than 2.7.
Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(vllm-cpp): pin MLX gate from mainline
The previous pin was a merge commit from the experimental C ABI v9 branch. Pin the same MLX prefill gate on upstream main so the backend build does not pull unrelated ABI v9 work into every platform variant.
Assisted-by: Codex:gpt-5 [systematic-debugging]
* fix(vllm-cpp): restore backend build portability
Keep the current master pin when enabling MLX so every backend variant builds against the known-good vllm.cpp revision. Suppress Apple clang’s GNU constant-folding diagnostic for Objective-C++ Metal compilation only, since upstream treats warnings as errors.
Assisted-by: Codex:gpt-5 [systematic-debugging]
* fix(vllm-cpp): demote MLX header VLA warning
MLX 0.29.3 headers trigger Apple clang's gnu-folding-constant diagnostic in the Objective-C++ provider. Keep the diagnostic visible while exempting only it from vllm.cpp's global warnings-as-errors policy.
Assisted-by: Codex:gpt-5 [systematic-debugging]
* fix(vllm-cpp): suppress MLX header VLA warning
Target-level Objective-C++ -Werror is appended after the directory flags, so a no-error demotion is re-promoted. Disable this single warning for the MLX header while keeping every other warning fatal.
Assisted-by: Codex:gpt-5 [systematic-debugging]
* fix(vllm-cpp): pin source-scoped MLX warning fix
Move the AppleClang warning exception into vllm.cpp where its target warning policy is defined, and pin LocalAI to that source-scoped fix.
Assisted-by: Codex:gpt-5
* fix(vllm-cpp): pin effective MLX warning suppression
The source-scoped no-error flag was overridden by the target warning policy. Pin the companion vllm.cpp change that disables only the MLX header diagnostic for its Objective-C++ translation unit.
Assisted-by: Codex:gpt-5
* fix(vllm-cpp): pin diagnostic pragma fix
Pin the companion vllm.cpp correction that scopes the AppleClang folding warning suppression inside the MLX translation unit, after command-line warning policy.
Assisted-by: Codex:gpt-5 [systematic-debugging]
* fix(vllm-cpp): pin remaining Darwin build fixes
Advance the MLX-enabled backend to the vllm.cpp revision already validated by the dependency update branch. This includes the feature guards and AppleClang pragma boundary needed by the Darwin build.
Assisted-by: Codex:gpt-5 [systematic-debugging]
* fix(vllm-cpp): pin MLX system dependency boundary
Pin the companion vllm.cpp change that models MLX as an imported system dependency, keeping third-party header diagnostics out of the project's warnings-as-errors policy while retaining fatal warnings for project sources.
Assisted-by: Codex:gpt-5 [Codex]
* fix(vllm-cpp): pin scoped MLX warning guard
Advance vllm.cpp to the companion fix that keeps MLX headers on a SYSTEM dependency and scopes AppleClang folding-constant suppression to the external includes.
Assisted-by: Codex:gpt-5 [systematic-debugging] [test-driven-development]
* fix(vllm-cpp): use available MLX wheel
MLX 0.29.3 is no longer available to the Darwin runner, so the backend build stopped before CMake. Pin the first available compatible wheel and keep the documented default in sync.
Assisted-by: Codex:gpt-5
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
|
||
|
|
137dfcf15a |
chore: ⬆️ Update antirez/ds4 to b7e9f0091139999b6c070a57590c447c5741da5c (#11333)
* ⬆️ Update antirez/ds4 Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> * fix(ds4): link upstream CUDA MMQ objects The updated ds4 CUDA object now calls into the vendored MMQ implementation. Build and link those objects into both the gRPC server and distributed worker. Assisted-by: Codex:gpt-5 [Codex] --------- Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
750ab91b2b |
test(advisorylock): replace fixed sleeps with signals (#11343)
Wait for observable loop events instead of budgeting hundreds of milliseconds for scheduler timing. Keep a short bounded overlap observation for the two-leader exclusion check. Assisted-by: Codex:gpt-5 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
08598a8611 |
chore(deps): bump the npm_and_yarn group across 1 directory with 5 updates (#11341)
Bumps the npm_and_yarn group with 5 updates in the /core/http/react-ui directory: | Package | From | To | | --- | --- | --- | | [hono](https://github.com/honojs/hono) | `4.12.25` | `4.12.34` | | [@hono/node-server](https://github.com/honojs/node-server) | `1.19.14` | `2.1.0` | | [fast-uri](https://github.com/fastify/fast-uri) | `3.1.4` | `3.1.5` | | [ip-address](https://github.com/beaugunderson/ip-address) | `10.2.0` | `10.4.0` | | [undici](https://github.com/nodejs/undici) | `7.28.0` | `7.29.0` | Updates `hono` from 4.12.25 to 4.12.34 - [Release notes](https://github.com/honojs/hono/releases) - [Commits](https://github.com/honojs/hono/compare/v4.12.25...v4.12.34) Updates `@hono/node-server` from 1.19.14 to 2.1.0 - [Release notes](https://github.com/honojs/node-server/releases) - [Commits](https://github.com/honojs/node-server/compare/v1.19.14...v2.1.0) Updates `fast-uri` from 3.1.4 to 3.1.5 - [Release notes](https://github.com/fastify/fast-uri/releases) - [Commits](https://github.com/fastify/fast-uri/compare/v3.1.4...v3.1.5) Updates `ip-address` from 10.2.0 to 10.4.0 - [Release notes](https://github.com/beaugunderson/ip-address/releases) - [Commits](https://github.com/beaugunderson/ip-address/compare/v10.2.0...v10.4.0) Updates `undici` from 7.28.0 to 7.29.0 - [Release notes](https://github.com/nodejs/undici/releases) - [Commits](https://github.com/nodejs/undici/compare/v7.28.0...v7.29.0) --- updated-dependencies: - dependency-name: hono dependency-version: 4.12.34 dependency-type: direct:production dependency-group: npm_and_yarn - dependency-name: "@hono/node-server" dependency-version: 2.1.0 dependency-type: indirect dependency-group: npm_and_yarn - dependency-name: fast-uri dependency-version: 3.1.5 dependency-type: indirect dependency-group: npm_and_yarn - dependency-name: ip-address dependency-version: 10.4.0 dependency-type: indirect dependency-group: npm_and_yarn - dependency-name: undici dependency-version: 7.29.0 dependency-type: indirect dependency-group: npm_and_yarn ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
211aa0a536 |
chore: ⬆️ Update mudler/vllm.cpp to a42b8187caff02c570c28e19e4dc2b1d7f55ed14 (#11174)
⬆️ Update mudler/vllm.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
c86b3b207b |
chore: ⬆️ Update ikawrakow/ik_llama.cpp to 60389410a1ff01f9d37dcc6261db33b3183bdea2 (#11331)
⬆️ Update ikawrakow/ik_llama.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
62316e52a9 |
chore: ⬆️ Update 0xShug0/audio.cpp to 4e3aea2fd99aeaa5924e71c51eb2793846045332 (#11332)
⬆️ Update 0xShug0/audio.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
3090101156 |
chore: ⬆️ Update CrispStrobe/CrispASR to fe3caf8e363b27572dbdd1a9d37083f25e6decda (#11334)
⬆️ Update CrispStrobe/CrispASR Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
8b667cd1ce |
chore: ⬆️ Update ggml-org/whisper.cpp to 64d57d3df5c8dacee098577257edcaa154bf5ef3 (#11326)
⬆️ Update ggml-org/whisper.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
93fe086798 |
chore(deps): bump the npm_and_yarn group across 1 directory with 2 updates (#11338)
Bumps the npm_and_yarn group with 2 updates in the /core/http/react-ui directory: [@hono/node-server](https://github.com/honojs/node-server) and [brace-expansion](https://github.com/juliangruber/brace-expansion). Updates `@hono/node-server` from 1.19.14 to 2.0.12 - [Release notes](https://github.com/honojs/node-server/releases) - [Commits](https://github.com/honojs/node-server/compare/v1.19.14...v2.0.12) Updates `brace-expansion` from 1.1.12 to 1.1.18 - [Release notes](https://github.com/juliangruber/brace-expansion/releases) - [Commits](https://github.com/juliangruber/brace-expansion/compare/v1.1.12...v1.1.18) --- updated-dependencies: - dependency-name: "@hono/node-server" dependency-version: 2.0.12 dependency-type: indirect dependency-group: npm_and_yarn - dependency-name: brace-expansion dependency-version: 1.1.18 dependency-type: indirect dependency-group: npm_and_yarn ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
2e14511fe2 |
docs(blog): add release write-ups for 3.10 through 4.3 (#11330)
The blog has a deep post for 4.8 and a history post that covers the earlier
releases at summary altitude, but nothing in between. These five fill that
gap in the same shape as what-landed-in-localai-4-8: what the release was
for, runnable examples, and the limits that apply.
Every endpoint, CLI flag, env var and gallery entry is verified against the
matching release tag rather than taken from the release notes. That caught
two paths the published 3.10.0 notes got wrong: tracing is /api/traces, not
/api/v1/trace, and a stored response is fetched from /v1/responses/:id, not
/api/v1/responses/{response_id}.
Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
88fdda6211 |
chore(model-gallery): ⬆️ update checksum (#11327)
⬆️ Checksum updates in gallery/index.yaml Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
f447faf08d |
chore: ⬆️ Update ggml-org/llama.cpp to 221f0f6356efe2260023208365705ec5d5a7c8f5 (#11303)
⬆️ Update ggml-org/llama.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
6e7c0a4df8 |
blog, website: edit out the AI writing tells readers called out on HN (#11324)
* blog: rewrite the engines post without the AI tells
The HN thread on this post (item 49125065) spent most of its comments on the
writing rather than the engines. Readers quoted specific lines back as tells.
This is the same post with the same numbers, edited against the updated
no-ai-slop skill.
Every figure, table and link is unchanged, except that "27% of the memory"
is now the underlying 363 MB against 1328 MB from the table.
Two substantive framing fixes, both from the reply draft in
hn-reply-engines-post.md:
- vllm.cpp is no longer implied to be a speed win. The table is a tie, the
result is the install size, and the post now says so before a reader has to
work it out and post about it.
- Added one line on the language mix. Readers took the C++/Python/Go tree as
incoherence rather than as a Go core with per-ecosystem backends.
Cut throughout: the ledger metaphor ("what those ports buy", "not paid for in
throughput"), unearned framing ("the honest reading is", "has nothing to do
with"), the shape summary ("that is the general shape of these wins"),
confident deference ("people who are better at those models than we are"),
self-grading numbers ("a good result for a 66 MiB binary"), verbless
comparisons, three of the four exactness idioms, and the aphoristic headings
and verdicts. The double-tricolon summary is one plain clause now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* blog, website: same anti-slop sweep over the rest of the site
One-by-one pass over the other four posts and the site templates, with the
same rules used on the engines post. All figures, tables, links and PR
numbers are unchanged everywhere; the edits are to prose only.
apex-moe-quantization: ledger metaphors were the main issue, eight uses of
buy/cost/pay/spend for things that are not money. Also "the honest reading
is", "that is the comparison that matters", and two section-ending aphorisms
("Size is a speed knob as much as a memory knob", "Q6_K is the ceiling worth
paying for").
localai-since-march-2023: light touch, this one already reads like a person.
Removed "the curve is not the point", a "not the feature list, but the four
decisions" contrast, and two "X is what made / is the piece that" forms.
parakeet-cpp-asr-on-cpu: six exactness idioms across one post, "byte for
byte" twice, "character for character" twice, "byte-identical" twice and
"bit-identical" once, including in the title. Down to one, kept where the
precision is load-bearing. Also the "what end-of-utterance detection buys
you" heading and the "we say so rather than averaging it away" flex.
what-landed-in-localai-4-8: no changes. It is dense, flat and ends every
section on a PR number or a plain fact, which is the shape the other posts
should look like.
Site templates: "Most backends wrap somebody else's engine. These do not."
was the same contrast the engines post opened with. Also "Not a degraded mode
that technically runs", "A port only ships once it matches the original",
"Speed is the part we then go and win ... not a marketing run", and the last
"byte for byte" on the landing page.
Hugo builds clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* website: it is eighteen engines, not nineteen
Three places said nineteen: the /engines/ page description, the JUL 2026
timeline entry on the landing page, and the header comment in
data/engines.yaml.
Eighteen is right, confirmed two ways. The "Backends built by us" table in
the README has exactly 18 rows, and data/engines.yaml has 19 entries of which
one is apex-quant, which is a quantization recipe rather than an engine. The
two lists otherwise match name for name.
The yaml comment is the likely origin: it read "the nineteen native engines
the LocalAI team wrote, and the one quantization recipe that feeds them",
which counts apex-quant twice.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
e2311045d3 |
fix(mcp): drop the duplicated scheduling methods on stubClient (#11323)
master does not compile:
vet: core/http/endpoints/mcp/localai_assistant_test.go:157:19:
method stubClient.ListScheduling already declared at
core/http/endpoints/mcp/localai_assistant_test.go:87:19
Two fixes for the same breakage landed. The four Scheduling methods were
already present at lines 87-99, in interface order after ListNodes, by
the time #11318 merged; #11318 appended its own copy after
GetRouterDecisions. The two blocks sit in different parts of the file, so
git merged both without a conflict and nothing flagged it.
Remove the appended copy and keep the one in interface order. Pure
deletion, no behaviour change.
Verified: go vet clean on ./core/http/endpoints/mcp/, and
go test ./core/http/endpoints/mcp/ passes.
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
6bdb04ab5d |
docs: point the News page at the blog instead of a stale highlights list
The News page kept a hand-maintained "Highlights" list that had drifted: it was missing all of 2025, duplicated the README's own news list, and linked /features/middleware/ for a page that lives at operations/. Both of its jobs already have owners. website/content/blog/ carries the release write-ups and engineering notes, and GitHub Releases carries the full changelog. Replace the list with a pointer at those two, so there is one place to update instead of three. The page keeps its url and front matter, so /docs/basics/news/ and the root /basics/news/ redirect that .github/ci/gen-redirects.sh generates both keep resolving. Also drop the two contributor instructions in .agents that told authors to add a whats-new.md bullet per feature: announcing a capability is the release blog post's job, per .agents/preparing-a-release.md. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Write] [Bash] |
||
|
|
bd076376be |
fix(ci): install the Go the module asks for when building the site (#11322)
Deploy site to GitHub Pages failed on five of the last eight master
pushes, always in the build job before Hugo runs:
Setup go version spec 1.22
...
go: downloading go1.26.0 (linux/amd64)
go: download go1.26.0: golang.org/toolchain@v0.0.1-go1.26.0.linux-amd64:
Get "https://proxy.golang.org/...": connect: network is unreachable
##[error]Command failed: go env GOPATH
The workflow pinned setup-go to 1.22 while go.mod declares go 1.26.0, so
the `go run ./.github/ci/modelslist.go` step that generates the gallery
page had to fetch the real toolchain from proxy.golang.org first. That
fetch is not reliably reachable from the runner, which is why the deploy
alternated between passing and failing rather than failing outright.
Track go.mod instead of a literal. The version the module needs is then
installed directly and there is no toolchain download to fail.
This matters beyond CI noise: the docs and the site, including the
release blog post, ship through this workflow.
Scoped deliberately to gh-pages, the workflow with the observed failure.
test-extra.yml pins 1.25.4 in a dozen places and is below go.mod for the
same reason, so those jobs also download a toolchain, but they are
currently green and rewriting twelve pins on a hunch risks more than it
fixes. Worth a follow-up.
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
d28ccf32b5 |
gallery: add Qwen3.6 14B FableVibes variants (#11317)
Add Q4_K_M and Q8_0 llama.cpp entries with the shared Q8_0 multimodal projector. Assisted-by: Codex:gpt-5 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
95bd59d78e |
fix(mcp): teach the assistant test stub the scheduling methods (#11318)
#11228 added ListScheduling, GetScheduling, SetScheduling and DeleteScheduling to localaitools.LocalAIClient but did not update stubClient, the hand-written test double in the mcp endpoints package. The package therefore fails to typecheck, which takes out both lint and tests on master: cannot use stubClient{} as localaitools.LocalAIClient value in argument to h.Initialize: stubClient does not implement localaitools.LocalAIClient (missing method DeleteScheduling) Red on |
||
|
|
1741df0bf1 |
fix(ui): scale the chrome audit's timeout to the number of routes it walks (#11319)
chrome-audit.spec.js walks 25 routes in a single test, and has been the
UI E2E suite's failure on 5 of the last 6 master runs. It always dies the
same way, at the 30s per-test default:
Test timeout of 30000ms exceeded.
Error: page.waitForTimeout: Test timeout of 30000ms exceeded.
19 | await page.goto(route)
> 20 | await page.waitForTimeout(400)
The spec is new in 5cb0c1a8; the commit before it was green, and every
run since has been red on this file.
The failure is cumulative rather than one bad route. Across those runs
the clock runs out at line 19, 20 or 21 depending on where the loop
happens to be, and the timeout lands on waitForTimeout rather than on
goto, which is what running out of budget looks like as opposed to a
navigation that hangs. 30s over 25 routes is ~1.2s each, including a
deliberate 400ms settle, so there is very little headroom to begin with.
Give the test a budget proportional to its work: six seconds a route.
That absorbs a slow runner and still fails promptly if a route genuinely
hangs.
Verified: the spec passes on the current UI in 12.2s solo, and the full
suite passes 418 at 8 workers locally. What I could NOT do is reproduce
the CI timeout on this machine, which has 20 cores against the runner's
2 to 4; under synthetic CPU load it still finished in 13.5s. So the fix
is argued from the CI signature and the arithmetic, not from a local
repro, and the proof is this spec going green on the hosted runner.
Note test.setTimeout() has to be called inside the test body. At module
scope Playwright rejects it with "test.setTimeout() can only be called
from a test".
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
b6d2e94153 |
fix(sglang): bound cuda-tile below the 1.6 prereleases (#11320)
Every CUDA sglang image failed in the 2026-08-02 full-matrix rebuild:
-gpu-nvidia-cuda-12-sglang, -gpu-nvidia-cuda-13-sglang and
-nvidia-l4t-cuda-13-arm64-sglang, all with the same build error.
Building cuda-tile==1.6.0rc3
x Failed to build `cuda-tile==1.6.0rc3`
ModuleNotFoundError: No module named 'wheel_stub'
hint: `cuda-tile` (v1.6.0rc3) was included because `sglang` (v0.5.16)
depends on `flashinfer-python` (v0.6.14) which depends on `cuda-tile`
This is the failure mode requirements-cublas1{2,3}-after.txt already
carries an nvidia-modelopt bound for, arriving through a different
package. install.sh passes a global --prerelease=allow, which is
load-bearing for flash-attn-4, so an unbounded dependency resolves to a
prerelease; cuda-tile 1.6.0rc3's build backend imports wheel_stub without
declaring it in build-system.requires; --no-build-isolation means nothing
provides it, and the build dies.
Nothing in this repo changed. cuda-tile published 1.6.0rc1 and rc3 and
the weekly cron picked them up, which is the drift that job exists to
catch.
Bound the one package rather than dropping the global flag, matching the
existing precedent. 1.5.0 is the newest stable release, so <1.6 takes the
last good one. l4t13 gets the same bound: it installs plain sglang rather
than sglang[all], but flashinfer-python is a dependency of both.
NOT VERIFIED LOCALLY: reproducing this needs a CUDA docker build, which
this machine cannot run. The diagnosis is from the CI log and the
resolver's own hint, and the change follows a fix already proven in these
same files. CI on this PR is the check that matters.
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
|
||
|
|
a0f7faaa2a |
fix(sycl): stop building the ggml CPU variant matrix with icpx (#11321)
Since #11255 and #11276 every GPU image also builds ggml's CPU_ALL_VARIANTS matrix, so a partial offload uses the host's SIMD kernels. That works everywhere except SYCL, where the Makefile compiles the whole tree with icpx -fsycl: icpx never finishes ggml-cpu/arch/x86/repack.cpp at -march=sapphirerapids. In run 30765516644 both sycl_f16 and sycl_f32 stopped at that translation unit and sat there for 5h30m with a single compile in flight until GitHub killed the job at its 6h limit, and turboquant's f16 job lost its runner outright. gcc compiles the same file in seconds in the vulkan and CPU jobs of the same run, so the CPU variant matrix is only unbuildable under icpx. Route SYCL back to the portable fallback binary, which is what these images shipped before #11255. run.sh already prefers *-cpu-all when present and falls back otherwise, so nothing else has to change. Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
133c546c3f |
feat(api): add text moderation endpoint (#11316)
* feat(api): add text moderation endpoint Add an OpenAI-compatible /v1/moderations endpoint backed by constrained local text generation. Register its auth and discovery surfaces, document the text-only MVP, and cover response shaping and access control. Assisted-by: Codex:gpt-5 * test(mcp): update assistant client stub Keep the LocalAI Assistant holder test stub aligned with the scheduling methods added to LocalAIClient so repository-wide type checking succeeds.\n\nAssisted-by: Codex:gpt-5 --------- Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
8a68f3571c |
feat(api): add POST /v1/images/upscale endpoint (#10227)
* feat(api): add POST /v1/images/upscale endpoint Add a new image upscaling endpoint that accepts a source image and returns an upscaled version. Supports selectable upscaler models (e.g. realesrgan) and a configurable scale factor (2x or 4x). - backend.proto: add UpscaleImage RPC and UpscaleImageRequest message - pkg/grpc: implement UpscaleImage in Backend interface, client, server and embed shim - core/backend/upscale.go: new backend helper (mirrors ImageGeneration) - core/http/endpoints/openai/upscale.go: new multipart/form-data handler - core/http/routes/openai.go: register POST /v1/images/upscale - core/http/auth/features.go: gate upscale routes under FeatureImages - backend/python/diffusers/backend.py: implement UpscaleImage — uses diffusers upscale pipeline when loaded, falls back to Lanczos resize * fix(grpc): add UpscaleImage stub to Base backend All Go backends embedding Base now satisfy the AIModel interface without needing to implement UpscaleImage explicitly. * fix(images): complete upscale endpoint integration Store generated upscales under the served images directory, validate scale factors, document and advertise the endpoint, and add a functional Stable Diffusion x4 gallery model. Assisted-by: Codex:gpt-5 --------- Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
fd4ec083b9 |
feat(downloads): add resume-safe pause action (#11222)
Give gallery operations distinct pause and cancel paths. Pause preserves partial download data so reinstalling the same model or backend resumes through HTTP Range, while cancel keeps its destructive semantics. Surface the action in the Activity UI and document the API behavior. Assisted-by: Codex:gpt-5 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
8f74f74b10 |
feat(mcp): expose scheduling admin tools (#11228)
* feat(mcp): add scheduling client contracts Assisted-by: Hephaestus:openai/gpt-5.5 Signed-off-by: Owen Adirah <owenadira@gmail.com> * feat(mcp): add scheduling HTTP client support Assisted-by: Hephaestus:openai/gpt-5.5 Signed-off-by: Owen Adirah <owenadira@gmail.com> * feat(mcp): add in-process scheduling stubs Assisted-by: Hephaestus:openai/gpt-5.5 Signed-off-by: Owen Adirah <owenadira@gmail.com> * feat(mcp): register scheduling tools Assisted-by: Hephaestus:openai/gpt-5.5 Signed-off-by: Owen Adirah <owenadira@gmail.com> * test(mcp): map scheduling tools to REST routes Assisted-by: Hephaestus:openai/gpt-5.5 Signed-off-by: Owen Adirah <owenadira@gmail.com> * docs(mcp): document scheduling assistant tools Assisted-by: Hephaestus:openai/gpt-5.5 Signed-off-by: Owen Adirah <owenadira@gmail.com> * fix(mcp): wire in-process scheduling Use an explicit MCP scheduling DTO and route in-process scheduling calls through the distributed node registry so the embedded assistant matches the REST scheduling surface. Assisted-by: Hephaestus:openai/gpt-5.5 Signed-off-by: Owen Adirah <owenadira@gmail.com> * fix(mcp): narrow scheduling dto Assisted-by: Hephaestus:openai/gpt-5.5 [opencode] Signed-off-by: Owen Adirah <owenadira@gmail.com> --------- Signed-off-by: Owen Adirah <owenadira@gmail.com> |
||
|
|
cd62e8ff18 |
gallery: add Nemotron 3 embedding models (#11314)
Add multilingual 1B and 8B Q4_K_M GGUF embedding entries and link them as variants for automatic memory-aware selection. Assisted-by: Codex:gpt-5 [web] Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
af98e76f84 |
fix(gallery): remove broken DeepSeek V4 0731 entry (#11313)
fix(gallery): repair DeepSeek V4 0731 entry Use the official single-file ggml-org MXFP4 artifact with its verified SHA256 and route it through llama.cpp instead of treating an unsloth repository page as a ds4 model file. Assisted-by: Codex:gpt-5 [Hugging Face API] Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
7f9ffd9f54 |
gallery: add AMD Instella MoE 16B variants (#11308)
Add Q4_K_M and Q8_0 GGUF builds for the trending Instella-MoE-16B-A3B-Think model, with host-selectable variant metadata and verified Hugging Face LFS hashes. Assisted-by: Codex:gpt-5 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
5cb0c1a872 |
feat(ui): close the gap between the shipped UI and the design mocks (#11307)
* feat(ui): give the Operate overview real numbers and traces a latency shape First two items from a component-by-component comparison against the mocks. The pattern that audit found: everything newly built matched, everything pre-existing got the palette but not the layout, and an "absent rather than empty" rule hid most of the overview exactly when someone was looking at an idle installation. **The headline grid is always rendered**, including at zero, with a fourth cell for host memory. Hiding it removed the page's structure precisely when it was most likely to be read, and "0 failed" is information — an absent panel is not. The quiet case is now said in a line underneath instead of by showing nothing. **The sections state counts** rather than listing their destinations: backends, models, updates and running operations instead of the words "Usage and traces". That needed installed backend and model counts in the summary context, which are two more cheap reads on the poll that was already running. **Traces rows carry latency as a bar as well as a figure**, scaled against the slowest request currently in view and turning amber past two seconds. The table had no latency column at all — the number was buried in the expanded detail, so the shape of the tail was invisible while scanning. Scaling against the view rather than an absolute ceiling is deliberate: what matters when reading a page of traces is which of these are the outliers, and an absolute scale flattens every row on a fast installation into nothing. Full e2e suite: 409 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): name the engine on Home's resident models, and add jump-back-in Third item from the mock comparison. The mock showed each resident model with the engine serving it. /system carried only the id, so the audit recorded this as blocked on a server field — but the config loader is already in scope where that response is built, so it is one lookup. SysInfoModel gains an optional `backend`, resolved from the model's config and omitted rather than guessed when there is none (a loose file, or a config since removed). Home renders the column blank in that case; the test pins both halves of that. Memory per model stays out. It is not one lookup — it would mean asking each backend process — and inventing a number beside a real one is worse than leaving the column off. "Jump back in" is the block the mock had and Home did not. The quick-links row above it is a set of first-run actions; these are the three places someone returns to, each stated with what it currently holds rather than as a bare label. Go: routes suite passes. Full e2e suite: 412 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): rank the recommended models as lanes instead of equal cards The hardware recommendations were a grid of equally-weighted cards. The list is already sorted by fit, and a grid throws that order away: three cards side by side say "pick one", when the page has actually formed an opinion about which one. They are lanes now, read top to bottom in fit order, with the leader carrying the single amber "Best fit" label and the rest marked "Also fits". One opinion per page — the alternatives are alternatives, not runners-up each worth their own colour, which is how a strip of coloured badges ends up meaning nothing. Below 720px the size and VRAM columns drop and the lane keeps the name and the install action, which are the two things a narrow screen needs. The existing panel spec moves off .rec-models-item onto .lane rather than being deleted; dismissal, collapse, keyboard operation and install all still pass unchanged, and there is a new assertion that exactly one row is called out. Full e2e suite: 413 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): drop capsule chips app-wide, and un-break the empty voice library **Pills are gone.** A capsule radius reads as a tag floating on the surface, which fights a system whose structure is hairlines and square corners — and with chips on Discover, Host, Activity and the biometrics pages, "some pages have pills" was the real inconsistency rather than any one page. Sixteen selectors move to the small radius: filter buttons, tab pills, activity and biometrics chips, file and count badges, the jump-to-latest control, the nav badge. Round *buttons* keep their circle — .lightbox__nav and .home-send-btn are circles, not capsules — as do every progress track, status dot and avatar, which are round because they are round, not because they are tags. **The empty voice library was unusable.** `.voice-library-empty` sets min-height: 430px, border: 0 and background: transparent — a description of the empty PANEL — and it had been attached to the action instead. The create button was therefore a 430px transparent box that pushed itself out of the panel and could not be seen. Moved onto the container it describes, which now centres its action rather than letting it fall off the bottom. Same class-mangling shape as the Agents header fixed earlier. Full e2e suite: 416 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): put Host's headline figures on the shared hairline strip Host had shadowed, clickable StatCards above a page that already has a rail, a pane and a tab bar — a second dashboard language on one screen, and a different one again from the figures inside its own detail pane. The Operate overview's figure grid is generalised into a shared `.stat-strip` and Host adopts it, so the two pages read as one system: same cell, same figure scale, same tone vocabulary, and the same hairline grid the split-view StatGrid already uses. The cells stay clickable and still route into the tab and filter they describe, because a count is worth more when it is also the way to the thing counted. Tone is spent only where the number means something — running and updates when non-zero — since a strip where every cell is coloured has no emphasis left. Two bugs made on the way, both now covered: - The first version put `<button>` elements inside a `<dl>` with `<dt>`/`<dd>` inside the buttons. Neither is valid, the browser re-parents both, and the cells collapsed. These cells are a set of controls, so a plain container of buttons is also the honest markup. - Even correct, the strip rendered 2px tall: `.page--app` is a flex column whose split view takes flex:1, so a child with no intrinsic minimum is shrunk away. The old cards survived only because `.stat-card` carried min-height:96px. The strip now declines to shrink, with a test pinning it. The stat-card specs are retargeted rather than deleted: they were written to guard a class collision on a page that no longer uses cards, so they now guard the strip's labels and its height. Full e2e suite: 417 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): make Backends notices an edge rather than a filled card The install and upgrade banners were tinted cards with a full border. A filled panel makes every notice shout at the weight of an error, which is how notices stop being read — and Backends shows one on most visits, so it was shouting routinely. They are now a hairline with a coloured left edge, the same treatment the Operate overview gives rows that want a decision, so "this needs you" looks the same wherever it appears. Counts in the notice take the monospace tabular figures the rest of the console uses. Also drops the last inline style on the page, and refreshes the inline-style baseline, which has read 624 against a real count since #11288 landed. The gate exits 0 either way, so nothing was failing — but a baseline 86 above the truth would have let that many inline styles back in unnoticed. Now at 538, which tightens the ratchet rather than loosening it. The spec creates the upgrade it asserts on rather than skipping when the mock has no notice: a test that skips is a test that proves nothing. Full e2e suite: 418 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): finish the mock parity list, and stop hiding the recommendations The last two items from the audit, plus a correction. **Discover's use-case shelf is lanes.** These are a list of ways in, read in order; a grid of equal cards asks the reader to compare them, which is not the choice on offer. **The request panel reaches every generator.** Video, 3D, Sound and Audio FX join Images and Speech, so each one teaches its own endpoint rather than two of six doing it. Audio FX records the fields that shape the request rather than the bytes, since its payload is multipart. **Recommendations no longer collapse themselves.** They were folded away by default once anything was installed. That is the page's one opinion about this host, and an opinion hidden by default is one the reader never gets. Someone who disagrees can still collapse it and that choice is remembered — the difference is that we no longer make it for them. Three specs asserted the old default and now assert the new one. The use-case heading also sat a line's width from the text it introduces, so the two read as one paragraph. It has air under it now, and the shelf is separated from the recommendations above it. Two tests removed rather than kept: a generator loop whose only real assertion was `expect(endpoint.length).toBeGreaterThan(0)`, and an earlier card-gap guard that could only skip. A test that cannot fail is worse than no test, because it reads as coverage. Full e2e suite: 418 passed, 4 skipped. Inline styles at baseline. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): make the Host figures legible and give the strip its spacing back Three defects introduced by the Host redesign, all found by looking at the running app rather than by the suite. **The figures were invisible.** "Running now" and "Updates available" rendered pure black on the dark ground. Two causes compounding: the `--muted` tone alias never landed, because the source rule has extra spaces before its brace and the exact-match edit missed it silently; and a `<button>` does not inherit colour, so with no tone rule the value fell back to the user agent's `buttontext`. Both fixed, and a test now fails on any figure computing to pure black. **The strip sat flush against the resources panel.** `.stat-strip` declares `margin: 0 0 ...` and is declared later in the file than `.manage-summary`, so the shorthand quietly won and the top margin became zero. Raised to `.stat-strip.manage-summary` so it beats the shorthand on specificity rather than on declaration order, which is the kind of thing that breaks again the next time a rule moves. **Discover's use-case heading had a doubled gap.** `.zero-pane` is a flex column that already separates its children; adding a margin on top of the gap stacked the two. The margin is gone and the heading keeps only its own breathing room. Full e2e suite: 420 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): make Studio's tabs path segments rather than a query parameter `/app/studio?tab=images` reads like a filter applied to a page. It is navigation: a different generator, with its own state and its own deep link. It is now `/app/studio/images`, with the overview at `/app/studio`. Legacy `?tab=` links are redirected once to the path form, replacing the history entry so Back does not bounce between two spellings of the same place. Bookmarks and older links keep working and land on the canonical URL rather than a second version of it, which is the part worth having a test for. The nine `?tab=` references were all in specs, none in docs, so the migration is contained. They move to paths, and a new spec pins the redirect. Full e2e suite: 421 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): make the hardware recommendation a section, not a dismissable card It was a bordered card with a collapse control and a close button, sitting inside a pane that is otherwise hairline sections. Two problems: it read as something bolted onto the page rather than part of it, and treating it as an interruption to be shut is the wrong frame for the one thing the page has to say about the machine it is running on. It is now a plain section with the same heading treatment as the shelves below it. The collapse state, the dismissal, their storage keys and the legacy key read for backwards compatibility all go with it, along with the installedCount prop that existed only to pick a default collapse. Five specs described behaviour that no longer exists and are removed rather than adjusted — collapsing, dismissing, persistence of both, and the toggle's keyboard handling. One new spec asserts the replacement contract: no control with aria-expanded, no dismiss, and no card border. Full e2e suite: 416 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): restore every stripped icon and every default-chrome control You reported two broken icons. They were not two: an earlier automated edit had stripped the `fa-*` class from twenty `<i>` elements across eleven pages, and an `<i>` with no icon class renders nothing at all. Settings' save button, the voice-profile back link, and eighteen others — agent row actions, task and job buttons, import and create actions — were all drawing empty space. Each is restored from its own context rather than a blanket icon: the agent row gets pause/play, pen, comments, file-export and trash; the fine-tune toggle swaps plus for xmark as it opens; the P2P documentation link gets the external-link glyph. The same edit left controls without their classes. Fine-tune's "Import config" was rendering in the browser's own chrome, and `.p2p-cmd__copy` set a border but no background, so it fell back to `buttonface` — a pale grey chip on a dark command block. FineTune's "New job" also had its icon classes folded into the button's className, the same mangling already fixed on the Agents header. Rather than fix the reported two and wait for the next report, this adds a standing audit: twenty-five routes are walked and the test fails on any visible control rendering with user-agent chrome, or any `<i>` without an `fa-*` class. It found the three remaining cases after the first sweep, and it is the reason the next one cannot ship quietly. Full e2e suite: 417 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
cd890b6a26 |
chore: ⬆️ Update leejet/stable-diffusion.cpp to db99efdd6d2a43c7937fd55b3359206c680a75b0 (#11299)
⬆️ Update leejet/stable-diffusion.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
1c0380ad44 |
chore: ⬆️ Update 0xShug0/audio.cpp to 5a8312ef7b8aa7cf14e9a24ac568cabd8725d68a (#11302)
⬆️ Update 0xShug0/audio.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
cb6e4d4391 |
chore: ⬆️ Update CrispStrobe/CrispASR to fcb79282a6bc52e13d858026c42b24fb6e63c97a (#11304)
⬆️ Update CrispStrobe/CrispASR Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
cba54c5ea1 |
gallery: add grug-27b GGUF variants (#11311)
Assisted-by: Codex:gpt-5 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
f951419207 |
chore(model-gallery): propose variant groupings for review (#11312)
chore(model-gallery): propose variant groupings Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
58ea2f5d79 |
feat(ui): give Operate and Studio a front door, and fix two layout regressions (#11305)
* feat(ui): give Operate a front door and fold six rail groups into four Opening Operate ran firstVisiblePath() and landed on Backends, because Backends happens to be written first in operateConsole.groups. The section that should answer "is anything wrong" opened on a package manager, and nothing was reported until you visited it. Adds /app/operate. Its one irreplaceable block is "Needs attention", which is empty when nothing is wrong and says so in a line rather than rendering a reassuring green panel. It collects stale backends, failed operations and unhealthy nodes. Everything else on the page is a summary you could already assemble by visiting four others. The rail regroups from six headings to four: Inference and Activity were both "the runtime right now", Access and System were both administration. No destination is removed and no gate changes, so isConsoleItemVisible and consolePaths are untouched. Overview leads the first group, which is what makes firstVisiblePath() return it without knowing it exists. Rail entries now carry a signal beside the label. This does not replace the sidebar badge and is not built as if it does: the badge stays on the always-visible sidebar entry for the reason recorded in Sidebar.jsx, that the rail exists only on Operate routes and can be collapsed. The signals are orientation while inside Operate, so they are aria-hidden and nothing urgent depends on them alone. OperateSummaryContext polls once for the whole console, following OperationsContext, which exists because per-consumer setInterval against one endpoint was the defect it fixed. It is mounted by ConsoleLayout for the Operate console only, so "poll only while in Operate" needs no route check. Built on usePolling, so it pauses on a hidden tab. Operations are read from OperationsContext rather than polled a second time, and each source degrades to no-signal on its own so one dead endpoint cannot blank the rest. It reads the cached GET /api/backends/upgrades and never the POST that forces a real registry check. Traces and Usage get no signal yet: /api/traces returns the list, so a count would mean fetching every trace to render one number. A counts endpoint is the honest fix and is scoped separately. Full e2e suite green (369 passed, 4 skipped), including a render-smoke entry for the new route. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): open Studio on what this machine can actually make Studio was a tab strip over six generators that opened on Images, which was never a decision, only the first entry in BASE_TABS. Nothing said which modalities this installation could run, so the way to learn that video had no model was to pick the tab and find an empty select. Adds an overview tab and makes it the fallback. Explicit tabs still win, so existing deep links keep working; anything unrecognised or gated now lands on the overview rather than Images. Each tab carries a dot: filled when an installed model advertises that modality, hollow when nothing serves it. That is the feature in one detail, turning the strip from navigation into a report of what the machine can do before anything is clicked. The dot is aria-hidden because the overview states the same facts in words and the dots change as models load. Two kinds of unavailable, which had to stop looking alike: - switched off, via a permission: no tab and no lane, unchanged - available with no model: a lane, and a route to installing one Studio now owns one MODALITIES table so the tab strip and the overview cannot disagree about what exists, and calls useModels() once, unfiltered, grouping in the browser. useModels(capability) fetches the whole list and filters locally, so a hook per modality would have been six identical requests to /api/models/capabilities on every mount. There is a test for that. Recent outputs read every localStorage store through a new readAllMediaHistory(), which avoids mounting five hooks that carry save timers the overview has no use for. 3D is read separately through use3DHistory rather than folded in: its entries are GLB blobs in IndexedDB, so they cannot come from the same synchronous read. Typical cost is the median of this machine's own history, not a guess, and renders as a dash when there is nothing to go on. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): stop the stat cards and the console rail breaking on small screens Two unrelated causes behind one report that /app/manage looks wrong when the window is narrow. The stat cards were being laid out by the wrong rule at every width. Two different components both claimed `.stat-grid`: the dashboard card strip that holds .stat-card children, and the detail-pane StatGrid the split views introduced further down App.css. Being later, the second won every shared property, so the cards got its 120px columns and its 1px hairline gap in place of their own 180px columns and spacing-md. Four cards were packed onto a row that fits two, labels wrapped to three lines and clipped, and the icon crowded the value. Renamed the strip to `.stat-cards`, after the children it actually holds, which also removes the mismatch of a `.stat-grid` container full of `.stat-card`s. The split-view component keeps `.stat-grid` and its BEM parts. The expanded console rail had no bounded height. Thirteen destinations stacked in one column is taller than a phone, so opening the menu pushed the page's own heading past the fold: the menu replaced the page rather than annotating it. Capped at 55vh with internal scrolling below 768px, so the content behind stays reachable. Both are asserted on behaviour rather than markup: no stat-card label may be clipped, the card gap must not be the detail pane's hairline, and expanding the rail must leave the page heading on screen. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): retemper the palette to localai.io and add the lane primitive The token half of the style transfer, plus the shared list idiom the two overviews had each grown their own copy of. theme.css moves from Nord to the website's palette, variable names preserved so every consumer moves with it: ground #13171f -> #0d1117, accent frost cyan #88c0d0 -> action blue #4f8cff, success sage -> mint #56d6a4, warning -> the amber #f1b95d the site spends only on the thing asking for a decision. Eyebrows go mint. Dividers become an opaque #29384a hairline rather than alpha over a varying surface, which is what makes stacked surfaces read crisply on the site. Light is derived, not inverted. The site ships one theme and never had to answer this, but the app does: blue darkens to #2f62d8, mint to #0d8b60 and amber to #8a5d0b, all clearing 4.5:1 on a cool paper ground, where the dark-mode values sit near 2:1. Same three roles, different values. Three files restate the palette because CSS variables cannot reach them: cmTheme.js (the whole CodeMirror theme), VoiceVisualizer and WaveformPlayer (canvas). Left alone they would have quietly kept the app half-Nord. The `.lane` primitive replaces the near-identical row CSS that OperateOverview and StudioOverview had each written: a full-bleed row on a hairline that insets on hover, with no card and no shadow. Callers supply only the column template. Both pages now use it, along with `.lane-head` for section rhythm and a `.page-pad` container for top-level pages outside a console shell — without which Studio sat flush against the sidebar with its eyebrow clipped. Studio's tab strip wraps rather than running off the edge at narrow widths. Full e2e suite: 386 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): put Home's resident models on lanes and give the footer one line Home's status line was three chips saying a thing was true. It now reports figures: how many models are resident, how many nodes are healthy, what share of memory is in use, set in tabular monospace so the digits line up. A chip answers whether; a figure answers how much, which is what someone opening the page at a glance is after. Resident models move from status chips to lanes, with the id set in a new `.lane__name--id` because an id is something you might type or paste and the UI face makes it read as a label. /api/system-information carries only the id, so there is deliberately no backend or memory column: inventing one would mean a server change this does not make. The footer was three centred rows and cost the bottom sixth of every page for chrome. It is one line now, version left and links right, wrapping to centred when the viewport is too narrow to hold both. Every link it had, it keeps. Full e2e suite: 392 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): correct three contrast failures and stop a guaranteed-404 poll A contrast audit of the new palette found three values below WCAG AA, one of which the previous commit message claimed was fine: - White on the #4f8cff button is 3.22:1, which is large-text only. The website does exactly this, but a button label in an app is not large text, so the label goes to dark ink at 5.88:1. Light mode keeps white, which is 5.44:1 on its darker blue. - Light-mode success was 4.08:1 on paper, not the 4.5 claimed. Darkened to #0a734f, 5.56:1. - Nord red was already 4.28:1 on raised surfaces, a pre-existing miss carried over unexamined. Lifted to #c96f78, 5.02:1. Lanes gain the two states they were missing: a 44px target on coarse pointers, matching what EntityRail already does so the two list idioms feel the same under a thumb, and a reduced-motion variant that keeps the background feedback while dropping the hover inset, which is a position change. The Operate summary no longer asks for /api/nodes on a single-node install. The cluster API answers 503 when distributed mode is off, so it was a guaranteed miss every fifteen seconds; it is now gated on useDistributedMode, the same condition the rail already uses for the Nodes entry. Full e2e suite: 392 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): restore the gap between overview blocks, and stop claiming zero nodes Two defects a design review surfaced. `.lane-head:first-child { margin-top: 0 }` was meant to stop the first block on a page carrying a top margin. But every <section> makes its lane-head a first child, so the reset applied to all of them and the gap between blocks vanished: "Sections" sat flush against the attention row above it. The header supplies its own bottom margin, so a uniform top margin is correct everywhere. The Cluster summary read "0 nodes" on a single-node install, which looks like a fault when the cluster API is simply switched off. It now says "Single node". Full e2e suite: 392 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): open dark by default, and stop clipping the collapsed sidebar footer Dark is the identity rather than a preference: localai.io ships one theme and it is this one, so an install should look like LocalAI before anyone has chosen anything. The OS setting no longer selects light on first load. The toggle still does, and a stored choice wins forever after, which the tests assert both ways. The collapsed sidebar footer stacked its controls but kept the expanded row's inline padding, so their edges were clipped against the 51px rail. Full e2e suite: 394 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(api): count traces server-side and give the Operate overview real totals The overview's headline block had no source. /api/traces returns the trace list, so "37 errors in 24h" meant fetching every buffered exchange to count it in the browser — waste that grows with the buffer, to produce three integers. Adds GET /api/traces/summary: totals, failures, p95 and a bucketed series for sparklines, over a window that defaults to 24 hours and is capped at a week. Deliberate calls, each with a spec: - A 4xx is the caller getting it wrong, not the installation being unhealthy, so only 5xx and transport errors count as failures. - p95 is a nearest-rank percentile rather than the slowest request, which is what a max would report and what makes latency panels lie. - Buckets are oldest-first so a sparkline reads left to right, and the slice is never nil: nil serialises as null and breaks .map() on the other side, which is a silent runtime error rather than an empty chart. - Exchanges outside the window are not counted at all. The route is registered before /api/traces/:id so "summary" is not captured as a trace ID. On the client, Traces and Usage gain the rail signals they were shipped without, the Observability section summary now states counts instead of listing its destinations, and an installation that has served nothing says so rather than showing three zeroes dressed as telemetry. Sparkline is a bare stroke with an emphasised endpoint and no axes: the figure above it already states the value, so its only job is the shape. Go: 185 middleware specs pass. Full e2e suite: 396 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): stop the memory chart calling a trade-off an error The VRAM-by-context chart rendered any build over the limit in error red, and escalated the verdict to the error tone as soon as two context sizes crossed it. But an over-limit build still installs — #11288 keeps a test on exactly that — so red overstates what is happening. A model that fits at 32k and not 64k is a trade-off, not a fault. Over-limit bars and the limit line now use the warning tone, which is the constraint colour used everywhere else in this branch: know what you are doing, not you may not. The error tone is reserved for "fits nowhere", where the model genuinely cannot run on this host. Full e2e suite: 397 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): give the new surfaces orchestrated motion Uses the reveal system already in the codebase rather than adding a library: pageReveal, .reveal-stagger and staggerStyle() were built for exactly this, and anime.js would be ~17KB duplicating four lines of CSS for list reveals. The overview's headline figures, attention rows and section lanes stagger in, as do Studio's modality lanes and recent outputs, so a page assembles in the order it is read instead of appearing all at once. Two additions beyond stagger. Rail signals transition on opacity when a poll lands, so a number changing reads as an update rather than a jump cut, and it stays on the compositor so it cannot reflow the rail. The attention block animates its left edge in — the one thing on the page that should announce itself, and on the border rather than the text so nothing moves under a reader. Both are dropped entirely under prefers-reduced-motion, alongside the lane hover inset already handled. Full e2e suite: 397 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): put the generators on a hairline field stack and record the request The workbench treatment from the mocks, applied where it costs least: both changes land on shared surfaces, so all six generators get them at once rather than drifting apart page by page. The control column stops being a shadowed card of boxed groups and becomes a hairline field stack — the panel is the page's left half, not an object floating on it — with uppercase micro-labels matching the eyebrow treatment used elsewhere. Because .media-controls is shared, Images, Video, 3D, Speech, Sound and Audio FX all move together. RequestPanel shows the request the form actually built, with a copy-as-curl. LocalAI is API-first and Studio is the best place in the app to teach its own endpoints: the form stops being a black box, and a result worth keeping can be reproduced from a shell without reverse-engineering which fields the page sent. It records what was sent rather than what the form currently holds, and renders nothing until a request has been made — a panel describing a request nobody made is a tutorial, not a record. Wired into Images and Speech. Full e2e suite: 401 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): make Chat a transcript instead of a bubble thread Rounded, filled, asymmetric bubbles fight a system built on hairlines, and they carry the speaker in shape and side rather than in words. The assistant side had already given up its bubble; this finishes the job. Both roles now run full width down one column, separated by a rule, each with a mono role label. The user turn keeps a left edge in the action tone so the two are still told apart at a glance, without a fill or a corner radius. The avatars go: the accent and the label carry the speaker, so the glyph was decoration once neither side had a bubble. Saying who is speaking in words rather than in geometry is also what survives being read aloud, printed, or looked at by someone who cannot pick the sides apart by colour. Full e2e suite: 404 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * feat(ui): dress the API reference in LocalAI's palette The Swagger page was the last surface still shipping in someone else's colours, which is conspicuous now that everything it links from is dark. Swagger UI has no theming hook, so rather than fork it we serve our own index ahead of the library's wildcard and restate the palette over its stylesheet. The library's own bundle and assets are still what load, so a swagger-ui upgrade cannot silently break the page — this is a skin, not a fork. Two things needed real care. Swagger tints the entire operation row per method via .opblock.opblock-post and friends, so the palette had to match that specificity rather than reach for !important; the method now lives on one edge instead of washing across the row, because a page where every row is a status colour has no status colour left. And the filled method chip put white on pale green, which was the least readable thing on the page — it is an outlined mono chip now, carrying the method in its border and text. Palette values are copied from theme.css rather than referenced: this page is served by Go and never sees the app's CSS. The comment says so, and says to keep them in step. Go: routes and middleware suites pass. Full e2e suite: 405 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] * fix(ui): make tall split-view pages reachable, repair the Agents header, scale titles Three things found by actually using the app rather than measuring it. **Host was unusable.** The shell above a split view is overflow:hidden so the document cannot grow, which left anything taller than the viewport simply unreachable — and Host stacks a resources card, four stat cards and a tab bar above its split, so the bottom of the pane fell off at every window height with nothing to scroll. Every sweep I ran for this was horizontal, which is why it kept coming back clean. The page now scrolls inside the pinned shell. The pane keeps its own scroller: letting it grow instead pushes the document taller and stretches the rail to match, which is the regression e2e/discover-height.spec.js exists to catch, and which the first version of this fix duly caused. **The Agents header controls were unstyled** — "Create Agent" was rendering with the browser's default chrome. The markup had been mangled at some point: six unrelated classes merged into one string on the link, and the label and button left with none at all and empty icons. Repaired, with the inline flex replaced by a shared .header-actions class. **Page titles take the editorial scale from the site**: larger, tracked at -0.04em, on a line height near 1, so a two-word title reads as a statement rather than a label. The typeface is unchanged — DESIGN.md keeps the existing type system — so the whole difference is scale, tracking and leading, which is where the site gets its voice from. This was the biggest reason the running app still did not look like the mocks. Full e2e suite: 404 passed, 4 skipped. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5[1m] [Read] [Edit] [Bash] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
b89b0f73e5 |
chore(model-gallery): ⬆️ update checksum (#11306)
⬆️ Checksum updates in gallery/index.yaml Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
45cd47cb99 |
chore: ⬆️ Update ikawrakow/ik_llama.cpp to cb9147fd0d9c08a9a84eee5ac405a73f4e10e3e1 (#11300)
⬆️ Update ikawrakow/ik_llama.cpp Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: mudler <2420543+mudler@users.noreply.github.com> |
||
|
|
1aa97381f3 |
perf(gallery): warm variant descriptions alongside VRAM estimates (#11297)
Follow-up to #11288, which warmed the VRAM estimate caches at startup and left the variant picker paying its own way. Describing an entry's variants probes the weight files of every build it offers, so the first time a model is opened costs 1.2-1.9s against a cold cache. That is the same cost as an estimate wearing a different hat, and it lands in the same caches underneath, so it belongs in the same pass rather than in a second mechanism. The warm-up now describes variants for the entries it walks. Entries that declare none cost nothing: the call is gated on HasVariants rather than attempted and discarded. The host resolve env is derived once for the run, since it describes the machine rather than the entry. Failure handling matches the estimate half. An entry whose variants cannot be described is logged at debug and skipped, and the estimate for that same entry is unaffected, because neither half is allowed to fail the other. Measured against a live instance with 1,595 models, first ever call to /api/models/variants/:id after a cold boot: before 1.2-1.9s after 2ms The warm-up's own cost barely moves: 3m0s to 3m19s for 300 entries, of which 40 declared variants. It stays bounded by the same knobs, and LOCALAI_VRAM_WARM_LIMIT=0 still turns the whole thing off. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |