Compare commits

..

348 Commits

Author SHA1 Message Date
Ettore Di Giacinto
1df5a3ef7f fix(model-artifacts): keep the companion option present on a remote load, even unresolved
Follow-up to the companion-persistence fix. On a live distributed cluster the
managed base_model companion option still failed to reach the remote worker
(nvidia-thor) even with the companion resolved-and-persisted in the config:
the backend logged "Downloading required files for meituan-longcat/LongCat-Video"
and failed "base_model must point to a LongCat-Video checkpoint". Symlinking the
companion into place did not help (the option was simply absent from the worker's
LoadModel), while an explicit absolute base_model in options: worked as a control.

Trace of where the remote ModelOptions is built and whether the companion is
present there:

- The *pb.ModelOptions the worker's LoadModel consumes is built on the
  CONTROLLER by grpcModelOpts (core/backend/options.go) -> withCompanionArtifactOptions,
  set as gRPCOptions, and sent by direct gRPC via FileStagingClient.LoadModel. It
  is NOT rebuilt on the worker. The reconciler's replica scale-up instead replays
  a Postgres-stored proto blob, which already carries whatever grpcModelOpts
  produced.
- withCompanionArtifactOptions is the ONLY builder of ModelOptions.Options in the
  tree, and it emits base_model iff the config's companion artifact has
  Resolved != nil. Staging preserves the option and derives the worker ModelPath
  as the nested per-model staged root, so a resolved companion resolves under it
  without a download (verified end to end; the path-nesting angle is a red
  herring here).

So the option is absent only when the config the loader is serving from carries
the companion WITHOUT a resolved snapshot (its resolved state not reaching the
serving config, e.g. a peer-replica reload from local disk or a config loaded
before resolution). In that state the old code emitted NOTHING for the companion,
and longcat-video fell back to its OWN hardcoded default (BASE_MODEL_ID), which
is exactly the observed download-and-fail.

Fix: an unresolved-but-declared companion no longer vanishes. It now falls back
to its DECLARED source repository id, so the backend fetches the artifact the
config actually asked for instead of a hardcoded default; the resolved snapshot
path (the staged, no-download fast path) is still preferred whenever the
companion is resolved, so the single-node and healthy distributed paths are
unchanged. The fallback logs a warning naming the artifact and repo, and the
router now logs the exact option strings crossing to the worker at debug, so a
recurrence is diagnosable in one load instead of by inference.

Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-23 16:34:56 +00:00
Ettore Di Giacinto
bb2ed02cca fix(model-artifacts): persist companion artifacts, not just the primary
A managed model can declare companion artifacts (LongCat-Video-Avatar-1.5
pulls its tokenizer, text encoder and VAE from the separate LongCat-Video
base repo via a target: companion artifact). preloadOne resolves the whole
set in memory, but the binding written back to disk carried only the
primary: persistArtifactBinding marshalled []Spec{result.Spec} and replaced
the entire artifacts: list with it, silently dropping every companion.

In a single process the loss is invisible because the in-memory config keeps
the companion. It bites on the next controller restart: the config reloads
from the mangled file with the primary alone, so withCompanionArtifactOptions
finds no resolved companion and synthesizes no base_model option. The remote
longcat-video backend then never receives base_model, falls back to
BASE_MODEL_ID and downloads the repo itself ("Downloading required files for
meituan-longcat/LongCat-Video"), failing the load with "base_model must point
to a LongCat-Video checkpoint".

This is why an explicit base_model:<path> added to the config options works
where the managed companion does not: an explicit option lives in options:,
which is never rewritten, while the managed companion lives in artifacts:,
which the binding overwrote.

Persist the full resolved set (primary + every companion), and widen
bindingNeedsPersistence to compare the whole artifact list so a companion
resolving for the first time still triggers a write. The single-node path is
unaffected: there the in-memory config already carried the companion, and the
staging/ModelPath resolution for a remote worker (nested per-model staged
root, #10949) is unchanged and already correct once the option is generated.

Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-23 14:18:25 +00:00
Tai An
aae69b1163 fix(ace-step): drop nonexistent Get* proto accessors in SoundGeneration (#11069) (#11072)
The Python gRPC bindings expose message fields as plain attributes
(request.language, request.caption), not Go/Java-style Get*() accessors.
Because request.language is an empty string when unset, the

    request.language or request.GetLanguage() or "en"

expression falls through to request.GetLanguage(), which does not exist
on the generated Python message and raises AttributeError, surfaced to
clients as:

    rpc error: code = Unknown desc = Exception calling application: GetLanguage

Every /v1/sound-generation request without an explicit language field
failed. Drop the bogus accessor calls (TTS already uses the plain-field
form a few lines below).

Signed-off-by: Tai An <antai12232931@outlook.com>
2026-07-23 15:00:26 +02:00
mudler's LocalAI [bot]
8d6fdf22d3 fix(backends): derive the protoc generator from the protobuf runtime, regenerate stubs after late installs (#11057)
* fix(backends): choose the protoc generator from the protobuf runtime, and regenerate stubs after late installs

The vLLM backends still crash on startup with

  VersionError: Detected incompatible Protobuf Gencode/Runtime versions when
  loading backend.proto: gencode 7.35.0 runtime 6.33.6

despite #10735 and #10944. Three separate defects kept it alive.

1. runProtogen picked the generator from the installed *grpcio* version.
   grpcio-tools' version tracks grpcio, but the gencode its bundled protoc
   emits tracks *protobuf*, and the two move independently: grpcio-tools
   1.82.1 (the version #10735 pins to, matching grpcio 1.82.1) requires
   protobuf>=7.35.1 and stamps gencode 7.35.0. Pinning to grpcio could
   therefore never constrain the gencode. Constrain the install to the
   protobuf already in the venv instead and let the resolver pick the newest
   compatible grpcio-tools. That both selects a generator the runtime accepts
   and stops protogen from moving the runtime under the backend's other deps.
   This is self-correcting, so the hardcoded GRPCIO_TOOLS_VERSION=1.78.0
   escape hatch from #10944 is no longer needed and is removed.

2. The stubs were generated too early. Most branches of vllm/install.sh (and
   vllm-omni) install vllm *after* installRequirements, and vllm re-resolves
   the protobuf runtime as it lands. Stubs generated against the pre-vllm
   runtime can end up newer than the runtime that finally ships, which is the
   ROCm failure exactly. Regenerate once the dependency set is final.

3. rm -f of the .py sources left __pycache__ behind. CPython validates a .pyc
   against source mtime and size, both of which can be unchanged across a
   regeneration (the gencode triple is the same width whether it reads 7.35.0
   or 6.33.5), so a stale backend_pb2.pyc could shadow the stub just written.

Also fail the build when the generated stub cannot be imported, so a
gencode/runtime mismatch surfaces at image build time instead of reaching
users as an opaque "grpc service not ready".

Verified by driving the real runProtogen through the ROCm install sequence in
a venv harness: before, gencode 7.35.0 against runtime 6.33.6 (reproducing the
reported error verbatim); after, gencode 6.33.5 against runtime 6.33.6 and the
stub imports cleanly.

Closes #10940
Closes #10718

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit]

* fix(backends): regenerate protobuf stubs in the other backends that install after installRequirements

Same defect as the vllm change: installRequirements generates the stubs at the
end of its own run, so any backend that installs further packages afterwards can
have the protobuf runtime moved out from under stubs that were already written.
The gencode stamped into backend_pb2.py then exceeds the runtime that ships and
the backend dies at model load with "grpc service not ready".

fish-speech already had this bug and worked around the symptom: it forces
protobuf>=5.29.0 after installRequirements precisely because "transitive deps
(wandb, tensorboard) may downgrade protobuf to 3.x but our generated
backend_pb2.py requires protobuf 5+". Regenerating after the pin addresses the
cause rather than propping up the runtime to match stale stubs.

Applied to the backends whose post-installRequirements step resolves a
dependency graph and can therefore move protobuf:

  fish-speech             -e . plus an explicit protobuf install
  vibevoice               pip install . (with deps)
  llama-cpp-quantization  gguf / GGUF_PIP_SPEC
  trl                     gguf / GGUF_PIP_SPEC

Deliberately not applied to ace-step and chatterbox (both --no-deps, so the
dependency graph cannot change) or voxcpm (pins setuptools only). gguf does not
depend on protobuf today, but it resolves dependencies, and "this package does
not touch protobuf right now" is exactly the assumption that made the earlier
fix ineffective.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit]

* fix(backends): resolve the protoc generator in a throwaway env so it cannot edit the backend's pinned deps

Installing grpcio-tools into the backend's own venv to generate the stubs also
drags its dependencies in: grpcio-tools 1.82.1 requires grpcio>=1.82.1, so a
backend that pinned grpcio==1.78.1 silently shipped 1.82.1 instead. Caught by
building the llama-cpp-quantization image and reading the versions back out of
the artifact:

  before   grpcio 1.82.1   (requirements.txt pins grpcio==1.78.1)
  after    grpcio 1.78.1   grpcio-tools absent from the venv entirely

Resolve the generator in a throwaway environment instead, still constrained to
the protobuf the backend ships so the gencode stays compatible. The backend's
dependency set is then exactly what its requirements files declared. protoc's
output is plain Python and carries no dependency on the interpreter that
produced it, so generating from a different env is safe; the import check still
runs under the backend's python, since that is the interpreter that has to load
the stubs at model load.

Verified on the rebuilt image: gencode 7.35.0, runtime protobuf 7.35.1, grpcio
back at its pinned 1.78.1, and the shipped stub imports cleanly against 7.35.1.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit]

* fix(backends): bound the protoc generator by BOTH the installed grpcio and protobuf

The generated stubs impose two independent constraints, and every fix so far,
including the previous commit on this branch, satisfied one while violating the
other:

  backend_pb2.py       needs  protobuf runtime >= gencode
  backend_pb2_grpc.py  needs  installed grpcio >= grpcio-tools

Resolving the generator against protobuf alone picked grpcio-tools 1.82.1 for a
backend holding grpcio at 1.78.1, so the gencode was fine but the gRPC stub was
not:

  RuntimeError: The grpc package installed is at version 1.78.1, but the
  generated code in backend_pb2_grpc.py depends on grpcio>=1.82.1.

That is also why installing grpcio-tools into the backend venv appeared to work
earlier: it dragged grpcio up to match, which was load-bearing rather than the
regression it looked like. Isolating the generator removed the accidental fix
and exposed the missing constraint.

Bound grpcio-tools from both sides instead and let the resolver find the newest
version satisfying both. The protobuf ceiling makes it back off to an older
generator when the runtime trails, bounding the gencode; the grpcio ceiling
keeps the _grpc stub loadable. Resolved against the four real runtime pairs
observed in built images:

  grpcio 1.78.1 / protobuf 7.35.1  -> grpcio-tools 1.78.0, gencode 6.31.1  OK
  grpcio 1.78.0 / protobuf 6.33.6  -> grpcio-tools 1.78.0, gencode 6.31.1  OK
  grpcio 1.82.1 / protobuf 6.33.6  -> grpcio-tools 1.81.1, gencode 6.33.5  OK
  grpcio 1.82.1 / protobuf 7.35.1  -> grpcio-tools 1.82.1, gencode 7.35.0  OK

Also restore the import check to cover backend_pb2_grpc as well as backend_pb2.
Narrowing it to backend_pb2 is why the image build passed while CI failed: the
guard could not see the constraint that was actually broken.

Verified by running the CI sequence locally for llama-cpp-quantization, the
backend whose test failed:
  make -C backend/python/llama-cpp-quantization        -> exit 0
  make -C backend/python/llama-cpp-quantization test   -> exit 0, OK

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Bash] [Edit]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-23 10:56:15 +02:00
mudler's LocalAI [bot]
1919e293c5 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 4f33af825d66e6ef1cb185e87b4589cacf747291 (#11040)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-23 10:54:46 +02:00
mudler's LocalAI [bot]
12ee5249be chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to 82cd05b9f3a175612dc89fd6943e610fab096ef5 (#11039)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-23 10:50:29 +02:00
mudler's LocalAI [bot]
e1d7491703 chore: ⬆️ Update leejet/stable-diffusion.cpp to 8a51eb92848c1327a5aaeff5ad81a7a9a2435255 (#11038)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-23 10:50:05 +02:00
mudler's LocalAI [bot]
c4c5849cea chore: ⬆️ Update localai-org/ced.cpp to db5aae02973a745722d6fbd2157cab1999106777 (#11037)
⬆️ Update localai-org/ced.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-23 10:49:51 +02:00
mudler's LocalAI [bot]
fab647c23a chore: ⬆️ Update CrispStrobe/CrispASR to 3ab5f4ac13685966b47cc75dc7fd02f3c4a51beb (#11035)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-23 10:49:33 +02:00
mudler's LocalAI [bot]
9b8ce0ace4 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260722081849 (#11034)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-23 10:49:16 +02:00
mudler's LocalAI [bot]
7358833f52 chore: ⬆️ Update mudler/locate-anything.cpp to 77376ab332de918220f7a7e391542eefb5407c9f (#11062)
⬆️ Update mudler/locate-anything.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-23 10:49:02 +02:00
mudler's LocalAI [bot]
b57aa8142f chore: ⬆️ Update ikawrakow/ik_llama.cpp to e5357286c0d433cd4384e82ed7e2b6d655f57087 (#11063)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-23 10:48:40 +02:00
localai-org-maint-bot
9fbb8e89cf fix(turboquant): supersede stale dependency bump (#11064)
* ⬆️ Update TheTom/llama-cpp-turboquant

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(turboquant): refresh HIP compatibility patch

The updated fork now carries its own HIP-safe peer-copy path, so the old hunk no longer applies. Keep only the event-creation compatibility change that the fork still needs.

Assisted-by: Codex:gpt-5 [Codex]

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-23 10:48:24 +02:00
mudler's LocalAI [bot]
ec49548c8e fix(modelartifacts): resume interrupted materialization per-file, not from scratch (#11071)
materializeLocked built a download task for every file in the resolved
snapshot unconditionally. A completed file is promoted from
.downloads/<hash> into snapshot/<path> and its blob deleted, so on any
re-entry (a controller pod roll, a resubmit, a crash) the new pass built a
task whose .downloads blob no longer existed, re-downloaded the whole file
from Hugging Face, and its AfterDownload even removed the already-complete
snapshot copy first. The only resume that worked was the downloader's
per-file .partial resume for a file caught mid-transfer; completed files
were never skipped.

Production consequence: installing longcat-video-avatar-1.5 (~35 GB after
allow_patterns) on a cluster whose controller rolls hourly (Flux image
automation) never converged across ~14 hours. Each roll restarted from the
first shard; the completed bytes on disk were repeatedly deleted and
re-fetched, and the artifact never promoted. curl of the same files from
inside the pod ran fine, proving the loss was the materializer re-fetching,
not the network.

Before building a task, check whether the file is already materialized and
verified in this staging tree's snapshot/ and, if so, keep it and count it
complete instead of downloading. "Materialized" means a regular file of the
expected size that passes the same verifyDownloadedFile check the download
path uses, so the kept manifest entry is byte-for-byte identical to a fresh
one and integrity is re-checked. The manifest requires a SHA-256 for every
file and non-LFS files carry none to borrow, so a hash is unavoidable for
the manifest anyway; a full re-hash of local disk is still orders of
magnitude cheaper than re-downloading, and the downloader re-verifies any
file it does fetch. Manifest entries are now written at their snapshot index
rather than appended in completion order, so a mix of skipped and downloaded
files keeps the resolved order that committedResult and staging read. The
unconditional root.Remove(destination) now runs only on the fresh-download
path; a kept file survives. Skips are logged at INFO with count and bytes so
an operator can see resume working.

This is the resume-side counterpart to the sibling defects on this path:
read/write error conflation and transient retry (#10985), hash-verify
progress accounting and silent success on an expired deadline (#11026), and
the response-header hang (#11053). The download machinery resumed a single
in-flight file; the materializer above it still threw away every completed
file on restart. It also makes orphan-partial adoption worth its cost:
an adopted tree's completed files were re-downloaded anyway until now.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-23 10:48:01 +02:00
mudler's LocalAI [bot]
95afddd936 chore(model-gallery): ⬆️ update checksum (#11061)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-23 01:26:26 +02:00
mudler's LocalAI [bot]
6584db992f fix(nodes): never schedule a model onto a node that cannot store it (#11054)
* fix(nodes): never schedule a model onto a node that cannot store it

A worker whose models filesystem was 100% full kept advertising
`status: healthy`, stayed a scheduling candidate, was picked to host a
70 GB video model, accepted the staging request, transferred ~17 GB and
only then failed:

  staging .../whisper-large-v3/model.fp32-00001-of-00002.safetensors:
    upload to node b7bacbf4-... failed with status 500:
    writing file: /models/longcat-video-avatar-1.5/...: no space left on device

The node was at 937G/937G/0-avail. Total elapsed before the truth
surfaced: 16 minutes, for a decision that could never have succeeded.

The worker health signal only ever proved liveness. `/readyz`
(WorkerReadiness/NATSReadiness) checks the NATS link; `status: healthy`
in the registry is driven by heartbeat recency. Node capacity carried
VRAM and RAM but no disk figure at all, and the router compared model
size against VRAM only — nothing anywhere looked at free space on the
filesystem that staging actually writes to.

Report it, then use it:

- Workers now measure the filesystem backing their MODELS directory
  (not `/` -- staged weights land in the models path, and that mount is
  very often separate) and report `total_disk`/`available_disk` on
  registration and on every heartbeat. Free disk moves faster than VRAM
  under staging traffic, so the per-heartbeat refresh matters.
- The SmartRouter drops nodes that cannot store the model before it
  picks one. The requirement comes from `modelPayloadBytes` -- the same
  local paths `stageModelFiles` uploads, already computed for the
  size-derived load budget -- plus a 5% / 1 GiB margin, rather than a
  fixed percentage of the node's disk. A percentage threshold would take
  a small-but-usable node out of rotation for models it could hold, and
  on a homogeneous cluster would strand every node at once.
- When no node fits, scheduling fails immediately with an error naming
  the requirement and each node's free space, instead of picking one and
  discovering it mid-transfer.

Two deliberate non-changes. Low disk does not mark a node `unhealthy`:
the check is per model, so a node too small for one model stays a valid
target for smaller ones. And `total_disk == 0` means "does not report
disk" (pre-upgrade worker, or a failed stat), not "full" -- such nodes
pass through untouched so a rolling upgrade never empties the candidate
pool. A genuinely full node is distinguishable: non-zero total, zero
available. Registry read failures are logged and scheduling continues
unfiltered; a database hiccup must not wedge a cluster.

Free space is surfaced on the node detail page next to VRAM, since the
incident's signature was a node that looked entirely healthy.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

* feat(nodes): make the disk-headroom check operator-controllable

The admission check added in the previous commit had no off switch. A
scheduler-side veto with no escape hatch is a liability: our size
estimate can be wrong (deduplicating or compressing filesystems, a
backend that fetches its own weights rather than loading the staged
copy), and an operator who hits that has no way out but a downgrade.

Add one knob with two surfaces that share a single source of truth:

- `--distributed-disk-headroom-check` / `LOCALAI_DISTRIBUTED_DISK_HEADROOM_CHECK`
  (default true), following the `--distributed-prefix-cache` pattern for
  a default-on distributed feature.
- `distributed_disk_headroom_check` in the runtime-settings registry, so
  it can be flipped without a restart from `POST /api/settings` and from
  Settings -> Distributed in the WebUI.

Both write `DistributedConfig.DiskHeadroomDisabled`, and the SmartRouter
reads that member LIVE on every scheduling decision through a closure
over the application config rather than a value snapshotted at
construction. Env/CLI sets the boot value, the runtime setting overrides
it live, last write wins, and there is exactly one member to read.
Snapshotting would have made the runtime toggle a no-op until restart.

Disabled means WARN, not SKIP. Selection goes back to ignoring free disk
-- byte for byte the pre-check behaviour -- but the check still runs, and
when it would have rejected every node it says so, naming the knob that
suppressed it. Going quiet when switched off would reproduce the exact
condition that made the original incident expensive: a cluster doing
something that could not work and saying nothing. Disabling is also
logged once at startup. Warning only on the total-rejection case keeps
it actionable rather than chatty on a heterogeneous cluster.

Also fixes a false positive in the check itself: shared-models mode
(LOCALAI_DISTRIBUTED_SHARED_MODELS) stages nothing at all -- every node
already mounts this models directory at this path -- so demanding the
full checkpoint size of free space per node would have rejected a
cluster that needs no new bytes. The check is skipped there entirely.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-23 00:03:21 +02:00
mudler's LocalAI [bot]
f317da7c0f fix(galleryop): make admitted operations queryable and survive a failed op (#11044)
Two lifecycle defects observed on a 2-replica distributed cluster.

The install endpoints mint a job UUID, hand the operation to an unbuffered
channel, and answer HTTP 200 immediately. The gallery worker is a single
goroutine that processes operations serially, and the first status write
happens inside modelHandler/backendHandler — i.e. only once the worker
actually starts the work. An operation queued behind a running install
therefore had no status at all: GET /models/jobs/<uuid> answered
"could not find any status for ID" and GET /models/jobs did not list it,
so the endpoint reported success for work nothing could observe. On the
paths that sent directly rather than from a goroutine, the same unbuffered
channel blocked the HTTP handler for the whole duration of the in-flight
install, which is how a replica came to accept no /models/apply at all
while /readyz stayed green.

Admission now goes through EnqueueModelOp/EnqueueBackendOp, which publish a
"queued" status before handing the operation over, so a job ID is queryable
from the instant it is handed out. Delivery selects on the operation's
context, so cancelling a still-queued operation releases the delivery
goroutine instead of stranding it on a send that will never be received,
and an operation the worker never accepts becomes a terminal failure rather
than a silent leak.

The worker also had no panic containment. A panic in any handler propagated
out of the single consumer goroutine and killed the process, taking every
queued operation with it; it is now contained to the operation that caused
it. The two ignored galleryStore.Create errors are logged, and the model and
backend delete endpoints now run under the same ID they hand back — they
previously ran under an empty ID and returned a status URL for a job that
could never have a status.

Second, an operation orphaned by a controller replaced mid-download kept
reporting phase=downloading, processed=false, error=none while nothing was
downloading. The PostgreSQL side does recover on its own (FindDuplicate
ignores rows untouched for 30 minutes and CleanStale marks them failed), but
the reaper only ever corrected the database. The in-memory statuses map that
GET /models/jobs/<id> and /api/operations actually read was never corrected,
so every replica kept serving the frozen tick indefinitely. ReapStaleOperations
now reconciles the in-memory copy with the reap.

Note that operation ownership is still not tracked: gallery_operations has a
FrontendID column that nothing writes, so a live operation and one whose owner
died are distinguished only by a 30-minute staleness timeout. Narrowing that
window needs a lease/heartbeat mechanism and is out of scope here.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-23 00:02:23 +02:00
mudler's LocalAI [bot]
6cee8dee54 docs: ⬆️ update docs version mudler/LocalAI (#11060)
⬆️ Update docs version mudler/LocalAI

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-22 23:07:00 +02:00
mudler's LocalAI [bot]
ff299df453 perf(http): gzip responses, cache hashed assets, bound the trace endpoints (#11056)
Three measured HTTP-layer regressions on a live deployment, fixed together
because they all shape the bytes on the wire.

1. No compression. The server sent no Content-Encoding regardless of what
   the client asked for, confirmed with curl straight at 127.0.0.1:8080 so
   it was not an ingress artefact. Adds gzip middleware, on by default and
   configurable via LOCALAI_DISABLE_HTTP_COMPRESSION and
   LOCALAI_HTTP_COMPRESSION_MIN_LENGTH (default 1024 bytes so tiny bodies
   are not wastefully wrapped). Streaming routes are skipped explicitly:
   an SSE Accept header, a WebSocket upgrade, and the completion / SSE /
   log-tail path prefixes, because whether a completion request streams is
   decided by the request body, which the middleware runs too early to see.
   Already-compressed formats (woff2, png, mp4, ...) are skipped too; gzip
   made those marginally larger. Measured over the embedded React build:
   JS+CSS 2815 KB raw to 808 KB gzipped (3.48x).

2. No cache headers on content-hashed assets. Vite hashes the filenames,
   so a given /assets/ URL can never change content, yet they shipped with
   no Cache-Control, ETag or Last-Modified, and the browser re-fetched the
   whole bundle on every navigation with no conditional request available.
   /assets/* now carries public, max-age=31536000, immutable. index.html
   stays no-cache so a deploy is picked up, and the unhashed locale JSONs
   get a short TTL rather than the immutable one.

3. Unbounded trace endpoints. /api/traces returned 21,033,606 bytes in
   4.65s and /api/backend-traces 3,471,682 bytes in 1.50s, and the admin
   UI polls both every few seconds. The ring buffer holds up to 1024
   entries, each embedding full input_text payloads. Both list endpoints
   now take limit / offset / full, default to 50 entries, and strip the
   heavy fields (request and response bodies plus headers for API traces,
   body and data for backend traces) unless full=true. Every trace gets a
   process-lifetime ID and GET /api/traces/{id} and
   /api/backend-traces/{id} serve the full record, which is what the UI
   fetches when a row is expanded. The list body stays a JSON array;
   paging metadata rides in X-Total-Count, X-Trace-Offset and
   X-Trace-Limit. Reproducing the live shape in a test, the polled payload
   goes from 21,131,097 bytes to 7,201 bytes.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-22 22:51:25 +02:00
mudler's LocalAI [bot]
8eb8376596 fix(ci): dedup the three workflows that stack runs on every PR push (#11058)
build-test.yaml, yaml-check.yml and secscan.yaml had no concurrency block at
all, so every push to a PR stacked another full batch instead of superseding
the previous one. build-test carries a macos-latest job, the scarcest runner
class we use, and secscan fires on every push to every branch because its
`push:` trigger is unfiltered.

build-test and yaml-check use the same group idiom as lint.yml and the other
eleven workflows that already have one: key on the PR number so pushes to a PR
share a group, and cancel only on pull_request. On a master push the key falls
back to github.sha and cancel-in-progress is false, so master runs never cancel
each other -- that is deliberate, since backend.yml builds only the backends a
given commit touched and superseding would drop those builds.

secscan needs a different key: it has no pull_request trigger, so the shared
idiom would fall back to the unique-per-commit sha and dedup nothing. It groups
on github.ref instead, and excludes master from cancellation for the same
per-commit reason. Cancelling a superseded feature-branch scan is safe because
the only output is a SARIF upload and code scanning keeps the latest result
per ref.

No behaviour change on master for any of the three.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-22 22:51:11 +02:00
mudler's LocalAI [bot]
16033d562a fix(downloader): bound the wait for response headers so a wedged origin cannot hang an install forever (#11053)
A gallery model install hung for 94 minutes with zero bytes transferred, no
error, no retry and no abort, leaving a partial tree frozen at 18G. The last
log line was the download starting, then silence:

    14:06:19 INFO Downloading url=".../LongCat-Video-Avatar-1.5/resolve/<rev>/base_model/diffusion_pytorch_model-000..."

The retry machinery from #10985 was working (two retries fired at 14:05:01 and
14:06:14); the third attempt simply never returned. The install never
completed, the model config was never written, and nothing surfaced the
failure.

The stall watchdog added earlier wraps the response *body*, so it only starts
guarding once downloadClient.Do() has returned. The transport had no
ResponseHeaderTimeout, so a peer that completes the dial and TLS handshake,
reads the request, and then never sends a status line parks Do() for the
process lifetime. IdleConnTimeout governs pooled idle connections, not an
in-flight request. Both the body request and the HEAD that probes for Range
support were unguarded.

Bound the header wait at the transport, not the client: a client-level Timeout
would also bound the body and truncate multi-tens-of-GB downloads. The knob is
opt-in (WithResponseHeaderTimeout) rather than a default in HardenedTransport,
because a streaming endpoint may legitimately withhold headers until it has
something to say, and capping that would break the streaming clients that share
this constructor.

Also fix a classification trap this exposed: net/http reports a
ResponseHeaderTimeout as an error satisfying errors.Is(err,
context.DeadlineExceeded), which IsRetryable read as "the caller gave up" and
refused to retry. An explicit transient marking now outranks the cancellation
sentinels; a caller who genuinely gave up is still caught by the ctx.Err()
check. The resume probe's error is likewise marked transient, so a momentarily
wedged origin no longer turns a resumable download into a hard install failure.

Third defect found in this download path, after #10985 (read vs write errors
conflated) and #11026 (hash verification emitted no progress and an expired
deadline returned success).


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-22 18:28:51 +02:00
mudler's LocalAI [bot]
248e1ef9a2 fix(worker): never reuse a backend process whose directory a reinstall replaced (#11029)
A backend reinstall could poison every subsequent model load on a worker
node until the worker process was restarted.

gallery.InstallBackend (and gallery.UpgradeBackend) replace a backend by
renaming the live directory to `<name>.install-backup`, moving the staged
directory into place, then deleting the backup. A working directory
follows the inode across a rename, so a backend process that outlives
that swap ends up with a deleted inode as its CWD, and every getcwd(2)
in it fails with ENOENT.

Observed on a Jetson Thor worker in distributed mode after two
successive reinstalls of cuda13-nvidia-l4t-arm64-longcat-video-development.
A later model load failed with:

    rpc error: code = Internal desc = failed to load LongCat model: [Errno 2] No such file or directory

The backend's own traceback shows it dying while importing torch, before
touching any model file:

    backend.py line 142 in LoadModel
    backend.py line 300 in _import_torch
      torch/_library/custom_ops.py  lib._register_fake(...)
      torch/library.py:183          caller_module = inspect.getmodule(frame)
      inspect.py:1013               f = getabsfile(module)
      inspect.py:983                return os.path.normcase(os.path.abspath(_filename))
      <frozen posixpath>, line 415, in abspath
    FileNotFoundError: [Errno 2] No such file or directory

os.path.abspath calls os.getcwd() for a relative path. Scanning /proc
inside the worker container found the deleted CWD directly:

    pid 23467 CWD DELETED: /backends/cuda13-nvidia-l4t-arm64-longcat-video-development.install-backup (deleted)

Restarting the worker container cleared it (dead CWD count 1 -> 0).

Python backends import torch lazily inside LoadModel, so such a survivor
still answers HealthCheck and keeps its gRPC port. It looks healthy and
only detonates when a model is actually loaded through it.

The install paths already stop running processes before replacing the
directory (installBackend's force branch, upgradeBackend, backend.delete),
but they resolve them by name. That bookkeeping reaps nothing whenever
the recorded name no longer resolves into the install's identity set: a
legacy entry with an empty backendName, backendIdentity degraded to
name-only matching after a ListSystemBackends failure, or an earlier
reinstall having already rewritten the metadata.json that carries the
alias. Any of those leaves a live process whose directory is about to be
unlinked, and nothing downstream notices, because the reuse gate checks
liveness and name -- and the name is precisely what does not change
across a reinstall.

Record the directory each supervised process runs out of, plus that
directory's identity at spawn time, and compare with os.SameFile before
reusing the process. This needs none of the name bookkeeping to have
been correct. Both reuse gates are covered: processMatchesBackend (the
install fast path) and startBackend's own already-running branch, which
now force-stops such a survivor so the fresh spawn chdirs into the newly
installed directory. Processes with no recorded directory are accepted,
so a rollout does not restart every running backend once.

This matters more with #11024 pending: making GPU backends visible to
the upgrade checker will have AutoUpgradeBackends fan upgrades out to
worker nodes at scale, and every one of those is a reinstall. Left as
is, a rare manual-upgrade footgun becomes a fleet-wide one.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-22 17:29:59 +02:00
mudler's LocalAI [bot]
7a8db9b1f1 fix(ollama): set ContextSize via the embedded LLMConfig so the package builds (#11049)
The num_ctx clamping specs added in #11032 construct their fixture with
`config.ModelConfig{ContextSize: &existing}`, but ContextSize is not a
direct field of ModelConfig: it belongs to LLMConfig, which ModelConfig
embeds inline. Go allows reading a promoted field but not setting one in
a composite literal, so the test file has never compiled:

  helpers_internal_test.go:33:31: unknown field ContextSize in struct
    literal of type "github.com/mudler/LocalAI/core/config".ModelConfig

This broke `make lint` on master from bf19758e0 onward, and because the
typecheck failure takes down the whole package it also reds tests-linux
and tests-apple on every PR branched after that commit.

Use the same literal form the rest of the tree already uses for this
field (see core/backend/options_internal_test.go).

Worth noting the specs were not merely uncompiled but inert: #11032 is a
DoS fix (an unauthenticated client raising the context ceiling drives
KV-cache allocation), and its regression guard was never actually
running. Verified the restored specs are functional by stubbing out the
ceiling clamp, which fails the "does not let an oversized num_ctx raise
an existing context ceiling" spec as intended.

Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-22 16:25:56 +02:00
walcz-de
6d1bbb74c4 fix(backend/python): don't await sync servicer behaviors in AsyncModelIdentityInterceptor (#10980)
* fix(backend/python): don't await sync servicer behaviors in AsyncModelIdentityInterceptor

The model-identity interceptor (added for #10952) is installed on every Python
backend's gRPC server. Its grpc.aio variant invokes the wrapped servicer
behavior itself and awaits the result unconditionally:

    result = await original(request, context)                 # LoadModel
    return await original_unary(request, context)             # guarded RPCs
    async for response in original_stream(request, context):  # streaming

But a backend's servicer methods may be plain sync functions. The transformers
backend, for one, defines `def LoadModel` and `def Embedding` (not `async def`).
grpc.aio's own dispatch adapts both shapes, but this interceptor calls the
behavior directly and bypasses that. For a sync method `original(...)` returns a
message object, not a coroutine, so the `await` raises:

    TypeError: object Result can't be used in 'await' expression

The model loads, then the LoadModel RPC dies on return; the guarded sync
Embedding fails the same way. It happens on every platform, not just one backend
build. CI never caught it because AsyncModelIdentityInterceptor had no
behavioral test -- only an "is it installed" assertion.

Fix: await only when the behavior actually returned an awaitable
(inspect.isawaitable), mirroring grpc.aio's own sync/async adaptation. The
streaming guard iterates a sync generator with `for` and an async one with
`async for`.

Adds async-path coverage to model_identity_test.py exercising both sync and
async LoadModel / guarded-unary / streaming behaviors. The sync cases fail on
the current code with the TypeError above and pass with this fix.

Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>

* fix(backend/python): dispatch sync servicer behaviors off the event loop

Addresses review feedback: awaiting only awaitable results removed the
TypeError, but still ran a sync LoadModel/Embedding -- and stepped a sync stream
via next() -- on the asyncio event-loop thread, so a slow load/inference/stream
could freeze all aio RPC handling.

Route sync behavior through run_in_executor (a worker thread) while awaiting
native async behavior directly. A callable wrapper that returns an awaitable is
run in the thread and its awaitable awaited back on the loop. Sync streaming
pulls each item via the executor with a done sentinel, so StopIteration cannot
escape through a Future.

Adds regression tests that record the handler thread id and assert it differs
from the event-loop thread, for LoadModel, a guarded unary RPC and a sync stream.

Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>

---------

Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>
2026-07-22 16:03:46 +02:00
dependabot[bot]
f92410b20b chore(deps): bump fast-uri from 3.1.2 to 3.1.4 in /core/http/react-ui in the npm_and_yarn group across 1 directory (#11043)
chore(deps): bump fast-uri

Bumps the npm_and_yarn group with 1 update in the /core/http/react-ui directory: [fast-uri](https://github.com/fastify/fast-uri).


Updates `fast-uri` from 3.1.2 to 3.1.4
- [Release notes](https://github.com/fastify/fast-uri/releases)
- [Commits](https://github.com/fastify/fast-uri/compare/v3.1.2...v3.1.4)

---
updated-dependencies:
- dependency-name: fast-uri
  dependency-version: 3.1.4
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-22 15:55:10 +02:00
mudler's LocalAI [bot]
54f531f452 fix(mcp): bound MCP session connect so an unreachable server can't hang the widget (#10880) (#10884)
Establishing an MCP session held the session-cache mutex across
client.Connect with no per-connect timeout. An unreachable remote server
(bounded only by the 360s httpClient timeout) or a stdio server whose
initialize handshake never completes therefore blocked the caller and,
because the mutex was held, every other MCP request for that model too.
In the UI this shows up as the MCP "Servers" widget spinning forever.

It is most visible for cloud-proxy models: their chat path bails out
before the MCP tool block, so it never warms the session cache in the
background. The widget's /v1/mcp/servers/<model> call is then the first
and only code that connects synchronously, in the request foreground.

The session, once established, stays bound to the shared context (it is
cancelled later via the cached cancel func on eviction/shutdown), so we
can't pass a WithTimeout context to Connect: firing the timeout would tear
a healthy session down, and cancelling the shared context would also kill
sibling servers that already connected. Instead connectMCP runs Connect on
the shared context in a goroutine and stops waiting after the discovery
timeout, returning an error for that one server without disturbing the
others. A stalled goroutine is reaped when the model's sessions are
cancelled. Applied to both SessionsFromMCPConfig and
NamedSessionsFromMCPConfig.


Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-22 15:49:40 +02:00
mudler's LocalAI [bot]
01fca9c9b2 fix(distributed): scale the remote model-load deadline with checkpoint size (#11030)
The gRPC deadline for the remote LoadModel call was a fixed 5m. It starts
only after the backend install and file staging have completed, so it
covers the worker's checkpoint read and pipeline init alone - work whose
duration is proportional to the bytes on disk. A fixed value is therefore
a model-size cliff, not a timeout.

Measured in production: a 70 GB video checkpoint (longcat-video-avatar-1.5)
on an NVIDIA Jetson Thor worker failed reproducibly with
"rpc error: code = DeadlineExceeded" after 953.5s of wall clock. Backend
install plus staging consumed ~11m, then LoadModel got its 5m and expired.
The load never had a chance, and the operator saw only a generic
DeadlineExceeded with no hint that a config value was the cause.

Raising the constant does not fix this. It moves the cliff to the next
larger model - the cluster has to support 600 GB checkpoints - and it makes
a genuinely wedged SMALL model hang for the whole inflated duration before
anyone notices, which is a real regression in failure latency.

So derive the budget from the checkpoint size instead:

    budget = 5m + 20s/GiB, capped at 6h

2 GiB -> 5m40s, 70 GiB -> 28m20s, 600 GiB -> 3h25m. The per-GiB rate is
deliberately pessimistic (~54 MB/s of weight read) because the errors are
not symmetric: too long costs only failure latency on a load that was going
to fail anyway, too short is a guaranteed false failure on a healthy load.

The size is measured from the frontend's local model files, over the same
path set stageModelFiles uploads. When those files are not present locally -
a backend handed a bare HuggingFace repo id fetches its own weights on the
worker - there is nothing to measure and the budget stays at today's 5m.

An explicit LOCALAI_NATS_MODEL_LOAD_TIMEOUT still wins outright, in both
directions: a shorter override is honoured, so an operator who wants fast
failure is not silently extended by the heuristic.

The cold-load hold needed widening to match. It extends on staging progress,
but LoadModel reports none, so once the last byte lands the hold expires a
stall window later and would cancel a load still well inside its own budget.
scheduleAndLoad now extends the hold by the load budget plus the staging
margin as it enters the load phase; ModelLoadCeilingFor stays the hold's
starting budget rather than its maximum.

Finally, a deadline that does expire now names the budget, the checkpoint
size it was derived from, and the knob that overrides it, instead of
surfacing a bare "context deadline exceeded".


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-22 09:30:23 +02:00
Tai An
d7020708f2 fix(completions): reject empty PromptStrings in streaming to avoid index-out-of-range panic (#11028)
* fix(completions): reject empty PromptStrings in streaming to avoid index-out-of-range panic

The streaming branch of CompletionEndpoint only guarded len(config.PromptStrings) > 1
before unconditionally reading config.PromptStrings[0]. A completion request whose
prompt field is an empty array, an array of non-strings, or omitted leaves
PromptStrings with length 0, so PromptStrings[0] panics with index out of range and
crashes the handler goroutine.

Guard for exactly one prompt string instead, returning a clean error for the 0-length
case as well as the pre-existing multi-prompt case.

Signed-off-by: Tai An <antai12232931@outlook.com>

* fix(completions): return 400 for malformed streaming prompt

Reject streaming completion requests whose prompt does not resolve to
exactly one string (omitted prompt, empty array, or a multi-element
array) with an HTTP 400 before writing any SSE headers, instead of
returning a plain error that Echo surfaces as a 500. Extract the guard
into validateStreamingPromptStrings and cover the three reported
payloads with a regression test.

Fixes #11021

Signed-off-by: Tai An <antai12232931@outlook.com>

---------

Signed-off-by: Tai An <antai12232931@outlook.com>
2026-07-22 09:29:11 +02:00
Tai An
bf19758e05 fix(ollama): cap num_ctx so it cannot wrap negative when cast to int32 (#11032)
* fix(ollama): cap num_ctx so it cannot wrap negative when cast to int32

applyOllamaOptions copied a client-supplied options.num_ctx straight into
cfg.ContextSize with only a > 0 check. That value is later cast to int32
before it reaches the backend (core/backend/options.go), so a num_ctx
above math.MaxInt32 silently wrapped into a negative context size that
was then sent to the LoadModel gRPC call. Both /api/chat and /api/generate
share applyOllamaOptions, so both endpoints were affected.

Cap num_ctx at math.MaxInt32 so the later cast stays positive, and add
internal regression coverage for the overflow, in-range, and unset cases.

num_ctx remains an intentional user override, so this does not re-impose
the hardware-aware auto context clamp; that policy choice is left to
maintainers.

Fixes #11022

Signed-off-by: Tai An <antai12232931@outlook.com>

* fix(ollama): clamp num_ctx to model context ceiling, not just int32

Per review on #11032: capping only at math.MaxInt32 still let an
unauthenticated request replace the hardware/model-derived context
limit with ~2.1B tokens, so a real backend could attempt a catastrophic
KV-cache allocation. Treat any existing positive cfg.ContextSize as the
server ceiling and clamp num_ctx down to it (smaller values still
honored), while retaining the int32-safe bound when no smaller ceiling
exists. Shared by /api/chat and /api/generate via applyOllamaOptions.

Add regression coverage proving num_ctx=2,000,000,000 cannot replace an
existing 4096/8192 ceiling.

Signed-off-by: Tai An <antai12232931@outlook.com>

---------

Signed-off-by: Tai An <antai12232931@outlook.com>
2026-07-22 09:28:02 +02:00
mudler's LocalAI [bot]
48b7d6d8fd docs: ⬆️ update docs version mudler/LocalAI (#11033)
⬆️ Update docs version mudler/LocalAI

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-22 09:26:29 +02:00
mudler's LocalAI [bot]
3154bec357 chore(model-gallery): ⬆️ update checksum (#11036)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-22 08:25:24 +02:00
localai-org-maint-bot
47c0e06198 fix(ci): authenticate the nightly dependency-bump API calls (#11042)
The "Bump Backend dependencies" workflow has failed every night for the
last two weeks. #11012 fixed one cause (repos renamed under localai-org);
what is left is rate limiting.

bump_deps.sh fans out to ~25 parallel matrix jobs that each query
api.github.com anonymously. Anonymous calls are capped at 60/hour per
source IP and GitHub-hosted runners egress through shared NAT addresses,
so a random handful of jobs draw HTTP 403 and die at curl exit 22 with an
empty response. Last night that hit ggml-org/whisper.cpp and
mudler/depth-anything.cpp -- both public and resolvable, nothing wrong
with either pin.

Route every bump script through a shared gh_curl helper that sends
GITHUB_TOKEN when present (1000/hour instead of 60) and retries transient
failures, including the 403s that plain --retry ignores. The helper
suppresses xtrace around the call so the Authorization header cannot land
in a public job log.

bump_docs.sh had a sharper version of the same bug: it piped an
unchecked response into `jq -r .tag_name`, so a throttled request
resolved to the string "null" and would have been published as the docs
version. It now refuses to write anything it cannot resolve to a tag.

Verified locally by running all four scripts end to end against their
real upstreams: correct SHAs/tags written, exit 0; a nonexistent repo now
fails with a named diagnostic instead of a bare exit 22 and leaves the
pinned file untouched; the token is absent from the xtrace output; and
the scripts still work unauthenticated.

Assisted-by: Claude:opus-4.8 [Claude Code]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-22 08:25:05 +02:00
mudler's LocalAI [bot]
5c96e097ba feat(gallery): fix stale DFlash drafters and add the APEX families as variant ladders (#11027)
* fix(gallery): repoint qwen3-4b/qwen3.5-9b dflash drafters at post-rename GGUFs

The drafters both entries referenced were converted from the pre-merge DFlash
PR branch and carry dflash.target_layer_ids. llama.cpp reads dflash.target_layers
and refuses the load. The stored values are offset by +1 relative to the HF-side
field, so the files cannot be repaired by renaming the key and must be replaced.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ci): add apexentries HuggingFace client

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* ci(apexentries): build the HF client via pkg/httpclient

The apexentries HuggingFace client was constructed as a raw
&http.Client{Timeout: 60s}. The repo convention (documented in
.golangci.yml, which cannot express this as a forbidigo pattern) is that
all outbound HTTP goes through pkg/httpclient, which refuses redirects by
default and sets a TLS 1.2 floor. The std client follows redirects and
forwards custom credential headers to the redirect target on a cross-host
hop (GHSA-3mj3-57v2-4636). Only a User-Agent is sent today, but this
calls an external API and an HF_TOKEN header added later would leak.

Switch to httpclient.NewWithTimeout, preserving the 60 second timeout.
No behaviour change for the current header set.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ci): discover APEX tiers by filename suffix

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ci): resolve unsloth counterparts and sharded quants

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ci): render APEX child entries with the dflash/mtp tag rule

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* ci(apexentries): set backend, known_usecases and cross-repo drafters

RenderChild left three gaps against the hand-written gallery entries.

The generated entries reference gallery/virtual.yaml, which supplies no
backend, so every generated entry named no engine at all. All comparable
hand-written entries set backend: llama-cpp in overrides; do the same.
Set known_usecases to [chat] alongside it: LocalAI falls back to the
backend defaults when it is absent, so this is convention rather than
breakage, but generated entries should not read differently from their
neighbours.

The drafter was also assumed to live in the repo publishing the weights.
Speculative pairings routinely cross repos, and a drafter URI built from
the weights repo 404s at install time. Add ChildInput.DraftRepo, used for
both the drafter URI and its local path, falling back to Repo when empty
so pairings that do ship the drafter alongside the weights are unchanged.

The dflash/mtp tagging rule is untouched: the tag still follows SpecType
and nothing else.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ci): dedupe generated entries against the existing gallery

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* apexentries: canonicalize HF URIs and dedup the generated batch

Merge exists to stop a second gallery entry being added for weights the
gallery already ships, but two gaps let duplicates through on a bulk run.

The URI key was compared as an exact string while render.go only ever emits
https://huggingface.co/{repo}/resolve/main/{file} and the gallery records
1038 of its URIs in huggingface://{repo}/{file} shorthand. A generated
unsloth rung whose weights are already shipped in shorthand was therefore
not recognised. canonicalURI reduces both spellings to one key and is
applied on both sides, taking care that the repo is exactly the first two
path segments so sharded quants in a subdirectory still match. A URI in
neither form is returned untouched so other hosts dedup on their literal
string.

Merge also never accounted for entries it had just accepted, so two
generated entries sharing a name or a primary URI both landed in add.
Several APEX repos share one base model and resolve to the same unsloth
counterpart, so the identical rungs are generated twice under the same
name. Batch state is tracked locally rather than written back into the
caller's ExistingIndex, which a caller may reasonably reuse.

Name is still checked before URI: a name collision must block the add
regardless of the weights.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ci): verify variant and tagging invariants in the gallery index

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ci): scope the apex-entries verifier to what it can actually judge

The verifier reported 60 problems against the real gallery, 57 of which were
llama.cpp assumptions meeting entries from other backends. A gate that is wrong
57 times out of 60 cannot gate anything.

- The weight-count check catches a quant label collision in llama-cpp quant
  discovery, so it now runs only for overrides.backend: llama-cpp. Entries with
  no declared backend are skipped because their weights are declared in the
  referenced url: template, which the verifier never reads.
- The dflash/mtp tag check now implements the per-backend table in
  .agents/adding-gallery-models.md instead of assuming llama.cpp's spec_type:
  vocabulary. ds4 declares mtp_path:/mtp_draft:; sglang declares
  speculative_algorithm: in a file this verifier cannot follow, so sglang
  entries are not judged in either direction. The check stays bidirectional
  within the backends it does judge.
- sha256 is now required on .gguf files only, since every non-GGUF asset in the
  index belongs to a hand-curated entry outside this generator's scope.

Against the current gallery this leaves exactly the three genuine problems:
two entries setting spec_type:draft-mtp without the mtp tag, and one entry
whose overrides.mmproj names a file it does not download.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* ci(apexentries): anchor quant matching and invert the sha256 rule

UnaccountedQuants matched files to wanted quants with strings.Contains, which
reproduces the substring collision it was written to warn about: Q8_0 is a
substring of UD-Q8_0, so a repo publishing only UD-Q8_0 was reported as
publishing an unbuilt Q8_0. Subdirectory-sharded UD quants are the normal
unsloth layout for large repos, so this fired on realistic input.

Match on the quant label as an anchored token instead, the way
DiscoverUnslothQuants does, so the diagnostic and the discovery it audits
cannot disagree about what a file is. Root-level shards, the layout the
diagnostic mainly exists to catch, stay detected.

The sha256 requirement was scoped to .gguf, which exempted seven real model
weights: wan_2.1_vae.safetensors and clip_vision_h.safetensors across the
wan-2.1-*-ggml entries, both load-bearing weights named by gallery/wan-ggml.yaml.
Invert the rule so a checksum is required on everything except metadata
extensions, which keeps a future weight format covered by default rather than
silently exempt.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ci): wire the apexentries command

Adds the generation path to the apexentries command: list the mudler APEX
repos, discover each one's quality ladder and its unsloth counterpart's quant
rungs from the filenames actually published, render a child entry per build
plus a family parent carrying the variants list, dedup against the gallery,
and write the additions to -out or append them with -apply.

Discovery shortfalls are reported at discovery time rather than left to the
verifier. A quant or a tier that discovery drops leaves no trace in a finished
gallery file, and because an empty imatrix ladder falls back to the plain one,
a repo whose imatrix filenames all fail to match downgrades the whole family
silently instead of erroring.

Merge's single reused map is split into two reported categories. A URI match
means the gallery already ships exactly these weights and referencing the
existing entry is correct; a name collision means an unrelated entry owns the
name and referencing it would substitute a different build.

Multimodal children now declare known_usecases [chat, vision]. An explicit
known_usecases suppresses the backend-default fallback, so a chat-only entry
carrying an mmproj never matches the vision or multimodal gallery filters.

.github/ci is invisible to go list ./..., so a workflow names both generator
packages explicitly and their specs finally run on pull requests.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ci): gather APEX builds under the base model entry

The hub for a family is the BASE model entry, never a generated *-apex
parent. Somebody looking for qwen3.6-35b-a3b has to find every build of
those weights under that one name, so a competing qwen3.6-35b-a3b-apex
hub would split the family and leave half of it invisible.

When the gallery already ships the base entry, a variants block is
spliced into it textually, leaving its description, icon, tags,
overrides and files untouched. Only a family whose base model the
gallery does not ship gets a new hub, still named for the base model and
carrying one of the discovered builds as its own payload so it declares
a backend the verifier can judge.

The line editing is factored into .github/ci/galleryedit, shared with
the variantproposals job, so the two cannot drift apart on where a
variants block belongs.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(apexentries): treat an unreadable optional counterpart repo as absent

HuggingFace answers 401 Unauthorized, not 404, for a repository that does
not exist when the request carries no credentials. FetchRepoFiles treated
only 404 as absence, so probing for the OPTIONAL unsloth counterpart hard
failed for every family that legitimately has none: 27 of the 45 APEX
families are community merges that will never have an unsloth build, and a
full run failed all of them.

Split the fetch so the two call sites can apply different policies to the
same response. The APEX repo itself stays strict: a 401 or 403 on a repo
the run requires is a real failure and still errors. Only the optional
probe tolerates it, because without a token 401 cannot be told apart from
absence.

That collapse is lossy in one direction, since a private or gated repo also
answers 401, so the skipped candidates are named in the run summary
alongside the other silent-shortfall counters instead of being dropped in
silence.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* ci(apexentries): report full-precision sources as a known exclusion

The 45 APEX repos publish their unquantized F16 sources next to the
imatrix ladder, flat or sharded. Discovery correctly emits nothing for
them, but they were landing in the unclassified total, leaving a
permanent baseline of 24 benign lines on every run.

That baseline is what the unclassified check exists to prevent: a
standing count of known-benign files is exactly what hides the one file
that ever genuinely matters. Count full-precision sources separately and
give them their own summary line, so unclassified returns to 0 and stays
loud when something really is an unknown shape.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(apexentries): namespace local paths by owner and enable MTP builds

localPath namespaced downloads by the repo basename alone, so two repos
publishing the same filename under different owners collapsed to one local
path. LiquidAI/LFM2.5-8B-A1B-GGUF and unsloth/LFM2.5-8B-A1B-GGUF collided that
way, and both were offered from the same hub, so installing the second either
overwrote the first model's weights or was skipped as already present while
recording a sha256 that did not match the bytes on disk. The owner is now its
own path segment: owner/repo is globally unique on HuggingFace and neither half
can contain a separator, so uniqueness holds by construction.

Verify gains a check for the whole class, that no local filename may map to two
different upstream URIs. It surfaces seven pre-existing collisions in the
gallery, which are left alone here.

Entries built from the *-APEX-MTP-GGUF repos now configure MTP rather than
shipping the heads inert, matching the pattern the hand-written MTP entries
already use: spec_type:draft-mtp with spec_n_max and spec_p_min, tagged mtp, and
no draft_model because the heads live in the weights. RenderChild no longer
requires a separate drafter file before it will configure a spec type, while the
cross-repo drafter path is unchanged.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add the APEX GGUF families as variant ladders

Adds the imatrix quality ladder from each mudler/*-APEX-GGUF repo, a fixed
subset of unsloth quant rungs where a counterpart repo exists, and the MTP
builds, then attaches them to the base model entry so one entry offers every
build of the same weights and LocalAI picks the one that fits the hardware.

Ten existing base model entries gain a variants list; twenty-seven families that
the gallery had no base entry for get one. Builds are discovered from the
filenames each repo actually publishes rather than derived from its name, since
six repos ship a stem that differs from their repo name. Every file carries a
sha256 taken from the HuggingFace API.

Assisted-by: Claude Opus 4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-21 21:40:51 +02:00
mudler's LocalAI [bot]
a4a181d2f7 fix(distributed): count staging verification as progress, not as a stall (#11026)
Testing the progress-based cold-load deadline on the live cluster surfaced a
false positive. The stall window observed UPLOAD bytes only, but the staging
path has a phase that does real work while moving zero upload bytes: the
resumable-upload verify phase.

When a shard is already present on the worker from an earlier attempt, the
frontend HEADs it, hashes the local copy to confirm it matches, and skips the
transfer. Staging a 70 GB model with 56 GB already staged:

  17:27:34 INFO Upload skipped (file already exists with matching hash) ...
  17:28:20 INFO Upload skipped (file already exists with matching hash) ...
  17:29:07 INFO Upload skipped (file already exists with matching hash) ...
  ... six-plus consecutive minutes, no bytes uploaded at all

~45s per skipped ~4 GB shard. That is correct and desirable - it is what makes
resume work - but it was indistinguishable from a stall. At 45s per shard it
sits inside the 5m window, so the run in flight was fine; the problem is the
600 GB scale this machinery exists to enable, where one shard can plausibly hash
for longer than the window. The guard would then fire during verification of a
transfer that is working perfectly.

Verified mechanism: probeExisting() HEADs the worker and then calls
downloader.CalculateSHA(). The staging progress callback is only consulted
inside doUpload(), which the skip path never reaches, so observeLoadProgress was
called zero times for the whole verify phase.

Verification exposed a second, worse bug in the same path: CalculateSHA consults
no context at all. An expired cold load kept hashing to completion, compared the
hashes, and returned success - reporting a file as staged on a dead load. The
failure only surfaced on the NEXT file, whose HEAD died immediately. That is
exactly the shape of the red test here, which fails on shard 3.

Fix: hash in 1 MiB chunks via hashFileWithActivity(), ticking the cold-load
deadline per chunk and checking ctx per chunk. A successful HEAD also counts,
since a 200 with a content hash proves the worker is serving right now.

Counting hash progress does not make a dead transfer look alive: hashing is
bounded, terminating work proportional to file size, in probeExisting it runs
only after a HEAD proved the worker was up, and the 24h absolute cap still
bounds the whole hold. The alternative of simply widening the window was
rejected - it would reintroduce the size cliff this work removes.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-21 21:38:54 +02:00
mudler's LocalAI [bot]
2b61e4bc1d fix(upgrade-check): don't filter upgrade candidates by controller capability (#11024)
CheckUpgradesAgainst resolved gallery entries through AvailableBackends,
which drops every entry the *local* host cannot run. In distributed mode
the host running the check is a CPU-only controller while the GPU
backends live on worker nodes, so FindGalleryElement returned nil for
every cuda/rocm/l4t entry and those backends were silently skipped.

Measured on a live cluster: GET /backends reported 48 installed
backends, POST /backends/upgrades/check evaluated 5 — all of them plain
or cpu-prefixed. The 43 skipped were all hardware-specific builds. As a
result cuda13-nvidia-l4t-arm64-longcat-video-development stayed at
sha256:0b8dc851 while the registry tag held sha256:38dae6ff, and a cuDNN
packaging fix sat unnoticed on a GPU worker for two days.

Every name looked up here is already installed somewhere in the cluster,
so hardware compatibility was decided at install time; re-deciding it
against the controller is wrong. Switch both CheckUpgradesAgainst and
UpgradeBackend to AvailableBackendsUnfiltered.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-21 19:28:34 +02:00
mudler's LocalAI [bot]
3584e0776d chore: ⬆️ Update antirez/ds4 to efdadd41e20134af4f3381e1ed90e96fe4faef6f (#11010)
* ⬆️ Update antirez/ds4

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(ds4): link new tensor parallel objects

The updated ds4 revision split tensor-parallel transport and layer placement into separate translation units. Build and link those objects on CPU, CUDA, and Metal builds.

Assisted-by: Codex:gpt-5

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-21 16:06:30 +00:00
dependabot[bot]
4ce67ccb84 chore(deps): bump body-parser from 2.2.2 to 2.3.0 in /core/http/react-ui in the npm_and_yarn group across 1 directory (#11016)
chore(deps): bump body-parser

Bumps the npm_and_yarn group with 1 update in the /core/http/react-ui directory: [body-parser](https://github.com/expressjs/body-parser).


Updates `body-parser` from 2.2.2 to 2.3.0
- [Release notes](https://github.com/expressjs/body-parser/releases)
- [Changelog](https://github.com/expressjs/body-parser/blob/master/HISTORY.md)
- [Commits](https://github.com/expressjs/body-parser/compare/v2.2.2...v2.3.0)

---
updated-dependencies:
- dependency-name: body-parser
  dependency-version: 2.3.0
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-21 15:35:01 +02:00
mudler's LocalAI [bot]
b700a78ae4 fix(distributed): make the cold-load hold scale with progress, not wall-clock (#11019)
A 70 GB video checkpoint (longcat-video-avatar-1.5) could not be loaded on a
distributed cluster. The request failed with HTTP 500 after 1499.98s - exactly
the 25m00s cold-load ceiling - while staging was demonstrably healthy: 26 of 57
files and 39 GB transferred at a sustained ~26 MB/s, zero errors, no stalls. It
was not wedged, it was killed by a timer.

ModelLoadCeilingFor covers node selection, backend install, file staging and the
remote LoadModel. Install and load carry their own budgets; staging was covered
only by a FIXED 5-minute margin. But staging time is bytes over bandwidth, not a
constant: 70 GB at 26 MB/s needs ~45m against a 25m ceiling, so the failure is
deterministic for any sufficiently large model rather than a flake. Simply
raising the constant moves the cliff to the next model size - the deployment
target here is checkpoints of 600 GB and beyond.

The ceiling's real purpose is that "a wedged worker can never pin the lock
indefinitely". Progress, not elapsed time, is what distinguishes a wedged worker
from a large one. The hold is now a deadline that extends whenever the transfer
reports bytes and expires a 5-minute stall window after they stop:

- A large model transferring fine continues, for hours if needed.
- A worker that died mid-transfer still fails within the stall window.

Progress is observed at byte level on the transfer itself, via the existing
staging progress callback. Per-file completion would be too coarse - a single
600 GB shard would be indistinguishable from a stall for hours. The observation
point is back-pressured by the socket, so it reflects the network rather than
local disk reads. Observation is coarsened to one timer touch per stall/20 so
the per-read callback stays cheap.

The base budget (unchanged, and still derived from the install and load
timeouts) continues to cover the steps that report no progress, so
LOCALAI_NATS_MODEL_LOAD_TIMEOUT keeps working exactly as before. An absolute
cap of 24h bounds the hold even while progress keeps arriving, so a peer
trickling bytes forever cannot pin the advisory lock; 600 GB at the measured
26 MB/s is ~6.5h, so the cap sits far above any legitimate transfer.

Also fixes the incoherent layering the same error exposed: the resumable upload
carried a 1h retry budget nested inside the 25m ceiling, so the inner budget was
unreachable and the message still blamed it ("failed after 1 attempts within
1h0m0s budget") while the 25m parent was the actual killer. The upload now
adopts the caller's deadline when there is one, and applies its fixed budget
only when nothing above bounded it - which also stops a fixed 1h from
reintroducing the size cliff under the now-extendable parent.

This is the successor to #10968, where a hardcoded 5-minute LoadModel gRPC
timeout was replaced by this derived ceiling. Fixing the inner timeout exposed
the outer ceiling as the new binding constraint.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-21 15:34:33 +02:00
Richard Palethorpe
0d2124894e docs(realtime): fix Opus backend installation (#11018)
The Realtime guide incorrectly sent the Opus backend through the model gallery endpoint. Point users to the backend gallery API and document the UI and CLI alternatives.

Assisted-by: Codex:gpt-5

Signed-off-by: Richard Palethorpe <io@richiejp.com>
2026-07-21 15:33:55 +02:00
mudler's LocalAI [bot]
54d5c18bfb fix(model): only announce a load at INFO when a load actually happens (#11017)
backendLoader logged "BackendLoader starting" at INFO as its very first
statement, unconditionally. That reads as "a model is being loaded", but
backendLoader is not only a load path: in distributed mode Load()
deliberately bypasses the local cache and calls backendLoader on every
inference request so SmartRouter can re-pick a replica per request. The
model is already resident, no process is spawned, and nothing is loaded,
yet the banner fires at request rate.

On a live cluster this produced ~5 "BackendLoader starting" lines per
second for a single embedding model, sustained, starting 22 seconds
after the load had already completed. The model was state=loaded with
in_flight=0 and exactly one backend process on the worker. It looked
exactly like a retry storm and cost real debugging time during an
unrelated production investigation. The adjacent "effective runtime
tuning" banner, documented as "logged once per load", had the same
problem for the same reason.

Emit both banners at INFO only when the model is not already resident,
and keep the per-call trace at DEBUG for anyone following the routing
path. isResident is a plain store lookup with no health probe and no
eviction, so it is safe on the per-request hot path (unlike
checkIsLoaded, which probes and can evict).

Same class of defect as #10985: a log line that sends the reader after
the wrong thing.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-21 15:33:29 +02:00
mudler's LocalAI [bot]
49fcdd921f chore: ⬆️ Update CrispStrobe/CrispASR to 644a8b1b31ca42e26f641df38e323ed9d698a1ff (#11007)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-21 15:33:00 +02:00
mudler's LocalAI [bot]
e887a1ccf3 chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to ba4c7f7838ecb24a75b0ac94e14fdbebb6bb138c (#11006)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-21 15:32:46 +02:00
mudler's LocalAI [bot]
5fbd79a4bb chore: ⬆️ Update PrismML-Eng/llama.cpp to 7529fdaaf99ffdc5ca71ace9c7409a56b27ad92f (#11009)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-21 15:32:26 +02:00
dependabot[bot]
2a8eb5a04b chore(deps): bump the npm_and_yarn group across 1 directory with 5 updates (#11011)
Bumps the npm_and_yarn group with 5 updates in the /core/http/react-ui directory:

| Package | From | To |
| --- | --- | --- |
| [dompurify](https://github.com/cure53/DOMPurify) | `3.4.0` | `3.4.11` |
| [hono](https://github.com/honojs/hono) | `4.12.18` | `4.12.31` |
| [qs](https://github.com/ljharb/qs) | `6.15.0` | `6.15.3` |
| [react-router](https://github.com/remix-run/react-router/tree/HEAD/packages/react-router) | `7.13.1` | `7.18.1` |
| [undici](https://github.com/nodejs/undici) | `7.25.0` | `7.28.0` |



Updates `dompurify` from 3.4.0 to 3.4.11
- [Release notes](https://github.com/cure53/DOMPurify/releases)
- [Commits](https://github.com/cure53/DOMPurify/compare/3.4.0...3.4.11)

Updates `hono` from 4.12.18 to 4.12.31
- [Release notes](https://github.com/honojs/hono/releases)
- [Commits](https://github.com/honojs/hono/compare/v4.12.18...v4.12.31)

Updates `qs` from 6.15.0 to 6.15.3
- [Changelog](https://github.com/ljharb/qs/blob/main/CHANGELOG.md)
- [Commits](https://github.com/ljharb/qs/compare/v6.15.0...v6.15.3)

Updates `react-router` from 7.13.1 to 7.18.1
- [Release notes](https://github.com/remix-run/react-router/releases)
- [Changelog](https://github.com/remix-run/react-router/blob/react-router@7.18.1/packages/react-router/CHANGELOG.md)
- [Commits](https://github.com/remix-run/react-router/commits/react-router@7.18.1/packages/react-router)

Updates `undici` from 7.25.0 to 7.28.0
- [Release notes](https://github.com/nodejs/undici/releases)
- [Commits](https://github.com/nodejs/undici/compare/v7.25.0...v7.28.0)

---
updated-dependencies:
- dependency-name: dompurify
  dependency-version: 3.4.11
  dependency-type: direct:production
  dependency-group: npm_and_yarn
- dependency-name: hono
  dependency-version: 4.12.31
  dependency-type: indirect
  dependency-group: npm_and_yarn
- dependency-name: qs
  dependency-version: 6.15.3
  dependency-type: indirect
  dependency-group: npm_and_yarn
- dependency-name: react-router
  dependency-version: 7.18.1
  dependency-type: indirect
  dependency-group: npm_and_yarn
- dependency-name: undici
  dependency-version: 7.28.0
  dependency-type: indirect
  dependency-group: npm_and_yarn
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-21 09:40:42 +02:00
mudler's LocalAI [bot]
9c8f510021 chore(model gallery): 🤖 add 1 new models via gallery agent (#11013)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-21 09:40:24 +02:00
localai-org-maint-bot
1e0baec2a7 fix(ci): repair nightly backend dep bumps for renamed localai-org repos (#11012)
The "Bump Backend dependencies" workflow has failed every night for over
ten days. Four upstreams — ced.cpp, moss-transcribe.cpp, voice-detect.cpp
and rf-detr.cpp — moved from the mudler org to localai-org, so the GitHub
API answers 301 for the old slugs. ced.cpp additionally renamed its
default branch to main.

bump_deps.sh fetched without -L or -f and never checked the response, so
the redirect's JSON body was passed straight to sed, which died with
"unterminated `s' command". The loud failure was luck: an error body
without slashes would have been substituted into the Makefile as the new
pin, silently corrupting the version and shipping it in a bump PR.

Point the matrix at the new slugs and branch, and harden the script so a
bad response can never reach sed: follow redirects, fail on HTTP errors,
and require a bare 40-hex SHA before rewriting anything. Also refresh the
now-stale repository URLs in the backend Makefiles, test scripts,
backend/index.yaml and the docs.

Verified all 25 matrix entries resolve to a commit SHA and that the four
previously-failing jobs run end to end against the real API.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-21 09:40:10 +02:00
mudler's LocalAI [bot]
7bda73fd66 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to f39cc4a3af988091f662313b336dddf8c83a3fb5 (#11002)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-21 09:39:35 +02:00
mudler's LocalAI [bot]
a2c87947a9 chore(model-gallery): ⬆️ update checksum (#11005)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-21 08:48:14 +02:00
mudler's LocalAI [bot]
7c984f5c81 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260720105820 (#11003)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-21 08:48:02 +02:00
mudler's LocalAI [bot]
f8755997cc feat(swagger): update swagger (#11001)
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-21 08:47:02 +02:00
mudler's LocalAI [bot]
d0401f9bb4 chore: bump go-processmanager to pick up the concurrent-Run fix (#11004)
Picks up mudler/go-processmanager#7, which closes the check-then-act
window in Run's "already started" guard. Concurrent Run calls on one
handle could each observe a nil p.proc, each start a process and each
launch a monitor goroutine; every monitor does `defer close(p.done)`
against a channel created once in New, so the second monitor to see its
process exit panicked on close of a closed channel. The same window
raced on p.proc itself.

LocalAI does not hit this today: the only New/Run pair
(pkg/model/process.go:178-191) runs a freshly created handle, and
core/services/worker/supervisor.go only ever calls Stop on handles it
holds. The bump is defensive, and keeps the dependency from drifting
further behind a fix in the process lifecycle we rely on.

No exported signature changes upstream; ErrProcessAlreadyRun is
additive and keeps the historical "command already started" prefix, so
any caller matching on that text is unaffected. Nothing in this repo
matches it.

Verified: go build ./core/... ./pkg/... clean; go vet clean;
go test -race ./pkg/model/ ./core/services/worker/ both ok.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Bash] [Edit]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 23:32:29 +02:00
mudler's LocalAI [bot]
0cdd781c2d ci(gallery): propose variant groupings for review instead of letting them decay (#10992)
ci(gallery): propose variant groupings for review on a schedule

A gallery entry may declare `variants:`, references to other entries that
are alternative builds of the same weights, and auto-selection then installs
the best build for the host. Those families exist only because humans curated
them in two manual sweeps.

The gallery agent dedupes on the HuggingFace repo URL and picks one
quantization per model, so it never adds a second build of a repo it already
has, and consequently never creates a family and never joins one. A model
published across two repos lands as two unrelated standalone entries. The
grouping decays as the gallery grows and nothing notices.

Add a scheduled job, in the same shape as the checksum checker: compute
offline, edit the index textually, open a pull request against ci-forks. It
proposes and never decides. Grouping is a judgement call that has gone wrong
in both directions, so the value is catching drift and surfacing candidates
with their evidence.

Three grouping signals, taken from the manual sweeps: same name once
quantization markers are stripped, the `:` config-suffix convention, and the
same primary weight filename once quantization markers are stripped. The
third requires the same upstream repository. Excluding auxiliary files is not
enough on its own: bert-embeddings, an ultravox audio model and a roleplay
finetune all declare a primary file called llama-3.2-1b-instruct-q4_k_m.gguf,
and grouping on that is the same error that linked four wan-2.1 entries and
Z-Image-Turbo to qwen3-4b.

Add gallery/variant-exclusions.yaml, a checked-in rejection ledger. A job
that re-proposes declined candidates every night becomes noise and gets
ignored. Declining a proposal is one flow-mapping line a reviewer adds inside
the proposal pull request itself. Seeded with the six -abliterated pairs whose
base is also in the gallery, the mistral-small multimodal pair, the whisper-1
alias, the kokoros language set, and the recurring finetune tokens. qat and
apex are deliberately not on it: they are quantization techniques.

Proposals refuse to nest, to let two parents claim one target, to target an
entry that installs nothing, and to touch a merge anchor, since a variants key
added to an anchor is inherited by every merging child. The anchor refusal
names every entry that would inherit, which is the worklist a human needs.

Run against the pre-sweep gallery, the job rediscovers 12 of the 19 groupings
the second manual sweep made, with no false positives. The rest it reports as
refusals or ledger declines rather than missing silently.

Assisted-by: Claude:claude-opus-4-8

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 23:08:00 +02:00
mudler's LocalAI [bot]
f01038f479 fix(modelartifacts): stage each writer's artifact in its own partial tree (#10995)
Every writer used to stage into the same `.artifacts/.partial/<cacheKey>`.
That was safe only because the artifact lock held: two writers that both
believed they had it opened the same blob with O_APPEND and interleaved
their bytes into one file, while the resume probe read the other writer's
in-flight size. SHA verification caught the damage only after both had
burned the entire download.

#10986 restored the lock's precondition on CIFS but left the dependency in
place. Suffix the staging tree with a writer identity drawn once per
process run, so concurrent writers cannot corrupt each other whatever the
lock does. The lock stops being a correctness dependency and becomes a
pure efficiency optimisation: a lock failure now costs a duplicated
download, not a corrupted one.

Commit stays an atomic rename. The loser of a commit race reconciles onto
the winner's tree instead of surfacing a bare ENOTEMPTY for work that
actually succeeded, since the artifact is content-addressed and both trees
hold the same verified bytes.

Writer-unique staging means a crashed writer's tree is no longer
overwritten by its successor, so two things are added to keep it from
becoming a disk leak and a resume regression:

- A sweep reclaims trees whose contents have been untouched for 24h,
  matching the window the startup reaper already uses for stray *.partial
  files. It reads the newest mtime anywhere inside the tree, because
  writing a blob never touches an ancestor, and refuses any name this
  package did not write. A live download writes continuously, and the
  downloader's stall watchdog aborts a silent one long before it could
  look abandoned.

- Adoption lets a restarted process claim a dead predecessor's tree for
  the same artifact and resume from its bytes, which a tens-of-gigabytes
  repo depends on. The claim is an atomic rename, so racing adopters
  cannot both win. It runs only under the artifact lock - which is
  released exactly when the owning process dies - and only on a tree idle
  for 5 minutes as a second line of defence for when the lock does not
  exclude.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 23:07:43 +02:00
mudler's LocalAI [bot]
0eb8a1188d fix(worker): give the worker a real health endpoint and a mode-aware HEALTHCHECK (#10999)
fix(worker): give the worker a real health endpoint (#10987)

The image bakes in a single HEALTHCHECK that curls
http://localhost:8080/readyz, but the same image also runs `local-ai
worker`, which serves HTTP on the gRPC base port minus one and never
binds 8080. Every worker container was therefore permanently
`unhealthy` (43 consecutive failures observed on a production node),
which is worse than having no healthcheck: a genuinely broken worker and
a perfectly good one both report `unhealthy`, so the signal carries no
information and orchestration that keys on it misbehaves.

The worker already served /readyz on that port via the file-transfer
server, but as a constant 200 — it only proved the listener was bound,
which is precisely the failure mode at issue. Readiness now tracks the
live NATS connection: all of a worker's actual work (backend lifecycle
events, inference dispatch, file staging) arrives over NATS, so a worker
whose link is dead is up and useless. Registration is already implied,
since the server only starts after registration succeeds.

This reports something the controller cannot already see. The node
registry's status/last_heartbeat is fed by an HTTP heartbeat to the
frontend, a different network path from NATS — a worker can keep
heartbeating while its NATS connection is dead and still look healthy in
the registry. /healthz stays a constant 200: liveness must not follow
readiness, or a NATS blip becomes a cluster-wide restart storm.

The HEALTHCHECK is now a script that derives its endpoint from the mode
the container is actually running plus the env vars that configure the
bind address, so a frontend moved off 8080 with LOCALAI_ADDRESS (broken
the same way) and a worker on a non-default base port are both probed
correctly. Modes with no HTTP surface (agent-worker, one-shot commands)
report healthy rather than false-unhealthy. HEALTHCHECK_ENDPOINT remains
as an explicit override, so the workaround shipped in
docker-compose.distributed.yaml keeps working; both overrides in that
file are now unnecessary and have been removed.

Also fixes the latent --start-period gap. Since #10949 a frontend's
startup preload materializes HuggingFace artifacts before the HTTP
server binds (31 GB observed on a live cluster), so a healthy replica
can legitimately fail probes for a long time. --start-period is Docker's
knob for exactly this: failures inside it leave the container `starting`
instead of burning retries, and it ends early on the first success, so a
generous 60m costs a fast-starting container nothing. --timeout drops
from 10m to 10s — it is a per-probe deadline, and a localhost curl that
has not answered in 10s is itself the fault being detected.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 23:07:27 +02:00
mudler's LocalAI [bot]
d7e04dcc32 fix(openresponses): make responses visible and cancellable across replicas (#11000)
In distributed mode the Open Responses store is process-local: a
sync.OnceValue over a map behind an RWMutex. With several frontend
replicas behind a round-robin load balancer, every request that lands on
a replica other than the creator misses.

Measured on a live 2-replica cluster (#10993): the same response id
returns 200 on the creating replica and 404 on its peer, and a cancel on
the peer returns 404 without ever invoking CancelFunc, so generation runs
to completion on the other replica while the caller is told the response
does not exist. previous_response_id chaining fails through the same
lookup.

Split the state by what can actually cross a process boundary:

- Replicated: response metadata (request, response resource, owner,
  expiry, stream/background flags) via syncstate.SyncedMap, the same
  component finetune, quantization and agent tasks already use. A local
  miss in Get/FindItem now falls back to it and returns a read-only
  remote view, so polling and chaining resolve on any replica.

- Delegated: cancellation. context.CancelFunc is a function pointer and
  exists only in the creating process, so a cancel that lands elsewhere
  is broadcast on responses.<id>.cancel and applied by whichever replica
  holds the function. The broadcast is fire-and-forget rather than
  request/reply: if the owner crashed or was scaled down nobody answers,
  and the handler must not block on a reply that will never come. The
  replicated status moves to cancelled either way, which is truthful,
  since a dead owner's generation died with its process.

- Refused: streaming resume. The resume buffer is a byte log plus a live
  notification channel and cannot be replicated without shipping every
  token over the bus. A resume that reaches the wrong replica now returns
  HTTP 409 naming the owning replica via the new ErrResponseNotLocal,
  instead of an empty event list that looks like a finished stream. It is
  deliberately distinct from ErrOffsetLost, which means the owner's
  buffer evicted the requested events.

Standalone deployments never call EnableDistributed and keep exactly the
previous process-local behaviour.

Fixes #10993


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 23:06:33 +02:00
mudler's LocalAI [bot]
a4bab71f27 gallery: remove duplicated entries and lint against them recurring (#10996)
gallery: remove duplicated entries and lint against them coming back

gallery/index.yaml declared eight names twice: deepseek-r1-distill-llama-8b,
llama3.2-3b-enigma, qwen3-asr-0.6b, qwen3-asr-1.7b, qwopus-glm-18b-merged,
voice-en-us-kathleen-low, whisper-large-q5_0 and whisper-small-q5_1.

FindGalleryElement resolves a reference by returning the first match, so in
every pair the second copy was unreachable: it could not be installed, could
not be selected as a variant target, and could not be corrected, because any
edit to it went to a copy nobody reads. A reference to such a name is also
ambiguous to anything reasoning over the catalog, which is why the variant
proposal job refuses to propose against them.

Each pair was compared both as parsed entries and as raw text, and all eight
were byte-identical apart from position. None of the sixteen blocks defines a
YAML anchor or pulls one in with a merge key, so nothing was reachable only
through a deleted block, and no entry named a removed copy as a variant target.
Removing the second copy of each therefore changes no behaviour: the parsed set
loses exactly eight entries and every surviving entry is field-for-field
unchanged.

The removal is textual, by line range, so the diff is pure deletions rather than
a reflow of forty thousand lines.

checkNoDuplicateEntryNames is the rule that keeps them out, added beside the
existing gallery invariants and reporting in the same style.

checkSingleVariantClaim closes the adjacent gap in the same place. VariantParents
resolves a build claimed by two parents by taking the first in gallery order and
calls that deterministic "for a gallery the linter would reject", but nothing
rejected it: the invariant was held by curation alone. Now it is a rule, and the
comment describes something real. No target is doubly claimed today, so the rule
is green on arrival.

Assisted-by: Claude:claude-opus-4-8

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 22:58:48 +02:00
mudler's LocalAI [bot]
2f33ad6669 fix(modelartifacts): treat CIFS EACCES as lock contention, not failure (#10986)
flock(2) on CIFS/SMB returns EACCES when another client holds the lock:
the kernel maps STATUS_LOCK_NOT_GRANTED and STATUS_FILE_LOCK_CONFLICT to
-EACCES and never produces EWOULDBLOCK on that path. gofrs/flock only
recognises EWOULDBLOCK as contention, so TryLockContext returned a bare
"permission denied" and Ensure aborted. Both replicas then fell back to
legacy loading, which makes the worker download the whole repo in-band
inside LoadModel and blow the remote-load deadline.

Replace TryLockContext with an explicit wait loop over a new Locker
interface, classifying EWOULDBLOCK/EAGAIN/EACCES/EBUSY as contention.
EACCES is ambiguous at the syscall boundary but not here: the lock file
is already open O_CREATE|O_RDWR, so a real permission problem would have
failed the open with an *fs.PathError, and flock(2) documents no EACCES
on Linux at all. The wait is bounded (DefaultLockWait, overridable via
WithLockWait), so even a misclassification degrades to a delay. On
timeout the committed result is re-checked before reporting the new
ErrLockContended, so a peer that finished the work still wins.

Locker also exists so the contention path is testable without a network
filesystem: nothing in CI can make flock(2) return EACCES on demand.

Raise the fallback to error for a managedArtifactBackends backend, via a
shared config.LogArtifactFallback used by both call sites. For those
backends the legacy path is not graceful degradation, and the operator
otherwise sees only a timeout with no causal link. The fallback stays
non-fatal.

Drop the os.Chmod(layout.Lock, 0o600) after acquisition: flock.New
already creates the file 0600, and the chmod was gratuitous risk on a
nounix mount that ignores modes.

Fixes #10981


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 22:12:39 +02:00
mudler's LocalAI [bot]
1cd7d63c7b fix(distributed): reject wrong-model requests on the remaining modalities (#10990)
#10970 gave the four PredictOptions RPCs a model-identity check so a
backend reached through a stale distributed route rejects the request
instead of answering from whatever model it holds (#10952). Every other
modality shares that exposure: the route is cached by host:port, a worker
can recycle a stopped backend's port for another model's backend, and a
liveness-only probe cannot tell a stale row from a valid one.

Extends the same mechanism to the 21 remaining request messages that reach
a backend through the router, using the pattern #10970 established rather
than a parallel one:

- proto: ModelIdentity on each modality request message.
- controller: populated from ModelConfig.Model at the call site that also
  builds ModelOptions, so load-time and request-time values are equal by
  construction.
- backends: one generic guard in pkg/grpc/server.go (27 Go backends), the
  method set in backend/python/common (36 Python backends), llama-cpp
  (AudioTranscription/Stream, Rerank, Score) and privacy-filter
  (TokenClassify).
- reconcile already drops the stale row on IsModelMismatch; no change.

TTSRequest and SoundGenerationRequest get a SEPARATE ModelIdentity field
rather than reusing their existing `model`: FileStagingClient rewrites
`model` to a worker-local path, so comparing it would reject valid
requests in exactly the configuration this guards.

AudioEncode/AudioDecode are deliberately left unguarded: the opus codec
backend is loaded from a literal rather than a ModelConfig, so no value
carries the equality guarantee the comparison depends on. The four
bidirectional stream RPCs are out of scope; they bypass reconcile.

Empty means skip on both sides, so an old controller, an old backend, and
the bare request structs in tests/e2e-backends all keep working.


Assisted-by: Claude Code:claude-opus-4-8 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 21:58:19 +02:00
mudler's LocalAI [bot]
a784cf669f fix(ci): rebuild Go backends on linked pkg/ changes and on matrix entry edits (#10988)
PR #10975 taught the backend matrix filter about shared build inputs, but
left two paths that still rebuild nothing.

Go backends link code from the main tree. `go list -deps ./backend/go/...`
resolves to exactly six pkg subtrees (audio, grpc incl. base/grpcerrors/proto,
httpclient, sound, store, utils), identical for GOOS/GOARCH in
{linux,darwin} x {amd64,arm64}. Editing any of them changes the shipped
binary, but they sit outside every backend directory so the prefix match
never saw them. Enumerating those six rather than taking all of pkg/ is the
point: all of pkg/ changes in ~8.6% of commits, these six in 2.0% — the same
order as the already accepted scripts/build/ rule (1.9%). Blast radius
199/417 Linux, 26/56 Darwin; the ~21 core-server-only pkg subtrees still
rebuild nothing, and neither do _test.go files.

.github/backend-matrix.yml was excluded wholesale because matching its path
would rebuild all 417 entries on every new-backend PR. That hid a real hole:
editing an existing entry's base-image, build-type or cuda version changes
the image it produces while touching no file the filter can see. Since the
change is within a structured file, compare it against the base revision and
rebuild only the entries whose fields actually differ — 1 entry for a
base-image edit, 0 for a comment or whitespace edit, and all 417 only when
the previous revision cannot be resolved. This also closes a third hole: a
new matrix entry for an existing backend (a new CUDA variant, say) touches
nothing under that backend's directory and previously rebuilt nothing.

changed-backends.js fetches the base revision via the contents API, and only
when the changed-file list actually names the matrix file, so the common path
costs no extra request.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 21:46:37 +02:00
mudler's LocalAI [bot]
65bdbc4ee3 fix(http): make /readyz reflect startup readiness, plus gitignore and coverage-ratchet fixes (#10989)
* fix(http): make /readyz reflect startup readiness instead of always 200

/readyz was registered as a static handler returning 200 unconditionally,
so it carried no information: it was green whenever it could be reached at
all. Readiness could not distinguish "serving" from "still starting", and
any future change that started the HTTP listener earlier would silently
turn the probe into a lie.

Track startup completion on the Application (atomic flag, flipped at the
very end of New() on the success path only) and have the readiness handler
consult it per request, returning 503 with a small JSON body while startup
is in progress. A nil readiness source fails open so embedders keep the
historical behaviour.

/healthz is deliberately left readiness-independent. Liveness and readiness
answer different questions, and failing liveness during a long preload makes
an orchestrator restart the pod mid-download so the preload never finishes.

This matters because since #10949 the startup preload materializes
HuggingFace artifacts for managed backends: tens of GB for a large model
(31 GB observed on a live cluster). Both probes stay in quietPaths and stay
exempt from auth.

Note the listener is still started only after New() returns, so today the
not-ready state is not observable over HTTP. Moving the listener earlier is
a separate, deliberate decision and is not made here.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* chore(gitignore): anchor the mock-backend pattern so its source dir is traversable

The bare `mock-backend` pattern matched the *directory*
tests/e2e/mock-backend/, not just the binary built into it. Git will not
descend into an ignored directory even for tracked files, so
`git add tests/e2e/mock-backend/main.go` required -f. This was hit while
working on #10970.

Anchor it to the artifact's full path. The built binary stays ignored (it is
also covered by tests/e2e/mock-backend/.gitignore) while the source directory
becomes traversable again.

Verified with `git check-ignore -v`: a new source file under
tests/e2e/mock-backend/ is no longer ignored, and the binary produced by
`make build-mock-backend` still is.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* chore(coverage): raise the coverage ratchet from 48.5% to 54.2%

The committed baseline had drifted well below reality: it still read 48.5%
while a full instrumented run measures 54.2%. A stale-low baseline makes the
gate meaningless — coverage could regress by more than 5 percentage points
and still pass.

Raising a ratchet is a deliberate act, not something to fold into an
unrelated fix, so it gets its own commit. The headroom was earned by tests
landed in #10946, #10947, #10948, #10949, #10956, #10967, #10968, #10970 and
#10975.

Measured with `make test-coverage` on this branch (the same instrumented run
`make test-coverage-baseline` uses: ginkgo over ./pkg and ./core plus the
in-process tests/e2e suite, --covermode=atomic, --coverpkg over core/... and
pkg/..., generated protobuf excluded). The run completed with exit 0 and zero
spec failures; the total was then written with the exact command the
test-coverage-baseline target uses:

  go tool cover -func=coverage/coverage.out \
    | awk '/^total:/{gsub(/%/,"",$NF); print $NF}' > coverage-baseline.txt

Verified afterwards with scripts/coverage-check.sh, which reports OK.

Note the measured figure includes the readiness specs added earlier on this
branch, so it is a demonstrated floor rather than an estimate.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 21:45:27 +02:00
mudler's LocalAI [bot]
0d9d07d3a5 fix(downloader): distinguish read from write failures and retry transient ones (#10985)
Two independent defects in the download path, both surfaced by the same
incident (#10982).

A failed `io.Copy` was always reported as "failed to write file", because
`io.Copy` folds read and write errors into a single return value. A
peer-cancelled HTTP/2 stream therefore presented as a filesystem failure and
sent an investigation after mount permissions while the disk was healthy. The
source is now wrapped in a recorder so the error names the side that actually
broke, and a write failure names the `.partial` it was writing rather than the
final blob path.

The plan runner returned on the first task error with no retry, so one
transient stream cancel discarded every file already downloaded in a
multi-file materialization. The `.partial` resume machinery already existed
but was unreachable because nothing made a second attempt. Transient failures
(dropped transport, mid-stream read failure, stall, 5xx, 429) are now retried
with bounded exponential backoff and resume from the partial; permanent ones
(4xx, checksum mismatch, local write failure, caller cancellation) fail
immediately.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 19:56:09 +02:00
mudler's LocalAI [bot]
f381844403 gallery: group QAT, APEX and cross-backend builds under their base entry (#10983)
feat(gallery): group 21 more model families under variants

Second variant-grouping sweep. QAT and APEX builds are now treated as
quantization techniques rather than distinct weights, per maintainer
ruling, so they group with their base entry instead of standing alone.

Adds 15 new families: quantization and serving-config pairs for
llama-3.2-1b/3b-instruct, dolphin-2.9-llama3-8b, phi-2-chat, ideogram-4,
meta-llama-3.1-8b-instruct, omnivoice-cpp and qwen3-tts-cpp; the gemma-3
4b/12b/27b QAT families; and three cross-backend pairs (silero-vad plus
its sherpa-onnx build, and the vibevoice TTS and ASR builds shared
between the vibevoice-cpp and crispasr backends). The cross-backend
pairs are the first entries that meaningfully exercise engine-preference
ranking during auto-selection.

Restructures four gemma-4 families (31b-it, 26b-a4b-it, e2b-it, e4b-it).
Those bare entries were skipped by the first sweep, which left a QAT
build as parent by default. The bare entry is what installs when nothing
else fits, so it reclaims the parent slot and the former parent becomes
a plain target. Every pre-existing relationship is preserved; nothing is
dropped and nothing nests. gemma-4-12b-it has no bare entry, so it is
left as is.

qwen3-tts-cpp is a YAML anchor with nine merging children, so the five
children that did not already override variants get an explicit empty
list to stop them inheriting the parent's.

Abliterated builds stay excluded: abliteration edits the weights to
remove refusal behaviour, which makes them a different model rather than
another build of the same one.

Assisted-by: Claude:claude-opus-4-8

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 19:34:03 +02:00
mudler's LocalAI [bot]
83a0f16a21 feat(gallery): let one gallery entry offer several builds of the same model (#10943)
* feat(system): expose raw detected capability for model meta resolution

Model meta gallery entries express hardware fallback through candidate
ordering rather than a capability map, so they need the undecorated
detected capability string without Capability's default/cpu fallback
chain.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(system): drop duplicate capability accessor, cover DetectedCapability

ReportedCapability was added with a body identical to the existing
DetectedCapability. Keep one accessor and move the specs onto it, since
DetectedCapability had no direct coverage of its no-fallback behavior.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): parse IEC binary size suffixes (KiB..PiB)

ParseSizeString accepted only SI suffixes, so a "20GiB" floor was rejected
outright. Model and VRAM sizes are conventionally quoted in IEC units, and
silently reading GiB as GB would understate a floor by about 7%.

Purely additive: these inputs previously returned an unknown-suffix error.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add Candidate type for meta model entries

Candidate is one option in a meta entry's ordered variant list. It names a
concrete gallery entry and declares when that entry suits the host.

EffectiveMinVRAM resolves the VRAM floor, letting an authored min_vram win
over a nightly-inferred one. An unparseable floor errors instead of being
treated as absent: swallowing a typo would turn a constrained candidate into
an unconstrained one and select a too-large variant rather than fail loudly.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add hardware-aware model variant resolver

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): allow gallery model entries to declare variant candidates

A gallery entry with a non-empty candidates list is a meta entry: it names
an ordered list of concrete entries and resolves to the first one the host
can satisfy, instead of describing model files directly.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): resolve meta model entries to hardware-appropriate variants at install

Meta gallery entries carry an ordered candidate list; at install time the
first candidate the host satisfies is resolved and its payload installed
under the meta's name, so the model keeps a stable name regardless of which
variant backs it. The resolution is recorded in the installed gallery
config so a reinstall honors a prior pin and operators can see the backing
variant.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): key meta pin recall on the installed name and detach resolved entries

Six review findings on the meta-entry install path.

Pin recall was keyed on the gallery entry name while applyModel writes the
record under the install name (req.Name when supplied), so a meta installed
under a custom name with a pin lost that pin on reinstall and was silently
re-resolved onto a different variant, possibly swapping its backend. Compute
the install name with applyModel's own precedence before the recall.

ResolveMetaModel returned a shallow struct copy, so the resolved entry's
Overrides aliased the gallery entry's map and the install path's in-place
mergo merge wrote the caller's request into the shared catalog. Detach
Overrides, ConfigFile, AdditionalFiles, URLs and Tags. Not exploitable today
only because this path re-unmarshals the gallery per call, which is a
property nobody should have to rely on.

Also: overlay the meta's name onto the persisted config for meta installs so
the gallery file no longer records the variant's name; move the pinned-VRAM
warning below the variant validation so a pin naming a nonexistent entry does
not warn about VRAM before failing for an unrelated reason; and stop seeding
config.URLs in the config_file branch, which duplicated every declared URL.

Add seven network-free specs driving InstallModelFromGallery with a meta
entry: variant payload wins over the meta's legacy url fallback, the
resolution record round-trips to disk, a pin is recorded and honored on
reinstall including under a custom install name, and the resolved entry does
not alias the gallery's maps.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): deep-copy meta overrides and make two specs functional

ResolveMetaModel detached the resolved entry's Overrides and ConfigFile with
maps.Clone, which only copies the top level. Gallery overrides are nested in
practice (parameters.model is near-universal) and the install path merges the
caller's request with mergo.WithOverride, which recurses into nested maps and
overwrites them in place, so the gallery entry's own inner maps were still
reachable and still got rewritten by the last caller to install.

Copy both maps all the way down instead, recursing through the container shapes
a YAML decoder produces. ConfigFile is not mutated on the install path today,
but it carries the same kind of nested payload and leaving it shallowly cloned
would invite the bug back.

Also fix two specs that passed whether or not their target fix was present:

- "does not write the caller's overrides back into the gallery entry" re-read
  the catalog from disk, which re-unmarshals fresh structs and so cannot
  observe in-memory aliasing. It now asserts against the in-memory gallery
  entry and drives the real mergo merge.
- "round-trips the resolution record to disk under the meta's name" asserted a
  name that is already correct in the config_file branch. It now drives the url
  branch via a file:// fixture, where the meta-name overlay actually applies.

Both were verified red by reverting their fix.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(gallery): lint meta model entry invariants in index.yaml

Adds Ginkgo specs that parse the shipped gallery/index.yaml and enforce
the invariants that keep meta entries safe: a legacy url fallback equal
to the final candidate's url, references only to existing non-meta
entries, a min_vram floor on every candidate but the last-resort one,
a capability drawn only from the vocabulary the system can report, and
descending VRAM floors within a capability group.

The capability check is the only compensating control for a typo there.
Candidate matching is a case-sensitive exact comparison against
SystemState.DetectedCapability(), so an unknown value never matches and
falls through silently instead of erroring. The vocabulary therefore
mirrors the raw return set of getSystemCapabilities(), which notably
excludes "cpu": that is a fallback key inside Capability(capMap) on the
meta backend path, never a reported capability. A CPU-only host reports
"default".

These pass vacuously until the pilot meta entry lands; the guard is
intentionally in place before the thing it guards.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(gallery): close coverage gaps in the meta entry lint

The ordering invariant grouped candidates by capability and asserted floors
descend within a group. A candidate with an EMPTY capability matches every
host, so it does not belong in its own group: it dominates every later
candidate whose floor is at or above its own, across capability groups.
Track a running minimum floor over the unconditional candidates instead,
which subsumes the old same-group check for the empty capability.

Every spec skipped non-meta entries, so with zero meta entries in the index
all five bodies were no-ops. Aligning GalleryModel.IsMeta() with
GalleryBackend.IsMeta(), whose semantics are deliberately opposite, would
have made all of them pass while checking nothing. Extract each invariant
into a helper over a slice of entries returning the violations it finds, and
cover those helpers with synthetic fixtures so the logic stays tested at zero
meta entries. The index-driven specs are now a thin application of already
proven logic.

Also assert the index parses non-empty, report every violation in one run
rather than aborting on the first, and parse the index once for the suite.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* ci(gallery): add nightly denormalization of meta model candidates

Fills the read-only backend, quantization and inferred_min_vram fields on
meta gallery candidates and opens a PR, modeled on the existing
checksum_checker job. Computing these needs network access, so it happens
nightly rather than at install time.

An authored min_vram is never modified: a human who measured a real load
knows more than a pre-download estimate does.

The index is rewritten via yaml.Node rather than a document round-trip. A
full round-trip reflows all ~26k lines of gallery/index.yaml, which would
bury the computed values and make the nightly PR unreviewable. The rewrite
touches only the three derived keys, so authored styling survives and a run
that computes nothing leaves the file untouched.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ci): keep the gallery denormalize diff reviewable and self-healing

The nightly denormalization job edits YAML nodes instead of round-tripping
structs so its PR stays small enough for a human to review, but the write
path undid that: yaml.Marshal re-encoded the node tree at yaml.v3's default
4-space indent and dropped the leading document marker, reflowing roughly
6000 lines around the handful of real changes. Encode through
yaml.NewEncoder at the index's authored 2-space indent and restore the
header. A write that changes three fields now changes three lines.

Stale inferred_min_vram values were also never cleared. Both skip paths
(an authored min_vram is present, or the candidate is the last resort)
returned before touching the field, so a candidate that gained a floor or
became the last resort after a reorder kept an inferred value that
EffectiveMinVRAM reported as a real constraint, failing the meta lint with
no way for the job to self-heal. Clear the field before both skips.

The workflow discarded a whole night's work on any single failure: the
program exits 1 when a candidate cannot be estimated, which aborted the job
before the PR step, so one unreachable candidate blocked every other
refresh indefinitely. Capture the status, open the PR with what was
computed, mark the PR body as partial, and fail the run afterwards so the
problem still surfaces.

Also preserve the index's existing file mode instead of forcing 0644, and
drop the redundant //go:build ignore tag, since Go already skips dot
directories and the sibling modelslist.go carries no tag.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add nanbeige4.1-3b meta entry with hardware-resolved variants

Adds the first real meta entry to the gallery index. It resolves to the
Q8_0 build on hosts with at least 6GiB of VRAM and to the Q4_K_M build
everywhere else, installing either payload under the stable name
nanbeige4.1-3b.

The entry carries a url equal to its final candidate's url. LocalAI
releases that predate candidates support parse the index non-strictly
and drop the key silently, so without that url they would list the entry
and install nothing. A regression spec parses the index the way those
releases do and asserts every meta entry stays installable for them.

Also teaches core/schema/gallery-model.schema.json about candidates. The
schema sets additionalProperties: false at the top level, so an author
following CONTRIBUTING.md and adding the yaml-language-server comment
would otherwise get a validation error on this entry.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): make candidate entries complete, installable entries

Reworks hardware-resolved gallery variants after a design pivot. There is no
longer a separate "meta" entry kind. A gallery entry is a normal, complete
entry that may additionally carry candidates:, a list of hardware-gated
upgrades over itself, and the entry is itself the last-resort candidate.

The previous design relied on a bare url: as the fallback for LocalAI releases
that predate candidates support. That fallback is empty in practice: none of
the 80 gallery/*.yaml files carry a top-level files:, and 1216 of 1281 index
entries carry their payload in the index entry itself, so a url alone yields a
config template with nothing to download. Since every released LocalAI reads
gallery/index.yaml live from master, merging a payload-less entry would have
shown every existing user a model that installs to a broken state. Making the
entry its own base candidate removes the problem at the root: old clients drop
the candidates key and install the entry exactly as they do today.

Resolution order is now explicit pin, then capability plus VRAM over the
declared upgrades, then the entry itself. The entry ALWAYS installs: when its
own min_vram or capability is unmet the installer warns and installs it
anyway, because there is nothing below it and refusing would make the gallery
behave worse the newer the client is. A pin naming the entry's own name is
valid and is how an operator declines an upgrade.

IsMeta() becomes HasCandidates(), ResolveMetaModel becomes ResolveVariant, and
the persisted meta_name record key becomes entry_name. GalleryBackend.IsMeta()
is a separate concept and is untouched.

The lint drops the three rules the pivot makes wrong (url equality with the
final candidate, no inline payload, unconstrained final candidate) and gains
one: the entry's own floor must sit strictly below every candidate's, since a
base that outranks a candidate makes that candidate unreachable.

The pilot entry is now the existing nanbeige4.1-3b-q4, which gains a 2GiB
floor of its own and a single 6GiB upgrade to nanbeige4.1-3b-q8, replacing the
separate nanbeige4.1-3b entry added in d0d441bb4.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): select model variants by hardware fit, not authored order

Gallery entries could already carry a list of alternatives, but selection was
an authored, ordered, first-match policy: every candidate declared a
`capability` string and the VRAM floors had to descend in a hand-tuned order.
That pushed hardware knowledge onto whoever edits the gallery and made ordering
load-bearing, so a reordered list silently changed what users installed.

None of it was necessary. SystemState.IsBackendCompatible already derives
hardware support from a backend name alone: it knows MLX and metal are
Darwin-only, CUDA is NVIDIA-only, ROCm AMD-only, SYCL Intel-only. Selection can
read that instead of asking authors to restate it.

Authoring is now just a list of names:

    - name: qwen3.6-27b
      min_memory: 4GiB
      variants:
        - model: qwen3.6-27b-mlx-8bit
        - model: qwen3.6-27b-gguf-q8
          min_memory: 28GiB

and all the intelligence moved into the selector. Given a host it drops the
variants whose backend cannot run here, drops those whose known memory
requirement exceeds what the host has, and takes the LARGEST of what is left,
because a bigger footprint is a higher quality quantization of the same model.
A variant of unknown size is kept, since nothing proves it does not fit, but it
ranks last so a proven fit always beats a guess. An explicit pin still wins
outright, and if nothing survives the entry installs its own payload: the base
always installs, this never refuses.

Available memory is VRAM when a GPU was detected and system RAM otherwise, read
through xsysinfo so a cgroup limit is honored and a container gets its own
limit rather than the node's RAM.

Capability disappears entirely, from the types, the schema and the lint. VRAM
and RAM collapse into one `min_memory`, because a model's footprint is roughly
the same wherever it lives and one figure is compared against whichever applies.
The lint rules about ordering, the capability vocabulary and floor
relationships are deleted with the hazards they described; what remains is that
every variant names an entry that exists and does not itself declare variants,
plus that any memory figure actually parses.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* gallery: size model variants with a live probe, drop the nightly denormalizer

Selection needs each variant's size to decide whether it fits and to rank
largest-first. That figure was written into the index by a nightly job, which
made the gallery carry a derived value that could drift from the entry it was
derived from. Derive it at install time instead.

pkg/vram already sizes a model without downloading it, and the gallery UI
already uses it: a remote GGUF header range-fetch, then an HTTP HEAD for the
content length, then any declared size:. It caches its results, so reuse it
rather than writing a second probing path.

A probe failure must never fail an install, so an unprobeable variant is
treated as unknown: it survives the memory filter, because nothing proves it
does not fit, and it ranks last, so a known-good fit always beats a guess. If
every probe fails, selection still terminates on the base entry.

The probe is injected through ResolveEnv rather than called directly, for the
same reason the backend compatibility check is: specs pin an exact size, or an
exact failure, without reaching the network.

With that in place three things are dead weight and go:

- The nightly job and the fields it populated. Variant.Backend was redundant
  because the backend is resolved live from the referenced entry during
  selection, and Quantization was display-only that nothing read.
- min_memory on the base entry. The base always installs and its floor could
  only warn, so it could not change any outcome.
- The lint rules and schema entries for both.

min_memory on individual variants stays, as the override for when the probed
size is wrong. An authored figure now suppresses the probe entirely rather
than merely outranking it, so it costs no round trip.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): expose model variants for selection over API, CLI and MCP

A gallery entry may carry `variants:`, alternative builds of the same model.
Selection already worked at install time, but nothing could see what an entry
offered or ask for a specific build, so the feature was undrivable.

Listing: `GET /api/models` now reports `variants` and `auto_variant` for the
entries that declare variants. Each variant carries its resolved backend, its
measured size and whether it fits this host. `auto_variant` is what installing
without a choice would pick right now.

The new gallery.DescribeVariants runs the same variantOptions + SelectVariant
pass the installer runs, so the reported default cannot drift from what
installing actually does, and HostResolveEnv is extracted so both derive the
host and share pkg/vram's probe cache from one place.

Performance: an entry that declares no variants returns early without touching
the probe, so the ~1280 ordinary entries cost exactly what they cost before.

Selection: `variant` is accepted on POST /models/apply, as a query param on
POST /api/models/install/:id, on the gallery apply file/string request, as
`local-ai models install --variant`, and as a parameter on the install_model
MCP tool (both the httpapi and inproc clients). Empty means auto-select.

An unknown variant name now fails the install naming what was requested. This
closes a real hole: an entry declaring no variants short-circuits before
selection runs, so a requested variant was previously dropped silently and the
install reported success.

startup.InstallModels ends in a variadic model list, so install options could
not be appended to it; InstallModelsWithOptions is added alongside and
InstallModels delegates to it. No caller signature changed.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): drop the redundant variant min_memory field

Variant.MinMemory was an authored override for when the live probe misreads
a variant's footprint. It duplicated an existing field: probeEntryMemory
already passes the entry's declared size: into EstimateModelMultiContext,
whose cascade prefers that declared size over its own guesswork. Correcting
size: on the referenced entry fixes the figure for every consumer rather
than only for variant selection, so min_memory shadowed the right answer.

A variant is now nothing but a name. Its effective size is exactly the probe
result, and an unknown stays unknown: it survives the filter and ranks last.

EffectiveMemory loses its error return along with the field. The authored
string was the only thing that could fail to parse, so the error had no
remaining source and was propagating dead nil-checks through SelectVariant,
DescribeVariants and the pin warning.

Selection behaviour is unchanged. The specs covering probe-derived sizing,
ranking, filtering, the unknown-size path, pin recall, entry/variant
metadata split and deep-copy isolation all survive; the three install specs
that needed a definite size now declare it through the referenced entry's
own size:, which exercises the documented escape hatch directly.

gallery/index.yaml is untouched: no entry ever carried the key.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): rank the entry's own build against its variants

Variant selection pulled the declaring entry's own payload, the base, out
of the candidate set and consulted it only once every declared variant had
been rejected. Two real failures followed.

A variant whose size the probe cannot determine deliberately survives the
memory filter, because nothing proves it does not fit. As the only survivor
it then won outright on any host, however small: a 2GiB machine installed an
unmeasured variant in preference to the 4GiB build the entry itself ships,
with no warning. 241 of the 1280 current index entries carry no files and no
size, which is exactly that shape.

"Largest wins" also broke whenever the base was the largest. An author
writing a Q8 entry that offers a Q4 downgrade for small hosts, a natural
shape that nothing in the lint, schema or docs discourages, had the Q4
installed on every large host instead.

Make the base an ordinary participant. It is still exempt from both filters,
so selection always terminates on something installable, but it is now
ranked against the variants: a proven fit first and largest, then the base,
then any variant whose size nothing could measure. Both failures disappear
together. The base is probed for its size accordingly, which it was not
before, because an unsized base would lose every contest to an unmeasurable
variant.

FellBackToBase is kept but narrowed to "no declared variant survived",
rather than "the base was chosen", since the base now also wins on merit and
that is not worth warning about.

A recalled variant pin also became a permanent install failure. A pin the
caller supplies on this request must stay fatal, but one recalled from
._gallery_<name>.yaml can be invalidated by any later gallery edit, and
failing on it turned one rename into a model that could never be reinstalled
or upgraded again short of deleting a dotfile the user has never heard of.
A stale recalled pin is now dropped with a warning naming it, and selection
runs as if it had never been recorded.

Also drop the last textual reference to two abandoned designs from the
DetectedCapability comment, correct the documented variants JSON example,
which showed a memory_bytes of 0 that omitempty makes impossible, and remove
an em dash from the install skill.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): budget variant memory from RAM when a GPU reports no VRAM

Variant selection read its memory budget from VRAM whenever a GPU
capability was detected, and from system RAM only when none was. Apple
Silicon satisfies the first branch and fails the premise: arm64 macs report
the metal capability unconditionally, without probing anything, while
TotalAvailableVRAM has no discrete VRAM pool to find and returns zero. The
budget therefore came out as zero on every Mac.

Zero drops every variant carrying a known size, so the base build was
installed on all of them however much memory the machine had. The feature
was inert on the platform, and silently: falling back to the base is a
legitimate outcome, so nothing looked wrong.

Take VRAM only when it is actually a number, and fall back to RAM
otherwise. On a unified-memory host RAM is not an approximation of the
budget, it is the budget, since the GPU shares it. A discrete GPU whose
VRAM could not be read also lands on RAM, which overstates what the card
holds but understates nothing the host has; the previous zero understated
both.

An unreadable RAM figure still yields zero and still installs the base, so
a genuinely unknown host is not talked into a larger download.

This is what turned tests-apple red: "installs a fitting variant's payload
under the entry's own name" asserts on selection, and the runner resolved
to the base because its budget was zero. The added specs pin the branch
directly rather than relying on a macOS runner to notice again.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): add a model variant picker to the models gallery

PR #10943 shipped the server side: a gallery entry may declare `variants:`,
`GET /api/models` attaches `variants` and `auto_variant` to declaring
entries, and `POST /api/models/install/:id` accepts a `variant` query
parameter. Nothing in the UI consumed any of it, so the feature was not
reachable from the browser. This wires it up.

modelsApi.install takes an optional second argument and appends an encoded
`?variant=` only when one is given, so every existing call site keeps
sending exactly the request it sent before.

On the models table, an entry that declares variants gets a split button.
The primary Install still installs the auto-selected build, because auto is
the default and the point of the feature; the chevron opens a menu for a
deliberate override. It follows the Backends.jsx precedent: one shared
Popover re-anchored per row, rendering .action-menu items, which brings
Escape, outside-click and focus return along with it. An entry that
declares no variants renders exactly as it did before.

A variant that does not fit is dimmed but stays selectable, since the server
honors an explicit choice with a warning rather than refusing it.

memory_bytes is omitempty on the wire, so an absent key means the size is
unknown and never zero. A single helper guards both the menu and the detail
row, because formatBytes would otherwise render a falsy value as "0 B",
which reads as "needs nothing".

The expanded detail row gains a Variants section listing each build's
backend, size, whether it fits, which is the entry's own build, and which
one auto-selection would pick, built from the existing DetailRow helper and
.badge classes.

Eight Playwright specs cover the picker, including that plain Install sends
no variant parameter and that choosing one sends it. One pre-existing
assertion was scoped with .first(): the Variants section legitimately adds
more llama-cpp badges to the detail row, which tripped strict mode.

UI line coverage 49.42% -> 49.36% against a 40.0 baseline and 0.8pp
tolerance; branch coverage rose 72.04% -> 72.66%.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): describe model variants from a companion endpoint

Variant description probes each referenced entry's weight files over the
network: an HTTP HEAD plus a ranged GET, serial, five seconds per probe
with no aggregate deadline. Running it inline in GET /api/models made one
listing cost (entries x variants) round trips. The Manage page fetches
with items=9999, so at 200 declaring entries that is ~1000 serial probes,
minutes of a blocked handler and gigabytes of range traffic for a single
page load. Only one entry declares variants today, but the feature exists
so that many will.

Follow the precedent already set for VRAM estimates. The listing now
reports only has_variants, a length check on loaded metadata that touches
nothing, and GET /api/models/variants/:id returns the description for one
entry, mirroring estimate/:id in route shape, auth and error handling.
DescribeVariants itself is unchanged; only its caller moved.

The picker fetches lazily at the two points where a user asks to see
variants, opening the split-button menu and expanding the detail row, and
caches per entry for the page session. An entry declaring no variants
issues no request at all.

A spec counts real HTTP hits on the weight files, so it goes red if
description becomes reachable from the listing path again through any
caller.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): filter the model gallery to entries that declare variants

The gallery is heading towards showing parent entries and hiding the
individual builds they reference, so a user sees one row per model
rather than six quantizations of it.

Adoption is a single entry today, so defaulting to that would leave a
one-row gallery. This ships the migration-phase inverse instead: the
default is untouched, and a toggle narrows the list to only the entries
that declare variants. It previews the end state and changes nothing
until someone asks for it.

The filter is server-side, next to term/tag/backend/capability and above
the pagination arithmetic. The listing paginates at 9 items, so
narrowing on the client would leave totalPages and availableModels
describing the unfiltered set and hand the user empty pages. It selects
on HasVariants(), which reads already-loaded metadata, so it issues no
variant probes.

The parameter is named has_variants after the listing field it selects
on, and is compared against "true" like the other boolean query params
(all_users, save_checkpoint), so has_variants=false reads as absent.
With it omitted the response is byte-for-byte what it was before.

The control is the shared Toggle component, matching the fitsFilter
toggle already on this page: same wrapper class, same icon and label
shape, same localStorage persistence. Unlike fitsFilter it resets to
page 1 on change, which a server-side filter has to do.

Stacking the toggle with a tag or backend filter easily yields nothing
while one entry declares variants, so the empty state now names the
variants filter as the cause rather than leaving a user to conclude the
gallery is broken.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): render gallery model descriptions as Markdown

Gallery descriptions are Markdown, but the React UI dumped them raw, so a
model whose description opens with an ATX heading showed a literal
"# Qwen3.6-27B [](https://chat.qwen.ai)" in the list.

Full-description areas now render through renderMarkdown (marked +
DOMPurify), matching how Backends.jsx and the Manage detail panels already
handle the same content:

  - Models.jsx expanded detail row
  - VoiceLibrary.jsx voice detail header

The truncated one-line previews must not render block Markdown: a leading
"#" would become an <h1> and wreck the row height and rhythm. They get a new
stripMarkdown() helper instead, which reduces Markdown to a single line of
readable plain text. It is used for the cell text and for the title tooltip,
since a tooltip full of "[](url)" is no better than a cell full of it:

  - Models.jsx gallery table description cell
  - Manage.jsx model and backend resource-row descriptions

stripMarkdown walks marked's lexer output rather than running regexes over
the source, so what it strips is by construction what renderMarkdown would
have rendered, and it needs no new dependency. Output lands in JSX text
nodes, so React escapes it; no new dangerouslySetInnerHTML beyond the two
full-description sites, both of which run DOMPurify.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): strip Markdown from the backends table description cell

Commit b35d630cf fixed this for gallery models but left the Backends admin
page with the same asymmetry: its detail panel renders the description
through renderMarkdown, while the collapsed table row dumped the raw gallery
string into both the cell body and the title tooltip.

That is user-visible. 40 of the 949 entries in backend/index.yaml carry
Markdown - insightface uses inline code backticks, others use lists and
links - and backend descriptions also contain embedded newlines, so the
one-line cell showed literal syntax.

The cell now runs stripMarkdown over the description once and uses the
result for the text and the title, matching Models.jsx and the
ResourceRowDesc component in Manage.jsx. The '-' placeholder is preserved,
and now also fires when a description reduces to nothing after stripping.
The detail panel is untouched and no new dangerouslySetInnerHTML is
introduced: stripMarkdown output lands in a JSX text node, so React escapes
it.

Three Playwright specs cover it: a description with a heading, inline code
and a link renders as clean text with no literal syntax and no block
element in the cell, the title tooltip carries the same stripped text, and
a backend without a description still shows the placeholder.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* ui(models): polish the variant detail view and scope rendered Markdown

The gallery detail pane rendered every field through the same two-column
label/value row, including the description. Multi-paragraph prose in a value
cell ran eight rows tall at the top of the pane on a ~1200px measure, breaking
the grid's rhythm exactly where the eye enters. Move it into its own full-width
block above the table, capped at a 68ch measure, keeping the label.

Rendered Markdown had no scoped typography anywhere in the app, so a
description opening with `#` inherited the browser default 2em inside a 13px
surface while a `##` further down was indistinguishable from body text. Add a
reusable .markdown-body block mapping h1-h6, paragraphs, lists, links, code,
blockquotes, images and tables onto the existing type scale, and apply it to
every renderMarkdown() consumer: the models detail, the backends detail, both
Manage details and the voice library detail.

Rebalance the variants list so the name leads. Backend and size drop from
badge/secondary weight to muted metadata; the FITS badge goes entirely, since
it was true of nearly every row and so said nothing, while the variant that
does not fit keeps a warning badge and a dimmed name. AUTO-SELECTED stays
marked because it answers what a plain Install produces. Rows share the
parent's grid tracks via subgrid so name, backend, size and status line up
down the list instead of raggedly following name length.

Finally, make each variant row actionable. It looked like a list of choices
but was inert text, with per-variant install hidden behind the split-button
chevron elsewhere; each row is now a button onto the existing
handleInstall(modelId, variant) path, with hover, keyboard focus and a
disabled state while an install is in flight.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): collapse the listing to one row per model

The listing supported has_variants=true, which narrowed to entries that
DECLARE variants. With adoption at three entries that showed three rows,
which is useless; it was always a placeholder.

Replace it with the view that is actually useful: the deduplicated
gallery. Show every entry installable in its own right and nothing twice,
which means the parents plus every entry nobody references, and hide only
the builds another entry already offers as a variant, since those are
reachable through their parent.

The parameter is renamed to collapse_variants accordingly: the filter is
no longer a predicate on a row's own metadata but a view over the whole
gallery. Default stays off, so the response with the parameter absent is
unchanged.

VariantReferencedIDs never reports an entry that declares variants of its
own, so parents are always visible. That guarantees every hidden entry
has a visible entry offering it, and no chain can strand a row. Variant
resolution already refuses to install such a reference, but the listing
has to stay coherent in the presence of a gallery that has one rather
than silently swallowing entries. Self-references and dangling references
hide nothing.

The referenced set is computed over the whole gallery rather than over
what the other filters left, so an entry is hidden because a parent
offers it and never because of what the user searched for. The pass is
over metadata already in memory: it resolves nothing over the network and
triggers no variant description or size probe, so the listing's zero-probe
contract still holds.

The UI toggle keeps its behaviour (persistence, page reset, clear
filters) and becomes "One row per model", which says what the user gets.
Its localStorage key moves too, since the stored value meant a different
filter.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): show the collapsed model listing by default

The gallery listing is what a user reaches for to answer "what can I
install". Answering that with several rows for the same model, one per
build, makes the reader do the deduplication the collapsed view already
does, so the collapsed view is the one to land on.

The UI now asks for collapse_variants=true unless the toggle says
otherwise. The server default is deliberately untouched: a request with
the parameter absent still returns the full listing, because other API
clients depend on that response and collapsing it under them would be a
breaking change. Opting out omits the parameter rather than sending
false, so it asks for exactly the listing everyone else gets.

The stored preference changes vocabulary from '1'/'0' to 'on'/'off'. The
previous build wrote it from an effect that runs on mount, so a stored
'0' recorded that the page had been opened rather than that anyone chose
the expanded view, and honouring it would pin every earlier visitor to a
default they never picked. Only the new vocabulary counts as a choice;
a legacy '1' meant the collapsed view and is what the new default gives
anyway, so no earlier deliberate choice is lost.

Collapsing being the default also changes what the empty state may say
about it. An opted-into filter can be named as the cause of an empty
result; a default cannot, so the filters keep the top line and the
collapsed view drops to a hint below it, shown only once filters are
narrowing the set. For the same reason "Clear filters" now restores the
collapsed default instead of switching it off, and the toggle alone no
longer counts as a filter worth offering to clear.

The label stays "One row per model": it describes the view the user is
looking at rather than an action, so it reads the same whether it is
opted into or out of.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* gallery: group alternative builds of the same weights under variants

Sweep the gallery for entries that are alternative builds of the same
weights (different quantization, precision, or runtime format) and declare
them as variants of a single parent row, so the listing offers one row per
model instead of one row per quantization and the installer picks the
largest build that this host can actually run.

41 families over 95 entries, turning 54 entries into variants.

The parent is the bare-named entry wherever one exists, so nothing changes
about what any existing entry installs. Ranking already selects the largest
fitting build regardless of which entry is nominally the parent, so the
parent only decides the pathological case where nothing fits. For the ten
families that have no bare-named entry, the smallest build is the parent,
since that is the one that has to install when nothing fits.

Grouping was verified against the actual model filenames rather than the
entry names alone. Different parameter sizes, languages, finetunes, and
products that merely share a name prefix are left as separate rows: the
qwen3.6 APEX and pi-tune finetunes, the DFlash and MTP speculative-decoding
pairings, English-only versus multilingual Whisper, the QAT versus non-QAT
Gemma 4 weights, and the abliterated FLUX build are all distinct models.

Six parents define YAML anchors that other entries pull in with a merge key,
which would have handed their variants to every merging child. For the two
depth-anything anchors that would have made fourteen unrelated entries
advertise the base model's builds as their own. All 26 merging children
therefore carry an explicit empty variants list, which overrides the merged
key and is equivalent to the key being absent.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): rank model variants by host backend preference

Variant auto-selection filtered candidates by whether their backend can
run on the host, then ranked the survivors by size alone. The backend
never influenced the choice beyond that gate, so a Mac offered both an
MLX build and a llama.cpp build kept neither filtered and installed
whichever was larger, leaving the native accelerated runtime unused. The
same held for CUDA against CPU on NVIDIA and ROCm against Vulkan on AMD.

Rank by the host's backend preference between the fit tier and size: fit
stays a filter, preference decides among the builds the host can equally
hold, and size still separates builds on equally preferred runtimes.

The preference data stays in one declarative table in pkg/system, now
read by a prefix lookup instead of a switch, so adding a capability or
reordering one host's runtimes is a one-line edit and the gallery's
ranking code carries no per-backend branching. MLX joins the metal rule
ahead of metal itself, which is inert for the existing alias-resolution
consumer because no alias group holds a candidate named for mlx.

An unrecognised backend, an unrecognised capability and an absent
preference list all collapse to the previous size-only ordering rather
than erroring or dropping candidates.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): rank variants by engine name, not backend build tag

Variant auto-selection ranked candidates with
SystemState.BackendPreferenceTokens, but that function and the variant
ranker speak different vocabularies.

BackendPreferenceTokens returns BUILD TAGS ("cuda", "rocm", "sycl",
"vulkan", "metal", "cpu"). It exists to match installed backend build
directory names like "llama-cpp-cuda-12" during alias resolution in
ListSystemBackends. Variant ranking instead matches a gallery entry's
`backend:` value, which is an ENGINE NAME: "llama-cpp", "vllm",
"vllm-omni", "sglang", "mlx" and the rest. No engine name in
gallery/index.yaml contains "cuda", "rocm", "sycl" or "vulkan".

preferenceRank matches by substring, so on an NVIDIA host the tokens
[cuda, vulkan, cpu] matched neither "llama-cpp" nor "vllm", every
candidate scored identically and size alone decided. The NVIDIA, AMD,
Intel, darwin-x86 and vulkan rules were all inert. Only metal appeared
to work, and only because the token "mlx" happens to equal an engine
name. The mismatch does not error, it silently deletes the feature.

Separate the two vocabularies. backendBuildTagPreferenceRules keeps the
build tags and its original output for every capability, including
metal, whose "mlx" token is removed again; its alias-resolution consumer
is byte-identical to before. engineNamePreferenceRules is new, holds
engine names, and is read by the new EnginePreferenceTokens, which
HostResolveEnv wires into the renamed ResolveEnv.EnginePreference. Both
tables sit adjacent under one block comment naming each vocabulary and
each consumer, and share one lookup helper so their semantics cannot
drift.

On NVIDIA the order is vLLM, then SGLang, then llama-cpp: vLLM is the
throughput engine and a model published with a vLLM build is published
that way because that build is the one worth running. AMD and Intel get
the same order, since rocm and intel builds of both serving engines
ship. Metal prefers mlx over llama-cpp. Vulkan prefers llama-cpp, the
only LLM engine with a Vulkan build. darwin-x86 and unknown
capabilities are deliberately absent rather than guessed at, degrading
to the size-only ordering that predates preference.

preferenceRank stays generic and names no engine and no capability, so
adding a runtime remains a one-line table edit.

Specs pin the NVIDIA and metal rules through the live table and the real
HostResolveEnv wiring, so emptying the engine table or wiring the build
tag source back in both go red. A regression table asserts
BackendPreferenceTokens' original output per capability, and mirrored
locks assert neither table carries the other's vocabulary.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: record that variant selection ranks by engine before size

A gallery entry can now declare variants, and selection ranks the builds a
host can run by engine preference before size. Nothing told a contributor
adding a backend that engineNamePreferenceRules exists, so a new engine would
silently rank below every known one and lose to whatever build happened to be
larger on hosts where it should have won.

Document the step where a backend is added, warn against the sibling
backendBuildTagPreferenceRules table (build tags, not engine names: the wrong
table matches nothing, scores every candidate equally and disables the
preference without erroring), and index it from AGENTS.md.

Fix the authoring and user docs, which still claimed the largest surviving
build wins. An author grouping builds under one entry has to be able to
predict what a user gets, and size alone no longer decides it.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(cli,mcp): describe variant auto-selection as preference before size

The CLI flag help and the install_model tool schema both still said
auto-selection takes the largest build that runs. Ranking now puts engine
preference ahead of size, so on NVIDIA a vLLM build wins over a larger
llama.cpp one. An assistant reading the old schema would tell users the
wrong thing.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): prefer llama.cpp over GPU serving engines on hosts with no GPU

engineNamePreferenceRules had no row for the "default" capability, which
getSystemCapabilities() returns both when no GPU is detected and when a GPU
is present but under the 4 GiB VRAM floor. A missing row yields an empty
preference list, which preferenceRank reads as "score everything equally",
collapsing variant selection to size alone.

That would be harmless if the hardware filter dropped GPU serving engines on
such a host, but it does not. IsBackendCompatible derives support from the
engine NAME, and "vllm" and "sglang" contain none of the darwin, cuda, rocm
or sycl tokens it keys on, so they fall through to its closing "return true".
A vLLM variant therefore survives on a CPU-only box and wins whenever its
build is the larger of the two on offer: the machine installs vLLM in
preference to llama.cpp.

darwin-x86 had the identical hole. It was documented as a deliberate omission
because nothing accelerates on an Intel Mac, which is true about acceleration
and wrong about consequence: with every engine tied, download size decides.

Add rows for both putting llama-cpp first. The GPU engines are enumerated
behind it rather than left unmatched: an unmatched engine already ranks below
every listed one, so llama.cpp would win either way, but unmatched engines
also tie with each other and let size decide among them. Naming them fixes
that order. MLX is left off the darwin-x86 row on purpose so it ranks last,
since IsBackendCompatible admits darwin-tokened engines on that capability
even though MLX needs Apple silicon.

Preference orders survivors and never filters, so a model published only as a
vLLM build is still installed on a host with no GPU; there is a spec for it.

Surveyed every other value getSystemCapabilities() can return. nvidia, amd,
intel and vulkan have rows; the l4t and cuda-refined values reach the nvidia
row by prefix; "apple" and "" cannot reach the vendor fallthrough because the
darwin and no-GPU branches return earlier. default and darwin-x86 were the
only live holes.

BackendPreferenceTokens and its build-tag table are untouched, and
preferenceRank stays generic, naming no engine and no capability.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* gallery: prefer speculative-decoding builds when they fit

Rank serving features between engine preference and size, so a host that can
hold a DFlash or MTP build of a model's weights installs it instead of the
plain build. Both answer faster for the same output, so whenever one survives
the filters there is no reason to take the plain build.

Precedence is now fit, then engine, then serving feature, then size. Engine
outranks the feature deliberately: a serving feature makes the right engine
faster, it does not make a wrong engine right, so a plain vLLM build still
beats a DFlash llama.cpp build on NVIDIA. Fit outranks both, and a drafter
pairing is strictly larger than the plain build, so the existing size filter
drops it on a host too small for it before this axis is consulted.

The order lives in a third preference table in pkg/system, alongside the build
tag and engine name tables. It is the odd one of the three: not keyed by
capability, because no hardware prefers a plain build over an equivalent
faster one, and matched against whole segments of a gallery ENTRY NAME rather
than as a substring of a backend value. Nothing on a gallery entry declares a
serving feature, and tags are not a usable substitute: gemma-4-e2b-it:sglang-mtp
carries an mtp tag while ornith-1.0-9b-mtp and qwen3.6-27b-nvfp4-mtp carry
none. Entry names are author-supplied free text, unlike the closed engine
vocabulary, so a short marker can turn up inside an unrelated word and whole
segment matching is what keeps smtp-assistant from ranking as an MTP build.
The block comment over the tables now documents all three together and states
what each is matched against; the ranking code names no feature, so adding one
stays a one-line edit to the table.

29c49203b rejected these entries as serving configurations rather than
alternative builds of the same weights. The definition is now "alternative ways
to serve the same model", which includes them, so regroup 14 entries under 12
parents. Judged by the files each entry points at: the qwen3.6, qwen3.5, qwen3
and deepseek pairings are the base GGUF plus a drafter, the gemma-4 QAT MTP
entries are the same QAT weights at a different quantization plus an MTP
drafter, and the two sglang MTP entries describe themselves as the same model
served with speculative decoding. Left separate: qwen3.6-27b-mtp-pi-tune, a
finetune with its own weights, and every entry whose base model LocalAI does
not ship as its own row, which is the whole Qwopus line plus gemmable-4-12b-mtp,
mimo-7b-mtp:sglang and qwen3.5-4b-dflash.

None of the twelve parents defines a YAML anchor, so no variants key can leak
through a merge key and no empty override was needed this time. The index was
edited by line insertion only.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test: check env restore errors in capability and variant specs

errcheck flagged ten unchecked os.Setenv and os.Unsetenv returns in the
specs added while the pre-commit hook was being skipped. Restoring an env
var is exactly the place a silent failure leaks state into the next spec,
so assert on it rather than suppressing the linter.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): make the mtp tag authoritative for serving-feature ranking

Variant auto-selection ranks survivors by fit, then engine, then serving
feature, then size. The serving-feature lookup read only whole alphanumeric
segments of a variant's entry name, because tags were inconsistent: every
dflash entry carried a dflash tag, but only 7 of 20 MTP entries carried an
mtp tag.

Tag the 13 untagged MTP entries, then teach the lookup to read tags as well
as names. A tag is now the authoritative signal and is compared whole and
case-insensitively, which is safe precisely because a tag is a deliberate
declaration rather than free text: there is no word-inside-a-word failure
mode, so the segment splitting the name half needs is unnecessary there.

The name check stays as a fallback rather than being replaced. Switching to
tags only would have regressed the six already-grouped entries on the day it
shipped, and would depend on tagging discipline that does not exist yet.

The lookup still names no feature, so adding one remains a one-line edit to
servingFeaturePreferenceTokens.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): make a declared tag the sole serving-feature signal

Variant auto-selection ranks survivors by fit, then engine, then serving
feature, then size. The serving-feature lookup recognised a speculative build
by either a declared tag or a whole segment of its entry name. Drop the name
half: a tag is now the only signal.

A name is author-supplied free text and a naming convention is not a contract,
so reading a marker out of one infers a capability nobody declared. The gallery
already had the failure in it: the four NVFP4 entries name MTP-bearing weights
while setting no option that enables speculative decoding, and being live
variants they were winning the feature axis without answering any faster.

overrides.options was considered as the replacement and rejected. It carries
spec_type:draft-mtp / spec_type:draft-dflash, which is what actually turns the
feature on, but that spelling is llama.cpp's config vocabulary: ds4 spells the
same feature mtp_path and sglang spells it speculative_algorithm in a
referenced config. Keying a cross-backend ranking decision on one backend's
option syntax would rank the other backends' builds as plain. Options are the
curation-time check instead, and never reach the selection logic.

With no fallback left, tag correctness is load bearing, so audit every entry
against the rule "tagged when the entry configures that feature, in whatever
vocabulary its backend uses". Three entries configure MTP untagged and gain the
tag (hy3, glm-5.2, qwythos-9b-claude-mythos-5-1m, all spec_type:draft-mtp with
no marker in their names). Four carry the tag while configuring nothing and
lose it: qwen3.6-27b-nvfp4-mtp, qwen3.6-35b-a3b-nvfp4-mtp,
qwopus3.6-27b-coder-mtp-nvfp4 and qwopus3.6-27b-v2-mtp-nvfp4, whose only option
is use_jinja:true. The dflash side was checked independently rather than assumed
consistent: all five dflash entries declare spec_type:draft-dflash and all five
are tagged, so it needed no edits.

Four entries keep a tag that a literal spec_type-only reading would strip,
because they configure MTP through a different backend: deepseek-v4-flash-q2-mtp
via ds4's mtp_path/mtp_draft, and the three sglang entries via
speculative_algorithm in their referenced configs. Stripping those would
contradict the reason spec_type was rejected as the signal and would demote four
genuinely faster builds to plain.

The index was edited by line insertion and deletion only, never round-tripped
through a serializer. A resolved-tag diff across all 1272 named entries, taken
after merge keys are applied, shows exactly these 7 changing and no entry
gaining or losing a tag through an anchor.

The two specs that pinned the name fallback are inverted rather than deleted,
since a name silently promoting a build is the regression worth guarding. The
whole-token guard survives on the tag path, where smtp must still not match mtp.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): make deepseek-v4-flash variant targets installable

Clicking install on deepseek-v4-flash failed with "invalid gallery model".
The parent entry is fine, but all four entries it was grouped with declared
neither url: nor config_file:, and applyModel needs one of the two to have
anything to build a config from. They carry urls: (plural), the informational
HuggingFace link list, which is a different field. None of the four was ever
independently installable, so grouping them routed a previously-working
install into a broken entry.

Give each the url: the parent already resolves through. virtual.yaml is a
no-op base, and applyModel passes overrides to InstallModel separately from
the fetched config, so backend: ds4, the parameters and the ssd/mtp options
all still land exactly as authored. This is the same pattern the parent and
many other GGUF entries in the index already use.

Add the lint rule that should have caught this. checkVariantReferences only
proved a target exists and is not itself a parent, which is structural
validity: an entry can exist, declare no variants, and still be
uninstallable. checkVariantTargetsInstallable mirrors applyModel's
precondition instead, and names the parent, the target and the missing
fields, because whoever hits it is reading a gallery entry and has no reason
to know applyModel exists.

The two index-driven resolution specs live in their own Ordered container:
an Ordered container stops at its first failure, so sharing one with the lint
rules let a lint breach skip them silently.

Nine further entries gallery-wide have the same defect and are unrelated to
variants, so they are broken installs that predate this branch. They are left
alone here rather than buried in a regression fix, and widening the rule to
cover every entry is deferred with them so the gate can ratchet up in one
step instead of needing a skip list.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): install entries with no url or config_file on an empty base

applyModel had three branches: fetch a base config from url:, build one from
an inline config_file:, or fail with "invalid gallery model". An entry
declaring neither is now installed on an empty base config, with overrides:
and files: supplying everything.

This is what the ~345 entries pointing at gallery/virtual.yaml were already
getting. That stub is five lines carrying name, description and license.
description and license are overwritten from the gallery entry immediately
after the fetch, and the name never reaches disk because InstallModel prefers
the install name. Crucially applyModel passes model.Overrides to InstallModel
as a separate argument rather than merging it into the fetched config, so
nothing an author writes depends on that base existing. The fetch bought a
round trip to GitHub and nothing else.

That makes f4ef80173 the wrong fix, so it is unwound. The four url: lines it
added to the deepseek-v4-flash variants are reverted: they are a pointless
network fetch now, and the family installs without them.

Relaxing the branch would hide a real authoring mistake, so a payload rule
replaces the base-config rule. An entry with no url, no config_file, no
overrides and no files installs nothing and would leave an empty model
directory while reporting success, so it is refused by name. The caller's
request counts toward the payload, because its overrides and files are merged
into the install exactly as the entry's own are. urls: (plural) is the
informational link list and does not count, which is what the four entries
that shipped broken had and why they were still uninstallable.

checkVariantTargetsInstallable asserted every variant target declares a url:
or a config_file:, which is no longer true and would now reject correct
authoring. checkEntriesInstallSomething pins what survives instead, and covers
every entry rather than only variant targets: the hazard is a half-written
stanza and a parent can be one as easily as a target. The old rule was scoped
to targets precisely because nine unrelated entries would have failed a
gallery-wide version; those nine are valid now, so the deferred ratchet
happens here in one step. 1280 entries, zero violations.

Those nine (aurore-reveil_koto-small-7b-it, lfm2-1.2b, the six liquidai_lfm2
entries and deepseek-v4-pro-q2-ssd) become installable for free. Each carries
overrides: and files:, and one of them is driven through the real install path
in a spec.

The no-fetch spec is paired rather than bare: an assertion that nothing was
fetched proves nothing unless something could have been, so a control runs the
same fixture with a url: pointing at a base config that is not there and
asserts the install fails. Only then does the identical fixture without the
url passing mean the read was skipped.

Follow-up, deliberately not here: the ~345 entries still naming virtual.yaml
can drop their url:. That is 345 index edits with their own risk, and mixing
them in would bury this change.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* ui(models): let search bypass the variant collapse, drop the toggle

The models page collapsed the gallery to one row per model by default and
offered a toggle to see every individual build. Because the collapse composed
with the search term, a build another entry offers as a variant could not be
found by typing its name, so the toggle was the only way to reach those builds
in the UI. A user who typed a name they knew existed got "no models found",
which reads as "that model does not exist".

Collapse is for browsing; search is for finding. An explicit search term now
bypasses the collapse in the listing handler, so a name lookup returns matching
entries whether or not a parent offers them. The term is trimmed once at the
top of the handler, so whitespace is neither a search nor a bypass; previously
an untrimmed blank term also narrowed the listing to whatever contained a
space. Tag and backend deliberately do not bypass: they refine a listing the
user is still reading rather than name an entry already known to exist.

That makes the toggle redundant, so it goes, along with its i18n strings in all
six locales, its localStorage persistence, its participation in "Clear filters"
and the empty-state hint telling users to turn it off. The hint was doubly
stale: it pointed at a control that no longer exists, and it was untrue exactly
when a user has a search term, since searching now sees every build. The page
always requests the collapsed listing.

The stored preference key is left inert rather than cleaned up: nothing reads
it, so a user who had the toggle off simply gets the collapsed view.

collapse_variants stays on the API, off by default, because other clients want
either view and the UI dropping its control is no reason to remove a working
parameter.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): give the models gallery filter form a deliberate structure

The filter area had accreted controls into one undifferentiated flow. The
"Fits in GPU" toggle and the backend select were direct children of
.filter-bar, the same wrapping container as the 18 taxonomy chips, so their
position was decided by how many chips happened to wrap at the current width
rather than by any layout intent. At narrow widths they were pushed past the
right edge of that container's horizontal scroll and became unreachable
entirely.

Restructure into three bands inside the house .filter-bar-group wrapper that
components/FilterBar.jsx already uses on Backends and the System tabs:

  1. query scope: search plus the backend select
  2. taxonomy: the chip row, alone, free to wrap
  3. refinements: fits-in-GPU and context size, under a hairline rule

The backend select leads the chips rather than trailing them because picking a
backend disables the use cases that backend cannot serve, so it gates the row
below it. Fits-in-GPU and context size share a band because they are one
control group: the context size is the length the VRAM estimate is computed at,
and that estimate is what the fits filter tests against.

Chips had no visible keyboard focus indicator. The global focus ring is wrapped
in :where(), so it carries the specificity of a bare :focus-visible, ties with
.filter-btn and loses on source order, leaving focused chips showing their
resting drop shadow. Restate the ring where it outranks both resting and hover.

Also: aria-pressed on the chips, a real label association and aria-valuetext on
the context slider (it steps over an index, so it announced "2"), disabled chip
styling moved off inline styles, a prefers-reduced-motion block for the chip
transition, and the hard-coded English "Context:" moved into all seven locales.

No behaviour change: same filters, same state, same requests. Page reset on
change, localStorage persistence and "Clear filters" verified unchanged.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): let the models recommendations panel fade into the background

The "Recommended for your hardware" strip rendered at full height on every
visit regardless of how many models were already installed, costing 186px at
1600px wide (287px at 1100px, where its cards wrapped to two rows) and pushing
the first gallery row to y=554 / y=703.

Make its prominence track how much the user still needs it. The panel now
defaults to a one-line summary once anything is installed, and both the
collapse choice and the existing dismissal persist:

  collapsed = explicit user choice, if one exists
            : installedCount > 0

The preference is three-valued on purpose. A boolean cannot tell "the user
expanded it" apart from "the user has never chosen", and those need opposite
handling when the installed count later crosses zero: someone who deliberately
opened the panel on an empty instance should not have it collapse out from
under them when their first model finishes installing.

Collapsed keeps the card, icon, title and a suggestion count, so the panel is
recovered by clicking what you are already looking at rather than by hunting.
Expanded is unchanged, because for a user with nothing installed it was never
the problem. Collapsed reclaims 145px at 1600 and 420, and 246px at 1100.

Models.jsx gains a statsLoaded flag: stats initializes to installed:0, so
reading it before the fetch resolves would render expanded and collapse a frame
later, which is exactly the layout shove this removes.

The dismissal key moves to the page's localai-models-* convention; the old
localai_rec_models_dismissed is still read, never written, so an existing
dismissal is honoured rather than resurrected by the rename.

Accessibility: the disclosure is a real button whose accessible name is the
visible title alone, with state on aria-expanded and aria-controls resolving in
both states, because the grid is hidden via the hidden attribute rather than
unmounted. That also keeps the four install buttons out of the tab order while
collapsed. The app's global focus ring applies; no per-component outline is
added, per the warning in App.css. Reveal animates opacity and transform only,
never height, and both it and the chevron rotation are disabled under
prefers-reduced-motion.

Only en had a recommended block, so the other six locales were falling back to
English for the whole panel. Translated the complete block rather than adding
one orphaned key to files that would still render the title in English.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(downloader): recover from a leftover .partial on non-HTTP URIs

An interrupted download leaves a `<file>.partial` behind. The partial
handling in DownloadFileWithContext gated resume on `err == nil &&
uri.LooksLikeHTTPURL()`, so for any URI that is not literally http(s)
the branch fell through to `else if !errors.Is(err, os.ErrNotExist)`,
which with a nil err is true. The download then failed with an error
wrapping nil:

  failed to check file ".../Ternary-Bonsai-27B-Q2_g64.gguf" existence: <nil>

Every gallery file URI uses `huggingface://`, so a single interrupted
download made that model permanently uninstallable until someone
deleted the partial by hand. The `<nil>` in the message compounded it
by pointing debugging at a filesystem failure that never happened.

Restructure the handling as an explicit switch over the four real
states: partial exists and is resumable, partial exists and is not
resumable (discard and restart, as already done for an HTTP server
without range support), no partial, and a genuine stat failure. The
error branch is now only reachable with a non-nil error, names the
path that was actually stat'd, and wraps with %w.

Discarding is required for correctness and not merely convenience: the
writer opens the partial with O_APPEND, so an un-resumed download would
concatenate a fresh body onto stale bytes.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): tell the models gallery's variant rows apart, and let browsing see every build

Both variant surfaces rendered name, backend and size. For two builds of one
model that is close to no information: a variant exists precisely because the
same weights are offered another way, so the backend usually matches and the
sizes usually land within a few hundred megabytes. Comparing
ternary-bonsai-27b-pq2 against ternary-bonsai-27b-q2-g64 meant reading two names
that differ by a suffix nobody has defined anywhere in the UI.

Report the quantization and the serving features on VariantView, and derive both
server-side from the referenced entry rather than parsing names in the browser,
so every client reads the same format out of the same file the installer will
hand the backend.

Quantization comes from overrides.parameters.model first, falling back to the
file list. That order is load bearing: entries routinely ship a vision tower
alongside the language model at a different quantization, so reading the file
list first reports the mmproj's format. Matching walks `-` and `.` delimited
segments right to left; `_` deliberately does not split, because it separates the
parts INSIDE a quant token and splitting on it reports Q4 for a Q4_K_M build. A
second, looser pass takes a segment's `_`-delimited tail, which catches the
gemma-4-E2B_q4_0-it.gguf style; it runs second so a precise match can never lose
to a fuzzy one further right in the name. An entry naming no format reports
nothing, which is the honest answer for a backend served from a directory of
weights.

Features are the same tag-against-vocabulary match servingFeatureRank already
ranks on, over the same host preference list. A build can therefore never be
shown as faster than one selection did not actually reward, nor rewarded without
being shown; a spec pins that agreement rather than trusting it.

The compact dropdown gets the quantization on its meta line and the bare feature
token. The detail row, which has the room, gets the quantization as its own
monospaced column so precision lines up down the list, and the feature spelled
out, because DFLASH names nothing to a user who has not met it. The referenced
entry's description stays out of both: the detail row already renders the
parent's prose above the table, and a second block per variant would push a
three-variant list past a screen to restate what the columns now say precisely.

The collapse toggle comes back. 462583f38 dropped it once search bypassed the
collapse, on the reasoning that nothing was unreachable any more. That holds for
finding a build whose name you know and does not hold for browsing: no sequence
of actions enumerated the 68 builds the default view hides. Collapse is for
browsing and search is for finding, and the toggle was the browsing half.

It goes in the refinements band 0d4823362 established, not back among the
taxonomy chips where its position depended on how many chips happened to wrap. It
leads that band because it decides how many rows the other two refine over, and
because unlike fits-in-GPU it is unconditional: a host with no GPU still browses.

The search bypass is untouched and re-checked by a spec in the toggle's default
state, since restoring the control must not restore the dead end it replaced. The
empty-state hint returns but only without a search term, because a term bypasses
the collapse and the hint would otherwise point at a control that cannot change
the result. The stored preference reads 'on'/'off' only: an older build wrote
'1'/'0' from an effect that ran on mount, so those record that the page was
opened, not that anyone chose a view.

Also fixes a latent flake it exposed. The collapse_variants spec compared whole
response bodies byte for byte, and the listing envelope carries live host
telemetry that drifts between two calls milliseconds apart, so it was asserting
on the machine's memory pressure. It now compares everything the parameter
governs -- the entries, their serialization and the paging -- and is green 25/25
where it was failing about one run in three.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): let the models gallery show a variant's full details

The variant list in an entry's expanded detail row says how the builds
differ: name, backend, quantization, size, and the auto-selected, base
and serving-feature markers. It cannot say what any one of them is. A
variant's own description, tags, license, source links and file list are
unreachable anywhere in the UI, because while the collapse is on a
variant has no gallery row of its own.

Give each variant row an info control that reveals its entry, rendered by
the same ModelDetail a top-level row gets, so a field added to the detail
view appears here too. variantData is withheld from the nested render: a
variant may declare variants of its own, and recursing would nest a
picker inside a picker two levels deep already.

An inline disclosure rather than a modal. The control sits inside a table
row that is already expanded, inside a variant list within that; a dialog
opened from there stacks a dismissal on a dismissal for a handful of
extra fields about the entry the user is already reading, and breaks the
page's own expand idiom. The third level is carried by an inset and a
left rule instead of another card.

The entry is fetched by exact name from the listing, once, on first use.
The listing already returns every field the detail view renders, and a
search term bypasses the variant collapse server-side, so no new endpoint
is needed and neither the listing nor DescribeVariants gains any work.
Expanding a row costs nothing; a variant nobody opens costs nothing. A
name the listing no longer returns is stated, not blanked: an empty panel
reads as a rendering fault rather than as a lookup that came back empty.

The control is a sibling of the install button, not a descendant, so
asking about a build can never install it.

The variant list keeps its content-sized columns via a trailing filler
track instead of max-content sizing, so the rows are unchanged while the
panel spanning them gets the pane width its file table needs.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): let search respect the collapse instead of switching it off

The models listing collapsed to one row per model, and an explicit search term
turned that off wholesale. Searching while collapsed therefore answered with the
individual builds a parent already offers, which are exactly the rows the view
the user asked for has no place for: typing "mtp" returned
qwen3.6-27b-nvfp4-mtp, a row that is invisible the moment the box is cleared.
The bypass was the right shape of fix for the wrong half of the problem. What a
search must not do is answer "no models found" for a build the gallery does
hold; that does not require abandoning the grouping the user asked for.

So the term is now matched against every entry either way, hidden builds
included, and the collapse decides how a match is reported rather than which
matches exist. Collapsing stops being a filter that drops rows and becomes a
substitution: a match on a build another entry offers is reported as that entry,
the one installable in its own right. Nothing becomes unfindable and nothing
comes back that the requested view cannot show.

Substitution happens after search, tag and backend, so every filter is judged
against the build that really carries the name, tag or backend rather than
against a parent that merely offers it; the other order would let backend=vllm
match a parent whose own backend is something else. The price is that the
surfaced row shows the parent's own metadata while the match was on a variant,
which is what grouping means, and the alternative is claiming the gallery holds
no such build. It happens before the count and the page math, so both describe
the rows actually handed out rather than the matches that produced them.

A parent already in the result keeps its own position and absorbs its matching
variants there, which is what leaves the browsing listing ordered exactly as it
was; a parent surfaced only by a variant takes the position of the first variant
that surfaced it. Either way it appears once, however many of its builds matched
and whether or not it matched itself. Search preserves gallery order rather than
scoring, so a surfaced parent has a real position rather than an invented one.

VariantParents never reports an entry that declares variants of its own, so a
parent is never itself hidden and one hop always lands on a visible row. The
handler follows exactly one anyway: refusing the second is what makes a gallery
the linter would have rejected terminate rather than loop.

The empty-state hint pointing at the toggle goes with it for every server-side
filter. Substitution means a match is always reported as some row, so the
collapse can no longer be why a term, a chip or a backend came back empty, and
naming it there sends the user to a control that cannot change the result. It
survives for the fits filter alone, which runs in the browser after the
substitution and judges the surfaced entry's own size: there the build that fits
really can be filtered out along with a parent that does not.

Searching a build's exact name while collapsed now answers with its parent, so
the result no longer contains the string the user typed. That is intended, and
the row is the one they can act on, but it is a real rough edge: nothing on the
row explains the connection. Closing it properly means reporting which variant
matched so the UI can say so, which the listing does not do today.

ResetGalleryModelCache is added for tests. The model cache is a package global
keyed by nothing, so a background refresh one spec triggers can land in the
middle of the next and answer it with the previous spec's gallery; the extra
specs here made that fail about one run in five. It waits for the in-flight
refresh to publish before clearing, since clearing alone only narrows the
window.

Assisted-by: Claude:claude-opus-4-8
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 18:43:02 +02:00
mudler's LocalAI [bot]
6e52d0c2ef fix(ci): rebuild backends when shared build inputs change (#10975)
The backend matrix path filter only matched files under a backend's own
directory, so a change to shared build infrastructure rebuilt nothing at
all: an empty matrix, every job green, and the change reaching no image.

PR #10946 fixed scripts/build/package-gpu-libs.sh shipping a partial
4-of-8 cuDNN library set, which mixed versions with the venv's pip cuDNN
and produced CUDNN_STATUS_SUBLIBRARY_VERSION_MISMATCH at inference time.
It merged 1h48m after the weekly full-matrix cron had already run, so no
backend image ever received the fix and nothing signalled that it had
been un-shipped.

Add a SHARED_BUILD_INPUTS table mapping each shared path to the narrowest
set of matrix entries it can honestly invalidate, plus a generic rule for
backend/Dockerfile.<x> (which each entry already names). A full matrix is
417 Linux + 56 Darwin builds, so package-gpu-libs.sh now rebuilds the 176
Python entries rather than everything. Unclassified files under
scripts/build/ fall back to a full rebuild deliberately: over-building is
recoverable, silently shipping nothing is not.

Extract the filtering logic to scripts/lib/backend-filter.mjs so it can be
unit-tested without bun, js-yaml or a GitHub API round-trip, and run those
tests from the existing lint workflow via `make test-ci-scripts`.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 13:48:12 +02:00
mudler's LocalAI [bot]
465d488c90 fix(distributed): reject wrong-model requests at the backend (#10970)
fix(distributed): reject wrong-model requests at the backend (#10952)

In distributed mode the controller caches a NodeModel row naming a backend's
host:port. A worker can recycle a stopped backend's gRPC port for a different
model's backend, and probeHealth verifies liveness rather than identity, so the
probe succeeds against whatever now occupies the port and the request is
dispatched to the wrong backend. The caller gets a silent wrong-model answer.

Nothing in the request could catch this: PredictOptions had no model field, so
model identity crossed the wire only in ModelOptions.Model at LoadModel time,
and the cached-hit path issues no LoadModel. Every backend's "model not loaded"
guard checks a nil handle, which a process holding a different model passes, so
the stale row was never dropped either.

Add PredictOptions.ModelIdentity and enforce it at the point of use:

  - The controller populates it in gRPCPredictOpts from ModelConfig.Model, the
    same expression ModelOptions feeds to model.WithModel and therefore the
    same value the backend received as ModelOptions.Model. Both are read from
    one config value in one function, so they are equal by construction and the
    comparison cannot false-reject.
  - Backends compare it against what they loaded and return NOT_FOUND with a
    fixed sentinel. Enforced in pkg/grpc/server.go (27 Go backends), an
    interceptor in backend/python/common (all 36 Python backends, no
    per-backend change), and the llama-cpp / ik-llama-cpp / ds4 C++ servers.
    That is every backend with real exposure: kokoros answers all four RPCs
    with unimplemented and privacy-filter implements none of them.
  - The router's reconcile drops the stale replica row on a mismatch, so the
    next request reloads somewhere correct.

Empty means "skip the check" on both sides: a controller that predates the
field sends nothing, a backend loaded by such a controller has nothing to
compare, and the C++ server synthesizes PredictOptions internally for ASR. That
keeps upgrades working in both directions.

Scoped to the four PredictOptions RPCs. TTSRequest.model and
SoundGenerationRequest.model are deliberately NOT validated: FileStagingClient
already rewrites them to worker-local absolute paths, so in distributed mode
they already differ from the load-time value and comparing them would reject
valid requests.

IsModelMismatch requires both the NOT_FOUND code and the sentinel, unlike the
neighbouring helpers which accept either. insightface's Embedding returns
NOT_FOUND "no face detected" on a PredictOptions RPC, and a code-only check
would drop a healthy replica row on every faceless image.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 13:05:47 +02:00
mudler's LocalAI [bot]
1618c2e445 chore(model gallery): 🤖 add 1 new models via gallery agent (#10971)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-20 08:34:55 +02:00
dependabot[bot]
9043cbc786 chore(deps): bump torch CPU wheels to 2.12.1 (#10969)
* chore(deps): bump the pip group across 6 directories with 1 update

Bumps the pip group with 1 update in the /backend/python/ace-step directory: torch.
Bumps the pip group with 1 update in the /backend/python/llama-cpp-quantization directory: torch.
Bumps the pip group with 1 update in the /backend/python/longcat-video directory: torch.
Bumps the pip group with 1 update in the /backend/python/sglang directory: torch.
Bumps the pip group with 1 update in the /backend/python/trl directory: torch.
Bumps the pip group with 1 update in the /backend/python/vllm-omni directory: torch.


Updates `torch` from 2.10.0+rocm7.0 to 2.12.1+cpu

Updates `torch` from 2.10.0 to 2.12.1+cpu

Updates `torch` from 2.12.1 to 2.12.1+cu130

Updates `torch` from 2.9.0 to 2.12.1+cpu

Updates `torch` from 2.10.0 to 2.12.1+cpu

Updates `torch` from 2.7.0 to 2.12.1+cu130

---
updated-dependencies:
- dependency-name: torch
  dependency-version: 2.12.1+cpu
  dependency-type: direct:production
  dependency-group: pip
- dependency-name: torch
  dependency-version: 2.12.1+cpu
  dependency-type: direct:production
  dependency-group: pip
- dependency-name: torch
  dependency-version: 2.12.1+cu130
  dependency-type: direct:production
  dependency-group: pip
- dependency-name: torch
  dependency-version: 2.12.1+cpu
  dependency-type: direct:production
  dependency-group: pip
- dependency-name: torch
  dependency-version: 2.12.1+cpu
  dependency-type: direct:production
  dependency-group: pip
- dependency-name: torch
  dependency-version: 2.12.1+cu130
  dependency-type: direct:production
  dependency-group: pip
...

Signed-off-by: dependabot[bot] <support@github.com>

* fix(deps): preserve platform-specific torch requirements

Keep the 2.12.1 CPU bump only where uv resolves it cleanly, and restore ROCm, CUDA, MPS, and unrelated transformers constraints that Dependabot rewrote to incompatible wheel variants.

Assisted-by: Codex:gpt-5 [uv]

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-20 08:34:33 +02:00
localai-org-maint-bot
0406741a8c fix(vibevoice): install diffusers from PyPI instead of git main (#10972)
Every vibevoice requirements file pulled diffusers straight from
git+https://github.com/huggingface/diffusers. That branch now reports
itself as 0.40.0.dev0 and requires huggingface-hub>=1.23.0,<2.0, while
transformers>=4.51.3,<5.0.0 (which upstream VibeVoice mandates) still
caps huggingface-hub at <1.0. Because a git URL offers the resolver
exactly one candidate version, uv has nothing to backtrack to and the
install fails outright:

  Because only diffusers==0.40.0.dev0 is available and diffusers==0.40.0.dev0
  depends on huggingface-hub>=1.23.0,<2.0 [...] we can conclude that your
  requirements are unsatisfiable.

This broke the vibevoice build on every variant - cpu (amd64/arm64),
cuda 12/13, l4t 12/13, intel and rocm, plus the darwin metal job - and
has been failing the weekly full-matrix rebuild for three weeks. It is
not caught by master pushes because backend builds are path-filtered
there, so it only surfaces on the Sunday cron and on release tags.

Use the PyPI package instead. That is what upstream VibeVoice declares
in its own pyproject.toml, and what every other LocalAI backend already
does - vibevoice was the only one tracking the git branch. With a real
release series available the resolver settles on diffusers 0.39.0 with
huggingface-hub 0.36.2 and transformers 4.57.6, and it can keep
backtracking on its own if upstream shifts again.

Verified with uv pip compile against cpu, cublas12, cublas13, hipblas,
intel, mps and l4t13: all resolve to that same coherent set. l4t12 only
resolves on aarch64, since its Jetson index ships no x86_64 torch wheel.


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Bash] [uv]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 08:27:25 +02:00
Nandana Dileep
b5e4413eab feat: add MiniMax-M3 model support (#10837)
Adds inference parameter defaults for the minimax-m3 model family and
includes a vendored patch of upstream llama.cpp PR #24523 to recognize
the minimax-m3 architecture. Once the upstream PR merges, the patch can
be removed and LLAMA_VERSION bumped normally.

Changes:
- backend/cpp/llama-cpp/patches/0001-add-minimax-m3-support.patch:
  vendored patch from ggml-org/llama.cpp#24523 (Preliminary MiniMax-M3
  support). Applied by prepare.sh during the build; keeps the pinned
  LLAMA_VERSION pointing at the latest upstream tag.
- core/config/inference_defaults.json: add minimax-m3 family entry
  (temperature=1.0, top_p=0.95, top_k=40, min_p=0.01,
  repeat_penalty=1.0, matching the existing minimax defaults) and
  register it in the patterns list before the shorter minimax-m2.7
  entry for correct longest-match-first ordering.

Upstream: depends on ggml-org/llama.cpp#24523
Closes: https://github.com/mudler/LocalAI/issues/10820

Signed-off-by: Nandana Dileep <110280757+nandanadileep@users.noreply.github.com>
2026-07-20 08:26:51 +02:00
mudler's LocalAI [bot]
e55cc3e2a7 fix(worker): bound the gRPC port allocator and stop leaking dead backends' ports (#10968)
The worker's gRPC port allocator grew monotonically with no upper bound:
nextPort started at the base port and incremented whenever freePorts was
empty, and nothing checked 65535. Past that it handed out integers that
cannot be bound, surfacing as an opaque "backend won't start".

#10961 estimated this needed ~15,000 concurrent-peak allocations, i.e.
effectively unreachable. It is not, because of a second defect: the
"process died unexpectedly" branch in startBackend deleted the process
map entry without releasing its port at all. That port was leaked, never
quarantined and never reused. A crash-looping backend leaks one port per
restart, so a backend dying every 30s walks 50051 to 65535 in about five
days. The leak, not concurrent peak, is the realistic route to exhaustion.

Fixing the leak alone would have been wrong. Releasing that port makes it
re-bindable, and the death path is the one teardown path with no
request/reply to carry StoppedProcessKeys back to the controller (#10952's
eager row removal), so a stale NodeModel row could then resolve to a live
listener belonging to a different backend. probeHealth verifies liveness,
not identity, so the request is silently misrouted. The 15s port
quarantine does not cover this: the only reaper is the per-model health
check at ~45s, and it can be disabled outright. The residual was masked
only because the port was never rebound.

So both are fixed together:

- The allocator takes an explicit [basePort, LOCALAI_GRPC_MAX_PORT] range
  and returns ErrNoFreePort naming the range, the live backend count, the
  quarantined count, and the knob to raise. Exhaustion is now diagnosable
  instead of surfacing as an unbindable port.

- Released ports carry per-key affinity: a port is offered back to the
  process key that last held it before any other key. Process keys
  (modelID#replica) and NodeModel rows (nodeID, modelName, replicaIndex)
  are isomorphic, so a port that can only be re-bound by its previous
  owner can only ever be named by that owner's row, which that key's
  re-registration overwrites. Misrouting to a different model becomes
  impossible by construction rather than by racing the quarantine timer.

Affinity is a preference, not a reservation: under range pressure an owned
port is stolen with a warning, because a guaranteed outage is worse than a
rare misroute window on a port long out of quarantine. Claiming a port
evicts its previous owner's entry, keeping ownership injective over ports
so the affinity map can never exceed the range width regardless of how
many distinct model keys the worker sees.

Ownership also expires. It is only load-bearing while a controller row
could still name the port, which the per-model reaper bounds at roughly
45s, so it lapses after five minutes and the port becomes ordinary free
space again. Holding it indefinitely would have made every distinct model
the worker ever served consume a port permanently: every release path is
keyed, so nothing would ever be unowned, the allocator would climb to the
end of its range on distinct-key count rather than concurrency, stealing
would become routine, and the steal warning would tell operators to widen
a range that was not the constraint. With expiry, reaching the steal
branch means the worker is genuinely out of concurrent capacity, so that
advice is correct when it appears.

Closes #10961
Closes #10952


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-19 23:56:37 +00:00
mudler's LocalAI [bot]
9d82c37f98 fix(distributed): backend discovery hid worker-installed backends behind the controller's filesystem (#10967)
fix(distributed): backend discovery hid worker-installed backends

Backend discovery endpoints filter on installed-state, which on a
distributed controller derives from the controller's own filesystem. A
backend lives on the worker node that runs it, so every backend an admin
installed on a GPU worker read as "not installed" and vanished from the
listing. #10947 fixed the sibling capability filter on the same endpoints,
so a fine-tuning-capable GPU worker now made the backend listable while
the installed-state filter still dropped it: the dropdown stayed empty.

The controller cannot derive this locally, but it already aggregates the
per-node view that GET /backends renders, so discovery reuses the active
BackendManager rather than growing a second path. Three surfaces shared the
root cause and route through the same helper now:

  - GET /backends/available (Installed is now cluster-wide)
  - GET /api/fine-tuning/backends
  - GET /api/quantization/backends

The response stays a boolean rather than an installed-on-N-of-M count:
per-node install state is already served by GET /backends nodes[], and
per-node control by POST /api/nodes/:id/backends/install, so a summary is
all these dropdowns need.

A nil provider (single-node) leaves the local filesystem as the only source
and reproduces today's listing exactly, and a registry error degrades to
that same listing instead of blanking the catalog.


Assisted-by: Claude:claude-opus-4-8 golangci-lint

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 01:09:37 +02:00
mudler's LocalAI [bot]
f735cb24c0 fix(worker): reap deleted backends and stop models that live on a worker (#10956)
* fix(worker): reap deleted backends and stop models that live on a worker

Three related backend-lifecycle defects, all reachable from the same
production incident on a Jetson/Thor worker: a deleted backend's gRPC
process survived ~40 minutes with its directory removed from disk, a later
model load was routed to that orphan and failed with a certifi path pointing
into the deleted directory, and the admin could not stop the model because
the frontend reported it as not loaded.

1. backend.delete orphaned the process it claimed to delete
------------------------------------------------------------
s.processes is keyed by `modelID#replicaIndex` (buildProcessKey), so the
backend name never appeared in a key and was recorded nowhere on the
process. backend.delete resolved its target via isRunning/stopBackend, whose
prefix path only matches a bare *modelID* - a delete keyed on a backend name
resolved to zero keys, the stop silently no-op'd, and the files were removed
out from under a live process.

The install fast path then handed that orphan back out: it returns any live
process for the (model, replica) slot without checking which backend started
it, so a reinstalled variant inherited the deleted backend's port.

- Record backendName on backendProcess, threaded installBackend ->
  startBackend.
- Add resolveProcessKeysForBackend, matching the recorded name and resolving
  alias <-> concrete via ListSystemBackends *before* DeleteBackendFromSystem
  erases the metadata that carries the alias. Alias resolution failure
  degrades to name-only matching so a delete never fails on it.
- backend.stop goes through resolveStopTargets, which accepts a backend
  name, a model name, or an exact modelID#replica key. Its payload field is
  named "backend" but is published with all three meanings: the admin UI
  sends a backend name, UnloadRemoteModel sends a model name, and the
  router's abandoned-load reap (#10948) sends an exact replica key.
  Narrowing it to backend names alone would strand the latter two.
  backend.delete stays strict - its identifier is unambiguously a backend.
- Gate the install fast path on processMatchesBackend so a slot held by a
  different backend is restarted rather than reused. Processes with no
  recorded name (pre-upgrade) are accepted, so rollout does not restart
  every running backend.
- stopBackendExact reports a real stop failure - the process still being
  alive afterwards, which is precisely what finishBackendStop already
  detects to keep the entry and its port reserved - and backend.delete no
  longer replies success when it knew about a process and could not kill it.
  "No process was running" stays a success but is logged, so the orphan case
  is visible rather than silent.

2. /backend/shutdown reported a running model as missing
---------------------------------------------------------
ModelLoader.deleteProcess short-circuits on a miss in this replica's
in-memory store. In distributed mode the authoritative record of "is this
model loaded" is the shared node registry: a frontend replica that never
served the model itself (load balancer picked a peer, or the replica
restarted) has no local entry. The remote unload path that pkg/model
documents ("when ShutdownModel is called for a model with no local process,
UnloadRemoteModel is called") sat behind that short-circuit, unreachable in
exactly the case it exists for. #10865 reworked this function but kept the
short-circuit at the top, so the gap survived that refactor.

- deleteProcess consults the remote unloader on a local-store miss, via a
  shared unloadRemote helper so this branch and the existing
  no-local-process branch both prefer #10865's RemoteModelContextUnloader,
  preserving force propagation across the distributed boundary.
- UnloadRemoteModelContext reports ErrRemoteModelNotLoaded when no node has
  the model; it previously returned nil, making a no-op stop
  indistinguishable from a real one. The converse case (nodes have it, none
  could be stopped) already errors since #10865 joined the per-node
  failures, so that half of the original fix was dropped as redundant.
- Only when the model is absent locally AND cluster-wide does the endpoint
  report not-found, now 404 naming both scopes rather than a bare 500.
- modelNotFoundErr becomes the exported ErrModelNotFound so the HTTP layer
  can map it without string matching; watchdog's identity comparison becomes
  errors.Is.

3. Coverage for the bounded Free() that #10865 shipped untested
----------------------------------------------------------------
The original branch also bounded the pre-stop Free(), but #10865 landed that
fix first (workerBackendFreeTimeout, applied in both stopBackendExact and
handleModelUnload). That production change is therefore DROPPED here as
superseded - master's version is strictly better, since it also releases the
supervisor mutex across the call and keeps the port reserved until
termination completes.

What #10865 did not ship is a test, and the bound is load-bearing: the
router-side reap in #10948 sends backend.stop for an abandoned load, and
against a wedged backend an unbounded Free would swallow that stop before it
reached the process. Nothing failed if the bound regressed.

The spec stands up a real gRPC backend server whose Free handler never
returns - what a Python backend looks like when its single worker thread
(PYTHON_GRPC_MAX_WORKERS=1 on 37 backends) is occupied by a stuck LoadModel.
A stub socket is not sufficient and was tried first: without a completed
HTTP/2 handshake, gRPC's own ~20s connect timeout ends the call, so that
version passed against the very bug it targets. With the connection READY,
only the caller's deadline can end it, so the spec hangs to its 60s limit if
the timeout is removed and passes with it.

Its fixture process is deliberately never started. go-processmanager v0.1.1
writes Process.pid from readPID() without synchronization, so a live process
races its own monitor goroutine under -race - reproducible with a bare
Run()+Stop() and unrelated to this spec. Since
scripts/model-lifecycle-conformance.sh runs this package with -race and is
fail-closed, starting one would turn that gate red on an upstream defect. An
unstarted process still proves the point: the stop is reached and the slot
released, which is exactly what an unbounded Free prevents.

Verified: make lint (new-from-merge-base origin/master) reports 0 issues;
scripts/model-lifecycle-conformance.sh passes all three stages including the
FizzBee liveness check (1458 states, IsLive: true).

Assisted-by: Claude:claude-opus-4-8 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): keep remote unload idempotent, ask presence separately

2035a4d25 made UnloadRemoteModel return ErrRemoteModelNotLoaded when no node
holds the model, so ShutdownModel could answer 404 instead of a misleading
500. That narrowed a shared adapter contract to serve one caller and broke
the documented idempotent-unload guarantee, which CI caught on PR #10956:

  [FAIL] Node Backend Lifecycle (NATS-driven) > NATS backend.stop events
         should be no-op for models not on any node [Distributed]
         Expected success, but got: model not loaded on any node

The spec name states the contract outright. The matching unit assertion was
updated in that commit; this e2e one was missed because it lives under
tests/e2e/ with no build tags and does not run in package-scoped test runs.

Caller audit - who breaks when an idempotent unload becomes an error:

- pkg/model/watchdog.go:902 (LRU memory reclaimer) is the serious one. It
  untracks a model ONLY when shutdown returns nil or ErrModelNotFound. A new
  error type means the model is never untracked, so the reclaimer keeps
  re-selecting the same entry and never reclaims - a live wedge whenever a
  local store entry outlives the remote model.
- core/services/galleryop/managers_local.go:43 (DeleteModel) would warn on
  every deletion of an already-unloaded model.
- core/services/modeladmin/{state,config,remote_sync}.go stop instances
  best-effort against models that are frequently not loaded.
- deleteProcess itself: the no-local-process branch returns the unload result
  directly, so a stale local entry for a model no longer on any node turned a
  previously-successful cleanup into a failure.

Only ShutdownModel wants the distinction, and only on the local-store-miss
path. So the distinction moves to the caller instead of the contract:

- UnloadRemoteModel/UnloadRemoteModelContext return nil again when no node
  has the model, and ErrRemoteModelNotLoaded is removed.
- New optional RemoteModelPresenceChecker (HasRemoteModel) answers the
  question directly. deleteProcess consults it BEFORE unloading, because an
  idempotent unload cannot report afterwards whether anything was stopped.
  Absent locally AND cluster-wide is the only case that reports 404.
- A failed registry lookup is surfaced rather than reported as absence: an
  unreachable registry is not evidence a model is gone, and answering a
  confident 404 off a failed lookup is how an operator gets told a running
  model does not exist.
- Unloaders that predate the extension keep working - deleteProcess attempts
  the unload rather than refusing it - and compile-time assertions in the
  nodes package now pin all three optional interfaces, since both are
  consumed by runtime type assertion where drift degrades behavior silently
  instead of failing the build.

The contract is now pinned at both levels that disagreed, each spec pointing
at the other: "with no nodes returns nil" in unloader_test.go and "should be
no-op for models not on any node" in node_lifecycle_test.go.

Verified: full distributed e2e suite 233 passed / 0 failed (the suite that
failed 232/1 in CI); pkg/model and core/services/nodes green; make lint
new-from-merge-base reports 0 issues.

Assisted-by: Claude:claude-opus-4-8 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): drop replica rows when a worker stops a backend

A worker returns a stopped backend's gRPC port to its allocator as soon as
the process is confirmed dead, and hands it to the next backend that starts.
The controller's NodeModel row for the old address survives, and both
SmartRouter.probeHealth and the HealthMonitor per-model probe verify
liveness, not identity, so once an unrelated backend binds the recycled port
the stale row passes every check and the request is served by the wrong
backend instead of failing.

backend.delete is newly able to trigger this: before #10956 a delete never
actually stopped a process, so it never recycled a port. backend.upgrade has
the identical gap and always did — upgradeBackend force-stops every process
using the binary and starts none back up, while
DistributedBackendManager.UpgradeBackend never removes rows. model.unload is
the one path that gets this right today: it calls RemoveAllNodeModelReplicas
straight after StopBackend.

Report the process keys the worker terminated on the delete and upgrade
replies, and drop the matching rows in RemoteUnloaderAdapter, which already
holds a ModelLocator with RemoveNodeModel. All three call sites funnel
through that adapter, so no new interface, DB migration, or proto change is
needed. A key is reported only once its process is confirmed gone, so the
list stays trustworthy on the partial-failure replies too.

Old workers never populate the new fields. ReportsStoppedProcesses tells
"stopped nothing" apart from "does not report", so an old worker's silence
falls back to the pre-existing probe-based staleness recovery instead of
being mistaken for a completed cleanup.

Quarantine released ports for a short window as an interlock covering the
NATS round-trip between the worker freeing the port and the controller
dropping the row. It is deliberately not derived from HealthCheckInterval:
that cadence is operator-tunable and the per-model reaper can be disabled
outright, so coupling a worker-local constant to it would be silently wrong
on some clusters. Eager row removal is the fix; the delay only closes the
handoff gap.

Identity verification in probeHealth was considered and rejected: Health and
Status carry no backend identity, so it needs a proto change plus an
implementation in 36 Python and 4 C++ Health servicers, it is fail-open for
any backend not yet rebuilt, and the probeCache short-circuit means it would
not even execute during the 30s window where the misroute happens.

Fixes #10952
Refs #10954, #10956

Assisted-by: Claude:claude-opus-4-8 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(deps): bump go-processmanager, assert real backend termination

go-processmanager wrote Process.PID from readPID() with no synchronization
while its own monitor goroutine cleared the same field on exit, so a bare
Run()+Stop() tripped the race detector without any concurrent access from
the caller. LocalAI hit this on every backend stop.

Upstream fixed it in a94e2b7 by guarding PID with a mutex and adding
CurrentPID() as a race-safe accessor. The exported field was kept to avoid
a breaking change but is now deprecated: a direct read still races the
monitor. No tag carries the fix yet, so pin the pseudo-version.

GetGRPCPID reads through CurrentPID() instead of the field. The accessor
returns the same string under an RLock, so the empty-PID and strconv error
paths are unchanged; it is the only direct field read in the tree.

With the race gone, the Free-timeout spec no longer has to leave its
fixture process unstarted. It now runs a real child and asserts the child
genuinely exits, which is exactly what the earlier workaround gave up: the
spec could show the stop was reached and the slot released, but not that
SIGTERM ever landed. Termination is observed through Done(), which closes
only once the library has waited on the child. The pidfile-based liveness
helpers cannot serve here, because Stop() deletes the pidfile while
releasing the handle and so reports "not alive" even if no signal was sent.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 01:09:22 +02:00
mudler's LocalAI [bot]
5c607c09d5 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 339e3d7fc7161f8ae61d22c291ff40f68b690266 (#10962)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-20 00:38:25 +02:00
mudler's LocalAI [bot]
2f7b292143 chore: ⬆️ Update CrispStrobe/CrispASR to 5fca47ecf05cd68bb0075f8a00fe04da06f208d0 (#10963)
* ⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(crispasr): initialize only declared submodules

The latest upstream commit contains an undeclared CrispASR gitlink that makes a blanket recursive submodule update fail. Limit initialization to the two submodules used by the backend build.

Assisted-by: Codex:gpt-5 [Codex]

* fix(crispasr): resolve vendored WebRTC from project root

CrispASR now builds a vendored WebRTC VAD, but its include paths assume CrispASR is the top-level CMake project. Extend the existing embedded-project rewrite to the shared third_party root.

Assisted-by: Codex:gpt-5 [Codex]

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-20 00:38:13 +02:00
zjuzhongwen
864c84f48b chore: fix some comments to improve readability (#10960)
Signed-off-by: zjuzhongwen <zjuzhongwen@outlook.com>
2026-07-20 00:37:40 +02:00
mudler's LocalAI [bot]
8cef340659 chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to e93292bee1778854ab7dcb2d325ffe531fef910f (#10964)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-20 00:37:23 +02:00
mudler's LocalAI [bot]
c2704dba5b fix(gpu): detect GPUs via sysfs when no pci.ids database is present (#10966)
* fix(gpu): detect GPUs via sysfs when no pci.ids database is present

ghw.GPU() calls pci.New() before it reads /sys/class/drm and fails
outright when it cannot find a pci.ids database file. jaypipes/pcidb
embeds no database and has network fetch disabled by default, so on an
image that ships no pci.ids, GPU enumeration returns an error and every
detection path downstream goes dark.

The Dockerfile installs pciutils only in the vulkan and cublas branches,
so the Intel image had no pci.ids. A correctly passed-through Arc A310
was reported as "No GPU detected" with zero VRAM even though clinfo and
sycl-ls both enumerated it inside the same container. NVIDIA and AMD
images were shielded by their nvidia-smi / rocm-smi binary fallbacks;
Intel has no equivalent, leaving it fully exposed.

Read PCI vendor IDs directly from /sys/class/drm/card*/device/vendor,
which needs no database, and consult that from DetectGPUVendor. The
same scan replaces the ghw-only guard in getIntelGPUMemory, which is
what had been blocking the working clinfo path and keeping VRAM at
zero. Install hwdata in the base image stage as well, so ghw stops
failing for every image variant rather than only Intel.

Also apply the documented NVIDIA > AMD > Intel priority to the ghw
path, which previously returned whichever card DRM enumerated first
and so reported "intel" on a machine with an Intel iGPU at card0 and
an NVIDIA dGPU at card1.

HasGPU() carried the same blindness plus one of its own: it matched
the requested vendor against ghw's card description with a
case-sensitive Contains, so "nvidia" never matched the pci.ids
spelling "NVIDIA Corporation". It only worked because that same
description embeds the lowercase kernel driver name ("nvidia",
"amdgpu"), and it returned false outright whenever ghw errored. Route
it through the shared vendor lookup so it matches case-insensitively
and falls back to sysfs. It feeds the GPU option and NGPULayers
defaults in core/config/gguf.go.

Fixes #10941

Assisted-by: Claude:claude-opus-4-8 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* refactor(gpu): key vendor detection off the numeric PCI ID in both paths

The ghw and sysfs legs were identifying vendors by different means: ghw
by substring-matching the pci.ids vendor name, sysfs by the numeric PCI
vendor ID. ghw already exposes that same numeric ID via
DeviceInfo.Vendor.ID, read from the kernel's modalias rather than from
the database, so the name matching was both a duplicate mechanism and
the weaker of the two.

It is weaker because a card absent from an outdated pci.ids gets
Name: "unknown" while its ID is still correct. Detection then failed
even though ghw had enumerated the card successfully. Verified in a
container with a vendor-less pci.ids and an Arc's modalias: before,
DetectGPUVendor returned ""; after, "intel".

Both legs now resolve through the same pciVendorIDs table and share the
hex parsing, with the vendor name kept only as a fallback for devices
exposing no parseable ID.

ghwHasVendor is deliberately not a priority pick, unlike vendorFromGHW:
HasGPU("intel") must stay true on a hybrid-graphics host whose discrete
NVIDIA card outranks the integrated Intel one.

Assisted-by: Claude:claude-opus-4-8 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gpu): silence the gosec G304 on the sysfs attribute read

gosec flags os.ReadFile with a non-literal path. The path here is the
DRM root (a package constant in production, a temp dir under test)
joined with a ReadDir entry name and a fixed attribute filename, so no
external input reaches it.

gosec's suggested autofix, os.Root, cannot be used: /sys/class/drm/cardN
is a symlink into the PCI device tree, and os.Root refuses to traverse
it ("path escapes from parent"), which would disable the whole scan.

Assisted-by: Claude:claude-opus-4-8 gosec golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-20 00:37:06 +02:00
mudler's LocalAI [bot]
92dc326606 chore(model-gallery): ⬆️ update checksum (#10965)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-19 23:35:30 +02:00
Tai An
217fdd2234 fix(qwen-asr): map ISO language codes to the names Qwen3-ASR expects (#10959)
request.language usually carries an ISO 639-1 code (e.g. "de"), which
OpenAI-compatible clients such as Home Assistant / wyoming_openai send,
but qwen_asr.validate_language() only accepts full English names
("German") and raises ValueError otherwise. Normalize the requested
language: accept full names case-insensitively, translate ISO codes
(with optional region suffix like "de-DE") to the expected name, and
pass anything unrecognised through so qwen_asr still reports it clearly.

Fixes #10958

Signed-off-by: Tai An <antai12232931@outlook.com>
2026-07-19 22:00:23 +02:00
mudler's LocalAI [bot]
0e0221b0f5 fix(vision): probe the media marker for pinned llama.cpp backend variants (#10955)
llama.cpp picks a random per-process media marker (ggml-org/llama.cpp#21962),
so LocalAI renders the prompt with a "<__media__>" sentinel and swaps in the
backend's real marker after probing ModelMetadata.

That probe was gated on an exact match against "llama-cpp", the gallery's meta
backend name. A model config pinning a concrete build ("vulkan-llama-cpp",
"cuda12-llama-cpp", "rocm-llama-cpp", ... and their -development counterparts)
runs the same llama.cpp gRPC server but skipped the probe, so MediaMarker
stayed empty, no substitution happened, and the prompt reached mtmd still
carrying the sentinel. mtmd_tokenize then counted zero markers against one
bitmap and every image request failed with "Failed to tokenize prompt".

The same early return also skipped thinking-mode detection and tool-format
marker extraction, so a pinned variant silently lost reasoning and native
tool-call parsing too.

Add IsLlamaCppBackend, which recognises the whole variant family (plus the
empty auto-detect name, which resolves to llama.cpp) while excluding
ik-llama.cpp, a separate engine that merely shares the suffix.

Fixes #10945


Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-19 12:46:50 +02:00
mudler's LocalAI [bot]
fb4c61d1c9 fix(distributed): configurable remote model-load timeout, and reap the load when it times out (#10948)
* fix(distributed): make the remote LoadModel deadline configurable

The router hardcoded a 5 minute gRPC deadline for the remote LoadModel
call. Staging finishes before the timer starts, so those five minutes
cover only the worker backend's own checkpoint load and pipeline init.
A cold load of meituan-longcat/LongCat-Video-Avatar-1.5 (~83 GB) on an
ARM64 Thor worker fails at exactly 302s with DeadlineExceeded while the
backend process is still making progress (CPU time accumulating, RSS
moving as weights are mapped), so the load was cut short rather than
wedged.

Add LOCALAI_NATS_MODEL_LOAD_TIMEOUT / --model-load-timeout mirroring the
existing backend-install timeout knob, defaulting to 5m so unset
clusters keep today's behaviour.

The cold-load hold ceiling (which bounds how long one load may hold the
per-model advisory lock) was derived from the install timeout alone, so
raising the load deadline past it would have been silently clipped.
Derive it from both budgets via ModelLoadCeilingFor:

    max(install + load + 5m staging margin, 25m)

With the defaults that is 15m + 5m + 5m = 25m, identical to the previous
constant, and the 25m floor means shrinking either budget can never
tighten the ceiling below what clusters relied on before.

Assisted-by: Claude:claude-opus-4-8 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): reap the abandoned replica when a remote load times out

The gRPC deadline on the remote LoadModel call only cancels the client
side. A backend blocked in a synchronous weight load never observes its
cancelled handler context, so when scheduleAndLoad gave up it left the
worker loading with nobody waiting for the result.

Observed on an ARM64 Thor worker loading LongCat-Video-Avatar-1.5: the
client returned DeadlineExceeded at 302s, and the backend process was
still alive 30 minutes later having pulled ~57GB from HuggingFace. Every
retry stacked another multi-GB loader on the worker; they had to be
reaped by hand via POST /api/nodes/:id/models/unload.

Send backend.stop for the exact `modelID#replicaIndex` process key we
just abandoned. The exact key matters: a bare model ID stops every
replica on that node, including healthy ones serving traffic.

Only a deadline or cancellation triggers the reap. Any other LoadModel
failure is the backend answering, which means its handler returned and
the process is idle - stopping it there would discard a warm process and
its downloaded weights. The reap is best-effort and never replaces the
load error the caller is waiting on.

The `modelID#replicaIndex` format was already hand-rolled in two places
(the worker's buildProcessKey and pkg/model's log store). Rather than add
a third, export model.BackendProcessKey from pkg/model, the lowest common
dependency of both sides.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 golangci-lint

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-19 12:01:48 +02:00
mudler's LocalAI [bot]
626ae4d51e fix(model-artifacts): materialize longcat-video on the controller, and support companion repos (#10949)
* fix(model-artifacts): materialize longcat-video checkpoints on the controller

longcat-video loads a checkpoint directory: its backend.py takes
request.ModelFile when os.path.isdir(request.ModelFile) and otherwise
falls back to snapshot_download. That places it in the same class as
transformers/vllm/diffusers/sglang, but the allow-list added in #10910
did not enumerate it, so PrimaryArtifactSpec returned no managed
artifact for a bare HuggingFace repo id.

The consequence in distributed mode: nothing was acquired on the
controller, ModelFileName fell through to the raw repo id, and staging
skipped the resulting phantom /models/<owner>/<repo> path. The worker
received a blank ModelFile, fell back to request.Model, and downloaded
~83GB from HuggingFace inside the remote LoadModel deadline - so the
load could only ever fail with DeadlineExceeded while an abandoned
backend process kept downloading.

Note this materializes the full repository. The backend restricts its
own snapshot_download with allow_patterns, and the avatar repo ships
both base_model/ and base_model_int8/ where only one is ever loaded;
inferred specs have no way to carry patterns today. Tracked separately.

Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): warn when staging skips a non-existent model path

stageModelFiles logs "Staging model files for remote node" up front, then
silently drops any path field that does not exist on the controller. The
skip itself is legitimate and must stay: a backend outside
managedArtifactBackends that takes a bare HuggingFace repo id gets an
optimistically constructed path (ModelFileName falls through to the raw
model reference) that was never materialized, and sources its own weights
on the worker. Erroring would break those configs.

But at debug level the operator is left with a reassuring staging line and
no trace of the skip, so a genuine controller-side acquisition gap is
indistinguishable from a healthy pass-through - it surfaces much later as
a remote LoadModel timeout, on a worker that is quietly downloading tens
of gigabytes. Raise the skip to warn and name the field, path, node and
tracking key. Behavior is unchanged.

Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(model-artifacts): allow a config to declare companion artifacts

A composed pipeline needs more than one HuggingFace snapshot.
LongCat-Video-Avatar-1.5 loads its own transformer but takes the
tokenizer, text encoder and VAE from the separate LongCat-Video base
repo, so a single-artifact config cannot express it and the backend is
left to fetch the second repo itself at load time.

Widen the artifact model to target: model plus any number of named
target: companion entries. Normalize accepts the new target and
constrains a companion name to [a-z0-9][a-z0-9_-]{0,63} because that
name is the option key the backend later receives; a companion may not
claim primary_file, which only means anything for a load target.
ModelConfig.Validate requires exactly one primary and requires it first,
since Artifacts[0] is what ModelFileName, size estimation and staging all
resolve from.

Both acquisition paths now loop instead of touching index 0 alone:
preloadOne for an already-installed config, bindPrimaryArtifact for a
gallery install. Failure policy differs by provenance. An inferred
primary keeps its warn-and-fall-back, because the legacy download path
still exists for it. Companions are explicit by construction, so they are
all-or-nothing: a config naming one is asserting the backend needs it,
and failing at the acquisition boundary is far more legible than a
missing-weights error surfacing later inside the backend.

The cache key is deliberately unchanged. It hashes source identity only,
never name or target, so every already-installed managed model still hits
its existing snapshot instead of silently re-downloading. Two specs pin
that: one proving a companion and a primary with identical sources agree
on the key, and one pinning the digest of a known primary outright.

Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(model-artifacts): hand resolved companion snapshots to the backend

A materialized companion is useless until the backend can find it, and
its location is a content-addressed cache key that does not exist until
the artifact resolves. A static gallery override cannot carry that, and
persisting it into the config YAML would rot the moment a re-resolve
produced a new key.

Synthesize it instead at load time: each resolved companion becomes
"<artifact name>:<snapshot path>" in ModelOptions.Options, reusing the
key:value convention backends already parse for options like
attention_backend. The value stays relative to the models directory so a
remote worker can resolve it under its own ModelPath once staging has
rewritten the model root. An option the author set explicitly always
wins, so pinning a companion to a local checkout still beats the managed
snapshot.

longcat-video resolves base_model through ModelPath, the same convention
qwen-tts, voxcpm, outetts and ace-step already use for companion assets.
Its sibling-directory heuristic is deleted: it looked for a LongCat-Video
directory next to the model, which cannot exist under the content
addressed .artifacts/huggingface/<key>/snapshot layout, so it was dead
code the moment the model became managed.

The gallery entry declares both repositories and restricts each with
allow_patterns. The avatar repo ships base_model/ and base_model_int8/
and only ever loads one, so fetching the whole repo would roughly double
the download. The patterns match the entry's own options (use_distill
true, use_int8 default false); enabling use_int8 here also requires
adding base_model_int8/**, which is called out in the entry.

Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): stage managed artifact trees from the models root

Staging anchored the worker's models directory on the primary snapshot
whenever a model was managed, so a companion snapshot could not reach the
worker at all.

frontendModelsDir was derived by stripping the Model relative path off
the end of ModelFile. For a managed artifact nothing matches: ModelFile
is .artifacts/huggingface/<key>/snapshot while Model stays a bare
HuggingFace repo id, so the strip was a no-op and the "models directory"
came out as the snapshot itself. Two consequences, both silent. Staging
keys lost the .artifacts/huggingface/<key>/snapshot prefix, so two
snapshots of one model were indistinguishable on the worker. And a
companion, which lives in a sibling snapshot directory outside the
primary, fell outside that directory entirely: StagingKeyMapper.Key
collapsed its files to bare basenames and resolveOptionPath could not
resolve the relative option at all, so it was skipped without a word.

Derive the models root from the artifact tree instead when the path runs
through it, and compute the worker's ModelPath from the file's path
relative to that root rather than from the Model field. The legacy layout
is unaffected: where Model really is the relative path, the new
derivation reduces to the old one, which a regression spec pins.

This deliberately changes an invariant that router_dirstage_test.go
pinned: for a managed primary, ModelFile and ModelPath were both the
snapshot directory, and staging keys were relative to it. Now ModelFile
is the snapshot, ModelPath is the models root above it, and keys keep the
full relative path. That spec is updated rather than accommodated, with
the reasoning recorded inline, because the old invariant is exactly what
made a sibling companion unreachable.

Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-19 12:01:36 +02:00
localai-org-maint-bot
09b85ee00e fix(ci): build Bonsai backend images (#10939) (#10951)
fix(ci): build Bonsai backend images

Register the Bonsai C++ source path with the backend matrix filter so changes select its image jobs. Also make shared llama.cpp changes rebuild the Bonsai and Turboquant fork images in the actual matrix, not only their test flags.\n\nAssisted-by: Codex:gpt-5 [Codex]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-19 11:49:07 +02:00
mudler's LocalAI [bot]
b19afb192a fix(distributed): backend discovery hid GPU-only backends behind the controller's capability (#10947)
* fix(backends): list backends runnable on worker nodes in distributed mode

GET /backends/available filtered the gallery against the system state of
the host serving the request. In a distributed deployment that host is the
controller, which typically has no GPU, while the GPUs live on worker
nodes. Any meta backend whose capabilities map lacks a "default" (or "cpu")
key was therefore dropped from the listing entirely — longcat-video,
vllm-omni, ltx-video, parakeet, edgetam and qwentts were invisible in the
UI even though installing them by name on a GPU worker worked fine.

Workers now report their own meta-backend capability at registration and
the controller persists it on the node row. The controller cannot derive
it: OS-dependent capabilities (metal, darwin-x86, nvidia-l4t) and the CUDA
runtime refinements are only observable on the worker. Nodes registered
before this field existed fall back to a coarse capability derived from
their GPU vendor and VRAM.

Backend discovery then evaluates compatibility as the union over healthy
backend nodes, so a backend runnable on any node is offered while one no
node can run stays hidden. Each remote capability is evaluated through a
capability-pinned system state, otherwise a forced capability on the
controller image (LOCALAI_FORCE_META_BACKEND_CAPABILITY or
/run/localai/capability) would silently override every worker's verdict.
With no registered nodes the listing is byte-for-byte what it was, so
single-node deployments are unaffected.

Also fixes the same-root-cause misclassification in /api/operations, which
used the capability-filtered listing to decide whether an operation was a
backend or a model install. A GPU-only backend installing on a worker is
still a backend operation on the controller, so that lookup is now
unfiltered.

Assisted-by: Claude:claude-opus-4-8 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(backends): union worker capabilities in backend discovery

Implementation for the specs added in the previous commit, plus the two
remaining discovery endpoints.

Capability-filtered backend discovery evaluated compatibility against the
system state of the host serving the request. In a distributed deployment
that host is the controller, which typically has no GPU, while the GPUs
live on worker nodes. Any meta backend whose capabilities map lacks a
"default" (or "cpu") key was dropped entirely — longcat-video, vllm-omni,
ltx-video, parakeet, edgetam and qwentts were invisible in the UI even
though installing them by name on a GPU worker worked fine.

Workers now report their own meta-backend capability at registration and
the controller persists it on the node row. The controller cannot derive
it: OS-dependent capabilities (metal, darwin-x86, nvidia-l4t) and the CUDA
runtime refinements are only observable on the worker. Nodes registered
before this field existed fall back to a coarse capability derived from
their GPU vendor and VRAM.

Discovery then evaluates compatibility as the union over healthy backend
nodes, so a backend runnable on any node is offered while one no node can
run stays hidden. Each remote capability is evaluated through a
capability-pinned system state, otherwise a forced capability on the
controller image (LOCALAI_FORCE_META_BACKEND_CAPABILITY or
/run/localai/capability) would silently override every worker's verdict.
With no registered nodes the listing is byte-for-byte what it was, so
single-node deployments are unaffected.

Four surfaces shared this root cause and are all routed through the same
helper now:

  - GET /backends/available
  - GET /api/fine-tuning/backends
  - GET /api/quantization/backends
  - /api/operations backend-vs-model classification, which additionally
    had no reason to filter by capability at all: a GPU-only backend
    installing on a worker is still a backend operation on the
    controller, so that lookup is now unfiltered.

Assisted-by: Claude:claude-opus-4-8 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-19 07:53:46 +00:00
mudler's LocalAI [bot]
963c637130 fix(gpu-libs): bundle cuDNN only where it is used, and complete it when it is (#10946)
cuDNN 9 is a dispatcher (libcudnn.so.9) plus seven sublibraries the dispatcher
dlopen()s by bare soname. Only the dispatcher is ever a DT_NEEDED, so ldd finds
it and never the seven. The allowlist force-copied three of them
(libcudnn.so*, libcudnn_ops.so*, libcudnn_cnn.so*) into every CUDA backend,
which is wrong in both directions at once: too few libraries for a backend that
uses cuDNN, and too many for one that does not.

On an L4T fleet, ten of the eleven backends carrying cuDNN were in a broken end
state; the one that was correct was correct by accident, being BUILD_TYPE=cpu
so package_cuda_libs never ran for it.

  longcat-video bundled 4 of 8 at 9.24.0 over a complete pip set at 9.20.0.48
  in its venv. libbackend.sh puts lib/ on LD_LIBRARY_PATH, searched before
  DT_RUNPATH, so the bundle won and the rest still came from the venv:
  CUDNN_STATUS_SUBLIBRARY_VERSION_MISMATCH.

  Nine others bundled 3 of 8 and had no venv cuDNN. None bundled
  libcudnn_graph, which libcudnn_cnn has a hard DT_NEEDED on, so it resolved
  out of the runtime image and the process ran bundled 9.22.0 against system
  9.23.2.

Five of those nine - llama-cpp, whisper, rfdetr-cpp, sam3-cpp,
stablediffusion-ggml - do not reference cuDNN at all. ggml goes through cuBLAS.
They were carrying ~57 MB of cuDNN with no consumer, and completing the family
for them would have taken that to ~576 MB for nothing.

Sizes overall: backends with no cuDNN consumer shed ~57 MB each (seven
instances on the fleet measured, plus longcat's ~60 MB), while the ones that
genuinely use cuDNN grow from ~57 MB to ~576 MB, because the five missing
sublibraries are ~517 MB, dominated by libcudnn_engines_precompiled. Net on
that fleet is an increase of roughly 570 MB. That growth is the bug being paid
off, not a regression: those backends only work today by silently borrowing the
missing five from the runtime image. Whether the engines set can be trimmed is
an open question, not addressed here.

So bundle per backend, by what that backend actually needs:

  - venv has a complete pip cuDNN -> bundle nothing; $ORIGIN resolves the pip
    set, which is the one its torch was built against            (longcat-video)
  - venv has no pip cuDNN         -> bundle the complete family. Stays
    conservative rather than detecting consumers: for a Python backend they sit
    inside the venv (torch, ctranslate2, onnxruntime) where the sweep does not
    look                                                                 (vllm)
  - no venv, nothing references cuDNN -> bundle nothing    (llama-cpp, whisper,
                                     rfdetr-cpp, sam3-cpp, stablediffusion-ggml)
  - no venv, something references it   -> bundle the complete family
                                                  (face-detect, voice-detect)

The no-venv case needs no new machinery. Go backends stage their own shared
object into package/lib, which IS the target dir, so sweep_transitive_deps
already pulls the dispatcher when it is a genuine dependency - that is exactly
how libcudnn_graph reached longcat. cuDNN simply comes off the force-copy list,
and complete_cudnn_family fills in the seven dlopen'd sublibraries around
whatever the sweep found. Detection is a string scan rather than ldd, so a
consumer that only dlopen()s cuDNN is seen too; over-matching costs an unused
library, under-matching costs a backend that cannot load.

Keeping bundled and pip versions in agreement instead is not viable: nothing
here pins nvidia-cudnn (zero occurrences), torch is unpinned for l4t13 except
longcat-video, and the fleet already runs five concurrent cuDNN versions -
9.19.0.56, 9.20.0.48, 9.22.0, 9.23.2, 9.24.0.

verify_cudnn_bundle asserts the end state: exactly one complete cuDNN visible to
whoever needs one - never both, never partial, and never zero for a backend that
references it. Zero is correct and common otherwise. It deliberately does not
accept the build image's system cuDNN as completing a partial bundle, which is
the shape that had been shipping silently; the build image is not the runtime
image. A version check alone would have missed longcat too, whose four bundled
libs were all 9.24.0 and mutually consistent.

Match per family for the other components for the same dlopen reason: TensorRT
(libnvinfer_plugin, libnvinfer_builder_resource), cuBLAS, cuFFT, cuSPARSE,
cuSOLVER, nvRTC. Exclusions bind inside copy_lib so they cover the sweep.

The packaging scripts' shell tests ran nowhere in CI. Add make
test-build-scripts and a lint workflow job so they gate every PR.

Fixes #10905


Assisted-by: Claude:claude-opus-4-8 golangci-lint shellcheck

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-19 07:48:51 +00:00
localai-org-maint-bot
71e98c13a3 fix(vllm): generate protobuf 6 compatible stubs (#10944)
Pin vLLM protogen to grpcio-tools 1.78.0 so its generated code remains importable by protobuf 6.33.x, and remove stale generated artifacts before regeneration.

Closes #10940

Assisted-by: Codex:gpt-5 [Codex]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-07-19 08:56:57 +02:00
mudler's LocalAI [bot]
10211948b5 chore(model gallery): 🤖 add 1 new models via gallery agent (#10942)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-19 08:45:41 +02:00
mudler's LocalAI [bot]
139470cca0 chore: ⬆️ Update ggml-org/llama.cpp to 571d0d540df04f25298d0e159e520d9fc62ed121 (#10935)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-19 08:45:08 +02:00
mudler's LocalAI [bot]
078614c701 chore: ⬆️ Update CrispStrobe/CrispASR to 1e6f3ad962dc46d86422c3baa4f3c1110d037e4d (#10934)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-19 08:44:57 +02:00
mudler's LocalAI [bot]
c1efdbeb9e chore: ⬆️ Update leejet/stable-diffusion.cpp to ea4e566ccffa10f853ecc3f29e74b1820bc91beb (#10936)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-19 08:44:45 +02:00
mudler's LocalAI [bot]
7f72dc3412 chore: ⬆️ Update PrismML-Eng/llama.cpp to 9fcaed763ccda38ea81068ad9d7f991aaddca451 (#10937)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-19 08:44:27 +02:00
mudler's LocalAI [bot]
81c407bc40 chore(model-gallery): ⬆️ update checksum (#10938)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-19 08:44:02 +02:00
Richard Palethorpe
9c43b2da8f fix(model): make backend shutdown model-scoped (#10865)
Avoid holding the global loader lock across backend lifecycle waits and propagate forced shutdown through distributed workers. Track parallel requests with in-flight counters and reserve worker ports until process termination.

Add focused race tests and an authoritative FizzBee lifecycle model with a fail-closed conformance target.

Assisted-by: Codex:GPT-5 [FizzBee] [Ginkgo]

Signed-off-by: Richard Palethorpe <io@richiejp.com>
2026-07-19 08:43:17 +02:00
mudler's LocalAI [bot]
27955e0a33 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 9d07d8681ece159a89fb4e16a1f9c9f3a5fac20f (#10933)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-19 00:43:49 +02:00
dependabot[bot]
036eccc32d chore(deps): bump actions/setup-node from 6 to 7 (#10915)
Bumps [actions/setup-node](https://github.com/actions/setup-node) from 6 to 7.
- [Release notes](https://github.com/actions/setup-node/releases)
- [Commits](https://github.com/actions/setup-node/compare/v6...v7)

---
updated-dependencies:
- dependency-name: actions/setup-node
  dependency-version: '7'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-18 22:36:37 +02:00
dependabot[bot]
a15b23b775 chore(deps): bump torch from 2.12.1+xpu to 2.13.0+xpu in /backend/python/common/template (#10917)
chore(deps): bump torch in /backend/python/common/template

Bumps torch from 2.12.1+xpu to 2.13.0+xpu.

---
updated-dependencies:
- dependency-name: torch
  dependency-version: 2.13.0+xpu
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-18 22:36:16 +02:00
dependabot[bot]
a4a14c6263 chore(deps): bump grpcio from 1.80.0 to 1.82.1 in /backend/python/common/template (#10918)
chore(deps): bump grpcio in /backend/python/common/template

Bumps [grpcio](https://github.com/grpc/grpc) from 1.80.0 to 1.82.1.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.80.0...v1.82.1)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.82.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-18 22:35:55 +02:00
dependabot[bot]
00cbfc369b chore(deps): bump grpcio from 1.80.0 to 1.82.1 in /backend/python/rerankers (#10921)
chore(deps): bump grpcio in /backend/python/rerankers

Bumps [grpcio](https://github.com/grpc/grpc) from 1.80.0 to 1.82.1.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.80.0...v1.82.1)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.82.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-18 22:35:34 +02:00
dependabot[bot]
cee6780ea7 chore(deps): bump grpcio from 1.80.0 to 1.82.1 in /backend/python/coqui (#10922)
Bumps [grpcio](https://github.com/grpc/grpc) from 1.80.0 to 1.82.1.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.80.0...v1.82.1)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.82.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-18 22:35:17 +02:00
dependabot[bot]
79113c7f90 chore(deps): update transformers requirement from >=5.9.0 to >=5.14.1 in /backend/python/transformers (#10926)
chore(deps): update transformers requirement

Updates the requirements on [transformers](https://github.com/huggingface/transformers) to permit the latest version.
- [Release notes](https://github.com/huggingface/transformers/releases)
- [Commits](https://github.com/huggingface/transformers/compare/v5.9.0...v5.14.1)

---
updated-dependencies:
- dependency-name: transformers
  dependency-version: 5.14.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-18 22:35:01 +02:00
dependabot[bot]
e7520af5d7 chore(deps): bump grpcio from 1.81.0 to 1.82.1 in /backend/python/transformers (#10925)
chore(deps): bump grpcio in /backend/python/transformers

Bumps [grpcio](https://github.com/grpc/grpc) from 1.81.0 to 1.82.1.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.81.0...v1.82.1)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.82.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-18 22:34:29 +02:00
dependabot[bot]
a9456bbce9 chore(deps): bump sentence-transformers from 5.5.1 to 5.6.0 in /backend/python/transformers (#10927)
chore(deps): bump sentence-transformers in /backend/python/transformers

Bumps [sentence-transformers](https://github.com/huggingface/sentence-transformers) from 5.5.1 to 5.6.0.
- [Release notes](https://github.com/huggingface/sentence-transformers/releases)
- [Commits](https://github.com/huggingface/sentence-transformers/compare/v5.5.1...v5.6.0)

---
updated-dependencies:
- dependency-name: sentence-transformers
  dependency-version: 5.6.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-18 22:33:46 +02:00
dependabot[bot]
24c16c9bb5 chore(deps): bump vllm from 0.25.0 to 0.25.1 in /backend/python/vllm (#10929)
Bumps [vllm](https://github.com/vllm-project/vllm) from 0.25.0 to 0.25.1.
- [Release notes](https://github.com/vllm-project/vllm/releases)
- [Changelog](https://github.com/vllm-project/vllm/blob/main/RELEASE.md)
- [Commits](https://github.com/vllm-project/vllm/compare/v0.25.0...v0.25.1)

---
updated-dependencies:
- dependency-name: vllm
  dependency-version: 0.25.1
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-18 21:37:41 +02:00
localai-org-maint-bot
0389495388 fix(webui): use relative asset base so fonts and lazy chunks honor X-Forwarded-Prefix (#10889) (#10904)
The Vite build emitted path-absolute asset URLs (base: '/'). index.html
entry scripts and the favicon were rewritten to include the reverse-proxy
prefix in serveIndex, but two reference kinds are not in index.html and so
bypassed that rewrite:

  - CSS `url()` font references (e.g. Font Awesome .woff2), which the browser
    resolves relative to the stylesheet and which `<base href>` never affects
  - lazily-imported route chunks, whose preload base came from the absolute
    Vite base

Under a subpath mount (X-Forwarded-Prefix: /llm/) both were fetched from the
origin root, 404ing — missing-glyph "tofu" icons and broken lazy-loaded pages.

Switch Vite to a relative base ('./') so every generated URL resolves against
the file that references it: CSS fonts and route chunks now load from
`/llm/assets/...`, and index.html's now-relative entry refs resolve via the
`<base href>` serveIndex already injects on every response. Root deployments
are unaffected. The existing path-absolute rewrite in app.go still covers the
public `/favicon.svg`.


Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-18 08:38:26 +02:00
dependabot[bot]
2f011094d9 chore(deps): bump torch from 2.8.0 to 2.12.1+xpu in /backend/python/common/template in the pip group across 1 directory (#10911)
chore(deps): bump torch

Bumps the pip group with 1 update in the /backend/python/common/template directory: torch.


Updates `torch` from 2.8.0 to 2.12.1+xpu

---
updated-dependencies:
- dependency-name: torch
  dependency-version: 2.12.1+xpu
  dependency-type: direct:production
  dependency-group: pip
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-18 08:37:03 +02:00
localai-org-maint-bot
bc653c9b09 ci(dependabot): ignore torch/transformers for diffusers to fix Jetson-index auth failure (#10913)
The weekly "Dependabot Updates" pip job for /backend/python/diffusers has been
failing with `private_source_authentication_failure` against the Jetson pip
index (https://pypi.jetson-ai-lab.io/jp6/cu129/), referenced by that backend's
requirements-l4t12.txt. diffusers is the only dependabot-configured pip
directory that pulls from that private index, so it is the only update job that
fails; the other backends update cleanly.

torch and transformers are deliberately pinned in this backend for
reproducibility (see backend/python/diffusers/requirements-*.txt and #9979), so
we do not want dependabot bumping them anyway. Ignoring both dependencies for
this directory stops dependabot from resolving them against the unreachable
Jetson index and keeps the weekly update job green, without removing update
coverage for the rest of the backend's dependencies.


Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-18 08:36:23 +02:00
mudler's LocalAI [bot]
f9a2d9be32 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260717051959 (#10903)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-18 08:35:41 +02:00
mudler's LocalAI [bot]
f40e07d72e chore: ⬆️ Update CrispStrobe/CrispASR to c96281d6d409a7f97edbce62c12a6dd2f4da6a92 (#10900)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-18 08:35:28 +02:00
mudler's LocalAI [bot]
911fb754a6 chore: ⬆️ Update ggml-org/llama.cpp to 6bdd77f13cf11b264b4231d320afc404f48d576e (#10898)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-18 08:35:15 +02:00
mudler's LocalAI [bot]
2dade4a9f9 fix(model-artifacts): gate inferred artifact materialization by backend (#10910)
The managed-artifact materializer stages a HuggingFace snapshot into a
directory (.artifacts/huggingface/<key>/snapshot/). That is the right load
target for directory-consuming backends (transformers, vLLM, diffusers, ...),
but PrimaryArtifactSpec inferred a managed artifact from ANY HuggingFace-shaped
model reference regardless of backend. A single-file backend such as llama.cpp
or whisper was therefore handed the snapshot directory instead of the weight
file and failed to load it.

The /import-model importer already guards this with a backend allow-list
(managedArtifactBackends), but the loader-side inference did not. Move the
allow-list into core/config as IsManagedArtifactBackend and apply it in
PrimaryArtifactSpec: only directory-consuming backends may have an artifact
inferred from a bare reference; every other backend stays on the legacy
download-to-file path. An explicit artifacts: block still bypasses the gate,
where single-file snapshot resolution handles the load path.

The importer now shares the same predicate, so both paths agree on which
backends auto-materialize.

Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 23:26:23 +00:00
mudler's LocalAI [bot]
c0a20d6ab1 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 73a88bf6323f7b9dfff8dde76b4fffcc2fd618ce (#10896)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-18 01:05:26 +02:00
mudler's LocalAI [bot]
78775c77d8 chore: ⬆️ Update mudler/parakeet.cpp to 1da853421de9710cbe894a0110711de5a0516486 (#10899)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-18 01:05:14 +02:00
mudler's LocalAI [bot]
525af1df1b chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to 95b4840ad3722b0b67acb945cd57682aae1ac9ca (#10902)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-18 00:46:10 +02:00
localai-org-maint-bot
279f5b8a93 fix(model-artifacts): load single-file HF snapshots from the file, not the directory (#10909)
fix(model-artifacts): load single-file HF snapshots from the file, not the dir

The managed Hugging Face artifact materializer (#10825) always pointed
backends at the snapshot *directory*
(.artifacts/huggingface/<key>/snapshot). For a single-file model
reference such as huggingface://nomic-ai/nomic-embed-text-v1.5-GGUF/nomic-embed-text-v1.5.f16.gguf,
the GGUF lives *inside* that directory, so llama.cpp was handed a
directory and failed with "gguf_init_from_reader: failed to read magic".
This has kept the tests-aio job red on master since the feature merged
(the embeddings e2e tests could not load text-embedding-ada-002).

Record the single file of a one-file snapshot as Resolved.PrimaryFile and
have ModelFileName() resolve to snapshot/<PrimaryFile> when it is set.
Multi-file snapshots (e.g. transformers repos consumed as a directory)
keep pointing at the snapshot directory. PrimaryFile is derived from the
resolved contents and is deliberately excluded from the artifact cache
key. estimateModelSizeBytes now derives the snapshot directory from the
cache key instead of ModelFileName(), so its manifest lookup is unaffected
by the file-vs-directory resolution.


Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 22:42:50 +00:00
mudler's LocalAI [bot]
9edb08ea94 chore: ⬆️ Update ikawrakow/ik_llama.cpp to fbcc743c70391e63fba74a16740f8157b469feeb (#10897)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-18 00:13:02 +02:00
mudler's LocalAI [bot]
2ad3b5088b chore: ⬆️ Update PrismML-Eng/llama.cpp to 79697f23a2c8f3aa2ccb2fd7406095a8dbfbb454 (#10901)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-18 00:02:20 +02:00
mudler's LocalAI [bot]
a89d780707 fix(gallery): keep multi-file HF install progress proportional during verify (#10908)
The artifact progress bridge mapped every PhaseVerifying event to a flat
95%. The materializer emits PhaseVerifying once per file (from each file's
AfterDownload hook) and downloads run sequentially, so the first small file
to finish pinned the bar at 95% - and, because progress is monotonic, it
stayed at 95% for the entire remaining download (e.g. a 70GB checkpoint
reporting 95% at 410MB / 69.7GB).

Track per-file verify proportionally to the running aggregate bytes, the
same way downloading does. CurrentBytes already reflects "completed files +
this file", so the percentage advances honestly. The flat 95%/99% is now
reserved for the genuinely once-per-install Committing/Persisting phases.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-18 00:01:24 +02:00
mudler's LocalAI [bot]
55e2726958 chore(model-gallery): ⬆️ update checksum (#10906)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-17 23:47:43 +02:00
localai-org-maint-bot
4be6e22b5f feat(webui): surface user, client IP and user agent in API traces (#10886, #10887) (#10907)
The Operate → Traces "API Traces" panel already recorded who made each
request (user_id/user_name) but never showed it, and did not capture the
caller's network identity at all. Operators asked to see the requesting
user (#10886) and the client IP + user agent (#10887) so a trace can be
attributed to who/what issued it.

Backend: add ClientIP and UserAgent to APIExchange and populate them from
echo's c.RealIP() (honours X-Forwarded-For / X-Real-IP behind a trusted
proxy) and the request's User-Agent header. Both are omitempty and the
/api/traces swagger response is map[string]any, so this is additive.

UI: add a sortable "User" column to the API traces table and a metadata
block (User / Client IP / User Agent) at the top of the expanded row
detail. Fields render only when present, so older buffered traces and
unauthenticated/local requests degrade cleanly.

Adds an e2e spec covering the new column value and the expanded metadata.


Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 23:47:31 +02:00
localai-org-maint-bot
bf484c5181 feat(webui): show date alongside time in the Traces view (#10888) (#10905)
The Operate -> Traces table rendered the request time with the time of day
only, so entries that span more than one day were ambiguous. Add a
formatDateTime helper (localized date + existing time-with-millis) and use it
for the Traces "Time" column, keeping the cell on a single line. The shared
formatTimestamp used by the log views is unchanged.


Assisted-by: Claude:opus-4.8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 23:47:10 +02:00
mudler's LocalAI [bot]
40d35c0385 docs: onboarding overhaul, dedup, and error docs (#7711) (#10895)
* docs: fix CPU image tag (latest, not latest-cpu)

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: use canonical localai/localai registry in models guide

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: replace dead llama-stable backend with llama-cpp

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: correct mitm-proxy intercept config and redaction tier

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: fix text-to-audio endpoint and broken notice block

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: fix VAD example, stale FAQ, broken link, CLI list, whats-new dump

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: render advanced/reference section indexes (consolidate _index)

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: remove duplicate getting-started build/kubernetes pages

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: fold container image reference into installation/containers

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: remove stale advanced fine-tuning page (superseded by features/fine-tuning)

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: fold distribution/longcat/sound pages into their parents

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: make getting-started index accurate and complete

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: carry one concrete model through the getting-started path

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: add end-to-end 'build your first agent' walkthrough

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: add runtime errors reference; consolidate troubleshooting from FAQ

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: add agent actions catalog

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: agent-scoped MCP, skills walkthrough, agentic disambiguation

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: add concrete gallery install lines to media feature pages

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: merge installation into getting-started (URLs preserved via aliases)

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: add Operations section; move operator pages and P2P API reference

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: journey-ordered top nav and grouped feature sections

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: add docs-with-code process gate (PR template + agent instructions)

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: remove em/en dashes from documentation prose

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 22:08:20 +02:00
futurehua
d3ea65a112 refactor: replace Split in loops with more efficient SplitSeq (#10879)
Signed-off-by: futurehua <futurehua@outlook.com>
2026-07-17 22:07:16 +02:00
mudler's LocalAI [bot]
4f592c8734 chore(model-gallery): ⬆️ update checksum (#10891)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-17 20:06:32 +02:00
walcz-de
6ccb1130d8 fix(agent-ui): reset streamed text at generation boundaries in agent chat (#10664)
One agent turn runs several internal LLM generations (tool selection,
reasoning, final answer) that all emit stream_event deltas over the same
per-agent SSE channel. The chat page accumulated every 'content' delta
into a single live bubble and ignored the 'done' boundary events, so the
internal generations' text (e.g. the English tool-selection rationale)
merged with — and visually corrupted — the streamed final answer.

Reset the accumulated content/reasoning on 'done': each generation gets
a clean live bubble, and the authoritative full answer still arrives via
the final json_message event as before.

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-07-17 15:26:50 +02:00
LocalAI [bot]
3f8806b0b2 chore(model gallery): 🤖 add 1 new models via gallery agent (#10881)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-17 15:22:11 +02:00
LocalAI [bot]
14c7c04feb chore(model gallery): 🤖 add 1 new models via gallery agent (#10874)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-17 12:58:47 +02:00
LocalAI [bot]
ec933b837d chore(moss-transcribe-cpp): bump pin to CUDA K-quant embed fix (#10862) (#10878)
chore(moss-transcribe-cpp): bump pin to CUDA K-quant embed fix

Bumps the moss-transcribe.cpp pin to 190a569c, which merges the
host-side embed-lookup fallback for K-quant token_embd tensors
(localai-org/moss-transcribe.cpp#2).

Before this, running a q5_K/q4_K/q6_K moss-transcribe GGUF on CUDA
(or any non-CPU backend) aborted in getrows.cu with
"unsupported src0 type: q5_K" because ggml's GET_ROWS op has no
K-quant implementation on GPU, killing the backend process on the
first request. The engine now dequantizes the needed rows on the
host when the backend cannot run GET_ROWS for the tensor type,
producing bit-identical results.

Fixes #10862


Assisted-by: Claude:claude-opus-4-8 [Bash] [Edit]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 12:47:13 +02:00
LocalAI [bot]
cc26083423 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260716042225 (#10849)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-17 12:46:56 +02:00
Nicholas Ciechanowski
8aa8e0fac0 fix(distributed): setup script (#10551)
Assisted-by: OpenCode:GPT-5.5 [Read] [Edit]

Signed-off-by: Nicholas Ciechanowski <nicholas@ciech.anow.ski>
2026-07-17 10:14:19 +02:00
LocalAI [bot]
e9056399a7 feat(gallery): add MOSS-TTS-Local v1.5 models for the moss-tts-cpp backend (#10877)
Add the q8_0 (default) and f16 gallery entries for the moss-tts-cpp backend, each
pulling the MOSS-TTS-Local v1.5 GGUF plus the MOSS-Audio-Tokenizer-v2 codec and
the text tokenizer from mudler/MOSS-TTS-Local-Transformer-v1.5-GGUF. The backend
auto-discovers the codec and tokenizer siblings; output is 48 kHz stereo with
reference-audio voice cloning.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 10:02:01 +02:00
LocalAI [bot]
3bb0d1cb49 feat(backend): add moss-tts-cpp text-to-speech backend (#10860)
* feat(backend): add moss-tts-cpp text-to-speech backend

Add a Go + purego backend wrapping the moss-tts.cpp ggml port of the OpenMOSS
MOSS-TTS-Local v1.5 text-to-speech model (GPT-J local transformer decoded through
MOSS-Audio-Tokenizer-v2), producing 48 kHz stereo audio with optional
reference-audio voice cloning. Mirrors the qwen3-tts-cpp backend: dlopen the
static-ggml shared library, bind the moss-tts.cpp C-API via purego, and serve
the gRPC TTS method. A thin C shim holds the pipeline handle and copies engine
PCM into a Go-freeable buffer.

Wires the CI registration: backend-matrix.yml (CPU, CUDA 12/13, Intel SYCL
f16/f32, Vulkan, ROCm, NVIDIA L4T, plus Darwin metal), backend/index.yaml metas
and image entries pointing at mudler/MOSS-TTS-Local-Transformer-v1.5-GGUF, the
root Makefile build targets, and the changed-backends.js path mapping.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: list the moss-tts-cpp backend among the LocalAI-maintained engines

Add moss-tts.cpp to the README "Backends built by us" table, the
Text-to-Speech compatibility table, and the reference-audio voice-cloning
backend list, so the new backend is documented alongside its peers.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(moss-tts-cpp): pin moss-tts.cpp to the squashed single-commit release

moss-tts.cpp history was collapsed to a single commit; repoint MOSSTTS_CPP_VERSION
to ee722b8e9205ee9b1b1c398a4e87e4e393e9be41.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(moss-tts-cpp): add the moss-tts-cpp-development gallery meta

The gallery had the -development image entries but no matching -development
meta anchor (as locate-anything-cpp and depth-anything-cpp have), so the master
build was not installable as a gallery backend. Add moss-tts-cpp-development
mirroring the production meta with the -development capability image names.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 09:26:12 +02:00
LocalAI [bot]
0bd7a29f31 feat(gallery): add Gemma 4 llama.cpp MTP variants; fix gemmable-4-12b-mtp (#10876)
Google shipped the Gemma 4 MTP drafter heads and llama.cpp merged native
support in ggml-org/llama.cpp#23398. LocalAI's pinned llama.cpp already
carries it, and the config plumbing (draft_model + core/config/mtp.go)
was built for exactly this path, but no official Gemma 4 gallery entry
wired it up.

Add llama.cpp draft-mtp speculative-decoding variants for the dense
sizes, sourced from the unsloth QAT GGUF repos (target UD-Q4_K_XL +
mtp-*.gguf drafter + BF16 mmproj):

  - gemma-4-e2b-it-qat-mtp
  - gemma-4-e4b-it-qat-mtp
  - gemma-4-12b-it-qat-mtp
  - gemma-4-31b-it-qat-mtp

These replace the previously commented-out attempts, which were disabled
because the Janvitos/boxwrench drafter GGUFs declared the architecture as
`gemma4_assistant` (underscore) and failed to load on stock llama.cpp.
The unsloth drafters use the upstream `gemma4-assistant` (hyphen) spelling
that mtp.go's isDraftOnlyAssistantArch expects, so they load without any
backend patch. The 26B-A4B MoE is intentionally omitted (the upstream PR
reports no meaningful MTP speedup for it).

Also fix gemmable-4-12b-mtp: it loaded the draft-only `-mtp` GGUF as the
main model with no draft_model set, which cannot run standalone. It now
loads the target as the model, wires the drafter via draft_model, enables
spec_type:draft-mtp, and downloads both files.

All sha256 pins were taken from the HuggingFace API lfs.oid (reliable
content hash even for Xet-backed repos).


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 09:05:44 +02:00
Tai An
45b8047736 fix(p2p): serialize access to p2pCtx/p2pCancel (#10839) (#10861)
StopP2P() read and wrote a.p2pCtx/a.p2pCancel without holding
a.p2pMutex, and StartP2P() reassigned both fields with no lock at
all -- including when RestartP2P() calls it from a background
goroutine after releasing the mutex. Both paths are reachable from
POST /api/settings (empty p2p_token -> StopP2P, non-empty ->
RestartP2P), so concurrent requests race on the same fields.

Take a.p2pMutex in StopP2P and around the field publication in
StartP2P, factor the shared teardown into stopP2PLocked() so
RestartP2P reuses it, and route the goroutine error path through
StopP2P instead of touching a.p2pCancel unlocked.

Signed-off-by: Anai-Guo <antai12232931@anaiguo.com>
Co-authored-by: Anai-Guo <antai12232931@anaiguo.com>
2026-07-17 09:02:05 +02:00
LocalAI [bot]
fd0d1b946d chore: ⬆️ Update PrismML-Eng/llama.cpp to 62061f91088281e65071cc38c5f69ee95c39f14e (#10869)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-17 09:00:44 +02:00
LocalAI [bot]
6dfda9c4b6 chore: ⬆️ Update ggml-org/llama.cpp to e8f19cc0ad70a243c8012bf17b4be601abfc8ea2 (#10870)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-17 09:00:30 +02:00
LocalAI [bot]
dffcbd7e5d chore(model-gallery): ⬆️ update checksum (#10871)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-17 09:00:17 +02:00
LocalAI [bot]
7c542fb979 chore: ⬆️ Update leejet/stable-diffusion.cpp to b2906939774dc73453467215c80390404d0a2701 (#10872)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-17 09:00:02 +02:00
LocalAI [bot]
cbf232e5fe chore: ⬆️ Update CrispStrobe/CrispASR to a38cb89f7b9a743db2e8e50869fa646f91dc7f08 (#10873)
* ⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(crispasr): rewrite c2pa-audio submodule path for subproject builds

CrispASR a38cb89 adds a crispasr_c2pa_native static library whose sources
live in the new third_party/c2pa-audio git submodule, located via
CMAKE_SOURCE_DIR in src/CMakeLists.txt. That variable assumes CrispASR is
the top-level CMake project; LocalAI embeds it via add_subdirectory, so
the path resolved to backend/go/crispasr/third_party/c2pa-audio and every
build variant failed at CMake generate with 'Cannot find source file:
c2pa_native.cpp'.

Extend the existing talk-llama sed workaround to also rewrite the
c2pa-audio reference to PROJECT_SOURCE_DIR, which is correct both
standalone and as a subproject. The submodule itself is already checked
out by the recursive submodule init. Verified locally: the exact CI error
reproduces with CMAKE_SOURCE_DIR, and with the rewrite CMake configure,
crispasr_c2pa_native, and crispasr-lib all build cleanly on a CPU-only
fallback configuration.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-4-8

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 08:59:48 +02:00
LocalAI [bot]
1f53dff436 fix(turboquant,bonsai): do not apply vendored llama.cpp patches to fork trees (#10866)
The turboquant and bonsai backends copy backend/cpp/llama-cpp/ wholesale
into their build directories and reuse its Makefile/prepare.sh against
their own llama.cpp forks. When PR #10837 added
backend/cpp/llama-cpp/patches/0001-add-minimax-m3-support.patch, the
copied patches/ directory was mis-applied to the fork checkouts: the
fork trees diverge from upstream, hunks rejected, and because the
patch-apply loop in prepare.sh ran before set -e took effect the build
kept going and died much later with a confusing compile error
("'LLM_ARCH_MINIMAX_M3' was not declared in this scope"). This broke
tests-turboquant-grpc on that PR.

Two hardening changes:

- turboquant/bonsai Makefiles: delete the copied patches/ directory
  right after the cp -rf of backend/cpp/llama-cpp/. Patches vendored
  for upstream llama.cpp must never be applied to the forks; each fork
  carries its own patch series under backend/cpp/<backend>/patches/,
  applied by its apply-patches.sh.

- llama-cpp prepare.sh: run the patch-apply loop under set -e so a
  rejecting patch fails fast and loudly at apply time instead of
  surfacing as a downstream compile error. A missing or empty patches/
  directory remains a no-op success, so all existing callers (the
  llama-cpp Makefile targets and the turboquant/bonsai copies) are
  unaffected when no patches ship.

Exposed by PR #10837.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 00:28:10 +02:00
LocalAI [bot]
c1a891662c refactor(settings): single declarative registry for runtime settings (fixes the #10845 bug class) (#10864)
* feat(settings): add declarative runtime-settings field registry

One fieldSpec row per RuntimeSettings field, with a reflection
completeness spec so a field added without a registry row is a red
test instead of a silently-dropped setting (the #10845 bug class).

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-fable-5

* refactor(settings): drive ToRuntimeSettings/ApplyRuntimeSettings from the field registry

Behavior-preserving: ~350 hand-written per-field lines become two loops
over runtimeSettingsFields, gated by a To->Apply->To round-trip spec.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-fable-5

* feat(settings): baseline-driven startup merge for persisted runtime settings

ApplyRuntimeSettingsAtStartup compares the live config against
DefaultRuntimeBaseline (option-less-run defaults incl. kong-injected
flag defaults) instead of per-field == 0 guards. Fixes persisted
lru_eviction_max_retries, tracing_max_items, agent_job_retention_days,
memory_reclaimer_threshold, galleries and autoload flags being
silently ignored at boot.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-fable-5

* fix(settings): registry-driven startup merge, applied before consumers

loadRuntimeSettingsFromFile becomes a thin wrapper over
ApplyRuntimeSettingsAtStartup and runs at the top of New(), before
model configs capture app-level defaults. WithThreads stops eagerly
resolving 0 so a persisted thread count survives restart while
LOCALAI_THREADS still wins (#10845); the physical-core fallback moves
after the merge.

Also: run.go now injects the memory-reclaimer threshold unconditionally
so the option-less boot matches DefaultRuntimeBaseline and a UI-saved
threshold survives restart.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-fable-5

* refactor(settings): file watcher delegates to the registry merge; shared API-key merge

Manual edits to runtime_settings.json now behave like a boot-time load
(env still wins) instead of the inverted diverged-from-startup guard
that ignored most manual edits. MergeAPIKeys dedups env keys in one
place for the endpoint and the watcher.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-fable-5

* docs(settings): document unified runtime-settings precedence

Document the single env/CLI > runtime_settings.json > defaults rule,
applied identically at boot, on POST /api/settings, and on manual file
edits, plus the two known limitations (default-valued env vars are
indistinguishable from unset; API-changed fields hot-apply on the next
restart only). Also add a completion debug log when the watcher applies
runtime_settings.json.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-fable-5

* test(settings): reset the global VRAM cap leaked by the round-trip spec

The round-trip spec applies vram_budget=12GiB, whose post-loop hook
installs a process-global default cap; without a reset every spec
ordered after it runs under that phantom budget. Also drop a stale
enumeration in the ApplyRuntimeSettings doc comment.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-fable-5

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-16 22:39:59 +02:00
pos-ei-don
06b4a29387 docs(config): document grpc.attempts timing + tuning guidance (#10868)
The gRPC configuration table only listed the two fields with a one-line
description each, without defaults, without explaining what the total
load window looks like, and without hinting when a user should adjust
them. In practice the default 20 attempts x 2 s = 40 s window is way
too tight for large NVFP4 / FP8 models on slow storage or first-run
CUDA-graph capture, and the resulting kill (exitCode=120, 'context
canceled') looks like a backend crash even though the backend is still
making legitimate forward progress.

Extend the section with:
- Defaults column (20 and 2) added to the table
- Prose explaining that these govern the readiness handshake between
  LocalAI and a freshly spawned backend (Health polling loop)
- Total-load-window formula
- Concrete failure signature so users can recognize a timeout-kill
  vs. a real backend crash
- Example configuration for a ~10 min cold-load window (grpc.attempts
  140, attempts_sleep_time 5), with a note that inference-timeouts and
  the watchdog are unaffected.
2026-07-16 22:18:47 +02:00
pos-ei-don
e62221b020 fix(sglang): implement Status RPC to unblock backend-monitor polling (#10867)
The sglang Python backend inherits the default Status RPC from
backend_pb2_grpc.BackendServicer, which raises NotImplementedError.
LocalAI's backend-monitor polls /backend.Backend/Status periodically on
every registered backend; when the call fails, /backend/monitor returns
HTTP 500 and downstream inference requests to the sglang backend are
blocked even though the model is loaded and answering directly via the
gRPC endpoint.

Add a minimal Status shim that mirrors the existing Health method and
returns StatusResponse{state=READY} unconditionally. This unblocks the
monitor path; a state-aware follow-up (UNINITIALIZED during load, BUSY
under active inference) is left for a subsequent change.

Reproduced on DGX Spark (GB10, arm64-l4t-cuda-13 image) with the sglang
v0.5.15 backend and Qwen3-Coder-Next-NVFP4-GB10; verified locally that
patching the shim in place immediately restores /backend/monitor and
inference across the sglang slot.
2026-07-16 22:18:12 +02:00
Tai An
dc2cc4da43 fix(audio-transform): serialize WebSocket writes to avoid concurrent-write panic (#10857)
* fix(audio-transform): serialize WebSocket writes to avoid concurrent-write panic

AudioTransformStreamEndpoint writes to the same Gorilla WebSocket connection
from two goroutines: the backend-forwarding goroutine emits binary PCM frames
(and can call sendWSError on a backend recv error), while the read loop calls
sendWSError for malformed mid-stream JSON or a backend send failure. Gorilla
WebSocket permits only one concurrent writer, so these writers race and can
panic with "concurrent write to websocket connection", resetting the client
session; a -race build reports the data race directly.

Wrap the connection in a lockedConn that serializes WriteMessage behind a
mutex, mirroring the existing lockedConn used by the openresponses WebSocket
endpoint. Reads stay on the single read loop, so only writes need the lock.

Fixes #10844

Signed-off-by: Tai An <antai12232931@outlook.com>

* chore: empty commit to re-trigger checks

Signed-off-by: Anai-Guo <antai12232931@anaiguo.com>

---------

Signed-off-by: Tai An <antai12232931@outlook.com>
Signed-off-by: Anai-Guo <antai12232931@anaiguo.com>
Co-authored-by: Anai-Guo <antai12232931@anaiguo.com>
2026-07-16 16:25:31 +01:00
Nandana Dileep
ab7b58fc85 fix(watchdog): force-kill stuck-busy backends instead of deadlocking the loader (#10578)
When the watchdog's busy-killer decides a backend has been busy past the
busy timeout, it shuts it down via ModelLoader.ShutdownModel -> deleteProcess,
which grabs ml.mu and then waits for IsBusy() to clear BEFORE stopping the
process. But a backend that exceeds the busy timeout is, by definition,
stuck on an in-flight gRPC call, so the graceful wait never returns, ml.mu
is held forever, and every other ml.Load blocks — including the shared
opus backend load at the start of every realtime (WebRTC) session. New
realtime connections then hang at "Connected, waiting for session..."
whenever the watchdog is enabled, while logs repeatedly print the
watchdog's busy / "active connection" line.

Fix: add a force shutdown path (ShutdownModelForce / deleteProcess(s,
force=true)) that stops the process FIRST — dropping the stuck call's
gRPC connection and unblocking it — instead of waiting on it. Route the
watchdog's busy-killer and busy LRU / group / memory evictions through
the force path; keep the graceful wait for idle and user-initulated
unloads. Graceful/unforced kills are unchanged.

Regression test: the watchdog busy-killer uses ShutdownModelForce.

Fixes #10391


Assisted-by: opencode:glm-5.2 [opencode]

Signed-off-by: Nandana Dileep <110280757+nandanadileep@users.noreply.github.com>
2026-07-16 12:36:23 +00:00
LocalAI [bot]
bcdb8debfe chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to 9e11ce41b90a2238ca1ec09e0c71fcc913544f2a (#10850)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-16 10:10:55 +02:00
LocalAI [bot]
bbe018c1a0 feat(bonsai): PrismML llama.cpp fork backend + Bonsai/Ternary-Bonsai gallery models (#10834)
feat(bonsai): add PrismML llama.cpp fork backend + Bonsai gallery models

Adds a new `bonsai` backend that runs the PrismML fork of llama.cpp
(github.com/PrismML-Eng/llama.cpp, `prism` branch), which ships the Q1_0
(1-bit) and Q2_0 (ternary / 1.58-bit) weight-quantization kernels used by the
Bonsai and Ternary-Bonsai models. Stock llama.cpp cannot decode these quants.

Modeled on the turboquant backend: reuses backend/cpp/llama-cpp/grpc-server.cpp
against the fork's libllama via a thin wrapper Makefile, so the sub-2-bit models
are served with the same OpenAI-compatible API. No grpc-server allow-list patch
is needed (bonsai adds weight quants, transparent to the server, not KV-cache
types), and the reused server compiles cleanly against the fork with no skew
patches (validated locally via a CPU docker build; patches/ is present but empty
for any future re-pin skew).

Backend wiring: backend/cpp/bonsai/, .docker/bonsai-compile.sh,
backend/Dockerfile.bonsai, top-level Makefile targets, backend-matrix.yml build
rows (CPU, CUDA 12/13, L4T, SYCL f32/f16, Vulkan, ROCm/hipblas), backend/index.yaml
meta-backend + per-platform images, and a nightly bump_deps entry tracking the
`prism` branch.

Gallery: 8 entries across 4 families - bonsai-8b-1bit, ternary-bonsai-8b (+g64,
+pq2), bonsai-27b-1bit (vision), ternary-bonsai-27b (+pq2, +g64, vision). The 27B
models wire the mmproj vision tower; the DSpark speculative drafter GGUFs are not
wired (custom semi-autoregressive drafter, not a standard llama.cpp draft model).


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-16 10:09:14 +02:00
LocalAI [bot]
3880812ed6 chore: ⬆️ Update CrispStrobe/CrispASR to 5b38179a4a3281fcdba4220ff285f32e80df43a8 (#10851)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-16 10:02:09 +02:00
Tai An
808312b4b9 fix(watchdog): guard StopWatchdog with watchdogMutex to prevent double close (#10841) (#10859)
fix(watchdog): guard StopWatchdog with watchdogMutex to prevent double close

StopWatchdog checked, closed and cleared a.watchdogStop without holding
a.watchdogMutex, while startWatchdog and RestartWatchdog reassign and close the
same channel under that lock.

POST /api/settings dispatches to StopWatchdog or RestartWatchdog depending on
ApplicationConfig.WatchdogShouldRun(), so both are reachable concurrently. Two
callers can observe a non-nil watchdogStop and both close it, which panics with
'close of closed channel' and takes the server down.

Take the mutex, matching the other two writers. StopWatchdog is only called from
the settings handler, which holds no lock, so this cannot deadlock.

Fixes #10841

Co-authored-by: Anai Guo <antai12232931@anaiguo.com>
2026-07-16 09:40:29 +02:00
LocalAI [bot]
6a985d13ea chore: ⬆️ Update ikawrakow/ik_llama.cpp to 1fddd12ba861c4815a8633f14d9c5670692099cc (#10762)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-16 09:06:24 +02:00
Tai An
5fe48e4910 fix(backend): don't crash the whole process on an invalid cutstrings/extract_regex (#10855)
Finetune() compiled every model cutstrings/extract_regex entry via regexp.Compile
and called xlog.Fatal on failure, which terminates the entire local-ai process.
A single model config with an invalid regex (e.g. cutstrings: ["("]) turns one
/v1/chat/completions request into a process-level denial of service.

Log the compile error and skip the offending pattern instead. The mutex is
released before continuing, and skipping avoids dereferencing the nil regexp
that removing the fatal would otherwise leave behind.

Fixes #10843

Signed-off-by: Tai An <antai12232931@outlook.com>
2026-07-16 08:54:35 +02:00
LocalAI [bot]
ff8774327f feat(swagger): update swagger (#10847)
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-16 08:53:02 +02:00
Tai An
688f904a10 fix(runtime-settings): apply persisted threads/context_size/f16 at startup (#10853)
ApplyRuntimeSettings persists the performance settings (threads,
context_size, f16) on the live /api/settings path, but the startup
loader loadRuntimeSettingsFromFile never read them back, so a value
saved via the Middleware UI was silently ignored on the next restart:
the model booted with the CLI/physical-core default and GET /api/settings
echoed that default instead of the saved value (#10845).

Threads needs special handling: unlike context_size/f16, WithThreads
eagerly resolves an unset (0) value to xsysinfo.CPUPhysicalCores() at
option-apply time, so options.Threads is never 0 in the loader and the
usual "== default" heuristic cannot tell an env/CLI value from the
physical-core fallback. Detect LOCALAI_THREADS/THREADS explicitly so the
env still wins over the persisted file value.

Signed-off-by: Anai-Guo <Anai-Guo@users.noreply.github.com>
Co-authored-by: Anai-Guo <Anai-Guo@users.noreply.github.com>
2026-07-16 08:52:11 +02:00
LocalAI [bot]
8c9b3b2e33 chore(model-gallery): ⬆️ update checksum (#10854)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-16 08:41:35 +02:00
LocalAI [bot]
e488884b20 chore: ⬆️ Update ggml-org/llama.cpp to 505b1ed15ca80e2a19f12ff4ac365e40fb374053 (#10848)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-16 08:36:59 +02:00
LocalAI [bot]
e062179d4d chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 11c67198db58f75bf1bafc9051c2b018aaf1a3da (#10852)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-16 08:36:41 +02:00
Richard Palethorpe
b9d6d49e31 fix(cloud-proxy): publish backend gallery entries (#10858)
Add stable and development gallery variants for Linux and Darwin, and wire the backend build matrix so the referenced images are published.

Assisted-by: Codex:gpt-5 [yq]

Signed-off-by: Richard Palethorpe <io@richiejp.com>
2026-07-16 08:36:23 +02:00
LocalAI [bot]
a23fcc90c3 feat(gallery): add Qwen3.5-4B DFlash speculative-decoding model (#10842)
Pairs unsloth/Qwen3.5-4B-GGUF (Q4_K_M target) with the
AtomicChat/Qwen3.5-4B-DFlash-GGUF Q8_0 drafter (quantized from
z-lab/Qwen3.5-4B-DFlash, upstream GGUF arch `dflash`), same shape as
the existing DFlash entries.

Assisted-by: Claude Code:claude-fable-5 [Bash] [Read] [Edit]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-15 15:45:27 +02:00
LocalAI [bot]
d19c9875ed chore: ⬆️ Update ggml-org/llama.cpp to 00fa7cb284cbf133fc426733bd64238a3588a33e (#10814)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-15 09:59:46 +02:00
LocalAI [bot]
8cec22c3b7 feat(vram): per-node VRAM allocation budget (LOCALAI_VRAM_BUDGET) (#10833)
* feat(vram): add vrambudget primitive for per-node VRAM caps

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): apply default VRAM budget in xsysinfo aggregate getters

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): wire LOCALAI_VRAM_BUDGET flag to xsysinfo default budget

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): persist VRAM budget via runtime settings with live apply

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(vram): reset process-global VRAM budget after runtime-settings spec

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): add VRAM budget field to Settings page

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): store and enforce per-node VRAM budget in the node registry

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): apply per-node VRAM budget in router hardware defaults

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): report worker VRAM budget in node registration

The distributed worker now reports its operator-set VRAM budget string
(LOCALAI_VRAM_BUDGET) to the server on registration. The worker keeps
reporting RAW total/available VRAM and never sets the xsysinfo
process-global budget (that stays standalone-only); the server resolves
and enforces the budget uniformly (Task 6).

Also closes a Task 6 gap: on re-registration, a struct Updates zero-skips
an empty budget, so a worker that dropped LOCALAI_VRAM_BUDGET left the
stale cap in place. For non-admin-override nodes the budget columns are
now force-written (map Updates) even when empty, so removing the env var
clears the cap; admin overrides are preserved unchanged.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* style(vram): drop em dash from worker-clear comment

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): add node VRAM budget admin endpoints

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): add node VRAM budget control to the node UI

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vram): expose set_node_vram_budget MCP admin tool

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs(vram): document LOCALAI_VRAM_BUDGET and node VRAM budget UI

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vram): avoid double-applying VRAM budget in GetResourceAggregateInfo

The GPU-branch aggregate returned by GetResourceInfo is sourced from
GetGPUAggregateInfo, which already caps total/free/used against the
process-wide VRAM budget. GetResourceAggregateInfo then applied the
budget a second time. For an absolute budget this is idempotent, but for
a percentage budget b.Apply resolves the ceiling as a fraction of its
input total, so a second pass yields P*(P*T) instead of P*T and distorts
UsagePercent (read by the memory reclaimer in pkg/model/watchdog.go).

Remove the redundant second application so the budget is applied exactly
once, against the raw physical totals, upstream in GetGPUAggregateInfo.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vram): implement SetNodeVRAMBudget on mcp assistant test stub

The LocalAIClient interface gained SetNodeVRAMBudget; the stubClient in
core/http/endpoints/mcp used by the assistant tests is a separate
implementer and needs the method too (broke golangci-lint typecheck and
both test jobs).

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-15 09:58:45 +02:00
LocalAI [bot]
3601174ce0 fix(distributed): make per-node backend upgrade actually upgrade (#10838)
* test(core/http): make the suite's HTTP port overridable

app_test.go and openresponses_test.go hardcoded 127.0.0.1:9090. When
another service already listens on 9090 the suite does not fail fast:
the server goroutine logs the bind error and the specs then poll
whatever is squatting the port until Eventually times out. On machines
where 9090 is permanently taken this makes the pre-commit coverage gate
impossible to pass.

Introduce testHTTPAddr, defaulting to 127.0.0.1:9090 (what CI has
always used) and overridable via LOCALAI_TEST_HTTP_PORT for local runs.

Assisted-by: Claude:claude-fable-5 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): make per-node backend upgrade actually upgrade

The node detail page's Upgrade button reused the node-scoped install
path (POST /api/nodes/:id/backends/install). That fires NATS
backend.install with force=false, and the worker's install handler is
deliberately "ensure installed": when the backend binary already exists
on disk it short-circuits without touching the gallery. Since only an
installed backend can be upgraded, the whole chain was a guaranteed
successful no-op - the UI then toasted "backend upgraded" without even
waiting for the async job.

Route upgrades through the real force-reinstall path instead:

- BackendManager.UpgradeBackend now receives the ManagementOp (like
  InstallBackend already did) so implementations can honor
  op.TargetNodeID.
- DistributedBackendManager.UpgradeBackend scopes the backend.upgrade
  fan-out to op.TargetNodeID when set, and errors when the target node
  does not report the backend as installed.
- New POST /api/nodes/:id/backends/upgrade endpoint enqueues an
  Upgrade=true node-scoped op (async 202 + jobID, mirroring install).
- NodeDetail UI calls the new endpoint and reports the dispatch
  ("Upgrading ... on this node...") instead of claiming success; the
  Operations panel tracks the actual job.

Verified against a live local cluster (NATS + Postgres + two workers):
the target worker stops the running process, force-reinstalls from the
gallery and re-downloads the OCI image; the second worker receives no
backend.upgrade event; upgrading a backend missing from the target node
fails the job with a clear error.

Assisted-by: Claude:claude-fable-5 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-15 09:16:55 +02:00
LocalAI [bot]
40763d1181 chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to 7bb91886f613f4f54407604f4284e5b6ecd2acdf (#10832)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-15 09:02:14 +02:00
LocalAI [bot]
afbed9d49b chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 98a5d5fb43268fb85c637ac0a29ed67cc6a1f7d9 (#10830)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-15 01:09:46 +02:00
LocalAI [bot]
bcc41219f7 feat: materialize Hugging Face model artifacts (#10825)
* feat(config): add model artifact source contract

Assisted-by: Codex:GPT-5 [Codex]

* feat(downloader): add authenticated raw-byte progress

Assisted-by: Codex:GPT-5 [Codex]

* feat(huggingface): resolve immutable snapshot manifests

Assisted-by: Codex:GPT-5 [Codex]

* feat(models): add artifact storage primitives

Assisted-by: Codex:GPT-5 [Codex]

* feat(models): materialize pinned Hugging Face snapshots

Assisted-by: Codex:GPT-5 [Codex]

* feat(models): bind managed snapshots at runtime

Assisted-by: Codex:GPT-5 [Codex]

* feat(gallery): materialize model artifacts during install

Assisted-by: Codex:GPT-5 [Codex]

* feat(gallery): declare managed Hugging Face artifacts

Assisted-by: Codex:GPT-5 [Codex]

* feat(models): preload managed model artifacts

Assisted-by: Codex:GPT-5 [Codex]

* fix(gallery): retain shared artifact caches on delete

Assisted-by: Codex:GPT-5 [Codex]

* feat(models): report artifact acquisition progress

Assisted-by: Codex:GPT-5 [Codex]

* refactor(backends): load managed models from ModelFile

Assisted-by: Codex:GPT-5 [Codex]

* refactor(backends): load staged speech model snapshots

Assisted-by: Codex:GPT-5 [Codex]

* refactor(backends): use staged snapshots in engine backends

Assisted-by: Codex:GPT-5 [Codex]

* test(distributed): cover staged artifact snapshots

Assisted-by: Codex:GPT-5 [Codex]

* docs: explain managed model artifacts

Assisted-by: Codex:GPT-5 [Codex]

* docs: add product design context

Assisted-by: Codex:GPT-5 [Codex]

* feat(ui): show model artifact download progress

Assisted-by: Codex:GPT-5 [Codex]

* Eagerly materialize Hugging Face artifacts

Materialize HF-backed model references as managed GGUF artifacts during load, with lazy download retained only as fallback.

Assisted-by: Codex:GPT-5 [shell]

* Refactor HF
  downloads through a shared executor

Assisted-by: Codex:GPT-5 [shell]

* drop

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-15 01:09:33 +02:00
LocalAI [bot]
d82c38ee77 chore: ⬆️ Update leejet/stable-diffusion.cpp to a8a91b24cdf18a3e415d7f2a28f69b5be8a17700 (#10828)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-15 01:09:18 +02:00
LocalAI [bot]
64124f3fa1 chore: ⬆️ Update CrispStrobe/CrispASR to 40d508096bb52850862edafc9741da509c5ede97 (#10829)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-15 00:53:17 +02:00
LocalAI [bot]
88cc80ee3d chore(model-gallery): ⬆️ update checksum (#10831)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-14 23:53:45 +02:00
LocalAI [bot]
bed5e7417c docs: ⬆️ update docs version mudler/LocalAI (#10826)
⬆️ Update docs version mudler/LocalAI

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-14 23:53:28 +02:00
LocalAI [bot]
ba1d0f5507 chore: ⬆️ Update vllm-project/vllm cu130 wheel to 0.25.1 (#10827)
⬆️ Update vllm-project/vllm cu130 wheel

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-14 23:53:16 +02:00
LocalAI [bot]
2bed6f65ba fix(kokoro): pin compatible Intel XPU runtime (#10823)
PyTorch 2.13 XPU pulls oneAPI 2026 libraries that conflict with the oneAPI 2025.3 backend image. Pin torch and torchaudio to the matching 2.11 XPU pair so the build resolves a coherent 2025.3 runtime.

Assisted-by: Codex:GPT-5 [uv]

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-14 18:58:13 +02:00
LocalAI [bot]
b224c96db6 fix(config): only inject llama.cpp serving options on the llama.cpp path (#10822)
SetDefaults injected the llama.cpp server options cache_reuse
(ApplyServingDefaults) and parallel (ApplyHardwareDefaults, re-applied
per selected node by the distributed router) onto every model config
regardless of backend. Every other backend ignores options it does not
understand, so this was harmless until longcat-video, which strictly
validates its options and fails LoadModel with
"unknown model option(s): cache_reuse, parallel".

Gate both injections behind a new UsesLlamaCppServingOptions allow-list
(llama-cpp plus the empty/auto-detect case that resolves to llama.cpp
from a GGUF file, mirroring how llamaCppDefaults is registered). This
follows the existing UsesLlamaSamplerDefaults precedent for llama-only
defaults. The typed NBatch field is deliberately left alone: it is a
proto field every backend simply ignores, which is why batch never
triggered the error.

Also harden the longcat-video backend to warn-and-ignore unknown model
options and request params through a testable select_known_options
helper, matching the other LocalAI Python backends, so a future
server-injected option cannot break loading again.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-14 17:46:15 +02:00
LocalAI [bot]
9f14571397 chore(model-gallery): ⬆️ update checksum (#10816)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-14 11:25:00 +02:00
Ijas
a5aa56db81 fix: preserve uploaded file content when regenerating a non-last answer (#10819)
handleRegenerate rebuilt the outbound message from the display-only
message.files metadata ({name, type: 'file'|'image'|..., content}),
which doesn't carry the base64/textContent payload sendMessage's
file-building loop expects. As a result, regenerating any answer whose
own question had an attachment silently dropped that attachment from
the resent message. This wasn't fork-specific, but forking a chat and
then regenerating an earlier (now non-last) answer is the natural way
to hit it.

Fix by reusing the original message's already-assembled `content`
verbatim (it already has the file text / image_url / audio_url /
video_url parts embedded from the first send) instead of trying to
reconstruct it from lossy display metadata.

Fixes #10806

Assisted-by: Claude:claude-sonnet-5

Signed-off-by: ajuijas <189517297+ajuijas@users.noreply.github.com>
Co-authored-by: ajuijas <189517297+ajuijas@users.noreply.github.com>
2026-07-14 11:24:42 +02:00
LocalAI [bot]
05b8d8aafe chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to bbdf4be03aa5bc3c188b5db778b6f3fd63ceff6c (#10794)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-14 08:38:20 +02:00
LocalAI [bot]
3c1e583985 chore: ⬆️ Update CrispStrobe/CrispASR to d76cce027e3b183fc3d8c72e976e69d11f71bc8b (#10813)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-14 08:38:05 +02:00
LocalAI [bot]
2609848d80 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260713103604 (#10812)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-14 08:36:49 +02:00
dependabot[bot]
cdb6702ca6 chore(deps): bump actions/stale from 10.3.0 to 10.4.0 (#10807)
Bumps [actions/stale](https://github.com/actions/stale) from 10.3.0 to 10.4.0.
- [Release notes](https://github.com/actions/stale/releases)
- [Changelog](https://github.com/actions/stale/blob/main/CHANGELOG.md)
- [Commits](eb5cf3af3a...1e223db275)

---
updated-dependencies:
- dependency-name: actions/stale
  dependency-version: 10.4.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-13 22:50:01 +02:00
dependabot[bot]
0d8bea0158 chore(deps): update charset-normalizer requirement from >=3.4.7 to >=3.4.9 in /backend/python/vllm (#10809)
chore(deps): update charset-normalizer requirement

Updates the requirements on [charset-normalizer](https://github.com/jawah/charset_normalizer) to permit the latest version.
- [Release notes](https://github.com/jawah/charset_normalizer/releases)
- [Changelog](https://github.com/jawah/charset_normalizer/blob/master/CHANGELOG.md)
- [Commits](https://github.com/jawah/charset_normalizer/compare/3.4.7...3.4.9)

---
updated-dependencies:
- dependency-name: charset-normalizer
  dependency-version: 3.4.9
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-13 22:49:44 +02:00
dependabot[bot]
0b0f52bedc chore(deps): bump grpcio from 1.81.1 to 1.82.1 in /backend/python/vllm (#10808)
Bumps [grpcio](https://github.com/grpc/grpc) from 1.81.1 to 1.82.1.
- [Release notes](https://github.com/grpc/grpc/releases)
- [Commits](https://github.com/grpc/grpc/compare/v1.81.1...v1.82.1)

---
updated-dependencies:
- dependency-name: grpcio
  dependency-version: 1.82.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-13 22:49:28 +02:00
dependabot[bot]
48b1ab28b7 chore(deps): bump vllm from 0.24.0 to 0.25.0 in /backend/python/vllm (#10811)
Bumps [vllm](https://github.com/vllm-project/vllm) from 0.24.0 to 0.25.0.
- [Release notes](https://github.com/vllm-project/vllm/releases)
- [Changelog](https://github.com/vllm-project/vllm/blob/main/RELEASE.md)
- [Commits](https://github.com/vllm-project/vllm/compare/v0.24.0...v0.25.0)

---
updated-dependencies:
- dependency-name: vllm
  dependency-version: 0.25.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-13 22:48:35 +02:00
Dedy F. Setyawan
b10e330590 feat(react-ui): localize MediaHistory, AudioTransform, and Sound components (#10802)
- Replace hardcoded text with useTranslation hook in UI components
- Add localization support for both English (en) and Indonesian (id) locales

Signed-off-by: Dedy F. Setyawan <dedyfajars@gmail.com>
2026-07-13 11:51:23 +00:00
LocalAI [bot]
4056283aa4 [voice] feat: add managed voice cloning profiles (#10799)
* feat(ui): add voice library workflow

Give administrators a production-ready flow to record or upload consented reference audio, manage reusable profiles, inspect API usage, discover compatible models, and hand a saved voice directly to text-to-speech.

Assisted-by: Codex:gpt-5

* feat(voice): add managed voice cloning profiles

Make reusable reference voices manageable through the admin API instead of requiring model-directory and YAML edits. Discover compatible installed and gallery models from server-side backend capabilities, retain explicit model configuration controls, and stage saved references for supported backends.

Expose profile management through REST and MCP, document backend-specific behavior, and cover the workflow from profile creation through real Qwen3-TTS synthesis. Harden the agent-job HTTP test against completion racing cancellation.

Assisted-by: Codex:gpt-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-13 09:54:46 +02:00
LocalAI [bot]
b90e1cae73 chore: ⬆️ Update ggml-org/llama.cpp to 6b4dc2116a92c5c8f2782bfe51fabe5ee66fb5ef (#10797)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-13 09:05:44 +02:00
LocalAI [bot]
c43ee40eb7 chore: ⬆️ Update CrispStrobe/CrispASR to 841281c46ce5b34323b2861ae5714c09aaa2542e (#10795)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-13 01:10:16 +02:00
LocalAI [bot]
659b9f02e0 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260711221406 (#10796)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-13 01:10:03 +02:00
LocalAI [bot]
67c14e1b7e chore(model-gallery): ⬆️ update checksum (#10798)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-13 01:09:13 +02:00
LocalAI [bot]
0e6241a5aa chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to d17c33d4ee2f56d15f9ca8a1bb82f7389305f838 (#10793)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-13 01:09:03 +02:00
LocalAI [bot]
b00422e45f feat(backends): add LongCat video and avatar generation (#10792)
* feat(backends): add LongCat video and avatar generation

Assisted-by: Codex:GPT-5 [apply_patch] [exec_command] [web]

* refactor(config): declare model I/O modalities

Make model configs declare input and output modalities so capability discovery no longer branches on backend or checkpoint names. Complete the LongCat gallery and user documentation, make the SDPA patch apply to the pinned upstream revision, and stabilize the Agent Jobs race exposed by the required hook.

Assisted-by: Codex:GPT-5 [web]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-12 23:58:46 +02:00
LocalAI [bot]
af8f74cba2 chore: ⬆️ Update CrispStrobe/CrispASR to 4beda42f63bf5c813c1e1bb55249efb36f918c80 (#10784)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-12 11:27:12 +02:00
LocalAI [bot]
cdd9582653 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to bfc447c47a592698b539e0b819b1a4e3c91c4730 (#10786)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-12 11:14:22 +02:00
LocalAI [bot]
fb0f5e4bdd feat(gallery): add Qwen DFlash speculative-decoding models (#10791)
Add four ready-to-run DFlash speculative-decoding entries for the
llama.cpp backend, now that upstream DFlash support (draft-dflash) is in
the pinned llama.cpp. Each entry bundles a full target model with its
small z-lab block-diffusion drafter and sets spec_type:draft-dflash,
spec_n_max:15, and flash attention (required by DFlash):

- qwen3-4b-dflash          (Qwen3-4B + Qwen3-4B-DFlash drafter)
- qwen3.5-9b-dflash        (Qwen3.5-9B + Qwen3.5-9B-DFlash drafter)
- qwen3.6-27b-dflash       (Qwen3.6-27B dense + drafter)
- qwen3.6-35b-a3b-dflash   (Qwen3.6-35B-A3B MoE + drafter)

The 4B pair uses the base Qwen3-4B target (not Qwen3.5-4B): its drafter
reports general.name "Qwen3 4B DFlash" and is the canonical pairing
documented upstream. All drafters were downloaded and verified to carry
GGUF architecture "dflash" (not the fork-only "dflash-draft" /
"DFlashDraftModel") so they load in the upstream backend, and every
drafter SHA256 was confirmed against the downloaded bytes.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-12 11:14:06 +02:00
LocalAI [bot]
5013d53a1c chore: ⬆️ Update ggml-org/llama.cpp to e3546c7948e3af463d0b401e6421d5a4c2faf565 (#10787)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-12 10:07:58 +02:00
LocalAI [bot]
e8d8f5b0b8 chore: ⬆️ Update ggml-org/whisper.cpp to 080bbbe85230f624f0b52127f1ae1218247989f9 (#10785)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-12 10:07:45 +02:00
LocalAI [bot]
459ffb3054 chore: ⬆️ Update leejet/stable-diffusion.cpp to b5d812008eb7082a238fc589444544b3278187ae (#10774)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-12 10:07:34 +02:00
LocalAI [bot]
cad07be2fc chore(model-gallery): ⬆️ update checksum (#10789)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-11 23:41:48 +02:00
hogeheer499-commits
2634b13a5d fix(ds4): bundle transitive runtime dependencies (#10783)
Bundle the resolved dependency closure for grpc-server and ds4-worker, then validate that packaged dependencies resolve only from package/lib.

Assisted-by: Codex:gpt-5 shellcheck

Signed-off-by: JS van Dijk <267467744+hogeheer499-commits@users.noreply.github.com>
Co-authored-by: JS van Dijk <267467744+hogeheer499-commits@users.noreply.github.com>
2026-07-11 23:41:27 +02:00
LocalAI [bot]
8786eace97 chore: ⬆️ Update vllm-project/vllm cu130 wheel to 0.25.0 (#10788)
⬆️ Update vllm-project/vllm cu130 wheel

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-11 23:39:51 +02:00
Ettore Di Giacinto
a1cefe862d Rename voxtral.c to voxtral-tts.c in README
Signed-off-by: Ettore Di Giacinto <mudler@users.noreply.github.com>
2026-07-11 16:24:44 +02:00
LocalAI [bot]
50cd897719 docs: refresh LocalAI homepage (#10780)
* docs: refresh LocalAI homepage

Reframe the homepage around LocalAI's modular multimodal runtime, native inference engines, deployment range, and built-in platform capabilities. Remove the outdated video in favor of current product visuals and clearer paths into the documentation.

Assisted-by: Codex:gpt-5

* docs: give homepage a full-width canvas

Let the product homepage opt out of Relearn's persistent sidebar and duplicate title while preserving the documentation shell on interior pages. Tighten the responsive bounds for narrow screens.

Assisted-by: Codex:gpt-5

* docs: fit homepage to the Relearn content flow

Remove the full-width shell exception and use a single-column homepage inside the standard documentation layout. This avoids competing scroll containers and the compressed split hero.

Assisted-by: Codex:gpt-5

* docs: hide homepage scroll rail

Preserve Relearn's content scrolling while removing the visible scrollbar beside the landing-page hero.

Assisted-by: Codex:gpt-5

* docs: contain homepage sections within docs column

Prevent landing-page headings, figures, and section grids from widening Relearn's content pane or exposing overflow rails.

Assisted-by: Codex:gpt-5

* docs: remove nested homepage scrollbars

Wrap the quick-start command within its column and suppress component-level scrollbar tracks across the landing page.

Assisted-by: Codex:gpt-5

* docs: remove outdated gallery screenshot

Drop the stale Model Gallery image from the homepage until a current product visual is available.

Assisted-by: Codex:gpt-5

* docs: fix homepage architecture link

Point the homepage CTA at the generated reference/architecture route.

Assisted-by: Codex:gpt-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-11 09:33:41 +02:00
LocalAI [bot]
8ff3c8c466 chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to bb250f57e48d7dbaff8f66b8125c42a14ddabbe7 (#10759)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-11 09:17:48 +02:00
Richard Palethorpe
1f9fda7138 fix(backends): refuse foreign model loads in opus and local-store (#10769)
When a model config has no explicit backend, the model loader greedily
probes every installed backend and binds to the first Load that
succeeds. opus and local-store were the only in-tree backends with no
model artefact to validate, so they accepted anything — an LLM
installed after them could silently bind to the audio codec or the
vector store and then fail at inference with "unimplemented"
(see #9287).

opus now accepts only its own name (what the realtime WebRTC path
sends) or none. local-store namespaces are arbitrary (router caches,
biometrics, user-named stores), so core's StoreBackend now marks
genuine store loads with a store:// prefix on the gRPC model name and
the backend refuses names without it; core and backend ship from the
same release, so the convention upgrades in lockstep.

Also repair the bit-rotted 'make test-stores' bootstrap (the suite
never registered external backends, so BACKENDS_PATH was dead weight)
and add the Load-validation rule to the adding-backends checklist.

Related: #9287

Assisted-by: Claude:claude-fable-5 golangci-lint

Signed-off-by: Richard Palethorpe <io@richiejp.com>
2026-07-11 09:17:34 +02:00
LocalAI [bot]
23a044ee0b chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260710141421 (#10775)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-11 09:16:12 +02:00
LocalAI [bot]
921a8ffc8b chore: ⬆️ Update mudler/moss-transcribe.cpp to 92a923dca88a41a34e47a364d55ee25731a9a0a2 (#10771)
⬆️ Update mudler/moss-transcribe.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-11 09:15:58 +02:00
LocalAI [bot]
a8f1c92a24 chore: ⬆️ Update CrispStrobe/CrispASR to 74efaf24d457e34cd200d138f9724987e7ececc3 (#10773)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-11 09:15:45 +02:00
LocalAI [bot]
6084497da1 chore: ⬆️ Update ggml-org/llama.cpp to 4f37f519722aa3242eecb7649466b4a4a2d6d6da (#10772)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-11 09:15:32 +02:00
LocalAI [bot]
2482f075a2 chore: ⬆️ Update ggml-org/whisper.cpp to 7695a5331230c585f5ce92291c4256973985ae5a (#10776)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-11 09:14:19 +02:00
LocalAI [bot]
9b4f373bc4 chore(model-gallery): ⬆️ update checksum (#10778)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-11 09:14:06 +02:00
LocalAI [bot]
185956154a chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to b297bc36d5e54ac37263fc8f2766e96179bf1e8a (#10761)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-10 10:43:11 +02:00
LocalAI [bot]
d3d5488dc7 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260708043308 (#10760)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-10 10:42:56 +02:00
LocalAI [bot]
ae58115ee6 chore: ⬆️ Update leejet/stable-diffusion.cpp to cc734292286f85f9c48305d94d7fd22f42838522 (#10738)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-10 10:42:42 +02:00
LocalAI [bot]
c7a9db29a6 chore: ⬆️ Update CrispStrobe/CrispASR to aa23c61c8ba483f539fa669708ea7ddbba72f293 (#10758)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-10 10:42:26 +02:00
LocalAI [bot]
c46a224b44 chore: ⬆️ Update ggml-org/llama.cpp to 049326a00025d00b08cc188ed716b681e984a3f8 (#10757)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-10 10:42:12 +02:00
LocalAI [bot]
7e8542ba32 fix(tests): make e2e backend model downloads resumable and stall-based (#10766)
The vibevoice transcription e2e hangs until the go test timeout when
the HF CDN is slow: the ASR Q4_K model is >10 GB, downloadFile capped
every curl attempt at --max-time 600 (needs a sustained ~17 MB/s to
fit), and curl's --retry restarts from byte zero, so no attempt ever
makes forward progress. This killed the job twice on PR #10764 and
previously forced skipping it on release tags (#10567).

Replace the wall-clock cap with stall detection (--speed-limit 1 MiB/s
over --speed-time 120s) and resume from the bytes already on disk with
-C -, retrying from Go because curl does not re-evaluate the resume
offset on its internal retries. Resume against the HF Xet CDN was
verified by killing a transfer mid-flight and confirming the next
invocation appended (114 MB -> 235 MB, GGUF magic intact).

Also parameterize the suite timeout (BACKEND_TEST_TIMEOUT, default
30m) and raise it to 120m for the vibevoice transcription wrapper: a
10 GB download plus 25 specs does not fit in 30m even on a good day,
and the job-level GHA timeout there is already 150m.


Assisted-by: Claude:claude-fable-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-10 09:19:17 +02:00
LocalAI [bot]
3c2d85aae4 feat(vibevoice-cpp): true streaming TTS (TTSStream via vv_capi_tts_stream) (#10764)
* feat(vibevoice-cpp): true streaming TTSStream via vv_capi_tts_stream

Replaces the synth-to-tempfile TTSStream hack with a real streaming path:
binds the new vv_capi_tts_stream callback ABI via a single reusable purego
callback (CGO_ENABLED=0-safe, no runtime/cgo), copies each int16 PCM window
into the gRPC results channel after the streaming WAV header.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* test(vibevoice-cpp): real-model streaming integration test with TTFA measurement

Gated behind VIBEVOICE_IT=1, this Ginkgo spec dlopens the engine .so and
drives the exact Go->purego->C TTSStream/TTS path against the real
vibevoice-realtime-0.5B model. It measures time-to-first-audio for the
streaming path versus the batch path and asserts the streaming win:
44-byte WAV header first, >=2 PCM windows, non-silent audio, and
TTFA < total_stream. Without the env var the spec skips so CI and
normal go test are unaffected.

Measured: TTFA 2.38s vs batch deliver-time 39.96s (first audio in 5.9%
of the batch time, ~17x faster), 18 stream chunks, non-silent 24kHz PCM.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* chore(vibevoice-cpp): pin streaming-decoder engine build

Bumps VIBEVOICE_CPP_VERSION to the streaming-decoder engine commit that
adds vv_capi_tts_stream (localai-org/vibevoice.cpp#8). Re-pin to the
merged master commit once that PR lands.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* chore(vibevoice-cpp): re-pin to merged streaming-decoder commit

localai-org/vibevoice.cpp#8 merged to master as 000e372; move the pin
off the PR branch commit onto the merged master commit.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* test(vibevoice-cpp): check writer errors in TTFA report (errcheck)

golangci-lint errcheck flagged the unchecked fmt.Fprintf calls that
print the streaming TTFA headline. Build the report once with
fmt.Sprintf and write it per destination with an explicitly discarded
error, matching the GinkgoWriter reporting idiom used by the other
backend tests.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-fable-5 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-09 23:48:44 +00:00
LocalAI [bot]
6ceb2f86a7 chore(model-gallery): ⬆️ update checksum (#10763)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-10 00:41:19 +02:00
LocalAI [bot]
c5b36639d4 chore(model gallery): 🤖 add 1 new models via gallery agent (#10755)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-09 23:28:44 +02:00
LocalAI [bot]
94bdc825dc feat(backend): add moss-transcribe-cpp backend (MOSS-Transcribe-Diarize) (#10756)
C++/ggml transcription + speaker diarization + timestamps backend. Purego
dlopens libmoss-transcribe.so (ggml statically linked) from moss-transcribe.cpp
and serves offline AudioTranscription, parsing the [start][Sxx]text[end] output
into segments with nanosecond timestamps. Adds the importer (surfaces in
GET /backends/known), backend-matrix (Linux + Darwin/metal), backend/index.yaml,
and a gallery entry (default q5_k GGUF from mudler/moss-transcribe.cpp-gguf).

Local L0 smoke (go build + go test ./... = 16 pass, golangci-lint 0 issues)
passed against the real libmoss-transcribe.so. The pre-commit coverage gate
(full pkg/core + tests/e2e) could not run in the authoring sandbox (no live
models, port 9090 held); CI must enforce it before merge.

Assisted-by: Claude:claude-opus-4-8 golangci-lint

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-09 23:27:11 +02:00
Dedy F. Setyawan
294487eb61 fix(ui): prevent table container from breaking flexbox layout (#10754)
Signed-off-by: Dedy F. Setyawan <dedyfajars@gmail.com>
2026-07-09 23:21:07 +02:00
Ettore Di Giacinto
fa3e139540 docs(readme): add face-detect.cpp, voice-detect.cpp and free-splatter.cpp to native engines list
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-09 11:51:28 +00:00
LocalAI [bot]
35024338a6 feat(crispasr): add F5-TTS support and gallery model (#10753)
Link the f5-tts library into the crispasr backend so CrispASR's native
F5-TTS runtime (SWivid F5-TTS, 22-layer DiT flow-matching + built-in Vocos
vocoder) is compiled in. The single self-contained GGUF auto-detects as
f5-tts through the session router, so no explicit backend selector is
needed. Add the f5-tts-crispasr gallery entry (cstr/f5-tts-GGUF) and an
env-gated e2e synthesis spec.

F5-TTS is voice-cloning only and has no baked speaker: it clones from a
reference WAV plus its transcript, supplied via the voice/voice_text
options. The gallery description documents this bring-your-own-reference
requirement.

Verified e2e on the pinned engine (278fb79): the GGUF auto-detects as
f5-tts, the reference voice loads, and synthesis produces a valid 24 kHz
mono WAV.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-09 08:09:00 +00:00
LocalAI [bot]
b987f39de8 feat(swagger): update swagger (#10745)
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-09 09:35:22 +02:00
LocalAI [bot]
16d028a127 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to e03824a9f874fe648b8f26bf293703af60afe936 (#10715)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-09 09:05:40 +02:00
LocalAI [bot]
70e15679cd chore: ⬆️ Update ggml-org/llama.cpp to a646006f09d2f76f2d62d6c0d5e8e8490d570720 (#10747)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-09 09:03:51 +02:00
LocalAI [bot]
5569b2de56 feat(config): context_size: -1 to auto-use model's full trained context (#10752)
* feat(config): clamp negative context_size to default in EffectiveContextSize

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* feat(config): resolve context_size=-1 to model trained max with VRAM warn

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* fix(config): treat negative context_size as unset when GGUF is unparseable

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* docs(config): document context_size=-1 auto-max

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* docs(backend): drop em dashes from EffectiveContextSize comment

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-09 09:03:40 +02:00
LocalAI [bot]
c9f73f40ff chore: ⬆️ Update CrispStrobe/CrispASR to 278fb7927633d47fa0aeb6f81491b9952ed17c35 (#10737)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-09 01:13:06 +02:00
LocalAI [bot]
4cd3dbe931 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 6198a356a85ed71534c02a9c1026203389f341e5 (#10746)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-09 01:12:48 +02:00
LocalAI [bot]
1a04b670f4 chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to e93b3fedd53504c29c1c1a9bed4fbe722bbb1df5 (#10748)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-09 01:12:35 +02:00
LocalAI [bot]
e948f27965 chore(model-gallery): ⬆️ update checksum (#10749)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-09 01:12:22 +02:00
LocalAI [bot]
40dae953f4 feat: interleaved thinking with tool calls (reasoning_content alias + Anthropic thinking blocks) (#10744)
* feat(schema): accept reasoning_content as inbound alias for reasoning

Interleaved-thinking clients (cogito, vLLM/DeepSeek-style) emit reasoning_content
on assistant turns. Accept it as an inbound alias so reasoning survives the
tool-result loop; canonical reasoning wins when both are present. Emission is
unchanged (still reasoning).

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(schema): pin interleaved reasoning+tool_calls round-trip

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(openai): pin reachedTokenBudget truncation detection

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(anthropic): add thinking and signature fields to content blocks

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(anthropic): parse inbound thinking blocks into reasoning

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(anthropic): emit thinking blocks with synthetic signature on tool turns

Extract buildAnthropicContentBlocks so non-streaming content assembly is
unit-testable, and prepend a thinking block (with an opaque synthetic
signature) before text/tool_use blocks when the request opts into thinking.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(anthropic): stream thinking_delta and signature_delta before tool_use

Extract anthropicStreamSequence so the streaming block order is unit-testable,
and emit content_block_start(thinking) -> thinking_delta -> signature_delta ->
content_block_stop before the tool_use block sequence when thinking is enabled.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: add interleaved thinking with tool calls guide

Add a features guide describing interleaved thinking: an assistant turn
carrying reasoning and tool_calls together, the reasoning-round-trip
contract (including the reasoning_content inbound alias and Anthropic
thinking blocks with a synthetic signature), per-backend enablement
(reasoning_format for llama.cpp, reasoning_parser/tool_call_parser for
vLLM/SGLang plus the vLLM auto-config hook), a worked request/response
example, and known limitations. Cross-link from model-configuration,
text-generation, and openai-functions.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-08 16:45:43 +00:00
LocalAI [bot]
8671c8adac chore(model gallery): 🤖 add 1 new models via gallery agent (#10743)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-08 17:43:02 +02:00
LocalAI [bot]
d5d659bb65 fix(backends): pin grpcio-tools to installed grpcio in runProtogen (protobuf gencode mismatch) (#10735)
fix(backends): pin grpcio-tools to the installed grpcio in runProtogen

runProtogen installed grpcio-tools unpinned, so the protoc it bundles
stamped backend_pb2.py with the newest Protobuf gencode (7.35.0). When a
backend caps the protobuf runtime lower -- vLLM pins protobuf to 6.33.6 --
the import-time guarantee runtime >= gencode fails:

  google.protobuf.runtime_version.VersionError: Detected incompatible
  Protobuf Gencode/Runtime versions ... gencode 7.35.0 runtime 6.33.6

The backend crashes on `import backend_pb2` before it can serve, which
surfaces to the user as "grpc service not ready". It was mis-reported as a
ROCm/gfx1201 failure in #10718 but is not GPU-specific and affects every
vLLM variant (and any backend that caps protobuf below the latest gencode).

Pin grpcio-tools to the grpcio version the backend already installed --
they release in lockstep -- so the generated gencode stays in step with
the protobuf runtime. Falls back to unpinned when grpcio isn't present.

Closes #10718


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-08 15:24:42 +02:00
LocalAI [bot]
a0ed395cf9 fix(auth): accept EC/PS/EdDSA-signed OIDC ID tokens, not just RS256 (#10736)
The OIDC verifier was built with a bare oidc.Config{ClientID: ...}, so
go-oidc applied its default of accepting RS256-signed ID tokens only. An
identity provider configured with an EC signing key (e.g. Authentik) issues
ES256-signed tokens, and the callback failed verification with:

  failed to verify ID token: oidc: malformed jwt: unexpected signature
  algorithm "HS256"; expected ["RS256"]

surfacing to the user as HTTP 500 "failed to fetch user info" (#10677; the
underlying cause became visible after the logging fix in #10679).

Set SupportedSigningAlgs to the standard asymmetric algorithms
(RS256/384/512, ES256/384/512, PS256/384/512, EdDSA). All are verified
against the provider's published JWKS. HS256 is intentionally excluded: it
is symmetric and would validate against the client secret, a different and
security-sensitive trust model.

Tested with a functional spec that signs an ES256 ID token and confirms it
verifies with the configured algorithms and is rejected under go-oidc's
RS256-only default (using oidc.StaticKeySet, no network).

Closes #10677


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-08 15:24:29 +02:00
LocalAI [bot]
40c29db8c4 fix(logs): capture backend logs by default in single mode (#10742)
Backend log capture into the per-model BackendLogStore (which feeds the
UI "Backend Logs" page and /api/backend-logs) was opt-in and off by
default in single mode, while worker/distributed mode force-enables it
via SetBackendLoggingEnabled(true). There was no CLI flag either, so the
only way to populate the store was the Settings UI toggle - and the page
was silently empty out of the box. Distributed "just worked"; single
mode looked broken.

Default EnableBackendLogging to true in NewApplicationConfig so single
mode matches worker mode. The store is a small in-memory ring buffer, so
the cost is negligible.

Now that the default is on, loadRuntimeSettingsFromFile's usual
"only flip false->true" merge would ignore a persisted false and revert
the UI toggle-off on every restart. There is no env var/CLI flag for
this setting, so an explicit persisted value is now authoritative in
both directions, letting the toggle-off survive a restart.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-08 12:13:52 +00:00
LocalAI [bot]
0aaf7cce76 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 5c2552b6e820cc791ba9cfa6ffdfd98002b4cd82 (#10732)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-08 08:50:00 +02:00
LocalAI [bot]
731bf04668 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260707023344 (#10733)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-08 08:49:48 +02:00
LocalAI [bot]
d829e818d0 chore: ⬆️ Update ggml-org/llama.cpp to bec4772f6a2527d371557b5d2032641e5ff7619c (#10739)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-08 08:49:34 +02:00
github-actions[bot]
0ae84be362 chore: bump inference defaults from unsloth (#10741)
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-08 08:49:22 +02:00
LocalAI [bot]
d521608e6d chore: ⬆️ Update mudler/locate-anything.cpp to ade2634f7f79b56121125e5885628744795a478f (#10734)
⬆️ Update mudler/locate-anything.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-07 22:03:16 +00:00
LocalAI [bot]
7dde5a4225 fix(diffusers,vllm-omni,tinygrad): save generated images as PNG explicitly (#10729)
The image backends call PIL Image.save(request.dst) without a format, so
Pillow infers the encoder from the file extension. The core passes an
absolute staging path ending in .tmp (e.g. /staging/localai-output-*.tmp),
which Pillow can't map to a format, raising "unknown file extension: .tmp"
and crashing the worker right after a successful GPU inference.

Pass format="PNG" explicitly. LocalAI serves generated images as PNG
regardless of the temporary path, so this is always correct and no longer
depends on the extension of the destination the core happens to allocate.

diffusers is the reported backend (#10727); vllm-omni and tinygrad carry
the identical latent crash for any .tmp staging destination.

Closes #10727


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-07 21:24:58 +00:00
LocalAI [bot]
cd65a1f645 fix(transcription): honor model-config language/translate + OpenAI language form field (#10731)
fix(transcription): honor model-config language/translate, add language form field

The /v1/audio/transcriptions endpoint read only input.Language /
input.Translate from the parsed request, and the request middleware never
populates those from a multipart upload -- nor did it read a `language`
form field. As a result the model config's parameters.language /
parameters.translate (a valid PredictionOptions field under `parameters:`)
were silently ignored, and multilingual models like canary defaulted to
translating into English even when the YAML set language: ru,
translate: false (#10655).

Resolve both with clear precedence: the request form field wins, then any
language on the parsed request, then the model config default. This also
makes the endpoint honor OpenAI's `language` form parameter, which was
not read before.

Applies to both the streaming and non-streaming paths (the resolved
values are built into the shared TranscriptionRequest). Note this ensures
the language/translate flags reach the backend; whether a given engine
acts on them is up to the backend.

Closes #10655


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-07 21:16:25 +00:00
LocalAI [bot]
97175f4b5a feat(model): debounce model loads after a failure to stop retry-storms (#10728)
A client that keeps polling a model whose load fails (e.g. a backend that
crashes deterministically on init) triggered a fresh backend start on
every request: request -> load -> crash in ~10s -> 500, repeat on the
next poll. Each attempt could leak GPU/CUDA state, and under
LOCALAI_SINGLE_ACTIVE_BACKEND it kept stealing the active slot from
healthy models. The existing loading-coalesce map only dedups
*concurrent* loads, so sequential polls were never covered.

Track load failures per modelID in ModelLoader. After a load fails,
refuse fresh load triggers for that model until a cooldown elapses,
returning a typed ModelLoadCooldownError that the HTTP layer maps to 503
with a Retry-After header. The cooldown grows exponentially per
consecutive failure (base, doubling, capped at 5m) and resets on a
successful load. The coalesced follower-retry of an in-flight burst
bypasses the gate, so a genuinely concurrent burst still gets its one
retry -- only new, independent triggers are refused, matching the
report's "refuse new load-triggers" wording.

Configurable via --model-load-failure-cooldown /
LOCALAI_MODEL_LOAD_FAILURE_COOLDOWN (default 10s, 0 disables), plumbed
through ApplicationConfig and applied unconditionally at startup.

Closes #10719


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-07 20:55:24 +00:00
LocalAI [bot]
1d5139f0a0 fix(gallery): make backend (re)install a clean replace instead of an overlay (#10726)
InstallBackend extracted the artifact directly into the target directory
with no pre-clean, so a reinstall overlaid the new files onto the old
ones. Files present in a previous version but absent in the new artifact
(a stale .so, an orphaned package dir) survived and could shadow the new
build at import time -- e.g. an old vllm shared object lingering next to
a freshly pulled one. Only a failed download cleaned the directory.

Stage the download/extraction into a `<name>.install-tmp` dir, validate
run.sh is present, write metadata, then atomically swap it into place
(rename current -> .install-backup, staging -> current, drop backup),
rolling back on failure. This mirrors the atomic swap UpgradeBackend
already performs, so install and upgrade now leave identical on-disk
state with no orphaned files.

Reported as part of #10720.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-07 20:23:13 +00:00
LocalAI [bot]
8565febe45 fix(vllm): pin L4T arm64 backend to vllm==0.24.0 for GB10 stability (#10725)
The nvidia-l4t-cuda-13-arm64 vLLM backend left `vllm` unpinned, so the
prebuilt image drifted onto whatever aarch64 wheel was latest at build
time (0.23.x). On GB10 / DGX Spark (Grace Blackwell, unified memory),
0.23 crashes deterministically during cold model loads with an empty
"Engine core initialization failed" set and pins GPU memory until a host
reboot.

vLLM 0.24.0 carries vllm-project/vllm#45179 ("release cached device
memory under pressure on UMA GPUs during weight loading"), which the
reporter verified fixes the crash on GB10. Pin the L4T requirements to
0.24.0 to match the already-pinned cublas13 build
(requirements-cublas13-after.txt) and keep the image deterministic.

Editing this file also re-triggers the single-arch L4T image build via
the path filter, republishing the gallery image with 0.24.0 (the
single-arch matrix builds again after #10703).

Closes #10722


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-07 20:18:23 +00:00
Roman Mazurenko
a3fdfbc0d1 feat(llama-cpp): add device selection option (#10724)
Allow llama.cpp model configs to select the backend devices used for offload, matching upstream --device behavior so users can exclude a display or debug GPU.

Signed-off-by: rvmzes <rvmzes@rvmzess-MacBook-Pro.local>
Co-authored-by: rvmzes <rvmzes@rvmzess-MacBook-Pro.local>
2026-07-07 20:09:05 +00:00
Tai An
2f33cc7bc4 fix(vram): report largest GGUF quant instead of whole HF repo for gallery size (#10700) (#10707)
* fix(vram): report largest GGUF quant, not whole repo, for HF gallery size (#10700)

Signed-off-by: Tai An <antai12232931@outlook.com>

* test(vram): cover multi-GGUF quant repo size estimation (#10700)

Signed-off-by: Tai An <antai12232931@outlook.com>

---------

Signed-off-by: Tai An <antai12232931@outlook.com>
Co-authored-by: Ettore Di Giacinto <mudler@users.noreply.github.com>
2026-07-07 11:50:27 +00:00
LocalAI [bot]
22225217e0 docs: ⬆️ update docs version mudler/LocalAI (#10709)
⬆️ Update docs version mudler/LocalAI

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-07 08:55:01 +02:00
LocalAI [bot]
c1fd12a506 chore: ⬆️ Update ggml-org/llama.cpp to f36e5c348bc8795c34f9a038e58876e7a8423d4d (#10710)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-07 08:54:33 +02:00
LocalAI [bot]
d01b2c4f46 chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to 0725d2e53b7d2749a99ff33d3a460b954ffa7805 (#10711)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-07 08:41:46 +02:00
LocalAI [bot]
fb9ff061f1 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260706123649 (#10712)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-07 08:41:34 +02:00
LocalAI [bot]
40f847745e chore: ⬆️ Update ikawrakow/ik_llama.cpp to a8cf53fd69bada5450bd653eb0e32d1113fbd7fe (#10713)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-07 08:41:21 +02:00
LocalAI [bot]
ba9327b9f8 chore: ⬆️ Update CrispStrobe/CrispASR to 0a7643f1006c6bf2d1f37f5c63e9726f5eb4f364 (#10716)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-07 08:41:00 +02:00
LocalAI [bot]
fd467c5b3b chore: ⬆️ Update leejet/stable-diffusion.cpp to bb84971129d2a094ab8051c6feed5406d3b4409d (#10684)
* ⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(stablediffusion-ggml): pass chroma knobs via model_args after upstream API change

Upstream stable-diffusion.cpp bb849711 removed the dedicated
chroma_use_dit_mask / chroma_use_t5_mask / chroma_t5_mask_pad fields from
sd_ctx_params_t and now reads them from the generic model_args key=value
spec (parse_key_value_args). Assigning the old struct members broke the
gosd.cpp build. Emit the three options into model_args instead so the
existing chroma controls keep working. Verified by building
libgosd-fallback.so against the pinned upstream commit.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-06 23:57:50 +00:00
dependabot[bot]
fa0622604a chore(deps): bump actions/cache from 4 to 6 (#10704)
Bumps [actions/cache](https://github.com/actions/cache) from 4 to 6.
- [Release notes](https://github.com/actions/cache/releases)
- [Changelog](https://github.com/actions/cache/blob/main/RELEASES.md)
- [Commits](https://github.com/actions/cache/compare/v4...v6)

---
updated-dependencies:
- dependency-name: actions/cache
  dependency-version: '6'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-06 21:39:37 +02:00
Dennis Huang
ff5758113b chore(model gallery): add MiniCPM series models (#10699)
Add 9 MiniCPM models to the gallery:
- MiniCPM-V 4.6 (1.3B multimodal, edge-optimized)
- MiniCPM-V 4.6 Thinking (1.3B multimodal with reasoning)
- MiniCPM-V 4 (multimodal)
- MiniCPM-o 4.5 (8B omni-modal, vision+speech)
- MiniCPM-o 2.6 (7.6B omni-modal)
- MiniCPM5-1B (text)
- MiniCPM4.1-8B (text)
- MiniCPM4-8B (text)
- MiniCPM3-4B (text)

All sha256 checksums sourced from HuggingFace LFS metadata.

Signed-off-by: Dennis Huang <huangsiyuan20060408@hotmail.com>
2026-07-06 21:39:12 +02:00
LocalAI [bot]
29db4ab414 fix(ci): shard single-arch backend matrix under GitHub's 256-job limit (#10703)
GitHub Actions refuses to instantiate a matrix that would generate more
than 256 jobs. It does so silently: the job hangs forever at "Waiting for
pending jobs" and the whole run is marked `failure` while every other job
stays green. This is exactly what happened on the v4.6.1 tag build
(run 28786533892): the single-arch build matrix had grown to 268 entries,
so `backend-jobs-singlearch` (and its downstream merge) never produced a
single job, and the release build "failed" with no failing job to point at.

The single-arch list is the one that grows unbounded as backends are added,
so shard it across a fixed number of matrix jobs (SINGLEARCH_SHARDS=4,
~67 entries each today, headroom to ~1020 backends). Each merge shard
`needs:` only its matching build shard, preserving the "merge waits only on
its own build" property that keeps slow CUDA/ROCm builds from gating
multi-arch manifest assembly.

changed-backends.js now emits per-shard matrix/has-* outputs and throws
loudly if a shard ever reaches the 256 limit (telling the maintainer to
bump SINGLEARCH_SHARDS and add matching job blocks) instead of letting
GitHub drop the overflow silently. backend.yml and backend_pr.yml define
the four build + four merge shard jobs; multi-arch and darwin groups are
untouched.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-06 21:38:12 +02:00
weifanglab
a6cf67cc6b refactor: use slices.Contains to simplify code (#10702)
Signed-off-by: weifanglab <weifanglab@outlook.com>
2026-07-06 19:33:28 +02:00
LocalAI [bot]
85f5267ed2 fix(llama-cpp): cap single-pass embedding batch to fit VRAM (#10695)
* fix(llama-cpp): cap single-pass embedding batch to fit VRAM

Embedding/score/rerank all decode or pool the whole input in one physical
batch, so EffectiveBatchSize sized the batch to the full context window. For
a large context that makes n_ubatch huge, and the per-device CUDA compute
buffer (forward-graph scratch, ~n_ubatch * n_ctx, NOT split across GPUs)
balloons into multi-GiB: a large-context embedding model then aborts on load
(exitCode=-1) even with plenty of free VRAM. Reproduced with qwen3-embedding-4b
(context 40960 -> n_batch 40960 -> abort) and qwen3-embedding-0.6b
(n_batch 8192); pinning batch:512 avoided it.

This is the same root cause as issue #10485 (a large context turns the batch
into multi-GiB of scratch that must fit on a SINGLE card), but the single-pass
path bypassed the VRAM headroom guard the config layer already had — it
returned the unbounded context as the batch with no GPU awareness.

Make the single-pass batch VRAM-aware: cap it to the largest batch whose
compute buffer fits the per-device VRAM headroom, clamped to
[DefaultPhysicalBatch, ctx], reusing the existing computeBufferBytesPerCell and
headroom-divisor math (no duplication). Unknown per-device VRAM (0) stays
conservative (DefaultPhysicalBatch, not the context) so a detection gap can't
OOM. The GPU is resolved through an injectable package var (config.LocalGPU,
backed by sync.Once-cached xsysinfo detection) so the per-request router call
stays cheap and tests inject a deterministic device. Explicit batch: still
wins. An input longer than the cap can no longer be pooled in one pass — the
accepted tradeoff, since a batch that OOMs the device processes nothing.

Assisted-by: Claude:claude-opus-4-8 golangci-lint go-test
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): single-pass batch follows context on unknown VRAM

The single-pass (embedding/score/rerank) batch cap must only shrink the batch
when the per-device VRAM ceiling is KNOWN. On unknown VRAM (CPU-only or a GPU
detection gap) SinglePassBatchForContext returned DefaultPhysicalBatch, which
under-sized the batch below the context — over-trimming score/embed/rerank
inputs (the modelTokenTrim middleware regression) with no OOM benefit on CPU
where the compute buffer lives in system RAM. Return the full context instead,
preserving the original single-pass behavior; the VRAM cap stays a downward
safety that only engages when VRAM is known.

Assisted-by: Claude:claude-opus-4-8 [go-test go-vet]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-06 12:56:09 +02:00
LocalAI [bot]
ed3b59baf1 fix(config): cap auto-derived context to fit VRAM (#10696)
When a model is imported without an explicit context_size, the GGUF
importer defaulted the model's context to its full trained window
(n_ctx_train). For long-context models (128k / 256k / 1M) that KV cache
cannot fit a consumer GPU, so the backend aborts on load (exitCode=-1)
even though the model file is perfectly fine. Reproduced live:
gemma-4-26b-a4b-it-qat-q4_0 defaulted to context=262144 and
qwythos-9b-claude-mythos-5-1m to 1048576, both aborting on a 20 GB card.

Instead of chasing the trained max, auto-derive a conservative default:
min(trainedMax, DefaultAutoContextSize=8192). A small model keeps its
trained window; a long-context model caps at 8k and users opt into more
via context_size. This cap applies always, including CPU / unknown-VRAM
hosts, so it never regresses those paths.

Per-device VRAM is used only as a DOWNWARD safety: when a per-device
ceiling is detected (xsysinfo.MinPerGPUVRAM) and even the 8k cap would
not fit it with headroom, step down through candidate contexts to the
largest that fits, floored at DefaultContextSize. When VRAM is unknown
(0) or no GPU is detected we do NOT clamp — the bug is GPU OOM and the
8k cap is already safe, so detection gaps must not shrink the window.

The footprint estimate reuses gpustack/gguf-parser-go's
EstimateLLaMACppRun at a given context with all layers offloaded, taking
the per-device NonUMA VRAM figure. The estimate and VRAM detection are
package vars so tests inject deterministic values. Explicit context_size
always wins (guessGGUFFromFile only acts when it is nil).

Assisted-by: Claude:claude-opus-4-8 [golangci-lint go-test]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-06 12:53:45 +02:00
LocalAI [bot]
461ae84732 fix(startup): scope generated-content and upload dirs to the current user (#10698)
The `--generated-content-path` and `--upload-path` defaults were the fixed
shared locations `/tmp/generated/content` and `/tmp/localai/upload`. On any
multi-user host these collide across accounts: macOS routes `/tmp` to the
shared `/private/tmp` for every user, so whichever account starts LocalAI
first creates the parent with 0750 perms and every other account then fails
startup with:

    unable to create ImageDir: "mkdir /tmp/generated/content: permission denied"
    unable to create UploadDir: "mkdir /tmp/localai/upload: permission denied"

The same happens on Linux once a stale root-owned `/tmp/generated` (e.g. from
a prior `sudo` run) is left behind. This bites the desktop launcher and any
app embedding the raw binary (Wingman, nib-desktop), which start `local-ai
run` with no path flags.

Default both paths under the OS temp dir (`os.TempDir()`, honoring `$TMPDIR`;
already per-user on macOS) namespaced by the current UID
(`TMPDIR/localai-<uid>/...`), so accounts never collide while the paths stay
ephemeral. Wired via new kong vars in main.go so every consumer of the raw
binary inherits the fix. All content subdirs (audio, images) derive from
`GeneratedContentDir`, so they are fixed transitively.

As defense in depth, the launcher also anchors these two paths under its own
per-user data directory (mirroring the #10610 fix for data/config), extracted
into a testable `BuildRunArgs`.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-06 12:53:25 +02:00
Tai An
2a4426c5ec fix(reasoning): don't persist request-scoped reasoning_effort as an operator disable (#10622) (#10623)
* fix(reasoning): don't persist request-scoped reasoning_effort into model config

When a model sets `reasoning_effort: none` (or any default) in its YAML
without an explicit `reasoning.disable`, ApplyReasoningEffort resolves that
default at request time and sets ReasoningConfig.DisableReasoning on the
request-scoped config copy. The post-load thinking/marker probe then wrote
that request-scoped value back into the loader's persistent config via
UpdateModelConfig, making it look as though the operator had explicitly set
reasoning.disable=true. From then on, per-request `reasoning_effort` overrides
were silently ignored (an explicit operator disable wins over a request
asking to think).

DetectThinkingSupportFromBackend only fills reasoning slots that are still
nil, so a slot already set here came from ApplyReasoningEffort, not the probe.
Snapshot which slots were nil before the probe and only persist those, so the
probe's genuine backend detection is still saved while request-time reasoning
effort never leaks into the persistent config.

Fixes #10622

Signed-off-by: Tai An <antai12232931@outlook.com>

* test(reasoning): cover persist-guard added in this PR, extract for testability

ModelInference's post-probe persistence of ReasoningConfig.DisableReasoning /
DisableReasoningTagPrefill had no test: the guard logic lived inline in a
closure only reachable through a live gRPC backend. Extract it into
persistProbedReasoning (pure refactor, no behavior change) so it can be
exercised directly against a ModelConfigLoader, then add specs covering:

- a probe-filled slot (nil beforehand) gets persisted
- a slot that already carried a request-scoped value (e.g. from
  reasoning_effort: none) is left alone, i.e. the #10622 regression stays
  fixed
- an operator's explicit persisted disable is preserved when the guard is
  false
- the media marker still persists unconditionally

Verified red/green: reverting persistProbedReasoning to the old unconditional
copy fails exactly the two guard specs.

Assisted-by: Claude:claude-sonnet-5 go vet
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* test(reasoning): ignore os.Remove error in temp file cleanup (errcheck)

Signed-off-by: Tai An <antai12232931@outlook.com>

* chore: empty commit to re-trigger flaky Agent Jobs CI test

Signed-off-by: Tai An <antai12232931@outlook.com>

---------

Signed-off-by: Tai An <antai12232931@outlook.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@users.noreply.github.com>
2026-07-06 09:23:10 +02:00
LocalAI [bot]
2348bdc16d chore: ⬆️ Update ggml-org/llama.cpp to 2da668617612d2df773f966e3b0ee22dc2beef7b (#10694)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-06 01:46:47 +02:00
walcz-de
2ccc67bc7f feat(agents): native Prometheus metrics for agent chat runs (#10689)
Operators need a scrape-friendly signal for agent-turn health (completing,
erroring, cancelled, duration) — log-derived counters proved brittle (ANSI/
timezone parsing, restart gaps). Adds localai_agent_runs_total{agent,outcome}
and localai_agent_run_seconds histogram, recorded at the Chat() response
handoff (single choke point of the local execution path). Lazy meter init,
same pattern as the PII events counter (#10641).

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
2026-07-06 01:06:15 +02:00
LocalAI [bot]
0a6c62bb59 chore: ⬆️ Update ServeurpersoCom/qwentts.cpp to 73fe0c67bbf0898ba2999535e0680a02a7f8537d (#10683)
⬆️ Update ServeurpersoCom/qwentts.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-06 01:05:43 +02:00
LocalAI [bot]
1297356e29 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to daedb763fd442e0916eb130a479fdd74947291c0 (#10682)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-06 01:05:25 +02:00
LocalAI [bot]
3f36b1dbed chore: ⬆️ Update CrispStrobe/CrispASR to 09df654e304947f7521e1f52992ceacccf03c300 (#10693)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-06 00:32:28 +02:00
LocalAI [bot]
783222baf4 docs: ⬆️ update docs version mudler/LocalAI (#10680)
⬆️ Update docs version mudler/LocalAI

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-06 00:32:00 +02:00
LocalAI [bot]
bd3f2588fd fix(ui): center the home empty-state wizard (#10691)
The no-models getting-started wizard (`.home-wizard`) rendered
left-aligned instead of centered. `.home-page` is a column flexbox with
the default `align-items: stretch`; a child with `max-width: 48rem`
cannot be stretched past its max-width, so it falls back to the
cross-start (left) edge. The populated home branch never exposed this
because its children are full-width.

Add `margin: 0 auto` to `.home-wizard` so the max-width block centers
horizontally, for both the admin getting-started wizard and the
non-admin no-models hero.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-05 13:12:09 +02:00
LocalAI [bot]
40e659974d chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260704102955 (#10668)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-05 10:20:19 +02:00
LocalAI [bot]
deb43e56c0 chore(model-gallery): ⬆️ update checksum (#10686)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-05 10:20:02 +02:00
LocalAI [bot]
33869da527 chore: ⬆️ Update ggml-org/llama.cpp to 665892536dfb1b7532161e3182304bd35c33e768 (#10681)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-05 10:19:36 +02:00
LocalAI [bot]
8059117c2d chore: ⬆️ Update CrispStrobe/CrispASR to 1109cb3fcae2e242c2b3d42ec0e3fd6e813f2ce7 (#10685)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-05 09:29:18 +02:00
LocalAI [bot]
b0959d4756 feat(api): add GET /v1/models/capabilities endpoint (#10687)
Additive superset of /v1/models that enriches each model entry with the
capabilities it supports plus its input/output modalities
(text / image / audio / video). Clients that only understand /v1/models
are unaffected -- they simply never call the new route.

Audio and video *input* are derived from the model's multimodal limits
(vLLM limit_mm_per_prompt), which no single usecase FLAG expresses. That
gap is exactly why a plain capability list is insufficient and this
enriched endpoint exists: an attachment router can now decide whether an
image/audio/video file can go to the active model directly, or must be
converted/transcribed first.

Capability derivation lives in core/config as the single source of truth
(ModelConfig.Capabilities / InputModalities / OutputModalities /
VisionSupported / ...); the Ollama capability surface now delegates to
it instead of keeping a parallel copy. Vision is gated on
chat/completion capability so a MediaMarker hydrated onto a non-chat
model (e.g. a pure ASR/TTS backend) no longer reports a false vision
capability.

Read-only listing: no new FLAG_* flag, reuses the existing `models`
swagger tag, and intentionally exposes no MCP admin tool (there is
nothing to manage conversationally).

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-05 08:51:55 +02:00
LocalAI [bot]
9e41be4bfb fix(auth): log the real cause of OIDC/OAuth user-info failures (#10679)
The OAuth callback discarded the error returned by user-info resolution
before sending the generic 500, so real failures were completely opaque
in the logs: ID-token verification errors (e.g. issuer/audience mismatch
behind a reverse proxy), a missing id_token, claim-parse errors, or a
rejecting GitHub userinfo endpoint all collapsed into
"failed to fetch user info" with nothing logged.

Log the wrapped cause with xlog.Error (provider + error), matching the
code-exchange step just above it. The client-facing message is unchanged,
so no internal detail leaks to the browser.

Refs #10677


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-04 19:33:53 +02:00
LocalAI [bot]
38350d363e fix(backends): enable ROCm/HIP GPU offload for ggml audio backends (#10666) (#10667)
qwen3-tts-cpp, omnivoice-cpp, acestep-cpp and vibevoice-cpp shipped
rocm-* variants that silently ran on CPU ([Load] backend: CPU). Two
coupled defects:

- The Makefiles passed -DGGML_HIPBLAS=ON, but the vendored ggml only
  understands -DGGML_HIP=ON (GGML_HIPBLAS was removed upstream), so the
  ggml-hip backend target was never created and no GPU code was built.
- The CMake foreach that links the ggml GPU backends into the module
  listed blas/cuda/metal/vulkan but not hip, so even a built ggml-hip
  would not have been linked and its static backend registration would
  never run.

CUDA users were unaffected because cublas passes the correct GGML_CUDA=ON
and the foreach already links cuda. Mirror the proven llama-cpp hipblas
block (ROCm clang CC/CXX + AMDGPU_TARGETS) and add hip to each foreach.
Upstream picks the best device via ggml_backend_init_best(), so no
runtime flag is needed once HIP is compiled and linked.


Assisted-by: Claude:claude-opus-4-8[1m] [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-04 09:08:20 +02:00
LocalAI [bot]
817136c20e chore: ⬆️ Update CrispStrobe/CrispASR to f35185b876fc482fcb2053a81a2697936ed5fcc0 (#10670)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-04 08:17:02 +02:00
LocalAI [bot]
8396ce1388 chore: ⬆️ Update ggml-org/llama.cpp to d4cff114c0084f1fbc9b4c62717eca8fb2ae494a (#10671)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-04 08:16:41 +02:00
LocalAI [bot]
348f3c87c0 fix(gpu-libs): bundle hipBLASLt TensileLibrary data so ROCm backends stop falling back (#10660) (#10672) the
The ROCm packager copied rocBLAS kernel data (rocblas/library/*.dat) into the
bundled lib/ dir and run.sh pointed ROCBLAS_TENSILE_LIBPATH at it, but the
parallel hipBLASLt data dir (hipblaslt/library/TensileLibrary_lazy_gfx*.dat)
was never packaged and no HIPBLASLT_TENSILE_LIBPATH was set. The bundled
libhipblaslt.so therefore resolved its per-arch kernel data relative to itself,
found nothing, and silently fell back to slow generic kernels, logging:

    rocblaslt error: Cannot read "TensileLibrary_lazy_gfx1201.dat": No such file or directory
    rocblaslt error: Could not load "TensileLibrary_lazy_gfx1201.dat"

Fix, mirroring the existing rocBLAS handling:
- package-gpu-libs.sh: extract the rocblas data-dir copy into a reusable
  copy_rocm_data_dir helper and call it for both rocblas and hipblaslt.
- llama-cpp/turboquant run.sh: export HIPBLASLT_TENSILE_LIBPATH when the
  bundled hipblaslt/library dir exists.

The helper takes an optional ROCM_BASE_DIRS override so the copy is unit
testable without a real ROCm install; add a regression test that runs
package_rocm_libs against a fabricated ROCm tree and asserts both data dirs
are bundled.

Note: this bundles whatever gfx*.dat the build image's ROCm provides. If a
given arch's tensile data is absent from the shipped ROCm, that arch still
needs a ROCm bump; the packaging gap itself is fixed for every supported arch.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-04 08:14:12 +02:00
LocalAI [bot]
13310905a3 chore: ⬆️ Update ikawrakow/ik_llama.cpp to bbc7de475178dd0535c16ad85f204a2529806c9d (#10669)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-03 23:35:41 +02:00
LocalAI [bot]
2cbb3c96b3 fix(gallery): block SSRF in gallery config URL fetch (#10665) (#10673)
POST /models/apply with an empty "id" fetches the attacker-supplied
"url" gallery config directly via http.Client, with no check that the
URL resolves to a public IP. In the default Docker deployment no API key
is configured, so any network-reachable client can coerce LocalAI into
issuing requests to internal services or cloud-metadata endpoints (and
exfiltrate a small slice of the response through the job error message).

Guard the config fetch chokepoints (GetGalleryConfigFromURL and
GetGalleryConfigFromURLWithContext, which back both the /models/apply
worker and gallery installs) with utils.ValidateExternalURL, matching
the protection already applied to the CORS proxy and image/video/audio
download paths. Only plain http(s) URLs are validated; non-network
schemes (huggingface://, github:, oci://, ollama://, file://) resolve to
fixed public services or local files and are left untouched.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-03 21:32:42 +00:00
Ettore Di Giacinto
1152acc167 Revert "feat(config): default swa_full:true for sliding-window-attention models" (#10674)
Revert "feat(config): default swa_full:true for sliding-window-attention mode…"

This reverts commit 02b007a31e.
2026-07-03 22:46:44 +02:00
walcz-de
cc8ee62db0 feat(pii): export PII/audit events as a Prometheus counter (#10641)
The PII EventStore ring buffer is capacity-bound and meant for
recent-audit browsing via /api/pii/events; operators also want a
monotonic, scrape-friendly signal on /metrics — how many
detections/masks/blocks per hour, per origin, and whether the filter
stopped firing after a deploy (silent-failure class).

EventStore.Record is the single choke point every producer already goes
through (request middleware, response scrubbing, MITM proxy
connects/intercepts), so one lazily-initialised counter there covers all
paths without touching any producer:

  localai_pii_events_total{kind, origin, action, direction}

Same lazy otel.Meter pattern as core/services/routing/billing, so the
counter lands on the Prometheus-backed global MeterProvider installed by
the monitoring service. No behaviour change; label cardinality is
bounded (enum-like fields only, no pattern IDs or user IDs).

Assisted-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>
2026-07-03 20:36:15 +00:00
LocalAI [bot]
bfd6c09d88 chore(model gallery): 🤖 add 1 new models via gallery agent (#10663)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-03 18:02:09 +02:00
Richard Palethorpe
eb32cd9073 feat(realtime): eager blocking pipeline warm-up + /backend/load API (#10662)
Realtime sessions previously lazy-loaded each pipeline sub-model (VAD,
transcription, LLM, TTS) on first use, so every cold session paid a
per-request model-load stall and load errors only surfaced mid-stream.

Warm the whole pipeline eagerly and blockingly at session start
(including the voice-gate speaker-recognition model, which an enforced
gate blocks each utterance on; compaction's summary_model stays lazy
since it only runs off the response path):
- Add backend.PreloadModel / PreloadModelByName as the single load path
  for every modality (no transcription special-case; backend-omitted
  configs are deprecated).
- The realtime session blocks on Model.Warmup and returns a
  model_load_error to the client if any stage fails to load;
  updateSession warms in the background. Opt out per pipeline with
  pipeline.disable_warmup, exposed as a UI toggle via the
  config-metadata registry.

Add a LocalAI-native POST /backend/load (and /v1/backend/load) that
pre-loads a model -- expanding realtime pipelines into their sub-models
-- as the inverse of /backend/shutdown. There is one preload engine
(backend.PreloadStages): the realtime Warmup methods, /backend/load and
the --load-to-memory startup flag all use it, so --load-to-memory now
also expands pipeline models and records load-failure traces. Pipeline
sub-model alias resolution is likewise shared
(ModelConfigLoader.LoadResolvedModelConfig). Surface the endpoint
everywhere an admin manages models:
- MCP admin tool load_model (httpapi + inproc clients, safety/catalog
  prompts, catalog/dispatch tests).
- "Load into memory" action in the React models UI.
- Swagger regenerated; docs moved to the general backend-monitor page
  since it is not realtime-specific.

Fix a Traces UI crash ("json: unsupported value: -Inf"): audio-snippet
RMS/peak now floor at a finite dBFS, and backend-trace data is sanitized
to drop non-finite floats before marshaling. The sanitizer is
copy-on-write -- it runs on every RecordBackendTrace, so containers are
only re-allocated on the paths that actually changed.

Migrate core/http/openresponses_test.go onto the prebuilt mock-backend
the rest of the http suite already uses -- it was the last spec still
pointing at a real HuggingFace model, so it 404'd wherever no vision
backend was built -- and fix its item_reference specs to send the
spec's "id" field instead of "item_id", which the handler never
accepted.

Assisted-by: Claude:claude-opus-4-8 Claude Code

Signed-off-by: Richard Palethorpe <io@richiejp.com>
2026-07-03 18:00:37 +02:00
alaningtrump
80ec22945a refactor: use the built-in max/min to simplify the code (#10657)
Signed-off-by: alaningtrump <alaningtrump@outlook.com>
2026-07-03 17:59:26 +02:00
LocalAI [bot]
7a3583b52c fix(python-backends): parse tool-call arguments for chat templates and split implicit reasoning blocks (#10658)
Two bugs broke OpenAI-style tool calling on the MLX backend (and any
Python backend sharing backend/python/common), reproduced end-to-end on
LocalAI v4.5.5 with the metal-mlx backend and
mlx-community/Qwen3.5-2B-MLX-8bit.

messages_to_dicts left each tool call's function.arguments as the raw
OpenAI-wire JSON string. HuggingFace chat templates (e.g. Qwen3.5)
iterate arguments as a mapping (.items()), so any request whose history
contained a prior assistant tool_calls message failed with HTTP 500
"Generation failed: Can only get item pairs from a mapping." — breaking
every agent loop on its second turn. Decode the string back into a dict
so the template sees a mapping.

split_reasoning returned ("", text) whenever the opening think tag was
absent. Models like Qwen3.5 open the assistant turn already inside
thinking, so the generated text carries only the closing </think>; the
whole chain-of-thought leaked into content. When the opener is missing
but the closer is present, treat everything before the closer as
reasoning.

Adds platform-independent unit tests under backend/python/common
(stdlib-only, no MLX/venv required, following parent_watch_test.py).

Assisted-by: Claude Code:claude-opus-4-8

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-03 12:13:37 +02:00
LocalAI [bot]
715d4ed8e5 chore: ⬆️ Update ggml-org/llama.cpp to fdb1db877c526ec90f668eca1b858da5dba85560 (#10647)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-03 00:46:56 +02:00
LocalAI [bot]
9fcc9c0d43 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 87fc8701ff4da81a7d2a91ec0695f95eb3066a47 (#10649)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-03 00:46:41 +02:00
LocalAI [bot]
3c67b5b746 chore: ⬆️ Update CrispStrobe/CrispASR to 9a26976a8c8cf5af0afcdd04463cf8ba91e96a54 (#10648)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-03 00:46:25 +02:00
LocalAI [bot]
bea66fd84e chore: ⬆️ Update leejet/stable-diffusion.cpp to 2574f5936571645f784b77623e1f09bad97d948a (#10650)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-03 00:46:10 +02:00
LocalAI [bot]
f7a5dfd5ae chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260701212152 (#10646)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-03 00:45:36 +02:00
LocalAI [bot]
6bcaf30c14 chore: ⬆️ Update localai-org/privacy-filter.cpp to 735a6c28607ee82afc3a670383f41b55266a3b9a (#10628)
⬆️ Update localai-org/privacy-filter.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-03 00:45:17 +02:00
LocalAI [bot]
ef15b4bfda fix(vllm): install ROCm vLLM from the AMD wheel index on Python 3.12 (#10651)
* fix(vllm): install ROCm vLLM from the AMD wheel index on Python 3.12

The rocm-vllm backend crashed at load with "No module named 'vllm'".
requirements-hipblas-after.txt requested a bare `vllm`, which resolves to
the CUDA-only PyPI wheel; that wheel is unusable on an AMD GPU. vLLM's
prebuilt ROCm wheels live on a dedicated index (https://wheels.vllm.ai/rocm/)
and are published only for CPython 3.12, so on the backend's default 3.10
the installer silently falls back to the CUDA wheel.

Add a hipblas branch to backend/python/vllm/install.sh that pins Python to
3.12 and installs vllm from the ROCm wheel index, hiding the bare-`vllm`
after-file so installRequirements installs only the base ROCm
torch/transformers first and does not pull the CUDA wheel.

Fixes #10642

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* chore(vllm): drop the dead hipblas-after requirement and its hide dance

requirements-hipblas-after.txt (a bare `vllm`) is never installed for
hipblas: installRequirements only adds requirements-${BUILD_PROFILE}-after.txt
when BUILD_TYPE != BUILD_PROFILE, and for hipblas they are equal. So the file
was dead and the install.sh hide/restore of it was a no-op. Remove both. The
hipblas branch already installs vllm explicitly from the ROCm wheel index, so
deleting the bare-`vllm` file also removes a latent CUDA-wheel trap should the
installRequirements gap ever be closed.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-03 00:44:55 +02:00
LocalAI [bot]
237bce48e8 feat(ui): forking chat - retry any answer, copy, duplicate, branch (#10645) (#10654)
* feat(ui): clone a chat into a new conversation (#10645)

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): retry any assistant answer, not just the last (#10645)

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): copy an entire chat to the clipboard (#10645)

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(ui): branch a new chat from any assistant answer (#10645)

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(ui): send truncated history on mid-conversation retry (#10645)

Mid-conversation retry regenerated an answer with the downstream turns
still in the model's context. handleRegenerate truncated the DOM history
via updateChatSettings (a scheduled state update), but the synchronous
sendMessage that followed read the stale, pre-truncation history from its
closure to build the outbound API payload. Thread the intended base
history explicitly through sendMessage's options.baseHistory so the
request body matches the truncated view. Backward compatible: the normal
send path (no baseHistory) is unchanged.

Also guard two minor issues in Chat.jsx: the "Branch from here" button now
renders under !isStreaming to match the retry button, and the duplicate
toast only fires when forkChat returns a chat (not on a null result).

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-03 00:04:44 +02:00
LocalAI [bot]
a4e6e01e4d fix(process): give backend workers a parent-death safety net (#10639)
* fix(grpc): self-terminate backend workers when LocalAI dies non-gracefully

Symptom: a backend model-worker subprocess (the per-model gRPC server LocalAI
spawns) can be orphaned and linger — holding VRAM and its listen port — if the
LocalAI process is killed non-gracefully (e.g. a supervisor's graceful-shutdown
grace period elapses and LocalAI is SIGKILLed) before its own teardown runs.

Root cause: LocalAI's graceful teardown (pkg/signals/handler.go installs the
SIGINT/SIGTERM handler; core/cli/run.go registers app.Shutdown ->
ModelLoader.StopAllGRPC -> process.Stop in pkg/model/process.go) only runs when
LocalAI receives a catchable signal and survives long enough to run its
handlers. Backends are spawned via github.com/mudler/go-processmanager v0.1.1,
whose getSysProcAttr() sets Setpgid:true (own process group, so the group can be
signalled) but never PR_SET_PDEATHSIG/Pdeathsig, and exposes no Config field or
option for a caller to inject/extend SysProcAttr. LocalAI fully delegates
spawning to that library (it never builds the exec.Cmd itself), so it cannot set
a kernel parent-death signal at the spawn site. If LocalAI is SIGKILLed, nothing
tells the backend to exit and it is reparented to init.

Fix: add a best-effort, backend-side safety net at the one shared choke point
every out-of-process Go backend routes through — grpc.StartServer / RunServer in
pkg/grpc. On startup it captures getppid() and polls; when the process is
reparented (getppid changes / becomes 1 — the standard POSIX signal the original
parent died) it logs and self-terminates. getppid() reparent detection is
portable (Linux + macOS), unlike Linux-only PR_SET_PDEATHSIG. Toggle via
LOCALAI_BACKEND_PARENT_WATCH (default on; off on Windows) and
LOCALAI_BACKEND_PARENT_WATCH_INTERVAL. This is strictly a backstop alongside the
existing graceful SIGTERM->grace->SIGKILL teardown, which is unchanged.

Scope/limitations: covers Go-based backends (everything using pkg/grpc). The
C++ backends (e.g. llama-cpp) and Python backends do not route through
pkg/grpc and are not covered by this mechanism — they would each need an
equivalent parent-death check (follow-up). The fully general fix is for
go-processmanager to expose SysProcAttr injection so LocalAI can set Pdeathsig
at spawn for every backend regardless of language (suggested upstream follow-up;
out of scope for this LocalAI-only PR).

Test: pkg/grpc/parentwatch_test.go builds a real test -> middle -> grandchild
process tree, lets the middle process exit to orphan the grandchild running the
real watchParentDeath, and asserts it detects the reparent and self-terminates.
Unix-only (build-tagged), runs in CI (Linux).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(process): extend parent-death backstop to C++ and Python backends

The Go parent-death watcher (pkg/grpc/parentwatch.go, commit 772b435d5)
only protects backends that route through pkg/grpc. C++ and Python
backends don't, so the originally-reported case — the llama.cpp gRPC
worker surviving a non-graceful LocalAI death — was still uncovered.

Extend the same best-effort backstop to both languages, reusing the
exact mechanism and semantics:

- capture getppid() at startup, skip if already orphaned (<=1)
- a background thread polls getppid() and self-exits on reparenting
  (getppid() != orig || == 1), portable across Linux/macOS, no-op on
  Windows
- same env vars: LOCALAI_BACKEND_PARENT_WATCH (default on; falsy
  false/0/no/off disable) and LOCALAI_BACKEND_PARENT_WATCH_INTERVAL
  (default 2s; accepts Go-style durations like 500ms/2s/1m)

C++: implemented in backend/cpp/llama-cpp (the reported, most-used C++
backend) as a dependency-free header parent_watch.h, wired into
grpc-server.cpp's main() and copied at build time via prepare.sh. C++
backends have no shared server scaffolding, so other C++ backends
(ds4, ik-llama-cpp, privacy-filter, ...) are not yet covered and would
each need the same one-line include+call as follow-ups.

Python: implemented once in the shared common/parent_watch.py and armed
from common/grpc_auth.py's get_auth_interceptors() — the single helper
every one of the 35 Python backends invokes while building its gRPC
server — so all Python backends (and future ones) are covered with no
per-backend edits and no duplicated implementation.

Tests (real process-tree reparent detection, mirroring the Go test):
- backend/cpp/llama-cpp/parent_watch_test.cpp (via run-unit-tests.sh)
- backend/python/common/parent_watch_test.py (python -m unittest)

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-02 19:16:48 +02:00
LocalAI [bot]
6eea3ef2ac fix(backends): make backend install ops idempotent unless forced (#10643)
* fix(backends): make backend install ops idempotent unless forced

POST /backends/apply hardcoded force=true through
LocalBackendManager.InstallBackend, so applying an already-installed
backend re-downloaded and re-extracted the whole artifact every time.
API clients that ensure a backend exists at startup paid a full OCI
image pull on every boot.

Backend install ops now default to non-forced — an installed, runnable
backend short-circuits (the orphaned-meta reinstall path in
InstallBackendFromGallery is preserved) — and reinstall stays available:

- ManagementOp gains a Force field; the local manager passes it through
  instead of hardcoding true.
- /backends/apply accepts an optional "force" boolean in the body.
- The React UI install route keeps forcing, since its button doubles as
  the explicit "Reinstall backend" action.

Distributed installs already behaved this way (workers skip when the
binary exists unless force is set); this aligns single-node behavior.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(backends): don't force-reinstall LOCALAI_EXTERNAL_BACKENDS on boot

The startup loop for LOCALAI_EXTERNAL_BACKENDS runs
InstallExternalBackend for each listed backend on every boot, and its
gallery-name path hardcoded force=true — so every start re-downloaded
and re-extracted each listed backend's OCI image even when it was
installed and runnable. Supervising apps that list several backends
paid several full OCI pulls per launch.

Give InstallExternalBackend an explicit force parameter (it only
affects the gallery-name fallback; URI installs always write) and pass:

- false from the boot loop and `local-ai backends install` (idempotent
  ensure — `backends upgrade` is the refresh path),
- op.Force from the local manager's external-URI op,
- the request's force on the worker install path and true on its
  upgrade path (behavior unchanged).

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-02 19:16:29 +02:00
LocalAI [bot]
ad97bcbbdd chore(model gallery): 🤖 add 1 new models via gallery agent (#10644)
chore(model gallery): 🤖 add new models via gallery agent

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-02 19:16:09 +02:00
walcz-de
9d8ff90941 fix(cloud-proxy): parameter compatibility with newest reasoning models (#10640)
Newest cloud reasoning models reject two parameters the cloud-proxy
backend currently sends:

- Anthropic (claude-opus-4-x) and OpenAI (gpt-5.x) return 400 when
  temperature is present: "'temperature' is deprecated for this model".
  OpenAI-compatible clients typically send only the server-side DEFAULT
  sampling values rather than user intent, so the translators now forward
  neither temperature nor top_p and let the upstream apply its own
  defaults.
- OpenAI gpt-5.x rejects max_tokens ("Unsupported parameter: 'max_tokens'
  ... Use 'max_completion_tokens' instead"). The OpenAI translator now
  serializes the token limit as max_completion_tokens, which current
  chat-completions models accept.

Verified live against claude-opus-4-8, gpt-5.5 and gemini-3.1-pro
(Gemini OpenAI-compat endpoint). Tests updated to the new contract.

Assisted-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>
2026-07-02 19:15:43 +02:00
LocalAI [bot]
29001a88c1 fix(distributed): don't let a dead worker pin the model-load advisory lock (#10600)
* fix(distributed): don't let a dead worker pin the model-load advisory lock

In distributed mode a chat request could fail with:

  failed to route model with internal loader: routing model ...:
  loading model ...: advisorylock: acquiring lock <id>:
  ERROR: canceling statement due to lock timeout (SQLSTATE 55P03)

Root cause is two independent defects in the cross-replica model-load path:

1. SmartRouter.Route holds a per-model PostgreSQL advisory lock for the whole
   cold-load sequence, which includes installBackendOnNode -> InstallBackend,
   a NATS request-reply with a 15m deadline (DefaultBackendInstallTimeout) that
   ignored ctx. When the chosen worker died mid-install, the holder sat on the
   lock for up to 15m. The detached loadCtx (WithoutCancel) had no deadline, so
   nothing capped the hold.

2. The acquiring statement, pg_advisory_lock(), is subject to any deployment
   global lock_timeout. A common operator setting (e.g. 10s) aborts the wait
   with SQLSTATE 55P03, so every other replica's request for that model hard
   -errored instead of waiting for the in-progress load and reusing it. For the
   ~15m window the model was effectively unroutable.

Fixes:

- advisorylock.WithLockCtx (postgres): SET lock_timeout = 0 on its dedicated
  connection (RESET before it returns to the pool) so the Go context, not a
  deployment-wide GUC, governs how long we wait. Waiters now block and then
  re-check, reusing the model another replica just loaded.

- SmartRouter: bound the detached loadCtx with a single ModelLoadCeiling so the
  lock is always released in bounded time even if a sub-step wedges. Default is
  the configured backend.install deadline + 10m (staging + LoadModel margin),
  so a legitimately slow load is never cut.

- installBackendOnNode: use singleflight.DoChan + select on ctx.Done() so the
  install wait honors cancellation; the ceiling can then actually free a caller
  pinned behind a dead worker. The shared install still coalesces via
  singleflight.

Reproduced both defects as failing tests first (a real 55P03 against a
testcontainer with a short lock_timeout; a wedged install that blocks Route)
and confirmed green.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

* fix(distributed): bound advisory-lock wait instead of disabling lock_timeout

Setting lock_timeout = 0 to override a deployment's short global lock_timeout
meant "wait forever" server-side. Safe for SmartRouter.Route (its loadCtx now
carries the model-load ceiling) but unsafe for the schema-migration callers
that pass context.Background(): a holder whose session never releases would
hang them indefinitely.

Derive the server-side lock_timeout from the caller's context instead: its
remaining budget plus a margin (so the Go context's cancellation still wins
with a clean error and the server bound is only a backstop), or a finite
30m backstop when the context has no deadline. Never zero - "wait forever"
is no longer possible, while a deployment's hostile short lock_timeout is
still overridden so legitimate cross-replica waits don't fail with 55P03.

Added a spec proving a deadline-less waiter gives up at the (shrunk) backstop
rather than hanging.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@users.noreply.github.com>
2026-07-02 09:52:51 +02:00
LocalAI [bot]
b0bfa0852e chore: ⬆️ Update CrispStrobe/CrispASR to fcbc8718e654995e3bd2d0c98bcb8e55e297d23c (#10634)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-02 09:48:20 +02:00
LocalAI [bot]
39a93e91cf chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260701132215 (#10633)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-02 09:48:08 +02:00
LocalAI [bot]
26e0c98967 chore: ⬆️ Update leejet/stable-diffusion.cpp to 3590aa8d626e671a1b1dc84506ea2932a243a480 (#10631)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-02 09:47:54 +02:00
LocalAI [bot]
9acca54b25 chore: ⬆️ Update mudler/parakeet.cpp to e8acc6172a94e20a952cf1843decace5d771a94b (#10629)
⬆️ Update mudler/parakeet.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-02 09:47:41 +02:00
LocalAI [bot]
2728e6000e chore: ⬆️ Update ikawrakow/ik_llama.cpp to 068b173649f2fd8dc96b35ada5a0b76d8985105d (#10632)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-02 09:47:28 +02:00
LocalAI [bot]
006310d746 chore: ⬆️ Update ggml-org/llama.cpp to 4fc4ec5541b243957ae5099edb67372f8f3b550e (#10630)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-02 09:47:15 +02:00
LocalAI [bot]
05acdb1778 chore: ⬆️ Update ggml-org/whisper.cpp to 6fc7c33b4c3a2cec83e4b65abd5e96a890480375 (#10635)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-02 09:47:01 +02:00
LocalAI [bot]
5e68b5700c chore(model-gallery): ⬆️ update checksum (#10637)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-02 09:26:32 +02:00
pos-ei-don
7910018249 fix(vllm): non-streaming tool-call regression after #10351 (#10638)
fix(vllm): non-streaming tool-call regression after #10351 (native_streaming is a capability flag, not a state flag)

#10351 introduced native streaming via `parser.extract_tool_calls_streaming`
and gated the post-loop `extract_tool_calls` block on `native_streaming and
not native_streaming_error`. That works for streaming requests, but for
non-streaming requests the same flag is still True (it only means "the
parser can stream", not "we actually streamed"), so the block was skipped
and the `elif` cleared `content = ""` — the tool call was silently lost.

Symptom: non-streaming chat.completions with `tools=[...]` returns
`finish_reason: "stop"` with `content: ""` and no `tool_calls`. Streaming
requests are unaffected.

Fix: gate both branches on `streaming` too, so the extract_tool_calls
block runs for non-streaming requests (and for streaming requests that
fell back to the buffered path).

Reproduction (vLLM 0.24, Qwen3-Coder-Next-NVFP4, qwen3_coder parser):

    curl -s -X POST http://localhost:8080/v1/chat/completions \
      -H 'Content-Type: application/json' \
      -d '{"model":"coder","stream":false,
           "messages":[{"role":"user","content":"7*8 via calc"}],
           "tools":[{"type":"function","function":{"name":"calc",
             "parameters":{"type":"object",
               "properties":{"expression":{"type":"string"}}}}}]}'

Before: finish_reason: "stop", content: "", tool_calls: []
After:  finish_reason: "tool_calls", tool_calls[0].function.name: "calc"

Streaming path re-verified in the same setup: delta.tool_calls arrives
token-by-token, finish_reason: "tool_calls", no raw XML in content.

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
2026-07-02 09:26:14 +02:00
LocalAI [bot]
1a03712a6f fix(hipblas): symlink amdgpu.ids so ROCm backends find the ASIC ID table (#10627)
* fix(hipblas): symlink amdgpu.ids so ROCm backends find the ASIC ID table

ROCm's bundled libdrm_amdgpu looks up the GPU ASIC ID table at a
hardcoded fallback path, /opt/amdgpu/share/libdrm/amdgpu.ids, which is
only populated by AMD's full amdgpu-install (graphics/DKMS) stack. The
hipblas image is compute-only and doesn't have it, so every model load
logs "No such file or directory" and the GPU can't be identified.
Symlink it to the equivalent file already shipped by Ubuntu's
libdrm-amdgpu1 package.

Fixes #10624

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(hipblas): correct amdgpu.ids source package name in comment

Verified against the real rocm/dev-ubuntu-24.04:7.2.1 image with
hipblas-dev/hipblaslt-dev/rocblas-dev installed: /usr/share/libdrm/amdgpu.ids
is owned by libdrm-common, not libdrm-amdgpu1 as the comment said.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-02 09:25:14 +02:00
LocalAI [bot]
703ea32de6 chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260630095652 (#10616)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-01 21:56:59 +02:00
LocalAI [bot]
751db06e35 chore: ⬆️ Update CrispStrobe/CrispASR to 8fd9db8fec8cb5e929d23d3267ed5817794feb1a (#10615)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-01 21:56:41 +02:00
LocalAI [bot]
f46c0e9c83 docs: ⬆️ update docs version mudler/LocalAI (#10614)
⬆️ Update docs version mudler/LocalAI

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-01 21:56:21 +02:00
LocalAI [bot]
0d8adfc59a chore: ⬆️ Update ggml-org/llama.cpp to 0eca4d490e591d4e93058d07540cf47278a72577 (#10617)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-01 09:31:50 +02:00
LocalAI [bot]
43f2615e19 chore: ⬆️ Update vllm-project/vllm cu130 wheel to 0.24.0 (#10618)
⬆️ Update vllm-project/vllm cu130 wheel

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-01 08:53:03 +02:00
LocalAI [bot]
875c539ad5 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 29431b31c89e79c10f8736e8f2742485ba1713d6 (#10620)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-01 08:52:36 +02:00
LocalAI [bot]
d641ded194 chore: ⬆️ Update ggml-org/whisper.cpp to 0874de3e8e8e48361dba85c7fe6d176f008bf158 (#10621)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-07-01 08:43:40 +02:00
LocalAI [bot]
40445fff05 chore: ⬆️ Update leejet/stable-diffusion.cpp to 484baa41e5e006c52dcd4addc38c830b9489745f (#10619)
* ⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(stablediffusion-ggml): adapt to new generate_image() out-param signature

leejet/stable-diffusion.cpp@484baa4 changed generate_image() from
returning sd_image_t* to returning bool with images_out/num_images_out
out-parameters (same pattern already used by generate_video()).

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-01 08:32:57 +02:00
Tai An
057dee956a fix(launcher): keep data/config under ~/.localai (#10610) (#10613)
The launcher starts the server with run --models-path/--backends-path but
leaves --data-path and the dynamic config dir unset, so the server falls
back to its /data and /configuration defaults.
 is kong.ExpandPath("."), i.e. the launcher process CWD
(commonly the user's home root), producing ~/data and ~/configuration
outside ~/.localai and an agent-pool stateDir under ~/data.

Pass --data-path and --localai-config-dir explicitly, rooted at the
launcher's own data directory (GetDataPath() -> ~/.localai), so data and
config stay consistent with --models-path/--backends-path.
2026-06-30 22:14:59 +02:00
Adira
4ec39bb776 fix(watchdog): don't log optional Free() as an error when backend returns Unimplemented (#10602) (#10607)
* fix(watchdog): don't log optional Free() as an error when backend returns Unimplemented (#10602)

When the watchdog evicts a model, deleteProcess calls the backend's gRPC
Free() to release VRAM before stopping the process. Free is optional:
backends that don't override it -- the generated UnimplementedBackendServer
stub, many Python/external backends, or a federation proxy in distributed
mode -- return gRPC Unimplemented. That is expected, not a failure: VRAM is
reclaimed when the local process is stopped, or by the remote unloader for
remote backends. Logging it as "WARN Error freeing GPU resources" made a
benign, optional RPC look like a fault (the alarming line in #10602, seen
in distributed mode where the model is remote and Free hits a stub).

Treat gRPC Unimplemented from Free() as a no-op logged at Debug; genuine
failures still Warn. Free() is still attempted for every backend, so any
backend that does implement it is unaffected.

Add a reusable grpcerrors.IsUnimplemented helper following the package's
existing code-based detection idiom (prefer the typed status code, fall
back to the message across non-gRPC boundaries), with table tests.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com>

* fix(watchdog): log a non-Unimplemented Free() failure at error level

Per review: now that the expected gRPC Unimplemented case is split out and
logged at Debug, any remaining Free() error is a genuine failure to release
VRAM, so surface it at error level instead of warn.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com>

---------

Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com>
2026-06-30 22:14:01 +02:00
Ettore Di Giacinto
25ecb9f015 fix(gallery): use Q8_0 for lfm2.5-8b-a1b to fix poor tool-call quality
The Q4_K_M quant degraded tool-call reliability for LFM2.5-8B-A1B.
Switch the gallery entry to the Q8_0 GGUF (sha256 verified via HF
x-linked-etag) while keeping the native jinja tool-parsing config.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]
2026-06-30 17:46:20 +00:00
LocalAI [bot]
2be495f9c0 fix(kokoros): implement AudioTranscriptionLive trait stub (#10612)
The backend.proto AudioTranscriptionLive bidirectional streaming RPC added
new required trait items (AudioTranscriptionLiveStream + audio_transcription_live)
on the generated Backend trait. The kokoros (TTS) backend did not implement
them, breaking its release build with E0046 (missing trait items).

kokoros is text-to-speech and has no live-ASR support, so stub the method to
return UNIMPLEMENTED, mirroring the existing audio_transcription_stream stub.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-30 19:38:41 +02:00
LocalAI [bot]
02b007a31e feat(config): default swa_full:true for sliding-window-attention models (#10611)
LocalAI enables a cross-request prompt-prefix cache (cache_reuse, see
core/config/serving_defaults.go) so repeated prefixes — system prompts,
RAG context, agent scaffolds, multi-turn chat — are not reprocessed every
turn. For sliding-window-attention (SWA) models (Gemma 2/3, Cohere2,
Llama 4, ...) this silently does nothing: llama.cpp defaults to a reduced
SWA KV cache sized to the sliding window, and that reduced cache cannot
preserve a prompt prefix across requests, so every turn reprocesses the
whole prompt anyway.

llama.cpp's --swa-full (params.swa_full, already wired through the
LocalAI llama.cpp backend's `swa_full` option) keeps the full KV cache so
the shared prefix is reused. Enable it automatically, but only for models
that are actually SWA: detection reads the gguf-parser-normalized
`<arch>.attention.sliding_window` metadata (which also applies llama.cpp's
family rules, e.g. Phi-3 → not SWA), right where the GGUF is already
parsed for defaults. It is never applied to dense models (pure memory
waste) and never overrides an explicit user `swa_full`/`n_swa` choice.

Tradeoff: the full SWA cache scales with context_size, so it costs more
memory at large contexts — hence the SWA gating and the documented
`swa_full:false` opt-out.

Assisted-by: Claude:claude-opus-4-8 [Claude Code] golangci-lint

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-30 17:58:17 +02:00
LocalAI [bot]
fd8cebd0b3 fix(watchdog): persist UI-saved Check Interval across restarts (#10601) (#10605)
fix(watchdog): persist a UI-saved Check Interval across restarts (#10601)

The watchdog Check Interval saved via /api/settings reverted to 500ms on
every restart, while the idle/busy timeouts persisted correctly.

Root cause: NewApplicationConfig baseline-defaulted WatchDogInterval to
500ms, whereas the idle/busy timeouts default to 0. The startup loader
(loadRuntimeSettingsFromFile) applies a persisted runtime_settings.json
value only when the field is still at its zero default - its heuristic
for "this wasn't set by an env var". Because the interval was always
500ms at that point, the loader never read the persisted value back, so
the saved interval was silently discarded on each boot.

Fix: drop the non-zero baseline default so the interval behaves like the
sibling timeouts (0 = unset). The effective 500ms default is now supplied
at the watchdog layer: WithWatchdogInterval ignores a non-positive value
so DefaultWatchDogOptions' 500ms is preserved (and a 0 interval can never
turn the watchdog loop into a busy spin). Also mirror the interval in the
live config file watcher alongside idle/busy, and report the real 500ms
default (not the stale "2s") from ToRuntimeSettings.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-30 17:48:14 +02:00
LocalAI [bot]
dd625921ff fix(macos): staple the notarization ticket to the .app, not just the dmg (#10606)
Stapling only the dmg leaves the LocalAI.app bundle with no embedded
notarization ticket. Gatekeeper then falls back to an online notarization
check on first launch, so the app fails to open on a Mac that is offline or
behind a firewall, or once it has been copied out of the dmg — while it keeps
working on the (online) build host, which masks the problem.

Notarize and staple the .app before packaging it into the dmg so the bundle
verifies offline. Adds a `notarize-app` subcommand to
contrib/macos/sign-and-notarize.sh (zips the bundle for notarytool, then
staples + validates) and invokes it from dmg-launcher-darwin. Stays a no-op
when notary secrets are unset, so unsigned local/fork builds are unaffected.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: mudler <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-30 17:38:47 +02:00
LocalAI [bot]
d74f88357e fix(tests): align openresponses test model name with GGUF-derived naming (#10589) (#10609)
PR #10589 changed repo-root HuggingFace URI imports to name the model after
the selected GGUF file rather than the repository. The Open Responses API
integration test still requested the old repo-derived name
("Qwen3-VL-2B-Instruct-GGUF"), so every request 404'd on an unknown model and
the suite has failed on master since 1a4f68ed4.

Update testModel to the name the importer now registers for the default
q4_k_m quant ("Qwen3-VL-2B-Instruct-Q4_K_M") so the specs resolve the model
again. The #10589 behaviour change is intentional; only the stale test needed
updating.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-30 15:41:44 +02:00
Adira
dfaec3bd51 fix(import): strip file:// scheme from model path for local imports (#10599)
Importing a model from a local directory (e.g. a HuggingFace checkout or an
LM Studio store) via a file:// URI produced a config whose model field kept
the scheme verbatim, e.g. model: file:///Users/u/.../Qwen3-4bit. The mlx and
vllm backends treat that field as a HuggingFace repo id or local path and
reject the file:// form with "Repo id must be in the form 'repo_name' or
'namespace/repo_name'", so the model imported fine but failed to load (issue
#7461).

Add a shared LocalModelPath helper that reduces a file:// URI to the bare
filesystem path it points at and leaves HuggingFace/HTTP URIs untouched, and
route the mlx, vllm, transformers and diffusers importers (all of which pass
details.URI straight into the model field for from_pretrained-style loading)
through it.

Cover the helper directly plus end-to-end file:// import specs for the mlx and
vllm importers.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com>
2026-06-30 10:21:08 +02:00
LocalAI [bot]
0e381897b5 chore: ⬆️ Update ikawrakow/ik_llama.cpp to f74a6fb87b315b2c3154166e075360e15021a61d (#10598)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-30 09:17:48 +02:00
LocalAI [bot]
b1af37257d chore: ⬆️ Update CrispStrobe/CrispASR to 3b93758f9725d400eca82976f895e4cec3f31260 (#10597)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-30 09:17:11 +02:00
LocalAI [bot]
ebefa6dcca chore: ⬆️ Update localai-org/privacy-filter.cpp to 595f59630c69d361b5196f2aba2c71c873d0c13c (#10596)
⬆️ Update localai-org/privacy-filter.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-30 09:16:52 +02:00
LocalAI [bot]
605348925d chore: ⬆️ Update ggml-org/llama.cpp to 6f4f53f2b7da54fcdbbecaaa734337c337ad6176 (#10595)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-30 09:16:37 +02:00
LocalAI [bot]
686ce10b54 chore: ⬆️ Update leejet/stable-diffusion.cpp to 3b6c9ca97cfcda8e68e719e6670d06379fcbe943 (#10594)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-30 09:16:21 +02:00
pos-ei-don
2cee318fad fix(functions): avoid quadratic-time debug logging in CleanupLLMResult / ParseFunctionCall (#10592)
fix(functions): avoid quadratic-time debug logging in CleanupLLMResult/ParseFunctionCall

The streaming chat path (core/http/endpoints/openai/chat_stream_workers.go)
calls CleanupLLMResult / ParseFunctionCall once per delta chunk with the
*full accumulated* LLM result so far. Both functions xlog.Debug the entire
argument on entry and exit, so a single N-chunk stream emits roughly
chunk_size * N^2 bytes of debug output.

Under LOG_LEVEL=debug this was observed in a recent SGLang-via-LocalAI
session on a DGX Spark host (about 50K tokens, long streaming generation)
to drive container logs to ~96 GiB, which interacted with the streaming
hot loop on the same filesystem and contributed to a host-wide hard hang
once disk pressure built up. Workaround was setting LOG_LEVEL=info, but
the quadratic shape remains a foot-gun for anyone intentionally enabling
debug.

Replace the four result-content debug arguments with len(...) plus a
fixed-size head (200 bytes via a new truncForLog helper), bounding per-
call output to a constant. The debug signal stays useful: the first 200
chars are enough to identify which generation is in flight, and the
length lets you observe growth without paying for the payload itself.

No API change. No behaviour change for LOG_LEVEL != debug.

Signed-off-by: Poseidon <philipp.wacker@ibf-solutions.com>
Co-authored-by: Poseidon <philipp.wacker@ibf-solutions.com>
2026-06-30 09:16:03 +02:00
Adira
1a4f68ed4a fix(import): derive model name from selected GGUF for repo-root URIs (#10589)
When importing a HuggingFace GGUF model from a repository-root URI (no file
component, e.g. hf://owner/repo) with the Model Name field left blank, the
importer named the model after the repository (filepath.Base(details.URI))
instead of the GGUF file it actually selected from the repo listing (issue
#10587).

Track whether the user supplied an explicit name; the URI base is now only a
fallback. In the HuggingFace branch, once the model group is picked, re-derive
the name from the selected GGUF via a new modelNameFromShardGroup helper that
uses ShardGroup.Base minus the .gguf extension. For sharded models this yields
a clean logical name (e.g. Qwen3-30B-A3B-Q4_K_M) rather than a shard filename
like ...-00001-of-00002. An explicit name preference still always wins, and the
.gguf/URL/OCI paths are unchanged.

Add network-free unit specs covering name-from-GGUF, clean-name-from-shard-base,
and explicit-name precedence, and update the live integration specs that had
encoded the previous repo-name behaviour.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Adira Denis Muhando <dennisadira@gmail.com>
2026-06-30 09:03:27 +02:00
Adira
28d7397743 fix(openai): stop max_tokens streaming retry loop on reasoning models (#9716) (#10448)
fix(openai): stop max_tokens streaming retry loop on reasoning models

When a thinking model spends its entire max_tokens budget on the reasoning
block, the C++ autoparser clears the raw Response and delivers reasoning-only
ChatDeltas (no content, no tool calls). ComputeChoices' empty-response retry
then fires and regenerates from scratch up to maxRetries times, each
re-consuming the whole budget, instead of terminating with finish_reason
"length" (issue #9716).

Add a reachedTokenBudget helper and suppress both the built-in and
caller-driven retries when the completion count has reached the configured
max_tokens ceiling. Report finish_reason "length" instead of "stop" in the
streaming and non-streaming chat paths when the budget was exhausted.

Adds a deterministic regression test that counts backend invocations
(previously 6, now 1) plus boundary tests for the helper.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Dennisadira <dennisadira@gmail.com>
2026-06-30 09:01:53 +02:00
Richard Palethorpe
5d0c43ec6e feat(realtime): Semantic VAD EOU token (#10444)
* feat(realtime): EOU-driven semantic_vad turn detection

Add a `semantic_vad` turn-detection mode to the realtime API that feeds
the transcription model live and decides "the user finished speaking"
from the `<EOU>` end-of-utterance token rather than from silence alone.
When EOU fires the turn commits immediately (~0.3s); otherwise it falls
back to an eagerness-scaled silence threshold (low/med/high = 8/4/2s).

Plumbing, bottom to top:

- proto: `AudioTranscriptionLive` bidirectional RPC (config-first oneof,
  mono float PCM @16k, ready-ack / Unimplemented degrade signal) plus
  `TranscriptResult.eou` for the unary retranscribe gate.
- pkg/grpc: client/server/base/embed scaffolding for the bidi stream,
  modeled on AudioTransformStream; release stream conns on terminal Recv.
- parakeet-cpp: live transcription RPC with per-C-call engine locking
  (one live stream per turn, finalize+free at commit); bump parakeet.cpp
  to ABI v5 — incremental StreamingMel (no more quadratic per-feed mel
  recompute that delayed EOU on long turns) and the <EOU>/<EOB> split;
  strip the literal <EOU>/<EOB> from offline text and set Eou.
- core/backend: LiveTranscriptionSession wrapper + pipeline
  `turn_detection:` config block (type/eagerness/retranscribe).
- realtime: semantic_vad integration — live input captions streamed as
  transcription deltas while the user speaks, EOU-immediate commit with
  eagerness fallback, optional retranscribe gate (batch re-decode must
  also end in <EOU> to confirm), clause synthesis off the LLM token
  callback, and per-turn live-transcription / model_load telemetry.
- UI: show the realtime pipeline components as a vertical list.

Docs and tests included; opt-in via the pipeline YAML or per-session
`session.update`. Non-streaming STT backends degrade to silence-only.

Assisted-by: Claude Code:claude-opus-4-8 [Read] [Edit] [Write] [Bash]
Assisted-by: Claude Code:claude-fable-5 [Read] [Edit] [Bash]
Signed-off-by: Richard Palethorpe <io@richiejp.com>

* feat(realtime): explicit formally-verified state machines + parakeet streaming driver

The realtime API had several implicit state machines whose state was inferred
from scattered booleans, channels, and five separate mutexes, leaving
illegal/inconsistent states reachable. Make them explicit and keep the
implementation in step with a formal design; rework the parakeet streaming
backend along the same lines.

Realtime state machines (M1-M5). Each is a sealed sum-type State/Event/Effect
with a total, pure Next(state,event)->(state,[]effect) behind a single-writer
Coordinator:

  M1 conncoord    connection lifecycle: VAD toggle + once-only teardown
                  (replaces vadServerStarted + a `done` channel closed from
                  two sites).
  M2 turncoord    turn detection: collapses speechStarted and the live-stream
                  "turn open" flag into one state, so discardTurn can no longer
                  desync them and suppress the next onset.
  M3 respcoord    response coordination: serializes the dual-writer
                  start/cancel so at most one response is live; one
                  response.done per response.create.
  M4 compactcoord conversation compaction: single-flight (replaces the
                  `compacting atomic.Bool` CAS).
  M5 ttscoord     TTS pipeline: open->closing->closed, idempotent wait(),
                  rejects enqueue-after-close (was a silent drop).

The Coordinator/Sink/Next plumbing — only the sealed types and Next differed
per machine — is extracted once into core/http/endpoints/openai/coordinator as
a generic Coordinator[S,E,F]; each machine keeps its public API via type
aliases, so no sink, call-site, or test moved.

Hierarchy. session_lifecycle.fizz models M1 as the parent region with its
children (M2/M3/M4) as one statechart and asserts ChildrenDieWithParent (conn
torn => all children terminal, none start after teardown). respcoord and
compactcoord gain an absorbing Terminated state + Shutdown event; conncoord's
teardown drives the children terminal. This closes a compaction teardown gap: a
fire-and-forget compaction could outlive a torn session — compactionSink now
takes a session-scoped cancellable context + WaitGroup and joins the in-flight
summarize+evict on shutdown.

Formal verification. formal-verification/ holds one authoritative FizzBee spec
per machine plus the composition spec, each with an always-assertion and a
documented one-line edit that makes the checker fail (verified non-vacuous).
scripts/realtime-conformance.sh is fail-closed: all Go conformance suites under
-race AND a model-check of every .fizz spec; a missing FizzBee is a hard error
(only the loud REALTIME_CONFORMANCE_SKIP_FIZZBEE=1 bypasses it, never in CI).
FizzBee is pinned by sha256 and installed via scripts/install-fizzbee.sh into
.tools/ (gitignored). Wired as make test-realtime-conformance, a CI workflow,
and a pre-commit path filter. Go conformance tests are Ginkgo/Gomega (per the
repo's forbidigo lint): transition tables + fixed-seed property walks +
concurrent/-race specs, no rapid dependency. Design map:
docs/design/realtime-state-machines.md.

Parakeet streaming backend. The same treatment applied to the parakeet-cpp
streaming paths:
- AudioTranscriptionStream returns codes.Unimplemented for non-streaming models
  instead of decoding offline and emitting it as one delta + final. A client
  that asked for streaming learns the model cannot stream rather than receiving
  a batch result shaped like a stream. New grpcerrors.StreamTranscriptionUnsupported
  carries that signal; the HTTP /v1/audio/transcriptions stream path surfaces it
  as an SSE error event. Mirrors AudioTranscriptionLive, which already did this.
- utteranceBoundary (boundary.go): a single definition of the end-of-utterance
  latch, replacing three open-coded finalEou toggles. Modelled as a two-valued
  type so illegal states are unrepresentable.
- Shared decode driver (driver.go): streamFeedResult (one per-feed event) +
  feedChunk (hides the ABI v4 JSON vs text-only split) + feedSlices + flushTail.
  The feed loop is written once.
- AudioTranscriptionLive becomes a bidi adapter: it streams the per-feed
  {delta,eou,eob,words} the realtime turn detector consumes and a terminal
  FinalResult carrying only Text. Segments/duration/eou are offline-only and no
  longer produced (nor read) on the live path; liveTraceState drops the terminal
  eou and keeps the per-feed eou_events count.
- AudioTranscriptionStream + streamJSON merge into one driver-based function;
  streamSegmenter is generalized to the unified event with a text-only fallback
  that preserves the legacy (no-words) library's per-utterance segmentation.

Verified: build/vet/gofumpt clean, golangci-lint 0 issues, all coordinator and
parakeet packages under -race, the fail-closed conformance gate green, and
make test-realtime (12 e2e WS+WebRTC).

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Richard Palethorpe <io@richiejp.com>

---------

Signed-off-by: Richard Palethorpe <io@richiejp.com>
2026-06-30 09:01:22 +02:00
pos-ei-don
6ab29ec8b9 fix(sglang): parse tool_call function arguments before applying the chat template (#10558)
OpenAI wire format carries `function.arguments` as a JSON-encoded string,
but chat templates (e.g. Qwen3-Coder) iterate over it as a mapping. The
vllm backend already parses arguments before applying the chat template
(PR #10256); this mirrors that fix in the sglang backend.

Without this fix the second turn of any tool-using session (assistant
returns tool_calls, user posts `role:"tool"` result, model is invoked
with arguments still as a string) crashes inside transformers' Jinja
chat-template rendering with:

  TypeError: Can only get item pairs from a mapping.
  File ".../transformers/utils/chat_template_utils.py", in render_jinja_template
  File ".../jinja2/filters.py", in do_items
      raise TypeError("Can only get item pairs from a mapping.")

Reproduced on `lmsysorg/sglang:v0.5.14` via LocalAI v4.5.4 with
`saricles/Qwen3-Coder-Next-NVFP4-GB10` (W4A4 NVFP4 / compressed-tensors)
on NVIDIA DGX Spark (GB10, sm_121).

After the patch, a tool-call roundtrip (assistant tool_calls -> tool
result -> assistant final answer) returns http=200 with the expected
follow-up content; no behaviour change on requests that don't carry
tool_calls.

Signed-off-by: Poseidon <philipp.wacker@ibf-solutions.com>
Co-authored-by: Poseidon <philipp.wacker@ibf-solutions.com>
2026-06-30 09:00:51 +02:00
dependabot[bot]
036f950b1b chore(deps): bump actions/cache from 4 to 6 (#10593)
Bumps [actions/cache](https://github.com/actions/cache) from 4 to 6.
- [Release notes](https://github.com/actions/cache/releases)
- [Changelog](https://github.com/actions/cache/blob/main/RELEASES.md)
- [Commits](https://github.com/actions/cache/compare/v4...v6)

---
updated-dependencies:
- dependency-name: actions/cache
  dependency-version: '6'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-29 22:31:10 +02:00
LocalAI [bot]
5b7b914b4f chore(recon): re-pin voice/face-detect to squashed release commits (+ graph-cache fix) (#10591)
chore(recon): re-pin voice/face-detect to squashed release commits

The voice-detect.cpp and face-detect.cpp engine repos were squashed to a single
release commit, which orphaned the previous pins (voice 3d51077, face 06914b0).
Re-pin to the new single-commit SHAs (voice 1db1759, face e22260d).

These also fold in a real correctness fix: the persistent graph-cache fingerprint
now includes op_params, so two structurally identical GGML_OP_CUSTOM graphs (a
blocked 3x3 vs a blocked 1x1 strided conv) can no longer false-hit the cache and
replay the wrong kernel. voice CI was failing test_blocked/conv1x1_s2 with an
out-of-bounds write on the GGML_NATIVE=OFF build; both engine repos are now green
and WeSpeaker embed parity is 1.0 vs golden.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-29 18:48:47 +02:00
LocalAI [bot]
d1cee4c52a chore: ⬆️ Update vllm-metal (darwin) to v0.3.0.dev20260628073537 (#10562)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-29 09:13:22 +02:00
LocalAI [bot]
baaa0fe94f chore: ⬆️ Update mudler/face-detect.cpp to 06914b077d52f90d5421299138e7be6bdd06b5e8 (#10580)
⬆️ Update mudler/face-detect.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-29 08:04:22 +02:00
LocalAI [bot]
c3b5c7c3fa chore: ⬆️ Update mudler/voice-detect.cpp to 3d510772357538c5182808ac7de2278b84824e24 (#10581)
⬆️ Update mudler/voice-detect.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-29 08:03:43 +02:00
LocalAI [bot]
bd1ec8f2c2 chore: ⬆️ Update ggml-org/llama.cpp to dbdaece23de9ac63f2e7ca9e6bfcdc4fc156a3fa (#10582)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-29 08:03:20 +02:00
LocalAI [bot]
135debf9af chore: ⬆️ Update CrispStrobe/CrispASR to 6b50f76e59700665358a1aabf5295597fa318e06 (#10583)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-29 08:03:06 +02:00
LocalAI [bot]
e8c18ae28e chore: ⬆️ Update leejet/stable-diffusion.cpp to c1790754d31bec0731ed5fddc9d5b9ff22ee19cd (#10584)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-29 08:02:52 +02:00
LocalAI [bot]
c4d302e1ab chore(model-gallery): ⬆️ update checksum (#10585)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-06-28 23:26:28 +02:00
LocalAI [bot]
323b57a4bc fix(oci): retry layer downloads on transient network errors (#10579)
Installing large backend images (e.g. vLLM/vLLM-omni, several GiB) over
the Web UI could fail with "failed to download layer 0: unexpected EOF"
when a single connection to the registry dropped mid-stream. The whole
install then failed with no recovery, and since the download is not
resumable, retrying from the UI restarted from zero and usually hit the
same blip again - so users saw it as a consistent, size-correlated
failure (issue #10577).

The registry transport already retries manifest/digest fetches via
defaultRetryPredicate (GetImage/GetImageDigest), but the per-layer data
stream in DownloadOCIImageTar bypassed it entirely: layer.Compressed()
+ xio.Copy ran exactly once.

Extract the per-layer copy into downloadLayerToFile, which retries on the
same transient errors (unexpected EOF, EOF, EPIPE, ECONNRESET, connection
refused) with exponential backoff, truncating any partial data before
each retry. Non-retryable errors and context cancellation still fail
fast.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-28 21:21:08 +02:00
LocalAI [bot]
3d2f639213 fix(fish-speech): allow invalid_reference_casting so tokenizers builds on darwin (#10573)
On darwin arm64 the fish-speech editable install (pip install
--no-build-isolation -e) compiles the transitive `tokenizers` Python
package's Rust extension from source, because there is no prebuilt
manylinux wheel for that platform (Linux builds never compile it, so this
only breaks on macOS). The pinned tokenizers crate fish-speech's stack
resolves to contains a `&T` -> `&mut T` cast that the macOS CI runner's
newer Rust toolchain rejects via the now-deny-by-default
`invalid_reference_casting` lint:

    error: casting `&T` to `&mut T` is undefined behavior ...
    error: could not compile `tokenizers` (lib) due to 1 previous error
    ERROR: Failed building wheel for tokenizers

This failed the fish-speech darwin/metal (mps) backend image build in the
v4.5.5 release CI while all Linux variants built fine.

Fix: export RUSTFLAGS with `-A invalid_reference_casting` (appended to any
existing value, not clobbering) before installRequirements so the
unchanged third-party crate compiles as it did under the older toolchain.
Version-agnostic and harmless on Linux, where no Rust compile happens.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-28 19:10:27 +02:00
Nicholas Ciechanowski
be1ae9338b fix(distributed): missing agent NATS permissions (#10571)
Signed-off-by: Nicholas Ciechanowski <nicholas@ciech.anow.ski>
2026-06-28 12:58:13 +02:00
LocalAI [bot]
923c47020d fix(launcher): robust binary download/upgrade (resume, rate-limit, UX) (#10575)
* fix(launcher): resume flaky downloads, drop redundant percent, fit dialogs

The binary upgrade/download flow had three rough edges:

- The status label printed "Downloading... N%" right next to a progress
  bar already showing the percent. Replace it with a human-readable byte
  readout ("Downloading... 12.3 MB / 45.6 MB").
- A failed download (GitHub releases are flaky) had no recourse and always
  restarted from byte 0. Stream to "<dest>.part" and resume via a
  "Range: bytes=N-" request (handling 206/200/416), renaming to the final
  path only after checksum verification; on checksum failure the file is
  discarded so the next attempt starts clean. Add a Retry button that
  appears on failure and resumes from the partial file.
- Progress/install dialogs were hardcoded to oversized dimensions, leaving
  a blank gap below "View Release Notes". Size each window to its content
  with a sane minimum width.

Also unify the three near-identical download-progress popups into one
Launcher.showDownloadProgressWindow helper (and delete a dead unused copy
in ui.go) so the behaviour stays consistent across every entry point.

The progress callback now reports (downloaded, total) byte counts instead
of a single fraction. Resume/retry behaviour is covered by httptest-backed
unit tests in release_manager_test.go.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(launcher): resolve latest version via redirect to dodge GitHub API 403

On a fresh Linux start with no LocalAI installed, the download failed with
"failed to fetch latest release: status 403". The cause is the unauthenticated
api.github.com rate limit (60 requests/hour, per IP): on shared/NAT/CGNAT/cloud
addresses it is exhausted almost immediately and every request 403s.

Resolve the latest version by following the github.com "releases/latest"
redirect instead, reading the tag from the final ".../releases/tag/<tag>" URL.
That endpoint is not subject to the API rate limit. Only the version is ever
consumed by callers, so the tag is sufficient. The JSON API is kept as a
fallback, now honoring GITHUB_TOKEN and reporting rate-limit 403/429 clearly
instead of an opaque status code.

Covered by an httptest-backed unit test that asserts the redirect path is used.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-28 12:57:32 +02:00
LocalAI [bot]
b7a1dec773 fix(kokoro): add explicit click dep so spacy CLI works on intel build (#10572)
The kokoro install.sh ends with `python -m spacy download en_core_web_sm`.
spaCy's CLI imports typer -> click, so click must be present at that point.

On the intel build profile, install.sh adds `--upgrade --index-strategy=unsafe-first-match`
against the Intel pip index. With that resolution strategy, click is not
resolved/installed, so the spacy CLI import fails with:

    ModuleNotFoundError: No module named 'click'
    make: *** [Makefile:3: kokoro] Error 1

Other profiles (cpu/cublas) pull click in transitively and build fine; only
the intel profile breaks. This surfaced in the v4.5.5 release CI as the
gpu-intel-kokoro backend image build failure.

Make click an explicit dependency in the base requirements.txt (installed for
every profile) so it is always present before `python -m spacy download` runs,
regardless of index resolution. Unpinned: spacy constrains the version.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-28 11:29:17 +02:00
1071 changed files with 83240 additions and 56225 deletions

View File

@@ -34,7 +34,7 @@ The build matrix is data-only YAML at `.github/backend-matrix.yml` (not inside `
**Without an entry here no image is ever built or pushed, and the gallery entry in `backend/index.yaml` will point at a tag that does not exist.** The `dockerfile:` field must point at `./backend/Dockerfile.<lang>` matching the language bucket from step 1 (e.g. `Dockerfile.python`, `Dockerfile.golang`, `Dockerfile.rust`). The `tag-suffix` must match the `uri:` in the corresponding `backend/index.yaml` image entry exactly.
**`scripts/changed-backends.js` registration — REQUIRED for any new dockerfile suffix.** This is the single most common omission, because it has no effect on the PR that adds the backend (when no prior path filter could catch it anyway) — it only breaks the *next* PR that touches your backend's directory, which then gets zero CI jobs and looks broken for unrelated reasons. Edit `scripts/changed-backends.js:inferBackendPath` and add a branch BEFORE the more-generic suffixes:
**Path-filter registration — REQUIRED for any new dockerfile suffix.** This is the single most common omission, because it has no effect on the PR that adds the backend (when no prior path filter could catch it anyway) — it only breaks the *next* PR that touches your backend's directory, which then gets zero CI jobs and looks broken for unrelated reasons. Edit `scripts/lib/backend-filter.mjs:inferBackendPath` and add a branch BEFORE the more-generic suffixes:
```js
if (item.dockerfile.endsWith("<your-dockerfile-suffix>")) {
@@ -54,7 +54,9 @@ for (const e of m.include.filter(e => e.backend === '<your-backend>')) {
}"
```
A quick way to find the right insertion point: `grep -n 'item.dockerfile.endsWith' scripts/changed-backends.js`.
A quick way to find the right insertion point: `grep -n 'item.dockerfile.endsWith' scripts/lib/backend-filter.mjs`.
If your backend consumes a *shared* build input that lives outside its own directory (a new script under `scripts/build/`, a new file copied into every image), add a rule to `SHARED_BUILD_INPUTS` in the same file — the per-backend prefix match cannot see those, and a miss ships your change to no image at all. See `scripts/lib/backend-filter_test.mjs` for the pattern; `make test-ci-scripts` runs it.
**`bump_deps.yaml` registration — REQUIRED for any backend pinning an upstream commit.** If your backend's Makefile has a `*_VERSION?=<sha>` pin to a third-party repo, the daily auto-bump bot at `.github/workflows/bump_deps.yaml` won't notice it unless you register the backend in its matrix. The bot runs `.github/bump_deps.sh` which `grep`s for `^$VAR?=` in the Makefile you list — so the pin MUST live in the Makefile (not in a separate shell script). The bump for ds4 (#9761) had to walk this back because the original landed the pin in `prepare.sh`, which the bot can't see. Pattern (for `antirez/ds4`):
@@ -115,7 +117,7 @@ Wiring a backend into `includeDarwin:` is more than the matrix entry:
1. **`includeDarwin:` entry** — `tag-suffix: "-metal-darwin-arm64-<backend>"`, `build-type: "metal"`, `lang: "go"` for go+ggml backends; omit `build-type` for the bespoke C++ ones (llama-cpp / ds4 / privacy-filter). Match an existing entry of the same shape.
2. **`backend/index.yaml`** — add `metal:` to the backend's `capabilities` map (main and `-development`) and concrete `metal-<backend>` / `metal-<backend>-development` image entries pointing at the `-metal-darwin-arm64-<backend>` images.
3. **C/C++ backends only** — add an `inferBackendPathDarwin` case in `scripts/changed-backends.js` returning `backend/cpp/<backend>/` (the generic fallthrough assumes `backend/<lang>/`, which is wrong for a C++ source tree driven with `lang: go`), and give `run.sh` a Darwin branch that exports `DYLD_LIBRARY_PATH` instead of `LD_LIBRARY_PATH`. If the build is bespoke (single `grpc-server` + dylib bundling), model it on `scripts/build/ds4-darwin.sh` and add a `backends/<backend>-darwin` make target plus a gated step in `.github/workflows/backend_build_darwin.yml`.
3. **C/C++ backends only** — add an `inferBackendPathDarwin` case in `scripts/lib/backend-filter.mjs` returning `backend/cpp/<backend>/` (the generic fallthrough assumes `backend/<lang>/`, which is wrong for a C++ source tree driven with `lang: go`), and give `run.sh` a Darwin branch that exports `DYLD_LIBRARY_PATH` instead of `LD_LIBRARY_PATH`. If the build is bespoke (single `grpc-server` + dylib bundling), model it on `scripts/build/ds4-darwin.sh` and add a `backends/<backend>-darwin` make target plus a gated step in `.github/workflows/backend_build_darwin.yml`.
4. **C++ proto gotcha** — if the backend compiles the generated gRPC/protobuf in a separate CMake target (e.g. `hw_grpc_proto`), that target must link `protobuf::libprotobuf` + `gRPC::grpc++` so the Homebrew include dirs propagate; otherwise macOS fails with `google/protobuf/runtime_version.h not found` (Linux hides this because apt headers sit in `/usr/include`).
The CI path filter only builds a backend on a PR when a file under its directory changes, so a darwin-only YAML edit builds nothing — touch a file under `backend/<lang>/<backend>/` (a one-line comment is enough) in the same PR.
@@ -216,6 +218,69 @@ docker-build-backends: ... docker-build-<backend-name>
- If the backend is in `backend/python/<backend-name>/` but uses `.` as context in the workflow file, use `.` context
- Check similar backends to determine the correct context
## Engine preference for gallery model variants
A gallery entry can declare `variants`, alternative builds of the same weights,
and LocalAI picks one per host: it drops builds whose backend cannot run here or
that do not fit memory, then ranks the survivors by **engine preference
first, serving feature second, size third** (`SelectVariant` in
`core/gallery/resolve_variant.go`).
Ask whether your backend should outrank another one on some hardware. If it
should, add it to `engineNamePreferenceRules` in `pkg/system/capabilities.go`,
best engine first for that capability:
```go
{Nvidia, []string{engineVLLM, engineSGLang, engineLlamaCpp}},
+ {Nvidia, []string{engineVLLM, engineSGLang, engineMyEngine, engineLlamaCpp}},
```
That is the ENGINE NAME table, matched as a substring of a gallery entry's
`backend:` value. Two sibling tables in the same file speak different
vocabularies and are matched against different things:
| Table | Vocabulary | Matched against | Consumer |
|-------|-----------|-----------------|----------|
| `backendBuildTagPreferenceRules` | build tags (`cuda`, `rocm`, `metal`) | installed build directory names, as a substring | alias resolution in `ListSystemBackends` |
| `engineNamePreferenceRules` | engine names (`vllm`, `llama-cpp`, `mlx`) | a gallery entry's `backend:`, as a substring | gallery variant ranking |
| `servingFeaturePreferenceTokens` | serving features (`dflash`, `mtp`) | a gallery entry's `tags:`, compared whole and case-insensitively, and nothing else | gallery variant ranking, one rank below the engine |
**Putting a token in the wrong table matches nothing and does not error**: every
candidate scores equal and the next sort key decides, so the preference silently
stops existing. The block comment above all three tables spells the contract out.
The serving feature table is the odd one: it is not keyed by capability, because
no hardware prefers a plain build over an equivalent faster build of the same
weights. It reads a declared tag and nothing else. The entry name was the
original signal and is gone: a naming convention is not a contract, and names
are author-supplied free text where a short marker like `mtp` turns up inside
unrelated words or on weights whose entry enables nothing.
`overrides.options` was rejected for the mirror-image reason: `spec_type:` is
llama.cpp's config vocabulary, whereas a cross-backend ranking decision must
work the same for `ds4`'s `mtp_path:` and `sglang`'s `speculative_algorithm:`.
**If your backend can serve the same weights faster** (speculative decoding,
multi-token prediction), say so in the docs for its gallery entries so curators
tag them: the tagging rule and the per-backend evidence table live in
[adding-gallery-models.md](adding-gallery-models.md). A backend never needs to
appear in the token table itself; it ranks builds, not engines.
Leaving your backend out is a valid choice when no ordering can be justified for
it. It then ranks below every known engine and selection falls back to size,
which is the behaviour that predates preference.
**Leaving a whole capability out is not.** A missing row gives that host an
empty preference list, so size alone decides among everything that survives the
filters, and the filter will not save you: `IsBackendCompatible` derives hardware
support from the engine NAME, so `vllm` and `sglang` carry no darwin, cuda, rocm
or sycl token and are never dropped on a host with no GPU. That is why `default`
(no usable accelerator, including a GPU under the 4 GiB VRAM floor) and
`darwin-x86` both have rows putting `llama-cpp` first. Every capability
`getSystemCapabilities()` can return needs a row unless every engine really is
equally at home there. When you add one, enumerate the engines you are demoting
rather than relying on them falling through unmatched: unmatched engines all tie
with each other, so size decides among them.
## Documenting the backend (README + docs)
A backend is not "added" until it is discoverable. Update the user-facing docs:
@@ -243,7 +308,7 @@ After adding a new backend, verify:
- [ ] Backend directory structure is complete with all necessary files
- [ ] Build configurations added to `.github/backend-matrix.yml` for all desired platforms (per-arch entries with `platform-tag` for multi-arch; `builder-base-image` for llama-cpp / ik-llama-cpp / turboquant)
- [ ] **OS coverage considered**: added to `includeDarwin:` (macOS/Apple Silicon) if the backend can build there — with the `backend/index.yaml` `metal:` capability + `metal-<backend>` image entries, a `run.sh` Darwin/DYLD branch and `inferBackendPathDarwin` case for C++ backends — or the PR explains why an OS is unsupported. Do not ship Linux-only by default.
- [ ] **OS coverage considered**: added to `includeDarwin:` (macOS/Apple Silicon) if the backend can build there — with the `backend/index.yaml` `metal:` capability + `metal-<backend>` image entries, a `run.sh` Darwin/DYLD branch and `inferBackendPathDarwin` case (in `scripts/lib/backend-filter.mjs`) for C++ backends — or the PR explains why an OS is unsupported. Do not ship Linux-only by default.
- [ ] Meta definition added to `backend/index.yaml` in the `## metas` section
- [ ] Image entries added to `backend/index.yaml` for all build variants (latest + development)
- [ ] Tag suffixes match between workflow file and index.yaml
@@ -251,6 +316,8 @@ After adding a new backend, verify:
- [ ] No YAML syntax errors (check with linter)
- [ ] No Makefile syntax errors (check with linter)
- [ ] Follows the same pattern as similar backends (e.g., if it's a transcription backend, follow `faster-whisper` pattern)
- [ ] **`Load` validates its input and refuses models it can't serve.** When a model config has no explicit `backend:`, the model loader greedily probes *every* installed backend with the model's name and binds to the first `Load` that succeeds — an accept-anything `Load` will capture arbitrary LLMs (issue #9287). Backends that load a real artefact get this for free (the load fails); backends with no artefact must gate on the name: `opus` accepts only its own name (or none), `local-store` requires the `store.NamespacePrefix` namespace marker sent by `core/backend/stores.go`.
- [ ] **Gallery variant ranking considered**: if this backend should be preferred over another on some hardware, it is listed in `engineNamePreferenceRules` (NOT `backendBuildTagPreferenceRules`, NOT `servingFeaturePreferenceTokens`) in `pkg/system/capabilities.go`. A missing entry silently ranks it last and lets the next sort key decide.
- [ ] Documented: added to the category list in `docs/content/features/backends.md` (and any new endpoint/realtime capability documented under `docs/content/`)
- [ ] If it is an in-house native C/C++/GGML engine, added to the maintained-engines table in the top-level `README.md`

View File

@@ -91,6 +91,108 @@ To add a variant (e.g., different quantization), use YAML merge:
uri: huggingface://<gguf-org>/<gguf-repo>/<filename>-Q8_0.gguf
```
## Offering several builds of one model (`variants`)
When the same model is published in more than one quantization, or is also
servable by another engine, add each build as its own ordinary gallery entry and
then point one of them at the others with `variants`:
```yaml
- !!merge <<: *chatml
name: "nanbeige4.1-3b-q4"
# ... the usual urls / overrides / files for the Q4 build ...
variants:
- model: nanbeige4.1-3b-q8
```
Rules:
- The declaring entry is a **complete, normal entry**. It keeps its own
`files`/`overrides` and stays installable on every host and by every older
LocalAI release, which simply ignore `variants`.
- A variant references another gallery entry **by name**. That entry must exist
and must not declare `variants` of its own.
- **A referenced entry keeps its own gallery row by default.** It is hidden only
in the collapsed listing (`collapse_variants=true`, which the web UI requests
by default), where the declaring entry stands in for it. Searching there still
matches the referenced entry and answers with the entry declaring it, so
referencing an entry never makes it unfindable; turning the collapse off
returns it under its own name.
- **Order carries no meaning.** Do not try to encode a preference; write the
list in whatever order reads best.
- **A variant may be smaller than the declaring entry.** Offering a downgrade
for small hosts is a normal shape: the declaring entry's own build competes
like every other candidate, so a large host keeps the large build.
- **Do not describe hardware.** At install time LocalAI drops variants whose
backend cannot run on the host, then drops those that do not fit available
memory. The declaring entry's own build is exempt from both filters, so
selection always terminates on something installable. Sizes are measured live
from the weights and cached, so nothing has to be written down.
- **Engine preference outranks size.** Among the builds that survive the
filters, the host's preferred engine wins first and only then does the larger
footprint win. On NVIDIA a vLLM build beats a larger llama.cpp one; on Apple
silicon an MLX build beats a larger GGUF one; on a host with no preference for
either engine the larger build wins, since a bigger footprint is a higher
quality quantization of the same weights. Predict what a user gets by asking
which engine the host prefers before asking which build is biggest. The
per-capability order lives in `engineNamePreferenceRules`
(`pkg/system/capabilities.go`); see
[adding-backends.md](adding-backends.md) for how a backend gets into it.
- **Serving feature preference sits between engine and size.** Among builds on
an equally preferred engine, one that speculates or predicts several tokens
per step beats the plain build of the same weights, because it answers faster
for the same output: a `dflash` build beats an `mtp` one, and either beats a
plain build. The order lives in `servingFeaturePreferenceTokens`
(`pkg/system/capabilities.go`) and is matched against the entry's `tags:` and
**nothing else**: not the entry name, not `overrides.options`. See
[the tagging rule](#the-dflash--mtp-tagging-rule) below. Engine deliberately
outranks it: a serving feature makes the right engine faster, it does not make
a wrong engine right. Fit still outranks both, so a drafter pairing (strictly
larger than the plain build, since it ships a drafter alongside it) is dropped
on a host too small for it before this order is ever consulted.
- A variant is nothing but a name; there is no per-variant memory field. When
the measured size for a build is wrong, correct it on the referenced entry by
setting that entry's own `size:` (e.g. `size: "20GiB"`). The estimator prefers
a declared size over its own guesswork, so the fix applies everywhere the size
is shown or compared rather than only to variant selection.
Users can override the automatic choice with `variant` on `POST /models/apply`,
`local-ai models install --variant`, or the `install_model` MCP tool. See
`docs/content/features/model-gallery.md`.
The gallery lint specs live in `core/gallery`, so run that suite after adding a
`variants` list.
### The `dflash` / `mtp` tagging rule
**Tag an entry `dflash` or `mtp` when the entry actually configures that
feature. Variant ranking reads the tag and nothing else.**
Decide by looking at what the entry configures, in whatever vocabulary its
backend uses:
| Backend | Configures the feature when it declares |
|---------|------------------------------------------|
| `llama-cpp` | `overrides.options` contains `spec_type:draft-dflash` or `spec_type:draft-mtp` |
| `ds4` | `overrides.options` contains `mtp_path:` / `mtp_draft:` |
| `sglang` | the referenced `gallery/*.yaml` sets `speculative_algorithm:` |
That check is curation-time only. `spec_type` is llama.cpp's config vocabulary,
and a cross-backend ranking decision must not depend on one backend's option
syntax, which is precisely why the ranker reads the tag instead of the options.
Two mistakes the rule exists to prevent:
- **Weights that carry the heads are not an entry that enables them.** The
NVFP4 GGUF entries ship MTP-bearing weights but set only `use_jinja:true`, so
they enable no speculative decoding and must NOT be tagged. Tagging them wins
them the feature axis without being any faster.
- **A name is not a declaration.** An entry whose name spells `-mtp` while
configuring nothing gets no tag, and an entry that configures the feature is
tagged even when its name says nothing (`hy3`, `glm-5.2`). Ranking never reads
the name, so an untagged build that does enable the feature is simply ranked
as plain rather than promoted on a marker nobody meant.
## Available template configs
Look at existing `.yaml` files in `gallery/` to find the right prompt template for your model architecture:

View File

@@ -114,6 +114,24 @@ Both `backend.yml` (push) and `backend_pr.yml` (PR) generate their matrix dynami
- **Tag pushes**: `FORCE_ALL=true` is set from the workflow side (`startsWith(github.ref, 'refs/tags/')`) — releases rebuild every backend regardless of diff.
- **Schedule / `workflow_dispatch`**: no `event.before`, falls through to "run everything" automatically.
### Shared build inputs
The per-backend prefix match only sees files under a backend's own directory, so a change to shared build infrastructure would rebuild *nothing* — an empty matrix, every job green, and the change reaching no image. That silently un-shipped PR #10946 (a partial-cuDNN packaging fix in `scripts/build/package-gpu-libs.sh`), which merged 1h48m after the weekly cron and so sat unbuilt for a week.
`SHARED_BUILD_INPUTS` in `scripts/lib/backend-filter.mjs` closes that hole. Each rule maps a shared path to the narrowest set of matrix entries it can honestly invalidate, since a full matrix is 417 Linux + 56 Darwin builds:
| Changed path | Rebuilds |
|---|---|
| `backend/backend.proto` | everything (all languages compile or copy it) |
| `backend/Dockerfile.<x>` | the Linux entries whose `dockerfile:` names it |
| `backend/python/common/` | Python, Linux + Darwin |
| `scripts/build/package-gpu-libs.sh` | Python, Linux only |
| `scripts/build/<lang>-darwin.sh` | the Darwin entries that build target routes to |
| `.github/workflows/backend_build[_darwin].yml` | everything on that OS |
| anything else under `scripts/build/` (except `*_test.sh`) | everything — conservative default for unclassified packaging inputs |
Deliberately excluded: `backend/index.yaml` (gallery metadata, never enters an image), `.github/backend-matrix.yml` (adding a backend would rebuild all of them), `backend/Dockerfile.base-grpc-builder` (owned by `base-images.yml`), and the root `Makefile` (touched in ~11% of commits, and its backend-relevant edits arrive alongside the backend directory anyway). `make test-ci-scripts` pins all of this.
The Sunday 06:00 UTC cron on `backend.yml` exists specifically because path filtering can leave Python backends frozen on stale wheels. `DEPS_REFRESH` (below) only fires when the build actually runs, so an untouched Python backend would never re-resolve its unpinned deps. The weekly cron is the safety net.
## The `DEPS_REFRESH` cache-buster (Python backends)

View File

@@ -65,6 +65,7 @@ This is enforced by `forbidigo` (see `.golangci.yml`): `http.DefaultClient` and
The project documentation is located in `docs/content`. When adding new features or changing existing functionality, it is crucial to update the documentation to reflect these changes. This helps users understand how to use the new capabilities and ensures the documentation stays relevant.
- **Docs-with-code rule**: When you change user-facing behavior (API endpoints, CLI flags, config keys, or features), update the corresponding page under `docs/content/` in the SAME change, not as a follow-up. A user-facing change without a matching docs update is incomplete. The PR template carries a checklist item for this.
- **Feature Documentation**: If you add a new feature (like a new backend or API endpoint), create a new markdown file in `docs/content/features/` explaining what it is, how to configure it, and how to use it.
- **Configuration**: If you modify configuration options, update the relevant sections in `docs/content/`.
- **Examples**: providing concrete examples (like YAML configuration blocks) is highly encouraged to help users get started quickly.

View File

@@ -1,143 +0,0 @@
# llama-cpp-localai-paged Backend (paged attention + Blackwell NVFP4 decode)
`llama-cpp-localai-paged` is LocalAI's **CUDA-only** paged-attention variant of the
llama.cpp backend. It targets high-concurrency decode for the Qwen3.6 hybrid
gated-DeltaNet (SSM) models on Blackwell (GB10 / DGX Spark). It reuses the stock
`llama-cpp` backend's sources and applies a vendored patch series on top at build
time. It is **not** a fork: a source-only `*.patch` stack plus one canonical doc.
**Canonical reference:** `backend/cpp/llama-cpp-localai-paged/README.md`
(architecture, the patch series 0001-0030, benchmarks, dev notes, generality,
pin/canary policy). Read it for any technical detail; this guide is the maintenance
how-to.
## Where things live
- `backend/cpp/llama-cpp-localai-paged/Makefile` - the thin wrapper. It copies the
stock `backend/cpp/llama-cpp/` build infra into a build dir, clones llama.cpp at
this backend's **own** pin (`LLAMA_VERSION`), applies the paged series via the
`apply-paged-patches` define (strict `git apply`), then builds `grpc-server`.
- `backend/cpp/llama-cpp-localai-paged/patches/paged/` - the source-only `.patch`
series (0001-0030), nothing else.
- `backend/cpp/llama-cpp-localai-paged/README.md` - the canonical doc. The
operational docs (`PAGED_BITEXACT_NOTE.md`, `UPSTREAM_LAYER2_SCOPE.md`) and
dev artifacts live in
`backend/cpp/llama-cpp-localai-paged/docs/`.
- `backend/Dockerfile.llama-cpp-localai-paged`, `.docker/llama-cpp-localai-paged-compile.sh`
- the CUDA build entry points.
- `backend/cpp/llama-cpp/` - the **stock** backend, pure upstream. It carries no
paged patches.
## Invariants (do not break these)
- **Stock stays pure.** The paged patches live ONLY in this backend. Never add a
`patches/paged/` dir or `LLAMA_PAGED` logic to `backend/cpp/llama-cpp/`.
- **CUDA-only.** Ship cublas/cuda targets only. Off-CUDA the fusions are gated off
(patch 0030) and NVFP4 falls back to dequant, so the backend is neutral-to-
slightly-negative there - non-CUDA users use the stock `llama-cpp`. Do not add
cpu/vulkan/sycl/metal rows for this backend in `.github/backend-matrix.yml`.
(Those builds also fail to link `grpc-server` on darwin/arm64 against upstream
`stream_*` server symbols - another reason it is CUDA-only.)
- **Source-only patches.** A `.patch` may touch only llama.cpp source - never a
dev doc or `*.md`. Strict `git apply` on a clean checkout must reach exit 0. (A
stray `SSM_DECODE_FIX_RESULTS.md` hunk in patch 0019 once broke the CI build.)
- **Bit-exact by default.** Every shipped patch is byte-identical to the f32
baseline. (The one opt-in precision trade, `ssm_bf16_tau` / patch 0026, was
DROPPED: it went flat once the decode fusions landed - forcing all gated-DeltaNet
heads to bf16 gave 780.6 vs 780.0 t/s, zero benefit - so the series is now
bit-exact end to end. Do not reintroduce a per-head SSM-precision lever; see the
rejected-levers note in the backend README section 5.)
## Fork-first workflow (MANDATORY)
The fork **`mudler/llama.cpp` branch `localai-paged`** is the CANONICAL source
of truth for ALL paged-backend kernel and patch work. The vendored
`patches/paged/*.patch` series is a **derivative**: the fork is the source, the
series is a generated mirror of it.
**Always update the fork FIRST, in this exact order:**
1. **Commit the change on the `localai-paged` branch and push it.** Every
kernel or patch change lands as a fork commit first.
2. **Then regenerate the LocalAI series from the fork** via `git format-patch`
(one patch per fork commit, source-only) into
`backend/cpp/llama-cpp-localai-paged/patches/paged/`, so the series stays a
**1:1, drift-free mirror** of the branch.
Hard rules, no exceptions:
- **NEVER edit the `patches/paged/*.patch` files directly.** They are generated
output, not source.
- **NEVER add a patch to the series that has no corresponding fork-branch
commit.** Every `.patch` must be the `git format-patch` of a real commit on
`localai-paged`.
- The fork branch is **where the build and the per-path bit-exact md5 gate
actually run**, so it is the **only** place a change is truly validated. A
patch living only in the LocalAI series has never been built or gated.
Verify the mirror by tree hash: applying the full on-disk series on the pin
must reproduce the fork branch tree byte-for-byte. (The patch maintenance
detail is in `backend/cpp/llama-cpp-localai-paged/docs/PATCH_MAINTENANCE.md`;
the hard-gate is section 2.5 of `docs/PARITY_HANDOFF.md`.)
## Maintaining the pin against new llama.cpp
The pin (`LLAMA_VERSION` in the wrapper Makefile) is advanced ONLY by the manual
pin-sync. It is deliberately **excluded from the nightly auto-bumper**
(`bump_deps.yaml`): a naive bump would shift the tree out from under the patches
and break `git apply` at build time.
1. **The canary tells you when to sync.** `.github/workflows/llama-cpp-paged-canary.yml`
runs weekly: it applies + builds the series against the latest upstream tip and
goes **red** when upstream drifts past the patches. Canary red -> run a pin-sync.
2. **The pin-sync** (recorded in the README section 7 and git history): rebase the series onto the new
tip (resolve conflicts; re-export **source-only** with a pathspec like
`-- src/ ggml/ common/ include/ tools/ tests/ cmake/`), rebuild on a CUDA box,
pass the bit-exact gate on **every** path + `test-backend-ops`, **and confirm
the full grpc-server build/link is green on CI**, then bump `LLAMA_VERSION`.
**Hard constraint: keep the pin == the stock `llama-cpp` pin.** `grpc-server.cpp`
is shared with the stock backend and tracks the stock pin. A paged pin that
diverges PAST an upstream server-API refactor breaks the grpc-server LINK even
when the patches are byte-for-byte bit-exact - the bit-exact gate alone does NOT
catch it. The `c299a92c` bump did exactly this (patches applied + greedy-md5
bit-exact, but `grpc-server.cpp` failed to link with undefined `stream_*` server
helpers the refactor pulled into its headers), so it was reverted to `9d5d882d`.
A pin bump is shippable only once the full CI grpc-server build is green, which in
practice means moving in lockstep with the stock pin (or vendoring a
pin-matched grpc-server.cpp, which we deliberately do not, to keep stock pure).
## The bit-exact gate (run for every change)
- greedy md5: `llama-completion -m MODEL -ngl 99 -fa on -p "The capital of France is" -n 48 --temp 0 --seed 1 </dev/null | md5sum`,
paged paths prefixed `LLAMA_KV_PAGED=1` (+ `LLAMA_MOE_FORCE_GRAPHS=1` for paged
MoE). Must match the recorded baseline. Redirect stdin from `/dev/null` or
`llama-completion` hangs in conversation mode.
- `test-backend-ops` (CUDA0 vs CPU oracle) for every touched op (`SSM_CONV*`,
`GATED_DELTA_NET`, `MUL_MAT`, `MUL_MAT_ID`).
- **The gate is per-path.** The paged-MoE md5 differs from the non-paged md5 - a
benign, KL-validated FP-accumulation-order difference (see `docs/PAGED_BITEXACT_NOTE.md`).
Compare a paged-MoE change to the **paged** reference, not the non-paged one.
## Encapsulating your work
- When you change a kernel, follow the **Fork-first workflow** above: commit and
push on the `localai-paged` branch first, then regenerate the `.patch`
(source-only) from the fork so this worktree mirrors the branch byte-for-byte.
Commit with sign-off.
- New optimization -> next patch number (gaps 0005/0027 are intentional). Update
the README's patch table and dev notes - keep the README the single doc; do not
scatter `*_RESULTS.md` files.
- Record rejected/flat levers in the README too (they stop the next person from
re-running dead ends).
## Follow-ups (Metal / SYCL / Vulkan)
The decode fusions are implemented for **CUDA + CPU only**. The base
gated-DeltaNet + SSM_CONV ops already exist upstream on Metal, SYCL, and Vulkan,
so the models **run** there via the non-fused path - what is missing is the
fusion speedup. Porting it (strictly mirroring the CUDA kernels, since we have no
Metal/SYCL/Vulkan hardware to test on here) is scoped in `docs/UPSTREAM_LAYER2_SCOPE.md`
(recommended order: Metal, then SYCL, then Vulkan; ops-first upstream PR, then one
PR per backend, each gated by `test-backend-ops` on the target hardware). The
methodology for that work is in [.agents/vllm-parity-methodology.md](vllm-parity-methodology.md).

View File

@@ -1,101 +0,0 @@
# Methodology: Closing the vLLM Decode-Throughput Gap in llama.cpp
This is the playbook that took the paged backend
([.agents/llama-cpp-localai-paged-backend.md](llama-cpp-localai-paged-backend.md))
from ~38% of vLLM decode to **parity-to-ahead on dense** (and a proven, honest
ceiling on MoE) on GB10. Use it for any "make llama.cpp match or beat engine X on
accelerator Y" effort. The *levers* are model- and hardware-specific; the
*discipline* below is not. The worked example, with all numbers, is the paged
backend README.
## The core loop
1. **Establish a bit-exact baseline and gate FIRST.** Record the greedy md5 (per
path) and an f32 reference. Every optimization must stay byte-identical to it -
or ship as an explicit, default-off precision opt-in. This is what lets you
optimize aggressively without silently regressing quality. Gate two ways:
greedy md5, and `test-backend-ops` against the CPU oracle.
2. **Profile - do not assume.** nsys the steady-state decode step, broken down per
*kernel* AND per *memcpy*. Find the dominant cost. "It's the GEMM" was wrong
here: on hybrid gated-DeltaNet models the bottleneck was the recurrent-state
**plumbing** (state memcpy + gathers, ~67% of the step), not the weight GEMM.
Also sanity-check GPU-busy %: an early "low utilization" reading was a profiling
window artifact (decode was 96-99% GPU-busy), not real idle.
3. **Ground-truth BOTH engines.** Decompose *your* decode step AND the
competitor's, side by side, per bucket, and compute the per-bucket delta. This
tells you WHERE the gap actually is - not where you would guess. It overturned
premises here: e.g. vLLM does NOT run the GDN/attn projections as NVFP4 (it
keeps them bf16, same as us); the MoE expert GEMM was a llama *win*, not the gap.
4. **Per-lever discipline.** For each candidate: implement -> bit-exact gate ->
same-harness A/B bench. Use a runtime env-toggle (flag off vs on) ONLY for
levers that are actually runtime-gated; a lever **compiled into** the binary
(e.g. the SSM decode fusions here) is NOT isolated by a runtime flag, so measure
it build-vs-build. The full-patchset "stock" baseline likewise needs a
**separately-built unpatched binary at the same pin** - toggling the runtime
flag on the patched binary does not reproduce stock (it measures only the gated
part; here that was ~neutral, which is exactly how this gotcha hides). Bank only
what lifts AND gates. **Record every rejected or flat lever with the reason** -
over time this is the most valuable part: it stops the next person re-running
dead ends.
5. **Name the structural floor.** Prove the bit-exact ceiling exhaustively (every
lever measured, not assumed). What remains is physical - the memory-bandwidth
floor, the irreducible serial-SSM host loop (sampling can't start until logits
land). Name it; do not claim more than you measured.
## Hard rules learned
- **Apples-to-apples, or label it.** Stock-vs-patched on the SAME harness
(`llama-batched-bench`) is exact - lead with it. But "stock" must be a
separately-built unpatched binary at the SAME pin, NOT the patched binary with
the runtime flag off (compiled-in wins survive the toggle). Cross-engine "% of vLLM"
(batched-bench vs vLLM server+client) is *indicative*; always caveat the harness
and config (context length alone shifted the MoE figure 76% <-> 86%).
- **Re-measure a "win" after later levers land - it may evaporate.** bf16 SSM
state (the `ssm_bf16_tau` lever) benched +12% early and failed the f32 KL gate
(vLLM keeps f32 too), so it was kept default-off opt-in. Once the decode fusions
(recurrent-state gather-fusion + block-table cache) landed, a clean re-measure
forcing ALL gated-DeltaNet heads to bf16 (`tau=100000`) went **flat** - 780.6 vs
780.0 t/s. The "+12%" was subsumed by the fusions: the lever bought nothing, so
it was **dropped** (precision trade + bug surface + extra CUDA template-instantiation
compile cost, zero benefit). A win measured before the rest of the series is not a
win after it.
- **Reject the obvious-but-wrong, with evidence.** A faster kernel that is off the
critical path benches FLAT (the freed time becomes idle). Quantizing the bf16
projections to NVFP4 cost ~6% PPL - and vLLM keeps them bf16 for the same reason.
Always measure before believing; a plausible mechanism is not a result.
- **The gate can be per-path.** Paged vs non-paged attention legitimately produces
different (equivalent) FP-reduction orders; validate the difference is benign
(KLD to f32) and then gate each path against its own reference.
## Orchestration (multi-agent)
- **One GPU profiler/bencher at a time** (the GPU-contention rule). Parallel
design/analysis/read agents are fine; concurrent GPU benches pollute each other's
numbers.
- **Adversarial verify.** Before banking a finding, spawn skeptics prompted to
*refute* it; majority-refute kills it. Prevents plausible-but-wrong results.
- **Anti-punt.** Use foreground, blocking ssh loops with short benches and a
progress-file checkpoint. Agents that background work and "wait for the monitor
event" stall - forbid that pattern.
- **GPU coexistence.** On a shared host, stop the user's deployments for a clean
benchmark window (with their OK) and ALWAYS restore them (wrap the bench so a
failure cannot strand them).
## What generalizes (and what doesn't)
The *speedups* may be hardware-specific (here: CUDA/Blackwell - the SSM fusions,
NVFP4 FP4-MMA, the occupancy tune), which is why other accelerators did not
benefit. But the *findings* often generalize and are worth upstreaming: the
"decode is plumbing-bound, not GEMM-bound" insight and the bit-exact, CPU-mirrored
fusion ops help any backend running these models. Separate "ship our tuned backend"
from "upstream the portable op" - they are different deliverables.
## The closing record
Write up the result HONESTLY: the shipped wins, the rejected levers (with reasons),
the structural ceiling, and the cross-backend / cross-quant generality. Negative
results are as valuable as wins. The paged backend README is the template.

View File

@@ -1,5 +1,5 @@
#!/usr/bin/env bash
# Shared compile logic for backend/Dockerfile.llama-cpp-localai-paged.
# Shared compile logic for backend/Dockerfile.bonsai.
# Sourced (via bind mount) from both builder-fromsource and builder-prebuilt stages.
set -euxo pipefail
@@ -14,10 +14,10 @@ if [[ -n "${CUDA_DOCKER_ARCH:-}" ]]; then
CUDA_ARCH_ESC="${CUDA_DOCKER_ARCH//;/\\;}"
export CMAKE_ARGS="${CMAKE_ARGS} -DCMAKE_CUDA_ARCHITECTURES=${CUDA_ARCH_ESC}"
echo "CMAKE_ARGS(env) = ${CMAKE_ARGS}"
rm -rf /LocalAI/backend/cpp/llama-cpp-localai-paged-*-build
rm -rf /LocalAI/backend/cpp/bonsai-*-build
fi
cd /LocalAI/backend/cpp/llama-cpp-localai-paged
cd /LocalAI/backend/cpp/bonsai
if [ -z "${BUILD_TYPE:-}" ]; then
# Pure CPU image: one ggml CPU_ALL_VARIANTS build replaces the per-microarch binaries.
@@ -26,14 +26,14 @@ if [ -z "${BUILD_TYPE:-}" ]; then
apt-get update -qq && apt-get install -y -qq gcc-14 g++-14
export CC=gcc-14 CXX=g++-14
fi
make llama-cpp-localai-paged-cpu-all
make bonsai-cpu-all
else
# GPU build (cublas/hipblas/sycl/vulkan/...): single fallback CPU build, the accelerator
# does the compute. Keeps the GPU compile from also building the CPU variant matrix and
# avoids the gcc-14 apt step on GPU base images such as nvidia l4t.
make llama-cpp-localai-paged-fallback
make bonsai-fallback
fi
make llama-cpp-localai-paged-grpc
make llama-cpp-localai-paged-rpc-server
make bonsai-grpc
make bonsai-rpc-server
ccache -s || true

View File

@@ -7,8 +7,11 @@
# Runs only the checks relevant to what's staged:
# - Go files -> make lint + make test-coverage-check
# - core/http/react-ui -> make test-ui-coverage-check (Playwright e2e + gate)
# A commit touching neither is skipped entirely (docs/YAML/etc. can't change
# lint findings, Go coverage, or the UI).
# - realtime state machines / specs -> make test-realtime-conformance
# (respcoord/**, turncoord/**, or formal-verification/** -- a pure .fizz
# spec edit must still re-verify the design, detected separately from Go)
# A commit touching none of these is skipped entirely (other docs/YAML can't
# change lint findings, Go coverage, the UI, or the realtime conformance gate).
#
# To bypass for a single commit (e.g. a WIP checkpoint): git commit --no-verify
set -eu
@@ -20,11 +23,13 @@ staged="$(git diff --cached --name-only --diff-filter=ACMRD)"
go_changed=0
ui_changed=0
rt_changed=0
if echo "$staged" | grep -qE '\.go$'; then go_changed=1; fi
if echo "$staged" | grep -qE '^core/http/react-ui/'; then ui_changed=1; fi
if echo "$staged" | grep -qE '^(core/http/endpoints/openai/(coordinator|respcoord|turncoord|conncoord|compactcoord|ttscoord)/|formal-verification/)'; then rt_changed=1; fi
if [ "$go_changed" -eq 0 ] && [ "$ui_changed" -eq 0 ]; then
echo "pre-commit: no Go or React UI changes staged — skipping."
if [ "$go_changed" -eq 0 ] && [ "$ui_changed" -eq 0 ] && [ "$rt_changed" -eq 0 ]; then
echo "pre-commit: no Go, React UI, or realtime-spec changes staged — skipping."
exit 0
fi
@@ -57,4 +62,11 @@ if [ "$ui_changed" -eq 1 ]; then
make test-ui-coverage-check
fi
if [ "$rt_changed" -eq 1 ]; then
echo "pre-commit ▶ realtime state-machine conformance (make test-realtime-conformance) —"
echo " Go transition/rapid tests under -race + FizzBee model check of the"
echo " authoritative specs. Fail-closed: needs FizzBee (make install-fizzbee)."
make test-realtime-conformance
fi
echo "pre-commit ✓ all relevant checks passed"

View File

@@ -7,6 +7,7 @@ This PR fixes #
**[Signed commits](../CONTRIBUTING.md#signing-off-on-commits-developer-certificate-of-origin)**
- [ ] Yes, I signed my commits.
- [ ] Documentation updated (docs/content/) for user-facing changes, or not applicable
<!--
Thank you for contributing to LocalAI!

View File

@@ -23,7 +23,7 @@
# checklist in .agents/adding-backends.md (includeDarwin entry, the index.yaml
# `metal:` capability + `metal-<backend>` image entries, a `run.sh` Darwin/DYLD
# branch for C/C++ backends, and the inferBackendPathDarwin case in
# scripts/changed-backends.js so the path filter actually builds it).
# scripts/lib/backend-filter.mjs so the path filter actually builds it).
# Linux matrix (consumed by backend-jobs).
include:
@@ -452,6 +452,22 @@ include:
dockerfile: "./backend/Dockerfile.turboquant"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-12-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-cuda-12-amd64'
# bigger-runner: same rationale as -gpu-nvidia-cuda-12-llama-cpp above
# (observed 6h5m wall-clock on v4.2.1, just past the 6h job timeout).
runs-on: 'bigger-runner'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
@@ -478,6 +494,19 @@ include:
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-12-longcat-video'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "longcat-video"
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
@@ -790,6 +819,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-12-moss-transcribe-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
@@ -816,6 +858,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-12-moss-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
@@ -1056,6 +1111,21 @@ include:
dockerfile: "./backend/Dockerfile.turboquant"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-cuda-13-amd64'
# bigger-runner: observed 6h5m wall-clock on v4.2.1 — at the GHA timeout.
runs-on: 'bigger-runner'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1084,6 +1154,20 @@ include:
backend: "turboquant"
dockerfile: "./backend/Dockerfile.turboquant"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-cuda-13-arm64-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-cuda-13-arm64'
base-image: "ubuntu:24.04"
runs-on: 'ubuntu-24.04-arm'
ubuntu-version: '2404'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1136,6 +1220,19 @@ include:
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13-longcat-video'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "longcat-video"
dockerfile: "./backend/Dockerfile.python"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1344,6 +1441,19 @@ include:
backend: "vllm-omni"
dockerfile: "./backend/Dockerfile.python"
context: "./"
- build-type: 'l4t'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/arm64'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-cuda-13-arm64-longcat-video'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
ubuntu-version: '2404'
backend: "longcat-video"
dockerfile: "./backend/Dockerfile.python"
context: "./"
- build-type: 'l4t'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1721,6 +1831,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13-moss-transcribe-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1760,6 +1883,19 @@ include:
backend: "parakeet-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-cuda-13-arm64-moss-transcribe-cpp'
base-image: "ubuntu:24.04"
ubuntu-version: '2404'
runs-on: 'ubuntu-24.04-arm'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1786,6 +1922,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13-moss-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1838,6 +1987,19 @@ include:
backend: "qwen3-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-cuda-13-arm64-moss-tts-cpp'
base-image: "ubuntu:24.04"
ubuntu-version: '2404'
runs-on: 'ubuntu-24.04-arm'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1905,6 +2067,20 @@ include:
dockerfile: "./backend/Dockerfile.llama-cpp"
context: "./"
ubuntu-version: '2404'
- build-type: 'hipblas'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-rocm-amd64'
runs-on: 'ubuntu-latest'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
skip-drivers: 'false'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
ubuntu-version: '2404'
- build-type: 'hipblas'
cuda-major-version: ""
cuda-minor-version: ""
@@ -2169,6 +2345,20 @@ include:
dockerfile: "./backend/Dockerfile.turboquant"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f32'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-intel-sycl-f32-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-intel-amd64'
runs-on: 'ubuntu-latest'
base-image: "intel/oneapi-basekit:2025.3.0-0-devel-ubuntu24.04"
skip-drivers: 'false'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f16'
cuda-major-version: ""
cuda-minor-version: ""
@@ -2197,6 +2387,20 @@ include:
dockerfile: "./backend/Dockerfile.turboquant"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f16'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-intel-sycl-f16-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-intel-amd64'
runs-on: 'ubuntu-latest'
base-image: "intel/oneapi-basekit:2025.3.0-0-devel-ubuntu24.04"
skip-drivers: 'false'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
ubuntu-version: '2404'
- build-type: 'intel'
cuda-major-version: ""
cuda-minor-version: ""
@@ -2649,6 +2853,21 @@ include:
dockerfile: "./backend/Dockerfile.turboquant"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-amd64'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
@@ -2664,6 +2883,21 @@ include:
dockerfile: "./backend/Dockerfile.turboquant"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-arm64'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
@@ -2806,6 +3040,20 @@ include:
dockerfile: "./backend/Dockerfile.turboquant"
context: "./"
ubuntu-version: '2204'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-arm64-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-l4t-cuda-12-arm64'
base-image: "nvcr.io/nvidia/l4t-jetpack:r36.4.0"
runs-on: 'ubuntu-24.04-arm'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
ubuntu-version: '2204'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
@@ -2852,6 +3100,22 @@ include:
context: "./"
ubuntu-version: '2404'
# Stablediffusion-ggml
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-vulkan-amd64'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
ubuntu-version: '2404'
# Stablediffusion-ggml
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
@@ -2868,6 +3132,22 @@ include:
context: "./"
ubuntu-version: '2404'
# Stablediffusion-ggml
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-bonsai'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-vulkan-arm64'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "bonsai"
dockerfile: "./backend/Dockerfile.bonsai"
context: "./"
ubuntu-version: '2404'
# Stablediffusion-ggml
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
@@ -3597,6 +3877,115 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# moss-transcribe-cpp
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-moss-transcribe-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-moss-transcribe-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f32'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-intel-sycl-f32-moss-transcribe-cpp'
runs-on: 'ubuntu-latest'
base-image: "intel/oneapi-basekit:2025.3.0-0-devel-ubuntu24.04"
skip-drivers: 'false'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f16'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-intel-sycl-f16-moss-transcribe-cpp'
runs-on: 'ubuntu-latest'
base-image: "intel/oneapi-basekit:2025.3.0-0-devel-ubuntu24.04"
skip-drivers: 'false'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-moss-transcribe-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-moss-transcribe-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-arm64-moss-transcribe-cpp'
base-image: "nvcr.io/nvidia/l4t-jetpack:r36.4.0"
runs-on: 'ubuntu-24.04-arm'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2204'
- build-type: 'hipblas'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-moss-transcribe-cpp'
base-image: "rocm/dev-ubuntu-24.04:7.2.1"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "moss-transcribe-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# ced
- build-type: 'cublas'
cuda-major-version: "12"
@@ -4179,6 +4568,35 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# moss-tts-cpp
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-moss-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-moss-tts-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# omnivoice-cpp
- build-type: ''
cuda-major-version: ""
@@ -4221,6 +4639,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f32'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-intel-sycl-f32-moss-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "intel/oneapi-basekit:2025.3.0-0-devel-ubuntu24.04"
skip-drivers: 'false'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f32'
cuda-major-version: ""
cuda-minor-version: ""
@@ -4247,6 +4678,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f16'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-intel-sycl-f16-moss-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "intel/oneapi-basekit:2025.3.0-0-devel-ubuntu24.04"
skip-drivers: 'false'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'sycl_f16'
cuda-major-version: ""
cuda-minor-version: ""
@@ -4274,6 +4718,20 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-moss-tts-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
@@ -4302,6 +4760,20 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-moss-tts-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
@@ -4329,6 +4801,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2204'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-arm64-moss-tts-cpp'
base-image: "nvcr.io/nvidia/l4t-jetpack:r36.4.0"
runs-on: 'ubuntu-24.04-arm'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2204'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "0"
@@ -4355,6 +4840,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'hipblas'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-rocm-hipblas-moss-tts-cpp'
base-image: "rocm/dev-ubuntu-24.04:6.4.4"
runs-on: 'ubuntu-latest'
skip-drivers: 'false'
backend: "moss-tts-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'hipblas'
cuda-major-version: ""
cuda-minor-version: ""
@@ -4652,7 +5150,6 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# rfdetr
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
@@ -4667,6 +5164,35 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# cloud-proxy
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-cloud-proxy'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "cloud-proxy"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-cloud-proxy'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "cloud-proxy"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# rfdetr
- build-type: ''
cuda-major-version: ""
@@ -5177,39 +5703,6 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# llama-cpp-localai-paged: the LocalAI paged-attention llama.cpp variant. Each
# row mirrors the corresponding llama-cpp row with backend/dockerfile/tag-suffix
# swapped; builder-base-image is left UNCHANGED so these reuse the same
# base-grpc-* prebuilt bases (same gRPC + same toolchain), needing no new
# base-images.yml variant.
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13-llama-cpp-localai-paged'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-cuda-13-amd64'
runs-on: 'bigger-runner'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "llama-cpp-localai-paged"
dockerfile: "./backend/Dockerfile.llama-cpp-localai-paged"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-cuda-13-arm64-llama-cpp-localai-paged'
builder-base-image: 'quay.io/go-skynet/ci-cache:base-grpc-cuda-13-arm64'
base-image: "ubuntu:24.04"
runs-on: 'ubuntu-24.04-arm'
ubuntu-version: '2404'
backend: "llama-cpp-localai-paged"
dockerfile: "./backend/Dockerfile.llama-cpp-localai-paged"
context: "./"
# Darwin matrix (consumed by backend-jobs-darwin).
includeDarwin:
@@ -5253,6 +5746,10 @@ includeDarwin:
tag-suffix: "-metal-darwin-arm64-parakeet-cpp"
build-type: "metal"
lang: "go"
- backend: "moss-transcribe-cpp"
tag-suffix: "-metal-darwin-arm64-moss-transcribe-cpp"
build-type: "metal"
lang: "go"
- backend: "ced"
tag-suffix: "-metal-darwin-arm64-ced"
build-type: "metal"
@@ -5273,6 +5770,10 @@ includeDarwin:
tag-suffix: "-metal-darwin-arm64-qwen3-tts-cpp"
build-type: "metal"
lang: "go"
- backend: "moss-tts-cpp"
tag-suffix: "-metal-darwin-arm64-moss-tts-cpp"
build-type: "metal"
lang: "go"
- backend: "omnivoice-cpp"
tag-suffix: "-metal-darwin-arm64-omnivoice-cpp"
build-type: "metal"
@@ -5398,6 +5899,10 @@ includeDarwin:
tag-suffix: "-metal-darwin-arm64-local-store"
build-type: "metal"
lang: "go"
- backend: "cloud-proxy"
tag-suffix: "-metal-darwin-arm64-cloud-proxy"
build-type: "metal"
lang: "go"
- backend: "llama-cpp-quantization"
tag-suffix: "-metal-darwin-arm64-llama-cpp-quantization"
build-type: "mps"

18
.github/bump_deps.sh vendored
View File

@@ -1,5 +1,8 @@
#!/bin/bash
set -xe
source "$(dirname "${BASH_SOURCE[0]}")/gh_curl.sh"
REPO=$1
BRANCH=$2
VAR=$3
@@ -9,7 +12,20 @@ if [ -z "$FILE" ]; then
FILE="Makefile"
fi
LAST_COMMIT=$(curl -s -H "Accept: application/vnd.github.VERSION.sha" "https://api.github.com/repos/$REPO/commits/$BRANCH")
# gh_curl follows redirects so a renamed/transferred upstream repo (GitHub
# answers 301) still resolves, and fails on HTTP errors rather than letting an
# error page reach sed below. `|| true` keeps a failed lookup from aborting the
# script at exit 22 with no context — the SHA guard below reports it instead.
LAST_COMMIT=$(gh_curl -H "Accept: application/vnd.github.VERSION.sha" "https://api.github.com/repos/$REPO/commits/$BRANCH" || true)
# Guard the sed input: anything that is not a bare 40-hex SHA (an API error
# body, an empty response) would otherwise be spliced into the Makefile pin —
# either corrupting it silently or blowing up sed with an unterminated
# expression, which is how this job failed for a renamed repo.
if ! [[ "$LAST_COMMIT" =~ ^[0-9a-f]{40}$ ]]; then
echo "Refusing to bump $VAR: expected a 40-char commit SHA for $REPO@$BRANCH, got: $LAST_COMMIT" >&2
exit 1
fi
# Read $VAR from Makefile (only first match)
set +e

13
.github/bump_docs.sh vendored
View File

@@ -1,7 +1,18 @@
#!/bin/bash
set -xe
source "$(dirname "${BASH_SOURCE[0]}")/gh_curl.sh"
REPO=$1
LATEST_TAG=$(curl -s "https://api.github.com/repos/$REPO/releases/latest" | jq -r '.tag_name')
LATEST_TAG=$(gh_curl -H "Accept: application/vnd.github+json" \
"https://api.github.com/repos/$REPO/releases/latest" | jq -r '.tag_name')
# jq prints the string "null" for a missing key, so a throttled or otherwise
# unexpected API response would otherwise be published as the docs version.
if [ -z "$LATEST_TAG" ] || [ "$LATEST_TAG" = "null" ]; then
echo "Refusing to bump docs version: could not resolve the latest release tag for $REPO." >&2
exit 1
fi
cat <<< $(jq ".version = \"$LATEST_TAG\"" docs/data/version.json) > docs/data/version.json

View File

@@ -11,6 +11,9 @@
# darwin build can only use the exact vLLM version vllm-metal supports, so it may
# lag the Linux pin (requirements-cublas13-after.txt) until vllm-metal catches up.
set -xe
source "$(dirname "${BASH_SOURCE[0]}")/gh_curl.sh"
REPO=$1 # vllm-project/vllm-metal
FILE=$2 # backend/python/vllm/install.sh
VAR=$3 # VLLM_METAL_VERSION (used for the workflow's output file names)
@@ -22,12 +25,12 @@ fi
# vllm-metal ships frequent dev releases, all flagged as non-prerelease, so
# /releases/latest returns the newest one (with its cp312 wheel asset).
LATEST_TAG=$(curl -sS -H "Accept: application/vnd.github+json" \
LATEST_TAG=$(gh_curl -H "Accept: application/vnd.github+json" \
"https://api.github.com/repos/$REPO/releases/latest" \
| python3 -c "import json,sys; print(json.load(sys.stdin)['tag_name'])")
# The coupled vLLM source version lives in vllm-metal's installer at that tag.
NEW_VLLM_VERSION=$(curl -fsSL \
NEW_VLLM_VERSION=$(gh_curl \
"https://raw.githubusercontent.com/$REPO/$LATEST_TAG/install.sh" \
| grep -oE 'vllm_v="[0-9]+\.[0-9]+\.[0-9]+"' | head -1 | cut -d'"' -f2)

View File

@@ -9,6 +9,9 @@
# vars in Makefiles; this script handles the two-value rewrite specific to the
# vLLM requirements file.
set -xe
source "$(dirname "${BASH_SOURCE[0]}")/gh_curl.sh"
REPO=$1 # vllm-project/vllm
FILE=$2 # backend/python/vllm/requirements-cublas13-after.txt
VAR=$3 # VLLM_VERSION (used for output file names so the workflow can read them)
@@ -19,7 +22,7 @@ if [ -z "$FILE" ] || [ -z "$REPO" ] || [ -z "$VAR" ]; then
fi
# /releases/latest returns the most recent non-prerelease tag.
LATEST_TAG=$(curl -sS -H "Accept: application/vnd.github+json" \
LATEST_TAG=$(gh_curl -H "Accept: application/vnd.github+json" \
"https://api.github.com/repos/$REPO/releases/latest" \
| python3 -c "import json,sys; print(json.load(sys.stdin)['tag_name'])")

194
.github/ci/apexentries/README.md vendored Normal file
View File

@@ -0,0 +1,194 @@
# apexentries
Generates gallery entries for the `mudler/*-APEX-GGUF` HuggingFace repositories.
Each APEX repo becomes one **family**: one entry per quality rung the repo
publishes and one per quantization rung its unsloth counterpart publishes, all
gathered under the **base model's** entry. LocalAI's variant selector then picks
the build that fits the hardware in front of it.
## The hub is the base model entry, never a generated `*-apex` parent
Somebody looking for `qwen3.6-35b-a3b` must find every build of those weights
under that one name: the APEX imatrix rungs, the unsloth quant rungs and any
speculative build. A separate `qwen3.6-35b-a3b-apex` hub competing with the base
entry would split the family in two and leave whichever half the user did not
search for effectively invisible.
So the generator resolves the hub by stripping the `-APEX`, `-MTP` and `-TQ`
markers and looking the result up in the index, trying both the repo-derived and
the stem-derived candidate the same way `CounterpartCandidates` does. Then:
- **The hub exists** (14 of the 45 repos, resolving to 10 distinct entries).
Nothing new is emitted for the family root. A `variants:` block is spliced into
the entry that is already there, textually, leaving its description, icon,
tags, overrides and files untouched. The line editing is shared with the
`variantproposals` job via `.github/ci/galleryedit`.
- **The hub is absent** (the other 31). A new hub is emitted, named for the base
model and never for the APEX repo. It carries one of the discovered builds as
its own payload so it is a complete installable entry rather than a bare index,
and that payload is what gives it an `overrides.backend`. Without a declared
backend the verifier would skip it, so a hub carrying feature tags would escape
the tagging check in silence.
Several APEX repos routinely resolve to one base model, so both paths accumulate
by hub name rather than assuming one family per hub.
Two references are always filtered out of a hub's list: anything the entry
already declares, and the hub's own name. The self reference is not merely
redundant. An unsloth rung whose weights the gallery already ships under the base
name resolves, through the merge, straight back to the hub, and the verifier
reads a self reference as a variant that declares variants of its own.
The four hand-written `*-apex` entries (`qwen3.6-35b-a3b-apex`,
`gemma-4-26b-a4b-it-apex`, `qwen3.5-35b-a3b-apex`,
`nemotron-3-nano-omni-30b-a3b-reasoning-apex`) are **ordinary builds**, not hubs.
They are referenced from their hub's variants list like any other rung, and are
never deleted or renamed.
## Flags
| Flag | Default | Meaning |
|------|---------|---------|
| `-index <path>` | `gallery/index.yaml` | Gallery index to dedup against. Read only, unless `-apply` is passed. |
| `-only <a,b,c>` | (all) | Comma-separated full repo names (`mudler/Foo-APEX-GGUF`) to restrict generation to. A name that matches nothing is reported as a warning, since it is a typo rather than an empty result. |
| `-out <path>` | (none) | Write the entries to add to this file. Nothing is written to the gallery. |
| `-apply` | `false` | Splice the variants into `-index` and append the new entries to it. |
| `-verify <path>` | (none) | Verify a gallery index and exit. Ignores every other flag. |
Either `-out` or `-apply` is required, otherwise the run has nothing to do.
`-apply` splices variant lines into existing entries and **appends** new ones. It
never re-serialises the index: it is roughly 40,000 lines, and a YAML round trip
would reflow the whole file, drop the anchors and merge keys the gallery relies
on, and produce a diff nobody can review. On the three-family sample the splice
is 24 added lines across 3 hunks with zero deletions.
## Discovery is by filename suffix, never by repo name
Builds come from the files a repo actually publishes. A filename is never
constructed from a repo name, because the two disagree:
`mudler/gemma-4-26B-A4B-it-APEX-GGUF` ships `gemma-4-26B-A4B-APEX-*.gguf`, and
five other repos likewise drop a suffix (`-it`, `-2603`) or a vendor prefix
(`NVIDIA-`) that the repo name carries. Composing a URL from the repo name would
produce a 404 for every one of them, and the 404 would only surface after the
entry shipped.
The quality ladder is matched on the trailing tier marker, `-(I-)?(Quality|
Balanced|Compact|Mini|Nano).gguf`. The `I-` prefix marks the imatrix ladder. The
imatrix ladder is emitted when it is non-empty and the plain ladder is used only
as a fallback, because two of the 45 repos publish no imatrix tiers at all and
must still contribute. Eleven repos carry a fifth `I-Nano` rung, so nothing
assumes a fixed number of rungs.
Every run prints, per repo, the counts that discovery accounted for. If the
number of classified files is short of the number of `.gguf` files the repo
publishes, the shortfall is printed as `UNCLASSIFIED`. That check is a set
difference on counts rather than a second pass over filenames: a second matcher
would duplicate the tier regex and the two copies would drift. The failure it
catches is quiet. A publishing-script typo that breaks every imatrix filename in
a repo does not produce a short ladder; it makes the imatrix ladder empty, and
the fallback then downgrades the whole family to the plain ladder with nothing
said. A downstream HTTP check cannot catch it either, because it validates the
URLs that were emitted, and an undiscovered tier emits none.
The same reasoning applies to `UNACCOUNTED QUANT`, printed when the unsloth
counterpart demonstrably publishes a wanted quant that produced no build. It is
reported at discovery time because a dropped quant leaves no trace at all in the
finished gallery file.
## sha256 always comes from the API
Every file stanza takes its `sha256` from the HuggingFace models API
(`lfs.sha256`). A GGUF the API describes without one is a fatal error for that
family: the repo is reported by name and the run ends non-zero. It is never
substituted from another field, because that is exactly how a Xet hash ends up
masquerading as a content hash.
## The dflash / mtp tagging rule
An entry is tagged `dflash` or `mtp` **if and only if** it configures the
matching `spec_type:draft-<feature>`. Variant ranking reads tags and nothing
else, so a tag that does not match the configuration either promotes a build
that is no faster or hides one that genuinely is.
A repo name is not configuration. `mudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF` ships
weights that carry MTP heads; an entry that does not enable them is not an MTP
entry and is not tagged as one.
A generated hub inherits the tags of the build it carries as its payload, rather
than rebuilding them from the base set, so a hub whose payload configures a
`spec_type` stays tagged consistently with the overrides copied alongside it.
## Reuse reporting: two categories, not one
Generated entries are deduped against the gallery and against the batch itself.
The run prints the result under two separate headings, because the two cases are
not equivalent:
- **URI MATCHES** mean the gallery, or an earlier entry in this batch, already
ships exactly these weights. Pointing the hub at the existing entry is correct
and needs no thought.
- **NAME COLLISIONS** mean an entry already owns the name but holds different
weights. Referencing it would point the hub at a build other than the one
generated. Every one of these must be inspected by hand.
The run then prints `HUBS SPLICED`, listing every reference that will be added to
an entry the gallery already ships along with the line it will be added at, and
`HUBS CREATED` for the families that get a new hub. The splices are the part a
review has to read closely, because they modify entries somebody else wrote.
Hubs are deliberately kept out of the merge. A new hub carries the family's top
rung as its own payload, so URI dedup would fold the hub into that rung and the
family would lose the very entry point this command exists to create.
## Workflow: sample first, then the full set
Never run the full generation straight into the gallery. Generate a small,
deliberately awkward sample, have it reviewed, then run the rest.
```bash
# 1. Sample three families that between them cover the awkward shapes:
# a standard four-rung repo, one with the extra I-Nano rung AND a file stem
# that differs from its repo name, and one whose unsloth counterpart shards
# its quants across subdirectories.
go run ./.github/ci/apexentries \
-index gallery/index.yaml \
-only mudler/Qwen3.6-35B-A3B-APEX-GGUF,mudler/gemma-4-26B-A4B-it-APEX-GGUF,mudler/Step-3.7-Flash-APEX-GGUF \
-out /tmp/sample.yaml
# 2. Verify the sample against the gallery it would join, splices included. Apply
# to a COPY, never to the real index, and check that the diff is only the
# intended variant lines. Compare the verifier output to the gallery's own
# baseline: what matters is that the sample adds no new problem, not that the
# total is zero.
cp gallery/index.yaml /tmp/index-copy.yaml
go run ./.github/ci/apexentries -index /tmp/index-copy.yaml -only <same list> -apply
diff -u gallery/index.yaml /tmp/index-copy.yaml # expect zero deletions
go run ./.github/ci/apexentries -verify gallery/index.yaml > /tmp/baseline.log 2>&1
go run ./.github/ci/apexentries -verify /tmp/index-copy.yaml > /tmp/spliced.log 2>&1
diff /tmp/baseline.log /tmp/spliced.log
# 3. Have a human review /tmp/sample.yaml and every reported name collision.
# 4. Only then, the full set.
go run ./.github/ci/apexentries -index gallery/index.yaml -apply
```
## Tests
```bash
go test ./.github/ci/apexentries/
```
The shared line editor has its own package:
```bash
go test ./.github/ci/galleryedit/
```
`.github/ci/` is invisible to `go list ./...`, so these specs are not covered by
`make lint` or the repository test run. `.github/workflows/ci-tools-tests.yaml`
names the package explicitly; keep that workflow in step with any package added
under `.github/ci/`.

70
.github/ci/apexentries/discover.go vendored Normal file
View File

@@ -0,0 +1,70 @@
package main
import (
"regexp"
"strings"
)
// tierRE matches the tier marker APEX repos put at the end of a weight
// filename. Discovery is by suffix because the stem is not predictable from
// the repo name: six of the 45 repos drop a suffix ("-it", "-2603") or a
// vendor prefix ("NVIDIA-") that the repo name carries.
var tierRE = regexp.MustCompile(`-(I-)?(Quality|Balanced|Compact|Mini|Nano)\.gguf$`)
// fullPrecisionRE matches the unquantized source weights an APEX repo publishes
// alongside its ladder, flat (-F16.gguf) or sharded across a numbered set
// (-F16-00001-of-00010.gguf). bf16 is accepted because some repos publish that
// instead, and the match is case-insensitive because the casing varies between
// publishing scripts.
//
// These are deliberately not tiers: they are the weights the ladder is quantized
// FROM, and generation is scoped to the ladder itself.
var fullPrecisionRE = regexp.MustCompile(`(?i)-b?f16(-\d{5}-of-\d{5})?\.gguf$`)
// IsFullPrecision reports whether a weight filename is an unquantized source.
func IsFullPrecision(name string) bool {
return fullPrecisionRE.MatchString(name)
}
// Tier is one discovered build of an APEX repo.
type Tier struct {
Label string
File GGUFFile
}
// DiscoverAPEXTiers splits a repo's weight files into the imatrix ladder and
// the plain ladder. mmproj files are never tiers.
func DiscoverAPEXTiers(files []GGUFFile) (imatrix, plain []Tier) {
for _, f := range files {
if strings.HasPrefix(f.Name, "mmproj") {
continue
}
m := tierRE.FindStringSubmatch(f.Name)
if m == nil {
continue
}
if m[1] != "" {
imatrix = append(imatrix, Tier{Label: "I-" + m[2], File: f})
continue
}
plain = append(plain, Tier{Label: m[2], File: f})
}
return imatrix, plain
}
// DiscoverMMProj returns the repo's projector file, if it publishes one. The
// name varies across repos (mmproj.gguf, mmproj-F16.gguf,
// mmproj-step3.7-flash-f16.gguf), so match the prefix rather than a fixed name.
func DiscoverMMProj(files []GGUFFile) (GGUFFile, bool) {
for _, f := range files {
if strings.HasPrefix(f.Name, "mmproj") {
return f, true
}
}
return GGUFFile{}, false
}
// FileStem returns a tier's filename with its tier suffix removed.
func FileStem(t Tier) string {
return tierRE.ReplaceAllString(t.File.Name, "")
}

68
.github/ci/apexentries/discover_test.go vendored Normal file
View File

@@ -0,0 +1,68 @@
package main
import (
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("DiscoverAPEXTiers", func() {
It("finds tiers regardless of how the stem relates to the repo name", func() {
// This repo is mudler/gemma-4-26B-A4B-it-APEX-GGUF but its files drop "-it".
files := []GGUFFile{
{Name: "gemma-4-26B-A4B-APEX-I-Quality.gguf", SHA256: "a"},
{Name: "gemma-4-26B-A4B-APEX-I-Nano.gguf", SHA256: "b"},
{Name: "gemma-4-26B-A4B-APEX-Quality.gguf", SHA256: "c"},
{Name: "mmproj-F16.gguf", SHA256: "d"},
}
imatrix, plain := DiscoverAPEXTiers(files)
Expect(labels(imatrix)).To(ConsistOf("I-Quality", "I-Nano"))
Expect(labels(plain)).To(ConsistOf("Quality"))
})
It("excludes mmproj from the tier list", func() {
files := []GGUFFile{{Name: "mmproj.gguf", SHA256: "d"}}
imatrix, plain := DiscoverAPEXTiers(files)
Expect(imatrix).To(BeEmpty())
Expect(plain).To(BeEmpty())
})
})
var _ = Describe("DiscoverMMProj", func() {
It("finds an mmproj whatever its suffix", func() {
files := []GGUFFile{
{Name: "Model-APEX-I-Mini.gguf", SHA256: "a"},
{Name: "mmproj-step3.7-flash-f16.gguf", SHA256: "b"},
}
got, ok := DiscoverMMProj(files)
Expect(ok).To(BeTrue())
Expect(got.Name).To(Equal("mmproj-step3.7-flash-f16.gguf"))
})
It("reports absence when the repo ships none", func() {
_, ok := DiscoverMMProj([]GGUFFile{{Name: "Model-APEX-Quality.gguf", SHA256: "a"}})
Expect(ok).To(BeFalse())
})
})
var _ = Describe("FileStem", func() {
It("strips the tier suffix", func() {
t := Tier{Label: "I-Quality", File: GGUFFile{Name: "gemma-4-26B-A4B-APEX-I-Quality.gguf"}}
Expect(FileStem(t)).To(Equal("gemma-4-26B-A4B-APEX"))
})
})
func labels(ts []Tier) []string {
out := make([]string, 0, len(ts))
for _, t := range ts {
out = append(out, t.Label)
}
return out
}

130
.github/ci/apexentries/hf.go vendored Normal file
View File

@@ -0,0 +1,130 @@
package main
import (
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"strings"
"time"
"github.com/mudler/LocalAI/pkg/httpclient"
)
// ErrNoSHA256 marks a GGUF the HuggingFace API describes without an
// lfs.sha256. Emitting an entry without a hash would ship an unverifiable
// download, and guessing one from another field is how a Xet hash ends up
// masquerading as a content hash, so this is fatal rather than skippable.
var ErrNoSHA256 = errors.New("gguf file has no lfs.sha256")
// GGUFFile is one .gguf sibling of a HuggingFace repo.
type GGUFFile struct {
Name string
Size int64
SHA256 string
}
type apiSibling struct {
RFilename string `json:"rfilename"`
Size int64 `json:"size"`
LFS *struct {
SHA256 string `json:"sha256"`
} `json:"lfs"`
}
type apiModel struct {
Siblings []apiSibling `json:"siblings"`
}
// ParseRepoFiles returns every .gguf sibling described by a models API body.
func ParseRepoFiles(body []byte) ([]GGUFFile, error) {
var m apiModel
if err := json.Unmarshal(body, &m); err != nil {
return nil, fmt.Errorf("decoding model response: %w", err)
}
var out []GGUFFile
for _, s := range m.Siblings {
if !strings.HasSuffix(s.RFilename, ".gguf") {
continue
}
if s.LFS == nil || s.LFS.SHA256 == "" {
return nil, fmt.Errorf("%s: %w", s.RFilename, ErrNoSHA256)
}
out = append(out, GGUFFile{Name: s.RFilename, Size: s.Size, SHA256: s.LFS.SHA256})
}
return out, nil
}
// FetchOptionalRepoFiles asks the models API for a repo the caller can do
// without, and reports separately whether the repo was merely unreadable.
//
// HuggingFace answers 401 Unauthorized, not 404, for a repo that does not exist
// when the request carries no credentials. Without a token there is therefore no
// way to tell "this repo was never published" from "this repo is private", so an
// optional probe has to treat 401 and 403 exactly like 404: whatever the reason,
// there is nothing here for us to read, so there is no counterpart.
//
// The second return value exists because that collapse is lossy in one
// direction: 401/403 can also mean a real, gated repo whose quants we would
// genuinely want. The caller reports those repos so a silently dropped
// counterpart is visible to a human rather than invisible.
func FetchOptionalRepoFiles(client *http.Client, repo string) ([]GGUFFile, bool, error) {
files, status, err := fetchRepoFiles(client, repo)
if err != nil && (status == http.StatusUnauthorized || status == http.StatusForbidden) {
return nil, true, nil
}
return files, false, err
}
// FetchRepoFiles asks the models API for one repo. A 404 yields (nil, nil) so
// that probing for an optional counterpart repo is not an error. Every other
// non-200, 401 and 403 included, is an error: for a repo the run REQUIRES there
// is no benign reading of "we cannot see it".
func FetchRepoFiles(client *http.Client, repo string) ([]GGUFFile, error) {
files, _, err := fetchRepoFiles(client, repo)
return files, err
}
// fetchRepoFiles does the request and returns the HTTP status alongside the
// result, so the optional and required callers can apply different policies to
// the same response without duplicating the request.
func fetchRepoFiles(client *http.Client, repo string) ([]GGUFFile, int, error) {
url := fmt.Sprintf("https://huggingface.co/api/models/%s?blobs=true", repo)
req, err := http.NewRequest(http.MethodGet, url, nil)
if err != nil {
return nil, 0, err
}
req.Header.Set("User-Agent", "localai-apexentries/1.0")
resp, err := client.Do(req)
if err != nil {
return nil, 0, err
}
defer resp.Body.Close()
if resp.StatusCode == http.StatusNotFound {
return nil, resp.StatusCode, nil
}
if resp.StatusCode != http.StatusOK {
return nil, resp.StatusCode, fmt.Errorf("%s: unexpected status %d", repo, resp.StatusCode)
}
body, err := io.ReadAll(resp.Body)
if err != nil {
return nil, resp.StatusCode, err
}
files, err := ParseRepoFiles(body)
return files, resp.StatusCode, err
}
// newHTTPClient builds the client used against the HuggingFace API. It goes
// through pkg/httpclient rather than a bare &http.Client{} because the std
// client follows redirects and forwards custom credential headers to the
// redirect target on a cross-host hop (GHSA-3mj3-57v2-4636). This caller sends
// only a User-Agent today, but it talks to an external API that could start
// redirecting, and an HF_TOKEN header here later would then leak.
func newHTTPClient() *http.Client {
return httpclient.NewWithTimeout(60 * time.Second)
}

142
.github/ci/apexentries/hf_test.go vendored Normal file
View File

@@ -0,0 +1,142 @@
package main
import (
"bytes"
"io"
"net/http"
"testing"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
func TestApexEntries(t *testing.T) {
RegisterFailHandler(Fail)
RunSpecs(t, "apexentries")
}
// stubTransport answers every request with one canned status and body, so the
// status handling of the fetchers can be exercised without reaching the real
// HuggingFace API.
type stubTransport struct {
status int
body string
}
func (t stubTransport) RoundTrip(req *http.Request) (*http.Response, error) {
return &http.Response{
StatusCode: t.status,
Body: io.NopCloser(bytes.NewBufferString(t.body)),
Header: make(http.Header),
Request: req,
}, nil
}
func stubClient(status int, body string) *http.Client {
return &http.Client{Transport: stubTransport{status: status, body: body}}
}
const oneGGUFBody = `{"siblings":[{"rfilename":"Model-APEX-I-Quality.gguf","size":10,"lfs":{"sha256":"aa","size":10}}]}`
var _ = Describe("FetchOptionalRepoFiles", func() {
// HuggingFace answers 401 rather than 404 for a repo that does not exist
// when the client carries no credentials, so an optional probe cannot tell
// "absent" from "unauthorized" and must treat both as "no counterpart".
It("treats a 401 as an absent repo and flags it as unavailable", func() {
files, unavailable, err := FetchOptionalRepoFiles(stubClient(http.StatusUnauthorized, ""), "unsloth/Nope-GGUF")
Expect(err).ToNot(HaveOccurred())
Expect(files).To(BeEmpty())
Expect(unavailable).To(BeTrue())
})
It("treats a 403 as an absent repo and flags it as unavailable", func() {
files, unavailable, err := FetchOptionalRepoFiles(stubClient(http.StatusForbidden, ""), "unsloth/Gated-GGUF")
Expect(err).ToNot(HaveOccurred())
Expect(files).To(BeEmpty())
Expect(unavailable).To(BeTrue())
})
// A clean 404 is an unambiguous absence, so it must NOT be reported as
// unavailable: the whole point of the flag is to separate the ambiguous
// case a human may need to look at from the settled one.
It("treats a 404 as an absent repo without flagging it as unavailable", func() {
files, unavailable, err := FetchOptionalRepoFiles(stubClient(http.StatusNotFound, ""), "unsloth/Nope-GGUF")
Expect(err).ToNot(HaveOccurred())
Expect(files).To(BeEmpty())
Expect(unavailable).To(BeFalse())
})
It("parses a 200 body as usual", func() {
files, unavailable, err := FetchOptionalRepoFiles(stubClient(http.StatusOK, oneGGUFBody), "unsloth/Real-GGUF")
Expect(err).ToNot(HaveOccurred())
Expect(unavailable).To(BeFalse())
Expect(files).To(HaveLen(1))
Expect(files[0].Name).To(Equal("Model-APEX-I-Quality.gguf"))
Expect(files[0].SHA256).To(Equal("aa"))
})
// Tolerating 401/403 must not widen into tolerating everything: a 500 is a
// broken API, not evidence about whether the repo exists.
It("still errors on a 500", func() {
_, _, err := FetchOptionalRepoFiles(stubClient(http.StatusInternalServerError, ""), "unsloth/Real-GGUF")
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("unexpected status 500"))
})
})
var _ = Describe("FetchRepoFiles", func() {
// The APEX repo itself is not optional. A 401 there means the repo the run
// was asked to publish cannot be read, which is a real failure and must not
// be quietly downgraded to "no files".
It("errors on a 401 for a required repo", func() {
_, err := FetchRepoFiles(stubClient(http.StatusUnauthorized, ""), "mudler/Model-APEX-GGUF")
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("unexpected status 401"))
})
It("errors on a 403 for a required repo", func() {
_, err := FetchRepoFiles(stubClient(http.StatusForbidden, ""), "mudler/Model-APEX-GGUF")
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("unexpected status 403"))
})
It("still treats a 404 as an absent repo", func() {
files, err := FetchRepoFiles(stubClient(http.StatusNotFound, ""), "mudler/Model-APEX-GGUF")
Expect(err).ToNot(HaveOccurred())
Expect(files).To(BeEmpty())
})
})
var _ = Describe("ParseRepoFiles", func() {
It("returns gguf siblings with their lfs sha256", func() {
body := []byte(`{"siblings":[
{"rfilename":"Model-APEX-I-Quality.gguf","size":10,"lfs":{"sha256":"aa","size":10}},
{"rfilename":"README.md"},
{"rfilename":"mmproj.gguf","size":5,"lfs":{"sha256":"bb","size":5}}
]}`)
files, err := ParseRepoFiles(body)
Expect(err).ToNot(HaveOccurred())
Expect(files).To(HaveLen(2))
Expect(files[0].Name).To(Equal("Model-APEX-I-Quality.gguf"))
Expect(files[0].SHA256).To(Equal("aa"))
Expect(files[1].Name).To(Equal("mmproj.gguf"))
})
It("reports a gguf that carries no lfs sha256", func() {
body := []byte(`{"siblings":[{"rfilename":"mmproj.gguf","size":5}]}`)
_, err := ParseRepoFiles(body)
Expect(err).To(MatchError(ErrNoSHA256))
})
})

143
.github/ci/apexentries/hub.go vendored Normal file
View File

@@ -0,0 +1,143 @@
package main
import (
"fmt"
"os"
"strings"
"gopkg.in/yaml.v3"
"github.com/mudler/LocalAI/.github/ci/galleryedit"
)
// IndexText is the gallery index seen as text: the entries it declares plus the
// exact lines each one occupies, which is what splicing a variants block into an
// entry the gallery already ships requires.
//
// It is a second, narrower read of the same file LoadExisting parses. The two
// answer different questions: LoadExisting answers "do these weights already
// exist anywhere", this one answers "where in the file does this entry live".
type IndexText struct {
Lines []string
Entries []*indexEntry
byName map[string]*indexEntry
}
// indexEntry is one entry of the index: its name, the variants it already
// declares, and its coordinates in the file.
type indexEntry struct {
Name string `yaml:"name"`
Variants []VariantRef `yaml:"variants"`
Pos galleryedit.Entry `yaml:"-"`
}
// LoadIndexText reads the gallery index for editing.
func LoadIndexText(path string) (*IndexText, error) {
raw, err := os.ReadFile(path)
if err != nil {
return nil, err
}
return ParseIndexText(string(raw))
}
// ParseIndexText pairs the decoded entries with the top level list items the
// text actually contains.
//
// If the two views disagree on how many entries there are then every line number
// a splice would compute is suspect, and the failure mode is writing a variants
// block into the wrong model. The parse refuses instead.
func ParseIndexText(text string) (*IndexText, error) {
var entries []*indexEntry
if err := yaml.Unmarshal([]byte(text), &entries); err != nil {
return nil, fmt.Errorf("decoding gallery index: %w", err)
}
lines, starts := galleryedit.Scan(text)
if len(starts) != len(entries) {
return nil, fmt.Errorf("gallery index has %d decoded entries but %d top level list items; refusing to edit by line number",
len(entries), len(starts))
}
ix := &IndexText{Lines: lines, Entries: entries, byName: map[string]*indexEntry{}}
for i, e := range entries {
if e == nil {
return nil, fmt.Errorf("gallery index list item %d is empty; refusing to edit by line number", i)
}
end := len(lines)
if i+1 < len(starts) {
end = starts[i+1]
}
e.Pos = galleryedit.Entry{Name: e.Name, StartLine: starts[i], EndLine: end}
// First occurrence wins, matching the gallery's own resolution.
key := strings.ToLower(e.Name)
if _, seen := ix.byName[key]; !seen {
ix.byName[key] = e
}
}
return ix, nil
}
// Find looks an entry up by name, case insensitively.
func (ix *IndexText) Find(name string) *indexEntry {
return ix.byName[strings.ToLower(name)]
}
// ResolveHub returns the gallery name of a family's hub and whether the gallery
// already ships an entry under it.
//
// The hub is the BASE model entry, never a generated *-apex parent. Somebody
// looking for qwen3.6-35b-a3b has to find every build of those weights under
// that one name: the APEX imatrix rungs, the unsloth quant rungs and any
// speculative build. A separate qwen3.6-35b-a3b-apex hub competing with the base
// entry would split the family in two and leave whichever half the user did not
// search for invisible.
//
// Both candidates are tried for the same reason CounterpartCandidates tries
// both. The repo name and the published file stem disagree for several of these
// repos, and either one may be what the base entry was named after.
func ResolveHub(ix *IndexText, repoBase, stem string) (name string, exists bool) {
candidates := CounterpartCandidates(repoBase, stem)
for _, c := range candidates {
if n := slug(c); ix.Find(n) != nil {
return n, true
}
}
// Nothing matched, so the family needs a hub of its own under the repo
// derived name, which is the more reliable of the two.
return slug(candidates[0]), false
}
// HubLabel is the human-cased base model name, for prose rather than lookup.
func HubLabel(repoBase, stem string) string {
return CounterpartCandidates(repoBase, stem)[0]
}
// filterVariants drops the references a hub must not carry: itself, and anything
// it already lists.
//
// The self reference is not merely redundant. A hub that names itself makes the
// verifier resolve the reference back to the hub, see that the hub declares
// variants, and report a variant that declares variants of its own. It arises
// for real rather than in theory: an unsloth rung whose weights the gallery
// already ships under the base model name resolves, through Merge, straight back
// to the hub that is about to reference it.
func filterVariants(hub string, already []VariantRef, want []string) []string {
seen := map[string]bool{strings.ToLower(hub): true}
for _, v := range already {
seen[strings.ToLower(v.Model)] = true
}
var out []string
for _, w := range want {
key := strings.ToLower(w)
if seen[key] {
continue
}
seen[key] = true
out = append(out, w)
}
return out
}

795
.github/ci/apexentries/main.go vendored Normal file
View File

@@ -0,0 +1,795 @@
// Command apexentries generates gallery entries for the mudler APEX GGUF
// repositories: one entry per imatrix tier and per unsloth quant rung, all
// gathered under the BASE model's entry. Builds off a *-APEX-MTP-GGUF repo turn
// speculative decoding on, because those weights retain the model's MTP heads
// and are only worth their extra size with the heads in use.
//
// The base model entry is the hub. Somebody looking for qwen3.6-35b-a3b must
// find every build of those weights under that one name, so when the gallery
// already ships the base entry this command splices a variants block into it
// rather than emitting a competing *-apex parent beside it. Only a family whose
// base model the gallery does not ship at all gets a new hub entry, and that one
// is still named for the base model.
//
// Builds are discovered by inspecting the filenames a repo actually publishes.
// Repo names do not reliably predict them: mudler/gemma-4-26B-A4B-it-APEX-GGUF
// ships gemma-4-26B-A4B-APEX-*.gguf, and six of the 45 repos drop a suffix or a
// vendor prefix in the same way.
package main
import (
"encoding/json"
"flag"
"fmt"
"io"
"net/http"
"os"
"path"
"sort"
"strings"
"gopkg.in/yaml.v3"
"github.com/mudler/LocalAI/.github/ci/galleryedit"
)
const (
// entryTemplate carries no backend and no parameters of its own, which is
// why RenderChild states everything inline.
entryTemplate = "virtual.yaml"
unslothOwner = "unsloth"
authorListURL = "https://huggingface.co/api/models?author=mudler&limit=300"
)
// rungRank orders the quality ladder from best to smallest. The HuggingFace API
// returns siblings alphabetically and DiscoverAPEXTiers preserves that order, so
// an unsorted variants list reads I-Balanced, I-Compact, I-Mini, I-Nano,
// I-Quality. Selection ignores authored order, so this is purely so the file a
// human reviews scans in a meaningful sequence.
var rungRank = map[string]int{
"I-Quality": 0, "I-Balanced": 1, "I-Compact": 2, "I-Mini": 3, "I-Nano": 4,
"Quality": 5, "Balanced": 6, "Compact": 7, "Mini": 8, "Nano": 9,
}
// baseTags are the tags every generated entry carries. dflash and mtp are never
// among them: RenderChild adds those if and only if the entry configures the
// matching spec_type.
var baseTags = []string{"llm", "gguf", "cpu", "gpu"}
// childBuild pairs a rendered entry with its position on the quality ladder, so
// the parent's variants list can be sorted without re-parsing entry names.
type childBuild struct {
entry GalleryEntry
rank int
}
// family is one APEX repo's full generated output.
type family struct {
repo string
repoBase string
stem string
hasMMProj bool
children []childBuild
// skippedRepos are counterpart candidates HuggingFace would not describe.
// Carried on the family rather than printed and forgotten so the run can
// summarize them next to everything else a reviewer has to eyeball.
skippedRepos []string
census fileCensus
unaccounted int
}
// fileCensus splits the files discovery emitted nothing for into the ones a
// reviewer must chase and the ones that are deliberately out of scope.
//
// Full-precision sources are the second kind: they are the unquantized weights
// the ladder is derived FROM, not a rung of it. Folding them into the
// unclassified total would leave a permanent benign baseline, and a permanent
// baseline is exactly what hides the one file that ever genuinely matters.
type fileCensus struct {
unclassified int
fullPrecision int
}
// add accumulates one repo's census into a running total.
func (c *fileCensus) add(o fileCensus) {
c.unclassified += o.unclassified
c.fullPrecision += o.fullPrecision
}
// sortedChildren returns the family's builds in ladder order, best first.
func (f *family) sortedChildren() []childBuild {
sorted := append([]childBuild{}, f.children...)
sort.SliceStable(sorted, func(i, j int) bool { return sorted[i].rank < sorted[j].rank })
return sorted
}
func main() {
verify := flag.String("verify", "", "verify a gallery index and exit")
index := flag.String("index", "gallery/index.yaml", "gallery index to dedup against")
only := flag.String("only", "", "comma-separated repo names to restrict generation to")
out := flag.String("out", "", "write the entries to add to this file")
apply := flag.Bool("apply", false, "append the entries to add to -index")
flag.Parse()
if *verify != "" {
problems := Verify(*verify)
for _, p := range problems {
fmt.Fprintln(os.Stderr, p)
}
if len(problems) > 0 {
fmt.Fprintf(os.Stderr, "%d problem(s)\n", len(problems))
os.Exit(1)
}
fmt.Println("index is sound")
return
}
if err := generate(*index, *only, *out, *apply); err != nil {
fmt.Fprintln(os.Stderr, "error:", err)
os.Exit(1)
}
}
func generate(indexPath, only, outPath string, apply bool) error {
if outPath == "" && !apply {
return fmt.Errorf("nothing to do: pass -out <file> or -apply")
}
client := newHTTPClient()
repos, err := listAPEXRepos(client)
if err != nil {
return err
}
if only != "" {
repos = restrict(repos, only)
}
if len(repos) == 0 {
return fmt.Errorf("no APEX repos selected")
}
fmt.Printf("repos selected: %d\n", len(repos))
var families []family
var failed []string
for _, repo := range repos {
f, err := buildFamily(client, repo)
if err != nil {
// A missing sha256 is fatal for the family rather than skippable: an
// entry without one ships an unverifiable download. Report which repo
// and keep going, so one bad repo does not hide the state of the rest.
fmt.Fprintf(os.Stderr, "FAILED %s: %v\n", repo, err)
failed = append(failed, repo)
continue
}
families = append(families, *f)
}
existing, err := LoadExisting(indexPath)
if err != nil {
return err
}
ixText, err := LoadIndexText(indexPath)
if err != nil {
return err
}
fmt.Printf("existing index: %d names, %d weight URIs, %d lines\n",
len(existing.ByName), len(existing.ByURI), len(ixText.Lines))
// Only the builds go through Merge. A hub is deliberately kept out of it: a
// new hub carries the family's top rung as its own payload, so Merge's URI
// dedup would fold the hub into that rung and the family would lose the very
// entry point this command exists to create. Hub names are checked against
// the index directly, by ResolveHub.
var generated []GalleryEntry
for _, f := range families {
for _, c := range f.children {
generated = append(generated, c.entry)
}
}
add, reused := Merge(existing, generated)
reportReuse(existing, generated, reused)
// Variant references are resolved from `reused`, never used to decide what to
// emit: on a within-batch name collision Merge records reused[name] = name
// while the first entry of that name is still in `add`, so treating presence
// in `reused` as "dropped" would silently emit nothing for it.
added := map[string]bool{}
for _, e := range add {
added[e.Name] = true
}
inserts, newHubs, err := planHubs(families, ixText, reused, added)
if err != nil {
return err
}
reportHubs(ixText, inserts, newHubs)
skipped, census, fullPrecisionRepos, unaccounted := reportSkipped(families)
add = append(add, newHubs...)
fmt.Printf("\nentries generated: %d\nentries to add: %d\nentries reused: %d\nhubs spliced: %d\nhubs created: %d\nrepos skipped: %d\nexcluded (full precision): %d files across %d repos\nunclassified: %d\nunaccounted: %d\n",
len(generated), len(add), len(reused), len(inserts), len(newHubs), len(skipped),
census.fullPrecision, fullPrecisionRepos, census.unclassified, unaccounted)
lines, err := galleryedit.Apply(ixText.Lines, inserts)
if err != nil {
return err
}
if err := writeEntries(add, lines, outPath, apply, indexPath); err != nil {
return err
}
if len(failed) > 0 {
return fmt.Errorf("%d repo(s) failed: %s", len(failed), strings.Join(failed, ", "))
}
return nil
}
// resolveVariant maps a generated child name onto whatever entry actually stands
// for it after the merge. `added` is consulted first because a within-batch name
// collision puts a name in BOTH add and reused, and the entry that was emitted
// is the one the parent must reference.
func resolveVariant(name string, reused map[string]string, added map[string]bool) string {
if added[name] {
return name
}
if target, ok := reused[name]; ok {
return target
}
return name
}
// SpecTypeForRepo reports the speculative decoding mechanism a repo's builds can
// turn on with no extra download.
//
// The *-APEX-MTP-GGUF repos republish the base weights with the model's own MTP
// heads retained, so those builds are only worth their extra size if the heads
// are actually used. Every other APEX repo drops them, and switching MTP on
// there would name a mechanism the weights cannot serve.
//
// The suffix is read off the repo the FILES come from, so nothing downstream has
// to infer a capability from an entry name.
func SpecTypeForRepo(repo string) string {
if strings.HasSuffix(path.Base(repo), "-APEX-MTP-GGUF") {
return "draft-mtp"
}
return ""
}
// buildFamily discovers everything one APEX repo and its unsloth counterpart
// publish, and renders it.
func buildFamily(client *http.Client, repo string) (*family, error) {
files, err := FetchRepoFiles(client, repo)
if err != nil {
return nil, err
}
if len(files) == 0 {
return nil, fmt.Errorf("no gguf files")
}
imatrix, plain := DiscoverAPEXTiers(files)
mmproj, hasMMProj := DiscoverMMProj(files)
census := reportUnclassified(repo, files, imatrix, plain)
// The imatrix ladder is preferred, but two of the 45 repos publish no
// imatrix tiers at all and must still contribute their plain ladder.
ladder := imatrix
ladderKind := "imatrix"
if len(ladder) == 0 {
ladder = plain
ladderKind = "plain"
}
if len(ladder) == 0 {
return nil, fmt.Errorf("no tiers discovered")
}
sortTiers(ladder)
var mm *GGUFFile
if hasMMProj {
mm = &mmproj
}
repoBase := strings.TrimSuffix(path.Base(repo), "-GGUF")
f := &family{repo: repo, repoBase: repoBase, hasMMProj: hasMMProj, census: census}
// Only the APEX ladder can carry MTP heads; the unsloth counterpart quantizes
// the plain weights and gets nothing from this.
specType := SpecTypeForRepo(repo)
for _, t := range ladder {
f.children = append(f.children, childBuild{
rank: rungRank[t.Label],
entry: RenderChild(ChildInput{
Name: slug(repoBase) + "-" + slug(t.Label),
Repo: repo,
Template: entryTemplate,
SpecType: specType,
Weights: []GGUFFile{t.File},
MMProj: mm,
BaseTags: baseTags,
}),
})
}
stem := FileStem(ladder[0])
f.stem = stem
fmt.Printf("%s: %d %s tier(s) [%s], stem %s, mmproj %v\n",
repo, len(ladder), ladderKind, tierLabels(ladder), stem, hasMMProj)
counterpart, cpFiles, skipped, err := resolveCounterpart(client, repoBase, stem)
f.skippedRepos = skipped
if err != nil {
return nil, err
}
if counterpart != "" {
builds := DiscoverUnslothQuants(cpFiles)
// Called here rather than inside Verify: a quant dropped at discovery
// leaves no trace at all in the finished gallery file, so the only place
// the shortfall is still visible is the moment of discovery.
unaccounted := UnaccountedQuants(cpFiles, builds)
f.unaccounted = len(unaccounted)
for _, p := range unaccounted {
fmt.Fprintf(os.Stderr, "UNACCOUNTED QUANT %s: %s\n", counterpart, p)
}
cpMMProj, hasCPMMProj := DiscoverMMProj(cpFiles)
var cpMM *GGUFFile
if hasCPMMProj {
cpMM = &cpMMProj
}
cpBase := strings.TrimSuffix(path.Base(counterpart), "-GGUF")
for i, b := range builds {
f.children = append(f.children, childBuild{
rank: 100 + i,
entry: RenderChild(ChildInput{
Name: slug(cpBase) + "-" + slug(b.Quant),
Repo: counterpart,
Template: entryTemplate,
Weights: b.Files,
MMProj: cpMM,
BaseTags: baseTags,
}),
})
}
fmt.Printf("%s: counterpart %s, %d quant build(s) %s\n", repo, counterpart, len(builds), quantLabels(builds))
} else {
fmt.Printf("%s: no unsloth counterpart\n", repo)
}
return f, nil
}
// planHubs decides, per family, whether the family's builds are spliced into a
// base model entry the gallery already ships or gathered under a new hub.
//
// Splicing is strongly preferred and is the measured majority-adjacent case. The
// existing entry keeps its description, icon, tags, overrides and files
// untouched; only variant lines are added to it.
func planHubs(families []family, ix *IndexText, reused map[string]string, added map[string]bool) ([]galleryedit.Insert, []GalleryEntry, error) {
// Several APEX repos can resolve to one base model, so both paths accumulate
// by hub name rather than assuming one family per hub.
wantByHub := map[string][]string{}
var spliceOrder []string
var newHubs []GalleryEntry
hubAt := map[string]int{}
for i := range families {
f := &families[i]
hubName, exists := ResolveHub(ix, f.repoBase, f.stem)
want := hubVariants(f, ix, reused, added)
if exists {
if _, seen := wantByHub[hubName]; !seen {
spliceOrder = append(spliceOrder, hubName)
}
wantByHub[hubName] = append(wantByHub[hubName], want...)
continue
}
if at, dup := hubAt[hubName]; dup {
for _, v := range filterVariants(hubName, newHubs[at].Variants, want) {
newHubs[at].Variants = append(newHubs[at].Variants, VariantRef{Model: v})
}
continue
}
builds := f.sortedChildren()
if len(builds) == 0 {
return nil, nil, fmt.Errorf("%s: no builds to hang a hub on", f.repo)
}
hubAt[hubName] = len(newHubs)
newHubs = append(newHubs, renderHub(hubName, f, builds[0], filterVariants(hubName, nil, want)))
}
var inserts []galleryedit.Insert
for _, name := range spliceOrder {
e := ix.Find(name)
items := filterVariants(name, e.Variants, wantByHub[name])
if len(items) == 0 {
continue
}
inserts = append(inserts, galleryedit.Insert{Entry: e.Pos, Variants: items})
}
return inserts, newHubs, nil
}
// hubVariants is a family's full build list, in ladder order, named as the hub
// must reference them after the merge.
func hubVariants(f *family, ix *IndexText, reused map[string]string, added map[string]bool) []string {
var out []string
// A hand-written *-apex entry is an ordinary build of these weights. It is
// never deleted, never renamed and never treated as a hub; it is simply
// referenced like any other rung.
if apex := slug(f.repoBase); ix.Find(apex) != nil {
out = append(out, apex)
}
for _, c := range f.sortedChildren() {
out = append(out, resolveVariant(c.entry.Name, reused, added))
}
return out
}
// renderHub builds the hub for a family whose base model the gallery does not
// ship at all. It is named for the BASE model, never for the APEX repo.
//
// It carries one of the discovered builds as its own payload so it is a complete
// installable entry rather than a bare index pointing at other entries. That
// payload is what supplies overrides.backend, which matters beyond installation:
// the verifier can only judge the tagging rule for a backend it can read, so a
// hub carrying feature tags and no backend would escape the check in silence.
//
// The payload's own tags are kept rather than rebuilt from baseTags, so a hub
// whose payload configures a spec_type stays tagged for it and consistent with
// the overrides copied alongside.
func renderHub(name string, f *family, payload childBuild, variants []string) GalleryEntry {
e := payload.entry
e.Name = name
e.Description = fmt.Sprintf(
"%s. Quality ladder and quantization rungs published by %s and its unsloth counterpart; LocalAI picks the build that fits the hardware.",
HubLabel(f.repoBase, f.stem), f.repo)
e.Tags = append([]string{}, payload.entry.Tags...)
if f.hasMMProj && !hasTag(e.Tags, "vision") {
e.Tags = append(e.Tags, "vision")
}
e.Variants = nil
for _, v := range variants {
e.Variants = append(e.Variants, VariantRef{Model: v})
}
return e
}
func hasTag(tags []string, want string) bool {
for _, t := range tags {
if t == want {
return true
}
}
return false
}
// resolveCounterpart probes the unsloth candidates in order and returns the
// first that publishes files.
//
// CounterpartCandidates is handed a BARE repo name: its cleaner does not strip
// an owner prefix, so passing "mudler/Foo-APEX-GGUF" would yield "mudler/Foo"
// and compose into the nonsense probe "unsloth/mudler/Foo".
//
// It also returns the candidates HuggingFace refused to describe. Those are
// indistinguishable from absent without credentials, so they are skipped, but
// they are named rather than dropped: one of them could be a real gated repo
// whose quants belong in the gallery.
func resolveCounterpart(client *http.Client, repoBase, stem string) (string, []GGUFFile, []string, error) {
var unavailable []string
for _, cand := range CounterpartCandidates(repoBase, stem) {
repo := unslothOwner + "/" + cand + "-GGUF"
files, unreadable, err := FetchOptionalRepoFiles(client, repo)
if err != nil {
return "", nil, unavailable, fmt.Errorf("probing %s: %w", repo, err)
}
if unreadable {
unavailable = append(unavailable, repo)
continue
}
if len(files) > 0 {
return repo, files, unavailable, nil
}
}
return "", nil, unavailable, nil
}
// reportUnclassified prints the files discovery turned into nothing.
//
// It is a set difference on COUNTS, not a re-match of filenames: re-matching
// would duplicate the tier regex from discover.go and the two copies would
// drift. The likeliest trigger is a typo or case change from a publishing script
// rather than a genuine sixth tier, and because generation falls back to the
// plain ladder when the imatrix one is empty, a repo whose imatrix files all
// fail to match silently downgrades the whole family instead of erroring. The
// downstream HTTP check cannot catch that: it validates URLs that were emitted,
// and an undiscovered tier emits none.
// It returns the census so the run can total it.
func reportUnclassified(repo string, files []GGUFFile, imatrix, plain []Tier) fileCensus {
mmprojCount, fullPrecision := 0, 0
for _, f := range files {
// The mmproj test comes first because projectors are themselves often
// published at f16 (mmproj-F16.gguf), and counting such a file in both
// buckets would understate the unclassified remainder.
if strings.HasPrefix(f.Name, "mmproj") {
mmprojCount++
continue
}
if IsFullPrecision(f.Name) {
fullPrecision++
}
}
classified := len(imatrix) + len(plain) + mmprojCount + fullPrecision
if classified >= len(files) {
return fileCensus{fullPrecision: fullPrecision}
}
fmt.Fprintf(os.Stderr, "UNCLASSIFIED %s: %d of %d .gguf files classified, %d unaccounted for\n",
repo, classified, len(files), len(files)-classified)
return fileCensus{unclassified: len(files) - classified, fullPrecision: fullPrecision}
}
// reportReuse splits Merge's single reused map into the two cases it conflates.
//
// A URI match means the gallery already ships exactly these weights, and
// pointing the parent at the existing entry is correct. A NAME match with a
// different URI means an unrelated entry happens to own the name, and
// referencing it would point the parent at different weights than were
// generated, substituting a build without saying so. Only the first is safe to
// wave through.
func reportReuse(existing *ExistingIndex, generated []GalleryEntry, reused map[string]string) {
byName := map[string]GalleryEntry{}
for _, e := range generated {
if _, seen := byName[e.Name]; !seen {
byName[e.Name] = e
}
}
var nameCollisions, uriMatches []string
for name, target := range reused {
gen := byName[name]
uri := ""
if len(gen.Files) > 0 {
uri = gen.Files[0].URI
}
switch {
case hasName(existing, name):
nameCollisions = append(nameCollisions,
fmt.Sprintf(" %s -> gallery entry of the same name (generated uri: %s)", name, orNone(uri)))
case target == name:
nameCollisions = append(nameCollisions,
fmt.Sprintf(" %s -> earlier entry of the same name in this batch (generated uri: %s)", name, orNone(uri)))
default:
uriMatches = append(uriMatches, fmt.Sprintf(" %s -> %s (same weights: %s)", name, target, orNone(uri)))
}
}
sort.Strings(nameCollisions)
sort.Strings(uriMatches)
fmt.Printf("\nNAME COLLISIONS (%d) - inspect each by hand, the target may hold different weights\n", len(nameCollisions))
for _, l := range nameCollisions {
fmt.Println(l)
}
fmt.Printf("\nURI MATCHES (%d) - the gallery or this batch already ships these exact weights\n", len(uriMatches))
for _, l := range uriMatches {
fmt.Println(l)
}
}
// reportHubs prints exactly what will be written where. The splices are the part
// a human has to read: they modify entries the gallery already ships, so the
// review needs the target, the line, and every added reference spelled out.
func reportHubs(ix *IndexText, inserts []galleryedit.Insert, newHubs []GalleryEntry) {
fmt.Printf("\nHUBS SPLICED (%d) - variants added to the EXISTING base model entry, nothing else touched\n", len(inserts))
for _, in := range inserts {
e := ix.Find(in.Entry.Name)
fmt.Printf(" %s (line %d, %d variant(s) already declared):\n", in.Entry.Name, in.Entry.StartLine+1, len(e.Variants))
for _, v := range in.Variants {
fmt.Printf(" + - model: %s\n", galleryedit.QuoteName(v))
}
}
fmt.Printf("\nHUBS CREATED (%d) - the gallery ships no base model entry, so one is emitted for it\n", len(newHubs))
for _, h := range newHubs {
fmt.Printf(" %s:\n", h.Name)
for _, v := range h.Variants {
fmt.Printf(" - model: %s\n", v.Model)
}
}
}
// reportSkipped names the counterpart repos HuggingFace would not describe, and
// totals the other two silent-shortfall counters alongside them.
//
// A skipped repo is not the same as a clean 404. HuggingFace answers 401 for a
// nonexistent repo to an unauthenticated client, so the overwhelmingly likely
// reading is "there is no such counterpart", which is the normal case for the
// community merges. But a private or gated repo answers 401 too, and that one
// WOULD have quants worth shipping. Printing the list is what keeps that
// possibility auditable instead of silently discarded.
func reportSkipped(families []family) ([]string, fileCensus, int, int) {
var skipped []string
var census fileCensus
fullPrecisionRepos, unaccounted := 0, 0
for _, f := range families {
skipped = append(skipped, f.skippedRepos...)
census.add(f.census)
if f.census.fullPrecision > 0 {
fullPrecisionRepos++
}
unaccounted += f.unaccounted
}
sort.Strings(skipped)
fmt.Printf("\nREPOS SKIPPED AS UNAVAILABLE (%d) - HuggingFace answered 401/403, which is indistinguishable from absent without a token; check none of these is a real gated repo\n", len(skipped))
for _, r := range skipped {
fmt.Printf(" %s\n", r)
}
return skipped, census, fullPrecisionRepos, unaccounted
}
func hasName(ix *ExistingIndex, name string) bool {
_, ok := ix.ByName[name]
return ok
}
func orNone(s string) string {
if s == "" {
return "(no files)"
}
return s
}
// writeEntries emits the additions.
//
// -apply does two things in one pass: it writes back the spliced lines, which
// differ from the original only by the variant lines galleryedit inserted, and
// then appends the new entries. New entries are APPENDED rather than merged into
// the structure, for the same reason the splice is textual: a YAML round trip
// over 40,000 lines would reflow the whole file into an unreviewable diff.
func writeEntries(add []GalleryEntry, lines []string, outPath string, apply bool, indexPath string) error {
if apply {
if err := os.WriteFile(indexPath, []byte(strings.Join(lines, "\n")), 0o644); err != nil {
return err
}
fmt.Printf("spliced %s\n", indexPath)
}
if len(add) == 0 {
fmt.Println("nothing to append")
return nil
}
blob, err := yaml.Marshal(add)
if err != nil {
return err
}
if outPath != "" {
if err := os.WriteFile(outPath, blob, 0o644); err != nil {
return err
}
fmt.Printf("wrote %d entries to %s\n", len(add), outPath)
}
if apply {
f, err := os.OpenFile(indexPath, os.O_APPEND|os.O_WRONLY, 0o644)
if err != nil {
return err
}
defer f.Close()
if _, err := f.Write(blob); err != nil {
return err
}
fmt.Printf("appended %d entries to %s\n", len(add), indexPath)
}
return nil
}
// listAPEXRepos returns the mudler repos whose name marks them as APEX builds.
func listAPEXRepos(client *http.Client) ([]string, error) {
req, err := http.NewRequest(http.MethodGet, authorListURL, nil)
if err != nil {
return nil, err
}
req.Header.Set("User-Agent", "localai-apexentries/1.0")
resp, err := client.Do(req)
if err != nil {
return nil, err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return nil, fmt.Errorf("listing models: unexpected status %d", resp.StatusCode)
}
body, err := io.ReadAll(resp.Body)
if err != nil {
return nil, err
}
var models []struct {
ID string `json:"id"`
}
if err := json.Unmarshal(body, &models); err != nil {
return nil, fmt.Errorf("decoding model list: %w", err)
}
var out []string
for _, m := range models {
if strings.Contains(m.ID, "APEX") {
out = append(out, m.ID)
}
}
sort.Strings(out)
return out, nil
}
func restrict(repos []string, only string) []string {
want := map[string]bool{}
for _, r := range strings.Split(only, ",") {
if r = strings.TrimSpace(r); r != "" {
want[r] = true
}
}
var out []string
for _, r := range repos {
if want[r] {
out = append(out, r)
delete(want, r)
}
}
// A name in -only that matched nothing is a typo, not an empty result.
for r := range want {
fmt.Fprintf(os.Stderr, "WARNING: -only names %s, which is not an APEX repo of this author\n", r)
}
return out
}
func sortTiers(tiers []Tier) {
sort.SliceStable(tiers, func(i, j int) bool { return rungRank[tiers[i].Label] < rungRank[tiers[j].Label] })
}
func tierLabels(tiers []Tier) string {
var out []string
for _, t := range tiers {
out = append(out, t.Label)
}
return strings.Join(out, ",")
}
func quantLabels(builds []QuantBuild) string {
var out []string
for _, b := range builds {
l := b.Quant
if b.Sharded {
l += fmt.Sprintf("(%d shards)", len(b.Files))
}
out = append(out, l)
}
return strings.Join(out, ",")
}
// slug turns a repo, tier or quant label into a gallery entry name component.
func slug(s string) string {
return strings.ReplaceAll(strings.ToLower(s), "_", "-")
}

344
.github/ci/apexentries/main_test.go vendored Normal file
View File

@@ -0,0 +1,344 @@
package main
import (
"strings"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"github.com/mudler/LocalAI/.github/ci/galleryedit"
)
func mustIndex(text string) *IndexText {
ix, err := ParseIndexText(text)
ExpectWithOffset(1, err).ToNot(HaveOccurred())
return ix
}
// buildOf renders a realistic child so the specs exercise the payload a hub
// actually inherits rather than a bare name.
func buildOf(name, repo, file string, rank int) childBuild {
return childBuild{
rank: rank,
entry: RenderChild(ChildInput{
Name: name,
Repo: repo,
Template: entryTemplate,
Weights: []GGUFFile{{Name: file, SHA256: "aa"}},
BaseTags: baseTags,
}),
}
}
var _ = Describe("ResolveHub", func() {
It("picks the base model name over the APEX name, even when both are in the gallery", func() {
// The hub is the entry a user searches for. If the *-apex entry were
// chosen the family would be gathered under a name nobody looks up, and
// the base entry would go on advertising only its own build.
ix := mustIndex("- name: qwen3.6-35b-a3b\n url: u\n- name: qwen3.6-35b-a3b-apex\n url: u\n")
name, exists := ResolveHub(ix, "Qwen3.6-35B-A3B-APEX", "Qwen3.6-35B-A3B-APEX")
Expect(name).To(Equal("qwen3.6-35b-a3b"))
Expect(exists).To(BeTrue())
})
It("falls back to the stem-derived candidate when the repo-derived one is absent", func() {
// gemma's repo says "-it" and its published files do not, so only one of
// the two candidates can match whatever the base entry was named after.
ix := mustIndex("- name: gemma-4-26b-a4b\n url: u\n")
name, exists := ResolveHub(ix, "gemma-4-26B-A4B-it-APEX", "gemma-4-26B-A4B-APEX")
Expect(name).To(Equal("gemma-4-26b-a4b"))
Expect(exists).To(BeTrue())
})
It("reports the base name as absent rather than settling for the APEX entry", func() {
ix := mustIndex("- name: qwen3.5-35b-a3b-apex\n url: u\n")
name, exists := ResolveHub(ix, "Qwen3.5-35B-A3B-APEX", "Qwen3.5-35B-A3B-APEX")
Expect(name).To(Equal("qwen3.5-35b-a3b"))
Expect(exists).To(BeFalse())
})
It("strips the MTP and TQ markers as well as APEX", func() {
ix := mustIndex("- name: qwen3.6-35b-a3b\n url: u\n")
name, exists := ResolveHub(ix, "Qwen3.6-35B-A3B-APEX-MTP", "Qwen3.6-35B-A3B-APEX-MTP")
Expect(name).To(Equal("qwen3.6-35b-a3b"))
Expect(exists).To(BeTrue())
})
})
var _ = Describe("planHubs", func() {
noReuse := map[string]string{}
allAdded := func(names ...string) map[string]bool {
out := map[string]bool{}
for _, n := range names {
out[n] = true
}
return out
}
It("splices into the existing base entry instead of emitting an *-apex parent", func() {
ix := mustIndex("- name: step-3.7-flash\n url: u\n- name: other\n url: u\n")
fams := []family{{
repo: "mudler/Step-3.7-Flash-APEX-GGUF",
repoBase: "Step-3.7-Flash-APEX",
stem: "Step-3.7-Flash-APEX",
children: []childBuild{buildOf("step-3.7-flash-apex-i-quality", "mudler/Step-3.7-Flash-APEX-GGUF", "a.gguf", 0)},
}}
inserts, newHubs, err := planHubs(fams, ix, noReuse, allAdded("step-3.7-flash-apex-i-quality"))
Expect(err).ToNot(HaveOccurred())
Expect(newHubs).To(BeEmpty())
Expect(inserts).To(HaveLen(1))
Expect(inserts[0].Entry.Name).To(Equal("step-3.7-flash"))
Expect(inserts[0].Variants).To(Equal([]string{"step-3.7-flash-apex-i-quality"}))
})
It("merges into an entry that already declares variants, without repeating one", func() {
// The gallery's qwen3.6-35b-a3b already lists its APEX build. Re-adding it
// would put a duplicate key's worth of noise in the diff and a duplicate
// reference in the entry.
ix := mustIndex("- name: qwen3.6-35b-a3b\n variants:\n - model: qwen3.6-35b-a3b-apex\n url: u\n" +
"- name: qwen3.6-35b-a3b-apex\n url: u\n")
fams := []family{{
repo: "mudler/Qwen3.6-35B-A3B-APEX-GGUF",
repoBase: "Qwen3.6-35B-A3B-APEX",
stem: "Qwen3.6-35B-A3B-APEX",
children: []childBuild{buildOf("qwen3.6-35b-a3b-apex-i-quality", "mudler/Qwen3.6-35B-A3B-APEX-GGUF", "a.gguf", 0)},
}}
inserts, newHubs, err := planHubs(fams, ix, noReuse, allAdded("qwen3.6-35b-a3b-apex-i-quality"))
Expect(err).ToNot(HaveOccurred())
Expect(newHubs).To(BeEmpty())
Expect(inserts[0].Variants).To(Equal([]string{"qwen3.6-35b-a3b-apex-i-quality"}))
out, err := galleryedit.Apply(ix.Lines, inserts)
Expect(err).ToNot(HaveOccurred())
Expect(strings.Count(strings.Join(out, "\n"), "variants:")).To(Equal(1))
Expect(out).To(HaveLen(len(ix.Lines) + 1))
})
It("never lets the hub reference itself", func() {
// An unsloth rung whose weights the gallery already ships under the base
// name resolves, through Merge, straight back to the hub. The verifier
// reads a self reference as a variant that declares variants of its own.
ix := mustIndex("- name: step-3.7-flash\n url: u\n")
fams := []family{{
repo: "mudler/Step-3.7-Flash-APEX-GGUF",
repoBase: "Step-3.7-Flash-APEX",
stem: "Step-3.7-Flash-APEX",
children: []childBuild{buildOf("step-3.7-flash-ud-q4-k-m", "unsloth/Step-3.7-Flash-GGUF", "a.gguf", 100)},
}}
inserts, _, err := planHubs(fams, ix, map[string]string{"step-3.7-flash-ud-q4-k-m": "step-3.7-flash"}, map[string]bool{})
Expect(err).ToNot(HaveOccurred())
Expect(inserts).To(BeEmpty())
})
It("emits a hub named for the base model when the gallery has none", func() {
ix := mustIndex("- name: qwen3.5-35b-a3b-apex\n url: u\n")
fams := []family{{
repo: "mudler/Qwen3.5-35B-A3B-APEX-GGUF",
repoBase: "Qwen3.5-35B-A3B-APEX",
stem: "Qwen3.5-35B-A3B-APEX",
hasMMProj: true,
children: []childBuild{
buildOf("qwen3.5-35b-a3b-apex-i-quality", "mudler/Qwen3.5-35B-A3B-APEX-GGUF", "a.gguf", 0),
buildOf("qwen3.5-35b-a3b-ud-q6-k", "unsloth/Qwen3.5-35B-A3B-GGUF", "b.gguf", 102),
},
}}
inserts, newHubs, err := planHubs(fams, ix, noReuse,
allAdded("qwen3.5-35b-a3b-apex-i-quality", "qwen3.5-35b-a3b-ud-q6-k"))
Expect(err).ToNot(HaveOccurred())
Expect(inserts).To(BeEmpty())
Expect(newHubs).To(HaveLen(1))
hub := newHubs[0]
Expect(hub.Name).To(Equal("qwen3.5-35b-a3b"))
Expect(hub.Name).ToNot(HaveSuffix("-apex"))
// A hand-written *-apex entry is an ordinary build, referenced like any
// other rung and never deleted or renamed.
Expect(hub.Variants).To(Equal([]VariantRef{
{Model: "qwen3.5-35b-a3b-apex"},
{Model: "qwen3.5-35b-a3b-apex-i-quality"},
{Model: "qwen3.5-35b-a3b-ud-q6-k"},
}))
// The verifier skips entries with no declared backend, so a hub without
// one would escape the tagging check in silence.
Expect(hub.Overrides).To(HaveKeyWithValue("backend", "llama-cpp"))
Expect(hub.Files).ToNot(BeEmpty())
Expect(hub.Tags).To(ContainElement("vision"))
})
It("gathers two APEX repos that share one base model under a single hub", func() {
ix := mustIndex("- name: unrelated\n url: u\n")
fams := []family{
{
repo: "mudler/Solo-APEX-GGUF",
repoBase: "Solo-APEX",
stem: "Solo-APEX",
children: []childBuild{buildOf("solo-apex-i-quality", "mudler/Solo-APEX-GGUF", "a.gguf", 0)},
},
{
repo: "mudler/Solo-APEX-MTP-GGUF",
repoBase: "Solo-APEX-MTP",
stem: "Solo-APEX-MTP",
children: []childBuild{buildOf("solo-apex-mtp-i-quality", "mudler/Solo-APEX-MTP-GGUF", "b.gguf", 0)},
},
}
_, newHubs, err := planHubs(fams, ix, noReuse, allAdded("solo-apex-i-quality", "solo-apex-mtp-i-quality"))
Expect(err).ToNot(HaveOccurred())
Expect(newHubs).To(HaveLen(1))
Expect(newHubs[0].Name).To(Equal("solo"))
Expect(newHubs[0].Variants).To(Equal([]VariantRef{
{Model: "solo-apex-i-quality"},
{Model: "solo-apex-mtp-i-quality"},
}))
})
})
var _ = Describe("hubVariants", func() {
It("orders builds by quality rung rather than discovery order", func() {
// DiscoverAPEXTiers preserves input order and the HF API returns siblings
// alphabetically, so an unsorted list reads I-Balanced, I-Compact, I-Mini,
// I-Nano, I-Quality. Selection ignores authored order; this is for the
// human reading the file.
f := family{repoBase: "X-APEX", stem: "X-APEX", children: []childBuild{
{rank: rungRank["I-Nano"], entry: GalleryEntry{Name: "x-i-nano"}},
{rank: 100, entry: GalleryEntry{Name: "x-ud-q4-k-m"}},
{rank: rungRank["I-Quality"], entry: GalleryEntry{Name: "x-i-quality"}},
{rank: rungRank["I-Compact"], entry: GalleryEntry{Name: "x-i-compact"}},
}}
got := hubVariants(&f, mustIndex("- name: x\n url: u\n"), map[string]string{}, map[string]bool{})
Expect(got).To(Equal([]string{"x-i-quality", "x-i-compact", "x-i-nano", "x-ud-q4-k-m"}))
})
})
var _ = Describe("ParseIndexText", func() {
It("refuses to edit by line number when the two views of the file disagree", func() {
_, err := ParseIndexText("- name: one\n url: u\n-\n")
Expect(err).To(MatchError(ContainSubstring("empty")))
})
It("records the line range of each entry", func() {
ix := mustIndex("- name: first\n url: u\n- name: second\n url: u\n")
Expect(ix.Find("FIRST").Pos.StartLine).To(Equal(0))
Expect(ix.Find("first").Pos.EndLine).To(Equal(2))
Expect(ix.Find("second").Pos.StartLine).To(Equal(2))
})
})
var _ = Describe("resolveVariant", func() {
It("keeps an entry that was emitted even when it is also in reused", func() {
// A within-batch name collision records reused[name] = name while the
// FIRST entry of that name is still in add. Treating presence in reused as
// "dropped" would emit nothing for it.
added := map[string]bool{"dup": true}
reused := map[string]string{"dup": "dup"}
Expect(resolveVariant("dup", reused, added)).To(Equal("dup"))
})
It("redirects a reused name at the entry that stands in for it", func() {
added := map[string]bool{}
reused := map[string]string{"generated": "already-in-gallery"}
Expect(resolveVariant("generated", reused, added)).To(Equal("already-in-gallery"))
})
})
var _ = Describe("slug", func() {
It("lowercases and turns quant underscores into hyphens", func() {
Expect(slug("UD-Q4_K_M")).To(Equal("ud-q4-k-m"))
Expect(slug("gemma-4-26B-A4B-it-APEX")).To(Equal("gemma-4-26b-a4b-it-apex"))
Expect(slug("I-Nano")).To(Equal("i-nano"))
})
})
var _ = Describe("sortTiers", func() {
It("puts the imatrix ladder in descending quality order", func() {
tiers := []Tier{
{Label: "I-Balanced"}, {Label: "I-Compact"}, {Label: "I-Mini"},
{Label: "I-Nano"}, {Label: "I-Quality"},
}
sortTiers(tiers)
Expect(tierLabels(tiers)).To(Equal("I-Quality,I-Balanced,I-Compact,I-Mini,I-Nano"))
})
})
var _ = Describe("restrict", func() {
It("keeps only the named repos", func() {
got := restrict([]string{"mudler/A-APEX-GGUF", "mudler/B-APEX-GGUF"}, "mudler/B-APEX-GGUF")
Expect(got).To(Equal([]string{"mudler/B-APEX-GGUF"}))
})
It("returns nothing when the filter matches nothing", func() {
Expect(restrict([]string{"mudler/A-APEX-GGUF"}, "mudler/typo")).To(BeEmpty())
})
})
var _ = Describe("reportUnclassified", func() {
// One real imatrix rung is always present so the specs measure how the
// remaining files are bucketed, not an empty-repo edge case.
tier := Tier{Label: "I-Quality", File: GGUFFile{Name: "Model-APEX-I-Quality.gguf"}}
censusOf := func(names ...string) fileCensus {
files := []GGUFFile{tier.File}
for _, n := range names {
files = append(files, GGUFFile{Name: n})
}
return reportUnclassified("mudler/Model-APEX-GGUF", files, []Tier{tier}, nil)
}
It("counts a flat full-precision source as excluded, not unclassified", func() {
got := censusOf("Carnice-MoE-35B-A3B-F16.gguf")
Expect(got.fullPrecision).To(Equal(1))
Expect(got.unclassified).To(Equal(0))
})
It("counts every shard of a sharded full-precision source as excluded", func() {
got := censusOf(
"MiniMax-M2.7-APEX-F16-00001-of-00003.gguf",
"MiniMax-M2.7-APEX-F16-00002-of-00003.gguf",
"MiniMax-M2.7-APEX-F16-00003-of-00003.gguf",
)
Expect(got.fullPrecision).To(Equal(3))
Expect(got.unclassified).To(Equal(0))
})
It("treats bf16 the same as f16, in either case", func() {
got := censusOf("Model-APEX-BF16.gguf", "Model-APEX-bf16-00001-of-00002.gguf", "Model-APEX-f16.gguf")
Expect(got.fullPrecision).To(Equal(3))
Expect(got.unclassified).To(Equal(0))
})
It("still reports a genuinely unknown filename as unclassified", func() {
got := censusOf("Model-APEX-Turbo.gguf")
Expect(got.unclassified).To(Equal(1))
Expect(got.fullPrecision).To(Equal(0))
})
It("separates the two kinds when a repo publishes both", func() {
got := censusOf("Model-APEX-F16.gguf", "Model-APEX-Turbo.gguf")
Expect(got.fullPrecision).To(Equal(1))
Expect(got.unclassified).To(Equal(1))
})
})

143
.github/ci/apexentries/merge.go vendored Normal file
View File

@@ -0,0 +1,143 @@
package main
import (
"fmt"
"os"
"strings"
"gopkg.in/yaml.v3"
)
const (
hfShorthandPrefix = "huggingface://"
hfResolvePrefix = "https://huggingface.co/"
hfResolveInfix = "/resolve/main/"
)
// canonicalURI reduces the two interchangeable spellings of a HuggingFace file
// to one key, so a generated resolve/main URI dedups against the shorthand the
// gallery uses for the majority of its entries.
//
// The repo is exactly the first two path segments; everything after is the file
// path, which may itself contain slashes because sharded quants live in a
// subdirectory. Anything that is not recognisably one of the two forms is
// returned unchanged rather than guessed at, so mirrors and other hosts still
// dedup on their literal string.
func canonicalURI(uri string) string {
switch {
case strings.HasPrefix(uri, hfShorthandPrefix):
rest := strings.TrimPrefix(uri, hfShorthandPrefix)
owner, after, ok := strings.Cut(rest, "/")
if !ok {
return uri
}
name, file, ok := strings.Cut(after, "/")
if !ok || owner == "" || name == "" || file == "" {
return uri
}
return hfShorthandPrefix + owner + "/" + name + "/" + file
case strings.HasPrefix(uri, hfResolvePrefix):
rest := strings.TrimPrefix(uri, hfResolvePrefix)
repo, file, ok := strings.Cut(rest, hfResolveInfix)
if !ok || file == "" {
return uri
}
// A repo is owner/name and nothing more; a longer prefix means this is
// some other huggingface.co URL that must not be rewritten.
owner, name, ok := strings.Cut(repo, "/")
if !ok || owner == "" || name == "" || strings.Contains(name, "/") {
return uri
}
return hfShorthandPrefix + repo + "/" + file
default:
return uri
}
}
// ExistingIndex is the lookup built from the current gallery: entry names, and
// which entry claims each weight URI.
type ExistingIndex struct {
ByName map[string]int
ByURI map[string]string
}
// LoadExisting reads the gallery index for dedup purposes only. It is
// deliberately not used to rewrite the file: the index is 40,000 lines, and a
// YAML round trip would reflow the whole thing into an unreviewable diff.
func LoadExisting(path string) (*ExistingIndex, error) {
raw, err := os.ReadFile(path)
if err != nil {
return nil, err
}
var entries []struct {
Name string `yaml:"name"`
Files []struct {
URI string `yaml:"uri"`
} `yaml:"files"`
}
if err := yaml.Unmarshal(raw, &entries); err != nil {
return nil, fmt.Errorf("parsing %s: %w", path, err)
}
ix := &ExistingIndex{ByName: map[string]int{}, ByURI: map[string]string{}}
for i, e := range entries {
ix.ByName[e.Name] = i
for _, f := range e.Files {
if f.URI != "" {
ix.ByURI[canonicalURI(f.URI)] = e.Name
}
}
}
return ix, nil
}
// Merge splits generated entries into those to add and those already covered.
// reused maps a generated name to the existing entry that stands in for it, so
// a parent can reference what is already there instead of duplicating weights.
// Several APEX repos share one base model, so the same counterpart rungs are
// generated more than once in a batch. The batch has to dedup against itself as
// well as against the gallery, tracked locally because the caller may reuse the
// ExistingIndex it passed in.
func Merge(existing *ExistingIndex, generated []GalleryEntry) (add []GalleryEntry, reused map[string]string) {
reused = map[string]string{}
batchNames := map[string]string{}
batchURIs := map[string]string{}
// Canonicalized into a local copy rather than in place: an ExistingIndex may
// be hand-built or reused by the caller, so Merge must not rewrite it.
existingURIs := make(map[string]string, len(existing.ByURI))
for uri, owner := range existing.ByURI {
existingURIs[canonicalURI(uri)] = owner
}
for _, e := range generated {
// Name is checked before URI: a name collision must block the add
// whatever the weights say, since duplicate names corrupt the index.
if _, clash := existing.ByName[e.Name]; clash {
reused[e.Name] = e.Name
continue
}
if claimant, clash := batchNames[e.Name]; clash {
reused[e.Name] = claimant
continue
}
if len(e.Files) > 0 {
uri := canonicalURI(e.Files[0].URI)
if owner, ok := existingURIs[uri]; ok {
reused[e.Name] = owner
continue
}
if claimant, ok := batchURIs[uri]; ok {
reused[e.Name] = claimant
continue
}
batchURIs[uri] = e.Name
}
batchNames[e.Name] = e.Name
add = append(add, e)
}
return add, reused
}

183
.github/ci/apexentries/merge_test.go vendored Normal file
View File

@@ -0,0 +1,183 @@
package main
import (
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("Merge", func() {
It("drops a generated entry whose weight URI already exists and reports the existing name", func() {
existing := &ExistingIndex{
ByName: map[string]int{"qwen3.6-35b-a3b-apex": 0},
ByURI: map[string]string{
"https://huggingface.co/mudler/X-APEX-GGUF/resolve/main/X-APEX-I-Quality.gguf": "qwen3.6-35b-a3b-apex",
},
}
gen := []GalleryEntry{{
Name: "x-apex-i-quality",
Files: []EntryFile{{URI: "https://huggingface.co/mudler/X-APEX-GGUF/resolve/main/X-APEX-I-Quality.gguf"}},
}}
add, reused := Merge(existing, gen)
Expect(add).To(BeEmpty())
Expect(reused).To(HaveKeyWithValue("x-apex-i-quality", "qwen3.6-35b-a3b-apex"))
})
It("keeps a generated entry whose weights are new", func() {
existing := &ExistingIndex{ByName: map[string]int{}, ByURI: map[string]string{}}
gen := []GalleryEntry{{
Name: "x-apex-i-mini",
Files: []EntryFile{{URI: "https://huggingface.co/mudler/X-APEX-GGUF/resolve/main/X-APEX-I-Mini.gguf"}},
}}
add, reused := Merge(existing, gen)
Expect(add).To(HaveLen(1))
Expect(reused).To(BeEmpty())
})
It("refuses to add an entry whose name collides with an existing one", func() {
existing := &ExistingIndex{
ByName: map[string]int{"x-apex-i-mini": 0},
ByURI: map[string]string{},
}
gen := []GalleryEntry{{
Name: "x-apex-i-mini",
Files: []EntryFile{{URI: "https://huggingface.co/mudler/X-APEX-GGUF/resolve/main/other.gguf"}},
}}
add, reused := Merge(existing, gen)
Expect(add).To(BeEmpty())
Expect(reused).To(HaveKeyWithValue("x-apex-i-mini", "x-apex-i-mini"))
})
// The gallery records most of its URIs in huggingface:// shorthand while
// render.go only ever emits the resolve/main form, so without
// canonicalization the majority of the file is invisible to the dedup.
It("matches a generated https URI against the shorthand form recorded in the gallery", func() {
existing := &ExistingIndex{
ByName: map[string]int{"foo-gguf-q8-0": 0},
ByURI: map[string]string{
"huggingface://unsloth/Foo-GGUF/Foo-Q8_0.gguf": "foo-gguf-q8-0",
},
}
gen := []GalleryEntry{{
Name: "foo-apex-q8-0",
Files: []EntryFile{{URI: "https://huggingface.co/unsloth/Foo-GGUF/resolve/main/Foo-Q8_0.gguf"}},
}}
add, reused := Merge(existing, gen)
Expect(add).To(BeEmpty())
Expect(reused).To(HaveKeyWithValue("foo-apex-q8-0", "foo-gguf-q8-0"))
})
It("matches a generated shorthand URI against the https form recorded in the gallery", func() {
existing := &ExistingIndex{
ByName: map[string]int{"foo-gguf-q8-0": 0},
ByURI: map[string]string{
"https://huggingface.co/unsloth/Foo-GGUF/resolve/main/Foo-Q8_0.gguf": "foo-gguf-q8-0",
},
}
gen := []GalleryEntry{{
Name: "foo-apex-q8-0",
Files: []EntryFile{{URI: "huggingface://unsloth/Foo-GGUF/Foo-Q8_0.gguf"}},
}}
add, reused := Merge(existing, gen)
Expect(add).To(BeEmpty())
Expect(reused).To(HaveKeyWithValue("foo-apex-q8-0", "foo-gguf-q8-0"))
})
// Sharded quants live under a subdirectory, so the file path carries slashes
// of its own and only the first two segments are the repo.
It("matches across both forms when the file path has a subdirectory", func() {
existing := &ExistingIndex{
ByName: map[string]int{"model-ud-q4-k-m": 0},
ByURI: map[string]string{
"huggingface://unsloth/Model-GGUF/UD-Q4_K_M/Model-UD-Q4_K_M-00001-of-00002.gguf": "model-ud-q4-k-m",
},
}
gen := []GalleryEntry{{
Name: "model-apex-ud-q4-k-m",
Files: []EntryFile{{URI: "https://huggingface.co/unsloth/Model-GGUF/resolve/main/UD-Q4_K_M/Model-UD-Q4_K_M-00001-of-00002.gguf"}},
}}
add, reused := Merge(existing, gen)
Expect(add).To(BeEmpty())
Expect(reused).To(HaveKeyWithValue("model-apex-ud-q4-k-m", "model-ud-q4-k-m"))
})
// Several APEX repos share one base model, so the same unsloth rungs are
// generated more than once in a single batch.
It("adds only the first of two generated entries sharing a name", func() {
existing := &ExistingIndex{ByName: map[string]int{}, ByURI: map[string]string{}}
gen := []GalleryEntry{
{
Name: "shared-rung-q8-0",
Files: []EntryFile{{URI: "https://huggingface.co/unsloth/Shared-GGUF/resolve/main/Shared-Q8_0.gguf"}},
},
{
Name: "shared-rung-q8-0",
Files: []EntryFile{{URI: "https://huggingface.co/unsloth/Other-GGUF/resolve/main/Other-Q8_0.gguf"}},
},
}
add, reused := Merge(existing, gen)
Expect(add).To(HaveLen(1))
Expect(add[0].Files[0].URI).To(Equal("https://huggingface.co/unsloth/Shared-GGUF/resolve/main/Shared-Q8_0.gguf"))
Expect(reused).To(HaveKeyWithValue("shared-rung-q8-0", "shared-rung-q8-0"))
})
It("adds only the first of two generated entries sharing a primary URI", func() {
existing := &ExistingIndex{ByName: map[string]int{}, ByURI: map[string]string{}}
gen := []GalleryEntry{
{
Name: "shared-rung-from-apex",
Files: []EntryFile{{URI: "https://huggingface.co/unsloth/Shared-GGUF/resolve/main/Shared-Q8_0.gguf"}},
},
{
Name: "shared-rung-from-apex-mtp",
Files: []EntryFile{{URI: "huggingface://unsloth/Shared-GGUF/Shared-Q8_0.gguf"}},
},
}
add, reused := Merge(existing, gen)
Expect(add).To(HaveLen(1))
Expect(add[0].Name).To(Equal("shared-rung-from-apex"))
Expect(reused).To(HaveKeyWithValue("shared-rung-from-apex-mtp", "shared-rung-from-apex"))
})
// Anything that is not a HuggingFace URI must survive untouched, so an
// unrecognised scheme still dedups against the very same string.
It("leaves a URI in neither recognised form alone and still dedups it exactly", func() {
existing := &ExistingIndex{
ByName: map[string]int{"mirrored-model": 0},
ByURI: map[string]string{
"https://mirror.example.com/weights/Model-Q8_0.gguf": "mirrored-model",
},
}
gen := []GalleryEntry{
{
Name: "mirrored-apex",
Files: []EntryFile{{URI: "https://mirror.example.com/weights/Model-Q8_0.gguf"}},
},
{
Name: "elsewhere-apex",
Files: []EntryFile{{URI: "https://mirror.example.com/weights/Other-Q8_0.gguf"}},
},
}
add, reused := Merge(existing, gen)
Expect(add).To(HaveLen(1))
Expect(add[0].Name).To(Equal("elsewhere-apex"))
Expect(reused).To(HaveKeyWithValue("mirrored-apex", "mirrored-model"))
})
})

175
.github/ci/apexentries/render.go vendored Normal file
View File

@@ -0,0 +1,175 @@
package main
import (
"fmt"
"path"
"strings"
)
// EntryFile is one downloadable file of a gallery entry.
type EntryFile struct {
Filename string `yaml:"filename"`
SHA256 string `yaml:"sha256"`
URI string `yaml:"uri"`
}
// GalleryEntry is the subset of a gallery entry this generator writes.
//
// Named GalleryEntry rather than Entry because the test files dot-import
// Ginkgo, whose table DSL exports an Entry that a package-level Entry would
// collide with. The yaml tags are what the gallery index sees, so the Go
// identifier is free to differ.
type GalleryEntry struct {
Name string `yaml:"name"`
URL string `yaml:"url"`
Description string `yaml:"description,omitempty"`
Tags []string `yaml:"tags,omitempty"`
Overrides map[string]any `yaml:"overrides,omitempty"`
Files []EntryFile `yaml:"files,omitempty"`
Variants []VariantRef `yaml:"variants,omitempty"`
}
// VariantRef mirrors the gallery's variant reference: a name and nothing else.
type VariantRef struct {
Model string `yaml:"model"`
}
// ChildInput is everything needed to render one non-parent entry.
type ChildInput struct {
Name string
Repo string
// DraftRepo is the repo publishing the drafter, when it is not the repo
// publishing the weights. Speculative pairings routinely cross repos, so
// the drafter cannot be assumed to sit next to the weights. Empty means
// same-repo, which is how the *-APEX-MTP-GGUF repos ship.
DraftRepo string
Template string
Weights []GGUFFile
MMProj *GGUFFile
SpecType string
DraftFile *GGUFFile
BaseTags []string
}
// specTuning is the acceptance-window tuning each spec type ships with, copied
// from the hand-written entries that already run these two mechanisms rather
// than invented here. The two differ because the drafters differ: self-drafted
// MTP heads produce a short, high-confidence proposal (15+ hand-written entries
// use 6 with a 0.75 floor), while a separate DFlash drafter is cheap enough to
// run far ahead unconditionally (the five hand-written dflash entries use 15 and
// set no floor).
var specTuning = map[string][]string{
"draft-mtp": {"spec_n_max:6", "spec_p_min:0.75"},
"draft-dflash": {"spec_n_max:15"},
}
func hfURI(repo, file string) string {
return fmt.Sprintf("https://huggingface.co/%s/resolve/main/%s", repo, file)
}
// localPath is where a downloaded file lands.
//
// The hand-written entries namespace by the repo's BARE name
// (llama-cpp/models/<repo>/<file>), which is not unique. LiquidAI/LFM2.5-8B-A1B-GGUF
// and unsloth/LFM2.5-8B-A1B-GGUF share a basename, so both claim
// llama-cpp/models/LFM2.5-8B-A1B-GGUF/, and installing the second after the first
// either overwrites weights whose recorded sha256 belongs to the other file or is
// skipped as already present. Two owners publishing the same model name is the
// normal case for quantizers, not an edge case, so the owner has to be in the path.
//
// The owner becomes its own path segment rather than being folded into the
// directory name: owner/repo is unique on HuggingFace and "/" cannot occur inside
// either half, so this is the only form that is collision-proof by construction.
// It still reads as the hand-written convention with the owner restored, and the
// extra depth is already present in the index for sharded builds.
func localPath(kind, repo, file string) string {
// path.Dir yields "." for a repo named without an owner, which path.Join
// drops, so such a caller keeps the historical two-segment layout.
return path.Join("llama-cpp", kind, path.Dir(repo), path.Base(repo), file)
}
// RenderChild builds one child entry.
//
// The dflash/mtp tag is added if and only if this entry sets a spec_type,
// because variant ranking reads tags and nothing else, and a tag that does not
// match what the entry configures either promotes a build that is no faster or
// hides one that is.
func RenderChild(in ChildInput) GalleryEntry {
e := GalleryEntry{
Name: in.Name,
URL: fmt.Sprintf("github:mudler/LocalAI/gallery/%s@master", in.Template),
Tags: append([]string{}, in.BaseTags...),
Overrides: map[string]any{},
}
// gallery/virtual.yaml carries no backend, so nothing else would name an
// engine for these entries. Matching the hand-written entries on
// known_usecases too: LocalAI would fall back to the backend defaults, but
// generated entries should not read differently from their neighbours.
e.Overrides["backend"] = "llama-cpp"
e.Overrides["known_usecases"] = []string{"chat"}
options := []string{"use_jinja:true"}
for _, w := range in.Weights {
e.Files = append(e.Files, EntryFile{
Filename: localPath("models", in.Repo, w.Name),
SHA256: w.SHA256,
URI: hfURI(in.Repo, w.Name),
})
}
e.Overrides["parameters"] = map[string]any{
"model": localPath("models", in.Repo, in.Weights[0].Name),
}
if in.MMProj != nil {
// An explicit known_usecases SUPPRESSES the backend-default fallback in
// core/gallery/models_types.go, so a multimodal entry left at chat-only
// never matches FilterGalleryModelsByUsecase(FLAG_VISION) or
// FilterGalleryModelsByMultimodal and vanishes from the UI's vision and
// multimodal filters. 19 of the 45 APEX repos ship an mmproj.
e.Overrides["known_usecases"] = []string{"chat", "vision"}
e.Overrides["mmproj"] = localPath("mmproj", in.Repo, in.MMProj.Name)
e.Files = append(e.Files, EntryFile{
Filename: localPath("mmproj", in.Repo, in.MMProj.Name),
SHA256: in.MMProj.SHA256,
URI: hfURI(in.Repo, in.MMProj.Name),
})
}
// A spec type is configured independently of a drafter FILE. Weights that
// carry their own MTP heads need no second download, and requiring one left
// the *-APEX-MTP-GGUF builds shipping the larger heads-bearing weights with
// the heads switched off: a strictly bigger download at the same speed,
// ranked identically to the plain rung at the same tier.
if in.SpecType != "" {
options = append(options, "spec_type:"+in.SpecType)
options = append(options, specTuning[in.SpecType]...)
// The tag is derived from the spec type this entry sets and from nothing
// else. Variant ranking reads tags only, so a tag taken from a repo or
// entry NAME would promote a build that is no faster whenever the name
// and the configuration disagree.
e.Tags = append(e.Tags, strings.TrimPrefix(in.SpecType, "draft-"))
}
if in.SpecType != "" && in.DraftFile != nil {
// Fall back to the weights repo so pairings that publish the drafter
// alongside the weights keep working without restating the repo.
draftRepo := in.DraftRepo
if draftRepo == "" {
draftRepo = in.Repo
}
draftPath := localPath("models", draftRepo, in.DraftFile.Name)
e.Overrides["draft_model"] = draftPath
e.Overrides["flash_attention"] = "on"
e.Files = append(e.Files, EntryFile{
Filename: draftPath,
SHA256: in.DraftFile.SHA256,
URI: hfURI(draftRepo, in.DraftFile.Name),
})
}
e.Overrides["options"] = options
return e
}

249
.github/ci/apexentries/render_test.go vendored Normal file
View File

@@ -0,0 +1,249 @@
package main
import (
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("RenderChild", func() {
It("tags an entry that configures draft-dflash", func() {
e := RenderChild(ChildInput{
Name: "qwen3.5-9b-dflash",
Repo: "mudler/Example-APEX-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "Example-APEX-I-Quality.gguf", SHA256: "a"}},
SpecType: "draft-dflash",
DraftFile: &GGUFFile{Name: "Example-DFlash.Q8_0.gguf", SHA256: "b"},
BaseTags: []string{"llm", "gguf"},
})
Expect(e.Tags).To(ContainElement("dflash"))
Expect(e.Tags).ToNot(ContainElement("mtp"))
Expect(e.Overrides["options"]).To(ContainElement("spec_type:draft-dflash"))
Expect(e.Overrides["draft_model"]).ToNot(BeNil())
})
It("does not tag an MTP-named repo that configures no speculation", func() {
// mudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF ships MTP-bearing weights. Weights
// that carry the heads are not an entry that enables them, and tagging it
// would win the feature axis without being any faster.
e := RenderChild(ChildInput{
Name: "qwen3.6-35b-a3b-apex-mtp-i-quality",
Repo: "mudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "Qwen3.6-35B-A3B-APEX-MTP-I-Quality.gguf", SHA256: "a"}},
BaseTags: []string{"llm", "gguf"},
})
Expect(e.Tags).ToNot(ContainElement("mtp"))
Expect(e.Tags).ToNot(ContainElement("dflash"))
Expect(e.Overrides).ToNot(HaveKey("draft_model"))
})
It("lists every shard of a sharded build and points the model at the first", func() {
e := RenderChild(ChildInput{
Name: "step-3.7-flash-ud-q4-k-m",
Repo: "unsloth/Step-3.7-Flash-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{
{Name: "UD-Q4_K_M/Step-3.7-Flash-UD-Q4_K_M-00001-of-00002.gguf", SHA256: "a"},
{Name: "UD-Q4_K_M/Step-3.7-Flash-UD-Q4_K_M-00002-of-00002.gguf", SHA256: "b"},
},
BaseTags: []string{"llm", "gguf"},
})
Expect(e.Files).To(HaveLen(2))
params, ok := e.Overrides["parameters"].(map[string]any)
Expect(ok).To(BeTrue())
Expect(params["model"]).To(HaveSuffix("00001-of-00002.gguf"))
Expect(e.Files[0].URI).To(Equal(
"https://huggingface.co/unsloth/Step-3.7-Flash-GGUF/resolve/main/UD-Q4_K_M/Step-3.7-Flash-UD-Q4_K_M-00001-of-00002.gguf"))
})
It("wires mmproj when the repo publishes one", func() {
e := RenderChild(ChildInput{
Name: "example-i-mini",
Repo: "mudler/Example-APEX-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "Example-APEX-I-Mini.gguf", SHA256: "a"}},
MMProj: &GGUFFile{Name: "mmproj-F16.gguf", SHA256: "c"},
BaseTags: []string{"llm", "gguf"},
})
Expect(e.Overrides["mmproj"]).ToNot(BeNil())
Expect(e.Files).To(HaveLen(2))
})
It("names the engine and the usecases the hand-written entries name", func() {
// gallery/virtual.yaml supplies no backend, so an entry that omits one
// names no engine at all and cannot load.
e := RenderChild(ChildInput{
Name: "example-i-mini",
Repo: "mudler/Example-APEX-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "Example-APEX-I-Mini.gguf", SHA256: "a"}},
BaseTags: []string{"llm", "gguf"},
})
Expect(e.Overrides["backend"]).To(Equal("llama-cpp"))
Expect(e.Overrides["known_usecases"]).To(ContainElement("chat"))
})
It("draws the drafter from DraftRepo when the pairing spans two repos", func() {
// unsloth/Qwen3-4B-GGUF pairs with a drafter published separately by
// AtomicChat, so a drafter URI built from the weights repo 404s.
e := RenderChild(ChildInput{
Name: "qwen3-4b-dflash",
Repo: "unsloth/Qwen3-4B-GGUF",
DraftRepo: "AtomicChat/Qwen3-4B-DFlash-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "Qwen3-4B-Q4_K_M.gguf", SHA256: "a"}},
SpecType: "draft-dflash",
DraftFile: &GGUFFile{Name: "Qwen3-4B-DFlash.Q8_0.gguf", SHA256: "b"},
BaseTags: []string{"llm", "gguf"},
})
Expect(e.Files[0].URI).To(Equal(
"https://huggingface.co/unsloth/Qwen3-4B-GGUF/resolve/main/Qwen3-4B-Q4_K_M.gguf"))
Expect(e.Files[1].URI).To(Equal(
"https://huggingface.co/AtomicChat/Qwen3-4B-DFlash-GGUF/resolve/main/Qwen3-4B-DFlash.Q8_0.gguf"))
Expect(e.Files[1].Filename).To(Equal(
"llama-cpp/models/AtomicChat/Qwen3-4B-DFlash-GGUF/Qwen3-4B-DFlash.Q8_0.gguf"))
Expect(e.Overrides["draft_model"]).To(Equal(
"llama-cpp/models/AtomicChat/Qwen3-4B-DFlash-GGUF/Qwen3-4B-DFlash.Q8_0.gguf"))
})
It("falls back to the weights repo for the drafter when DraftRepo is empty", func() {
// The *-APEX-MTP-GGUF repos ship the drafter alongside the weights.
e := RenderChild(ChildInput{
Name: "example-apex-dflash",
Repo: "mudler/Example-APEX-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "Example-APEX-I-Quality.gguf", SHA256: "a"}},
SpecType: "draft-dflash",
DraftFile: &GGUFFile{Name: "Example-DFlash.Q8_0.gguf", SHA256: "b"},
BaseTags: []string{"llm", "gguf"},
})
Expect(e.Files[1].URI).To(Equal(
"https://huggingface.co/mudler/Example-APEX-GGUF/resolve/main/Example-DFlash.Q8_0.gguf"))
Expect(e.Files[1].Filename).To(Equal(
"llama-cpp/models/mudler/Example-APEX-GGUF/Example-DFlash.Q8_0.gguf"))
})
})
var _ = Describe("RenderChild known_usecases", func() {
It("declares vision alongside chat when the entry carries an mmproj", func() {
// An explicit known_usecases suppresses the backend-default fallback, so a
// chat-only multimodal entry disappears from the UI's vision filter.
e := RenderChild(ChildInput{
Name: "example-i-quality",
Repo: "mudler/Example-APEX-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "Example-APEX-I-Quality.gguf", SHA256: "a"}},
MMProj: &GGUFFile{Name: "mmproj-F16.gguf", SHA256: "c"},
BaseTags: []string{"llm", "gguf"},
})
Expect(e.Overrides["known_usecases"]).To(ConsistOf("chat", "vision"))
})
It("leaves a text-only entry at chat", func() {
e := RenderChild(ChildInput{
Name: "example-i-quality",
Repo: "mudler/Example-APEX-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "Example-APEX-I-Quality.gguf", SHA256: "a"}},
BaseTags: []string{"llm", "gguf"},
})
Expect(e.Overrides["known_usecases"]).To(ConsistOf("chat"))
})
})
var _ = Describe("localPath", func() {
It("keeps two repos with the same basename but different owners apart", func() {
// LiquidAI and unsloth both publish LFM2.5-8B-A1B-GGUF. A path built from
// the bare repo name gives both the same local file, so installing the
// second overwrites or skips the first and one of them then serves bytes
// that do not match its recorded sha256.
liquid := RenderChild(ChildInput{
Name: "lfm2.5-8b-a1b-i-quality",
Repo: "LiquidAI/LFM2.5-8B-A1B-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "LFM2.5-8B-A1B-Q8_0.gguf", SHA256: "33ab3b8c"}},
BaseTags: []string{"llm", "gguf"},
})
unsloth := RenderChild(ChildInput{
Name: "lfm2.5-8b-a1b-q8-0",
Repo: "unsloth/LFM2.5-8B-A1B-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "LFM2.5-8B-A1B-Q8_0.gguf", SHA256: "ec11666b"}},
BaseTags: []string{"llm", "gguf"},
})
Expect(liquid.Files[0].Filename).ToNot(Equal(unsloth.Files[0].Filename))
Expect(unsloth.Files[0].Filename).To(Equal(
"llama-cpp/models/unsloth/LFM2.5-8B-A1B-GGUF/LFM2.5-8B-A1B-Q8_0.gguf"))
})
It("namespaces the mmproj by owner too", func() {
e := RenderChild(ChildInput{
Name: "example-i-quality",
Repo: "mudler/Example-APEX-GGUF",
Template: "virtual.yaml",
Weights: []GGUFFile{{Name: "Example-APEX-I-Quality.gguf", SHA256: "a"}},
MMProj: &GGUFFile{Name: "mmproj-F16.gguf", SHA256: "c"},
BaseTags: []string{"llm", "gguf"},
})
Expect(e.Overrides["mmproj"]).To(Equal(
"llama-cpp/mmproj/mudler/Example-APEX-GGUF/mmproj-F16.gguf"))
})
})
var _ = Describe("MTP builds", func() {
renderTier := func(repo string) GalleryEntry {
return RenderChild(ChildInput{
Name: "example-i-quality",
Repo: repo,
Template: "virtual.yaml",
SpecType: SpecTypeForRepo(repo),
Weights: []GGUFFile{{Name: "Example-I-Quality.gguf", SHA256: "a"}},
BaseTags: []string{"llm", "gguf"},
})
}
It("turns MTP on for a build off an APEX-MTP repo", func() {
// These weights retain the model's own MTP heads, so shipping them with
// speculation off is a strictly larger download at the same speed,
// ranked identically to the plain rung at the same tier.
e := renderTier("mudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF")
Expect(e.Overrides["options"]).To(ContainElements(
"spec_type:draft-mtp", "spec_n_max:6", "spec_p_min:0.75"))
Expect(e.Tags).To(ContainElement("mtp"))
})
It("needs no drafter file, because the heads travel with the weights", func() {
e := renderTier("mudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF")
Expect(e.Overrides).ToNot(HaveKey("draft_model"))
Expect(e.Files).To(HaveLen(1))
})
It("leaves a build off a plain APEX repo alone", func() {
e := renderTier("mudler/Qwen3.6-35B-A3B-APEX-GGUF")
Expect(e.Tags).ToNot(ContainElement("mtp"))
Expect(e.Overrides["options"]).To(ConsistOf("use_jinja:true"))
})
It("leaves an unsloth counterpart rung alone", func() {
// The counterpart quantizes the plain weights; nothing there carries heads.
e := renderTier("unsloth/Qwen3.6-35B-A3B-GGUF")
Expect(e.Tags).ToNot(ContainElement("mtp"))
Expect(e.Overrides["options"]).To(ConsistOf("use_jinja:true"))
})
})

71
.github/ci/apexentries/unsloth.go vendored Normal file
View File

@@ -0,0 +1,71 @@
package main
import (
"regexp"
"sort"
"strings"
)
// WantedQuants is the fixed unsloth subset this generator emits. It is a
// deliberate subset: unsloth publishes north of 20 quants per repo, and the
// selector needs useful fitness points rather than every rung.
var WantedQuants = []string{"UD-Q4_K_M", "UD-Q5_K_M", "UD-Q6_K", "Q8_0"}
var shardRE = regexp.MustCompile(`-(\d{5})-of-(\d{5})\.gguf$`)
// QuantBuild is one unsloth quantization, which may be a single file or an
// ordered set of shards.
type QuantBuild struct {
Quant string
Files []GGUFFile
Sharded bool
}
// CounterpartCandidates returns the unsloth repo base names worth probing, most
// likely first. Both derivations are needed: the repo name finds
// unsloth/gemma-4-26B-A4B-it-GGUF, while the file stem is what matches for
// repos whose stem is the canonical model name.
func CounterpartCandidates(repoName, fileStem string) []string {
clean := func(s string) string {
s = strings.TrimSuffix(s, "-GGUF")
s = regexp.MustCompile(`-(MTP|TQ)$`).ReplaceAllString(s, "")
s = strings.TrimSuffix(s, "-APEX")
return regexp.MustCompile(`-(MTP|TQ)$`).ReplaceAllString(s, "")
}
out := []string{clean(repoName)}
if stem := clean(fileStem); stem != out[0] {
out = append(out, stem)
}
return out
}
// DiscoverUnslothQuants returns the wanted quants a repo publishes, handling
// both the flat single-file layout and the sharded layout where a quant lives
// in its own subdirectory.
func DiscoverUnslothQuants(files []GGUFFile) []QuantBuild {
var out []QuantBuild
for _, q := range WantedQuants {
var flat []GGUFFile
var shards []GGUFFile
for _, f := range files {
switch {
case !strings.Contains(f.Name, "/") && strings.HasSuffix(f.Name, "-"+q+".gguf"):
flat = append(flat, f)
case strings.HasPrefix(f.Name, q+"/") && shardRE.MatchString(f.Name):
shards = append(shards, f)
}
}
switch {
case len(flat) > 0:
out = append(out, QuantBuild{Quant: q, Files: flat})
case len(shards) > 0:
sort.Slice(shards, func(i, j int) bool { return shards[i].Name < shards[j].Name })
out = append(out, QuantBuild{Quant: q, Files: shards, Sharded: true})
}
}
return out
}

75
.github/ci/apexentries/unsloth_test.go vendored Normal file
View File

@@ -0,0 +1,75 @@
package main
import (
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("CounterpartCandidates", func() {
It("offers both the repo-derived and stem-derived names", func() {
// mudler/gemma-4-26B-A4B-it-APEX-GGUF ships gemma-4-26B-A4B-APEX-*.gguf,
// and only the repo-derived name finds unsloth/gemma-4-26B-A4B-it-GGUF.
got := CounterpartCandidates("gemma-4-26B-A4B-it-APEX-GGUF", "gemma-4-26B-A4B-APEX")
Expect(got).To(Equal([]string{"gemma-4-26B-A4B-it", "gemma-4-26B-A4B"}))
})
It("strips the MTP marker", func() {
got := CounterpartCandidates("Qwopus3.6-35B-A3B-v1-APEX-MTP-GGUF", "Qwopus3.6-35B-A3B-v1-APEX-MTP")
Expect(got[0]).To(Equal("Qwopus3.6-35B-A3B-v1"))
})
It("strips the TQ marker", func() {
// This is the branch that folds mudler/Qwen3.5-35B-A3B-APEX-TQ-GGUF into
// the qwen3.5-35b-a3b hub. Without it the probe is
// unsloth/Qwen3.5-35B-A3B-TQ-GGUF, which does not exist, so the family
// silently loses every unsloth rung.
got := CounterpartCandidates("Qwen3.5-35B-A3B-APEX-TQ-GGUF", "Qwen3.5-35B-A3B-APEX-TQ")
Expect(got).To(Equal([]string{"Qwen3.5-35B-A3B"}))
})
It("does not repeat a candidate when both derivations agree", func() {
got := CounterpartCandidates("Qwen3.6-35B-A3B-APEX-GGUF", "Qwen3.6-35B-A3B-APEX")
Expect(got).To(Equal([]string{"Qwen3.6-35B-A3B"}))
})
})
var _ = Describe("DiscoverUnslothQuants", func() {
It("finds flat single-file quants", func() {
files := []GGUFFile{
{Name: "Qwen3.6-35B-A3B-UD-Q4_K_M.gguf", SHA256: "a"},
{Name: "Qwen3.6-35B-A3B-UD-IQ1_M.gguf", SHA256: "b"},
}
got := DiscoverUnslothQuants(files)
Expect(got).To(HaveLen(1))
Expect(got[0].Quant).To(Equal("UD-Q4_K_M"))
Expect(got[0].Sharded).To(BeFalse())
Expect(got[0].Files).To(HaveLen(1))
})
It("collects a sharded quant from its subdirectory in shard order", func() {
files := []GGUFFile{
{Name: "UD-Q4_K_M/Step-3.7-Flash-UD-Q4_K_M-00002-of-00002.gguf", SHA256: "b"},
{Name: "UD-Q4_K_M/Step-3.7-Flash-UD-Q4_K_M-00001-of-00002.gguf", SHA256: "a"},
}
got := DiscoverUnslothQuants(files)
Expect(got).To(HaveLen(1))
Expect(got[0].Quant).To(Equal("UD-Q4_K_M"))
Expect(got[0].Sharded).To(BeTrue())
Expect(got[0].Files).To(HaveLen(2))
Expect(got[0].Files[0].Name).To(HaveSuffix("00001-of-00002.gguf"))
})
It("ignores quants outside the wanted subset", func() {
files := []GGUFFile{{Name: "Model-UD-IQ2_XXS.gguf", SHA256: "a"}}
Expect(DiscoverUnslothQuants(files)).To(BeEmpty())
})
})

312
.github/ci/apexentries/verify.go vendored Normal file
View File

@@ -0,0 +1,312 @@
package main
import (
"fmt"
"os"
"strings"
"gopkg.in/yaml.v3"
)
type verifyEntry struct {
Name string `yaml:"name"`
Tags []string `yaml:"tags"`
Variants []VariantRef `yaml:"variants"`
Overrides struct {
// Backend scopes the checks that only hold for one engine. An entry that
// declares none takes its configuration from the referenced url: template,
// which this verifier never reads, so it cannot be judged either way.
Backend string `yaml:"backend"`
Options []string `yaml:"options"`
// MMProj and DraftModel name the files that are not weights. They are
// the only signal for it: a drafter lands in the same models/ prefix as
// the weights, so the path alone cannot tell them apart.
MMProj string `yaml:"mmproj"`
DraftModel string `yaml:"draft_model"`
} `yaml:"overrides"`
Files []struct {
Filename string `yaml:"filename"`
SHA256 string `yaml:"sha256"`
URI string `yaml:"uri"`
} `yaml:"files"`
}
// Verify checks the invariants the variants schema and the tagging rule
// require. It returns every problem rather than the first, so one run tells the
// author everything that needs fixing.
func Verify(path string) []string {
raw, err := os.ReadFile(path)
if err != nil {
return []string{fmt.Sprintf("reading %s: %v", path, err)}
}
var entries []verifyEntry
if err := yaml.Unmarshal(raw, &entries); err != nil {
return []string{fmt.Sprintf("parsing %s: %v", path, err)}
}
var problems []string
byName := map[string]verifyEntry{}
for _, e := range entries {
if _, seen := byName[e.Name]; seen {
problems = append(problems, fmt.Sprintf("duplicate entry name: %s", e.Name))
continue
}
byName[e.Name] = e
}
for _, e := range entries {
for _, v := range e.Variants {
target, ok := byName[v.Model]
if !ok {
problems = append(problems, fmt.Sprintf("%s: variant %q does not exist", e.Name, v.Model))
continue
}
if len(target.Variants) > 0 {
problems = append(problems, fmt.Sprintf("%s: variant %q declares variants of its own", e.Name, v.Model))
}
}
for _, f := range e.Files {
if requiresSHA256(f.Filename) && f.SHA256 == "" {
problems = append(problems, fmt.Sprintf("%s: file %s has no sha256", e.Name, f.Filename))
}
}
problems = append(problems, checkWeightCount(e)...)
problems = append(problems, checkFeatureTag(e, "dflash")...)
problems = append(problems, checkFeatureTag(e, "mtp")...)
}
problems = append(problems, checkPathCollisions(entries)...)
return problems
}
// checkPathCollisions catches two different upstream files claiming one local
// path. The install layer keys on the local filename, so whichever entry is
// installed second either overwrites weights the first entry recorded a
// different sha256 for or is skipped as already present. Either way some entry
// afterwards serves bytes that do not match its own checksum, and nothing at
// install time says so.
//
// This is an index-wide invariant rather than a per-entry one: neither entry is
// wrong on its own and the collision exists only in their pairing. The usual
// source is a path scheme built from the repo's BARE name, because two owners
// publishing the same model name is routine for quantizers.
//
// Sharing a path is fine when the uri is the same, which is how several entries
// legitimately reuse one projector. Files with no uri are skipped: there is
// nothing to compare.
func checkPathCollisions(entries []verifyEntry) []string {
type source struct{ uri, entry string }
first := map[string]source{}
reported := map[string]bool{}
var problems []string
for _, e := range entries {
for _, f := range e.Files {
if f.Filename == "" || f.URI == "" {
continue
}
prev, seen := first[f.Filename]
if !seen {
first[f.Filename] = source{uri: f.URI, entry: e.Name}
continue
}
if prev.uri == f.URI || reported[f.Filename] {
continue
}
// Reported once per path however many entries pile onto it, so one
// heavily reused filename cannot bury the rest of the report.
reported[f.Filename] = true
problems = append(problems, fmt.Sprintf(
"local path %s is claimed by two different uris: %s (%s) and %s (%s)",
f.Filename, prev.uri, prev.entry, f.URI, e.Name))
}
}
return problems
}
// auxiliaryExtensions are the metadata formats an entry ships beside its
// weights, where an unverified download is a nuisance rather than a hole.
//
// The exclusion is stated as a list of metadata formats on purpose. Requiring
// the checksum only on a blessed list of weight formats would silently exempt
// every format nobody has shipped yet, and it already exempted safetensors
// weights, which are downloaded and loaded exactly like GGUF ones.
var auxiliaryExtensions = []string{".json", ".txt", ".md"}
// requiresSHA256 reports whether an unverified download of this file would be
// a supply-chain hole rather than a cosmetic gap.
func requiresSHA256(filename string) bool {
for _, ext := range auxiliaryExtensions {
if strings.HasSuffix(filename, ext) {
return false
}
}
return true
}
// checkWeightCount catches an entry carrying two whole models. The flat-match
// branch in DiscoverUnslothQuants appends every match, so a quant label that is
// a suffix of another one (Q8_0 of UD-Q8_0) collects both files into one build
// while the rendered model: points at only the first. The result downloads
// twice the bytes and serves whichever file sorted first, silently.
//
// Shards are exempt because a sharded build is legitimately many files.
//
// The collision is a property of llama-cpp quant discovery, so the check is
// scoped to that backend. Multi-component TTS, ASR and diffusion engines ship an
// encoder, a decoder and a vocoder as one model, and there the second GGUF is
// the design rather than a bug.
func checkWeightCount(e verifyEntry) []string {
if e.Overrides.Backend != "llama-cpp" {
return nil
}
var weights []string
for _, f := range e.Files {
switch {
case !strings.HasSuffix(f.Filename, ".gguf"):
case shardRE.MatchString(f.Filename):
case f.Filename == e.Overrides.MMProj:
case f.Filename == e.Overrides.DraftModel:
default:
weights = append(weights, f.Filename)
}
}
if len(weights) > 1 {
return []string{fmt.Sprintf("%s: more than one weight file: %s", e.Name, strings.Join(weights, ", "))}
}
return nil
}
// checkFeatureTag enforces the rule in both directions. A tag without the
// configuration promotes a build that is no faster; configuration without the
// tag leaves a genuinely faster build ranked as plain.
//
// It only speaks about backends whose declaration it can actually read, because
// a rule applied where the evidence is invisible reports noise rather than bugs.
func checkFeatureTag(e verifyEntry, feature string) []string {
decl, configured, judgeable := featureDeclaration(e, feature)
if !judgeable {
return nil
}
tagged := false
for _, t := range e.Tags {
if t == feature {
tagged = true
break
}
}
switch {
case tagged && !configured:
return []string{fmt.Sprintf("%s: tagged %s but sets no %s", e.Name, feature, decl)}
case configured && !tagged:
return []string{fmt.Sprintf("%s: sets %s but is not tagged %s", e.Name, decl, feature)}
}
return nil
}
// featureDeclaration implements the per-backend table in
// .agents/adding-gallery-models.md. It returns the declaration the backend uses
// to configure the feature, whether the entry carries it, and whether this
// verifier is in a position to answer at all.
func featureDeclaration(e verifyEntry, feature string) (decl string, configured, judgeable bool) {
switch e.Overrides.Backend {
case "llama-cpp":
decl = "spec_type:draft-" + feature
for _, o := range e.Overrides.Options {
if strings.TrimSpace(o) == decl {
return decl, true, true
}
}
return decl, false, true
case "ds4":
// ds4 carries the MTP heads in the weights and turns them on with
// mtp_path / mtp_draft. It has no dflash counterpart, so dflash is not a
// question that can be asked of a ds4 entry.
if feature != "mtp" {
return "", false, false
}
decl = "mtp_path:"
for _, o := range e.Overrides.Options {
o = strings.TrimSpace(o)
if strings.HasPrefix(o, "mtp_path:") || strings.HasPrefix(o, "mtp_draft:") {
return decl, true, true
}
}
return decl, false, true
default:
// sglang configures the feature with speculative_algorithm: in the
// referenced gallery/*.yaml, and an entry that declares no backend takes
// its whole configuration from its url: template. Verify reads one index
// file and follows neither, so it must not judge these in either
// direction.
return "", false, false
}
}
// UnaccountedQuants reports a wanted quant the repo demonstrably publishes but
// that discovery produced no build for. The layout that triggers it today is
// root-level shards, which match neither branch of DiscoverUnslothQuants; no
// counterpart ships that way yet, but a batch generator must not drop a build
// with nothing said about it.
func UnaccountedQuants(files []GGUFFile, builds []QuantBuild) []string {
built := map[string]bool{}
for _, b := range builds {
built[b.Quant] = true
}
var problems []string
for _, q := range WantedQuants {
if built[q] {
continue
}
for _, f := range files {
if filePublishesQuant(f.Name, q) {
problems = append(problems, fmt.Sprintf("quant %s is published upstream (%s) but produced no build", q, f.Name))
break
}
}
}
return problems
}
// filePublishesQuant reports whether an upstream file is a publication of
// quant q. It anchors on the quant label the way DiscoverUnslothQuants does,
// as the trailing token of the base name or as the sharding subdirectory, so
// the diagnostic and the discovery it audits cannot disagree about what a file
// is.
//
// An unanchored match would reproduce the very collision this diagnostic warns
// about: Q8_0 is a substring of UD-Q8_0, so a repo publishing only UD-Q8_0
// would be reported as publishing an unbuilt Q8_0, which it does not, and
// UD-Q8_0 is not a wanted quant at all.
func filePublishesQuant(name, q string) bool {
if strings.HasPrefix(name, q+"/") {
return true
}
base := name[strings.LastIndex(name, "/")+1:]
// Shard numbering sits between the quant label and the extension, so it has
// to come off before the label can be read as the trailing token. Root-level
// shards are the layout that matches neither branch of
// DiscoverUnslothQuants, and so the layout this diagnostic mainly catches.
base = shardRE.ReplaceAllString(base, ".gguf")
if !strings.HasSuffix(base, "-"+q+".gguf") {
return false
}
// UD- is unsloth's dynamic-quant modifier, and UD-<q> is a distinct quant
// label rather than a publication of <q>.
return !strings.HasSuffix(base, "-UD-"+q+".gguf")
}

480
.github/ci/apexentries/verify_test.go vendored Normal file
View File

@@ -0,0 +1,480 @@
package main
import (
"os"
"path/filepath"
"strings"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("Verify", func() {
write := func(body string) string {
dir := GinkgoT().TempDir()
p := filepath.Join(dir, "index.yaml")
Expect(os.WriteFile(p, []byte(body), 0o600)).To(Succeed())
return p
}
It("passes a sound index", func() {
Expect(Verify(write(`
- name: parent
variants:
- model: child
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
- name: child
files:
- filename: b.gguf
sha256: bb
uri: https://example.com/b.gguf
`))).To(BeEmpty())
})
It("reports a variant pointing at a missing entry", func() {
Expect(Verify(write(`
- name: parent
variants:
- model: ghost
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
`))).To(ContainElement(ContainSubstring("ghost")))
})
It("reports a variant that itself declares variants", func() {
Expect(Verify(write(`
- name: parent
variants:
- model: child
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
- name: child
variants:
- model: grandchild
files:
- filename: b.gguf
sha256: bb
uri: https://example.com/b.gguf
- name: grandchild
files:
- filename: c.gguf
sha256: cc
uri: https://example.com/c.gguf
`))).To(ContainElement(ContainSubstring("declares variants of its own")))
})
It("reports duplicate entry names", func() {
Expect(Verify(write(`
- name: dup
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
- name: dup
files:
- filename: b.gguf
sha256: bb
uri: https://example.com/b.gguf
`))).To(ContainElement(ContainSubstring("duplicate entry name")))
})
It("reports a file with no sha256", func() {
Expect(Verify(write(`
- name: one
files:
- filename: a.gguf
uri: https://example.com/a.gguf
`))).To(ContainElement(ContainSubstring("no sha256")))
})
It("reports an entry tagged dflash without a matching spec_type", func() {
Expect(Verify(write(`
- name: liar
tags:
- dflash
overrides:
backend: llama-cpp
options:
- use_jinja:true
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
`))).To(ContainElement(ContainSubstring("tagged dflash")))
})
It("reports an entry configuring spec_type without the tag", func() {
Expect(Verify(write(`
- name: shy
overrides:
backend: llama-cpp
options:
- spec_type:draft-mtp
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
`))).To(ContainElement(ContainSubstring("not tagged mtp")))
})
// ds4 carries the MTP heads in the weights and names them with mtp_path, so
// the rule holds there in a different vocabulary rather than not at all.
It("reports a ds4 entry configuring mtp_path without the tag", func() {
Expect(Verify(write(`
- name: ds4-shy
overrides:
backend: ds4
options:
- mtp_path:model-mtp.gguf
- mtp_draft:2
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
`))).To(ContainElement(ContainSubstring("not tagged mtp")))
})
It("reports a ds4 entry tagged mtp that configures no mtp_path", func() {
Expect(Verify(write(`
- name: ds4-liar
tags:
- mtp
overrides:
backend: ds4
options:
- context_size:4096
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
`))).To(ContainElement(ContainSubstring("tagged mtp")))
})
It("accepts a ds4 entry that both configures mtp_path and carries the tag", func() {
Expect(Verify(write(`
- name: ds4-honest
tags:
- mtp
overrides:
backend: ds4
options:
- mtp_path:model-mtp.gguf
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
`))).To(BeEmpty())
})
// sglang declares speculative_algorithm in the referenced gallery/*.yaml,
// which Verify never reads, so it may not judge such an entry either way.
It("says nothing about an sglang entry tagged mtp", func() {
Expect(Verify(write(`
- name: sglang-mtp
tags:
- mtp
overrides:
backend: sglang
files: []
`))).To(BeEmpty())
})
It("says nothing about the tag on an entry with no declared backend", func() {
Expect(Verify(write(`
- name: templated
tags:
- mtp
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
`))).To(BeEmpty())
})
// The flat-match branch in unsloth.go appends every match, so a repo
// publishing both a plain and a UD Q8_0 renders one entry holding two full
// models while model: points at only the first.
It("reports an entry holding more than one non-shard weight file", func() {
Expect(Verify(write(`
- name: greedy
overrides:
backend: llama-cpp
options:
- use_jinja:true
parameters:
model: llama-cpp/models/repo/Model-Q8_0.gguf
files:
- filename: llama-cpp/models/repo/Model-Q8_0.gguf
sha256: aa
uri: https://example.com/a.gguf
- filename: llama-cpp/models/repo/Model-UD-Q8_0.gguf
sha256: bb
uri: https://example.com/b.gguf
`))).To(ContainElement(ContainSubstring("more than one weight file")))
})
It("accepts many shards alongside an mmproj and a drafter", func() {
Expect(Verify(write(`
- name: sharded
tags:
- mtp
overrides:
backend: llama-cpp
options:
- spec_type:draft-mtp
mmproj: llama-cpp/mmproj/repo/mm.gguf
draft_model: llama-cpp/models/repo/Model-draft.gguf
files:
- filename: llama-cpp/models/repo/Model-00001-of-00002.gguf
sha256: aa
uri: https://example.com/a.gguf
- filename: llama-cpp/models/repo/Model-00002-of-00002.gguf
sha256: bb
uri: https://example.com/b.gguf
- filename: llama-cpp/mmproj/repo/mm.gguf
sha256: cc
uri: https://example.com/c.gguf
- filename: llama-cpp/models/repo/Model-draft.gguf
sha256: dd
uri: https://example.com/d.gguf
`))).To(BeEmpty())
})
// Multi-component TTS and ASR engines legitimately ship an encoder, a
// tokenizer, a vocoder and so on as one model, so the collision the weight
// count catches does not exist for them.
It("accepts a multi-component non-llama-cpp entry declaring five weights", func() {
Expect(Verify(write(`
- name: multi
overrides:
backend: qwen3-tts-cpp
files:
- filename: talker.gguf
sha256: aa
uri: https://example.com/a.gguf
- filename: tokenizer.gguf
sha256: bb
uri: https://example.com/b.gguf
- filename: vocoder.gguf
sha256: cc
uri: https://example.com/c.gguf
- filename: encoder.gguf
sha256: dd
uri: https://example.com/d.gguf
- filename: vae.gguf
sha256: ee
uri: https://example.com/e.gguf
`))).To(BeEmpty())
})
It("says nothing about the weight count of an entry with no declared backend", func() {
Expect(Verify(write(`
- name: templated-weights
files:
- filename: model-Q4_K_M.gguf
sha256: aa
uri: https://example.com/a.gguf
- filename: model-mmproj-f16.gguf
sha256: bb
uri: https://example.com/b.gguf
`))).To(BeEmpty())
})
It("says nothing about an auxiliary metadata file carrying no sha256", func() {
Expect(Verify(write(`
- name: aux
files:
- filename: a.gguf
sha256: aa
uri: https://example.com/a.gguf
- filename: params.json
sha256: ""
uri: https://example.com/params.json
`))).To(BeEmpty())
})
// safetensors weights are downloaded and loaded exactly like GGUF weights,
// so an unverified one is the same supply-chain hole.
It("reports a safetensors weight carrying no sha256", func() {
Expect(Verify(write(`
- name: vae
files:
- filename: wan_2.1_vae.safetensors
sha256: ""
uri: https://example.com/vae.safetensors
`))).To(ContainElement(ContainSubstring("no sha256")))
})
It("says nothing about a txt or md file carrying no sha256", func() {
Expect(Verify(write(`
- name: docs
files:
- filename: notes.txt
sha256: ""
uri: https://example.com/notes.txt
- filename: README.md
sha256: ""
uri: https://example.com/README.md
`))).To(BeEmpty())
})
})
var _ = Describe("UnaccountedQuants", func() {
// A quant published only as root-level shards matches neither branch in
// DiscoverUnslothQuants, so without this diagnostic the build would vanish
// from a batch run with nothing said about it.
It("reports a wanted quant upstream publishes but discovery dropped", func() {
files := []GGUFFile{
{Name: "Model-UD-Q4_K_M-00001-of-00003.gguf", SHA256: "aa"},
{Name: "Model-UD-Q4_K_M-00002-of-00003.gguf", SHA256: "bb"},
{Name: "Model-UD-Q4_K_M-00003-of-00003.gguf", SHA256: "cc"},
}
Expect(UnaccountedQuants(files, DiscoverUnslothQuants(files))).
To(ContainElement(ContainSubstring("UD-Q4_K_M")))
})
It("says nothing when every published wanted quant produced a build", func() {
files := []GGUFFile{
{Name: "Model-UD-Q4_K_M.gguf", SHA256: "aa"},
{Name: "UD-Q6_K/Model-UD-Q6_K-00001-of-00002.gguf", SHA256: "bb"},
{Name: "UD-Q6_K/Model-UD-Q6_K-00002-of-00002.gguf", SHA256: "cc"},
}
Expect(UnaccountedQuants(files, DiscoverUnslothQuants(files))).To(BeEmpty())
})
It("says nothing about a wanted quant the repo does not publish at all", func() {
files := []GGUFFile{{Name: "Model-UD-Q4_K_M.gguf", SHA256: "aa"}}
Expect(UnaccountedQuants(files, DiscoverUnslothQuants(files))).To(BeEmpty())
})
// UD-Q8_0 is its own quant label and is not a wanted one. Reading it as a
// publication of Q8_0 is the substring collision this diagnostic exists to
// warn about, and subdirectory-sharded UD quants are the normal unsloth
// layout for large repos, so the false positive would fire on every batch.
It("does not read a subdirectory-sharded UD-Q8_0 as a published Q8_0", func() {
files := []GGUFFile{
{Name: "UD-Q8_0/Model-UD-Q8_0-00001-of-00002.gguf", SHA256: "aa"},
{Name: "UD-Q8_0/Model-UD-Q8_0-00002-of-00002.gguf", SHA256: "bb"},
}
Expect(UnaccountedQuants(files, DiscoverUnslothQuants(files))).To(BeEmpty())
})
// A quant in its own subdirectory but not shard-numbered matches neither
// branch of DiscoverUnslothQuants, so it is genuinely published and
// genuinely undiscovered.
It("reports a wanted quant published in its own subdirectory without shard numbering", func() {
files := []GGUFFile{{Name: "Q8_0/Model-Q8_0.gguf", SHA256: "aa"}}
Expect(UnaccountedQuants(files, DiscoverUnslothQuants(files))).
To(ContainElement(ContainSubstring("quant Q8_0 is published upstream")))
})
// builds is empty on purpose: it isolates the file-to-quant match from
// whatever DiscoverUnslothQuants would have made of the same file.
It("matches the flat single-file layout", func() {
files := []GGUFFile{{Name: "Model-Q8_0.gguf", SHA256: "aa"}}
Expect(UnaccountedQuants(files, nil)).
To(ConsistOf(ContainSubstring("quant Q8_0 is published upstream")))
})
})
var _ = Describe("Verify local path collisions", func() {
write := func(body string) string {
dir := GinkgoT().TempDir()
p := filepath.Join(dir, "index.yaml")
Expect(os.WriteFile(p, []byte(body), 0o600)).To(Succeed())
return p
}
It("reports one local path claimed by two different uris", func() {
// The shape that shipped: LiquidAI and unsloth both publish
// LFM2.5-8B-A1B-GGUF, so a path built from the bare repo name gives both
// entries the same local file under two different checksums.
Expect(Verify(write(`
- name: lfm2.5-8b-a1b
files:
- filename: llama-cpp/models/LFM2.5-8B-A1B-GGUF/LFM2.5-8B-A1B-Q8_0.gguf
sha256: 33ab3b8c
uri: https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-GGUF/resolve/main/LFM2.5-8B-A1B-Q8_0.gguf
- name: lfm2.5-8b-a1b-q8-0
files:
- filename: llama-cpp/models/LFM2.5-8B-A1B-GGUF/LFM2.5-8B-A1B-Q8_0.gguf
sha256: ec11666b
uri: https://huggingface.co/unsloth/LFM2.5-8B-A1B-GGUF/resolve/main/LFM2.5-8B-A1B-Q8_0.gguf
`))).To(ContainElement(SatisfyAll(
ContainSubstring("claimed by two different uris"),
ContainSubstring("lfm2.5-8b-a1b-q8-0"),
)))
})
It("accepts two entries reusing one file from the same uri", func() {
// Sibling builds of one repo legitimately share a projector.
Expect(Verify(write(`
- name: a
files:
- filename: llama-cpp/mmproj/mudler/Example-GGUF/mmproj-F16.gguf
sha256: cc
uri: https://huggingface.co/mudler/Example-GGUF/resolve/main/mmproj-F16.gguf
- name: b
files:
- filename: llama-cpp/mmproj/mudler/Example-GGUF/mmproj-F16.gguf
sha256: cc
uri: https://huggingface.co/mudler/Example-GGUF/resolve/main/mmproj-F16.gguf
`))).To(BeEmpty())
})
It("reports a collision once however many entries pile onto the path", func() {
problems := Verify(write(`
- name: a
files:
- filename: shared.gguf
sha256: aa
uri: https://example.com/a.gguf
- name: b
files:
- filename: shared.gguf
sha256: bb
uri: https://example.com/b.gguf
- name: c
files:
- filename: shared.gguf
sha256: cc
uri: https://example.com/c.gguf
`))
var collisions int
for _, p := range problems {
if strings.Contains(p, "claimed by two different uris") {
collisions++
}
}
Expect(collisions).To(Equal(1))
})
It("says nothing about files that carry no uri", func() {
// A hand-written entry may record only a checksum. There is no upstream
// to compare, so the check cannot conclude anything either way.
Expect(Verify(write(`
- name: a
files:
- filename: shared.gguf
sha256: aa
- name: b
files:
- filename: shared.gguf
sha256: bb
`))).To(BeEmpty())
})
})

152
.github/ci/galleryedit/edit.go vendored Normal file
View File

@@ -0,0 +1,152 @@
// Package galleryedit splices variant references into the LocalAI gallery index
// as TEXT.
//
// Re-serialising the index through a YAML marshaller would reflow 40,000 lines,
// drop the anchors and merge keys the gallery relies on, and produce a diff no
// reviewer could read, which makes a pull request worthless even when the
// content inside it is right. Every generator that adds variants to an entry the
// gallery already ships therefore edits lines, and they share this package so
// that two of them cannot drift apart on where a variants block belongs.
package galleryedit
import (
"fmt"
"regexp"
"sort"
"strings"
)
var (
entryStart = regexp.MustCompile(`^-(?: |$)`)
inlineName = regexp.MustCompile(`^- (?:&\S+ )?name:`)
keyName = regexp.MustCompile(`^ name:`)
keyVariants = regexp.MustCompile(`^ variants:\s*(.*)$`)
variantItem = regexp.MustCompile(`^ - `)
unsafeInName = regexp.MustCompile(`[:#{}\[\],&*?|>'"%@` + "`" + `]|^\s|\s$`)
)
// Entry is the positional view of one gallery entry: what it is called and
// which lines it occupies. Nothing about what the entry MEANS belongs here, so
// each caller keeps its own semantic decode and only hands over the coordinates.
type Entry struct {
Name string
// StartLine and EndLine bound the entry, zero based and half open.
StartLine int
EndLine int
}
// Insert is one entry's pending variants addition. The caller owns the contents
// of Variants: this package neither orders nor deduplicates them, because the
// right order and the right dedup rule differ between generators.
type Insert struct {
Entry Entry
Variants []string
}
// Scan splits index text into lines and reports the line each top level list
// item begins on.
func Scan(text string) (lines []string, starts []int) {
lines = strings.Split(text, "\n")
for i, line := range lines {
if entryStart.MatchString(line) {
starts = append(starts, i)
}
}
return lines, starts
}
// Apply splices every insert into the index lines and returns the new text.
func Apply(lines []string, inserts []Insert) ([]string, error) {
type edit struct {
at int
remove int
insert []string
}
var edits []edit
for _, in := range inserts {
if len(in.Variants) == 0 {
continue
}
items := make([]string, 0, len(in.Variants))
for _, v := range in.Variants {
items = append(items, " - model: "+QuoteName(v))
}
at, remove, err := insertionPoint(lines, in.Entry)
if err != nil {
return nil, err
}
block := items
if remove > 0 || !hasVariantsKey(lines, in.Entry) {
block = append([]string{" variants:"}, items...)
}
edits = append(edits, edit{at: at, remove: remove, insert: block})
}
// Applying from the bottom up keeps every line number computed against the
// original text valid while earlier edits are still pending.
sort.Slice(edits, func(i, j int) bool { return edits[i].at > edits[j].at })
out := append([]string(nil), lines...)
for _, e := range edits {
tail := append([]string(nil), out[e.at+e.remove:]...)
out = append(out[:e.at], append(append([]string(nil), e.insert...), tail...)...)
}
return out, nil
}
func hasVariantsKey(lines []string, e Entry) bool {
for i := e.StartLine; i < e.EndLine; i++ {
if keyVariants.MatchString(lines[i]) {
return true
}
}
return false
}
// insertionPoint reports where new variant items belong, and how many existing
// lines the insertion replaces.
//
// An entry with no variants key gets one right after its name, which is where
// the hand-written families put it. An entry with an empty "variants: []" has
// that line replaced by a block. An entry with a block gets its items appended.
func insertionPoint(lines []string, e Entry) (at int, remove int, err error) {
for i := e.StartLine; i < e.EndLine; i++ {
m := keyVariants.FindStringSubmatch(lines[i])
if m == nil {
continue
}
if strings.TrimSpace(m[1]) == "[]" {
return i, 1, nil
}
if strings.TrimSpace(m[1]) != "" {
return 0, 0, fmt.Errorf("entry %q writes its variants inline (%q); this job only edits block lists", e.Name, strings.TrimSpace(m[1]))
}
last := i
for j := i + 1; j < e.EndLine && variantItem.MatchString(lines[j]); j++ {
last = j
}
return last + 1, 0, nil
}
if inlineName.MatchString(lines[e.StartLine]) {
return e.StartLine + 1, 0, nil
}
for i := e.StartLine; i < e.EndLine; i++ {
if keyName.MatchString(lines[i]) {
return i + 1, 0, nil
}
}
return 0, 0, fmt.Errorf("entry %q has no name line to anchor the insertion to", e.Name)
}
// QuoteName quotes a variant reference when the name would otherwise change
// meaning as bare YAML. Config-suffixed names carry a ":" and always need it.
func QuoteName(name string) string {
if unsafeInName.MatchString(name) {
return `"` + strings.ReplaceAll(name, `"`, `\"`) + `"`
}
return name
}

133
.github/ci/variantproposals/body.go vendored Normal file
View File

@@ -0,0 +1,133 @@
package main
import (
"fmt"
"strings"
)
// RenderBody writes the pull request body.
//
// The body is the product of this job, not the diff. Grouping is a judgement
// call that has gone wrong in both directions before, so a reviewer has to be
// able to accept or reject each family from the body alone, without opening
// HuggingFace to work out whether two entries hold the same weights.
func RenderBody(r *Result, ledgerPath string) string {
var b strings.Builder
b.WriteString("## Proposed gallery variant groupings\n\n")
b.WriteString("This is a proposal, not a decision. The gallery agent adds one build per model and never joins an existing family, so entries that are alternative builds of the same weights drift apart as the gallery grows. This job re-applies the grouping heuristics from the manual sweeps and asks a human to confirm.\n\n")
b.WriteString("Each family below lists the parent, the variants, and the evidence that they are the same weights. **Reject anything whose evidence you do not believe.**\n\n")
b.WriteString(fmt.Sprintf("To decline a family permanently, add one line to `%s` in this pull request and close it:\n\n", ledgerPath))
b.WriteString("```yaml\npairs:\n - {parent: some-model, variant: some-model-thing, reason: \"different finetune\"}\n```\n\n")
b.WriteString(fmt.Sprintf("### Proposed families (%d)\n\n", len(r.Families)))
if len(r.Families) == 0 {
b.WriteString("None.\n\n")
}
for _, f := range r.Families {
b.WriteString(fmt.Sprintf("#### `%s`\n\n", f.Parent))
b.WriteString("| variant | signals | evidence |\n|---|---|---|\n")
for _, p := range f.Proposals {
b.WriteString(fmt.Sprintf("| `%s` | %s | %s |\n", p.Variant, joinSignals(p.Evidence.Signals), describeEvidence(p.Evidence)))
}
b.WriteString("\n")
}
b.WriteString(fmt.Sprintf("### Declined by the ledger (%d)\n\n", len(r.Suppressed)))
if len(r.Suppressed) == 0 {
b.WriteString("Nothing the heuristics found was already on the ledger.\n\n")
} else {
b.WriteString("Candidates the heuristics found and the ledger has already settled. They are listed so the ledger's effect stays visible rather than silently shrinking the job's output.\n\n")
for _, s := range r.Suppressed {
b.WriteString(fmt.Sprintf("- `%s` + `%s`: %s\n", s.A, s.B, s.Reason))
}
b.WriteString("\n")
}
if len(r.AliasSkipped) > 0 {
b.WriteString(fmt.Sprintf("### Aliases, not variants (%d)\n\n", len(r.AliasSkipped)))
b.WriteString("These entries install byte for byte the same payload. An alias exists so clients can send a particular name; folding it under another entry would hide that name.\n\n")
for _, s := range r.AliasSkipped {
b.WriteString(fmt.Sprintf("- `%s` + `%s`: %s\n", s.A, s.B, s.Reason))
}
b.WriteString("\n")
}
if len(r.Refusals) > 0 {
b.WriteString(fmt.Sprintf("### Found but refused (%d)\n\n", len(r.Refusals)))
b.WriteString("Candidates the heuristics found but the authoring rules would not let this job write. They need a human edit or a rule change.\n\n")
for _, ref := range r.Refusals {
b.WriteString(fmt.Sprintf("- %s: %s\n", codeList(ref.Members), ref.Reason))
}
b.WriteString("\n")
}
b.WriteString("---\n\nOpened by `.github/ci/variantproposals`. Heuristics and the rejection ledger live there and in the ledger file; a wrong proposal is a bug in one of the two.\n")
return b.String()
}
func joinSignals(signals []Signal) string {
if len(signals) == 0 {
return "inferred through another member of the family"
}
out := make([]string, 0, len(signals))
for _, s := range signals {
out = append(out, "`"+string(s)+"`")
}
return strings.Join(out, ", ")
}
func describeEvidence(e Evidence) string {
var parts []string
if e.SharedStem != "" {
parts = append(parts, fmt.Sprintf("same name once quantization markers are stripped: `%s`", e.SharedStem))
}
if e.SharedFile != "" {
parts = append(parts, fmt.Sprintf("same primary weight filename once quantization markers are stripped: `%s`", e.SharedFile))
}
if e.SharedRepo != "" {
parts = append(parts, fmt.Sprintf("same upstream repo `%s`", e.SharedRepo))
}
if len(e.QuantTokens) > 0 {
parts = append(parts, "differing quantization tokens: `"+strings.Join(e.QuantTokens, "`, `")+"`")
}
if len(parts) == 0 {
return "reached this family through another member"
}
return strings.Join(parts, "; ")
}
func codeList(names []string) string {
out := make([]string, 0, len(names))
for _, n := range names {
out = append(out, "`"+n+"`")
}
return strings.Join(out, " + ")
}
// RenderSummary is the terminal-facing digest of a run, so the workflow log
// says what happened without anyone opening the pull request.
func RenderSummary(r *Result) string {
var b strings.Builder
fmt.Fprintf(&b, "families proposed: %d\n", len(r.Families))
for _, f := range r.Families {
names := make([]string, 0, len(f.Proposals))
for _, p := range f.Proposals {
names = append(names, p.Variant)
}
fmt.Fprintf(&b, " %s <- %s\n", f.Parent, strings.Join(names, ", "))
}
fmt.Fprintf(&b, "declined by ledger: %d\n", len(r.Suppressed))
for _, s := range r.Suppressed {
fmt.Fprintf(&b, " %s\n", s)
}
fmt.Fprintf(&b, "aliases skipped: %d\n", len(r.AliasSkipped))
for _, s := range r.AliasSkipped {
fmt.Fprintf(&b, " %s\n", s)
}
fmt.Fprintf(&b, "refused: %d\n", len(r.Refusals))
for _, ref := range r.Refusals {
fmt.Fprintf(&b, " %s: %s\n", strings.Join(ref.Members, " + "), ref.Reason)
}
return b.String()
}

42
.github/ci/variantproposals/edit.go vendored Normal file
View File

@@ -0,0 +1,42 @@
package main
import (
"fmt"
"strings"
"github.com/mudler/LocalAI/.github/ci/galleryedit"
)
// ApplyFamilies writes the proposed variant lists into the index text.
//
// The line editing itself lives in galleryedit, shared with the apexentries
// generator. Both jobs add variants to entries the gallery already ships, and a
// second answer to "where does a variants block go" would drift from this one;
// see that package for why the edit is textual rather than a YAML round trip.
func ApplyFamilies(ix *Index, families []Family) ([]string, error) {
byName, _ := ix.ByName()
var inserts []galleryedit.Insert
for _, f := range families {
entry, ok := byName[strings.ToLower(f.Parent)]
if !ok {
return nil, fmt.Errorf("parent %q is not in the index", f.Parent)
}
variants := make([]string, 0, len(f.Proposals))
for _, p := range f.Proposals {
variants = append(variants, p.Variant)
}
inserts = append(inserts, galleryedit.Insert{
Entry: galleryedit.Entry{
Name: entry.Name,
StartLine: entry.StartLine,
EndLine: entry.EndLine,
},
Variants: variants,
})
}
return galleryedit.Apply(ix.Lines, inserts)
}

153
.github/ci/variantproposals/edit_test.go vendored Normal file
View File

@@ -0,0 +1,153 @@
package main
import (
"strings"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("ApplyFamilies", func() {
apply := func(ix *Index, families []Family) []string {
lines, err := ApplyFamilies(ix, families)
ExpectWithOffset(1, err).ToNot(HaveOccurred())
return lines
}
// insertedLines is what a reviewer would see in the diff. A textual editor
// that reflowed the file would show thousands here, which is the failure
// this whole approach exists to avoid.
insertedLines := func(before, after []string) int {
remaining := map[string]int{}
for _, l := range before {
remaining[l]++
}
n := 0
for _, l := range after {
if remaining[l] > 0 {
remaining[l]--
continue
}
n++
}
return n
}
It("adds a variants block right after the entry's name and touches nothing else", func() {
ix := indexOf(
entryYAML("foo-model", "acme/repo", "foo-model-Q4_K_M.gguf", "aa"),
entryYAML("foo-model-q8_0", "acme/repo", "foo-model-Q8_0.gguf", "bb"),
)
out := apply(ix, []Family{{Parent: "foo-model", Proposals: []Proposal{{Variant: "foo-model-q8_0"}}}})
Expect(out[0]).To(Equal("- name: foo-model"))
Expect(out[1]).To(Equal(" variants:"))
Expect(out[2]).To(Equal(" - model: foo-model-q8_0"))
Expect(len(out)).To(Equal(len(ix.Lines) + 2))
Expect(insertedLines(ix.Lines, out)).To(Equal(2))
})
It("appends to a variants block that already exists", func() {
ix := indexOf(`- name: partial
variants:
- model: partial-q8_0
url: u
overrides:
parameters:
model: partial-Q4_K_M.gguf
`, entryYAML("partial-f16", "acme/repo", "partial-f16.gguf", "cc"))
out := apply(ix, []Family{{Parent: "partial", Proposals: []Proposal{{Variant: "partial-f16"}}}})
Expect(out[1]).To(Equal(" variants:"))
Expect(out[2]).To(Equal(" - model: partial-q8_0"))
Expect(out[3]).To(Equal(" - model: partial-f16"))
Expect(out[4]).To(Equal(" url: u"))
})
It("replaces an explicit empty list rather than leaving two variants keys", func() {
ix := indexOf(`- name: emptied
variants: []
url: u
`, entryYAML("emptied-q8_0", "acme/repo", "emptied-Q8_0.gguf", "cc"))
out := apply(ix, []Family{{Parent: "emptied", Proposals: []Proposal{{Variant: "emptied-q8_0"}}}})
Expect(strings.Join(out[:4], "\n")).To(Equal("- name: emptied\n variants:\n - model: emptied-q8_0\n url: u"))
Expect(strings.Count(strings.Join(out, "\n"), "variants:")).To(Equal(1))
})
It("quotes a config-suffixed name so the reference stays a string", func() {
ix := indexOf(
entryYAML("phi-2-chat", "acme/repo", "phi-2-chat-Q4_K_M.gguf", "aa"),
entryYAML("phi-2-chat:Q8_0", "acme/repo", "phi-2-chat-Q8_0.gguf", "bb"),
)
out := apply(ix, []Family{{Parent: "phi-2-chat", Proposals: []Proposal{{Variant: "phi-2-chat:Q8_0"}}}})
Expect(out[2]).To(Equal(` - model: "phi-2-chat:Q8_0"`))
// The result has to still be a gallery, and the reference has to
// resolve to the entry it names.
reparsed, err := ParseIndex(strings.Join(out, "\n"))
Expect(err).ToNot(HaveOccurred())
Expect(reparsed.Entries[0].Variants).To(ConsistOf(VariantRef{Model: "phi-2-chat:Q8_0"}))
})
It("keeps line numbers correct when several entries are edited at once", func() {
ix := indexOf(
entryYAML("alpha", "acme/repo", "alpha-Q4_K_M.gguf", "aa"),
entryYAML("alpha-q8_0", "acme/repo", "alpha-Q8_0.gguf", "bb"),
entryYAML("beta", "acme/repo", "beta-Q4_K_M.gguf", "cc"),
entryYAML("beta-q8_0", "acme/repo", "beta-Q8_0.gguf", "dd"),
)
out := apply(ix, []Family{
{Parent: "alpha", Proposals: []Proposal{{Variant: "alpha-q8_0"}}},
{Parent: "beta", Proposals: []Proposal{{Variant: "beta-q8_0"}}},
})
reparsed, err := ParseIndex(strings.Join(out, "\n"))
Expect(err).ToNot(HaveOccurred())
Expect(reparsed.Entries).To(HaveLen(4))
Expect(reparsed.Entries[0].Variants).To(ConsistOf(VariantRef{Model: "alpha-q8_0"}))
Expect(reparsed.Entries[2].Variants).To(ConsistOf(VariantRef{Model: "beta-q8_0"}))
Expect(reparsed.Entries[1].Variants).To(BeEmpty())
Expect(reparsed.Entries[3].Variants).To(BeEmpty())
})
It("fails loudly rather than editing an entry it cannot find", func() {
ix := indexOf(entryYAML("only", "acme/repo", "only-Q4_K_M.gguf", "aa"))
_, err := ApplyFamilies(ix, []Family{{Parent: "missing", Proposals: []Proposal{{Variant: "x"}}}})
Expect(err).To(MatchError(ContainSubstring("not in the index")))
})
})
var _ = Describe("ParseIndex", func() {
It("records the anchor an entry defines and the anchor an entry merges", func() {
ix := indexOf(`- &anc
name: anchored
url: u
`, `- !!merge <<: *anc
name: child
`)
Expect(ix.Entries[0].AnchorName).To(Equal("anc"))
Expect(ix.Entries[1].MergesFrom).To(Equal("anc"))
Expect(ix.MergeChildren("anc")).To(HaveLen(1))
})
It("carries merged values into the child, so an inherited variants key is visible", func() {
ix := indexOf(`- &anc
name: anchored
url: u
variants:
- model: something
`, `- !!merge <<: *anc
name: child
`)
Expect(ix.Entries[1].HasVariants()).To(BeTrue())
})
It("refuses a list item that decodes to nothing", func() {
// Every line number the editor works from comes from pairing decoded
// entries with top level list items. If those two views can disagree,
// the editor writes into the wrong entry, so the parse refuses instead.
_, err := ParseIndex("- name: one\n url: u\n-\n")
Expect(err).To(MatchError(ContainSubstring("empty")))
})
})

281
.github/ci/variantproposals/index.go vendored Normal file
View File

@@ -0,0 +1,281 @@
package main
import (
"fmt"
"os"
"regexp"
"sort"
"strings"
"gopkg.in/yaml.v3"
"github.com/mudler/LocalAI/.github/ci/galleryedit"
)
// File is the subset of a gallery file entry the proposer reads.
type File struct {
Filename string `yaml:"filename"`
URI string `yaml:"uri"`
SHA256 string `yaml:"sha256"`
}
// VariantRef mirrors the gallery's variant reference.
type VariantRef struct {
Model string `yaml:"model"`
}
// GalleryEntry is one gallery entry, carrying both the semantics the heuristics need
// and the text range the editor needs.
//
// The two views are kept together deliberately. The editor must not round-trip
// the index through a YAML marshaller: the gallery is 40,000 lines and a
// reflowed diff cannot be reviewed, which defeats the entire point of a job
// whose output is a human decision.
type GalleryEntry struct {
Name string `yaml:"name"`
URL string `yaml:"url"`
ConfigFile map[string]any `yaml:"config_file"`
Overrides map[string]any `yaml:"overrides"`
Files []File `yaml:"files"`
Variants []VariantRef `yaml:"variants"`
// Index is the entry's position in gallery order.
Index int `yaml:"-"`
// StartLine and EndLine bound the entry's lines, zero based and half open.
StartLine int `yaml:"-"`
EndLine int `yaml:"-"`
// AnchorName is set when the entry defines a YAML anchor. Adding a variants
// key to such an entry is inherited by everything that merges it, which is
// why proposals involving anchors get special treatment.
AnchorName string `yaml:"-"`
// MergesFrom is the anchor this entry pulls in with "!!merge <<:".
MergesFrom string `yaml:"-"`
}
// Index is a parsed gallery index: entries plus the exact lines they came from.
type Index struct {
Lines []string
Entries []*GalleryEntry
}
var (
anchorStart = regexp.MustCompile(`^- &(\S+)`)
mergeStart = regexp.MustCompile(`^- !!merge <<: \*(\S+)`)
)
// LoadIndex reads and parses a gallery index file.
func LoadIndex(path string) (*Index, error) {
data, err := os.ReadFile(path)
if err != nil {
return nil, err
}
return ParseIndex(string(data))
}
// ParseIndex builds an Index from the raw text of a gallery index.
//
// The YAML decode and the textual scan are cross checked against each other: if
// they disagree on how many entries there are, every line number the editor
// would use is suspect, so the run fails rather than editing the wrong entry.
func ParseIndex(text string) (*Index, error) {
var entries []*GalleryEntry
if err := yaml.Unmarshal([]byte(text), &entries); err != nil {
return nil, fmt.Errorf("decoding gallery index: %w", err)
}
lines, starts := galleryedit.Scan(text)
if len(starts) != len(entries) {
return nil, fmt.Errorf("gallery index has %d decoded entries but %d top level list items; refusing to edit by line number", len(entries), len(starts))
}
for i, e := range entries {
if e == nil {
return nil, fmt.Errorf("gallery index list item %d is empty; refusing to edit by line number", i)
}
e.Index = i
e.StartLine = starts[i]
if i+1 < len(starts) {
e.EndLine = starts[i+1]
} else {
e.EndLine = len(lines)
}
if m := anchorStart.FindStringSubmatch(lines[e.StartLine]); m != nil {
e.AnchorName = m[1]
}
if m := mergeStart.FindStringSubmatch(lines[e.StartLine]); m != nil {
e.MergesFrom = m[1]
}
}
return &Index{Lines: lines, Entries: entries}, nil
}
// MergeChildren lists the entries that pull in the given anchor.
func (ix *Index) MergeChildren(anchor string) []*GalleryEntry {
var out []*GalleryEntry
for _, e := range ix.Entries {
if e.MergesFrom == anchor {
out = append(out, e)
}
}
return out
}
// ByName indexes entries by lowercased name. A name appearing twice keeps the
// first occurrence, matching the gallery's own first-match-wins resolution, and
// the duplicates are returned so the caller can refuse to touch them: a
// proposal naming an ambiguous entry cannot be reviewed.
func (ix *Index) ByName() (map[string]*GalleryEntry, map[string]int) {
byName := make(map[string]*GalleryEntry, len(ix.Entries))
counts := make(map[string]int, len(ix.Entries))
for _, e := range ix.Entries {
key := strings.ToLower(e.Name)
counts[key]++
if _, seen := byName[key]; !seen {
byName[key] = e
}
}
dupes := map[string]int{}
for name, n := range counts {
if n > 1 {
dupes[name] = n
}
}
return byName, dupes
}
// Installable reports whether installing this entry would put anything on disk.
// A variant target that installs nothing is a dead end for the selector, so it
// is never proposed as one.
func (e *GalleryEntry) Installable() bool {
return e.URL != "" || len(e.ConfigFile) > 0 || len(e.Overrides) > 0 || len(e.Files) > 0
}
// HasVariants reports whether the entry already offers builds of its own. Such
// an entry cannot be a variant target: nesting is what the gallery's own
// resolution refuses.
func (e *GalleryEntry) HasVariants() bool {
return len(e.Variants) > 0
}
// auxiliaryFile matches the shared side files that several unrelated models
// legitimately hand out the same copy of. Grouping on one of these is how an
// earlier sweep linked four wan-2.1 entries to each other and Z-Image-Turbo to
// qwen3-4b: they shared a text encoder, not weights.
var auxiliaryFile = regexp.MustCompile(`(?i)(mmproj|vae|clip|t5|umt5|text_?encoder|tokenizer|\bae\b|^ae\.|scheduler|config)`)
// IsAuxiliaryFile reports whether a filename is a side file rather than the
// model's own weights.
func IsAuxiliaryFile(filename string) bool {
base := filename
if i := strings.LastIndex(base, "/"); i >= 0 {
base = base[i+1:]
}
return auxiliaryFile.MatchString(base)
}
// PrimaryWeightFile returns the filename of the entry's own weights, and
// whether one could be identified unambiguously.
//
// The declared overrides.parameters.model wins because that is the file the
// backend is actually pointed at. Falling back to the file list only works when
// exactly one non-auxiliary file is present; anything else is ambiguous, and
// guessing is precisely the failure mode this heuristic has already had.
func (e *GalleryEntry) PrimaryWeightFile() (string, bool) {
if params, ok := e.Overrides["parameters"].(map[string]any); ok {
if model, ok := params["model"].(string); ok && model != "" && !IsAuxiliaryFile(model) {
return model, true
}
}
var candidates []string
for _, f := range e.Files {
if f.Filename == "" || IsAuxiliaryFile(f.Filename) {
continue
}
candidates = append(candidates, f.Filename)
}
if len(candidates) == 1 {
return candidates[0], true
}
return "", false
}
// SourceRepo returns the upstream repository the entry's files come from, as a
// coarse "host + owner + repo" key.
func (e *GalleryEntry) SourceRepo() string {
for _, f := range e.Files {
if f.URI == "" {
continue
}
return repoKey(f.URI)
}
return ""
}
func repoKey(uri string) string {
u := strings.ToLower(uri)
u = strings.TrimPrefix(u, "huggingface://")
u = strings.TrimPrefix(u, "https://huggingface.co/")
u = strings.TrimPrefix(u, "http://huggingface.co/")
parts := strings.Split(u, "/")
if len(parts) >= 2 {
return parts[0] + "/" + parts[1]
}
return u
}
// SameInstallPayload reports whether two entries install byte for byte the same
// thing.
//
// Entries like this are aliases, not variants. whisper-1 exists so a client
// speaking the OpenAI API can send that name and get whisper-base; folding it
// under whisper-base as a variant would hide the very name clients send.
func SameInstallPayload(a, b *GalleryEntry) bool {
if a.URL != b.URL {
return false
}
if !sameYAML(a.Overrides, b.Overrides) || !sameYAML(a.ConfigFile, b.ConfigFile) {
return false
}
return sameChecksums(a.Files, b.Files)
}
func sameChecksums(a, b []File) bool {
if len(a) != len(b) || len(a) == 0 {
return false
}
ha := make([]string, 0, len(a))
hb := make([]string, 0, len(b))
for _, f := range a {
if f.SHA256 == "" {
return false
}
ha = append(ha, f.SHA256)
}
for _, f := range b {
if f.SHA256 == "" {
return false
}
hb = append(hb, f.SHA256)
}
sort.Strings(ha)
sort.Strings(hb)
for i := range ha {
if ha[i] != hb[i] {
return false
}
}
return true
}
func sameYAML(a, b any) bool {
ba, err := yaml.Marshal(a)
if err != nil {
return false
}
bb, err := yaml.Marshal(b)
if err != nil {
return false
}
return string(ba) == string(bb)
}

180
.github/ci/variantproposals/ledger.go vendored Normal file
View File

@@ -0,0 +1,180 @@
package main
import (
"fmt"
"os"
"sort"
"strings"
"gopkg.in/yaml.v3"
)
// Ledger records the grouping decisions a human has already made against the
// proposer, so a declined candidate stays declined instead of coming back every
// night until reviewers stop reading the job's pull requests.
//
// It is checked in next to the gallery and is meant to be edited inside the
// proposal pull request itself: declining a family is adding one flow-mapping
// line under pairs or groups and closing the PR.
type Ledger struct {
// Tokens are name segments that mark a distinct model rather than another
// build of the same one: finetune names, language codes, product suffixes.
// A candidate whose two names differ by any of these is never proposed.
Tokens []LedgerToken `yaml:"tokens"`
// Pairs are individual candidates a human considered and declined. Order
// does not matter: the pair is matched both ways round.
Pairs []LedgerPair `yaml:"pairs"`
// Groups decline every pair drawn from a set at once, for families like a
// per-language release where listing each pair would be unreadable.
Groups []LedgerGroup `yaml:"groups"`
}
type LedgerToken struct {
Token string `yaml:"token"`
Reason string `yaml:"reason"`
}
type LedgerPair struct {
Parent string `yaml:"parent"`
Variant string `yaml:"variant"`
Reason string `yaml:"reason"`
}
type LedgerGroup struct {
Members []string `yaml:"members"`
Reason string `yaml:"reason"`
}
// LoadLedger reads a ledger file. A missing file is not an error: a gallery
// that has declined nothing yet is a legitimate state, and failing the job over
// it would only teach people to keep an empty file around.
func LoadLedger(path string) (*Ledger, error) {
data, err := os.ReadFile(path)
if os.IsNotExist(err) {
return &Ledger{}, nil
}
if err != nil {
return nil, err
}
return ParseLedger(data)
}
func ParseLedger(data []byte) (*Ledger, error) {
l := &Ledger{}
if err := yaml.Unmarshal(data, l); err != nil {
return nil, fmt.Errorf("parsing ledger: %w", err)
}
for i, t := range l.Tokens {
if strings.TrimSpace(t.Token) == "" {
return nil, fmt.Errorf("ledger tokens[%d] has an empty token", i)
}
}
for i, p := range l.Pairs {
if strings.TrimSpace(p.Parent) == "" || strings.TrimSpace(p.Variant) == "" {
return nil, fmt.Errorf("ledger pairs[%d] needs both parent and variant", i)
}
}
return l, nil
}
// Suppression is a ledger hit: why a candidate was not proposed, in words a
// reviewer can check against the ledger file.
type Suppression struct {
A string
B string
Reason string
}
func (s Suppression) String() string {
return fmt.Sprintf("%s + %s: %s", s.A, s.B, s.Reason)
}
// Suppresses reports whether the ledger has already declined pairing these two
// entries, and why.
//
// The token rule is applied to the segments the two names do not share. Two
// builds of the same weights differ only in quantization markers, so any
// ledgered token showing up in that difference is by construction a claim that
// the entries are different models.
func (l *Ledger) Suppresses(a, b string) (Suppression, bool) {
la, lb := strings.ToLower(a), strings.ToLower(b)
for _, p := range l.Pairs {
lp, lv := strings.ToLower(p.Parent), strings.ToLower(p.Variant)
if (lp == la && lv == lb) || (lp == lb && lv == la) {
return Suppression{A: a, B: b, Reason: p.Reason}, true
}
}
for _, g := range l.Groups {
var seenA, seenB bool
for _, m := range g.Members {
lm := strings.ToLower(m)
if lm == la {
seenA = true
}
if lm == lb {
seenB = true
}
}
if seenA && seenB {
return Suppression{A: a, B: b, Reason: g.Reason}, true
}
}
diff := differingSegments(la, lb)
for _, t := range l.Tokens {
token := strings.ToLower(strings.TrimSpace(t.Token))
if _, ok := diff[token]; ok {
reason := t.Reason
if reason == "" {
reason = fmt.Sprintf("names differ by %q", token)
}
return Suppression{A: a, B: b, Reason: fmt.Sprintf("%s (token %q)", reason, token)}, true
}
}
return Suppression{}, false
}
// segments splits a name into the atoms the token rules are written against.
func segments(name string) []string {
fields := strings.FieldsFunc(strings.ToLower(name), func(r rune) bool {
return r == '-' || r == '_' || r == '.' || r == ':' || r == '/'
})
return fields
}
// differingSegments returns the set of segments present in exactly one of the
// two names.
func differingSegments(a, b string) map[string]struct{} {
setA := map[string]int{}
for _, s := range segments(a) {
setA[s]++
}
setB := map[string]int{}
for _, s := range segments(b) {
setB[s]++
}
diff := map[string]struct{}{}
for s := range setA {
if setB[s] == 0 {
diff[s] = struct{}{}
}
}
for s := range setB {
if setA[s] == 0 {
diff[s] = struct{}{}
}
}
return diff
}
// SortedSuppressions gives the ledger's effect on one run in a stable order, so
// the pull request body reads the same way for the same gallery.
func SortedSuppressions(in []Suppression) []Suppression {
out := append([]Suppression(nil), in...)
sort.Slice(out, func(i, j int) bool {
if out[i].A != out[j].A {
return out[i].A < out[j].A
}
return out[i].B < out[j].B
})
return out
}

65
.github/ci/variantproposals/main.go vendored Normal file
View File

@@ -0,0 +1,65 @@
// Command variant-proposals looks for gallery entries that are alternative
// builds of the same weights but are not grouped under one another, and writes
// a proposal for a human to accept or reject.
//
// It never decides. Grouping has gone wrong repeatedly in both directions, so
// the job's value is catching drift and surfacing candidates with their
// evidence, not automating the call. The scheduled workflow feeds its output to
// a pull request in the same shape as .github/checksum_checker.sh.
package main
import (
"flag"
"fmt"
"os"
"strings"
)
func main() {
index := flag.String("index", "gallery/index.yaml", "path to the gallery index")
ledger := flag.String("ledger", "gallery/variant-exclusions.yaml", "path to the rejection ledger")
bodyOut := flag.String("body-out", "", "write the pull request body here")
apply := flag.Bool("apply", false, "write the proposed groupings back into the index")
flag.Parse()
if err := run(*index, *ledger, *bodyOut, *apply); err != nil {
fmt.Fprintln(os.Stderr, "variant-proposals:", err)
os.Exit(1)
}
}
func run(indexPath, ledgerPath, bodyOut string, apply bool) error {
ix, err := LoadIndex(indexPath)
if err != nil {
return err
}
ledger, err := LoadLedger(ledgerPath)
if err != nil {
return err
}
result := Propose(ix, ledger)
fmt.Print(RenderSummary(result))
if !result.HasProposals() {
// An empty pull request every night is how a proposal job gets muted.
fmt.Println("nothing to propose")
return nil
}
if bodyOut != "" {
if err := os.WriteFile(bodyOut, []byte(RenderBody(result, ledgerPath)), 0o644); err != nil {
return err
}
}
if !apply {
return nil
}
lines, err := ApplyFamilies(ix, result.Families)
if err != nil {
return err
}
return os.WriteFile(indexPath, []byte(strings.Join(lines, "\n")), 0o644)
}

615
.github/ci/variantproposals/propose.go vendored Normal file
View File

@@ -0,0 +1,615 @@
package main
import (
"fmt"
"regexp"
"sort"
"strings"
)
// Signal names the grouping heuristic that linked two entries.
type Signal string
const (
// SignalName is "same name once quantization markers are stripped".
SignalName Signal = "name-modulo-quant"
// SignalConfigSuffix is the ":" convention, foo:q8_0 as a build of foo.
SignalConfigSuffix Signal = "config-suffix"
// SignalWeightFile is "same primary weight filename once quantization
// markers are stripped", auxiliary files excluded.
SignalWeightFile Signal = "weight-filename"
)
// Evidence is what a reviewer needs in order to agree or disagree without
// opening HuggingFace: what the two entries share, and what differs.
type Evidence struct {
Signals []Signal
SharedStem string
SharedFile string
SharedRepo string
QuantTokens []string
}
// Proposal is one variant target offered to one parent.
type Proposal struct {
Variant string
Evidence Evidence
}
// Family is a complete proposal: one parent gaining one or more variants.
type Family struct {
Parent string
Proposals []Proposal
}
// Refusal is a family the heuristics found but the rules would not let through.
// Refusals are reported rather than dropped: a candidate the job keeps refusing
// is either a rule worth revisiting or a gallery bug worth fixing.
type Refusal struct {
Members []string
Reason string
}
// Result is one run of the proposer.
type Result struct {
Families []Family
Refusals []Refusal
Suppressed []Suppression
AliasSkipped []Suppression
}
// HasProposals reports whether the run found anything to open a pull request
// about. A job that opens an empty pull request every night is a job people
// filter out of their inbox.
func (r *Result) HasProposals() bool {
return len(r.Families) > 0
}
// sizeToken matches a parameter-count marker: 8b, 1.7b, a3b for an active
// expert count, e2b for the Gemma effective sizes, 8x7b for a mixture.
//
// This is a structural rule rather than a ledger entry because it is about the
// shape of the token, not about any one model. Different parameter sizes were
// mis-grouped by an earlier sweep and the failure is systematic.
var sizeToken = regexp.MustCompile(`^(?:[0-9]+(?:\.[0-9]+)?[bm]|[ae][0-9]+(?:\.[0-9]+)?b|[0-9]+x[0-9]+(?:\.[0-9]+)?b)$`)
func differsByParameterSize(a, b string) (string, bool) {
for seg := range differingSegments(a, b) {
if sizeToken.MatchString(seg) {
return seg, true
}
}
return "", false
}
// genericFileStem lists weight filenames too generic to be evidence of
// anything. Two entries both shipping "model.safetensors" share a convention,
// not a set of weights.
var genericFileStem = map[string]struct{}{
"model": {}, "weights": {}, "pytorch_model": {}, "diffusion_pytorch_model": {},
"consolidated": {}, "ggml-model": {}, "model-00001-of-00002": {},
}
// minFileStemLength keeps short, collision-prone filename stems from linking
// unrelated entries.
const minFileStemLength = 6
type pair struct {
a, b int
evidence Evidence
}
// Propose runs the grouping heuristics over a gallery index and returns what it
// would offer a human, what it refused, and what the ledger silenced.
//
// Nothing here touches the network or git, and the index is not modified.
func Propose(ix *Index, ledger *Ledger) *Result {
if ledger == nil {
ledger = &Ledger{}
}
result := &Result{}
byName, dupes := ix.ByName()
// Existing relationships. A target already claimed must not be claimed
// again, and two entries already in one family need no proposal.
claimedBy := map[string]string{}
familyOf := map[string]string{}
for _, e := range ix.Entries {
if !e.HasVariants() {
continue
}
familyOf[strings.ToLower(e.Name)] = strings.ToLower(e.Name)
for _, v := range e.Variants {
target := strings.ToLower(v.Model)
if _, taken := claimedBy[target]; !taken {
claimedBy[target] = strings.ToLower(e.Name)
}
familyOf[target] = strings.ToLower(e.Name)
}
}
candidates := map[[2]int]*Evidence{}
addPair := func(i, j int, sig Signal, apply func(*Evidence)) {
if i == j {
return
}
if i > j {
i, j = j, i
}
key := [2]int{i, j}
ev, ok := candidates[key]
if !ok {
ev = &Evidence{}
candidates[key] = ev
}
for _, s := range ev.Signals {
if s == sig {
apply(ev)
return
}
}
ev.Signals = append(ev.Signals, sig)
apply(ev)
}
// Signal 1 and 2: entries sharing a name stem.
byStem := map[string][]int{}
for _, e := range ix.Entries {
if e.Name == "" {
continue
}
byStem[NameStem(e.Name)] = append(byStem[NameStem(e.Name)], e.Index)
}
for stem, members := range byStem {
if len(members) < 2 {
continue
}
for i := 0; i < len(members); i++ {
for j := i + 1; j < len(members); j++ {
a, b := ix.Entries[members[i]], ix.Entries[members[j]]
sig := SignalName
if HasConfigSuffix(a.Name) || HasConfigSuffix(b.Name) {
sig = SignalConfigSuffix
}
// The bare parent carries no marker in its name, so the
// evidence would read "differs by q8_0" and say nothing about
// what the parent is. The weight filenames fill that in.
fa, _ := a.PrimaryWeightFile()
fb, _ := b.PrimaryWeightFile()
addPair(members[i], members[j], sig, func(ev *Evidence) {
ev.SharedStem = stem
ev.QuantTokens = quantDifference(a.Name, b.Name, fa, fb)
})
}
}
}
// Signal 3: entries whose own weight file is the same file at a different
// quantization. Auxiliary files never take part.
byFile := map[string][]int{}
for _, e := range ix.Entries {
primary, ok := e.PrimaryWeightFile()
if !ok {
continue
}
stem := FileStem(primary)
if len(stem) < minFileStemLength {
continue
}
if _, generic := genericFileStem[stem]; generic {
continue
}
byFile[stem] = append(byFile[stem], e.Index)
}
for stem, members := range byFile {
if len(members) < 2 {
continue
}
for i := 0; i < len(members); i++ {
for j := i + 1; j < len(members); j++ {
a, b := ix.Entries[members[i]], ix.Entries[members[j]]
// The filename alone is not evidence. Publishers reuse the
// upstream filename for finetunes and for models that merely
// embed the base weights: bert-embeddings, an ultravox audio
// model and a roleplay finetune all ship a file called
// llama-3.2-1b-instruct-q4_k_m.gguf. Requiring the same
// upstream repository turns the signal back into what it
// claims to be, one repo publishing one file at two
// quantizations. Two repos holding the same weights is a fact
// no filename proves, so it stays a human call.
repo := a.SourceRepo()
if repo == "" || repo != b.SourceRepo() {
continue
}
fa, _ := a.PrimaryWeightFile()
fb, _ := b.PrimaryWeightFile()
addPair(members[i], members[j], SignalWeightFile, func(ev *Evidence) {
ev.SharedFile = stem
ev.SharedRepo = repo
if len(ev.QuantTokens) == 0 {
ev.QuantTokens = quantDifference(fa, fb)
}
})
}
}
}
// Filter candidates. Everything dropped here is dropped for a reason a
// reviewer can read back off the ledger or the rules.
var kept []pair
for key, ev := range candidates {
a, b := ix.Entries[key[0]], ix.Entries[key[1]]
la, lb := strings.ToLower(a.Name), strings.ToLower(b.Name)
if la == lb {
continue
}
if dupes[la] > 0 || dupes[lb] > 0 {
result.Refusals = append(result.Refusals, Refusal{
Members: []string{a.Name, b.Name},
Reason: "one of these names appears more than once in the gallery, so a variant reference to it is ambiguous",
})
continue
}
if fa, fb := familyOf[la], familyOf[lb]; fa != "" && fa == fb {
continue
}
if seg, differs := differsByParameterSize(la, lb); differs {
result.Suppressed = append(result.Suppressed, Suppression{
A: a.Name, B: b.Name, Reason: fmt.Sprintf("different parameter sizes (segment %q)", seg),
})
continue
}
if s, ok := ledger.Suppresses(a.Name, b.Name); ok {
result.Suppressed = append(result.Suppressed, s)
continue
}
if SameInstallPayload(a, b) {
result.AliasSkipped = append(result.AliasSkipped, Suppression{
A: a.Name, B: b.Name,
Reason: "identical install payload; these are aliases of one build, not alternative builds",
})
continue
}
kept = append(kept, pair{a: key[0], b: key[1], evidence: *ev})
}
sort.Slice(kept, func(i, j int) bool {
if kept[i].a != kept[j].a {
return kept[i].a < kept[j].a
}
return kept[i].b < kept[j].b
})
// Components. A pair from either signal joins the same family, so a chain
// of alternative builds discovered by different signals stays one family
// rather than two overlapping ones that would double claim a target.
parent := map[int]int{}
var find func(int) int
find = func(x int) int {
if p, ok := parent[x]; ok && p != x {
parent[x] = find(p)
return parent[x]
}
if _, ok := parent[x]; !ok {
parent[x] = x
}
return parent[x]
}
union := func(x, y int) {
rx, ry := find(x), find(y)
if rx != ry {
parent[ry] = rx
}
}
evidenceFor := map[[2]int]Evidence{}
for _, p := range kept {
union(p.a, p.b)
evidenceFor[[2]int{p.a, p.b}] = p.evidence
}
components := map[int][]int{}
for _, p := range kept {
for _, m := range []int{p.a, p.b} {
root := find(m)
if !contains(components[root], m) {
components[root] = append(components[root], m)
}
}
}
roots := make([]int, 0, len(components))
for r := range components {
roots = append(roots, r)
}
sort.Ints(roots)
proposedTargets := map[string]string{}
for _, root := range roots {
members := components[root]
sort.Ints(members)
family, refusal := buildFamily(ix, members, evidenceFor, claimedBy, proposedTargets, byName)
if refusal != nil {
result.Refusals = append(result.Refusals, *refusal)
continue
}
if family == nil {
continue
}
for _, p := range family.Proposals {
proposedTargets[strings.ToLower(p.Variant)] = family.Parent
}
result.Families = append(result.Families, *family)
}
sort.Slice(result.Families, func(i, j int) bool { return result.Families[i].Parent < result.Families[j].Parent })
result.Suppressed = SortedSuppressions(result.Suppressed)
result.AliasSkipped = SortedSuppressions(result.AliasSkipped)
result.Refusals = dedupeRefusals(result.Refusals)
return result
}
// dedupeRefusals collapses the same refusal reached from both orderings of a
// pair, and sorts what is left. A reviewer reading the same complaint twice
// learns to skim the section.
func dedupeRefusals(in []Refusal) []Refusal {
seen := map[string]struct{}{}
var out []Refusal
for _, r := range in {
members := append([]string(nil), r.Members...)
sort.Strings(members)
key := strings.Join(members, "\x00") + "\x00" + r.Reason
if _, dup := seen[key]; dup {
continue
}
seen[key] = struct{}{}
out = append(out, r)
}
sort.Slice(out, func(i, j int) bool {
if a, b := strings.Join(out[i].Members, ","), strings.Join(out[j].Members, ","); a != b {
return a < b
}
return out[i].Reason < out[j].Reason
})
return out
}
func contains(xs []int, x int) bool {
for _, v := range xs {
if v == x {
return true
}
}
return false
}
// buildFamily turns a connected component into a proposal, or refuses it.
func buildFamily(ix *Index, members []int, evidenceFor map[[2]int]Evidence, claimedBy map[string]string, proposedTargets map[string]string, byName map[string]*GalleryEntry) (*Family, *Refusal) {
names := make([]string, 0, len(members))
for _, m := range members {
names = append(names, ix.Entries[m].Name)
}
parentIdx, err := selectParent(ix, members)
if err != nil {
return nil, &Refusal{Members: names, Reason: err.Error()}
}
parentEntry := ix.Entries[parentIdx]
parentName := strings.ToLower(parentEntry.Name)
// A parent that is itself somebody's variant would create a chain, which
// the gallery's own resolution refuses to install.
if owner, claimed := claimedBy[parentName]; claimed {
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("the natural parent %q is already a variant of %q; proposing it as a parent would nest variants", parentEntry.Name, owner)}
}
if owner, claimed := proposedTargets[parentName]; claimed {
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("the natural parent %q is already proposed as a variant of %q; proposing it as a parent would nest variants", parentEntry.Name, owner)}
}
// Adding a variants key to an anchor is inherited by every entry that
// merges it, silently grouping models nobody proposed. Handling that means
// editing each merging child too, which is a larger change than this job
// should make unsupervised, so it refuses and hands the reviewer the list.
if parentEntry.AnchorName != "" {
children := ix.MergeChildren(parentEntry.AnchorName)
if len(children) > 0 {
childNames := make([]string, 0, len(children))
for _, c := range children {
childNames = append(childNames, c.Name)
}
return nil, &Refusal{
Members: names,
Reason: fmt.Sprintf("the parent %q defines YAML anchor &%s, and a variants key added there is inherited by the %d entries that merge it (%s). Grouping this family by hand also means adding an explicit `variants: []` to each of those entries",
parentEntry.Name, parentEntry.AnchorName, len(children), strings.Join(childNames, ", ")),
}
}
}
existing := map[string]struct{}{}
for _, v := range parentEntry.Variants {
existing[strings.ToLower(v.Model)] = struct{}{}
}
family := &Family{Parent: parentEntry.Name}
for _, m := range members {
if m == parentIdx {
continue
}
target := ix.Entries[m]
lower := strings.ToLower(target.Name)
if _, already := existing[lower]; already {
continue
}
if target.HasVariants() {
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("%q already offers variants of its own, so it cannot itself be a variant target", target.Name)}
}
if !target.Installable() {
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("%q has no url, config_file, overrides or files, so it is not independently installable", target.Name)}
}
if owner, claimed := claimedBy[lower]; claimed && owner != parentName {
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("%q is already a variant of %q; a target claimed by two parents is not something the gallery resolves predictably", target.Name, owner)}
}
if owner, claimed := proposedTargets[lower]; claimed && owner != parentEntry.Name {
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("%q is already proposed as a variant of %q in this same run", target.Name, owner)}
}
family.Proposals = append(family.Proposals, Proposal{
Variant: target.Name,
Evidence: lookupEvidence(evidenceFor, parentIdx, m),
})
}
if len(family.Proposals) == 0 {
return nil, nil
}
sort.Slice(family.Proposals, func(i, j int) bool { return family.Proposals[i].Variant < family.Proposals[j].Variant })
return family, nil
}
func lookupEvidence(evidenceFor map[[2]int]Evidence, a, b int) Evidence {
if a > b {
a, b = b, a
}
if ev, ok := evidenceFor[[2]int{a, b}]; ok {
return ev
}
// The two entries reached the same family through a third one. Say so
// rather than inventing evidence that was never observed for this pair.
return Evidence{Signals: []Signal{SignalName}}
}
// selectParent picks the entry the others should hang off.
//
// The bare name wins when there is one: it is the name a user types and the one
// documentation links to. Otherwise the smallest build wins, judged by the
// quantization token in the entry's own weight filename, so the default install
// is the one most hosts can actually run.
func selectParent(ix *Index, members []int) (int, error) {
// The family's own stem: the one the most members reduce to, shortest name
// breaking a tie. An entry named exactly that is the bare entry.
stemCount := map[string]int{}
for _, m := range members {
stemCount[NameStem(ix.Entries[m].Name)]++
}
// Only a stem two or more members reduce to is the family's own stem. A
// stem reached by exactly one member is just that member's name, and
// treating it as the family stem would crown whichever name happens to be
// shortest rather than whichever build is the base one.
familyStem := ""
for stem, n := range stemCount {
if n < 2 {
continue
}
if familyStem == "" || n > stemCount[familyStem] ||
(n == stemCount[familyStem] && len(stem) < len(familyStem)) ||
(n == stemCount[familyStem] && len(stem) == len(familyStem) && stem < familyStem) {
familyStem = stem
}
}
var bare []int
for _, m := range members {
e := ix.Entries[m]
if HasConfigSuffix(e.Name) {
continue
}
if strings.ToLower(e.Name) == familyStem {
bare = append(bare, m)
}
}
if len(bare) == 1 {
return bare[0], nil
}
if len(bare) > 1 {
names := make([]string, 0, len(bare))
for _, m := range bare {
names = append(names, ix.Entries[m].Name)
}
return 0, fmt.Errorf("more than one entry is named exactly %q (%s), so which one is the base build is a judgement this job will not make", familyStem, strings.Join(names, ", "))
}
// No shared stem to be named after. An entry whose name every other member
// extends is still recognisably the base one, and this is the only handle
// left for families whose weights carry no readable quantization token at
// all, such as the ONNX builds.
if prefix, ok := uniquePrefixMember(ix, members); ok {
return prefix, nil
}
best := -1
bestWidth := 1 << 20
for _, m := range members {
e := ix.Entries[m]
width := unknownWidth
if primary, ok := e.PrimaryWeightFile(); ok {
width = BuildWidth(primary)
}
// Members are visited in gallery order, so a strict comparison leaves
// the earliest entry holding a tie and the choice is deterministic.
if width < bestWidth {
best, bestWidth = m, width
}
}
if best < 0 {
return 0, fmt.Errorf("no member could be identified as the smallest build")
}
if bestWidth == unknownWidth {
names := make([]string, 0, len(members))
for _, m := range members {
names = append(names, ix.Entries[m].Name)
}
return 0, fmt.Errorf("no member declares a weight file whose quantization can be read (%s), so the smallest build cannot be identified", strings.Join(names, ", "))
}
return best, nil
}
// uniquePrefixMember reports the single member whose name every other member's
// name starts with, if there is exactly one.
func uniquePrefixMember(ix *Index, members []int) (int, bool) {
found := -1
for _, m := range members {
name := strings.ToLower(ix.Entries[m].Name)
isPrefix := true
for _, other := range members {
if other == m {
continue
}
if !strings.HasPrefix(strings.ToLower(ix.Entries[other].Name), name) {
isPrefix = false
break
}
}
if !isPrefix {
continue
}
if found >= 0 {
return 0, false
}
found = m
}
return found, found >= 0
}
// quantDifference lists the quantization tokens that tell two names apart. It
// is the compact form of the evidence: "these differ only by q4_k_m vs q8_0".
func quantDifference(names ...string) []string {
var out []string
seen := map[string]struct{}{}
for _, name := range names {
// Filenames arrive here too, so the extension goes first and "/" counts
// as a separator. "_" deliberately does not: it holds "q4_k_m" together.
trimmed := weightExtension.ReplaceAllString(name, "")
for _, seg := range strings.FieldsFunc(strings.ToLower(trimmed), func(r rune) bool { return r == '-' || r == '/' }) {
if !IsQuantToken(seg) {
continue
}
if _, ok := seen[seg]; ok {
continue
}
seen[seg] = struct{}{}
out = append(out, seg)
}
}
sort.Strings(out)
return out
}

View File

@@ -0,0 +1,429 @@
package main
import (
"fmt"
"strings"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
// entryYAML writes one gallery entry with a single weight file, which is the
// shape almost every real entry has. Specs that need something else write the
// YAML out by hand.
func entryYAML(name, repo, filename, sha string) string {
return fmt.Sprintf(`- name: %s
url: github:mudler/LocalAI/gallery/virtual.yaml@master
overrides:
parameters:
model: %s
files:
- filename: %s
uri: huggingface://%s/%s
sha256: %s
`, name, filename, filename, repo, filename, sha)
}
func indexOf(entries ...string) *Index {
ix, err := ParseIndex(strings.Join(entries, ""))
ExpectWithOffset(1, err).ToNot(HaveOccurred())
return ix
}
// familyNames flattens a result into "parent <- variant, variant" strings, the
// form the specs assert against.
func familyNames(r *Result) []string {
out := make([]string, 0, len(r.Families))
for _, f := range r.Families {
names := make([]string, 0, len(f.Proposals))
for _, p := range f.Proposals {
names = append(names, p.Variant)
}
out = append(out, f.Parent+" <- "+strings.Join(names, ", "))
}
return out
}
func refusalReasons(r *Result) string {
var b strings.Builder
for _, ref := range r.Refusals {
b.WriteString(strings.Join(ref.Members, " + ") + ": " + ref.Reason + "\n")
}
return b.String()
}
func suppressionReasons(r *Result) string {
var b strings.Builder
for _, s := range r.Suppressed {
b.WriteString(s.String() + "\n")
}
return b.String()
}
var _ = Describe("Propose", func() {
Describe("the grouping signals", func() {
It("groups entries whose names differ only by a quantization marker", func() {
ix := indexOf(
entryYAML("foo-model", "acme/foo-GGUF", "foo-model-Q4_K_M.gguf", "aa"),
entryYAML("foo-model-q8_0", "acme/foo-GGUF", "foo-model-Q8_0.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(ConsistOf("foo-model <- foo-model-q8_0"))
Expect(r.Families[0].Proposals[0].Evidence.Signals).To(ContainElement(SignalName))
Expect(r.Families[0].Proposals[0].Evidence.SharedStem).To(Equal("foo-model"))
Expect(r.Families[0].Proposals[0].Evidence.QuantTokens).To(ContainElements("q4_k_m", "q8_0"))
})
It("groups entries that use the colon config-suffix convention", func() {
ix := indexOf(
entryYAML("bar-model", "acme/bar-GGUF", "bar-model-Q4_K_M.gguf", "aa"),
entryYAML("bar-model:grammar-functioncall", "acme/bar-GGUF", "bar-model-Q4_K_M-grammar.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(ConsistOf("bar-model <- bar-model:grammar-functioncall"))
Expect(r.Families[0].Proposals[0].Evidence.Signals).To(ContainElement(SignalConfigSuffix))
})
It("groups entries whose own weight file is the same file at another quantization", func() {
// The names share no stem, so only the filename signal can link
// these two.
ix := indexOf(
entryYAML("omni-cpp", "Serveurperso/Omni-GGUF", "omnivoice-base-Q8_0.gguf", "aa"),
entryYAML("omni-cpp-hq", "Serveurperso/Omni-GGUF", "omnivoice-base-BF16.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(ConsistOf("omni-cpp <- omni-cpp-hq"))
ev := r.Families[0].Proposals[0].Evidence
Expect(ev.Signals).To(ConsistOf(SignalWeightFile))
Expect(ev.SharedFile).To(Equal("omnivoice-base"))
Expect(ev.SharedRepo).To(Equal("serveurperso/omni-gguf"))
})
It("does not let a shared auxiliary file link unrelated models", func() {
// Both entries ship the same text encoder. That is a packaging
// convention, not evidence of shared weights: this is how an
// earlier sweep linked four wan-2.1 entries to each other.
ix := indexOf(`- name: wan-2.1-t2v
url: u
files:
- filename: wan-2.1-t2v-Q4_K_M.gguf
uri: huggingface://acme/wan/wan-2.1-t2v-Q4_K_M.gguf
sha256: aa
- filename: umt5-xxl-encoder-Q8_0.gguf
uri: huggingface://acme/wan/umt5-xxl-encoder-Q8_0.gguf
sha256: cc
`, `- name: z-image-turbo
url: u
files:
- filename: z-image-turbo-Q4_K_M.gguf
uri: huggingface://acme/wan/z-image-turbo-Q4_K_M.gguf
sha256: bb
- filename: umt5-xxl-encoder-Q8_0.gguf
uri: huggingface://acme/wan/umt5-xxl-encoder-Q8_0.gguf
sha256: cc
`)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
})
It("does not treat a shared filename in two different repos as evidence", func() {
// A finetune republished under the base model's filename is the
// most common way this signal misfires.
ix := indexOf(
entryYAML("llama-3.2-3b-instruct", "hugging-quants/Llama-3.2-3B-Instruct-GGUF", "llama-3.2-3b-instruct-q4_k_m.gguf", "aa"),
entryYAML("llama-3.2-3b-shiro-roleplay", "someone/Shiro-GGUF", "Llama-3.2-3B-Instruct.Q8_0.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
})
})
Describe("what must never be proposed", func() {
It("does not group different parameter sizes that share a prefix", func() {
ix := indexOf(
entryYAML("qwen3-tts-cpp-0.6b-base", "Serveurperso/Qwen3-TTS-GGUF", "qwen3-tts-talker-Q4_K_M.gguf", "aa"),
entryYAML("qwen3-tts-cpp-1.7b-base", "Serveurperso/Qwen3-TTS-GGUF", "qwen3-tts-talker-Q8_0.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
Expect(suppressionReasons(r)).To(ContainSubstring("different parameter sizes"))
})
It("does not group the Gemma effective sizes", func() {
ix := indexOf(
entryYAML("gemma-4-e2b-it", "google/gemma-GGUF", "gemma-4-it-Q4_K_M.gguf", "aa"),
entryYAML("gemma-4-e4b-it", "google/gemma-GGUF", "gemma-4-it-Q8_0.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
Expect(suppressionReasons(r)).To(ContainSubstring("different parameter sizes"))
})
It("does not group entries with a byte-identical install payload", func() {
// whisper-1 exists so OpenAI-compatible clients can send that name.
// Folding it under whisper-base would hide the name they send.
payload := ` url: github:mudler/LocalAI/gallery/whisper-base.yaml@master
overrides:
parameters:
model: ggml-whisper-base.bin
files:
- filename: ggml-whisper-base.bin
uri: huggingface://ggerganov/whisper.cpp/ggml-base.bin
sha256: aa
`
ix := indexOf("- name: whisper-base\n"+payload, "- name: whisper-1\n"+payload)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
Expect(r.AliasSkipped).To(HaveLen(1))
Expect(r.AliasSkipped[0].Reason).To(ContainSubstring("aliases"))
})
DescribeTable("declines the categories the ledger records",
func(nameA, nameB string, ledgerYAML string) {
ix := indexOf(
entryYAML(nameA, "acme/repo", "shared-weights-Q4_K_M.gguf", "aa"),
entryYAML(nameB, "acme/repo", "shared-weights-Q8_0.gguf", "bb"),
)
ledger, err := ParseLedger([]byte(ledgerYAML))
Expect(err).ToNot(HaveOccurred())
// Without the ledger these would be proposed, which is what
// makes the ledger load bearing rather than decorative.
Expect(familyNames(Propose(ix, nil))).ToNot(BeEmpty())
r := Propose(ix, ledger)
Expect(familyNames(r)).To(BeEmpty())
Expect(r.Suppressed).To(HaveLen(1))
},
Entry("a finetune", "base-model", "base-model-abliterated",
"tokens:\n - {token: abliterated, reason: finetune}\n"),
Entry("a distill", "base-model", "base-model-distilled",
"tokens:\n - {token: distilled, reason: distilled}\n"),
Entry("English-only versus multilingual ASR", "whisper-small", "whisper-small-en",
"pairs:\n - {parent: whisper-small, variant: whisper-small-en, reason: English-only versus multilingual}\n"),
Entry("two products sharing a prefix", "vibevoice-cpp", "vibevoice-cpp-asr",
"pairs:\n - {parent: vibevoice-cpp, variant: vibevoice-cpp-asr, reason: different products}\n"),
Entry("a per-language release", "kokoros-de", "kokoros-ja",
"groups:\n - {members: [kokoros, kokoros-de, kokoros-ja], reason: different languages}\n"),
)
It("reports the ledger's reason so its effect stays visible", func() {
ix := indexOf(
entryYAML("base-model", "acme/repo", "shared-weights-Q4_K_M.gguf", "aa"),
entryYAML("base-model-heretic", "acme/repo", "shared-weights-Q8_0.gguf", "bb"),
)
ledger, err := ParseLedger([]byte("tokens:\n - {token: heretic, reason: \"finetune, not a re-quantization\"}\n"))
Expect(err).ToNot(HaveOccurred())
r := Propose(ix, ledger)
Expect(suppressionReasons(r)).To(ContainSubstring("finetune, not a re-quantization"))
Expect(suppressionReasons(r)).To(ContainSubstring(`token "heretic"`))
})
})
Describe("parent selection", func() {
It("picks the bare-named entry when one exists", func() {
ix := indexOf(
entryYAML("base-model-q8_0", "acme/repo", "base-model-Q8_0.gguf", "aa"),
entryYAML("base-model", "acme/repo", "base-model-Q4_K_M.gguf", "bb"),
entryYAML("base-model-f16", "acme/repo", "base-model-f16.gguf", "cc"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(ConsistOf("base-model <- base-model-f16, base-model-q8_0"))
})
It("picks the smallest build when no entry is bare-named", func() {
ix := indexOf(
entryYAML("ced-base-f16", "acme/repo", "ced-base-f16.gguf", "aa"),
entryYAML("ced-base-q8", "acme/repo", "ced-base-Q8_0.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(ConsistOf("ced-base-q8 <- ced-base-f16"))
})
It("judges the smallest build by the quantization in the model filename, not the name", func() {
// The names carry no marker at all; only the filenames say which
// build is which.
ix := indexOf(
entryYAML("thing-hq", "acme/repo", "thing-weights-BF16.gguf", "aa"),
entryYAML("thing-lite", "acme/repo", "thing-weights-Q4_K_M.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(ConsistOf("thing-lite <- thing-hq"))
})
})
Describe("the rules a proposal has to respect", func() {
It("refuses to nest: a target that already offers variants of its own", func() {
ix := indexOf(
entryYAML("nest-model", "acme/repo", "nest-model-Q4_K_M.gguf", "aa"),
`- name: nest-model-q8_0
url: u
variants:
- model: nest-model-q8_0-mtp
overrides:
parameters:
model: nest-model-Q8_0.gguf
files:
- filename: nest-model-Q8_0.gguf
uri: huggingface://acme/repo/nest-model-Q8_0.gguf
sha256: bb
`,
entryYAML("nest-model-q8_0-mtp", "other/repo", "nest-model-mtp.gguf", "cc"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
Expect(refusalReasons(r)).To(ContainSubstring("already offers variants of its own"))
})
It("refuses to nest: a parent that is already somebody else's variant", func() {
ix := indexOf(
`- name: outer
url: u
variants:
- model: middle
overrides:
parameters:
model: outer-Q4_K_M.gguf
files:
- filename: outer-Q4_K_M.gguf
uri: huggingface://acme/repo/outer-Q4_K_M.gguf
sha256: aa
`,
entryYAML("middle", "acme/other", "middle-Q4_K_M.gguf", "bb"),
entryYAML("middle-q8_0", "acme/other", "middle-Q8_0.gguf", "cc"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
Expect(refusalReasons(r)).To(ContainSubstring("would nest variants"))
})
It("refuses to let two parents claim one target", func() {
ix := indexOf(
`- name: claimant
url: u
variants:
- model: contested-q8_0
overrides:
parameters:
model: claimant-Q4_K_M.gguf
files:
- filename: claimant-Q4_K_M.gguf
uri: huggingface://acme/repo/claimant-Q4_K_M.gguf
sha256: aa
`,
entryYAML("contested", "acme/other", "contested-Q4_K_M.gguf", "bb"),
entryYAML("contested-q8_0", "acme/other", "contested-Q8_0.gguf", "cc"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
Expect(refusalReasons(r)).To(ContainSubstring("already a variant of"))
})
It("refuses a target that is not independently installable", func() {
ix := indexOf(
entryYAML("stub-model", "acme/repo", "stub-model-Q4_K_M.gguf", "aa"),
"- name: stub-model-q8_0\n description: a stanza nobody finished\n",
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
Expect(refusalReasons(r)).To(ContainSubstring("not independently installable"))
})
It("refuses a family whose parent defines a merge anchor, naming the entries that would inherit", func() {
ix := indexOf(
`- &anchored
name: anchored-model
url: u
overrides:
parameters:
model: anchored-Q4_K_M.gguf
files:
- filename: anchored-Q4_K_M.gguf
uri: huggingface://acme/repo/anchored-Q4_K_M.gguf
sha256: aa
`,
`- !!merge <<: *anchored
name: anchored-child
variants: []
overrides:
parameters:
model: unrelated-child-Q4_K_M.gguf
files:
- filename: unrelated-child-Q4_K_M.gguf
uri: huggingface://other/repo/unrelated-child-Q4_K_M.gguf
sha256: cc
`,
entryYAML("anchored-model-q8_0", "acme/repo", "anchored-Q8_0.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
Expect(refusalReasons(r)).To(ContainSubstring("defines YAML anchor &anchored"))
Expect(refusalReasons(r)).To(ContainSubstring("anchored-child"))
Expect(refusalReasons(r)).To(ContainSubstring("variants: []"))
})
It("refuses an entry whose name is not unique in the gallery", func() {
ix := indexOf(
entryYAML("twin", "acme/repo", "twin-Q4_K_M.gguf", "aa"),
entryYAML("twin", "acme/repo", "twin-Q4_K_M.gguf", "aa"),
entryYAML("twin-q8_0", "acme/repo", "twin-Q8_0.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(BeEmpty())
Expect(refusalReasons(r)).To(ContainSubstring("appears more than once"))
})
It("says nothing about a pair that is already grouped", func() {
ix := indexOf(
`- name: settled
url: u
variants:
- model: settled-q8_0
overrides:
parameters:
model: settled-Q4_K_M.gguf
files:
- filename: settled-Q4_K_M.gguf
uri: huggingface://acme/repo/settled-Q4_K_M.gguf
sha256: aa
`,
entryYAML("settled-q8_0", "acme/repo", "settled-Q8_0.gguf", "bb"),
)
r := Propose(ix, nil)
Expect(r.HasProposals()).To(BeFalse())
Expect(r.Refusals).To(BeEmpty())
Expect(r.Suppressed).To(BeEmpty())
})
It("adds only the missing members to a family that already exists", func() {
ix := indexOf(
`- name: partial
url: u
variants:
- model: partial-q8_0
overrides:
parameters:
model: partial-Q4_K_M.gguf
files:
- filename: partial-Q4_K_M.gguf
uri: huggingface://acme/repo/partial-Q4_K_M.gguf
sha256: aa
`,
entryYAML("partial-q8_0", "acme/repo", "partial-Q8_0.gguf", "bb"),
entryYAML("partial-f16", "acme/repo", "partial-f16.gguf", "cc"),
)
r := Propose(ix, nil)
Expect(familyNames(r)).To(ConsistOf("partial <- partial-f16"))
})
})
It("does not modify the index it was given", func() {
text := entryYAML("foo-model", "acme/foo-GGUF", "foo-model-Q4_K_M.gguf", "aa") +
entryYAML("foo-model-q8_0", "acme/foo-GGUF", "foo-model-Q8_0.gguf", "bb")
ix, err := ParseIndex(text)
Expect(err).ToNot(HaveOccurred())
before := strings.Join(ix.Lines, "\n")
Propose(ix, nil)
Expect(strings.Join(ix.Lines, "\n")).To(Equal(before))
})
})

151
.github/ci/variantproposals/quant.go vendored Normal file
View File

@@ -0,0 +1,151 @@
package main
import (
"regexp"
"strconv"
"strings"
)
// Quantization and precision markers that distinguish one build of a set of
// weights from another build of the same weights. Stripping them from a name
// is what lets the proposer notice that two entries are the same model.
//
// qat and apex are in this list on a maintainer ruling: they are quantization
// techniques applied to published weights, not separate weights. Names that use
// "apex" to mean a finetune are handled by the rejection ledger instead, because
// no amount of pattern matching can tell the two uses apart.
const quantAlternation = `q[2-8](?:_[0-9a-z]+)*|pq[2-8](?:_[0-9a-z]+)*|iq[1-9][0-9a-z]*(?:_[0-9a-z]+)*|i1|` +
`f16|f32|bf16|fp16|fp32|fp8|fp4|nvfp4|mxfp4(?:_moe)*|awq|gptq|qat|apex|gguf|ggml|[0-9]+bit|g[0-9]+`
// quantSegment matches a whole hyphen-delimited segment of an entry name.
// Names separate their parts with "-" and keep quantization tokens internally
// joined with "_", so a segment is the right unit here: "q4_k_m" arrives whole.
var quantSegment = regexp.MustCompile(`^(?:` + quantAlternation + `)$`)
// quantFileSuffix matches a trailing quantization token in a weight filename.
// Filenames mix "-", "_" and "." as separators, so unlike entry names they
// cannot be split into segments up front without tearing "Q4_K_M" apart.
var quantFileSuffix = regexp.MustCompile(`(?i)[-_.](?:` + quantAlternation + `)$`)
var weightExtension = regexp.MustCompile(`(?i)\.(gguf|ggml|safetensors|bin|pt|pth|onnx)$`)
// IsQuantToken reports whether a single name segment is a quantization or
// precision marker rather than part of the model's identity.
func IsQuantToken(segment string) bool {
return quantSegment.MatchString(strings.ToLower(segment))
}
// NameStem reduces an entry name to the identity it shares with its alternative
// builds: the config suffix after ":" is dropped, then trailing quantization
// segments are stripped.
//
// It implements the first two grouping signals together because they answer the
// same question. "foo:q8_0" and "foo-q8_0" are both alternative builds of "foo",
// and the caller that needs to report which convention was used can compare the
// name against the stem itself.
//
// At least one segment always survives, so a name made entirely of quantization
// tokens does not collapse to the empty stem and swallow every other such name.
func NameStem(name string) string {
base := strings.ToLower(strings.TrimSpace(name))
if i := strings.Index(base, ":"); i >= 0 {
base = base[:i]
}
segments := strings.Split(base, "-")
for len(segments) > 1 && quantSegment.MatchString(segments[len(segments)-1]) {
segments = segments[:len(segments)-1]
}
return strings.Join(segments, "-")
}
// HasConfigSuffix reports whether a name uses the ":" convention for naming a
// config variant of another entry.
func HasConfigSuffix(name string) bool {
return strings.Contains(name, ":")
}
// FileStem reduces a weight filename to the identity shared by its other
// quantizations: directories, extension and trailing quantization tokens go.
//
// This is the third grouping signal. It is the one that has misfired before, so
// callers must filter auxiliary files out before handing a filename here: a
// shared text encoder is not evidence of shared weights.
func FileStem(filename string) string {
base := filename
if i := strings.LastIndex(base, "/"); i >= 0 {
base = base[i+1:]
}
base = weightExtension.ReplaceAllString(base, "")
for {
stripped := quantFileSuffix.ReplaceAllString(base, "")
if stripped == base {
break
}
base = stripped
}
return strings.ToLower(base)
}
// bitsPerWeight ranks quantization tokens so the smallest build of a family can
// be identified when no bare-named entry exists to be the parent.
//
// The figures are nominal bits per weight, not measured file sizes. Ranking is
// all that is asked of them, and a nominal figure is available from the name
// alone without downloading anything.
func bitsPerWeight(token string) (int, bool) {
t := strings.ToLower(token)
switch {
case t == "i1":
return 1, true
case strings.HasPrefix(t, "nvfp4"), strings.HasPrefix(t, "mxfp4"), t == "fp4":
return 4, true
case t == "fp8":
return 8, true
case t == "f16", t == "bf16", t == "fp16":
return 16, true
case t == "f32", t == "fp32":
return 32, true
case t == "awq", t == "gptq":
return 4, true
}
if m := regexp.MustCompile(`^p?q([1-9])`).FindStringSubmatch(t); m != nil {
n, _ := strconv.Atoi(m[1])
return n, true
}
if m := regexp.MustCompile(`^iq([1-9])`).FindStringSubmatch(t); m != nil {
n, _ := strconv.Atoi(m[1])
return n, true
}
if m := regexp.MustCompile(`^([0-9]+)bit$`).FindStringSubmatch(t); m != nil {
n, _ := strconv.Atoi(m[1])
return n, true
}
return 0, false
}
// unknownWidth sorts after every recognised quantization so an entry whose
// build cannot be read from its filename never wins the "smallest build" tie
// break by accident.
const unknownWidth = 1 << 10
// BuildWidth reports the nominal bits per weight of the build a filename holds.
// An unreadable filename gets unknownWidth.
func BuildWidth(filename string) int {
base := filename
if i := strings.LastIndex(base, "/"); i >= 0 {
base = base[i+1:]
}
base = weightExtension.ReplaceAllString(base, "")
best := unknownWidth
for {
m := quantFileSuffix.FindString(base)
if m == "" {
break
}
if bits, ok := bitsPerWeight(m[1:]); ok && bits < best {
best = bits
}
base = base[:len(base)-len(m)]
}
return best
}

View File

@@ -0,0 +1,92 @@
package main
import (
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("quantization markers", func() {
DescribeTable("NameStem strips the markers that distinguish builds, not models",
func(name, expected string) {
Expect(NameStem(name)).To(Equal(expected))
},
Entry("plain q4", "foo-model-q4_k_m", "foo-model"),
Entry("q8_0", "foo-model-q8_0", "foo-model"),
Entry("q5_1", "foo-model-q5_1", "foo-model"),
Entry("q2 with group size", "ternary-bonsai-8b-q2-g64", "ternary-bonsai-8b"),
Entry("iq variant", "ideogram-4-iq4nl-ggml", "ideogram-4"),
Entry("i1 imatrix", "orca-agent-v0.1-i1", "orca-agent-v0.1"),
Entry("f16", "ced-base-f16", "ced-base"),
Entry("bf16", "some-model-bf16", "some-model"),
Entry("fp8", "some-model-fp8", "some-model"),
Entry("nvfp4", "qwen3.6-27b-nvfp4", "qwen3.6-27b"),
Entry("mxfp4_moe", "huihui-qwen3-vl-30b-a3b-instruct-abliterated-mxfp4_moe", "huihui-qwen3-vl-30b-a3b-instruct-abliterated"),
Entry("pq2", "ternary-bonsai-8b-pq2", "ternary-bonsai-8b"),
Entry("awq", "some-model-awq", "some-model"),
Entry("gptq", "some-model-gptq", "some-model"),
Entry("Nbit", "qwen3-8b-mlx-4bit", "qwen3-8b-mlx"),
Entry("gguf", "some-model-gguf", "some-model"),
Entry("ggml", "flux.1-dev-ggml", "flux.1-dev"),
Entry("qat is a quantization technique", "gemma-3-27b-it-qat", "gemma-3-27b-it"),
Entry("apex is a quantization technique", "qwen3.6-35b-a3b-apex", "qwen3.6-35b-a3b"),
Entry("stacked markers", "gemma-4-e2b-it-qat-q4_0", "gemma-4-e2b-it"),
Entry("the config suffix is dropped", "phi-2-chat:Q8_0", "phi-2-chat"),
Entry("a non-quant config suffix is dropped too", "meta-llama-3.1-8b-instruct:grammar-functioncall", "meta-llama-3.1-8b-instruct"),
)
DescribeTable("NameStem leaves alone what identifies a different model",
func(name, expected string) {
Expect(NameStem(name)).To(Equal(expected))
},
Entry("parameter size", "qwen3-tts-cpp-0.6b-base", "qwen3-tts-cpp-0.6b-base"),
Entry("language suffix", "kokoros-de", "kokoros-de"),
Entry("English-only ASR", "whisper-small-en", "whisper-small-en"),
Entry("finetune", "qwen3-30b-a3b-abliterated", "qwen3-30b-a3b-abliterated"),
Entry("product suffix", "vibevoice-cpp-asr", "vibevoice-cpp-asr"),
)
It("never strips a name down to nothing", func() {
Expect(NameStem("q4_k_m")).To(Equal("q4_k_m"))
Expect(NameStem("f16-q8_0")).To(Equal("f16"))
})
DescribeTable("FileStem reduces a weight filename to the weights it holds",
func(filename, expected string) {
Expect(FileStem(filename)).To(Equal(expected))
},
Entry("directory and extension go", "bonsai/models/Ternary-Bonsai-8B-gguf/Ternary-Bonsai-8B-Q2_0.gguf", "ternary-bonsai-8b"),
Entry("underscored quant token stays whole", "Llama-3.2-1B-Instruct-Q4_K_M.gguf", "llama-3.2-1b-instruct"),
Entry("dot separated quant token", "Llama-3.2-3B-Instruct.Q4_K_M.gguf", "llama-3.2-3b-instruct"),
Entry("group size suffix", "Ternary-Bonsai-8B-Q2_0_g64.gguf", "ternary-bonsai-8b"),
Entry("bf16", "omnivoice-cpp-hq/omnivoice-base-BF16.gguf", "omnivoice-base"),
Entry("safetensors", "some/dir/Model-Name-fp8.safetensors", "model-name"),
)
DescribeTable("BuildWidth reads the nominal width out of a filename",
func(filename string, expected int) {
Expect(BuildWidth(filename)).To(Equal(expected))
},
Entry("q4", "foo-Q4_K_M.gguf", 4),
Entry("q8", "foo-Q8_0.gguf", 8),
Entry("q2", "foo-Q2_0.gguf", 2),
Entry("f16", "foo-f16.gguf", 16),
Entry("bf16", "foo-BF16.gguf", 16),
Entry("iq3", "foo-iq3_xxs.gguf", 3),
Entry("nothing readable sorts last", "foo.gguf", unknownWidth),
)
It("treats an auxiliary file as never being the model's own weights", func() {
for _, f := range []string{
"mmproj-model-f16.gguf",
"dir/vae-BF16.gguf",
"clip_l.safetensors",
"umt5-xxl-encoder-Q8_0.gguf",
"t5xxl_fp16.safetensors",
"ae.safetensors",
"omnivoice-tokenizer-Q8_0.gguf",
} {
Expect(IsAuxiliaryFile(f)).To(BeTrue(), "expected %q to be auxiliary", f)
}
Expect(IsAuxiliaryFile("gemma-3-27b-it-Q4_K_M.gguf")).To(BeFalse())
})
})

View File

@@ -0,0 +1,13 @@
package main
import (
"testing"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
func TestVariantProposals(t *testing.T) {
RegisterFailHandler(Fail)
RunSpecs(t, "gallery variant proposals")
}

View File

@@ -45,6 +45,16 @@ updates:
directory: "/backend/python/diffusers"
schedule:
interval: "weekly"
# torch and transformers are deliberately pinned in this backend (see
# backend/python/diffusers/requirements-*.txt and issue #9979), and the
# l4t12 variant resolves them from the Jetson pip index
# (https://pypi.jetson-ai-lab.io/jp6/cu129/). dependabot cannot authenticate
# against that index and fails the whole weekly update with a
# private_source_authentication_failure. Ignore the two pinned deps we don't
# want bumped anyway so the job stays green.
ignore:
- dependency-name: "torch"
- dependency-name: "transformers"
- package-ecosystem: "pip"
directory: "/backend/python/exllama"
schedule:

39
.github/gh_curl.sh vendored Executable file
View File

@@ -0,0 +1,39 @@
#!/bin/bash
# Shared curl wrapper for the nightly dependency-bump scripts.
#
# The bump workflow fans out to ~25 parallel matrix jobs, each querying
# api.github.com. Anonymous API calls are capped at 60/hour per source IP and
# GitHub-hosted runners egress through shared NAT addresses, so a random handful
# of jobs were getting rate-limited (HTTP 403 -> curl exit 22, empty response)
# every single night. Authenticating with GITHUB_TOKEN lifts the ceiling to
# 1000/hour; the retries absorb whatever transient blips remain.
# Wraps curl with GitHub auth (when a token is present) plus retry/timeout
# hardening. Callers pass their own headers and the URL.
gh_curl() {
# The bump scripts run under `set -x`; without this the Authorization header
# would be echoed into the job log on every call.
local had_xtrace=0
case "$-" in
*x*) had_xtrace=1; set +x ;;
esac
local args=(
--silent --show-error --location --fail
# --retry-all-errors so 403 rate-limit responses are retried too; plain
# --retry only covers 408/429/5xx. curl honours Retry-After when sent.
--retry 5 --retry-delay 3 --retry-all-errors
--connect-timeout 15 --max-time 60
)
if [ -n "${GITHUB_TOKEN:-}" ]; then
args+=(--header "Authorization: Bearer ${GITHUB_TOKEN}")
fi
curl "${args[@]}" "$@"
local rc=$?
if [ "$had_xtrace" -eq 1 ]; then
set -x
fi
return $rc
}

View File

@@ -1,77 +0,0 @@
#!/usr/bin/env bash
#
# paged-canary-apply.sh - apply the vendored paged-attention patch series
# (backend/cpp/llama-cpp-localai-paged/patches/paged/0001-0030) to a llama.cpp checkout, the
# same way the build does, but tolerating the ONE known-benign pre-existing
# quirk in the series. Used by the early-warning canary
# (.github/workflows/llama-cpp-paged-canary.yml) so it only goes red on a REAL
# upstream break, never on that quirk.
#
# Usage: paged-canary-apply.sh <llama.cpp-checkout-dir> <patches-dir>
# <patches-dir> is normally backend/cpp/llama-cpp-localai-paged/patches (it holds the
# top-level base series 0*.patch, currently empty, and the paged/ subseries).
#
# Exit 0 = the whole series applied -> patches still fit upstream.
# Exit !=0 = a patch failed to apply = the red signal: an upstream change moved
# the tree out from under the patches, so it is time to run a PIN_SYNC.
#
# Apply method MIRRORS backend/cpp/llama-cpp/Makefile's `llama.cpp` target:
# plain `git apply --verbose`, which natively tolerates @@ line-number offsets
# but NOT context-line changes. Matching the build's method is the point - the
# canary's apply result is exactly what the real build's apply would do.
#
# The ONLY tolerance, and it is path-scoped (not a blanket `|| true`): patch
# 0019 carries a stray *modify* hunk against the dev-only doc
# SSM_DECODE_FIX_RESULTS.md, a file that exists only on the DGX dev tree and is
# absent from any clean upstream checkout. `git apply` is atomic, so that single
# missing-file hunk rejects the whole patch - and because 0021/0022/0026/0028
# build on 0019's code, the rejection cascades to them too. This is a
# PRE-EXISTING shipped-series defect, present identically on every pin, NOT an
# upstream break (see backend/cpp/llama-cpp-localai-paged/README.md section 7,
# "Pin + maintenance policy"). We exclude ONLY that dev-doc path and still
# apply 0019's real code hunks atomically, so a genuine code-hunk break in 0019
# still fails the canary. prepare.sh tolerates the same hunk via
# `patch ... || true`; this mirrors that tolerance precisely.
set -euo pipefail
CHECKOUT="${1:?usage: paged-canary-apply.sh <llama.cpp-checkout> <patches-dir>}"
PATCHES="${2:?usage: paged-canary-apply.sh <llama.cpp-checkout> <patches-dir>}"
# The lone tolerated dev-doc, and the only patch allowed to carry it.
DEVDOC_GLOB='*SSM_DECODE_FIX_RESULTS.md'
DEVDOC_PATCH='0019-qwen35-ssm-decode-fused-gather.patch'
# Resolve to absolute paths so the apply works after we cd into the checkout.
PATCHES="$(cd "$PATCHES" && pwd)"
cd "$CHECKOUT"
shopt -s nullglob
apply_one() {
local p="$1"; shift
echo "paged-canary: applying $(basename "$p")"
if ! git apply --verbose "$@" "$p"; then
echo "::error::paged patch no longer applies to the upstream llama.cpp tip: $(basename "$p")"
echo "::error::upstream drifted past the vendored paged series - run a PIN_SYNC (see backend/cpp/llama-cpp-localai-paged/README.md section 7, Pin + maintenance policy), do NOT bump the pin blindly"
exit 1
fi
}
# Base series first (parity with the build: patches/0*.patch before
# patches/paged/0*.patch). Currently empty; nullglob makes this a no-op.
for p in "$PATCHES"/0*.patch; do
apply_one "$p"
done
# Paged series, in order.
for p in "$PATCHES"/paged/0*.patch; do
if [ "$(basename "$p")" = "$DEVDOC_PATCH" ]; then
# Apply 0019's real code hunks; exclude ONLY the benign dev-doc hunk.
apply_one "$p" --exclude="$DEVDOC_GLOB"
else
apply_one "$p"
fi
done
echo "paged-canary: the full paged patch series applied cleanly to the upstream tip"

View File

@@ -32,16 +32,30 @@ jobs:
if: github.repository == 'mudler/LocalAI'
runs-on: ubuntu-latest
outputs:
matrix-singlearch: ${{ steps.set-matrix.outputs['matrix-singlearch'] }}
matrix-multiarch: ${{ steps.set-matrix.outputs['matrix-multiarch'] }}
matrix-darwin: ${{ steps.set-matrix.outputs['matrix-darwin'] }}
merge-matrix-multiarch: ${{ steps.set-matrix.outputs['merge-matrix-multiarch'] }}
merge-matrix-singlearch: ${{ steps.set-matrix.outputs['merge-matrix-singlearch'] }}
has-backends-singlearch: ${{ steps.set-matrix.outputs['has-backends-singlearch'] }}
has-backends-multiarch: ${{ steps.set-matrix.outputs['has-backends-multiarch'] }}
has-backends-darwin: ${{ steps.set-matrix.outputs['has-backends-darwin'] }}
has-merges-multiarch: ${{ steps.set-matrix.outputs['has-merges-multiarch'] }}
has-merges-singlearch: ${{ steps.set-matrix.outputs['has-merges-singlearch'] }}
# Single-arch backends are sharded across SINGLEARCH_SHARDS matrix jobs to
# stay under GitHub's 256-jobs-per-matrix limit (see changed-backends.js).
matrix-singlearch-1: ${{ steps.set-matrix.outputs['matrix-singlearch-1'] }}
merge-matrix-singlearch-1: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-1'] }}
has-backends-singlearch-1: ${{ steps.set-matrix.outputs['has-backends-singlearch-1'] }}
has-merges-singlearch-1: ${{ steps.set-matrix.outputs['has-merges-singlearch-1'] }}
matrix-singlearch-2: ${{ steps.set-matrix.outputs['matrix-singlearch-2'] }}
merge-matrix-singlearch-2: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-2'] }}
has-backends-singlearch-2: ${{ steps.set-matrix.outputs['has-backends-singlearch-2'] }}
has-merges-singlearch-2: ${{ steps.set-matrix.outputs['has-merges-singlearch-2'] }}
matrix-singlearch-3: ${{ steps.set-matrix.outputs['matrix-singlearch-3'] }}
merge-matrix-singlearch-3: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-3'] }}
has-backends-singlearch-3: ${{ steps.set-matrix.outputs['has-backends-singlearch-3'] }}
has-merges-singlearch-3: ${{ steps.set-matrix.outputs['has-merges-singlearch-3'] }}
matrix-singlearch-4: ${{ steps.set-matrix.outputs['matrix-singlearch-4'] }}
merge-matrix-singlearch-4: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-4'] }}
has-backends-singlearch-4: ${{ steps.set-matrix.outputs['has-backends-singlearch-4'] }}
has-merges-singlearch-4: ${{ steps.set-matrix.outputs['has-merges-singlearch-4'] }}
steps:
- name: Checkout repository
uses: actions/checkout@v7
@@ -109,9 +123,9 @@ jobs:
# take their full ~6h cold without blocking manifest assembly for the
# multi-arch backends whose per-arch digests would otherwise sit untagged
# on quay long enough to be GC'd.
backend-jobs-singlearch:
backend-jobs-singlearch-1:
needs: generate-matrix
if: needs.generate-matrix.outputs['has-backends-singlearch'] == 'true'
if: needs.generate-matrix.outputs['has-backends-singlearch-1'] == 'true'
uses: ./.github/workflows/backend_build.yml
with:
tag-latest: ${{ matrix.tag-latest }}
@@ -138,7 +152,100 @@ jobs:
strategy:
fail-fast: false
max-parallel: 8
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch']) }}
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-1']) }}
backend-jobs-singlearch-2:
needs: generate-matrix
if: needs.generate-matrix.outputs['has-backends-singlearch-2'] == 'true'
uses: ./.github/workflows/backend_build.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
build-type: ${{ matrix.build-type }}
cuda-major-version: ${{ matrix.cuda-major-version }}
cuda-minor-version: ${{ matrix.cuda-minor-version }}
platforms: ${{ matrix.platforms }}
platform-tag: ${{ matrix.platform-tag || '' }}
runs-on: ${{ matrix.runs-on }}
builder-base-image: ${{ matrix.builder-base-image || '' }}
base-image: ${{ matrix.base-image }}
backend: ${{ matrix.backend }}
dockerfile: ${{ matrix.dockerfile }}
skip-drivers: ${{ matrix.skip-drivers }}
context: ${{ matrix.context }}
ubuntu-version: ${{ matrix.ubuntu-version }}
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
secrets:
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
max-parallel: 8
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-2']) }}
backend-jobs-singlearch-3:
needs: generate-matrix
if: needs.generate-matrix.outputs['has-backends-singlearch-3'] == 'true'
uses: ./.github/workflows/backend_build.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
build-type: ${{ matrix.build-type }}
cuda-major-version: ${{ matrix.cuda-major-version }}
cuda-minor-version: ${{ matrix.cuda-minor-version }}
platforms: ${{ matrix.platforms }}
platform-tag: ${{ matrix.platform-tag || '' }}
runs-on: ${{ matrix.runs-on }}
builder-base-image: ${{ matrix.builder-base-image || '' }}
base-image: ${{ matrix.base-image }}
backend: ${{ matrix.backend }}
dockerfile: ${{ matrix.dockerfile }}
skip-drivers: ${{ matrix.skip-drivers }}
context: ${{ matrix.context }}
ubuntu-version: ${{ matrix.ubuntu-version }}
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
secrets:
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
max-parallel: 8
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-3']) }}
backend-jobs-singlearch-4:
needs: generate-matrix
if: needs.generate-matrix.outputs['has-backends-singlearch-4'] == 'true'
uses: ./.github/workflows/backend_build.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
build-type: ${{ matrix.build-type }}
cuda-major-version: ${{ matrix.cuda-major-version }}
cuda-minor-version: ${{ matrix.cuda-minor-version }}
platforms: ${{ matrix.platforms }}
platform-tag: ${{ matrix.platform-tag || '' }}
runs-on: ${{ matrix.runs-on }}
builder-base-image: ${{ matrix.builder-base-image || '' }}
base-image: ${{ matrix.base-image }}
backend: ${{ matrix.backend }}
dockerfile: ${{ matrix.dockerfile }}
skip-drivers: ${{ matrix.skip-drivers }}
context: ${{ matrix.context }}
ubuntu-version: ${{ matrix.ubuntu-version }}
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
secrets:
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
max-parallel: 8
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-4']) }}
# Apply tags to per-arch digests via `imagetools create`. Split into two
# jobs that mirror the build split so each merge waits ONLY on its
@@ -174,10 +281,12 @@ jobs:
fail-fast: false
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-multiarch']) }}
backend-merge-jobs-singlearch:
needs: [generate-matrix, backend-jobs-singlearch]
# See note on backend-merge-jobs-multiarch above for !cancelled().
if: ${{ !cancelled() && needs.generate-matrix.outputs['has-merges-singlearch'] == 'true' }}
# One merge shard per build shard: backend-merge-jobs-singlearch-<n> needs only
# backend-jobs-singlearch-<n>, preserving the "merge waits only on its own
# build" property while staying under the 256-jobs-per-matrix limit.
backend-merge-jobs-singlearch-1:
needs: [generate-matrix, backend-jobs-singlearch-1]
if: ${{ !cancelled() && needs.generate-matrix.outputs['has-merges-singlearch-1'] == 'true' }}
uses: ./.github/workflows/backend_merge.yml
with:
tag-latest: ${{ matrix.tag-latest }}
@@ -189,7 +298,55 @@ jobs:
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch']) }}
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-1']) }}
backend-merge-jobs-singlearch-2:
needs: [generate-matrix, backend-jobs-singlearch-2]
if: ${{ !cancelled() && needs.generate-matrix.outputs['has-merges-singlearch-2'] == 'true' }}
uses: ./.github/workflows/backend_merge.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
secrets:
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-2']) }}
backend-merge-jobs-singlearch-3:
needs: [generate-matrix, backend-jobs-singlearch-3]
if: ${{ !cancelled() && needs.generate-matrix.outputs['has-merges-singlearch-3'] == 'true' }}
uses: ./.github/workflows/backend_merge.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
secrets:
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-3']) }}
backend-merge-jobs-singlearch-4:
needs: [generate-matrix, backend-jobs-singlearch-4]
if: ${{ !cancelled() && needs.generate-matrix.outputs['has-merges-singlearch-4'] == 'true' }}
uses: ./.github/workflows/backend_merge.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
secrets:
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-4']) }}
backend-jobs-darwin:
needs: generate-matrix

View File

@@ -82,7 +82,7 @@ jobs:
# as the Linux registry cache.
- name: Restore Homebrew cache
id: brew-cache
uses: actions/cache/restore@v4
uses: actions/cache/restore@v6
with:
path: |
~/Library/Caches/Homebrew/downloads
@@ -142,7 +142,7 @@ jobs:
- name: Save Homebrew cache
if: github.event_name != 'pull_request' && steps.brew-cache.outputs.cache-hit != 'true'
uses: actions/cache/save@v4
uses: actions/cache/save@v6
with:
path: |
~/Library/Caches/Homebrew/downloads
@@ -169,16 +169,16 @@ jobs:
# invalidates cleanly; restore-keys fall back to the latest entry for the
# same pin so unchanged TUs stay warm even when the cache is fresh.
- name: Compute llama.cpp version
if: inputs.backend == 'llama-cpp' || inputs.backend == 'llama-cpp-localai-paged'
if: inputs.backend == 'llama-cpp'
id: llama-version
run: |
version=$(grep '^LLAMA_VERSION' backend/cpp/llama-cpp/Makefile | head -1 | cut -d= -f2 | cut -d'?' -f1 | tr -d ' ')
echo "version=${version}" >> "$GITHUB_OUTPUT"
- name: Restore ccache
if: inputs.backend == 'llama-cpp' || inputs.backend == 'llama-cpp-localai-paged'
if: inputs.backend == 'llama-cpp'
id: ccache-cache
uses: actions/cache/restore@v4
uses: actions/cache/restore@v6
with:
path: ~/Library/Caches/ccache
key: ccache-llama-${{ runner.arch }}-${{ steps.llama-version.outputs.version }}-${{ github.run_id }}
@@ -186,7 +186,7 @@ jobs:
ccache-llama-${{ runner.arch }}-${{ steps.llama-version.outputs.version }}-
- name: Configure ccache
if: inputs.backend == 'llama-cpp' || inputs.backend == 'llama-cpp-localai-paged'
if: inputs.backend == 'llama-cpp'
run: |
mkdir -p "$HOME/Library/Caches/ccache"
ccache -M 2G
@@ -211,7 +211,7 @@ jobs:
- name: Restore Python wheel cache
if: inputs.lang == 'python'
id: pyenv-cache
uses: actions/cache/restore@v4
uses: actions/cache/restore@v6
with:
path: |
~/Library/Caches/pip
@@ -251,24 +251,19 @@ jobs:
BACKEND=${{ inputs.backend }} BUILD_TYPE=${{ inputs.build-type }} USE_PIP=${{ inputs.use-pip }} make build-darwin-${{ inputs.lang }}-backend
- name: ccache stats
if: inputs.backend == 'llama-cpp' || inputs.backend == 'llama-cpp-localai-paged'
if: inputs.backend == 'llama-cpp'
run: ccache -s
# Only stock llama-cpp persists the ccache: both backends share the same
# ccache-llama-<arch>-<version>-<run_id> key, so the paged job restores from
# the shared prefix (warm) but must NOT also save under the identical key in
# the same run (it would collide). The shared upstream TUs stay warm via the
# stock save; the paged-only patched TUs are a small recompile.
- name: Save ccache
if: inputs.backend == 'llama-cpp' && github.event_name != 'pull_request'
uses: actions/cache/save@v4
uses: actions/cache/save@v6
with:
path: ~/Library/Caches/ccache
key: ccache-llama-${{ runner.arch }}-${{ steps.llama-version.outputs.version }}-${{ github.run_id }}
- name: Save Python wheel cache
if: inputs.lang == 'python' && github.event_name != 'pull_request' && steps.pyenv-cache.outputs.cache-hit != 'true'
uses: actions/cache/save@v4
uses: actions/cache/save@v6
with:
path: |
~/Library/Caches/pip

View File

@@ -11,16 +11,30 @@ jobs:
generate-matrix:
runs-on: ubuntu-latest
outputs:
matrix-singlearch: ${{ steps.set-matrix.outputs['matrix-singlearch'] }}
matrix-multiarch: ${{ steps.set-matrix.outputs['matrix-multiarch'] }}
matrix-darwin: ${{ steps.set-matrix.outputs['matrix-darwin'] }}
merge-matrix-multiarch: ${{ steps.set-matrix.outputs['merge-matrix-multiarch'] }}
merge-matrix-singlearch: ${{ steps.set-matrix.outputs['merge-matrix-singlearch'] }}
has-backends-singlearch: ${{ steps.set-matrix.outputs['has-backends-singlearch'] }}
has-backends-multiarch: ${{ steps.set-matrix.outputs['has-backends-multiarch'] }}
has-backends-darwin: ${{ steps.set-matrix.outputs['has-backends-darwin'] }}
has-merges-multiarch: ${{ steps.set-matrix.outputs['has-merges-multiarch'] }}
has-merges-singlearch: ${{ steps.set-matrix.outputs['has-merges-singlearch'] }}
# Single-arch backends are sharded across SINGLEARCH_SHARDS matrix jobs to
# stay under GitHub's 256-jobs-per-matrix limit (see changed-backends.js).
matrix-singlearch-1: ${{ steps.set-matrix.outputs['matrix-singlearch-1'] }}
merge-matrix-singlearch-1: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-1'] }}
has-backends-singlearch-1: ${{ steps.set-matrix.outputs['has-backends-singlearch-1'] }}
has-merges-singlearch-1: ${{ steps.set-matrix.outputs['has-merges-singlearch-1'] }}
matrix-singlearch-2: ${{ steps.set-matrix.outputs['matrix-singlearch-2'] }}
merge-matrix-singlearch-2: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-2'] }}
has-backends-singlearch-2: ${{ steps.set-matrix.outputs['has-backends-singlearch-2'] }}
has-merges-singlearch-2: ${{ steps.set-matrix.outputs['has-merges-singlearch-2'] }}
matrix-singlearch-3: ${{ steps.set-matrix.outputs['matrix-singlearch-3'] }}
merge-matrix-singlearch-3: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-3'] }}
has-backends-singlearch-3: ${{ steps.set-matrix.outputs['has-backends-singlearch-3'] }}
has-merges-singlearch-3: ${{ steps.set-matrix.outputs['has-merges-singlearch-3'] }}
matrix-singlearch-4: ${{ steps.set-matrix.outputs['matrix-singlearch-4'] }}
merge-matrix-singlearch-4: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-4'] }}
has-backends-singlearch-4: ${{ steps.set-matrix.outputs['has-backends-singlearch-4'] }}
has-merges-singlearch-4: ${{ steps.set-matrix.outputs['has-merges-singlearch-4'] }}
steps:
- name: Checkout repository
uses: actions/checkout@v7
@@ -71,10 +85,10 @@ jobs:
fail-fast: true
max-parallel: 8
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-multiarch']) }}
backend-jobs-singlearch:
backend-jobs-singlearch-1:
needs: generate-matrix
if: needs.generate-matrix.outputs['has-backends-singlearch-1'] == 'true'
uses: ./.github/workflows/backend_build.yml
if: needs.generate-matrix.outputs['has-backends-singlearch'] == 'true'
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
@@ -98,7 +112,94 @@ jobs:
strategy:
fail-fast: true
max-parallel: 8
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch']) }}
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-1']) }}
backend-jobs-singlearch-2:
needs: generate-matrix
if: needs.generate-matrix.outputs['has-backends-singlearch-2'] == 'true'
uses: ./.github/workflows/backend_build.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
build-type: ${{ matrix.build-type }}
cuda-major-version: ${{ matrix.cuda-major-version }}
cuda-minor-version: ${{ matrix.cuda-minor-version }}
platforms: ${{ matrix.platforms }}
platform-tag: ${{ matrix.platform-tag || '' }}
runs-on: ${{ matrix.runs-on }}
builder-base-image: ${{ matrix.builder-base-image || '' }}
base-image: ${{ matrix.base-image }}
backend: ${{ matrix.backend }}
dockerfile: ${{ matrix.dockerfile }}
skip-drivers: ${{ matrix.skip-drivers }}
context: ${{ matrix.context }}
ubuntu-version: ${{ matrix.ubuntu-version }}
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
secrets:
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: true
max-parallel: 8
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-2']) }}
backend-jobs-singlearch-3:
needs: generate-matrix
if: needs.generate-matrix.outputs['has-backends-singlearch-3'] == 'true'
uses: ./.github/workflows/backend_build.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
build-type: ${{ matrix.build-type }}
cuda-major-version: ${{ matrix.cuda-major-version }}
cuda-minor-version: ${{ matrix.cuda-minor-version }}
platforms: ${{ matrix.platforms }}
platform-tag: ${{ matrix.platform-tag || '' }}
runs-on: ${{ matrix.runs-on }}
builder-base-image: ${{ matrix.builder-base-image || '' }}
base-image: ${{ matrix.base-image }}
backend: ${{ matrix.backend }}
dockerfile: ${{ matrix.dockerfile }}
skip-drivers: ${{ matrix.skip-drivers }}
context: ${{ matrix.context }}
ubuntu-version: ${{ matrix.ubuntu-version }}
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
secrets:
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: true
max-parallel: 8
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-3']) }}
backend-jobs-singlearch-4:
needs: generate-matrix
if: needs.generate-matrix.outputs['has-backends-singlearch-4'] == 'true'
uses: ./.github/workflows/backend_build.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
build-type: ${{ matrix.build-type }}
cuda-major-version: ${{ matrix.cuda-major-version }}
cuda-minor-version: ${{ matrix.cuda-minor-version }}
platforms: ${{ matrix.platforms }}
platform-tag: ${{ matrix.platform-tag || '' }}
runs-on: ${{ matrix.runs-on }}
builder-base-image: ${{ matrix.builder-base-image || '' }}
base-image: ${{ matrix.base-image }}
backend: ${{ matrix.backend }}
dockerfile: ${{ matrix.dockerfile }}
skip-drivers: ${{ matrix.skip-drivers }}
context: ${{ matrix.context }}
ubuntu-version: ${{ matrix.ubuntu-version }}
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
secrets:
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: true
max-parallel: 8
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-4']) }}
backend-merge-jobs-multiarch:
needs: [generate-matrix, backend-jobs-multiarch]
# backend_merge.yml's push-side steps are all gated on
@@ -118,9 +219,9 @@ jobs:
fail-fast: false
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-multiarch']) }}
backend-merge-jobs-singlearch:
needs: [generate-matrix, backend-jobs-singlearch]
if: ${{ !cancelled() && github.event_name != 'pull_request' && needs.generate-matrix.outputs['has-merges-singlearch'] == 'true' }}
backend-merge-jobs-singlearch-1:
needs: [generate-matrix, backend-jobs-singlearch-1]
if: ${{ !cancelled() && github.event_name != 'pull_request' && needs.generate-matrix.outputs['has-merges-singlearch-1'] == 'true' }}
uses: ./.github/workflows/backend_merge.yml
with:
tag-latest: ${{ matrix.tag-latest }}
@@ -130,7 +231,49 @@ jobs:
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch']) }}
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-1']) }}
backend-merge-jobs-singlearch-2:
needs: [generate-matrix, backend-jobs-singlearch-2]
if: ${{ !cancelled() && github.event_name != 'pull_request' && needs.generate-matrix.outputs['has-merges-singlearch-2'] == 'true' }}
uses: ./.github/workflows/backend_merge.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
secrets:
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-2']) }}
backend-merge-jobs-singlearch-3:
needs: [generate-matrix, backend-jobs-singlearch-3]
if: ${{ !cancelled() && github.event_name != 'pull_request' && needs.generate-matrix.outputs['has-merges-singlearch-3'] == 'true' }}
uses: ./.github/workflows/backend_merge.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
secrets:
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-3']) }}
backend-merge-jobs-singlearch-4:
needs: [generate-matrix, backend-jobs-singlearch-4]
if: ${{ !cancelled() && github.event_name != 'pull_request' && needs.generate-matrix.outputs['has-merges-singlearch-4'] == 'true' }}
uses: ./.github/workflows/backend_merge.yml
with:
tag-latest: ${{ matrix.tag-latest }}
tag-suffix: ${{ matrix.tag-suffix }}
secrets:
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
strategy:
fail-fast: false
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-4']) }}
backend-jobs-darwin:
needs: generate-matrix
uses: ./.github/workflows/backend_build_darwin.yml

View File

@@ -6,6 +6,14 @@ on:
- master
pull_request:
# Supersede an in-flight run when a PR gets a new push. Keyed on the PR number
# so every push to the same PR shares a group; on a master push the key falls
# back to github.sha (unique per commit) and cancel-in-progress is false, so
# master runs never cancel each other -- each commit is built on its own.
concurrency:
group: ci-build-test-${{ github.event.pull_request.number || github.sha }}-${{ github.repository }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
jobs:
build-test:
runs-on: ubuntu-latest

View File

@@ -9,23 +9,6 @@ jobs:
strategy:
fail-fast: false
matrix:
# NOTE: there is intentionally NO entry for the llama-cpp-localai-paged
# backend. It carries a vendored paged-attention patch series
# (backend/cpp/llama-cpp-localai-paged/patches/paged/) hand-verified bit-exact against
# ONE specific llama.cpp tip; a naive nightly bump would move the tip out
# from under the patches and break `git apply` at build time. Its pin is
# therefore decoupled (its own LLAMA_VERSION in
# backend/cpp/llama-cpp-localai-paged/Makefile) and advanced ONLY by the
# manual PIN_SYNC process. Do not add it here. (turboquant CAN be
# auto-bumped below because its fork branch carries the patches.)
#
# Excluding it from the auto-bumper removed the early warning of upstream
# drift; that signal is restored separately by the dedicated canary
# .github/workflows/llama-cpp-paged-canary.yml, which weekly applies +
# compiles the paged series against the latest llama.cpp tip and goes red
# when upstream breaks it (prompting a PIN_SYNC). The canary is
# signal-only - it never opens a bump PR and never moves the pin - so
# this dep-bump workflow and its PRs stay green regardless.
include:
- repository: "ggml-org/llama.cpp"
variable: "LLAMA_VERSION"
@@ -39,10 +22,18 @@ jobs:
variable: "TURBOQUANT_VERSION"
branch: "feature/turboquant-kv-cache"
file: "backend/cpp/turboquant/Makefile"
- repository: "PrismML-Eng/llama.cpp"
variable: "BONSAI_VERSION"
branch: "prism"
file: "backend/cpp/bonsai/Makefile"
- repository: "antirez/ds4"
variable: "DS4_VERSION"
branch: "main"
file: "backend/cpp/ds4/Makefile"
- repository: "meituan-longcat/LongCat-Video"
variable: "LONGCAT_VIDEO_VERSION"
branch: "main"
file: "backend/python/longcat-video/Makefile"
- repository: "localai-org/privacy-filter.cpp"
variable: "PRIVACY_FILTER_VERSION"
branch: "master"
@@ -59,11 +50,15 @@ jobs:
variable: "PARAKEET_VERSION"
branch: "master"
file: "backend/go/parakeet-cpp/Makefile"
- repository: "mudler/ced.cpp"
variable: "CED_VERSION"
- repository: "localai-org/moss-transcribe.cpp"
variable: "MOSS_VERSION"
branch: "master"
file: "backend/go/moss-transcribe-cpp/Makefile"
- repository: "localai-org/ced.cpp"
variable: "CED_VERSION"
branch: "main"
file: "backend/go/ced/Makefile"
- repository: "mudler/voice-detect.cpp"
- repository: "localai-org/voice-detect.cpp"
variable: "VOICEDETECT_VERSION"
branch: "master"
file: "backend/go/voice-detect/Makefile"
@@ -95,7 +90,7 @@ jobs:
variable: "SAM3_VERSION"
branch: "main"
file: "backend/go/sam3-cpp/Makefile"
- repository: "mudler/rf-detr.cpp"
- repository: "localai-org/rf-detr.cpp"
variable: "RFDETR_VERSION"
branch: "main"
file: "backend/go/rfdetr-cpp/Makefile"
@@ -120,6 +115,12 @@ jobs:
- uses: actions/checkout@v7
- name: Bump dependencies 🔧
id: bump
env:
# This job fans out to ~25 parallel matrix entries, all querying
# api.github.com from runner IPs that share the 60/hour anonymous
# rate limit. Authenticating raises it to 1000/hour, which is what
# kept a random handful of these red every night.
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
bash .github/bump_deps.sh ${{ matrix.repository }} ${{ matrix.branch }} ${{ matrix.variable }} ${{ matrix.file }}
{
@@ -156,6 +157,8 @@ jobs:
- uses: actions/checkout@v7
- name: Bump vLLM cu130 wheel pin 🔧
id: bump
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
bash .github/bump_vllm_wheel.sh vllm-project/vllm backend/python/vllm/requirements-cublas13-after.txt VLLM_VERSION
{
@@ -192,6 +195,8 @@ jobs:
- uses: actions/checkout@v7
- name: Bump vllm-metal pin 🔧
id: bump
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
bash .github/bump_vllm_metal.sh vllm-project/vllm-metal backend/python/vllm/install.sh VLLM_METAL_VERSION
{

View File

@@ -15,6 +15,10 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Bump dependencies 🔧
env:
# Authenticated API calls get 1000 req/hour instead of the 60/hour
# anonymous cap that is shared across every job on the runner's IP.
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
bash .github/bump_docs.sh ${{ matrix.repository }}
- name: Create Pull Request

39
.github/workflows/ci-tools-tests.yaml vendored Normal file
View File

@@ -0,0 +1,39 @@
---
# The packages under .github/ci/ are invisible to `go list ./...`, so neither
# `make lint` nor the repository test run ever touches them. Their specs are
# dead weight until a workflow names each package explicitly.
name: 'CI tool tests'
on:
pull_request:
paths:
- '.github/ci/**'
- '.github/workflows/ci-tools-tests.yaml'
push:
branches:
- master
paths:
- '.github/ci/**'
jobs:
ci-tools:
name: 'Test the .github/ci generators'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
cache: false
# The discovery heuristics are the risky part of these tools. A regression
# produces confident, wrong gallery entries, which is worse than no tool.
- name: 'Test the APEX entry generator'
run: go test ./.github/ci/apexentries/
- name: 'Test the variant proposer'
run: go test ./.github/ci/variantproposals/
# Shared by both generators above. Its behaviour is exercised through their
# specs; this step exists so a break in the shared package fails under its
# own name rather than as a puzzling failure in whichever caller ran first.
- name: 'Test the shared gallery editor'
run: go test ./.github/ci/galleryedit/

View File

@@ -0,0 +1,54 @@
name: Propose gallery variant groupings
on:
schedule:
- cron: 0 4 * * 1
workflow_dispatch:
jobs:
variant_proposals:
if: github.repository == 'mudler/LocalAI'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
cache: false
# The heuristics are the risky part of this job. A regression in them
# produces confident, wrong proposals, which is worse than no job at all.
- name: Test the proposer
run: go test ./.github/ci/variantproposals/
- name: Propose groupings 🔧
id: propose
run: |
rm -f /tmp/variant-proposals-body.md
go run ./.github/ci/variantproposals \
-index gallery/index.yaml \
-ledger gallery/variant-exclusions.yaml \
-body-out /tmp/variant-proposals-body.md \
-apply
if [ -s /tmp/variant-proposals-body.md ]; then
echo "have_proposals=true" >> "$GITHUB_OUTPUT"
{
echo 'body<<VARIANT_PROPOSAL_BODY_EOF'
cat /tmp/variant-proposals-body.md
echo VARIANT_PROPOSAL_BODY_EOF
} >> "$GITHUB_OUTPUT"
else
echo "have_proposals=false" >> "$GITHUB_OUTPUT"
fi
# No body file means the proposer found nothing. Opening an empty pull
# request every run is how a proposal job gets muted by its reviewers.
- name: Create Pull Request
if: steps.propose.outputs.have_proposals == 'true'
uses: peter-evans/create-pull-request@v8
with:
token: ${{ secrets.UPDATE_BOT_TOKEN }}
push-to-fork: ci-forks/LocalAI
commit-message: 'chore(model-gallery): propose variant groupings'
title: 'chore(model-gallery): propose variant groupings for review'
branch: "propose/variant-groupings"
body: ${{ steps.propose.outputs.body }}
signoff: true

View File

@@ -46,3 +46,23 @@ jobs:
touch core/http/react-ui/dist/index.html
- name: lint
run: make lint
build-scripts:
# The image packaging scripts encode invariants that only surface inside a
# container build (a missing transitive dep, a partial cuDNN family). Their
# shell tests need nothing but bash + gcc + ldd, so run them on every PR
# rather than waiting on a multi-GB cross-arch backend image build.
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- name: run packaging script tests
run: make test-build-scripts
# The backend matrix path filter fails silently: a miss emits an empty
# matrix, every job goes green, and the change reaches no image (#10946).
# Its tests need only node, so they ride along with this job.
- uses: actions/setup-node@v4
with:
node-version: '20'
- name: run CI script tests
run: make test-ci-scripts

View File

@@ -1,179 +0,0 @@
name: 'llama.cpp paged patches: upstream canary'
# EARLY-WARNING CANARY for the vendored paged-attention patch series
# (backend/cpp/llama-cpp-localai-paged/patches/paged/0001-0030).
#
# WHY THIS EXISTS
# The paged backend (backend/cpp/llama-cpp-localai-paged) pins its OWN verified
# llama.cpp tip (LLAMA_VERSION in backend/cpp/llama-cpp-localai-paged/Makefile)
# and is intentionally EXCLUDED from the nightly auto-bumper
# (.github/workflows/bump_deps.yaml), so a naive upstream bump can never silently
# break the shipped build. The cost of that safety: nobody finds out when
# upstream DRIFTS past the patches. This canary restores that signal WITHOUT
# touching the shipped pin - weekly it tries the patch series + a real compile
# against the LATEST llama.cpp master tip and goes red the moment upstream breaks
# the patches.
#
# RED HERE means: time to run a PIN_SYNC (rebase the patches onto the new tip,
# pass the bit-exact gate on the GPU, re-export the .patch files, THEN advance
# the pin in backend/cpp/llama-cpp-localai-paged/Makefile). See the backend README
# section 7 (Pin + maintenance policy):
# backend/cpp/llama-cpp-localai-paged/README.md.
#
# SIGNAL-ONLY: this workflow moves no pinned version, ships nothing, and is fully
# decoupled from bump_deps - so the main dep-bump PR stays green regardless. A
# green run means "the paged series still applies and compiles on upstream HEAD";
# a red run means "upstream moved - schedule a pin-sync".
on:
schedule:
# Weekly (Mondays 06:00 UTC), mirroring the weekly DEPS_REFRESH / bump_deps
# cadence. Offset from bump_deps' nightly 20:00 so the two never pile up.
- cron: '0 6 * * 1'
workflow_dispatch:
permissions:
contents: read
concurrency:
group: llama-cpp-paged-canary
cancel-in-progress: false
env:
# Upstream source of truth - the same repo/branch bump_deps tracks for the
# stock llama-cpp pin.
LLAMA_UPSTREAM: 'https://github.com/ggml-org/llama.cpp'
jobs:
apply-check:
# Cheap, fast, toolchain-free early warning: does the series still APPLY to
# the latest upstream tip? A patch no longer applying is by far the most
# common way upstream breaks a vendored series, so this runs first, is
# reliable on a free runner, and feeds the resolved tip to the compile job.
if: github.repository == 'mudler/LocalAI'
runs-on: ubuntu-latest
timeout-minutes: 20
outputs:
tip: ${{ steps.resolve.outputs.tip }}
steps:
- name: Checkout LocalAI
uses: actions/checkout@v7
- name: Resolve latest llama.cpp master tip
id: resolve
run: |
tip="$(git ls-remote "$LLAMA_UPSTREAM" refs/heads/master | cut -f1)"
if [ -z "$tip" ]; then
echo "::error::could not resolve llama.cpp master tip from $LLAMA_UPSTREAM"
exit 1
fi
pin="$(grep -m1 'LLAMA_VERSION?=' backend/cpp/llama-cpp-localai-paged/Makefile | cut -d= -f2)"
echo "latest llama.cpp master tip: $tip"
echo "shipped paged pin: $pin"
echo "tip=$tip" >> "$GITHUB_OUTPUT"
{
echo "## llama.cpp paged canary"
echo ""
echo "- upstream master tip: \`$tip\`"
echo "- shipped paged pin: \`$pin\`"
} >> "$GITHUB_STEP_SUMMARY"
- name: Checkout llama.cpp at latest tip (shallow)
run: |
mkdir -p /tmp/llama.cpp
cd /tmp/llama.cpp
git init -q
git remote add origin "$LLAMA_UPSTREAM"
git fetch -q --depth 1 origin "${{ steps.resolve.outputs.tip }}"
git checkout -q FETCH_HEAD
git log --oneline -1
- name: Apply paged patch series (build's git-apply method)
run: |
bash .github/scripts/paged-canary-apply.sh \
/tmp/llama.cpp \
"$PWD/backend/cpp/llama-cpp-localai-paged/patches"
echo "- apply: full paged series applies to the upstream tip :white_check_mark:" >> "$GITHUB_STEP_SUMMARY"
compile:
# Proves the patches still COMPILE against the latest tip, using the SAME
# toolchain + build target the shipped paged backend uses (the
# base-grpc-cuda-12 builder base + the Makefile `grpc-server` cublas target),
# so a failure means upstream drift, not toolchain noise. CUDA is compiled
# (nvcc; no GPU required) because most of the paged series is CUDA kernels.
# Runs only if the apply check passed, on the exact tip it validated.
#
# If a full CUDA compile on the hosted runner ever proves too heavy/flaky,
# switch `runs-on` to 'bigger-runner' (the runner class the real paged CUDA
# build uses), or drop to a CPU build (BUILD_TYPE='') which still compiles
# all host + CPU paged code, leaving CUDA-kernel coverage to the apply check
# plus the manual PIN_SYNC GPU gate.
needs: apply-check
if: github.repository == 'mudler/LocalAI'
runs-on: ubuntu-latest
timeout-minutes: 180
steps:
- name: Checkout LocalAI
uses: actions/checkout@v7
- name: Free disk space
uses: ./.github/actions/free-disk-space
with:
mode: hosted
- name: Login to Quay.io
uses: docker/login-action@v4
with:
registry: quay.io
username: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
password: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
- name: Compile paged backend against latest tip (cublas)
env:
TIP: ${{ needs.apply-check.outputs.tip }}
BUILDER_BASE_IMAGE: 'quay.io/go-skynet/ci-cache:base-grpc-cuda-12-amd64'
run: |
docker run --rm \
-v "$PWD":/LocalAI -w /LocalAI \
-e TIP -e LLAMA_UPSTREAM \
"$BUILDER_BASE_IMAGE" bash -euxo pipefail -c '
# Mirror the Dockerfile: gRPC lives at /opt/grpc in the base image;
# copy it to the prefix CMake find_package expects.
cp -a /opt/grpc/. /usr/local/
# Pre-populate the llama.cpp checkout at the latest tip with the
# paged series applied via the tolerant canary apply. Because
# backend/cpp/llama-cpp/llama.cpp now exists, the stock Makefile's
# llama.cpp target (clone + base-patch apply) is skipped and the
# now patch-free prepare.sh only copies the grpc-server sources -
# so we drive the REAL grpc-server build path on top of our paged
# apply. The stock llama-cpp backend no longer carries the paged
# series (it lives in backend/cpp/llama-cpp-localai-paged/patches/
# paged); we build it here in the stock dir only because that is
# where the shared build infra (Makefile / grpc-server.cpp /
# CMakeLists.txt / prepare.sh) lives.
cd backend/cpp/llama-cpp/
mkdir -p llama.cpp
cd llama.cpp
git init -q
git remote add origin "$LLAMA_UPSTREAM"
git fetch -q --depth 1 origin "$TIP"
git checkout -q FETCH_HEAD
cd /LocalAI
bash .github/scripts/paged-canary-apply.sh \
backend/cpp/llama-cpp/llama.cpp \
"$PWD/backend/cpp/llama-cpp-localai-paged/patches"
# Cheapest real CUDA build that proves the patches compile: one
# CUDA arch, cublas. CMAKE_ARGS is passed via the environment (not
# as a make arg) so the Makefile += flags are still appended,
# exactly like .docker/llama-cpp-localai-paged-compile.sh. The paged
# series is already applied to the checkout above, so the stock
# build just compiles the patched tree.
cd backend/cpp/llama-cpp/
BUILD_TYPE=cublas \
CMAKE_ARGS="-DCMAKE_CUDA_ARCHITECTURES=80" \
make grpc-server
test -x grpc-server
'
echo "- compile: paged series builds (cublas) against the upstream tip :white_check_mark:" >> "$GITHUB_STEP_SUMMARY"

View File

@@ -0,0 +1,69 @@
---
name: 'realtime-conformance'
# Verifies the realtime state-machine implementations conform to their formal
# designs (docs/design/realtime-state-machines.md, formal-verification/). BOTH
# layers are enforced and the gate is fail-closed: the Go conformance layer
# (respcoord + turncoord transition/rapid tests under -race) AND the FizzBee model check of
# the authoritative specs. FizzBee is pinned + checksum-verified
# (formal-verification/fizzbee.sha256), so a failed install fails the job rather
# than silently skipping verification.
on:
pull_request:
paths:
- 'core/http/endpoints/openai/coordinator/**'
- 'core/http/endpoints/openai/respcoord/**'
- 'core/http/endpoints/openai/turncoord/**'
- 'core/http/endpoints/openai/conncoord/**'
- 'core/http/endpoints/openai/compactcoord/**'
- 'core/http/endpoints/openai/ttscoord/**'
- 'formal-verification/**'
- 'scripts/realtime-conformance.sh'
- 'scripts/install-fizzbee.sh'
- '.github/workflows/realtime-conformance.yml'
push:
branches:
- master
paths:
- 'core/http/endpoints/openai/coordinator/**'
- 'core/http/endpoints/openai/respcoord/**'
- 'core/http/endpoints/openai/turncoord/**'
- 'core/http/endpoints/openai/conncoord/**'
- 'core/http/endpoints/openai/compactcoord/**'
- 'core/http/endpoints/openai/ttscoord/**'
- 'formal-verification/**'
- 'scripts/realtime-conformance.sh'
concurrency:
group: realtime-conformance-${{ github.event.pull_request.number || github.sha }}-${{ github.repository }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
jobs:
conformance:
runs-on: ubuntu-latest
strategy:
matrix:
go-version: ['1.26.x']
steps:
- name: Clone
uses: actions/checkout@v7
- name: Setup Go ${{ matrix.go-version }}
uses: actions/setup-go@v5
with:
go-version: ${{ matrix.go-version }}
cache: false
- name: Cache FizzBee
uses: actions/cache@v6
with:
path: .tools/fizzbee
key: fizzbee-v0.5.2-${{ runner.os }}-${{ hashFiles('formal-verification/fizzbee.sha256') }}
- name: Install FizzBee (pinned, checksum-verified)
# No `|| true`: a failed/forged download must fail the job, not silently
# drop the design verification. install-fizzbee.sh is a no-op if the
# cached binary is already present and valid.
run: ./scripts/install-fizzbee.sh
- name: Run conformance gate (fail-closed)
# No skip env: both the Go conformance and the FizzBee model check are
# required. The gate auto-detects .tools/fizzbee/fizz.
run: make test-realtime-conformance

View File

@@ -7,6 +7,19 @@ on:
schedule:
- cron: '0 0 * * 0'
# `push:` is deliberately unfiltered, so this fires on every push to every
# branch and there is no pull_request event to key on -- the usual
# `github.event.pull_request.number || github.sha` idiom used elsewhere would
# key on the unique-per-commit sha and dedup nothing. Group on the ref instead
# so successive pushes to the same feature branch supersede one another.
#
# Cancelling is safe here: the only output is a SARIF upload, and code scanning
# tracks the latest result per ref, so a superseded scan has nothing to lose.
# master is excluded anyway -- every commit on master gets its own scan.
concurrency:
group: ci-secscan-${{ github.ref }}-${{ github.repository }}
cancel-in-progress: ${{ github.ref != 'refs/heads/master' }}
jobs:
tests:
runs-on: ubuntu-latest

View File

@@ -11,7 +11,7 @@ jobs:
if: github.repository == 'mudler/LocalAI'
runs-on: ubuntu-latest
steps:
- uses: actions/stale@eb5cf3af3ac0a1aa4c9c45633dd1ae542a27a899 # v9
- uses: actions/stale@1e223db275d687790206a7acac4d1a11bd6fe629 # v9
with:
stale-issue-message: 'This issue is stale because it has been open 90 days with no activity. Remove stale label or comment or this will be closed in 5 days.'
stale-pr-message: 'This PR is stale because it has been open 90 days with no activity. Remove stale label or comment or this will be closed in 10 days.'

View File

@@ -587,7 +587,7 @@ jobs:
with:
go-version: '1.25.4'
- name: Setup Node.js
uses: actions/setup-node@v6
uses: actions/setup-node@v7
with:
node-version: '22'
- name: Build sherpa-onnx backend image and run realtime e2e tests

View File

@@ -48,7 +48,7 @@ jobs:
sudo apt-get update
sudo apt-get install curl ffmpeg libopus-dev
- name: Setup Node.js
uses: actions/setup-node@v6
uses: actions/setup-node@v7
with:
node-version: '22'
- name: Build React UI
@@ -100,7 +100,7 @@ jobs:
brew install protobuf grpc make protoc-gen-go protoc-gen-go-grpc libomp llvm opus ffmpeg
pip install --user --no-cache-dir grpcio-tools grpcio
- name: Setup Node.js
uses: actions/setup-node@v6
uses: actions/setup-node@v7
with:
node-version: '22'
- name: Build React UI

View File

@@ -47,7 +47,7 @@ jobs:
sudo apt-get update
sudo apt-get install -y build-essential libopus-dev
- name: Setup Node.js
uses: actions/setup-node@v6
uses: actions/setup-node@v7
with:
node-version: '22'
- name: Build React UI

View File

@@ -34,7 +34,7 @@ jobs:
go-version: ${{ matrix.go-version }}
cache: false
- name: Setup Node.js
uses: actions/setup-node@v6
uses: actions/setup-node@v7
with:
node-version: '22'
- name: Setup Bun

View File

@@ -1,6 +1,11 @@
name: 'Yamllint GitHub Actions'
on:
- pull_request
concurrency:
group: ci-yamllint-${{ github.event.pull_request.number || github.sha }}-${{ github.repository }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
jobs:
yamllint:
name: 'Yamllint'

32
.gitignore vendored
View File

@@ -9,15 +9,6 @@ prepare-sources
/backend/cpp/llama-cpp/llama.cpp
/backend/cpp/llama-*
!backend/cpp/llama-cpp
# llama-cpp-localai-paged is a tracked source dir (a thin wrapper Makefile over
# backend/cpp/llama-cpp). Re-include it like llama-cpp above; its sibling
# *-build dirs are still ignored by the /backend/cpp/llama-* rule, and its
# in-dir build artifacts (binaries, package output, collected ggml .so set) are
# re-ignored just below.
!backend/cpp/llama-cpp-localai-paged
/backend/cpp/llama-cpp-localai-paged/llama-cpp-localai-paged-*
/backend/cpp/llama-cpp-localai-paged/package
/backend/cpp/llama-cpp-localai-paged/ggml-shared-libs
/backends
/backend-images
/result.yaml
@@ -50,7 +41,12 @@ models/*
test-models/
test-dir/
tests/e2e-aio/backends
mock-backend
# The mock backend binary built by `make build-mock-backend`. Anchored to its
# full path: a bare `mock-backend` also matched the *directory* holding the
# source, so git would not descend into it and adding a file there needed -f.
# tests/e2e/mock-backend/.gitignore covers the same binary; kept here too so
# the artifact stays ignored if that scoped file is ever removed.
/tests/e2e/mock-backend/mock-backend
release/
@@ -106,3 +102,19 @@ core/http/react-ui/test-results/
# Local Apple signing material (never commit)
.certs/
# Pinned dev tools (e.g. FizzBee for the realtime-conformance gate)
.tools/
# FizzBee model-check artifacts: the parser emits <spec>.json next to each
# .fizz and the checker writes run dirs under out/. Both are regenerated by
# the realtime-conformance gate; only the .fizz sources are authoritative.
formal-verification/*.json
formal-verification/out/
# `go build ./.github/ci/apexentries` drops a binary of the package name into
# whatever directory it runs in, one `git add -A` away from being committed.
# Both paths are anchored: an unanchored `apexentries` would also match the
# package directory itself and untrack the source.
/apexentries
/.github/ci/apexentries/apexentries

View File

@@ -0,0 +1,6 @@
{
"files": ["core/http/react-ui/index.html"],
"insertBefore": "</body>",
"commentSyntax": "html",
"cspChecked": true
}

View File

@@ -23,8 +23,6 @@ LocalAI follows the Linux kernel project's [guidelines for AI coding assistants]
| [.agents/adding-backends.md](.agents/adding-backends.md) | Adding a new backend (Python, Go, or C++) — full step-by-step checklist, including importer integration (the `/import-model` dropdown is server-driven from `GET /backends/known`) |
| [.agents/coding-style.md](.agents/coding-style.md) | Code style, editorconfig, logging, documentation conventions |
| [.agents/llama-cpp-backend.md](.agents/llama-cpp-backend.md) | Working on the llama.cpp backend — architecture, updating, tool call parsing |
| [.agents/llama-cpp-localai-paged-backend.md](.agents/llama-cpp-localai-paged-backend.md) | Working on the CUDA-only paged-attention llama.cpp variant (Qwen3.6 hybrid-SSM / Blackwell NVFP4 decode) - patchset scope, the bit-exact gate, the manual pin-sync + weekly canary, CUDA-only invariants, stock-stays-pure, Metal/SYCL/Vulkan follow-up scope |
| [.agents/vllm-parity-methodology.md](.agents/vllm-parity-methodology.md) | The methodology for closing the vLLM decode-throughput gap in llama.cpp - bit-exact gating, profile-don't-assume, both-engine ground-truth, per-lever A/B discipline, recording rejected levers, multi-agent GPU orchestration |
| [.agents/vllm-backend.md](.agents/vllm-backend.md) | Working on the vLLM / vLLM-omni backends — native parsers, ChatDelta, CPU build, libnuma packaging, backend hooks |
| [.agents/sglang-backend.md](.agents/sglang-backend.md) | Working on the SGLang backend — `engine_args` validation against ServerArgs, speculative-decoding (EAGLE/EAGLE3/DFLASH/MTP) recipes, parser handling |
| [.agents/ds4-backend.md](.agents/ds4-backend.md) | Working on the ds4 backend - DSML state machine, thinking modes, KV cache, Metal+CUDA matrix |
@@ -39,12 +37,12 @@ LocalAI follows the Linux kernel project's [guidelines for AI coding assistants]
- **Git hooks & coverage gates**: Run `make install-hooks` once per clone so the pre-commit lint + coverage gates run. **Never bypass them with `git commit --no-verify`, and never lower a coverage baseline or widen a gate's tolerance to turn a red gate green** — the coverage ratchet only moves up. If a change drops coverage, add tests to raise it (e.g. render-smoke specs). See [.agents/building-and-testing.md](.agents/building-and-testing.md).
- **Logging**: Use `github.com/mudler/xlog` (same API as slog)
- **Paged llama.cpp backend**: `llama-cpp-localai-paged` is a CUDA-only variant that owns its own patch series + its own pinned llama.cpp (manual pin-sync, weekly canary); the stock `llama-cpp` backend stays patch-free. Read [.agents/llama-cpp-localai-paged-backend.md](.agents/llama-cpp-localai-paged-backend.md) before touching either, and [.agents/vllm-parity-methodology.md](.agents/vllm-parity-methodology.md) for the decode-parity methodology behind it.
- **Go style**: Prefer `any` over `interface{}`
- **Comments**: Explain *why*, not *what*
- **Docs**: Update `docs/content/` when adding features or changing config
- **Docs (docs-with-code rule)**: When you change user-facing behavior (API endpoints, CLI flags, config keys, or features), update the corresponding page under `docs/content/` in the SAME change, not as a follow-up. A user-facing change without a matching docs update is incomplete. See also the documentation conventions in [.agents/coding-style.md](.agents/coding-style.md).
- **New API endpoints**: LocalAI advertises its capability surface in several independent places — swagger `@Tags`, `/api/instructions` registry, auth `RouteFeatureRegistry`, React UI `capabilities.js`, docs. Read [.agents/api-endpoints-and-auth.md](.agents/api-endpoints-and-auth.md) and follow its checklist — missing any surface means clients, admins, and the UI won't know the endpoint exists.
- **Admin endpoints → MCP tool**: every admin endpoint that an admin would manage conversationally (install/list/edit/toggle/upgrade) MUST also be exposed as an MCP tool in `pkg/mcp/localaitools/`. The LocalAI Assistant chat modality and the standalone `local-ai mcp-server` consume that package; drift between REST and MCP is a real risk. Read [.agents/localai-assistant-mcp.md](.agents/localai-assistant-mcp.md) — the `TestToolHTTPRouteMappingComplete` test fails until you wire the new tool and update the route map.
- **Build**: Inspect `Makefile` and `.github/workflows/` — ask the user before running long builds
- **Backend OS coverage**: a new backend must target every OS it can build for, not just Linux. `.github/backend-matrix.yml` has two matrices — `include:` (Linux) and `includeDarwin:` (macOS / Apple Silicon). Most C/C++/GGML and many Python backends build on Darwin too — wire the `includeDarwin` entry + `backend/index.yaml` `metal:` entries, or say in the PR why an OS is unsupported. See the darwin checklist in [.agents/adding-backends.md](.agents/adding-backends.md).
- **Gallery variant ranking**: a gallery entry can declare `variants` (alternative builds of the same weights), and LocalAI ranks the ones a host can run by engine preference first, size second. A new backend that should be preferred on some hardware must be listed in `engineNamePreferenceRules` in `pkg/system/capabilities.go`; the sibling `backendBuildTagPreferenceRules` speaks build tags rather than engine names, and using the wrong table matches nothing without erroring. See [.agents/adding-backends.md](.agents/adding-backends.md).
- **UI**: The active UI is the React app in `core/http/react-ui/`. The older Alpine.js/HTML UI in `core/http/static/` is pending deprecation — all new UI work goes in the React UI

View File

@@ -12,12 +12,16 @@ ARG APT_MIRROR
ARG APT_PORTS_MIRROR
ENV DEBIAN_FRONTEND=noninteractive
# hwdata ships /usr/share/hwdata/pci.ids. Without it, the ghw library we use
# for hardware detection cannot resolve PCI vendor IDs and fails to enumerate
# GPUs at all, so the image reports "No GPU detected" (see issue #10941).
RUN --mount=type=bind,source=.docker/apt-mirror.sh,target=/usr/local/sbin/apt-mirror \
APT_MIRROR="${APT_MIRROR}" APT_PORTS_MIRROR="${APT_PORTS_MIRROR}" sh /usr/local/sbin/apt-mirror && \
apt-get update && \
apt-get install -y --no-install-recommends \
ca-certificates curl wget espeak-ng libgomp1 \
ffmpeg libopenblas0 libopenblas-dev libopus0 sox && \
ffmpeg libopenblas0 libopenblas-dev libopus0 sox \
hwdata && \
apt-get clean && \
rm -rf /var/lib/apt/lists/*
@@ -171,6 +175,17 @@ RUN if [ "${BUILD_TYPE}" = "hipblas" ]; then \
ln -s /opt/rocm-**/lib/llvm/lib/libomp.so /usr/lib/libomp.so \
; fi
# ROCm's bundled libdrm_amdgpu is built with a hardcoded fallback lookup path
# for the ASIC ID table (/opt/amdgpu/share/libdrm/amdgpu.ids), which only exists
# if AMD's full amdgpu graphics/DKMS stack is installed. This compute-only image
# doesn't have it, so hipblas/rocBLAS log "No such file or directory" on every
# model load and can fail to identify the GPU. Point it at the equivalent file
# Ubuntu's libdrm-common package already ships.
RUN if [ "${BUILD_TYPE}" = "hipblas" ] && [ -f /usr/share/libdrm/amdgpu.ids ] && [ ! -e /opt/amdgpu/share/libdrm/amdgpu.ids ]; then \
mkdir -p /opt/amdgpu/share/libdrm && \
ln -s /usr/share/libdrm/amdgpu.ids /opt/amdgpu/share/libdrm/amdgpu.ids \
; fi
RUN expr "${BUILD_TYPE}" = intel && echo "intel" > /run/localai/capability || echo "not intel"
# Cuda
@@ -378,7 +393,12 @@ RUN go install github.com/mikefarah/yq/v4@latest
# If you cannot find a more suitable place for an addition, this layer is a suitable place for it.
FROM requirements-drivers
ENV HEALTHCHECK_ENDPOINT=http://localhost:8080/readyz
# Optional override for the HEALTHCHECK target. Left empty so healthcheck.sh
# derives the endpoint from the mode the container is actually running — the
# same image runs `local-ai run` (HTTP on 8080) and `local-ai worker` (HTTP on
# the gRPC base port minus one), and a hardcoded default marked every worker
# permanently unhealthy (#10987). Set it to pin an explicit URL.
ENV HEALTHCHECK_ENDPOINT=""
ARG CUDA_MAJOR_VERSION=12
ENV NVIDIA_DRIVER_CAPABILITIES=compute,utility
@@ -388,6 +408,7 @@ ENV NVIDIA_VISIBLE_DEVICES=all
WORKDIR /
COPY ./entrypoint.sh .
COPY ./scripts/build/healthcheck.sh .
# Copy the binary
COPY --from=builder /build/local-ai ./
@@ -398,9 +419,22 @@ RUN --mount=from=builder,src=/build/,dst=/mnt/build \
# Make sure the models directory exists
RUN mkdir -p /models /backends /data
# Define the health check command
HEALTHCHECK --interval=1m --timeout=10m --retries=10 \
CMD curl -f ${HEALTHCHECK_ENDPOINT} || exit 1
# Define the health check command.
#
# --start-period is the knob for slow starts, not --timeout/--retries. Since
# #10949 a frontend's startup preload materializes HuggingFace artifacts before
# the HTTP server binds (31 GB observed on a live cluster), so a healthy replica
# can legitimately fail probes for a long time. Failures inside the start period
# leave the container `starting` instead of burning retries, and the period ends
# early on the first success — so a generous value costs a fast-starting
# container nothing. A process that actually died is handled by the restart
# policy, not by health.
#
# --timeout is a per-probe deadline: 10m meant a wedged probe could hang for ten
# minutes and stretch detection without bound. A localhost curl that has not
# answered in 10s is itself the fault being detected.
HEALTHCHECK --start-period=60m --interval=1m --timeout=10s --retries=3 \
CMD /healthcheck.sh
VOLUME /models /backends /configuration /data
EXPOSE 8080

103
Makefile
View File

@@ -1,5 +1,5 @@
# Disable parallel execution for backend builds
.NOTPARALLEL: backends/diffusers backends/llama-cpp backends/turboquant backends/outetts backends/piper backends/stablediffusion-ggml backends/whisper backends/crispasr backends/parakeet-cpp backends/faster-whisper backends/silero-vad backends/local-store backends/huggingface backends/rfdetr backends/rfdetr-cpp backends/insightface backends/speaker-recognition backends/kitten-tts backends/kokoro backends/chatterbox backends/llama-cpp-darwin backends/neutts build-darwin-python-backend build-darwin-go-backend backends/mlx backends/diffuser-darwin backends/mlx-vlm backends/mlx-audio backends/mlx-distributed backends/stablediffusion-ggml-darwin backends/vllm backends/vllm-omni backends/sglang backends/moonshine backends/pocket-tts backends/qwen-tts backends/faster-qwen3-tts backends/qwen-asr backends/nemo backends/voxcpm backends/whisperx backends/ace-step backends/acestep-cpp backends/fish-speech backends/voxtral backends/opus backends/trl backends/llama-cpp-quantization backends/kokoros backends/sam3-cpp backends/qwen3-tts-cpp backends/omnivoice-cpp backends/vibevoice-cpp backends/localvqe backends/tinygrad backends/sherpa-onnx backends/ds4 backends/ds4-darwin backends/liquid-audio backends/supertonic backends/depth-anything-cpp backends/privacy-filter backends/privacy-filter-darwin backends/llama-cpp-localai-paged
.NOTPARALLEL: backends/diffusers backends/llama-cpp backends/turboquant backends/bonsai backends/outetts backends/piper backends/stablediffusion-ggml backends/whisper backends/crispasr backends/parakeet-cpp backends/moss-transcribe-cpp backends/faster-whisper backends/silero-vad backends/local-store backends/cloud-proxy backends/huggingface backends/rfdetr backends/rfdetr-cpp backends/insightface backends/speaker-recognition backends/kitten-tts backends/kokoro backends/chatterbox backends/llama-cpp-darwin backends/neutts build-darwin-python-backend build-darwin-go-backend backends/mlx backends/diffuser-darwin backends/mlx-vlm backends/mlx-audio backends/mlx-distributed backends/stablediffusion-ggml-darwin backends/vllm backends/vllm-omni backends/longcat-video backends/sglang backends/moonshine backends/pocket-tts backends/qwen-tts backends/faster-qwen3-tts backends/qwen-asr backends/nemo backends/voxcpm backends/whisperx backends/ace-step backends/acestep-cpp backends/fish-speech backends/voxtral backends/opus backends/trl backends/llama-cpp-quantization backends/kokoros backends/sam3-cpp backends/qwen3-tts-cpp backends/moss-tts-cpp backends/omnivoice-cpp backends/vibevoice-cpp backends/localvqe backends/tinygrad backends/sherpa-onnx backends/ds4 backends/ds4-darwin backends/liquid-audio backends/supertonic backends/depth-anything-cpp backends/privacy-filter backends/privacy-filter-darwin
GOCMD=go
GOTEST=$(GOCMD) test
@@ -103,7 +103,7 @@ COVERAGE_E2E_LABELS?=!real-models
COVERAGE_EXCLUDE_RE?=grpc/proto/.*[.]pb[.]go
.PHONY: all test test-coverage test-coverage-baseline test-coverage-check test-backend-cpp test-ui test-ui-coverage-baseline test-ui-coverage-check install-hooks build vendor lint lint-all
.PHONY: all test test-coverage test-coverage-baseline test-coverage-check test-backend-cpp test-build-scripts test-ui test-ui-coverage-baseline test-ui-coverage-check install-hooks build vendor lint lint-all
all: help
@@ -208,6 +208,20 @@ test: prepare-test
test-backend-cpp:
bash backend/cpp/run-unit-tests.sh
## Runs the shell-level regression tests for the image packaging scripts
## (scripts/build/*_test.sh). These guard invariants that only ever break
## inside a container build - a missing transitive dep, a partial cuDNN
## family - and that no Go test can observe. Needs only bash + gcc + ldd.
test-build-scripts:
@set -e; for t in scripts/build/*_test.sh; do echo "== $$t"; bash "$$t"; done
## Runs the unit tests for the CI helper scripts under scripts/lib/. Currently
## the backend matrix path filter, whose failure mode is invisible in CI: it
## emits an empty matrix, every job goes green, and the change ships to no
## image at all (see PR #10946). Plain `node --test`, no dependencies.
test-ci-scripts:
@set -e; for t in scripts/lib/*_test.mjs; do echo "== $$t"; node --test "$$t"; done
## Runs the core suite ($(TEST_PATHS)) with statement-coverage instrumentation
## and writes a merged profile to $(COVERAGE_PROFILE). Deliberately omits
## --fail-fast so a single failure doesn't truncate the coverage number, and
@@ -405,6 +419,23 @@ test-realtime: build-mock-backend
@echo 'Running realtime e2e tests (mock backend)'
$(GOCMD) run github.com/onsi/ginkgo/v2/ginkgo --label-filter="Realtime && !real-models" --flake-attempts $(TEST_FLAKES) -v -r ./tests/e2e
# Verify the realtime state-machine implementations conform to their formal
# designs (Go transition/rapid tests under -race + FizzBee model check of the
# authoritative specs). See docs/design/realtime-state-machines.md (Part 6) and
# docs/design/specs/README.md.
test-realtime-conformance:
GOCMD=$(GOCMD) ./scripts/realtime-conformance.sh
# Verify the shared model-loader shutdown behavior independently of any API
# modality (focused loader/gRPC/distributed/worker tests under -race + FizzBee).
test-model-lifecycle-conformance:
GOCMD=$(GOCMD) ./scripts/model-lifecycle-conformance.sh
# Install the pinned, checksum-verified FizzBee model checker (into .tools/,
# gitignored) used by the conformance targets. Idempotent; no-op if present.
install-fizzbee:
./scripts/install-fizzbee.sh
# Container-based real-model realtime testing. Build env vars / pipeline
# definition kept here so test-realtime-models-docker can drive a fully wired
# pipeline (VAD + STT + LLM + TTS) from inside a containerised runner.
@@ -553,6 +584,7 @@ prepare-test-extra: protogen-python
$(MAKE) -C backend/python/chatterbox
$(MAKE) -C backend/python/vllm
$(MAKE) -C backend/python/vllm-omni
$(MAKE) -C backend/python/longcat-video
$(MAKE) -C backend/python/sglang
$(MAKE) -C backend/python/vibevoice
$(MAKE) -C backend/python/liquid-audio
@@ -582,6 +614,7 @@ test-extra: prepare-test-extra
$(MAKE) -C backend/python/chatterbox test
$(MAKE) -C backend/python/vllm test
$(MAKE) -C backend/python/vllm-omni test
$(MAKE) -C backend/python/longcat-video test
$(MAKE) -C backend/python/vibevoice test
$(MAKE) -C backend/python/liquid-audio test
$(MAKE) -C backend/python/moonshine test
@@ -633,6 +666,9 @@ test-extra: prepare-test-extra
## suite against it.
##
BACKEND_TEST_MODEL_URL?=https://huggingface.co/Qwen/Qwen3-0.6B-GGUF/resolve/main/Qwen3-0.6B-Q8_0.gguf
## Suite timeout for `go test`. Wrappers whose model download alone can eat
## most of the default (multi-GB models on a slow HF CDN day) override this.
BACKEND_TEST_TIMEOUT?=30m
## Generic target — runs the suite against whatever BACKEND_IMAGE points at.
## Depends on protogen-go so pkg/grpc/proto is generated before `go test`.
@@ -660,7 +696,7 @@ test-extra-backend: protogen-go
BACKEND_TEST_FACE_IMAGE_3_URL="$$BACKEND_TEST_FACE_IMAGE_3_URL" \
BACKEND_TEST_FACE_IMAGE_3_FILE="$$BACKEND_TEST_FACE_IMAGE_3_FILE" \
BACKEND_TEST_VERIFY_DISTANCE_CEILING="$$BACKEND_TEST_VERIFY_DISTANCE_CEILING" \
go test -v -timeout 30m ./tests/e2e-backends/...
go test -v -timeout $(BACKEND_TEST_TIMEOUT) ./tests/e2e-backends/...
## Convenience wrappers: build the image, then exercise it.
test-extra-backend-llama-cpp: docker-build-llama-cpp
@@ -671,15 +707,6 @@ test-extra-backend-llama-cpp: docker-build-llama-cpp
test-extra-backend-ik-llama-cpp: docker-build-ik-llama-cpp
BACKEND_IMAGE=local-ai-backend:ik-llama-cpp $(MAKE) test-extra-backend
## llama-cpp-localai-paged: the LocalAI paged-attention llama.cpp variant. Same
## GGUF surface as stock llama-cpp (the paged engine is runtime-gated by the
## LLAMA_KV_PAGED env the grpc-server option hooks set), so the standard
## llama-cpp capability set is what we exercise here.
test-extra-backend-llama-cpp-localai-paged: docker-build-llama-cpp-localai-paged
BACKEND_IMAGE=local-ai-backend:llama-cpp-localai-paged \
BACKEND_TEST_CAPS=health,load,predict,stream,logprobs,logit_bias \
$(MAKE) test-extra-backend
## turboquant: exercises the llama.cpp-fork backend with the fork's
## *TurboQuant-specific* KV-cache types (turbo3 for both K and V). turbo3
## is what makes this backend distinct from stock llama-cpp — picking q8_0
@@ -692,6 +719,16 @@ test-extra-backend-turboquant: docker-build-turboquant
BACKEND_TEST_CACHE_TYPE_V=turbo3 \
$(MAKE) test-extra-backend
## bonsai: exercises the llama.cpp-fork backend with a real Q1_0 (1-bit) model —
## the PrismML Bonsai-8B GGUF, whose weight quant is *only* decodable by the fork's
## Q1_0 kernels. Loading it is what makes this backend distinct from stock llama-cpp;
## a standard-quant model would only test the upstream code path the llama-cpp backend
## already covers.
test-extra-backend-bonsai: docker-build-bonsai
BACKEND_IMAGE=local-ai-backend:bonsai \
BACKEND_TEST_MODEL_URL=https://huggingface.co/prism-ml/Bonsai-8B-gguf/resolve/main/Bonsai-8B-Q1_0.gguf \
$(MAKE) test-extra-backend
## Audio transcription wrapper for the llama-cpp backend.
## Drives the new AudioTranscription / AudioTranscriptionStream RPCs against
## ggml-org/Qwen3-ASR-0.6B-GGUF (a small ASR model that requires its mmproj
@@ -1009,6 +1046,7 @@ test-extra-backend-vibevoice-cpp-tts: docker-build-vibevoice-cpp
## post-image disk budget.
test-extra-backend-vibevoice-cpp-transcription: docker-build-vibevoice-cpp
BACKEND_IMAGE=local-ai-backend:vibevoice-cpp \
BACKEND_TEST_TIMEOUT=120m \
BACKEND_TEST_MODEL_URL='https://huggingface.co/mudler/vibevoice.cpp-models/resolve/main/vibevoice-asr-q4_k.gguf#vibevoice-asr-q4_k.gguf' \
BACKEND_TEST_EXTRA_FILES='https://huggingface.co/mudler/vibevoice.cpp-models/resolve/main/tokenizer.gguf#tokenizer.gguf' \
BACKEND_TEST_AUDIO_URL=https://github.com/ggml-org/whisper.cpp/raw/master/samples/jfk.wav \
@@ -1036,7 +1074,19 @@ test-extra-backend-whisper-transcription: docker-build-whisper
## is reachable.
test-extra-backend-parakeet-cpp-transcription: docker-build-parakeet-cpp
BACKEND_IMAGE=local-ai-backend:parakeet-cpp \
BACKEND_TEST_MODEL_URL=https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/tdt_ctc-110m-f16.gguf \
BACKEND_TEST_MODEL_URL=https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/realtime_eou_120m-v1-f16.gguf \
BACKEND_TEST_AUDIO_URL=https://github.com/ggml-org/whisper.cpp/raw/master/samples/jfk.wav \
BACKEND_TEST_CAPS=health,load,transcription \
$(MAKE) test-extra-backend
## Audio transcription wrapper for the moss-transcribe-cpp (moss-transcribe.cpp
## ggml port) backend. Mirrors test-extra-backend-parakeet-cpp-transcription:
## drives the AudioTranscription RPC against a published MOSS GGUF using the JFK
## 11s clip from whisper.cpp's CI samples. Not part of the default test suite -
## run explicitly once the pinned model URL is reachable.
test-extra-backend-moss-transcribe-cpp-transcription: docker-build-moss-transcribe-cpp
BACKEND_IMAGE=local-ai-backend:moss-transcribe-cpp \
BACKEND_TEST_MODEL_URL=https://huggingface.co/mudler/moss-transcribe.cpp-gguf/resolve/main/moss-transcribe-q5_k.gguf \
BACKEND_TEST_AUDIO_URL=https://github.com/ggml-org/whisper.cpp/raw/master/samples/jfk.wav \
BACKEND_TEST_CAPS=health,load,transcription \
$(MAKE) test-extra-backend
@@ -1190,10 +1240,10 @@ BACKEND_IK_LLAMA_CPP = ik-llama-cpp|ik-llama-cpp|.|false|false
# turboquant is a llama.cpp fork with TurboQuant KV-cache quantization.
# Reuses backend/cpp/llama-cpp grpc-server sources via a thin wrapper Makefile.
BACKEND_TURBOQUANT = turboquant|turboquant|.|false|false
# llama-cpp-localai-paged = stock llama.cpp grpc-server + the LocalAI paged-attention
# patch series (vendored in this wrapper backend). Reuses backend/cpp/llama-cpp sources via a thin
# wrapper Makefile (same upstream pin as stock llama-cpp; no fork, no patch-grpc-server).
BACKEND_LLAMA_CPP_LOCALAI_PAGED = llama-cpp-localai-paged|llama-cpp-localai-paged|.|false|false
# bonsai is a llama.cpp fork (PrismML) adding the Q1_0 (1-bit) and Q2_0 (ternary)
# weight-quant kernels the Bonsai / Ternary-Bonsai models ship in. Reuses
# backend/cpp/llama-cpp grpc-server sources via a thin wrapper Makefile.
BACKEND_BONSAI = bonsai|bonsai|.|false|false
# ds4 is antirez/ds4, a DeepSeek V4 Flash-specific inference engine.
# Single-model; hardware-only validation lives at tests/e2e-backends/
# (BACKEND_BINARY mode); see docs/superpowers/plans/2026-05-11-ds4-backend.md.
@@ -1213,10 +1263,12 @@ BACKEND_STABLEDIFFUSION_GGML = stablediffusion-ggml|golang|.|--progress=plain|tr
BACKEND_WHISPER = whisper|golang|.|false|true
BACKEND_CRISPASR = crispasr|golang|.|false|true
BACKEND_PARAKEET_CPP = parakeet-cpp|golang|.|false|true
BACKEND_MOSS_TRANSCRIBE_CPP = moss-transcribe-cpp|golang|.|false|true
BACKEND_DEPTH_ANYTHING_CPP = depth-anything-cpp|golang|.|false|true
BACKEND_VOXTRAL = voxtral|golang|.|false|true
BACKEND_ACESTEP_CPP = acestep-cpp|golang|.|false|true
BACKEND_QWEN3_TTS_CPP = qwen3-tts-cpp|golang|.|false|true
BACKEND_MOSS_TTS_CPP = moss-tts-cpp|golang|.|false|true
BACKEND_OMNIVOICE_CPP = omnivoice-cpp|golang|.|false|true
BACKEND_VIBEVOICE_CPP = vibevoice-cpp|golang|.|false|true
BACKEND_LOCALVQE = localvqe|golang|.|false|true
@@ -1238,6 +1290,7 @@ BACKEND_NEUTTS = neutts|python|.|false|true
BACKEND_KOKORO = kokoro|python|.|false|true
BACKEND_VLLM = vllm|python|.|false|true
BACKEND_VLLM_OMNI = vllm-omni|python|.|false|true
BACKEND_LONGCAT_VIDEO = longcat-video|python|.|--progress=plain|true
BACKEND_SGLANG = sglang|python|.|false|true
BACKEND_DIFFUSERS = diffusers|python|.|--progress=plain|true
BACKEND_CHATTERBOX = chatterbox|python|.|false|true
@@ -1295,7 +1348,7 @@ endef
$(eval $(call generate-docker-build-target,$(BACKEND_LLAMA_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_IK_LLAMA_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_TURBOQUANT)))
$(eval $(call generate-docker-build-target,$(BACKEND_LLAMA_CPP_LOCALAI_PAGED)))
$(eval $(call generate-docker-build-target,$(BACKEND_BONSAI)))
$(eval $(call generate-docker-build-target,$(BACKEND_DS4)))
$(eval $(call generate-docker-build-target,$(BACKEND_PRIVACY_FILTER)))
$(eval $(call generate-docker-build-target,$(BACKEND_PIPER)))
@@ -1307,6 +1360,7 @@ $(eval $(call generate-docker-build-target,$(BACKEND_STABLEDIFFUSION_GGML)))
$(eval $(call generate-docker-build-target,$(BACKEND_WHISPER)))
$(eval $(call generate-docker-build-target,$(BACKEND_CRISPASR)))
$(eval $(call generate-docker-build-target,$(BACKEND_PARAKEET_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_MOSS_TRANSCRIBE_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_DEPTH_ANYTHING_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_VOXTRAL)))
$(eval $(call generate-docker-build-target,$(BACKEND_OPUS)))
@@ -1323,6 +1377,7 @@ $(eval $(call generate-docker-build-target,$(BACKEND_NEUTTS)))
$(eval $(call generate-docker-build-target,$(BACKEND_KOKORO)))
$(eval $(call generate-docker-build-target,$(BACKEND_VLLM)))
$(eval $(call generate-docker-build-target,$(BACKEND_VLLM_OMNI)))
$(eval $(call generate-docker-build-target,$(BACKEND_LONGCAT_VIDEO)))
$(eval $(call generate-docker-build-target,$(BACKEND_SGLANG)))
$(eval $(call generate-docker-build-target,$(BACKEND_DIFFUSERS)))
$(eval $(call generate-docker-build-target,$(BACKEND_CHATTERBOX)))
@@ -1340,6 +1395,7 @@ $(eval $(call generate-docker-build-target,$(BACKEND_WHISPERX)))
$(eval $(call generate-docker-build-target,$(BACKEND_ACE_STEP)))
$(eval $(call generate-docker-build-target,$(BACKEND_ACESTEP_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_QWEN3_TTS_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_MOSS_TTS_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_OMNIVOICE_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_VIBEVOICE_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_LOCALVQE)))
@@ -1359,7 +1415,7 @@ $(eval $(call generate-docker-build-target,$(BACKEND_SUPERTONIC)))
docker-save-%: backend-images
docker save local-ai-backend:$* -o backend-images/$*.tar
docker-build-backends: docker-build-llama-cpp docker-build-ik-llama-cpp docker-build-turboquant docker-build-llama-cpp-localai-paged docker-build-ds4 docker-build-rerankers docker-build-vllm docker-build-vllm-omni docker-build-sglang docker-build-transformers docker-build-outetts docker-build-diffusers docker-build-kokoro docker-build-faster-whisper docker-build-crispasr docker-build-coqui docker-build-chatterbox docker-build-vibevoice docker-build-liquid-audio docker-build-moonshine docker-build-pocket-tts docker-build-qwen-tts docker-build-fish-speech docker-build-faster-qwen3-tts docker-build-qwen-asr docker-build-nemo docker-build-voxcpm docker-build-whisperx docker-build-ace-step docker-build-acestep-cpp docker-build-voxtral docker-build-mlx-distributed docker-build-trl docker-build-llama-cpp-quantization docker-build-tinygrad docker-build-kokoros docker-build-sam3-cpp docker-build-rfdetr-cpp docker-build-qwen3-tts-cpp docker-build-omnivoice-cpp docker-build-vibevoice-cpp docker-build-localvqe docker-build-insightface docker-build-speaker-recognition docker-build-sherpa-onnx docker-build-cloud-proxy docker-build-supertonic docker-build-depth-anything-cpp docker-build-privacy-filter
docker-build-backends: docker-build-llama-cpp docker-build-ik-llama-cpp docker-build-turboquant docker-build-bonsai docker-build-ds4 docker-build-rerankers docker-build-vllm docker-build-vllm-omni docker-build-longcat-video docker-build-sglang docker-build-transformers docker-build-outetts docker-build-diffusers docker-build-kokoro docker-build-faster-whisper docker-build-crispasr docker-build-coqui docker-build-chatterbox docker-build-vibevoice docker-build-liquid-audio docker-build-moonshine docker-build-pocket-tts docker-build-qwen-tts docker-build-fish-speech docker-build-faster-qwen3-tts docker-build-qwen-asr docker-build-nemo docker-build-voxcpm docker-build-whisperx docker-build-ace-step docker-build-acestep-cpp docker-build-voxtral docker-build-mlx-distributed docker-build-trl docker-build-llama-cpp-quantization docker-build-tinygrad docker-build-kokoros docker-build-sam3-cpp docker-build-rfdetr-cpp docker-build-qwen3-tts-cpp docker-build-moss-tts-cpp docker-build-omnivoice-cpp docker-build-vibevoice-cpp docker-build-localvqe docker-build-insightface docker-build-speaker-recognition docker-build-sherpa-onnx docker-build-cloud-proxy docker-build-supertonic docker-build-depth-anything-cpp docker-build-moss-transcribe-cpp docker-build-privacy-filter
########################################################
### Mock Backend for E2E Tests
@@ -1484,8 +1540,13 @@ build-launcher-darwin:
mv cmd/launcher/LocalAI.app dist/LocalAI.app
bash contrib/macos/sign-and-notarize.sh sign dist/LocalAI.app
# Wrap the (signed) app into a drag-to-Applications DMG via hdiutil, then sign the DMG.
# Notarize + staple the .app itself, then wrap it into a drag-to-Applications
# DMG via hdiutil and sign the DMG. The app is stapled BEFORE packaging so the
# bundle carries its own ticket and verifies offline (a dmg-only staple leaves
# the app relying on an online Gatekeeper check, which fails offline / once the
# app is copied out of the dmg). No-op without notary secrets.
dmg-launcher-darwin: build-launcher-darwin
bash contrib/macos/sign-and-notarize.sh notarize-app dist/LocalAI.app
rm -rf dist/dmg dist/LocalAI.dmg
mkdir -p dist/dmg
cp -R dist/LocalAI.app dist/dmg/LocalAI.app
@@ -1497,7 +1558,7 @@ dmg-launcher-darwin: build-launcher-darwin
notarize-launcher-darwin: dmg-launcher-darwin
bash contrib/macos/sign-and-notarize.sh notarize dist/LocalAI.dmg
# Single entrypoint for CI: build -> sign app -> dmg -> sign dmg -> notarize -> staple.
# Single entrypoint for CI: build -> sign app -> notarize+staple app -> dmg -> sign dmg -> notarize+staple dmg.
release-launcher-darwin: notarize-launcher-darwin
@echo "dist/LocalAI.dmg is ready"

View File

@@ -177,7 +177,7 @@ For more details, see the [Getting Started guide](https://localai.io/basics/gett
## Latest News
- **June 2026**: New native biometric backends from the LocalAI team: [voice-detect.cpp](https://github.com/mudler/voice-detect.cpp) for speaker recognition and voice analysis (ECAPA-TDNN, WeSpeaker, ERes2Net, CAM++, wav2vec2 age/gender/emotion) and [face-detect.cpp](https://github.com/mudler/face-detect.cpp) for face detection, recognition, demographics and anti-spoofing (SCRFD/ArcFace, YuNet/SFace). Both are from-scratch C++/ggml engines with no Python or onnxruntime at inference, self-contained GGUF weights, bit-exact parity with the reference, and GPU cuDNN parity, replacing the heavier Python `insightface` and `speaker-recognition` backends ([PR #10441](https://github.com/mudler/LocalAI/pull/10441)).
- **June 2026**: New native biometric backends from the LocalAI team: [voice-detect.cpp](https://github.com/localai-org/voice-detect.cpp) for speaker recognition and voice analysis (ECAPA-TDNN, WeSpeaker, ERes2Net, CAM++, wav2vec2 age/gender/emotion) and [face-detect.cpp](https://github.com/mudler/face-detect.cpp) for face detection, recognition, demographics and anti-spoofing (SCRFD/ArcFace, YuNet/SFace). Both are from-scratch C++/ggml engines with no Python or onnxruntime at inference, self-contained GGUF weights, bit-exact parity with the reference, and GPU cuDNN parity, replacing the heavier Python `insightface` and `speaker-recognition` backends ([PR #10441](https://github.com/mudler/LocalAI/pull/10441)).
- **June 2026**: New [realtime voice assistant demo](https://github.com/localai-org/localai-realtime-demo) (a tiny Go client for the Realtime API with a full talk-back voice loop and tool calling), plus [streaming of the realtime LLM / TTS / transcription pipeline stages](https://github.com/mudler/LocalAI/pull/10176) and [configurable WebRTC ICE candidates](https://github.com/mudler/LocalAI/pull/10231).
- **June 2026**: Big speech push: the [parakeet.cpp](https://github.com/mudler/parakeet.cpp) ASR engine gains [NeMo-faithful segment timestamps](https://github.com/mudler/LocalAI/pull/10207), a [multilingual streaming Nemotron-3.5 model](https://github.com/mudler/LocalAI/pull/10199), [dynamic batching for concurrent transcription](https://github.com/mudler/LocalAI/pull/10112) and [CUDA graphs](https://github.com/mudler/LocalAI/pull/10273); the new [CrispASR backend](https://github.com/mudler/LocalAI/pull/10099) adds multi-architecture ASR + TTS, and [60 Piper TTS voices across 42 languages](https://github.com/mudler/LocalAI/pull/10296) land in the gallery (plus [per-request TTS instructions and params](https://github.com/mudler/LocalAI/pull/10172)).
- **June 2026**: New backends and models: [locate-anything.cpp](https://github.com/mudler/LocalAI/pull/10264) for open-vocabulary object detection via ggml, [Ideogram4 image generation](https://github.com/mudler/LocalAI/pull/10201) in stablediffusion-ggml, [llama.cpp video input](https://github.com/mudler/LocalAI/pull/10216), and the [Gemma 4 QAT family with MTP speculative-decoding pairs](https://github.com/mudler/LocalAI/pull/10215). Plus an [interactive CLI chat mode](https://github.com/mudler/LocalAI/pull/10226) and [RAG source citations in agent responses](https://github.com/mudler/LocalAI/pull/10228).
@@ -232,12 +232,17 @@ Most backends wrap a best-in-class upstream engine. A handful of them are native
| Backend | What it does |
|---------|-------------|
| [parakeet.cpp](https://github.com/mudler/parakeet.cpp) | C++/GGML port of NVIDIA NeMo Parakeet ASR (tdt/ctc/rnnt/hybrid), with cache-aware streaming transcription |
| [ced.cpp](https://github.com/mudler/ced.cpp) | C++/GGML port of the CED audio-tagging models: sound-event classification (527-class AudioSet) over REST and the realtime API for live recognition |
| [voxtral.c](https://github.com/mudler/voxtral.c) | Voxtral Realtime 4B speech-to-text in pure C |
| [moss-transcribe.cpp](https://github.com/localai-org/moss-transcribe.cpp) | C++/GGML port of OpenMOSS MOSS-Transcribe-Diarize: joint long-form transcription, speaker diarization and timestamping in a single pass |
| [moss-tts.cpp](https://github.com/mudler/moss-tts.cpp) | C++/GGML port of the OpenMOSS MOSS-TTS family: text-to-speech (MOSS-TTS-Local v1.5, 48 kHz stereo) with reference-audio voice cloning, through the MOSS-Audio-Tokenizer neural codec |
| [ced.cpp](https://github.com/localai-org/ced.cpp) | C++/GGML port of the CED audio-tagging models: sound-event classification (527-class AudioSet) over REST and the realtime API for live recognition |
| [voice-detect.cpp](https://github.com/localai-org/voice-detect.cpp) | Speaker recognition and voice analysis (ECAPA-TDNN, WeSpeaker, ERes2Net, CAM++, wav2vec2 age/gender/emotion), replacing the Python speaker-recognition backend |
| [voxtral-tts.c](https://github.com/mudler/voxtral-tts.c) | Voxtral Realtime 4B speech-to-text in pure C |
| [vibevoice.cpp](https://github.com/mudler/vibevoice.cpp) | Native port of Microsoft VibeVoice for TTS (voice cloning) and long-form ASR with speaker diarization |
| [rf-detr.cpp](https://github.com/mudler/rf-detr.cpp) | Native RF-DETR object detection and instance segmentation |
| [rf-detr.cpp](https://github.com/localai-org/rf-detr.cpp) | Native RF-DETR object detection and instance segmentation |
| [locate-anything.cpp](https://github.com/mudler/locate-anything.cpp) | Open-vocabulary object detection and visual grounding (LocateAnything-3B) |
| [depth-anything.cpp](https://github.com/mudler/depth-anything.cpp) | Depth Anything 3 monocular metric depth + camera pose estimation |
| [face-detect.cpp](https://github.com/mudler/face-detect.cpp) | Face detection, recognition, demographics and anti-spoofing (SCRFD/ArcFace, YuNet/SFace), replacing the Python insightface backend |
| [free-splatter.cpp](https://github.com/localai-org/free-splatter.cpp) | Pose-free 3D reconstruction (FreeSplatter): turns a handful of plain photos into 3D Gaussians, no camera poses or GPU required |
| [privacy-filter.cpp](https://github.com/localai-org/privacy-filter.cpp) | Standalone GGML PII/NER token-classification engine powering LocalAI's PII redaction tier |
| [LocalVQE](https://github.com/localai-org/LocalVQE) | Joint acoustic echo cancellation, noise suppression, and dereverberation |
| [local-store](https://github.com/mudler/LocalAI) | Local-first vector database for embeddings (shipped in-tree) |

View File

@@ -7,7 +7,7 @@ ARG BUILDER_BASE_IMAGE=${BASE_IMAGE}
# BUILDER_TARGET selects which builder stage the final scratch image copies
# package output from. Declared at global scope (before any FROM) so it's
# usable in `FROM ${BUILDER_TARGET}` below. Default keeps local
# `make backends/llama-cpp-localai-paged` on the from-source path.
# `make backends/bonsai` on the from-source path.
ARG BUILDER_TARGET=builder-fromsource
ARG APT_MIRROR=""
ARG APT_PORTS_MIRROR=""
@@ -18,7 +18,7 @@ ARG APT_PORTS_MIRROR=""
# Runs .docker/install-base-deps.sh (apt deps + cmake + protoc + gRPC +
# conditional CUDA/ROCm/Vulkan), copies /opt/grpc to /usr/local, then
# compiles the variant. Used when BUILDER_TARGET=builder-fromsource (the
# default; local `make backends/llama-cpp-localai-paged`).
# default; local `make backends/bonsai`).
#
# The install script is the same one that backend/Dockerfile.base-grpc-builder
# runs, so the result is bit-equivalent to the prebuilt-base path
@@ -84,22 +84,21 @@ RUN cp -a /opt/grpc/. /usr/local/
COPY . /LocalAI
# BuildKit cache mount for ccache. See Dockerfile.llama-cpp (commit 9228e5b4)
# for rationale. llama-cpp-localai-paged is the SAME upstream llama.cpp with
# the LocalAI paged patch series applied; it reuses backend/cpp/llama-cpp
# source via a thin wrapper Makefile, so MOST TUs are content-identical to the
# stock llama-cpp build. Sharing a cache id with llama-cpp could give
# cross-variant hits — but for now keep them separate (mirroring turboquant) so
# a regression in one doesn't poison the other. Revisit sharing after measuring
# the actual hit rate.
# for rationale. bonsai is a llama.cpp fork that reuses
# backend/cpp/llama-cpp source via a thin wrapper Makefile, so MOST TUs
# are content-identical to the upstream llama-cpp build. Sharing a cache
# id with llama-cpp could give cross-fork hits — but for now keep them
# separate so a regression in one doesn't poison the other. Revisit
# sharing after measuring the actual hit rate.
#
# The compile body is shared with builder-prebuilt via .docker/llama-cpp-localai-paged-compile.sh.
RUN --mount=type=bind,source=.docker/llama-cpp-localai-paged-compile.sh,target=/usr/local/sbin/compile.sh \
--mount=type=cache,target=/root/.ccache,id=llama-cpp-localai-paged-ccache-${TARGETARCH}-${BUILD_TYPE},sharing=locked \
# The compile body is shared with builder-prebuilt via .docker/bonsai-compile.sh.
RUN --mount=type=bind,source=.docker/bonsai-compile.sh,target=/usr/local/sbin/compile.sh \
--mount=type=cache,target=/root/.ccache,id=bonsai-ccache-${TARGETARCH}-${BUILD_TYPE},sharing=locked \
bash /usr/local/sbin/compile.sh
# Copy libraries using a script to handle architecture differences
RUN make -BC /LocalAI/backend/cpp/llama-cpp-localai-paged package
RUN make -BC /LocalAI/backend/cpp/bonsai package
# ============================================================================
@@ -108,9 +107,7 @@ RUN make -BC /LocalAI/backend/cpp/llama-cpp-localai-paged package
# That image already has gRPC at /opt/grpc + apt deps + CUDA/ROCm/Vulkan
# pre-installed, so we just copy gRPC to /usr/local and compile. Used when
# BUILDER_TARGET=builder-prebuilt (CI when the matrix entry sets
# builder-base-image). llama-cpp-localai-paged reuses the SAME base-grpc-* tags
# as the stock llama-cpp backend (same gRPC + same toolchain), so no new
# base-images.yml variant is required.
# builder-base-image).
# ============================================================================
FROM ${BUILDER_BASE_IMAGE} AS builder-prebuilt
@@ -121,9 +118,9 @@ ENV CUDA_DOCKER_ARCH=${CUDA_DOCKER_ARCH}
ARG CMAKE_ARGS
ENV CMAKE_ARGS=${CMAKE_ARGS}
# AMDGPU_TARGETS must be forwarded into the env here too — backend/cpp/llama-cpp/Makefile
# (which the llama-cpp-localai-paged Makefile reuses via a sibling build dir) errors out
# when the var is empty on a hipblas build, and the prebuilt path is what CI exercises most
# of the time. The builder-fromsource stage above already does this; mirror it here.
# (which the bonsai Makefile reuses via a sibling build dir) errors out when the var
# is empty on a hipblas build, and the prebuilt path is what CI exercises most of the
# time. The builder-fromsource stage above already does this; mirror it here.
ARG AMDGPU_TARGETS
ENV AMDGPU_TARGETS=${AMDGPU_TARGETS}
ARG TARGETARCH
@@ -136,11 +133,11 @@ RUN cp -a /opt/grpc/. /usr/local/
COPY . /LocalAI
RUN --mount=type=bind,source=.docker/llama-cpp-localai-paged-compile.sh,target=/usr/local/sbin/compile.sh \
--mount=type=cache,target=/root/.ccache,id=llama-cpp-localai-paged-ccache-${TARGETARCH}-${BUILD_TYPE},sharing=locked \
RUN --mount=type=bind,source=.docker/bonsai-compile.sh,target=/usr/local/sbin/compile.sh \
--mount=type=cache,target=/root/.ccache,id=bonsai-ccache-${TARGETARCH}-${BUILD_TYPE},sharing=locked \
bash /usr/local/sbin/compile.sh
RUN make -BC /LocalAI/backend/cpp/llama-cpp-localai-paged package
RUN make -BC /LocalAI/backend/cpp/bonsai package
# ============================================================================
@@ -160,4 +157,4 @@ FROM scratch
# Copy all available binaries (the build process only creates the appropriate ones for the target architecture)
COPY --from=builder /LocalAI/backend/cpp/llama-cpp-localai-paged/package/. ./
COPY --from=builder /LocalAI/backend/cpp/bonsai/package/. ./

View File

@@ -224,7 +224,11 @@ ARG DEPS_REFRESH=initial
RUN cd /${BACKEND} && PORTABLE_PYTHON=true make
# Package GPU libraries into the backend's lib directory
# Package GPU libraries into the backend's lib directory.
#
# Must stay after the venv is built above: package-gpu-libs.sh inspects
# /${BACKEND}/venv to decide whether this backend already carries a complete
# cuDNN from pip, and bundles one only when it does not (issue #10905).
RUN mkdir -p /${BACKEND}/lib && \
TARGET_LIB_DIR="/${BACKEND}/lib" BUILD_TYPE="${BUILD_TYPE}" CUDA_MAJOR_VERSION="${CUDA_MAJOR_VERSION}" \
bash /package-gpu-libs.sh "/${BACKEND}/lib"

View File

@@ -46,6 +46,7 @@ The backend system provides language-specific Dockerfiles that handle the build
- **vllm**: High-performance LLM inference
- **mlx**: Apple Silicon optimization
- **diffusers**: Stable Diffusion models
- **longcat-video**: CUDA text/image-to-video and speech-driven avatar generation
- **Audio**: coqui, faster-whisper, kitten-tts
- **Vision**: mlx-vlm, rfdetr
- **Specialized**: rerankers, chatterbox, kokoro

View File

@@ -18,6 +18,18 @@ service Backend {
rpc GenerateVideo(GenerateVideoRequest) returns (Result) {}
rpc AudioTranscription(TranscriptRequest) returns (TranscriptResult) {}
rpc AudioTranscriptionStream(TranscriptRequest) returns (stream TranscriptStreamResponse) {}
// AudioTranscriptionLive is the bidirectional live-microphone ASR RPC. The
// first message MUST carry a Config; subsequent messages carry Audio frames
// (mono float PCM at config.sample_rate, 16 kHz default). After a
// successful open the backend replies with a single ready ack
// (TranscriptLiveResponse{ready:true}); backends or models without
// cache-aware streaming support return UNIMPLEMENTED instead. Newly
// finalized text streams back as deltas; eou=true marks the model's
// end-of-utterance token. One stream spans many utterances (the decoder
// resets itself after each EOU). Closing the send side finalizes: the
// backend flushes the decoder tail and emits a terminal message carrying
// final_result. A second Config mid-stream resets the decode session.
rpc AudioTranscriptionLive(stream TranscriptLiveRequest) returns (stream TranscriptLiveResponse) {}
rpc TTS(TTSRequest) returns (Result) {}
rpc TTSStream(TTSRequest) returns (stream Reply) {}
rpc SoundGeneration(SoundGenerationRequest) returns (Result) {}
@@ -124,6 +136,10 @@ message MetricsResponse {
message TokenClassifyRequest {
string text = 1;
float threshold = 2;
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 3;
}
// TokenClassifyEntity is one detected entity span. Byte offsets are
@@ -161,6 +177,10 @@ message ScoreRequest {
// candidates differ in length and the consumer wants a per-token
// measure comparable across them (PMI-style scoring).
bool length_normalize = 4;
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 5;
}
// CandidateScore is one row in the ScoreResponse, matching by index
@@ -192,6 +212,10 @@ message RerankRequest {
string query = 1;
repeated string documents = 2;
int32 top_n = 3;
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 4;
}
message RerankResult {
@@ -303,6 +327,39 @@ message PredictOptions {
int32 TopLogprobs = 51; // Number of top logprobs to return per token (maps to OpenAI top_logprobs parameter)
map<string, string> Metadata = 52; // Generic per-request metadata (e.g., enable_thinking)
float MinP = 53; // Minimum probability sampling threshold (0.0 = disabled)
// ModelIdentity names the model this request is for, so a backend can reject
// a request that reached it by mistake instead of answering from whatever
// model it happens to hold. In distributed mode a worker can recycle a
// stopped backend's gRPC port for a different model's backend, and a
// liveness-only health probe cannot tell that apart from a valid cached
// route (#10952).
//
// The value is the controller's ModelConfig.Model, the SAME expression that
// produces ModelOptions.Model at LoadModel time, so the two are equal by
// construction rather than by convention.
//
// Empty means "no identity supplied": backends MUST skip the check. That
// keeps an old controller talking to a new backend working, and covers
// callers that legitimately synthesize a PredictOptions internally.
//
// Do NOT reuse TTSRequest.model or SoundGenerationRequest.model for this
// purpose. FileStagingClient already rewrites those to worker-local absolute
// paths (core/services/nodes/file_staging_client.go), so in distributed mode
// they already differ from the load-time value and comparing them would
// reject valid requests. Extending identity to those RPCs needs a separate
// field carrying the untranslated value - which is exactly what
// TTSRequest.ModelIdentity and SoundGenerationRequest.ModelIdentity are.
//
// Every other request message that reaches a backend through the distributed
// router now carries the same ModelIdentity field, populated from the same
// ModelConfig.Model. FileStagingClient rewrites Src/Dst/Voice/Model/
// StartImage/EndImage/Audio and never ModelIdentity, so what the backend
// compares is always what the controller sent.
string ModelIdentity = 54;
// 24 was never assigned; reserve it so it is not silently reused.
reserved 24;
}
// ToolCallDelta represents an incremental tool call update from the C++ parser.
@@ -472,6 +529,10 @@ message TranscriptRequest {
float temperature = 8;
repeated string timestamp_granularities = 9;
bool stream = 10;
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 11;
}
message TranscriptResult {
@@ -479,6 +540,10 @@ message TranscriptResult {
string text = 2;
string language = 3;
float duration = 4;
// True when the decode ended on the model's end-of-utterance special token
// (<EOU>/<EOB>, emitted by cache-aware streaming models such as
// parakeet_realtime_eou_120m-v1). The marker itself is stripped from text.
bool eou = 5;
}
message TranscriptStreamResponse {
@@ -486,6 +551,34 @@ message TranscriptStreamResponse {
TranscriptResult final_result = 2;
}
// === AudioTranscriptionLive messages =====================================
message TranscriptLiveRequest {
oneof payload {
TranscriptLiveConfig config = 1;
TranscriptLiveAudio audio = 2;
}
}
message TranscriptLiveConfig {
string language = 1; // "" => model default
int32 sample_rate = 2; // 0 => 16000; backends may reject others
map<string, string> params = 3; // backend-specific tuning
}
message TranscriptLiveAudio {
repeated float pcm = 1; // mono PCM in [-1,1] at config.sample_rate
}
message TranscriptLiveResponse {
bool ready = 1; // open ack: sent once, before any delta
string delta = 2; // newly-finalized text since previous response
bool eou = 3; // <EOU> fired during this feed (the user yielded the turn)
repeated TranscriptWord words = 4; // words finalized by this feed (stream-relative ns)
TranscriptResult final_result = 5; // terminal message only, after the send side closes
bool eob = 6; // <EOB> fired: a backchannel ("uh-huh") ended — NOT a turn boundary
}
message TranscriptWord {
int64 start = 1;
int64 end = 2;
@@ -518,6 +611,10 @@ message GenerateImageRequest {
// Reference images for models that support them (e.g., Flux Kontext)
repeated string ref_images = 12;
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 13;
}
message GenerateVideoRequest {
@@ -533,6 +630,14 @@ message GenerateVideoRequest {
float cfg_scale = 10; // Classifier-free guidance scale
int32 step = 11; // Number of inference steps
string dst = 12; // Output path for the generated video
string audio = 13; // Path to staged audio for audio-conditioned video
// Backend-specific per-request generation parameters. Values are strings
// and are validated/coerced by the selected backend.
map<string, string> params = 14;
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 15;
}
message TTSRequest {
@@ -550,10 +655,26 @@ message TTSRequest {
// (e.g. Chatterbox exaggeration/cfg_weight/temperature). Values are strings and
// coerced by the backend; unset leaves the backend's configured defaults.
map<string, string> params = 7;
// ModelIdentity is a SEPARATE field from `model` above and carries the
// UNTRANSLATED controller-side ModelConfig.Model, so a backend can reject a
// request that reached it through a stale distributed route (#10952).
//
// `model` cannot be reused for this: FileStagingClient.TTS/.TTSStream and the
// SoundGeneration path rewrite it into a worker-local absolute path
// (core/services/nodes/file_staging_client.go), while the load-time value is
// untranslated. In distributed mode - exactly the configuration this guards -
// the two already differ, so comparing them would reject valid requests.
//
// Empty means "no identity supplied" and backends MUST skip the check.
string ModelIdentity = 8;
}
message VADRequest {
repeated float audio = 1;
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 2;
}
message VADSegment {
@@ -585,6 +706,10 @@ message DiarizeRequest {
float min_duration_on = 8; // discard segments shorter than this (seconds); 0 = backend default
float min_duration_off = 9; // merge gaps shorter than this (seconds); 0 = backend default
bool include_text = 10; // when the backend can emit per-segment transcript for free, ask it to populate `text`
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 11;
}
message DiarizeSegment {
@@ -619,6 +744,18 @@ message SoundGenerationRequest {
optional string language = 14;
optional string timesignature = 15;
optional bool instrumental = 17;
// ModelIdentity is a SEPARATE field from `model` above and carries the
// UNTRANSLATED controller-side ModelConfig.Model, so a backend can reject a
// request that reached it through a stale distributed route (#10952).
//
// `model` cannot be reused for this: FileStagingClient.TTS/.TTSStream and the
// SoundGeneration path rewrite it into a worker-local absolute path
// (core/services/nodes/file_staging_client.go), while the load-time value is
// untranslated. In distributed mode - exactly the configuration this guards -
// the two already differ, so comparing them would reject valid requests.
//
// Empty means "no identity supplied" and backends MUST skip the check.
string ModelIdentity = 18;
}
message TokenizationResponse {
@@ -658,6 +795,10 @@ message DetectOptions {
repeated float points = 3; // Point coordinates as [x1, y1, label1, x2, y2, label2, ...] (label: 1=pos, 0=neg)
repeated float boxes = 4; // Box coordinates as [x1, y1, x2, y2, ...]
float threshold = 5; // Detection confidence threshold
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 6;
}
message Detection {
@@ -680,6 +821,10 @@ message SoundDetectionRequest {
string src = 1; // audio file path (LocalAI writes the upload to disk)
int32 top_k = 2; // number of top tags to return (0 = all classes)
float threshold = 3; // optional: drop tags scoring below this
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 4;
}
message SoundClass {
@@ -704,6 +849,10 @@ message DepthRequest {
bool include_points = 7; // back-project to a 3D point cloud (DualDPT)
float points_conf_thresh = 8; // keep points with confidence >= this threshold
repeated string exports = 9; // requested exports: "glb", "colmap"
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 10;
}
message DepthResponse {
@@ -735,6 +884,10 @@ message FaceVerifyRequest {
string img2 = 2; // base64-encoded image
float threshold = 3; // cosine-distance threshold; 0 = use backend default
bool anti_spoofing = 4; // run MiniFASNet liveness on each image; failed liveness forces verified=false
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 5;
}
message FaceVerifyResponse {
@@ -756,6 +909,10 @@ message FaceAnalyzeRequest {
string img = 1; // base64-encoded image
repeated string actions = 2; // subset of ["age","gender","emotion","race"]; empty = all-supported
bool anti_spoofing = 3;
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 4;
}
message FaceAnalysis {
@@ -788,6 +945,10 @@ message VoiceVerifyRequest {
string audio2 = 2; // path to second audio clip
float threshold = 3; // cosine-distance threshold; 0 = use backend default
bool anti_spoofing = 4; // reserved for future AASIST bolt-on
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 5;
}
message VoiceVerifyResponse {
@@ -802,6 +963,10 @@ message VoiceVerifyResponse {
message VoiceAnalyzeRequest {
string audio = 1; // path to audio clip
repeated string actions = 2; // subset of ["age","gender","emotion"]; empty = all-supported
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 3;
}
message VoiceAnalysis {
@@ -820,6 +985,10 @@ message VoiceAnalyzeResponse {
message VoiceEmbedRequest {
string audio = 1; // path to audio clip
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 2;
}
message VoiceEmbedResponse {
@@ -914,6 +1083,10 @@ message AudioTransformRequest {
string reference_path = 2; // optional auxiliary; empty => zero-fill
string dst = 3; // required, output file path
map<string, string> params = 4; // backend-specific tuning
// ModelIdentity names the model this request is for; see
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
// identity supplied" and backends MUST skip the check.
string ModelIdentity = 5;
}
message AudioTransformResult {
@@ -1212,4 +1385,3 @@ message ForwardReply {
repeated ForwardHeader headers = 2;
bytes body_chunk = 3;
}

105
backend/cpp/bonsai/Makefile Normal file
View File

@@ -0,0 +1,105 @@
# Pinned to the HEAD of the `prism` branch on https://github.com/PrismML-Eng/llama.cpp.
# Auto-bumped nightly by .github/workflows/bump_deps.yaml.
BONSAI_VERSION?=7529fdaaf99ffdc5ca71ace9c7409a56b27ad92f
LLAMA_REPO?=https://github.com/PrismML-Eng/llama.cpp
CMAKE_ARGS?=
BUILD_TYPE?=
NATIVE?=false
ONEAPI_VARS?=/opt/intel/oneapi/setvars.sh
TARGET?=--target grpc-server
JOBS?=$(shell nproc 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 1)
ARCH?=$(shell uname -m)
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))
LLAMA_CPP_DIR := $(CURRENT_MAKEFILE_DIR)/../llama-cpp
GREEN := \033[0;32m
RESET := \033[0m
# bonsai is a llama.cpp fork (PrismML) adding the Q1_0 (1-bit) and Q2_0 (ternary)
# weight-quantization kernels that the Bonsai / Ternary-Bonsai models ship in. Rather
# than duplicating grpc-server.cpp / CMakeLists.txt / prepare.sh we reuse the ones in
# backend/cpp/llama-cpp, and only swap which repo+sha the fetch step pulls. Each flavor
# target copies ../llama-cpp into a sibling ../bonsai-<flavor>-build directory, then
# invokes llama-cpp's own build with LLAMA_REPO/LLAMA_VERSION overridden to point at the
# fork.
#
# The Q1_0/Q2_0 additions are model *weight* types decoded inside libllama, transparent
# to the reused gRPC server, so (unlike turboquant's KV-cache types) no grpc-server.cpp
# allow-list patch is needed. The fork branched from upstream before a few API changes
# the shared grpc-server.cpp depends on; those are carried as patch files under
# backend/cpp/bonsai/patches/ and applied to the cloned fork by apply-patches.sh.
PATCHES_DIR := $(CURRENT_MAKEFILE_DIR)/patches
define bonsai-build
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build
cp -rf $(LLAMA_CPP_DIR) $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build
# Drop patches vendored for upstream llama.cpp: the fork tree diverges, so
# they reject there. Fork-specific patches live in backend/cpp/bonsai/patches/
# and are applied by apply-patches.sh below.
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build/patches
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build purge
$(info $(GREEN)I bonsai build info:$(1)$(RESET))
LLAMA_REPO=$(LLAMA_REPO) LLAMA_VERSION=$(BONSAI_VERSION) \
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build llama.cpp
bash $(CURRENT_MAKEFILE_DIR)/apply-patches.sh $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build/llama.cpp $(PATCHES_DIR)
CMAKE_ARGS="$(CMAKE_ARGS) $(2)" TARGET="$(3)" \
LLAMA_REPO=$(LLAMA_REPO) LLAMA_VERSION=$(BONSAI_VERSION) \
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build grpc-server
cp -rfv $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build/grpc-server bonsai-$(1)
endef
bonsai-avx2:
$(call bonsai-build,avx2,-DGGML_AVX=on -DGGML_AVX2=on -DGGML_AVX512=off -DGGML_FMA=on -DGGML_F16C=on,--target grpc-server)
bonsai-avx512:
$(call bonsai-build,avx512,-DGGML_AVX=on -DGGML_AVX2=off -DGGML_AVX512=on -DGGML_FMA=on -DGGML_F16C=on,--target grpc-server)
bonsai-avx:
$(call bonsai-build,avx,-DGGML_AVX=on -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server)
bonsai-fallback:
$(call bonsai-build,fallback,-DGGML_AVX=off -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server)
# Single-build CPU backend via ggml CPU_ALL_VARIANTS (mirrors llama-cpp-cpu-all).
# bonsai reuses backend/cpp/llama-cpp's CMakeLists.txt (hw_grpc_proto STATIC) and
# Makefile (SHARED_LIBS make-var + EXTRA_CMAKE_ARGS), so this passes the same overrides
# through to the copied build: SHARED_LIBS=ON, the DL flags, and --target ggml (which
# pulls in the per-microarch libggml-cpu-*.so via ggml's add_dependencies). The .so set
# is collected for package.sh to bundle into package/lib.
bonsai-cpu-all:
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build
cp -rf $(LLAMA_CPP_DIR) $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build
# Drop patches vendored for upstream llama.cpp: the fork tree diverges, so
# they reject there. Fork-specific patches live in backend/cpp/bonsai/patches/
# and are applied by apply-patches.sh below.
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/patches
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build purge
$(info $(GREEN)I bonsai build info:cpu-all-variants$(RESET))
LLAMA_REPO=$(LLAMA_REPO) LLAMA_VERSION=$(BONSAI_VERSION) \
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build llama.cpp
bash $(CURRENT_MAKEFILE_DIR)/apply-patches.sh $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/llama.cpp $(PATCHES_DIR)
SHARED_LIBS=ON EXTRA_CMAKE_ARGS="-DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON" TARGET="--target grpc-server --target ggml" \
LLAMA_REPO=$(LLAMA_REPO) LLAMA_VERSION=$(BONSAI_VERSION) \
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build grpc-server
cp -rfv $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/grpc-server bonsai-cpu-all
rm -rf ggml-shared-libs && mkdir -p ggml-shared-libs
find $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/llama.cpp/build \( -name '*.so*' -o -name '*.dylib' \) -exec cp -av {} ggml-shared-libs/ \;
@echo "Collected ggml shared backends:" && ls -la ggml-shared-libs/
bonsai-grpc:
$(call bonsai-build,grpc,-DGGML_RPC=ON -DGGML_AVX=off -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server --target rpc-server)
bonsai-rpc-server: bonsai-grpc
cp -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-grpc-build/llama.cpp/build/bin/rpc-server bonsai-rpc-server
package:
bash package.sh
purge:
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-*-build
rm -rf bonsai-* package
clean: purge

View File

@@ -0,0 +1,48 @@
#!/bin/bash
# Apply the bonsai patch series to a cloned PrismML llama.cpp (prism branch) checkout.
#
# The prism fork branched from upstream llama.cpp before a number of API changes that the
# shared backend/cpp/llama-cpp/grpc-server.cpp depends on. We carry those upstream commits
# as patch files under backend/cpp/bonsai/patches/ and apply them here so the reused
# grpc-server source compiles against the fork unmodified.
#
# Drop the corresponding patch from patches/ whenever the fork catches up with upstream —
# the build will fail fast if a patch stops applying, which is the signal to retire it.
set -euo pipefail
if [[ $# -ne 2 ]]; then
echo "usage: $0 <llama.cpp-src-dir> <patches-dir>" >&2
exit 2
fi
SRC_DIR=$1
PATCHES_DIR=$2
if [[ ! -d "$SRC_DIR" ]]; then
echo "source dir does not exist: $SRC_DIR" >&2
exit 2
fi
if [[ ! -d "$PATCHES_DIR" ]]; then
echo "no patches dir at $PATCHES_DIR, nothing to apply"
exit 0
fi
shopt -s nullglob
patches=("$PATCHES_DIR"/*.patch)
shopt -u nullglob
if [[ ${#patches[@]} -eq 0 ]]; then
echo "no .patch files in $PATCHES_DIR, nothing to apply"
exit 0
fi
cd "$SRC_DIR"
for patch in "${patches[@]}"; do
echo "==> applying $patch"
git apply --verbose "$patch"
done
echo "all bonsai patches applied successfully"

View File

@@ -11,7 +11,7 @@ REPO_ROOT="${CURDIR}/../../.."
# Create lib directory
mkdir -p $CURDIR/package/lib
cp -avrf $CURDIR/llama-cpp-localai-paged-* $CURDIR/package/
cp -avrf $CURDIR/bonsai-* $CURDIR/package/
cp -rfv $CURDIR/run.sh $CURDIR/package/
# Bundle the ggml shared backends from the CPU_ALL_VARIANTS build into package/lib. ggml

View File

@@ -0,0 +1,19 @@
# bonsai fork skew patches
The `bonsai` backend reuses `backend/cpp/llama-cpp/grpc-server.cpp` (written against
LocalAI's pinned *upstream* llama.cpp) but compiles it against the PrismML `prism` fork,
which branched from upstream some commits earlier. Any upstream API change that the shared
gRPC server depends on, but that the fork does not yet carry, is back-ported here as a
`*.patch` file and applied to the cloned fork checkout by `../apply-patches.sh`.
CI treats both this directory and `backend/cpp/llama-cpp/` as Bonsai inputs, since
the wrapper copies and builds the shared llama.cpp backend sources.
Rules:
- One upstream commit (or minimal hunk) per patch, named `NNNN-short-description.patch`.
- Patches are applied with `git apply` from the fork's checkout root.
- `apply-patches.sh` fails fast if a patch stops applying cleanly — that is the signal the
fork has caught up (or diverged), so re-cut or drop the patch.
- Keep this set as small as possible; the long-term fix is the fork rebasing onto a newer
upstream (or Q1_0/Q2_0 landing in mainline llama.cpp, retiring this backend entirely).

56
backend/cpp/bonsai/run.sh Executable file
View File

@@ -0,0 +1,56 @@
#!/bin/bash
set -ex
# Get the absolute current dir where the script is located
CURDIR=$(dirname "$(realpath "$0")")
cd /
echo "CPU info:"
grep -e "model\sname" /proc/cpuinfo | head -1
grep -e "flags" /proc/cpuinfo | head -1
BINARY=bonsai-fallback
# x86/arm64 ship a single bonsai-cpu-all built with ggml CPU_ALL_VARIANTS: ggml's
# backend registry dlopens the best libggml-cpu-*.so for this host, so no shell-side
# probing. ROCm ships only bonsai-fallback, so fall back to it when cpu-all is absent.
if [ -e "$CURDIR"/bonsai-cpu-all ]; then
BINARY=bonsai-cpu-all
fi
if [ -n "$LLAMACPP_GRPC_SERVERS" ]; then
if [ -e "$CURDIR"/bonsai-grpc ]; then
BINARY=bonsai-grpc
fi
fi
# Extend ld library path with the dir where this script is located/lib
if [ "$(uname)" == "Darwin" ]; then
export DYLD_LIBRARY_PATH="$CURDIR"/lib:$DYLD_LIBRARY_PATH
else
export LD_LIBRARY_PATH="$CURDIR"/lib:$LD_LIBRARY_PATH
# Tell rocBLAS where to find TensileLibrary data (GPU kernel tuning files)
if [ -d "$CURDIR/lib/rocblas/library" ]; then
export ROCBLAS_TENSILE_LIBPATH="$CURDIR"/lib/rocblas/library
fi
# Same for hipBLASLt (rocblaslt): the bundled libhipblaslt.so resolves its
# TensileLibrary_lazy_gfx*.dat kernel data relative to itself, so point it at
# the bundled data or it falls back to slow generic kernels (issue #10660).
if [ -d "$CURDIR/lib/hipblaslt/library" ]; then
export HIPBLASLT_TENSILE_LIBPATH="$CURDIR"/lib/hipblaslt/library
fi
fi
# If there is a lib/ld.so, use it
if [ -f "$CURDIR"/lib/ld.so ]; then
echo "Using lib/ld.so"
echo "Using binary: $BINARY"
exec "$CURDIR"/lib/ld.so "$CURDIR"/$BINARY "$@"
fi
echo "Using binary: $BINARY"
exec "$CURDIR"/$BINARY "$@"
# We should never reach this point, however just in case we do, run fallback
exec "$CURDIR"/bonsai-fallback "$@"

View File

@@ -76,12 +76,13 @@ elseif(DS4_GPU STREQUAL "cpu")
set(DS4_OBJS "${DS4_DIR}/ds4_cpu.o")
endif()
# ds4.c now references ds4_distributed.c (distributed inference) and ds4_ssd.c
# (SSD expert-cache), each split into its own translation unit upstream. Both
# are GPU-agnostic objects shared by every GPU mode, so link them in regardless
# of DS4_GPU.
# Upstream splits distributed inference, tensor-parallel transport, the SSD
# expert cache, and layer placement into GPU-agnostic translation units. Link
# them regardless of DS4_GPU.
list(APPEND DS4_OBJS "${DS4_DIR}/ds4_distributed.o")
list(APPEND DS4_OBJS "${DS4_DIR}/ds4_tp.o")
list(APPEND DS4_OBJS "${DS4_DIR}/ds4_ssd.o")
list(APPEND DS4_OBJS "${DS4_DIR}/ds4_layer_pack.o")
add_executable(${TARGET}
grpc-server.cpp

View File

@@ -1,10 +1,10 @@
# ds4 backend Makefile.
#
# Upstream pin lives below as DS4_VERSION?=80ebbc396aee40eedc1d829222f3362d10fa4c6c
# Upstream pin lives below as DS4_VERSION?=efdadd41e20134af4f3381e1ed90e96fe4faef6f
# (.github/bump_deps.sh) can find and update it - matches the
# llama-cpp / ik-llama-cpp / turboquant convention.
DS4_VERSION?=80ebbc396aee40eedc1d829222f3362d10fa4c6c
DS4_VERSION?=efdadd41e20134af4f3381e1ed90e96fe4faef6f
DS4_REPO?=https://github.com/antirez/ds4
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))
@@ -18,20 +18,19 @@ UNAME_S := $(shell uname -s)
CMAKE_ARGS ?= -DCMAKE_BUILD_TYPE=Release
# ds4_distributed.o and ds4_ssd.o are GPU-agnostic translation units that
# ds4.c/ds4_cpu.o now reference (upstream split distributed inference and the
# SSD expert-cache into their own .c files). Both objects are shared by every
# GPU mode, so they are appended unconditionally below.
# Upstream splits distributed inference, tensor-parallel transport, the SSD
# expert cache, and layer placement into GPU-agnostic translation units. They
# are shared by every GPU mode, so append them unconditionally below.
ifeq ($(BUILD_TYPE),cublas)
CMAKE_ARGS += -DDS4_GPU=cuda
DS4_OBJ_TARGET := ds4.o ds4_cuda.o ds4_distributed.o ds4_ssd.o
DS4_OBJ_TARGET := ds4.o ds4_cuda.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
else ifeq ($(UNAME_S),Darwin)
CMAKE_ARGS += -DDS4_GPU=metal
DS4_OBJ_TARGET := ds4.o ds4_metal.o ds4_distributed.o ds4_ssd.o
DS4_OBJ_TARGET := ds4.o ds4_metal.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
else
# CPU reference path (Linux only - macOS CPU path is broken by VM bug per ds4 README).
CMAKE_ARGS += -DDS4_GPU=cpu
DS4_OBJ_TARGET := ds4_cpu.o ds4_distributed.o ds4_ssd.o
DS4_OBJ_TARGET := ds4_cpu.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
endif
ifneq ($(NATIVE),true)
@@ -56,11 +55,11 @@ ds4:
# the right per-platform compile flags (Objective-C/Metal on Darwin, nvcc on Linux+CUDA).
ds4/ds4.o: ds4
ifeq ($(BUILD_TYPE),cublas)
+$(MAKE) -C ds4 ds4.o ds4_cuda.o ds4_distributed.o ds4_ssd.o
+$(MAKE) -C ds4 ds4.o ds4_cuda.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
else ifeq ($(UNAME_S),Darwin)
+$(MAKE) -C ds4 ds4.o ds4_metal.o ds4_distributed.o ds4_ssd.o
+$(MAKE) -C ds4 ds4.o ds4_metal.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
else
+$(MAKE) -C ds4 ds4_cpu.o ds4_distributed.o ds4_ssd.o
+$(MAKE) -C ds4 ds4_cpu.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
endif
grpc-server: ds4/ds4.o

View File

@@ -51,6 +51,11 @@ namespace {
// Global state - ds4 is single-engine-per-process by design.
std::mutex g_engine_mu;
// The ModelOptions.Model this process loaded, compared against
// PredictOptions.ModelIdentity so a request that arrived through a stale
// distributed route is rejected rather than answered from the wrong model
// (#10952). Guarded by g_engine_mu like the rest of the engine state.
std::string g_loaded_model_identity;
ds4_engine *g_engine = nullptr;
ds4_session *g_session = nullptr;
int g_ctx_size = 32768;
@@ -562,6 +567,24 @@ static void build_prompt(ds4_engine *engine, const backend::PredictOptions *requ
ds4_chat_append_assistant_prefix(engine, out, think);
}
// check_model_identity mirrors pkg/grpc/server.go and
// backend/python/common/model_identity.py. Either side empty means "skip": the
// request side is empty for a controller that predates the field, the loaded
// side when such a controller performed the load. A false rejection is worse
// than the miss it prevents. Callers must already hold g_engine_mu.
static GStatus check_model_identity(const backend::PredictOptions *request) {
if (request == nullptr || request->modelidentity().empty()) return GStatus::OK;
if (g_loaded_model_identity.empty() ||
g_loaded_model_identity == request->modelidentity()) {
return GStatus::OK;
}
// NOT_FOUND plus this exact sentinel is the cross-language contract the
// router matches on (grpcerrors.ModelMismatchSentinel).
return GStatus(StatusCode::NOT_FOUND,
"ds4: model identity mismatch: loaded \"" + g_loaded_model_identity +
"\", requested \"" + request->modelidentity() + "\"");
}
class DS4Backend final : public backend::Backend::Service {
public:
GStatus Health(ServerContext *, const backend::HealthMessage *,
@@ -716,6 +739,7 @@ public:
}
result->set_success(true);
g_loaded_model_identity = request->model();
result->set_message("loaded " + model_path);
return GStatus::OK;
}
@@ -724,6 +748,7 @@ public:
backend::TokenizationResponse *response) override {
std::lock_guard<std::mutex> lock(g_engine_mu);
if (!g_engine) return GStatus(StatusCode::FAILED_PRECONDITION, "ds4: model not loaded");
if (GStatus id = check_model_identity(request); !id.ok()) return id;
ds4_tokens out = {};
ds4_tokenize_text(g_engine, request->prompt().c_str(), &out);
for (int i = 0; i < out.len; ++i) response->add_tokens(out.v[i]);
@@ -738,6 +763,7 @@ public:
if (!g_engine || !g_session) {
return GStatus(StatusCode::FAILED_PRECONDITION, "ds4: model not loaded");
}
if (GStatus id = check_model_identity(request); !id.ok()) return id;
if (std::string route_err = wait_route_ready(lock); !route_err.empty()) {
return GStatus(StatusCode::UNAVAILABLE, route_err);
}
@@ -837,6 +863,7 @@ public:
if (!g_engine || !g_session) {
return GStatus(StatusCode::FAILED_PRECONDITION, "ds4: model not loaded");
}
if (GStatus id = check_model_identity(request); !id.ok()) return id;
if (std::string route_err = wait_route_ready(lock); !route_err.empty()) {
return GStatus(StatusCode::UNAVAILABLE, route_err);
}

View File

@@ -1,12 +1,14 @@
#!/bin/bash
set -e
set -euo pipefail
CURDIR=$(dirname "$(realpath "$0")")
REPO_ROOT="${CURDIR}/../../.."
PACKAGE_DIR="$CURDIR/package"
mkdir -p "$CURDIR/package/lib"
cp -avf "$CURDIR/grpc-server" "$CURDIR/package/"
cp -avf "$CURDIR/ds4-worker" "$CURDIR/package/"
cp -rfv "$CURDIR/run.sh" "$CURDIR/package/"
rm -rf "$PACKAGE_DIR"
mkdir -p "$PACKAGE_DIR/lib"
cp -avf "$CURDIR/grpc-server" "$PACKAGE_DIR/"
cp -avf "$CURDIR/ds4-worker" "$PACKAGE_DIR/"
cp -rfv "$CURDIR/run.sh" "$PACKAGE_DIR/"
UNAME_S=$(uname -s)
if [ "$UNAME_S" = "Darwin" ]; then
@@ -16,25 +18,54 @@ if [ "$UNAME_S" = "Darwin" ]; then
fi
if [ -f "/lib64/ld-linux-x86-64.so.2" ]; then
cp -arfLv /lib64/ld-linux-x86-64.so.2 "$CURDIR/package/lib/ld.so"
LIBDIR=/lib/x86_64-linux-gnu
cp -arfLv /lib64/ld-linux-x86-64.so.2 "$PACKAGE_DIR/lib/ld.so"
elif [ -f "/lib/ld-linux-aarch64.so.1" ]; then
cp -arfLv /lib/ld-linux-aarch64.so.1 "$CURDIR/package/lib/ld.so"
LIBDIR=/lib/aarch64-linux-gnu
cp -arfLv /lib/ld-linux-aarch64.so.1 "$PACKAGE_DIR/lib/ld.so"
else
echo "package.sh: unknown architecture" >&2; exit 1
fi
for lib in libc.so.6 libgcc_s.so.1 libstdc++.so.6 libm.so.6 libgomp.so.1 \
libdl.so.2 librt.so.1 libpthread.so.0; do
cp -arfLv "$LIBDIR/$lib" "$CURDIR/package/lib/$lib"
# Bundle the complete dependency closure for both executables. In particular,
# grpc-server links the distro gRPC/protobuf/absl stack; copying only the core
# C/C++ runtime libraries leaves the scratch image unable to start.
{
ldd "$CURDIR/grpc-server"
ldd "$CURDIR/ds4-worker"
} | awk '$2 == "=>" && $3 ~ /^\// { print $3 }' | sort -u | \
while read -r so; do
cp -arfLv "$so" "$PACKAGE_DIR/lib/"
done
GPU_LIB_SCRIPT="${REPO_ROOT}/scripts/build/package-gpu-libs.sh"
if [ -f "$GPU_LIB_SCRIPT" ]; then
source "$GPU_LIB_SCRIPT" "$CURDIR/package/lib"
# shellcheck source=/dev/null
source "$GPU_LIB_SCRIPT" "$PACKAGE_DIR/lib"
package_gpu_libs
fi
# Resolve every dependency through the same loader and library path used by
# the from-scratch image. The loader can still search host defaults, so reject
# any absolute dependency path that escapes the package instead of accepting a
# false-positive validation against a library that scratch will not contain.
validate_packaged_binary() {
local binary="$1"
local resolution
resolution=$("$PACKAGE_DIR/lib/ld.so" \
--library-path "$PACKAGE_DIR/lib" \
--list "$PACKAGE_DIR/$binary")
printf '%s\n' "$resolution" | awk -v prefix="$PACKAGE_DIR/lib/" '
$2 == "=>" && $3 ~ /^\// && index($3, prefix) != 1 {
print "package.sh: dependency resolved outside package: " $0 > "/dev/stderr"
invalid = 1
}
END { exit invalid }
'
}
for binary in grpc-server ds4-worker; do
validate_packaged_binary "$binary"
done
echo "ds4 package contents:"
ls -lah "$CURDIR/package/" "$CURDIR/package/lib/"
ls -lah "$PACKAGE_DIR/" "$PACKAGE_DIR/lib/"

View File

@@ -1,5 +1,5 @@
IK_LLAMA_VERSION?=f96eaddba8bed6a9a5e628bbf6a566775c70b49c
IK_LLAMA_VERSION?=e5357286c0d433cd4384e82ed7e2b6d655f57087
LLAMA_REPO?=https://github.com/ikawrakow/ik_llama.cpp
CMAKE_ARGS?=

View File

@@ -2412,7 +2412,33 @@ static void params_parse(const backend::ModelOptions* request,
// GRPC Server start
class BackendServiceImpl final : public backend::Backend::Service {
private:
// The ModelOptions.Model this process was loaded with. Compared against
// PredictOptions.ModelIdentity so a request that reached us through a stale
// distributed route is rejected instead of answered from the wrong model
// (#10952).
std::string loaded_model_identity;
public:
// checkModelIdentity mirrors pkg/grpc/server.go and
// backend/python/common/model_identity.py. Either side being empty means
// "skip": the request side is empty for a controller that predates the field,
// and the loaded side is empty when such a controller performed the load. A
// false rejection is worse than the miss it prevents.
grpc::Status checkModelIdentity(const backend::PredictOptions* request) {
if (request == nullptr || request->modelidentity().empty()) {
return grpc::Status::OK;
}
if (loaded_model_identity.empty() || loaded_model_identity == request->modelidentity()) {
return grpc::Status::OK;
}
// NOT_FOUND plus this exact sentinel is the cross-language contract the
// router matches on (grpcerrors.ModelMismatchSentinel).
return grpc::Status(grpc::StatusCode::NOT_FOUND,
"ik-llama-cpp: model identity mismatch: loaded \"" + loaded_model_identity +
"\", requested \"" + request->modelidentity() + "\"");
}
grpc::Status Health(ServerContext* context, const backend::HealthMessage* request, backend::Reply* reply) {
// Implement Health RPC
reply->set_message("OK");
@@ -2438,9 +2464,12 @@ public:
result->set_message("Loading succeeded");
result->set_success(true);
loaded_model = true;
loaded_model_identity = request->model();
return Status::OK;
}
grpc::Status PredictStream(grpc::ServerContext* context, const backend::PredictOptions* request, grpc::ServerWriter<backend::Reply>* writer) override {
auto identity = checkModelIdentity(request);
if (!identity.ok()) return identity;
json data = parse_options(true, request, llama);
const int task_id = llama.queue_tasks.get_new_id();
llama.queue_results.add_waiting_task_id(task_id);
@@ -2495,6 +2524,8 @@ public:
grpc::Status Predict(ServerContext* context, const backend::PredictOptions* request, backend::Reply* reply) {
auto identity = checkModelIdentity(request);
if (!identity.ok()) return identity;
json data = parse_options(false, request, llama);
const int task_id = llama.queue_tasks.get_new_id();
llama.queue_results.add_waiting_task_id(task_id);
@@ -2532,6 +2563,8 @@ public:
/// https://github.com/ggerganov/llama.cpp/blob/aa2341298924ac89778252015efcb792f2df1e20/examples/server/server.cpp#L2969
grpc::Status Embedding(ServerContext* context, const backend::PredictOptions* request, backend::EmbeddingResult* embeddingResult) {
auto identity = checkModelIdentity(request);
if (!identity.ok()) return identity;
json data = parse_options(false, request, llama);
const int task_id = llama.queue_tasks.get_new_id();
llama.queue_results.add_waiting_task_id(task_id);
@@ -2556,6 +2589,8 @@ public:
}
grpc::Status TokenizeString(ServerContext* context, const backend::PredictOptions* request, backend::TokenizationResponse* response){
auto identity = checkModelIdentity(request);
if (!identity.ok()) return identity;
json data = parse_options(false, request, llama);
std::vector<llama_token> tokens = llama.tokenize(data["prompt"],false);

View File

@@ -1,157 +0,0 @@
# llama-cpp-localai-paged is LocalAI's paged-attention llama.cpp variant. It
# builds upstream llama.cpp with the LocalAI paged-attention patch series
# (patches/paged/, vendored in THIS backend) applied on top. It reuses
# backend/cpp/llama-cpp's grpc-server.cpp / CMakeLists.txt / prepare.sh / Makefile
# sources verbatim via a thin wrapper - the stock llama-cpp backend is pure
# upstream and carries NONE of the paged patches; this backend OWNS them.
#
# Pin handling (mirrors the turboquant wrapper, the precedent this is modelled
# on): the paged patch series is hand-verified bit-exact against ONE specific
# llama.cpp tip and re-exported by the manual PIN_SYNC process
# (README section 7 + .agents/llama-cpp-localai-paged-backend.md). A naive
# pin bump would move the tip out from
# under the patches and break `git apply` at build time, so this backend OWNS
# its pin (LLAMA_VERSION below) instead of inheriting the auto-bumped stock pin
# from backend/cpp/llama-cpp/Makefile. The override is forced into every copied
# build via `LLAMA_VERSION=$(LLAMA_VERSION)`. There is deliberately NO
# bump_deps.yaml entry for it: it is advanced ONLY by PIN_SYNC, never nightly.
# (turboquant CAN auto-bump because its fork branch carries the patches; the
# paged series is vendored as .patch files here, so it cannot.)
#
# - NO patch-grpc-server.sh and NO apply-patches.sh: the shared grpc-server.cpp
# already carries the (runtime-gated) paged option hooks, and the paged patch
# series (patches/paged/) is applied by THIS Makefile's own apply step onto
# the freshly cloned tree, using the same strict `git apply` method the stock
# build uses for base patches. The stock llama-cpp Makefile applies only its
# own (currently empty) base patches/ series, never the paged one.
# Manually pin-synced llama.cpp tip the paged patch series is verified against.
# Decoupled from the auto-bumped stock pin in backend/cpp/llama-cpp/Makefile so
# the nightly llama.cpp bump cannot silently break the vendored paged patches.
# Advance ONLY via the PIN_SYNC process (rebase patches + bit-exact gate +
# re-export), then update this value. See:
# README section 7 + .agents/llama-cpp-localai-paged-backend.md
#
# This pin = the manual, verified sync. The signal telling you WHEN to do the
# next sync is the early-warning canary
# (.github/workflows/llama-cpp-paged-canary.yml): weekly it applies + compiles
# this patch series against the latest upstream llama.cpp tip and goes red the
# moment upstream drifts past the patches. Canary red -> run a PIN_SYNC, then
# bump this value. The canary never touches this pin; it is signal-only.
#
# HARD CONSTRAINT: keep this == the stock llama-cpp pin (backend/cpp/llama-cpp/
# Makefile). grpc-server.cpp is SHARED with the stock backend and tracks the
# stock pin; a paged pin that diverges PAST an upstream server-API refactor
# breaks the grpc-server LINK even when the patches are byte-for-byte bit-exact.
# The c299a92c bump did exactly this: patches applied + greedy-md5 bit-exact, but
# grpc-server.cpp failed to link with undefined references to stream_* server
# helpers that the refactor pulled into the headers grpc-server.cpp includes.
# Therefore a PIN_SYNC must pass the FULL grpc-server build/link on CI, not only
# the bit-exact gate. See README section 7 + .agents/llama-cpp-localai-paged-backend.md.
LLAMA_VERSION?=0ed235ea2c17a19fc8238668653946721ed136fd
CMAKE_ARGS?=
BUILD_TYPE?=
NATIVE?=false
ONEAPI_VARS?=/opt/intel/oneapi/setvars.sh
TARGET?=--target grpc-server
JOBS?=$(shell nproc 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 1)
ARCH?=$(shell uname -m)
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))
LLAMA_CPP_DIR := $(CURRENT_MAKEFILE_DIR)/../llama-cpp
# OUR vendored paged-attention patch series. Owned by this backend; the stock
# llama-cpp backend no longer carries it. Applied onto each freshly cloned
# llama.cpp tree by apply-paged-patches below (strict git apply).
PAGED_PATCHES_DIR := $(CURRENT_MAKEFILE_DIR)/patches/paged
GREEN := \033[0;32m
RESET := \033[0m
# Apply OUR vendored paged-attention patch series (patches/paged/0*.patch) onto a
# freshly cloned llama.cpp tree ($(1)) using the SAME strict git-apply method the
# stock build uses for its base patches (backend/cpp/llama-cpp/Makefile `llama.cpp`
# target). Strict: any patch that no longer applies aborts the build (exit 1) -
# that is the signal to run a PIN_SYNC, never to bump the pin blindly. The series
# is owned by THIS backend, not by the now-pure stock llama-cpp backend.
define apply-paged-patches
cd $(1) && \
for p in $(PAGED_PATCHES_DIR)/0*.patch; do \
[ -e "$$p" ] || continue; \
echo "applying llama.cpp PAGED patch: $$p"; \
git apply --verbose "$$p" || { echo "paged patch failed: $$p"; exit 1; }; \
done
endef
# Each flavor target:
# 1. copies backend/cpp/llama-cpp/ (grpc-server.cpp + prepare.sh +
# CMakeLists.txt + Makefile) into a sibling
# llama-cpp-localai-paged-<flavor>-build directory;
# 2. clones OUR pinned upstream llama.cpp into that copy via the copy's own
# `llama.cpp` target (which applies the stock base patches/ series, normally
# empty), then applies THIS backend's paged patch series (patches/paged/)
# onto the cloned tree with strict `git apply` (apply-paged-patches);
# 3. runs the copy's `grpc-server` target and copies the produced binary up as
# llama-cpp-localai-paged-<flavor>.
# We clone+patch only the *copy*, never the original under backend/cpp/llama-cpp/,
# so the stock llama-cpp build stays untouched and patch-free.
define paged-build
rm -rf $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build
cp -rf $(LLAMA_CPP_DIR) $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build purge
$(info $(GREEN)I llama-cpp-localai-paged build info:$(1)$(RESET))
LLAMA_VERSION=$(LLAMA_VERSION) $(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build llama.cpp
$(call apply-paged-patches,$(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build/llama.cpp)
CMAKE_ARGS="$(CMAKE_ARGS) $(2)" TARGET="$(3)" LLAMA_VERSION=$(LLAMA_VERSION) \
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build grpc-server
cp -rfv $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build/grpc-server llama-cpp-localai-paged-$(1)
endef
llama-cpp-localai-paged-avx2:
$(call paged-build,avx2,-DGGML_AVX=on -DGGML_AVX2=on -DGGML_AVX512=off -DGGML_FMA=on -DGGML_F16C=on,--target grpc-server)
llama-cpp-localai-paged-avx512:
$(call paged-build,avx512,-DGGML_AVX=on -DGGML_AVX2=off -DGGML_AVX512=on -DGGML_FMA=on -DGGML_F16C=on,--target grpc-server)
llama-cpp-localai-paged-avx:
$(call paged-build,avx,-DGGML_AVX=on -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server)
llama-cpp-localai-paged-fallback:
$(call paged-build,fallback,-DGGML_AVX=off -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server)
# Single-build CPU backend via ggml CPU_ALL_VARIANTS (mirrors llama-cpp-cpu-all).
# Reuses backend/cpp/llama-cpp's CMakeLists.txt (hw_grpc_proto STATIC) and
# Makefile (SHARED_LIBS make-var + EXTRA_CMAKE_ARGS), so this passes the same
# overrides through to the copied build: SHARED_LIBS=ON, the DL flags, and
# --target ggml (which pulls in the per-microarch libggml-cpu-*.so via ggml's
# add_dependencies). The .so set is collected for package.sh to bundle into
# package/lib.
llama-cpp-localai-paged-cpu-all:
rm -rf $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build
cp -rf $(LLAMA_CPP_DIR) $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build purge
$(info $(GREEN)I llama-cpp-localai-paged build info:cpu-all-variants$(RESET))
LLAMA_VERSION=$(LLAMA_VERSION) $(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build llama.cpp
$(call apply-paged-patches,$(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build/llama.cpp)
SHARED_LIBS=ON EXTRA_CMAKE_ARGS="-DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON" TARGET="--target grpc-server --target ggml" LLAMA_VERSION=$(LLAMA_VERSION) \
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build grpc-server
cp -rfv $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build/grpc-server llama-cpp-localai-paged-cpu-all
rm -rf ggml-shared-libs && mkdir -p ggml-shared-libs
find $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build/llama.cpp/build \( -name '*.so*' -o -name '*.dylib' \) -exec cp -av {} ggml-shared-libs/ \;
@echo "Collected ggml shared backends:" && ls -la ggml-shared-libs/
llama-cpp-localai-paged-grpc:
$(call paged-build,grpc,-DGGML_RPC=ON -DGGML_AVX=off -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server --target ggml-rpc-server)
llama-cpp-localai-paged-rpc-server: llama-cpp-localai-paged-grpc
cp -rf $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-grpc-build/llama.cpp/build/bin/ggml-rpc-server llama-cpp-localai-paged-rpc-server
package:
bash package.sh
purge:
rm -rf $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-*-build
rm -rf llama-cpp-localai-paged-* package
clean: purge

View File

@@ -1,699 +0,0 @@
# LocalAI paged-attention llama.cpp patch series
This backend vendors the patch series (in `patches/paged/`) that turns stock
llama.cpp into LocalAI's paged-attention variant (`llama-cpp-localai-paged`). The
patches are applied on top of a pinned upstream llama.cpp at build time; nothing
here is a fork - it is a source-only `*.patch` stack plus this canonical doc.
> One-file rule: this README is the canonical reference for the patch series. The
> only other docs are operational, kept in `docs/`, and linked below:
> - [`PAGED_BITEXACT_NOTE.md`](docs/PAGED_BITEXACT_NOTE.md) - the per-path bit-exactness gate (the canonical paged-MoE md5 reference).
> - [`LOCALAI_LLAMACPP_BACKEND_PLAN.md`](docs/LOCALAI_LLAMACPP_BACKEND_PLAN.md) - the design-of-record for shipping this as its own backend + the NVFP4 gallery items.
> - [`VLLM_PARITY_FINAL.md`](docs/VLLM_PARITY_FINAL.md) - the definitive, closed record of the GB10 vLLM-parity investigation: full benchmark, every lever + verdict, the structural floors, and the parity verdict (summarized in section 9 below). Read this before reopening any parity work.
> - [`EXECUTION_REARCH_SCOPE.md`](docs/EXECUTION_REARCH_SCOPE.md) - the reopened scope: ports vLLM's execution *architecture* (bf16-resident stream, expert-major fused MoE region, persistent-CTA GEMM, token-budget scheduler, blocked-solve GDN) into the fork additively, on the thesis that same-silicon 2-3x is software-architecture-conditional, not a hardware floor. Phased (P1-P6), each with a falsifiable P0 kill-gate. Read this to pick up parity work after `VLLM_PARITY_FINAL.md`.
---
## 1. What it is
`llama-cpp-localai-paged` is the LocalAI paged-attention llama.cpp backend: a
vendored patch series over upstream llama.cpp that adds
- a **paged KV cache** (vLLM-style block manager: on-demand fixed-size blocks,
free pool, ref-counted blocks) with a **block-table flash-attention** read so
the attention kernels index physical cells instead of a contiguous buffer;
- **cross-request prefix sharing** - concurrent requests that share a long
prefix physically reuse one committed copy of the prefix blocks and prefill
only their divergent suffix;
- a **decode-first prefill scheduler** - a dynamic per-step prefill-token budget
decoupled from `n_batch`, so a long prefill never freezes co-batched decode;
- **GB10 / Blackwell NVFP4 decode optimizations** for the Qwen3.6 hybrid
gated-DeltaNet (SSM) models, where the recurrent-state plumbing - not the FP4
GEMM - dominates the decode step.
It is **pinned to llama.cpp `0ed235ea2c17a19fc8238668653946721ed136fd`** (kept == the stock `llama-cpp` backend's
pin) and advanced only by a manual, bit-exact-gated pin-sync process (see
section 7, "Pin + maintenance policy"), decoupled from the nightly auto-bumper. The pin must stay aligned with the stock pin because
`grpc-server.cpp` is shared; an earlier bump to `c299a92c` was bit-exact but broke
the grpc-server link and was reverted to the then-current stock pin.
The build gate is `LLAMA_PAGED` (default on in this tree); the paged engine is
enabled per-model at runtime via the gallery `options:` knobs (`paged_kv:true`,
`max_batch_tokens:`, `kv_unified:false`, ...). Against unpatched llama.cpp the
runtime hooks are inert, so a single `grpc-server.cpp` is shared between the
clean and the paged build.
---
## 2. Architecture
The decode step on these models breaks into three cost centers; the patch series
attacks each one.
**Paged KV manager + block-table flash-attn.** A host-side `PagedKVManager`
(`FreeBlockQueue` / `BlockPool` / chained-hash content cache) hands out
fixed-size KV blocks on demand and reclaims them per-sequence (ref-counted, with
copy-on-write for shared prefixes). The attention path reads through a **block
table** - an `I32 [n_view, n_stream]` position-ordered physical-cell index passed
as `src[5]` of `ggml_flash_attn_ext` - so the CUDA fattn vec/tile kernels and the
CPU reference map logical KV index `j` to physical cell `block_table[seq*ne11+j]`
and read K/V in place. Token-position ordering keeps the flash-attn online-softmax
reduction order identical to stock. A null block table is the stock contiguous
read, byte-identical.
**The gated-DeltaNet (GDN / SSM) decode path.** The Qwen3.6 hybrid models are 48
gated-DeltaNet (linear-attention / SSM) layers + 16 full-attention layers. On
GB10 the recurrent-state plumbing, not the weight GEMM, is the dominant decode
cost. The series fuses that plumbing to mirror vLLM's
`fused_recurrent_gated_delta_rule`: the recurrent state is read from and written
to its cache slot in place (no copy-back, no `get_rows` materialization), the
conv state is updated in place, the output projection is reshaped to route to the
tensor-core MMQ GEMM, and the recurrence kernel is occupancy-retuned - all
bit-exact (md5-gateable) against the f32 baseline.
**NVFP4 native FP4-MMA on Blackwell.** The NVFP4 dense/expert weight GEMM uses
Blackwell's native FP4-MMA. The series removes a redundant activation-requantize
in the MoE broadcast projections (bit-exact byte copy of identical blocks) and
keeps CUDA graphs on for the grouped-MMQ MoE decode step. These are the only
NVFP4-specific optimizations; on non-Blackwell hardware the FP4 path falls back
to dequant.
**The prefill/decode scheduler.** `update_slots()` already emits one unified
mixed prefill+decode batch per step. The scheduler patches change only the *count*
of prefill tokens admitted per step: decode tokens are claimed first
(decode-first), then a dynamic budget `max(n_ubatch, T - D)` (where `D` is the
live decode load and `T` is `LLAMA_MAX_BATCH_TOKENS`) admits prefill, auto-
shrinking as decode load rises. Pure scheduler policy, byte-identical when off,
orthogonal to the paged allocator.
---
## 3. Patch series (0001-0063)
Source-only patches, with intentional numbering gaps (e.g. 0005, 0027). The
decode-serving graph-reuse levers are 0040-0041. "Bit-exact" = greedy md5 /
`test-backend-ops` byte-identical to the relevant baseline; the gate methodology
is in section 5.
### Paged-KV core (0001-0012)
| # | What it does | Bit-exact |
|---|---|---|
| 0001 | Vendor the host-side paged KV block manager (`FreeBlockQueue`, `BlockPool`, `PagedKVManager`, chained-hash prefix cache). Pure C++17, nothing uses it yet. | n/a (no behavior) |
| 0002 | Place each sequence at permuted, non-contiguous block positions in `find_slot` (proves attention is invariant to physical KV placement). | yes (token-identical) |
| 0003 | Gather K/V/mask down to each stream's non-empty cells before `build_attn_mha`, position-sorted so the FA reduction order matches stock. | yes |
| 0004 | Drive paged placement through the vendored manager: blocks popped on demand, returned on seq end. Core kv-cache struct untouched. | yes (stock path byte-identical) |
| 0006 | Host-side cross-request prefix caching: hash prefix blocks, reuse matching physical blocks (ref-count++), COW-privatise before a divergent write. | yes (default off) |
| 0007 | Wire the prefix cache into the engine so a new sequence physically shares cached prefix blocks and skips recomputing the shared prefix. | yes (verified byte-identical) |
| 0008 | Wire cross-request prefix share into the llama-server continuous-batch loop so concurrent shared-prefix requests prefill only the suffix (36x fewer prefill tokens at K=32). | within CUDA batch-shape non-determinism band |
| 0009 | Replace the per-step gather with an **in-kernel paged read** (block table as `src[5]`); the K/V `get_rows` is gone. Decode step at batch32 691->696ms (was 1279ms gathered). | yes on CPU/batch1; GPU batch>1 within vec-vs-mma band |
| 0010 | Graft the block-table read into the tile kernel; add a dispatch guard so a present block table routes ONLY to vec/tile (never the mma/wmma kernels that ignore it). | yes (CPU byte-identical; vec route) |
| 0011 | Route the GQA-grouped F16 decode to the **tile kernel** (native head-group reuse) by default; vec for everything else. Paged decode to within 1.8% of stock. | vs stock-mma: different-kernel rounding; bit-exact vs vec |
| 0012 | Defensive `GGML_ASSERT(n_view % 64 == 0)` so a future pad/tile change can't silently reintroduce a past-end KV leak on the tile route. | yes (additive assert) |
### Decode-first scheduler (0013, 0016)
| # | What it does | Bit-exact |
|---|---|---|
| 0013 | `LLAMA_PREFILL_BUDGET`: a static per-step prefill-token budget decoupled from `n_batch` (vLLM `--max-num-batched-tokens` analogue). Flattens the decode ITL spike a long prefill inflicts (8.5x smaller worst freeze). | yes (off/short = byte-identical; == `-b` chunking) |
| 0016 | Supersede 0013 with a **dynamic decode-first** budget: `max(n_ubatch, T-D)`, auto-shrinking as decode load `D` rises. Policy-only inside `update_slots()`, zero libllama changes. | yes (default-off byte-identical) |
(0014/0015 are the MoE token-tile levers: 0014 adds `LLAMA_MOE_MMQ_X` (opt-in
high-batch decode micro-opt, +4.8% on Qwen3-Coder-30B), 0015 makes it a
default-on, density-aware auto-select that is prefill-safe by construction. Both
bit-exact. 0017 is the dense FP4-GEMM occupancy-tune track: bit-exact gate green,
but every cheap occupancy lever regressed on GB10, so nothing is enabled - it
ships as the parity gate + default-off instrumentation only.)
### Decode-serving graph reuse (0040, 0041)
These two close the **continuous-serving** decode gap (distinct from the static
batched-bench decode kernel, which is already at vLLM parity - see
[`docs/DECODE_SERVING_SCOPE.md`](docs/DECODE_SERVING_SCOPE.md)). In serving the
host rebuilt the ggml graph on **every** decode step (layer-A graph reuse was 0%),
so the GPU idled while the host rebuilt - the host-bound -39% the static bench
hides.
| # | What it does | Bit-exact |
|---|---|---|
| 0040 | **S1 paged decode-graph reuse** - the paged decode inputs (`input_block_table` / `input_gather_idxs`) never overrode `can_reuse` (defaults to false), so any graph carrying a paged input could never be reused. Add a correct `can_reuse` keyed on the (256-bucketed) block-table dims + a live-mctx refresh from the owning attn input. `LLAMA_PAGED_NO_GRAPH_REUSE=1` forces the pre-S1 path. | yes (md5 byte-identical reuse on/off; dense `5951a5b4`, paged-MoE `8cb0ce23`) |
| 0041 | **S3 decode-shape-stable scheduling** - keep co-batched prefill OUT of decode steps so the pure-decode batch shape stays reuse-stable (S1 makes a pure-decode step reusable; S3 makes the scheduler emit them). Pure `update_slots()` policy on top of 0016; prefill admitted on a bounded cadence (`LLAMA_PAGED_PREFILL_PERIOD`, default 8). **Default OFF** (opt-in via `LLAMA_PAGED_DECODE_STABLE=1`): a measured end-to-end A/B proved default-on is a serving mistake - deferring prefill admission on the period-8 cadence gives **2.5x worse TTFT** (60s vs 24s at N=256) and **20-29% lower end-to-end throughput**, with no end-to-end win at any concurrency; its apparent `decode_agg` gain was a metric artifact (faster per-step decode bought by starving prefill). Default prefers prompt prefill admission for good TTFT; opt in only for decode-dominated, low-arrival traffic where TTFT is not a concern. | yes (byte-identical on/off; per-stream independent in serving) |
Measured (GB10, MoE Qwen3.6-35B-A3B-NVFP4, 128-client staggered streaming load):
graph reuse **0% -> 72.2%**, host window `hostproc` **15.98 -> 6.31 ms/step**,
decode **4.05 -> 5.52 tok/s/seq median (4.24 -> 5.96 mean, at vLLM's ~5.9
sustained)**. S1 is necessary but **not** sufficient alone (13.8% reuse - prefill
co-batching churns the shape nearly every step); S3 is the multiplier of that
per-step decode metric. **But those are per-step decode numbers, not an end-to-end
serving win**: a later end-to-end A/B showed S3-default-on regresses real serving
(2.5x worse TTFT, 20-29% lower end-to-end throughput, no win at any concurrency),
because the period-8 cadence defers prefill admission. So **only S1 (0040) ships
default-on; S3 (0041) now defaults OFF and is opt-in** (`LLAMA_PAGED_DECODE_STABLE=1`,
for decode-dominated low-arrival traffic). The static batched-bench A/B isolates the S1
mechanism: paged decode reuse 0% -> 95.5% (throughput flat there, since the static
regime is GPU-bound). **S2 (double-buffer `set_inputs`) was dropped**: the Phase-0
profile put `set_inputs` at ~0.05 ms/step (the cost is the rebuild, not the input
copy), so it has nothing to recover. The remaining ~28% serving rebuilds are
request-boundary D/seq-set churn + the prefill-cadence steps. A **padded/fixed-slot
decode shape** to capture them was then implemented and GPU-tested (2026-06-28) and
**REJECTED** - it is bit-exact/inert but regresses serving throughput at every
concurrency, because this serving decode is GPU-compute-bound (baseline reuse 0% ~=
S1+S3 reuse 72% on aggregate tok/s), so the dummy-row compute it adds costs more
than the reuse it recovers. Full record + numbers in `docs/DECODE_SERVING_SCOPE.md`
("Padded-shape lever - rejected").
### Prefill fusions (0042, 0044)
CUDA-family graph fusions of the pre-norm residual chain and the gated-DeltaNet
output norm: separate `rms_norm` / `mul` / `add` / `silu` launches collapse into
one kernel so the intermediate never round-trips to HBM. Bit-exact (the fused
kernel reproduces the unfused FP order; float multiply is commutative). Each is
env-gated default-ON (`LLAMA_FUSE_*=0` for a clean single-build A/B that reverts
to the byte- and kernel-identical unfused path).
| # | What it does | Bit-exact / effect |
|---|---|---|
| 0042 | **Fused residual-add + RMS norm + weight multiply** (`rms_norm_pre_add_mul_f32`) - the pre-norm residual `h = x + sub_out; n = rms_norm(h) * w` ran as a `k_bin_bcast` ADD feeding the fused rms_norm+mul; the residual ADD has a second consumer (the skip add) so it can't pass the single-use `ggml_can_fuse`. Recognized via `ggml_can_fuse_subgraph` (ADD + final MUL both outputs), folded into one launch that publishes `h` and emits `scale * h * w`. Gate `LLAMA_FUSE_ADD_RMSNORM`. | yes (dense `5951a5b4`, MoE `8cb0ce23`); dense S_PP +0.5% |
| 0044 | **Fused gated RMSNorm + SiLU gate multiply** (`rms_norm_gate_mul_f32`) - the gated-DeltaNet output norm `(rms_norm(x) * w) * silu(z)` (qwen35 / qwen35moe `build_norm_gated`) ran as rms_norm_mul + silu_mul, two launches with the normalized intermediate crossing HBM. The gate z-projection (a MUL_MAT) is scheduled between the weight MUL and the SILU, so the chain is not naturally consecutive; `build_norm_gated` emits the gate multiply as `mul(silu(z), normalized)` (commutative, bit-exact) so the graph lays out the consecutive subgraph `{ SILU, RMS_NORM, MUL, MUL }` that `ggml_cuda_can_fuse` folds into one `scale * x * w * silu(z)` launch. Gate `LLAMA_FUSE_GATE_RMSNORM`. Profile (dense npp512): 672 (rms_norm_mul + silu_mul) -> 336 fused launches. | yes (dense `5951a5b4`, MoE `8cb0ce23`, paged + non-paged; `test-backend-ops` 12979/12979); S_PP dense +1.1% (~+10 us/tok), MoE +0.9% |
### SSM (gated-DeltaNet) decode levers (0018-0022, 0028)
These are the dominant decode levers on the Qwen3.6 hybrid models. All bit-exact.
| # | What it does | Effect (dense q36-27b / MoE q36-35b-a3b @npl128) |
|---|---|---|
| 0018 | **In-place SSM state write-back** - the recurrence writes its final state directly into the cache slot, removing the ~225MB/copy D2D memcpy (18.9% of decode time). | dense +23.5% / MoE +18.9% |
| 0019 | **Fused recurrent-state gather** - the op reads each sequence's prior state directly from `cache[ids[seq]]` (no `get_rows` materialization); race-free in-place + ids read. | dense +37.8% / MoE +35.3% |
| 0020 | **o_proj MMVQ->MMQ reshape** - collapse the GDN output to 2D so the output projection routes to the M=128 tensor-core MMQ GEMM (was a batch<=8 MMVQ GEMV). The single biggest decode-parity lever. | dense +31.7% (->85.9% of vLLM) / MoE +23.3% |
| 0021 | **Conv-state in-place fusion** - one `ggml_ssm_conv_update_inplace` op replaces the 4-op conv chain (transpose+concat+conv+silu+ring-cpy), writing the shifted ring state in place. | dense +3.2% / MoE +3.5% |
| 0022 | **GDN recurrence occupancy/coalescing retune** - column-folding (NUM_WARPS/COLS_PER_WARP) raises memory-level parallelism on the bandwidth-bound B=128 recurrence kernel; per-column f32 FMA order unchanged. 73.4%->84.6% of GB10 peak BW. | dense +11.1% / MoE +8.3% |
| 0028 | **Recurrent conv-tap gather fusion** - the last `k_get_rows` in the GDN decode path (the conv-state tap gather) becomes an indexed in-kernel read. | dense ~377 t/s / MoE ~784 t/s |
### MoE NVFP4 quant (0023, 0025, 0043)
| # | What it does | Bit-exact |
|---|---|---|
| 0023 | **NVFP4 activation-quantize de-dup** - the broadcast up/gate projections re-quantize the same token activation once per expert; quantize the unique token activations once and byte-copy them into the expert-gathered layout. The only NVFP4-specific patch. | yes (byte-identical) |
| 0025 | **MoE decode re-graph** - keep CUDA graphs on for the grouped-MMQ MoE decode step (the upstream guard disables graphs conservatively; the grouped path has no host sync). Was env-gated `LLAMA_MOE_FORCE_GRAPHS`; now ON by default via 0043. | yes (graph replay re-issues identical kernels) |
| 0043 | **MoE decode graph default-on (D1)** - flip 0025 to ON by default: capture/replay the full-step decode CUDA graph (incl. the grouped-MMQ MoE dispatch) instead of re-issuing every kernel each step. Guard is `should_use_mmq()` (FALSE for the large-M NVFP4 prefill of 0034, so prefill keeps graphs disabled - its per-expert host-loop genuinely syncs). `LLAMA_MOE_NO_FORCE_GRAPHS=1` forces the conservative pre-0025 disable for A/B. D1 profiling: the per-expert host-loop (the only device->host MoE-routing readback) is never hit on the NVFP4 grouped path (sync count identical graphs on/off); steady decode is ~99% GPU-busy, so the cost removed is per-step host kernel RE-ISSUE, not a sync. | yes (md5 byte-identical default/off/forced; paged-MoE `8cb0ce23`, dense `5951a5b4`) |
### Pool reclaim, block-table cache, backend gate
| # | What it does | Bit-exact |
|---|---|---|
| 0024 | **Paged-pool burst-reclaim** - truncate trailing blocks on partial-tail `seq_rm`, defrag the free queue when idle, release blocks on slot completion. Fixes the long-server burst-degradation bug (post-burst prefill collapse 488->44 t/s, restored to 532). Host-side accounting only. | yes |
| 0029 | **Block-table within-step host cache** - the block table is fixed for the whole step; cache it on first build and memcpy it for the other full-attention layers (get_block_table -87%/-91%). | yes, per path (paged-MoE ref `8cb0ce23`) |
| 0030 | **Fused-op backend gate** - the fused GDN / discriminated SSM_CONV ops are CUDA-family + CPU only; force them off on any non-CUDA compute backend so a Vulkan/SYCL/Metal build can't silently run the wrong plain-conv kernel. | yes on CUDA (byte-identical pre-0030); safety gate elsewhere |
| 0031 | **Chunked parallel-scan GDN prefill kernel** (upstream TODO) - FLA-style chunked gated-delta-rule for prefill (non-KDA / f32 / final-state): intra-chunk delta rule solved in parallel (UT-transform + forward subst), inter-chunk recurrence over n_tokens/C steps. The scalar-serial form (`GDN_TC=0`) was bit-exact-benign but not faster than the tuned sequential scan at the GB10-forced C=16 (see section 5); **superseded for paged by the tensor-core M5 path of 0047**. | NEW per-path (`test-backend-ops` 91/91, <=1e-7 NMSE vs CPU ref) |
| 0047 | **GDN M5 tensor-core chunked-scan prefill, f32-only re-port, default-ON under paged KV** - the f32/tf32 tensor-core forms of 0031's scan (KK/QK Gram = M2, KS/QS state-boundary 3xtf32 = M3, P*U output = M4, full form-T solve + state-update mma = M5), single build, runtime-selected by `GDN_TC`. Ships **M5 default-on when `LLAMA_KV_PAGED` is set** (`GDN_TC=5` + `GDN_CHUNK_MIN=64`, both env-overridable; OFF/`INT_MAX` when not paged). `GDN_CHUNK_MIN` is the per-call engage threshold and stays > 1 so decode (1 tok/call) keeps the sequential recurrence (at 1 it swallows decode and drops S_TG ~25%); 64 tuned from a {1,32,64,128,256} sweep. The bf16/hybrid dev-tree machinery (STATE_BF16/HYBRID, the dropped 0026 ssm_bf16_tau) and the bf16 CONFIG-C (M8) plus register-resident M6/M7 variants are NOT part of this f32-only series. MoE prefill S_PP +3.5% @npp512 (3x A/B), +17.7% @npp2048; decode S_TG unchanged. | NEW per-path, benign (`test-backend-ops` GATED_DELTA_NET 46/46 default AND force-M5, incl. multi-chunk/tail-chunk/multi-seq; greedy md5 default-on == M5-forced == canonical on the gate prompt: paged-MoE `8cb0ce23`, dense `5951a5b4`; long MoE prompt = one benign greedy flip vs sequential, dense byte-identical) |
| 0046 | **GDN prefill geometry gated by scan length** - patch 0022's `(NUM_WARPS=16, COLS_PER_WARP=8)` column-fold of the GDN sequential-recurrence dispatch (`case 128`) is a decode win but was applied UNCONDITIONALLY, so it also hit dense prefill (~-6% vs stock): on a long sequential scan the launch `grid.z` collapses from `S_v/4 = 32` to `S_v/(16*8) = 1` and the SMs starve (profiled: `gated_delta_net` +54% GPU time = the whole dense-prefill regression). Gate the geometry by per-call scan length: long scans (prefill, `n_tokens >= GDN_PREFILL_NTOK`, default 256) take stock's high-grid.z `(4,1)` geometry; short scans (decode) keep the `(16,8)` retune. Recovers dense prefill +7.2% back to stock parity, keeps the decode win. `GDN_PREFILL_NTOK` tunes the crossover; an explicit `GDN_NW`/`GDN_CPW` sweep still overrides (gate yields when either is set), so the one-build %peak A/B harness is unchanged. | yes (patch 0022 proved every `{NW,CPW}` variant byte-identical, so switching geometry by scan length cannot move the md5) |
### Speculative / MTP investigation (0054, 0055)
| # | What it does | Bit-exact / effect |
|---|---|---|
| 0054 | **Disable backend sampling for MTP drafts** - forces server MTP draft generation through the target-side sampler acceptance path instead of letting the draft backend sample independently. This was required for the Phase 14 rollback/prefix safety gate. | yes for canonical non-MTP gates; Phase 14 MTP normalized greedy-prefix gate passed |
| 0055 | **Trace speculative batch shapes** - adds default-off `LLAMA_SPEC_SHAPE_TRACE=1` server logs around `server_slot::handle_last_sampled_token()`, reporting normal decode rows and MTP verification `K + 1` rows (`draft`, `outputs`, `spec_i_first`, `spec_i_last`). This is instrumentation only for Phase 18 shape-entropy measurement before any scheduler experiment. | yes (env unset is silent; DGX gates after patch: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`) |
| 0056 | **Trace MoE MMQ batch shapes** - adds default-off `LLAMA_MOE_MMQ_SHAPE_TRACE=<n>` logs from the grouped-MMQ host selector, reporting routed assignment count, estimated active experts, density, selected `mmq_x`, `mmq_y`, and stream-k. This is evidence-only instrumentation for sizing structural grouped-MMQ work after Phase 28 rejected launch-bounds/row-tile knobs. | yes (env unset and trace-enabled gates both green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`; trace cap verified with 4 lines) |
| 0057 | **Trace MoE MMQ launch shapes** - extends `LLAMA_MOE_MMQ_SHAPE_TRACE=<n>` with bounded `[LLAMA_MOE_MMQ_LAUNCH]` lines from `launch_mul_mat_q`, recording actual `ntiles_dst`, `stream_k_blocks`, tile efficiency, `fixup`, `ntx/nty/ntzw`, and compiled `mmq_x/mmq_y`. This is evidence-only instrumentation to distinguish real stream-k/fixup overhead from small-M kernel-shape cost. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`; Phase 31 n128 trace showed decode and prefill `fixup=0`, `stream_k_blocks == ntiles_dst`) |
| 0058 | **Trace MoE small-M MMQ candidates** - adds `LLAMA_MOE_MMQ_SMALL_M_TRACE=<n>` and a host-only classifier for decode-like low-density grouped-MMQ shapes (`ncols_max <= 128`, density `<=4`, `mmq_x_best <=64`). It only counts candidate calls for the next structural tile-policy A/B; no numeric branch is added. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`; Phase 32 n128 trace found 4096 candidates, mostly `mmq_x_best=64/48`) |
| 0059 | **Gate MoE small-M MMQ tile policy** - adds default-off `LLAMA_MOE_SMALL_M_TILE=<n>` to cap only classified small-M MoE grouped-MMQ calls. This was used to A/B vLLM-like smaller M blocks without changing default inference. | yes (default-off, tile16, tile8, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`; Phase 33 rejected tile16 and tile8 as slower) |
| 0060 | **Trace MoE MMID dispatch routes** - adds default-off `LLAMA_MOE_MMID_ROUTE_TRACE=<n>` around `MUL_MAT_ID` dispatch, classifying each call as `mmvq`, `mmvf`, grouped `mmq`, `mmf`, or host-sync `fallback`. This is evidence-only instrumentation to resolve whether serving hits the per-expert host-sync fallback. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`; Phase 34 n128 trace found `mmq=2776`, `mmvq=1320`, `host_sync=0/4096`) |
| 0061 | **Trace regular MUL_MAT dispatch routes** - adds default-off `LLAMA_MUL_MAT_ROUTE_TRACE=<n>` around regular `MUL_MAT`, classifying projection-heavy calls as `vec_f`, `mat_f`, `vec_q`, `mmq`, `batched_cublas`, `op_*`, `fp4_prefill`, or `fwht`. This is evidence-only instrumentation for the `bf16-proj` serving bucket. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT` `1146/1146`, `MUL_MAT_ID` `806/806`; Phase 35 n128 trace found BF16 routes `mat_f=2485`, `op_cublas=1330`) |
| 0062 | **Trace cuBLAS subroutes** - adds default-off `LLAMA_CUBLAS_ROUTE_TRACE=<n>` around the generic cuBLAS `MUL_MAT` path, classifying calls as `nvfp4_bf16_tc`, `bf16_tc`, `f16_tc_32f`, `f16_tc_16f`, or `sgemm`. This is evidence-only instrumentation for the Phase 35 `op_cublas` bucket. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT` `1146/1146`, `MUL_MAT_ID` `806/806`; Phase 36 n128 trace found `bf16_tc=5681`, `sgemm=2511`) |
| 0063 | **Trace cuBLAS tensor names** - extends `LLAMA_CUBLAS_ROUTE_TRACE=<n>` with `src0`, `src1`, and `dst` names so the `sgemm` bucket can be tied back to graph nodes. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT` `1146/1146`, `MUL_MAT_ID` `806/806`; Phase 37 n128 trace identified `sgemm` as `ffn_gate_inp* -> ffn_moe_logits/shared_expert_gate`) |
> **Dropped: patch 0026 (hybrid per-head bf16 SSM state, `ssm_bf16_tau`).** Once
> the decode fusions (0028 recurrent-state gather-fusion + 0029 block-table cache)
> landed, the bf16-SSM lever bought nothing: a clean re-measurement forcing **all**
> gated-DeltaNet heads to bf16 (`tau=100000`) gives **flat** decode (780.6 vs
> 780.0 t/s) - the mode engages but adds zero throughput because it is subsumed by
> the fusions. It was a precision trade (not bit-exact) plus extra bug surface and
> CUDA template-instantiation compile cost with no benefit, so it was removed. See
> section 5 ("rejected / flat levers") for the full record.
---
## 4. Benchmarks
Hardware: **GB10 / DGX Spark** (CUDA 13, sm_121). Models: dense
**Qwen3.6-27B-NVFP4** and MoE **Qwen3.6-35B-A3B-NVFP4**. Metric: `decode_agg`
S_TG (t/s) from `llama-batched-bench`, `-fa on -ngl 99`, `npp 128 / ntg 128`,
swept over serving width `npl` in {8, 32, 64, 128}. Plots:
[`qwen36_decode_overview.png`](docs/qwen36_decode_overview.png) (both models),
[`qwen36_dense_decode_vs_npl.png`](docs/qwen36_dense_decode_vs_npl.png),
[`qwen36_moe_decode_vs_npl.png`](docs/qwen36_moe_decode_vs_npl.png); raw data
[`final_benchmark.csv`](docs/final_benchmark.csv).
![NVFP4 decode throughput vs concurrency on GB10: llama.cpp standard vs vLLM vs LocalAI's llama.cpp patches](docs/qwen36_decode_overview.png)
> The plot above also shows a third "bf16-tau" llama curve. That was the opt-in
> `ssm_bf16_tau` lever (patch 0026), since **dropped** - a clean re-measurement
> showed it flat once the decode fusions landed (see section 5). The numbers below
> use only **stock** vs **patched** vs **vLLM**.
> **What was re-measured (2026-06-27).** The two llama columns - **stock** and
> **patched** - were re-measured this session on one consistent
> `llama-batched-bench` harness. The **vLLM** column is the **prior-session
> reference** (kept as-is, *not* re-run this session). Per-run peak
> VRAM was *not* re-captured: the GB10's unified Grace-Blackwell LPDDR5x reports
> `[N/A]` to `nvidia-smi --query-gpu=memory.used` and the bench does not print it
> (the memory-advantage note below is the prior-session finding).
### (a) + (b) Patched vs stock vs vLLM
The **stock** column is a separate, unpatched llama.cpp built at this backend's
**exact pin (`9d5d882d`)**; the **patched** column is
the paged binary, env/flag-toggled (`LLAMA_KV_PAGED=1`, plus
`LLAMA_MOE_FORCE_GRAPHS=1` for MoE). Both
run on the **same harness**, so "x over stock" is an apples-to-apples measure of
the patch series. (Note: the patch series' dominant SSM decode fusions are
compiled in, not env-gated - toggling `LLAMA_KV_PAGED` alone on the *patched*
binary does **not** reproduce stock; only the separately-built unpatched
`9d5d882d` binary does.) The **vLLM** column is a **different harness** (vLLM
server + client continuous batching) and a **prior-session reference**, so the
cross-engine "% of vLLM" is **indicative, not apples-to-apples**.
**Dense Qwen3.6-27B-NVFP4** (decode t/s):
| npl | stock | patched | vLLM (prior) | patched x over stock |
|----:|------:|--------:|-------------:|---------------------:|
| 8 | 68.3 | 85.3 | 70.4 | 1.25x |
| 32 | 119.9 | 211.9 | 211.8 | 1.77x |
| 64 | 142.8 | 305.2 | 309.1 | 2.14x |
| 128 | 155.1 | 382.1 | 418.8 | 2.46x |
Dense **patched** is parity-to-ahead of vLLM (121 / 100 / 99 / 91% of vLLM across
the widths).
**MoE Qwen3.6-35B-A3B-NVFP4** (decode t/s):
| npl | stock | patched | vLLM (prior) | patched x over stock |
|----:|------:|--------:|-------------:|---------------------:|
| 8 | 186.7 | 230.3 | 256.5 | 1.23x |
| 32 | 267.4 | 466.4 | 500.8 | 1.74x |
| 64 | 320.5 | 622.4 | 686.1 | 1.94x |
| 128 | 347.2 | 784.3 | 882.2 | 2.26x |
MoE **patched** is 90 / 93 / 91 / 89% of vLLM.
**Caveat on the vLLM column.** It is a **different harness** and a
**prior-session** measurement (not re-run this session), so the cross-engine "% of
vLLM" is **indicative, not apples-to-apples**. Memory (prior session): llama uses
**1.5-3x lower** memory than vLLM.
**Takeaway.** Re-measured this session, the patch series gives up to **2.46x
(dense) / 2.26x (MoE)** over true-stock `9d5d882d` on the same harness (close to,
slightly below, the prior 2.59x / 2.33x - llama was re-measured, vLLM kept).
Dense is parity-to-ahead of vLLM; MoE **patched** sits at ~89-93% of the
prior-session vLLM. The residual MoE gap is structural (see section 5).
### (c) Apple Silicon (M4, 16GB Metal) - does the patchset help here?
Short answer: **no - the wins are CUDA/Blackwell-specific.** Two facts first: the
24GB NVFP4 GGUF doesn't fit a 16GB M4 (SSD paging), and on Metal `supports_op`
**excludes NVFP4** from `MUL_MAT`/`MUL_MAT_ID`/`GET_ROWS` (FP4 matmuls fall back to
CPU - no Apple FP4-MMA). So NVFP4 Qwen3.6 is not a Mac fit; a Metal-native Q4_K is.
Measured **stock vs patched** (same pin `c299a92c`, both built `-DGGML_METAL=ON`;
the 28-patch series **compiles clean on Metal** - the CUDA code is `#if`-guarded),
on **Qwen3-8B Q4_K_M** (a dense GQA model that fits 16GB and exercises the *live*
Metal features; no Qwen3.6 hybrid GGUF fits 16GB, and the GDN fusions gate off on
Metal anyway), `llama-bench` pp512/tg128 t/s:
| config | pp512 | tg128 |
|---|---:|---:|
| stock | 226.7 | 20.4 |
| patched, paged **off** | 226.7 | 20.3 (= stock) |
| patched, paged **on** | 222.6 | 19.8 (~0.97x) |
Concurrency (`batched-bench`) scales identically to stock (S_TG ~20 -> ~137 at
npl32, from llama.cpp's existing batching). **Verdict: neutral-to-slightly-negative
on Metal.** Patched-paged-off equals stock; turning paged on is ~0-3% slower
decode / ~2-8% slower prefill, because the in-kernel block-table flash-attn read
that *recovers* the gather cost is CUDA-only (`fattn-*.cuh`) - on Metal the paged
path falls back to a host-side gather, pure overhead over stock's contiguous read.
Everything Blackwell-specific (NVFP4, GDN fusions via 0030, occupancy) is inert.
So **on Apple Silicon, prefer the stock `llama-cpp` backend.**
**Vulkan / SYCL** (source analysis): the gated-DeltaNet and SSM_CONV ops DO have
upstream kernels on Vulkan and SYCL (as on Metal), so the Qwen3.6 hybrids RUN on
all three via the non-fused path. The patchset's fusions are gated off there
(0030), so the outcome is the same neutral-to-slightly-negative as Metal - not
"won't run". This backend therefore ships **CUDA-only** (where the fusions are
live + verified); non-CUDA users should use the stock `llama-cpp` backend. See
[`UPSTREAM_LAYER2_SCOPE.md`](docs/UPSTREAM_LAYER2_SCOPE.md) for what native non-CUDA
fused kernels would take.
---
## 5. Dev notes - what we learned
**Bit-exact methodology.** Every bit-exact patch is gated two ways: (1) a greedy
md5 gate - `llama-completion -m MODEL -ngl 99 -fa on -p "The capital of France
is" -n 48 --temp 0 --seed 1 | md5sum`, paged paths prefixed with
`LLAMA_KV_PAGED=1` (+ `LLAMA_MOE_FORCE_GRAPHS=1` for paged MoE), on the default
chat-template path; and (2) `test-backend-ops` (CUDA0 vs CPU oracle) for every
touched op (`SSM_CONV*`, `GATED_DELTA_NET`, `MUL_MAT`, `MUL_MAT_ID`).
For DGX work, `paged-inference-gates.sh` runs the canonical MoE/dense transcript
md5 checks and selected `test-backend-ops` filters, and refuses to start while
docker, `local-ai-worker`, GPU compute processes, or a non-free GPU lock are
present.
For direct `llama-server` MTP serving A/B work, use
`paged-mtp-serving-bench.sh`. It runs the same pre/post inference gates, compares
baseline vs `--spec-type draft-mtp`, and captures the h2h client summaries plus
MTP acceptance lines. Phase 15 rejected current MTP serving on GB10 despite
passing safety gates; do not enable it by default.
**The gate is per-path** (see [`PAGED_BITEXACT_NOTE.md`](docs/PAGED_BITEXACT_NOTE.md)).
Dense is bit-exact across paged/non-paged (`5951a5b4`). The **paged MoE** md5
(`8cb0ce23`) does **not** byte-match the **non-paged MoE** md5 (`07db32c2`); this
is a benign FP-accumulation-order difference of the paged attention reduction,
**KL-validated** against the f16 reference: KLD(paged||f16) 0.13600 <=
KLD(nonpaged||f16) 0.13660, PPL within +/-0.29, ~zero probability bias - two
equivalent FP-reorderings of the same quantized model, not a regression. Future
paged-MoE regressions therefore compare to `8cb0ce23`, not `07db32c2`.
**MoE-parity conclusion** (the residual gap is structural). The two heaviest MoE
decode kernels - the GDN-SSM recurrence and the NVFP4-expert GEMM - are llama
**wins** after this series (the recurrence runs at 102.6% of vLLM's bandwidth;
the GEMM ties vLLM at the LPDDR5x BW floor). The residual gap is **bf16-projection
bandwidth + the host scheduling loop**, both at the LPDDR5x floor - not a kernel
llama is losing. The MoE GEMM kernel is *not* where the gap lives.
**Rejected / flat levers** (recorded so they are not re-tried):
- **Lever 2 - graph/stream coverage: FLAT.** Bit-exact graph coverage was
exhausted by 0025; more graph/stream overlap is a no-op or small regression on
this model.
- **D1 premise "static decode is host-sync-bound on the MoE-routing readback":
REFUTED.** The hypothesis was that the dominant decode cost is the device->host
readback of MoE routing before launching the per-expert GEMMs (mul_mat_id's
per-expert host-loop fallback). Profiling (GB10, q36-35b-a3b-nvfp4, batched-bench
npl128) shows the opposite: on NVFP4 the grouped stream-k MMQ id-path is what
runs (routing stays device-side), so the host-loop fallback is **never hit** -
`cudaStreamSynchronize` count is *identical* with CUDA graphs on vs off (1457
either way; only the kernel-launch count changes, ~100k vs ~229k). Steady-decode
GPU-busy is **~99%** (1% idle), i.e. static decode is GPU-bound, not idle waiting
on a sync. The one actionable residual the profile surfaced - per-step host
kernel **re-issue** when the step is not graph-captured - shipped as 0043
(default-on full-step decode graph), worth +2.6% (npl128) to +5-13% (npl32). The
larger continuous-serving host cost is the graph **rebuild** (0040/0041), and the
irreducible floor is the per-step logits-D2H-before-sampling serial point - none
of which is the MoE-routing readback.
- **Lever 3 - act-quant fusion: FLAT.** The W4A4 act-quant tax is removable only
by W4A16 (a precision change, rejected) or a structural kernel rewrite; no
further bit-exact lever clears it. 0023 already banks the de-dup.
- **Lever 4 - NVFP4 the bf16 GDN/attn projections: REJECTED (KL-gate fail).**
Quantizing the projections to NVFP4 costs ~+6% PPL; vLLM deliberately keeps the
same bf16 projections. No-ship.
- **W4A16-Marlin MoE GEMM: REJECTED.** It would be a precision upgrade nobody
needs bought with a ~5% slower kernel; both kernels are already at the BW floor.
(The "the win was NVFP4-dense-quant, not the Marlin kernel" dense verdict
carries over to MoE.)
- **Chunked parallel-scan GDN prefill (patch 0031): the scalar-serial form was
FLAT-to-SLOWER at C=16 - the tensor-core M5 form (patch 0047) is the win,
now DEFAULT-ON under paged KV.** 0031 implements the upstream "faster pre-fill"
TODO - the FLA-style chunked gated-delta-rule (intra-chunk delta rule solved in
parallel via the UT-transform + forward substitution, inter-chunk recurrence
over n_tokens/C steps), math validated equivalent (numpy f32 NMSE ~1e-13;
`test-backend-ops` within the 1e-7 NMSE gate, a NEW per-path result). **But
GB10's 99KB dynamic-smem opt-in forces C=16** (the 128x128 f32 state alone is
64KB of the all-shared layout); the scalar-serial scan (`GDN_TC=0`) was then
pinned to 1 block/SM with serial per-thread dk-reductions and measured **~761
t/s chunked vs ~971 t/s sequential (~22% slower)**, grid-starved at low n_seqs.
The lesson held: **at this head dim the win needs tensor cores, not just
chunking.** Patch 0047 builds those tensor-core forms (KK/QK Gram = M2, KS/QS
state-boundary 3xtf32 = M3, P*U output = M4, full form-T solve + state-update
mma = M5, all `GDN_TC`-selected in one build) and ships **M5** as the default
when `LLAMA_KV_PAGED` is set. It is an f32/tf32-only re-port: the bf16/hybrid
dev-tree machinery (from the dropped 0026 ssm_bf16_tau) and the bf16 CONFIG-C
(M8) plus register-resident M6/M7 variants are NOT part of this series. M5 is the
variant that beats the (already 84.7%-of-peak) sequential scan while staying on
the bit-exact gate: MoE prefill S_PP **+3.5% @npp512 (3x interleaved A/B), +17.7%
@npp2048**; decode S_TG unchanged (the tuned `GDN_CHUNK_MIN=64` engage threshold
is > 1, so the 1-tok decode steps never enter the chunked path - at
`GDN_CHUNK_MIN=1` the chunked path swallows decode and collapses S_TG ~25%, the
reason the threshold is the lever). Bit-exactness is per-path benign:
`test-backend-ops` GATED_DELTA_NET is **94/94** vs CPU with M5 forced (incl.
multi-chunk n_tokens up to 256); the greedy md5 default-on == M5-forced ==
canonical on the short gate prompt (paged-MoE `8cb0ce23`, dense `5951a5b4`); on
a long MoE prompt (where the default fires M5 at >=64 tokens) M5 and the
sequential path agree word-for-word until **one** benign greedy token-flip
("the User:" vs "the User's Request:"), the dense model not flipping at all -
the textbook reduction-order flip greedy amplifies, NMSE-validated. The chunk
geometry stays env-selectable (`GDN_TC`/`GDN_CHUNK_C`/`GDN_DV_TILE`) for further
tuning; M5 is the shipped default because it wins without losing the canonical gate.
- **GDN occupancy retune (patch 0022) was a decode win but an UNCONDITIONAL
dense-prefill regression - now gated by scan length (patch 0046).** Patch
0022's `(NUM_WARPS=16, COLS_PER_WARP=8)` column-fold of the GDN
sequential-recurrence dispatch (`case 128`) raises per-warp memory-level
parallelism on the short, wide DECODE scans (small `n_tokens`, large
`n_seqs`) - the measured +11.1% dense decode win. Applied unconditionally it
also hit the dense PREFILL path, where the scan is long and narrow: the launch
`grid.z` collapses from `S_v/4 = 32` to `S_v/(16*8) = 1`, the SMs starve, and
profiling attributed the whole ~-6% dense-prefill regression vs stock to
`gated_delta_net` (+54% GPU time at the (16,8) geometry). Patch 0046 gates the
geometry by per-call scan length: long scans (prefill,
`n_tokens >= GDN_PREFILL_NTOK`, default 256) take stock's high-grid.z `(4,1)`
geometry; short scans (decode) keep the `(16,8)` retune. That recovers dense
prefill +7.2% back to stock parity while keeping the decode win, and it is
bit-exact: patch 0022 already proved every selectable `{NUM_WARPS,
COLS_PER_WARP}` variant is byte-identical (the sweep cannot change the md5), so
switching geometry by scan length cannot move the greedy output. The explicit
`GDN_NW`/`GDN_CPW` one-build %peak sweep still overrides (the gate yields when
either is set), so the A/B harness is unchanged.
**Opt-in bf16-SSM fast mode - DROPPED (was patch 0026, `ssm_bf16_tau`).** The
design premise - that bf16 KL error concentrates in long-memory heads and can be
removed by keeping them f32 - was already shaky: the error scales with the bf16
head *count* and saturates (~0.06 MeanKLD / ~91% same-top-p) far below any useful
byte saving. The lever was then **removed entirely** once the decode fusions
(0028 recurrent-state gather-fusion + 0029 block-table cache) landed: a clean
re-measurement that forced **all** gated-DeltaNet heads to bf16 (`tau=100000`,
the most aggressive setting) gave **flat** decode throughput - **780.6 vs 780.0
t/s**. The mode engages but buys **zero** speed; the earlier "+12%" was subsumed
by the fusions. So bf16-tau was a precision trade (not bit-exact) plus extra bug
surface and CUDA template-instantiation compile cost with **no** offsetting
benefit, and patch 0026 was dropped from the series. Lesson recorded so it is not
re-tried: do not reintroduce a per-head SSM-precision lever - the bandwidth it
targeted is already recovered by the gather-fusion + block-table cache.
---
## 6. Architecture and quant generality
(From the arch-generality and quant-generality audits.)
- **15 of 16 optimizations are quant-AGNOSTIC.** Only **0023** (NVFP4
activation-quantize de-dup) is NVFP4-specific. The SSM/paged/MMQ optimizations
help **any quant** of these models (the GDN recurrence, conv, gather and
o_proj-MMQ levers operate on the f32 recurrent state and the routing layout,
not on the weight dtype).
- **Arch-safe to build everywhere.** NVFP4 use is Blackwell-gated and falls back
to dequant on other hardware; the GB10-tuned occupancy params (0022) are
perf-only and env-selectable (`GDN_NW` / `GDN_CPW`), so they never change
correctness on other GPUs. Patch 0030 makes the fused-op emission CUDA-family +
CPU only, so a non-CUDA paged build routes to the safe upstream non-fused path.
- **What generalizes beyond this backend (upstream candidates).** The *speedups*
are CUDA/Blackwell-specific (which is why Metal/Vulkan don't benefit - section
4c), but several *findings and ops* are portable and worth upstreaming:
- The headline is hardware-independent: on hybrid gated-DeltaNet models, decode
is bottlenecked by the recurrent-state **plumbing** (memcpy + gathers, ~67% of
the step), not the weight GEMM. The fusions for it (in-place state 0018, gather
0019/0028, conv 0021) are bit-exact and already have CPU reference kernels, so
they would speed up Qwen3.6 / Qwen3-Next / any hybrid-SSM decode on **every**
backend once the ggml ops gain the respective (Metal/Vulkan) kernels - the
highest-value upstream contribution.
- The o_proj GEMV->MMQ reshape (0020) is a model-graph fix (batch the projection
to hit the GEMM path) - arch-agnostic in principle, trivial to upstream.
- The paged KV + cross-request prefix sharing + decode-first scheduler align with
llama.cpp's own in-progress KV / chunked-prefill work and could inform it.
- The per-path bit-exact md5 gate + the weekly upstream-drift canary is a reusable
maintenance pattern for any vendored-patch backend.
---
## 7. Pin + maintenance policy
- **Canonical source = the fork branch `mudler/llama.cpp:localai-paged`.** The
vendored `patches/paged/*.patch` files are now generated (one `git format-patch`
per commit) from that branch, which is the pin commit plus the paged patch
commits in order, so there is no more hand-export drift between the dev tree and
the shipped series.
- **Pinned to llama.cpp `0ed235ea2c17a19fc8238668653946721ed136fd`** (kept == the stock `llama-cpp` pin). The pin
is advanced **only** by the manual pin-sync process (this section):
rebase the source-only patch series onto the new tip, rebuild on GPU, pass the
bit-exact gate on every path (dense + MoE, paged + non-paged) plus
`test-backend-ops`, **and confirm the full grpc-server build links on CI**.
- **The pin must track the stock pin.** `grpc-server.cpp` is shared with the stock
backend and tracks the stock pin, so a paged pin that diverges past an upstream
server-API refactor breaks the grpc-server LINK even when the patches are
bit-exact. A bump to `c299a92c` (23 commits ahead of stock) was greedy-md5
bit-exact but failed to link (undefined `stream_*` server helpers introduced by
the refactor), and was reverted to the then-current stock pin. The bit-exact gate alone does not
catch this; only the full CI grpc-server build does.
- **Decoupled from the nightly auto-bumper.** There is deliberately **no**
`bump_deps.yaml` entry for this backend - a naive `LLAMA_VERSION` bump could
silently shift the tree out from under the patches.
- **Weekly canary.** [`.github/workflows/llama-cpp-paged-canary.yml`](../../../.github/workflows/llama-cpp-paged-canary.yml)
(via [`.github/scripts/paged-canary-apply.sh`](../../../.github/scripts/paged-canary-apply.sh))
tries the patch series against the latest upstream tip with the build's own
strict `git apply`. **Red = upstream drifted past the series -> run a
PIN_SYNC** (do not bump the pin blindly), following the policy in this section.
---
## 8. Models
> **Build coverage: CUDA-only.** This backend ships only the CUDA/cublas build
> targets (cuda-12, cuda-13, and the nvidia-l4t arm64 cuda-12/cuda-13 Jetson
> rows). There are no cpu / vulkan / sycl / hipblas / metal-darwin builds: the
> patchset's wins are CUDA/Blackwell-specific (section 4c), so off-CUDA the
> backend is neutral-to-negative and non-CUDA users should run the stock
> `llama-cpp` backend instead. The `backend/index.yaml` meta-backend resolves
> `default`/`nvidia` to a CUDA variant accordingly.
The benchmarked NVFP4 GGUFs are published and wired into the LocalAI gallery:
| Gallery entry | Weights (HuggingFace) | Notes |
|---|---|---|
| `qwen3.6-27b-nvfp4-paged` | [`mudler/Qwen3.6-27B-NVFP4-GGUF`](https://huggingface.co/mudler/Qwen3.6-27B-NVFP4-GGUF) | Dense, native Blackwell NVFP4 (FP4-MMA). |
| `qwen3.6-35b-a3b-nvfp4-paged` | [`mudler/Qwen3.6-35B-A3B-NVFP4-GGUF`](https://huggingface.co/mudler/Qwen3.6-35B-A3B-NVFP4-GGUF) | MoE (256 experts, top-8), `file_type MOSTLY_NVFP4`. |
Both gallery entries set `backend: llama-cpp-localai-paged` and the paged serving config
(`paged_kv:true`, `max_batch_tokens`, `kv_unified:false`, `parallel`,
`flash_attention:on`, `context_size`). They are bit-exact. The full
backend-split + gallery plan is in
[`LOCALAI_LLAMACPP_BACKEND_PLAN.md`](docs/LOCALAI_LLAMACPP_BACKEND_PLAN.md).
---
## 9. vLLM parity - final state (CLOSED)
> 2026-07-01 follow-up: the investigation was reopened for MTP safety,
> MTP-serving, graph-shape tracing, and a current-stack serving snapshot. Phases
> 14-20 are recorded in
> [`docs/GB10_PARITY_PHASE0_RESULTS.md`](docs/GB10_PARITY_PHASE0_RESULTS.md) and
> [`docs/PARITY_HANDOFF.md`](docs/PARITY_HANDOFF.md). They did not change the
> GB10 conclusion: MTP/scheduler shortcuts are rejected, and the latest clean
> stack remains below vLLM serving parity.
The multi-week GB10 (DGX Spark, sm_121) vLLM-parity investigation is **closed**.
The standing, never-re-litigate record - full benchmark, every lever and verdict,
the structural floors, the parity verdict - is
[`docs/VLLM_PARITY_FINAL.md`](docs/VLLM_PARITY_FINAL.md). Summary:
- **Where we are (GB10, Qwen3.6 NVFP4, vs vLLM 0.23.0).** Decode: dense is
**ahead of vLLM at low concurrency (116.7% at N=8)** and both models are
bandwidth-floored at **~56-68% of vLLM at high concurrency**. Prefill is
**~36% (MoE) / ~43% (dense)** of vLLM. Memory: **1.5-3x lower** than vLLM
(NVFP4-resident; vLLM's peak is a fixed ~109-112 GB 0.85-util reservation,
paged grows with KV from ~50 GB). Output is bit-exact per-path
(`5951a5b4` dense, `8cb0ce23` paged-MoE).
- **Why the residual is a hardware ceiling, not missing work.** Decode kernels
are already **5.4x more GPU-efficient per token** than vLLM's; the gap is the
**LPDDR5x ~273 GB/s** floor. The prefill GEMM is **FP4-MMQ-optimal** (every
alternative - 0033 dequant->cuBLAS, 0034 native FP4-MMA, 0035/Marlin W4A16,
offline-repack and vLLM-verbatim Marlin - was rejected; bf16 TC peak is ~half
FP4 peak, and vLLM itself runs a bf16-Marlin fallback on sm_121). The GDN
chunked scan is at the tractable tensor-core win (**M5 tf32**, patch 0047);
its residual is the **O(C^2) intra-chunk solve + serial recurrence** (occupancy
and dtype proven not the bound: BV -1%, bf16-C64 -18.75%). The serving host
loop is **closed** (~0-1% of the wall; padded-decode built + rejected).
- **Shipped, bit-exact wins.** FP4-MMQ GEMM, M5 tensor-core GDN prefill (0047),
fused residual+RMSNorm (0042), fused GatedRMSNorm+SiLU (0044), GDN-prefill
geometry gate (0046), the SSM decode-fusion stack (0018-0022/0028, up to
2.46x/2.26x over stock), decode-graph reuse (0040/0043), the memory advantage,
and low-N decode lead.
- **The path to parity is different hardware.** Datacenter Blackwell (HBM,
native tcgen05/CUTLASS FP4) lifts the bandwidth floor and **restores exactly
the vLLM advantages that lose on GB10** (FLA blocked-solve GDN, Marlin/CUTLASS
grouped FP4, HBM-tuned full-cudagraph decode). Re-run the methodology on new
silicon; do not reopen the GB10 levers.
Latest current-stack MoE serving snapshot (`PTOK=128`, `GEN=64`, current clean
DGX mirror `f2521ab12`, artifact
`/home/mudler/bench/phase26_audited_snapshot/20260701_053650`). This run
includes `hardware.txt` and `gate_summary.tsv`; all pre/post gate rows are
`ok`:
| n | paged decode_agg | vLLM decode_agg | paged/vLLM decode | paged agg | vLLM agg | paged/vLLM agg |
|---|------------------|-----------------|-------------------|-----------|----------|----------------|
| 8 | 230.8 | 283.2 | 81.5% | 170.6 | 241.6 | 70.6% |
| 32 | 420.0 | 609.0 | 69.0% | 254.6 | 466.7 | 54.6% |
| 128 | 673.4 | 1025.0 | 65.7% | 324.0 | 656.5 | 49.4% |
Use `paged-current-serving-snapshot.sh` for future current-stack GB10 serving
snapshots. It targets the clean `~/llama-phase6-source` mirror, checks
docker/`local-ai-worker`/GPU-idle state, uses the owner-file lock, runs pre/post
inference gates, writes `hardware.txt`, emits `gate_summary.tsv`, and emits
paged/vLLM ratios.
`hardware.txt` records the GPU identity and hardware class so GB10/workstation
Blackwell evidence is not confused with a future datacenter-Blackwell rerun.
`gate_summary.tsv` records pre/post MoE md5, dense md5, and backend-op checks
so an artifact proves inferencing gates without reading full logs.
Do not use the stale DGX
`~/bench/combined_definitive.sh` without first porting it to the current mirror
and lock discipline.
Phase 28 challenged the remaining low-conflict NVFP4 grouped-MMQ occupancy
knobs on the same DGX mirror
(`/home/mudler/bench/phase28_mmq_occupancy/20260701_040450`). The only buildable
variant, `GGML_CUDA_FP4_MINBLOCKS=2`, was inference-safe before and after
serving (MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID 806/806`) but regressed
n128 decode serving (`705.1 -> 689.9` decode_agg_tps, `0.9784x`). The row-tile
knob `GGML_CUDA_FP4_MMQ_Y=64` failed the NVFP4 writeback compile-time
invariant. Do not promote these knobs; grouped-MMQ parity work now requires a
structural kernel change, not launch-bounds or row-tile tweaks.
Phase 29 added the default-off grouped-MMQ shape trace as patch `0056`
(`/home/mudler/bench/phase29_mmq_shape_trace/20260701_042428`). The helper was
added test-first (`test-cuda-mmq-shape-trace`), compiled under CUDA on DGX, and
kept inference stable with the trace disabled and enabled:
MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID 806/806`. Example trace line:
`[LLAMA_MOE_MMQ_SHAPE] type=40 moe=1 ncols_dst=104 nchannels_x=256 ncols_max=13 n_active_est=104 density=1 mmq_x_max=128 mmq_x_lim=64 mmq_x_best=16 mmq_y=128 stream_k=1`.
Phase 31 extended that trace as patch `0057`
(`/home/mudler/bench/phase31_mmq_launch_trace/20260701_064424`) with
`[LLAMA_MOE_MMQ_LAUNCH]` lines from `launch_mul_mat_q`. Default-off,
trace-enabled, and post-serving gates stayed stable: MoE `8cb0ce23`, dense
`5951a5b4`, `MUL_MAT_ID 806/806`. The n128 serving trace showed decode-like
`4800/4800` and prefill-like `4920/4920` launch lines with `fixup=0` and
`stream_k_blocks == ntiles_dst`, rejecting a no-fixup/no-stream-k shortcut for
this workload.
Phase 32 added the small-M classifier trace as patch `0058`
(`/home/mudler/bench/phase32_small_m_classifier/20260701_070127`). Default-off,
trace-enabled, and post-serving gates stayed stable: MoE `8cb0ce23`, dense
`5951a5b4`, `MUL_MAT_ID 806/806`. The n128 serving trace found 4096 small-M
candidate calls: `mmq_x_best=64` 1800, `48` 1096, `40` 360, `32` 360, `16`
360, `24` 120. This justifies Phase 33 as a default-off tile-policy A/B
(`mmq_x=16`, possibly `8`) rather than a broad kernel rewrite.
Phase 33 added default-off `LLAMA_MOE_SMALL_M_TILE=<n>` as patch `0059`
(`/home/mudler/bench/phase33_small_m_tile_policy/20260701_071136`). The knob is
md5/op safe, but both tested values were slower in same-session n128 serving:
baseline `672.1` decode_agg_tps, tile16 `640.3` (`0.953x`), tile8 `583.2`
(`0.868x`). Do not promote simple smaller `mmq_x` caps for this workload.
Phase 34 added default-off `LLAMA_MOE_MMID_ROUTE_TRACE=<n>` as patch `0060`
(`/home/mudler/bench/phase34_mmid_route_trace/20260701_072737`). Default-off,
trace-enabled, and post-serving gates stayed stable: MoE `8cb0ce23`, dense
`5951a5b4`, `MUL_MAT_ID 806/806`. Live n128 serving with trace cap 4096 produced
`mmq=2776`, `mmvq=1320`, and `host_sync=0/4096`; the top shapes were
`mmq ne2=12` (1096), `mmq ne2=18` (480), and `mmvq ne2=8` (360). This refutes
host-sync fallback as the current n128 `MUL_MAT_ID` problem; follow-up work should
target grouped-MMQ small-M kernel partitioning or another measured bucket.
Phase 35 added default-off `LLAMA_MUL_MAT_ROUTE_TRACE=<n>` as patch `0061`
(`/home/mudler/bench/phase35_mul_mat_route_trace/20260701_074359`). Default-off,
trace-enabled, and post-serving gates stayed stable: MoE `8cb0ce23`, dense
`5951a5b4`, `MUL_MAT 1146/1146`, `MUL_MAT_ID 806/806`. Live n128 serving with
trace cap 8192 produced route counts: `mat_f=2888`, `op_cublas=2292`,
`mmq=1328`, `vec_q=1214`, `vec_f=470`. BF16 (`type=30`) dominated the trace
with `mat_f=2485` and `op_cublas=1330`; top BF16 shapes were `mat_f ne1=12`
(775), `op_cublas ne1=18` (760), and `mat_f ne1=8` (570). Next projection work
should trace or optimize the BF16 `op_cublas`/`mat_f` split, not batched cuBLAS.

View File

@@ -1,374 +0,0 @@
# Accelerator-porting scope: bringing the paged backend's portable benefits to Metal / SYCL / Vulkan (+ a ROCm note)
Source-only analysis (no GPU, no build) of which `llama-cpp-localai-paged` benefits
are portable off the CUDA family, and what each port costs per accelerator. This is
the umbrella doc; it BUILDS ON, and does not repeat,
[`UPSTREAM_LAYER2_SCOPE.md`](UPSTREAM_LAYER2_SCOPE.md) (the GDN/SSM fusion kernel
scope) - that doc remains the authoritative reference for benefit #1 below.
The backend ships **CUDA-only** today (README sections 4c, 8): off-CUDA the fusions
gate off (patch 0030) and NVFP4 falls back to dequant, so it is
neutral-to-slightly-negative there and non-CUDA users run the stock `llama-cpp`.
"Porting the benefits" is the upstream-contribution track that would make these
wins real on the other accelerators. Methodology for the work itself is in
[`.agents/vllm-parity-methodology.md`](../../../../.agents/vllm-parity-methodology.md).
We have **no Metal / SYCL / Vulkan / ROCm hardware here**, so every port is gated
by `test-backend-ops` (backendX-vs-CPU) **on the target hardware** - the same gate
discipline the existing layer-2 doc sets out.
--------------------------------------------------------------------------------
## 0. The four benefits and their portability class
| # | Benefit (patches) | Portable off CUDA? | Where scoped |
|---|---|---|---|
| 1 | **GDN/SSM decode fusions** (0018-0022, 0028) - in-place state write-back, fused recurrent-state gather, conv-state in-place fusion, o_proj MMQ reshape, occupancy retune | YES - per-backend KERNEL work | [`UPSTREAM_LAYER2_SCOPE.md`](UPSTREAM_LAYER2_SCOPE.md) (consolidated in section 1 here) |
| 2 | **Paged KV in-kernel block-table flash-attn read** (0009-0011) | YES - per-backend KERNEL work | **Section 2 here (the new analysis)** |
| 3 | **Decode-first prefill scheduler** (0013/0016) | YES - FREE, host-side, zero kernel work | Section 3 here |
| 4 | **NVFP4 FP4-MMA + its decode levers** (0017/0023/0025) | NO (Blackwell FP4-MMA) - out of scope; two analogues flagged | Section 4 here |
The two kernel-bearing tracks (#1 and #2) share an identical port SHAPE - they touch
the same decode kernel(s), the same `supports_op`, the same dispatch guard, and
sequence the same way (ops-first PR, then one PR per backend). They should be
**bundled into one per-backend PR**, not pursued as two separate efforts; section 5
sequences them together. Tracks #3 (free) and #4 (out of scope) are independent.
--------------------------------------------------------------------------------
## 1. Benefit #1 - GDN/SSM decode fusions (consolidated; full scope is the layer-2 doc)
Do not re-derive this here. [`UPSTREAM_LAYER2_SCOPE.md`](UPSTREAM_LAYER2_SCOPE.md)
already establishes, and this doc adopts wholesale:
- The base `GGML_OP_GATED_DELTA_NET` + `GGML_OP_SSM_CONV` + `GGML_OP_SSM_SCAN`
kernels **already exist on Metal, Vulkan AND SYCL**, so the Qwen3.6 hybrids RUN
on all three today via the upstream non-fused path. Layer-2 is the decode
SPEEDUP, not "make it run." (NB: the README section 4c no longer carries the
stale "no Vulkan kernel" line that the layer-2 doc section 0 was written to
correct - that correction has since been folded into the README, so treat
layer-2 section 0 as historical context, not a live correction.)
- The four fusion ops (A in-place state 0018, B fused state gather 0019, C
conv-update in-place 0021, D conv-tap gather 0028) reuse the existing op enums
with extra `src[]` discriminators; only OP C is a genuinely new kernel, the rest
redirect the read source / write target of the EXISTING kernel. The builders,
CPU reference kernels, model graph and `test-backend-ops` cases are SHARED and
already done.
- Per-backend net-new work, effort and gotchas: **SYCL easiest** (near-verbatim
CUDA mirror, ~250-350 LOC, no shader-gen), **Metal medium** (~350-500 LOC,
fixed 32 simdgroup = simplest bit-exactness), **Vulkan hardest** (~450-650 LOC +
shaders-gen + descriptor growth + per-vendor subgroup validation).
- Bit-exactness is per-backend BY CONSTRUCTION (the fusions redirect addresses, not
the f32 reduce order); gated by `test-backend-ops` (backendX-vs-CPU).
- Upstream path: ops-first PR (incl. the capability-driven replacement for patch
0030's backend-name allow-list), then one PR per backend.
The value/effort ranking from that doc (**Metal 1st, SYCL 2nd, Vulkan 3rd**) is
adopted unchanged here and, as section 5 shows, coincides with benefit #2's ranking
- which is why the two bundle cleanly per backend.
--------------------------------------------------------------------------------
## 2. Benefit #2 - paged KV in-kernel block-table flash-attn read (NEW scope)
### 2.0 What it is, and why it is the lever that makes paged KV non-negative off-CUDA
On CUDA, patches 0009-0011 replaced the per-step host-side K/V gather (patch 0003)
with an **in-kernel paged read**. `ggml_flash_attn_ext` gained an optional
`src[5]` = an I32 block table `[n_view, n_stream]` in token-POSITION order; the
fattn vec/tile kernel maps logical KV index `j` to physical cell
`block_table[seq*ne11 + j]` and reads `K0 + cell*nb11` / `V0 + cell*nb21` in place,
so the `get_rows` of K and V (the bulk of the gather) is gone. A null block table is
the stock contiguous read, byte-identical. Position ordering keeps the online-softmax
reduction order identical to stock, so it is bit-exact (CPU/batch1) by construction.
The crucial point for portability: **the entire host side is already
backend-agnostic.** The block-table fill (`llama_kv_cache::get_block_table`), the
K/V views, the mask compaction, the `input_block_table` graph input, and the
`ggml.c` / `ggml.h` builder (`ggml_flash_attn_ext_set_block_table`) all live in
`src/` and `ggml/...` shared code. The ONLY per-backend work is, in each backend's
flash-attn kernel: (a) thread one extra source through to the kernel, and (b) do the
indexed read at the K/V load sites. The CPU reference already does it (patch 0009,
`ops.cpp`).
Off-CUDA today the paged path falls back to the **host-side gather** (patch 0003),
which the README section 4c measured as neutral-to-slightly-negative on the M4
(~0-3% slower decode / ~2-8% slower prefill vs stock's contiguous read - pure
overhead, because the in-kernel read that *recovers* the gather cost is CUDA-only).
**Porting the block-table read is exactly what flips paged KV from
"neutral-to-negative" to "neutral-to-positive" off CUDA** - it removes the gather
overhead so paged KV's memory-management and prefix-sharing wins come for free
instead of at a decode tax. (The big decode multipliers on the hybrids are still the
benefit-#1 GDN fusions; this benefit is what makes the paged *allocator* pay its own
way off CUDA.)
### 2.1 The cross-cutting finding (applies to all three backends)
The indexed per-cell read only fits the **vec / scalar decode kernel**. Every
backend's FAST attention path - CUDA mma, Metal `simdgroup_load` MM, Vulkan
coopmat2, SYCL tile - loads K/V as **contiguous tiles** (8-cell `simdgroup_load`,
`coopMatLoadTensorNV` over a linear stride, shared-memory tile loads) that cannot
express an arbitrary per-cell gather without a staging pre-pass. This is exactly why
the CUDA port (0009-0010) wired ONLY the vec kernel and added a dispatch guard
(`if (dst->src[5]) force vec`).
So each port mirrors that: **route any FA op carrying a block table onto the vec /
scalar kernel; leave the fast MM path contiguous-only**, and keep the null-table
contiguous read on the fast path untouched. The decode shape (1 query token/stream)
naturally lands on or near the vec/scalar kernel on all three, so this is a small
routing change, not a rewrite of the fast path.
### 2.2 SYCL - EASIEST (near line-for-line CUDA mirror)
- **Exists today:** `ggml-sycl/fattn-vec.hpp` is a DPCT-style near-verbatim mirror
of CUDA `fattn-vec.cuh`; the kernel signature ends in the same `nb11..nb33`
cluster the CUDA patch appends `const int* block_table` to (fattn-vec.hpp:65-76).
Args are passed by SYCL lambda value-capture - **no descriptor/binding/push-
constant bookkeeping at all** (strictly easier than CUDA). `supports_op`
(`fattn.cpp` -> `ggml_sycl_get_best_fattn_kernel`) needs no change to ACCEPT
`src[5]`.
- **Port shape (value: medium / effort: LOW):** append `const int* block_table`
to the kernel + `fattn_kernel_t` typedef + `lauch_kernel`/`launch_fattn`
(sourcing `dst->src[5]->data`); 3 read-site substitutions (K at line 318, V at
389 and 410): `K0 + block_table[seq*ne11 + k_VKQ_0 + i_KQ]*nb11`.
- **Two SYCL-specific gotchas:**
1. **Pointer pre-advance.** The vec kernel advances `K`/`V` by `k_VKQ_0` OUTSIDE
the inner read (fattn-vec.hpp:293-300), so `i_KQ`/`k` are tile-local. The port
must keep an UN-advanced base `K0`/`V0`, drop the per-iteration `K +=`/`V +=`
on the paged path, and reconstruct the absolute cell. Get this wrong and you
read the wrong cells with NO compile error.
2. **Dispatch guard is bigger than CUDA's.** f16-GQA decode routes to the TILE
kernel, not vec (`fattn.cpp:198-208` fall-through). Add
`if (dst->src[5]) return BEST_FATTN_KERNEL_VEC;` near the top of
`ggml_sycl_get_best_fattn_kernel`. The shared `fattn_kernel_t` typedef means
the tile kernel must gain a matching ignored `block_table` param (or split the
typedef) - a trivial chore.
- **Bit-exact:** sub-group width (16) is fixed and the indexed read does not touch
lane assignment, loop bounds, or the XOR-reduction stride - reduction order is
invariant, so the paged vec path is byte-identical to SYCL's own contiguous vec
path. Gate: `test-backend-ops` FLASH_ATTN_EXT (with a block-table case) on Intel
GPU.
### 2.3 Metal - EASY-MEDIUM (decode already routes to the vec kernel)
- **Exists today:** decode (1 query token/stream, GQA) dispatches to
`kernel_flash_attn_ext_vec` (`ggml-metal-ops.cpp` `..._use_vec`: `ne01 < 20`).
Metal IS a true vec-equivalent (not a single unified FA kernel), and the vec
kernel's quantized K/V branches ALREADY compute a per-cell base address
(`k + ((ic + NE*cc + ty)*nb11)`, ggml-metal.metal:6934 / V at :7045) - so a
per-cell indexed read is unambiguously admissible. `supports_op`
(`ggml-metal-device.m` FLASH_ATTN_EXT) inspects no src count, so `src[5]` is
accepted as-is.
- **Port shape (value: HIGH / effort: EASY-MEDIUM):** append a
`device const char * block_table` param after `dst` (**buffer index 8** for vec)
+ a kargs field + a `has_block_table` function-constant; reuse the existing
"bind dummy when null" idiom for a missing table; substitute the cell index with
`block_table[seq*ne11 + cell]` at the K reads (lines 6919/6934) and V reads
(7032/7045) - a localized rewrite of ~2 loops (the fast path must adopt the
per-cell base form the quantized branch already uses).
- **Gotcha:** the **non-vec MM kernel is a HARD blocker** -
`simdgroup_load(..., NS10, ...)` reads 8 physically-CONTIGUOUS KV cells as one
matrix tile (lines 6160 / 6339-6363); an arbitrary gather can't be a single
strided matrix load. Mitigate exactly as CUDA did: force any block-table op onto
the vec kernel in `..._use_vec` (ggml-metal-ops.cpp:2517); leave the MM path
contiguous-only. Also watch a NAME COLLISION: `kernel_flash_attn_ext_blk` is an
existing mask-skip optimization, NOT a paged block table.
- **Bit-exact:** fixed 32-wide simdgroup + address-only redirect = byte-identical to
Metal's own vec contiguous path. Gate: `test-backend-ops` on Apple Silicon.
### 2.4 Vulkan - MEDIUM (the fast NVIDIA decode path cannot do it)
- **Exists today:** three FA shaders - `flash_attn.comp` (scalar/vec),
`flash_attn_cm1.comp` (coopmat1, stages K/V through shared memory),
`flash_attn_cm2.comp` (coopmat2, the fast NVIDIA path). FA uses **7 descriptor
bindings (0-6)**; `supports_op` (`ggml-vulkan.cpp` FLASH_ATTN_EXT) checks
specific srcs only, no count check; but `src[5]` is **not even threaded today** -
`ggml_vk_flash_attn` stops at `src[4]` (ggml-vulkan.cpp:14537), so wiring it
through is part of the work.
- **Port shape (value: HIGHEST breadth / effort: MEDIUM):** add binding 7 in the
shader(s), bump `7`->`8` in the three `ggml_vk_create_pipeline` calls (:3997,
:4033, :4070) and the two dispatch subbuffer lists (passing a dummy when null),
and wrap the indexed read in one `phys_kv()` helper applied at the ~4 K + 2 V
load sites (flash_attn.comp; the logical index is the same `(j*Bc + ...)`
expression at every site).
- **Two gotchas, one structural:**
1. **Push constants are FULL.** `vk_flash_attn_push_constants` is exactly
128 bytes with a `static_assert(... <= 128)` (the Vulkan guaranteed minimum) -
**no room for a new field.** Signal "block-table enabled" via the existing
`Flags` spec constant (flash_attn_base.glsl, `constant_id=10`, already
bit-packed) - add a `BLOCK_TABLE_ENABLE` bit. The per-seq stride is already
`p.KV`; the seq index is derivable in `init_indices()`.
2. **coopmat2 (the fast NVIDIA GQA-decode path) is INCOMPATIBLE.** Its K/V load
is a hardware `coopMatLoadTensorNV` over a LINEAR stride
(flash_attn_cm2.comp:307-313/377-383); the decode callback only dequantizes,
it cannot remap the physical address. The indexed read drops cleanly into
**scalar** (which non-GQA decode already uses) and **cm1** (which stages
through shmem - remap the staging loop), but **not cm2**. With a block table
present, NVIDIA GQA decode falls back to scalar/cm1 (slower than cm2, still
correct); the **null-table path keeps using cm2 unchanged**. AMD/Intel (no
cm2) are fully covered by scalar/cm1.
- **Net positive?** Yes. Non-GQA decode already runs scalar (paged read ~free);
AMD/Intel covered; only NVIDIA GQA decode trades cm2 for scalar/cm1 *when a table
is supplied*, and paged KV's payoff is allocator/memory + prefix-sharing, not raw
FA throughput, so the trade is contained and the fast contiguous path is
untouched.
- **Bit-exact:** the read is a per-thread scalar load, subgroup-size agnostic
(already abstracted via the `SubGroupSize` spec constant); position ordering keeps
the reduction order identical, so byte-identical to the backend's own
scalar/cm1 contiguous path. **Build burden is low** - these are EXISTING shader
variants recompiling (no new `string_to_spv` shape), so no shaders-gen matrix
growth. Gate: `test-backend-ops` per vendor (AMD + Intel + NVIDIA).
### 2.5 Benefit-#2 ranking and the shared dispatch/supports_op pattern
| backend | value | author effort | structural risk | rank |
|---|---|---|---|---|
| SYCL | medium (Intel GPU) | **LOW** (line-for-line; no bindings) | low (pointer pre-advance; force-vec guard) | easiest |
| Metal | **HIGH** (largest non-CUDA base) | EASY-MEDIUM (decode = vec already) | medium (MM blocker -> force vec) | mid |
| Vulkan| **HIGHEST breadth** (AMD+Intel+NVIDIA) | MEDIUM (7->8 bindings; Flags bit) | medium (cm2 can't; full push-const) | hardest |
Common to all three (mirrors CUDA 0009-0010): (1) `supports_op` needs no change to
ACCEPT `src[5]`; (2) a **dispatch guard forces any block-table op onto the
vec/scalar kernel**; (3) the fast MM/coopmat2 path stays contiguous-only and the
null-table read on it is byte-identical to stock.
--------------------------------------------------------------------------------
## 3. Benefit #3 - decode-first prefill scheduler (FREE portable win, confirmed)
Patches 0013 (static `LLAMA_PREFILL_BUDGET`) and 0016 (dynamic decode-first
`max(n_ubatch, T-D)`) are **pure host-side scheduler policy inside `update_slots()`
with zero libllama / zero ggml-backend changes** (README sections 2, 3). They change
only the *count* of prefill tokens admitted per step; they touch no kernel, no
`supports_op`, no device code. They are therefore **already backend-portable with no
per-accelerator work** - they run identically on Metal, SYCL, Vulkan, ROCm, CPU.
Byte-identical when off (default-off / short prefill == upstream `-b` chunking).
This is the cheapest portable benefit: it needs no port at all, only the decision to
leave it enabled in the (currently CUDA-only) build, or to upstream the policy. The
only reason it is not "live everywhere" today is that the backend ships CUDA-only;
the code itself is accelerator-neutral. If the scheduler levers are upstreamed
independently of the kernels, they help any llama.cpp build on any accelerator at
once - the lowest-effort, broadest-reach contribution of the whole series.
--------------------------------------------------------------------------------
## 4. Benefit #4 - NVFP4 FP4-MMA (NOT portable) + two backend-agnostic analogues
The NVFP4 decode track is **Blackwell-specific and out of scope** for accelerator
porting: Metal, SYCL, Vulkan and ROCm/AMD lack native FP4-MMA (Metal `supports_op`
already excludes NVFP4 from `MUL_MAT`/`MUL_MAT_ID`/`GET_ROWS`; on non-Blackwell the
FP4 path dequants). Patch 0017 (dense FP4-GEMM occupancy tune) ships only as the
parity gate + default-off instrumentation even on CUDA, so there is nothing to port.
Two of the NVFP4 *decode levers*, however, have backend-agnostic analogues worth a
note (do not over-claim - these are observations, not scoped ports):
- **0023 (NVFP4 activation-quantize de-dup)** - the IDEA generalizes, the patch does
not. The MoE broadcast up/gate projections re-quantize the same token activation
once per expert; 0023 quantizes the unique activations once and byte-copies them
into the expert-gathered layout. Any backend whose MoE path requantizes a shared
activation per-expert (e.g. a Q8 activation-quant before an integer-dot MoE GEMM)
could dedup the same way. It is NOT NVFP4-specific in PRINCIPLE - but it IS the
one quant-specific patch in the series (README section 6), so a port is a
per-backend MoE-quant investigation, not a lift-and-shift. Low priority.
- **0025 (MoE decode re-graph / `LLAMA_MOE_FORCE_GRAPHS`)** - keeping the graph/
capture path on across the grouped-MMQ MoE decode step is a CUDA-graphs concept.
Metal/Vulkan/SYCL have their own command-buffer/graph reuse machinery; the
generalizable finding is "the grouped MoE decode step has no host sync, so it is
safe to keep in a captured/replayed command buffer." Whether each backend's graph
layer already covers this is a per-backend question. The methodology note (README
dev notes: graph/stream coverage was a FLAT lever beyond 0025 on CUDA) is the
more durable takeaway - do not expect a large graph-coverage win on any backend.
Neither analogue is on the critical path; both are recorded so the next person does
not mistake them for free ports.
--------------------------------------------------------------------------------
## 5. Combined sequencing and top recommendations
Benefits #1 (GDN fusions) and #2 (block-table FA read) share the port shape
(vec/scalar decode kernel + `supports_op`/dispatch guard + ops-first-then-per-backend
PR) and rank in the SAME order per backend. So sequence them TOGETHER, per backend,
behind one shared ops-first PR:
1. **PR #1 - OPS (largely done, upstreamable as-is):** the `ggml.h`/`ggml.c`
builders, the CPU reference kernels, the CUDA kernels, the `test-backend-ops`
cases (GDN fusions AND a FLASH_ATTN_EXT block-table case), and the
**capability-driven gate** replacing patch 0030's backend-name allow-list (make
`supports_op` + the dispatch guard authoritative, so routing falls out of the
normal scheduler fallback and no backend name is hard-coded). Independently
mergeable.
2. **PR #2 - Metal:** GDN fusion kernels (layer-2 doc) + block-table read into
`kernel_flash_attn_ext_vec` + the force-vec routing guard. Gate on Apple Silicon.
3. **PR #3 - SYCL:** the near-verbatim CUDA mirror of both tracks + the force-vec
guard. Gate on Intel GPU.
4. **PR #4 - Vulkan:** GDN fusion shaders + the scalar/cm1 block-table read (cm2
stays contiguous, falls back when a table is present) + the `Flags` spec-constant
bit + the 7->8 binding bump. Gate per vendor.
Do NOT bundle the backends into one PR (each needs its own hardware for
`test-backend-ops`; reviewers are backend-specialized; a regression in one must not
block the others).
### Top recommendations
1. **Metal first, both benefits together.** Largest non-CUDA LocalAI base; the
decode shape already routes to the Metal vec kernel (block-table read is
EASY-MEDIUM there) and the base GDN/conv kernels already exist (fusions are
MEDIUM); fixed 32-wide simdgroup makes bit-exactness the simplest of the three.
Highest value at moderate effort.
2. **SYCL second as the cheap mechanical follow-on.** Both tracks are near
line-for-line CUDA mirrors with no binding/shader-gen bookkeeping, so it is
low-cost insurance even though the Intel-GPU audience is smaller. Budget the
effort on the two SYCL gotchas (pointer pre-advance; the force-vec guard since
f16-GQA decode routes to tile), not on plumbing.
3. **Vulkan last as the high-breadth capstone.** Reaches AMD + Intel + NVIDIA, but
carries the most host glue and the coopmat2 limitation (NVIDIA GQA decode trades
the fast path for scalar/cm1 only when a table is present). Do it once the
pattern is proven on Metal + SYCL.
A cheaper variant (from the layer-2 doc, reaffirmed): ship **Metal + SYCL together**
right after the ops PR and treat Vulkan as a separate later effort.
--------------------------------------------------------------------------------
## 6. ROCm note
ROCm is in the **CUDA family**, not a separate port: patch 0030's allow-list already
admits `"CUDA"/"ROCm"/"MUSA"`, and the CUDA kernels compile for HIP, so benefits #1
and #2 are largely already-built or near-free on ROCm rather than a from-scratch
accelerator port. Two caveats:
- **FP4-MMA (benefit #4) stays NVIDIA-Blackwell-only** - AMD has no native FP4-MMA,
so the NVFP4 path dequants on ROCm exactly as elsewhere.
- **The block-table read's force-vec routing matters on AMD too.** The AMD fast FA
path is the wmma/mma kernel (`fattn-wmma-f16`), which - like CUDA mma, Metal MM
and Vulkan cm2 - ignores the block table; the CUDA dispatch guard already forces a
block-table op onto the vec kernel, so ROCm inherits correct routing, but the
perf trade (vec vs wmma for AMD GQA decode with a table present) should be
measured on AMD hardware before claiming a win. The GDN fusions, being plain
CUDA-C, port to HIP with the rest of the CUDA path.
Net: ROCm is a "validate, don't re-port" follow-up - confirm the HIP build picks up
the fusions + the force-vec block-table routing and gate it with `test-backend-ops`
on an AMD GPU. It is genuinely separate from, and lighter than, the Metal / SYCL /
Vulkan ports.
--------------------------------------------------------------------------------
## 7. Summary
- **Benefit #3 (decode-first scheduler) is free and already portable** - host-side
policy, zero kernel work; it only needs to be left enabled / upstreamed.
- **Benefits #1 (GDN fusions) and #2 (block-table FA read) are the real ports** -
both are vec/scalar-decode-kernel + `supports_op`/dispatch-guard changes, both
rank Metal-then-SYCL-then-Vulkan, and they bundle into one per-backend PR behind a
shared ops-first PR.
- **Benefit #2 is the lever that makes paged KV non-negative off CUDA** - it removes
the host-gather overhead the README measured as neutral-to-slightly-negative on
the Mac. Feasibility: SYCL EASY, Metal EASY-MEDIUM, Vulkan MEDIUM. The universal
constraint is that only the vec/scalar kernel admits the indexed read; the fast
MM/coopmat2 path is contiguous-only, so route block-table ops onto vec (as CUDA
already does) and leave the fast path's null-table read byte-identical.
- **Benefit #4 (NVFP4 FP4-MMA) is out of scope** (Blackwell only); 0023's de-dup and
0025's graph-coverage have backend-agnostic *ideas* but no lift-and-shift port.
- **ROCm rides the CUDA path** (validate, don't re-port); FP4-MMA stays Blackwell-only.
- Everything is bit-exact per-backend BY CONSTRUCTION (position-ordered table +
address-only redirect = identical reduction order), gated by `test-backend-ops`
(backendX-vs-CPU) **on the target hardware**, which we do not have here.
</content>
</invoke>

View File

File diff suppressed because it is too large Load Diff

View File

@@ -1,422 +0,0 @@
# DECODE_SERVING_SCOPE - the continuous-serving decode gap
**Status: S1 + S3 IMPLEMENTED, GPU-validated, bit-exact, shipped as patches
0040 (S1) + 0041 (S3). S2 DROPPED (measured non-target). See the results block
below; the rest of this doc is the design/rationale those patches implement.**
## Results (GB10, measured)
Phase 0 confirmed host-bound: serving graph reuse **0% over ~5k steps** (layer-A
rebuilds every step), `hostproc` 3.44 ms/step vs 1.59 static - the +1.85 ms IS the
graph rebuild; `set_inputs` 0.047 ms and block-table 0.002 ms are negligible.
- **S1 (patch 0040)** - root cause: the paged decode inputs never overrode
`can_reuse` (defaults false), so the graph could never be reused. Fixed with a
256-bucketed-shape `can_reuse` + live-mctx refresh. Static batched-bench A/B:
paged decode reuse **0% -> 95.5%**, bit-exact (md5 byte-identical reuse on/off).
Necessary but **not** sufficient in serving (13.8% reuse alone - prefill
co-batching churns the shape).
- **S3 (patch 0041)** - keeps prefill out of decode steps so the scheduler emits
reuse-stable pure-decode steps. **S1+S3 together (128-client staggered serving,
MoE Qwen3.6-35B-A3B-NVFP4): reuse 0% -> 72.2%, `hostproc` 15.98 -> 6.31 ms/step,
decode 4.05 -> 5.52 tok/s/seq median (4.24 -> 5.96 mean, at vLLM's ~5.9).**
- **S2 (double-buffer set_inputs) - DROPPED.** Phase 0 put `set_inputs` at
~0.05 ms/step: it is not the cost (the rebuild is), so S2 has nothing to recover.
- **Follow-up to ~100% reuse - PADDED/FIXED-SLOT DECODE SHAPE: IMPLEMENTED,
GPU-TESTED, REJECTED (not shipped).** See the "Padded-shape lever - rejected"
block below. Summary: it does NOT close the serving gap. Padding holds the
pure-decode width constant by emitting masked-inert dummy decodes for idle
slots, and it is provably inert (single-seq md5 bit-exact + per-stream
noise-floor determinism), but it **regresses throughput at every concurrency**
(catastrophically at low load) because the serving decode here is
**GPU-compute-bound, not host-rebuild-bound** - so the dummy-row compute it adds
costs more than the graph-reuse it recovers. The original "remaining ~28% is
request-boundary churn -> pad it" hypothesis stands mechanically, but the payoff
premise (closing reuse pulls decode toward vLLM) is **not supported by
measurement**.
---
## Padded-shape lever - rejected (implemented + GPU-tested, 2026-06-28)
The S1 section-(a) **padded / fixed-slot decode shape** was implemented in an
isolated worktree off the committed S1/S3/tail base (paged HEAD `05eceb4`), built
CUDA-only, and benched on GB10. **Verdict: REJECTED - it regresses serving
throughput and does not close the vLLM gap.** Recorded here so it is not re-tried.
**Implementation** (default-off, `LLAMA_PAGED_PAD_DECODE=1`; `LLAMA_PAGED_PAD_WIDTH`
caps the slot range): at the end of `pre_decode()`, on any step where no prompt
tokens were admitted (`n_prompt_budgeted == 0`) and there is decode load, emit a
masked-inert dummy decode for **every IDLE slot** (`batch.add(slot.id, 0,
pos_max+1, /*output=*/true)`; cold slot -> fresh pos-0). This holds `n_tokens`,
`n_seqs`, `n_seqs_unq`, `n_outputs` and the participating seq-id SET constant
across arrivals/completions. A `release()`-side guard keeps a finished slot warm
under padding (else patch 0024's reclaim-on-idle frees its KV and the next-step
pos-0 re-warm churns paged-block allocation, destroying reuse). Each dummy is its
OWN sequence, so its recurrent (gated-DeltaNet) state is private and its paged
attention reads only its own cells; its logits are computed but never read
(`post_decode()` only consumes `slot.i_batch` of GENERATING slots).
**Gates.** (1) Single-seq greedy md5 **bit-exact PASS** - dense
`5951a5b4d624ce891e22ab5fca9bc439`, paged-MoE `8cb0ce23777bf55f92f63d0292c756b0`
(the lever lives only in `llama-server`'s `update_slots()`, never in
`llama-completion`). (2) **Per-stream serving determinism**: the literal
"ON-vs-OFF token sequences identical" gate is **unachievable** - concurrent
cuBLAS/FA decode is **not bit-reproducible run-to-run** even with padding OFF
(OFF-vs-OFF diverging streams: dense 3/16, MoE 8/16, lockstep K=16). The
**achievable inertness gate PASSED**: per-stream prefix-agreement ON-vs-OFF equals
the OFF-vs-OFF noise floor exactly (MoE 0.940/0.940, dense 0.812/0.812), i.e. the
dummy slots inject no systematic divergence beyond the pre-existing concurrent FP
noise. So padding is provably inert; it just does not help.
**Bench (MoE Qwen3.6-35B-A3B-NVFP4, GB10).** Burst h2h, decode tok/s/seq:
| n | S1+S3 | PAD | vLLM |
|-----|-------|------|------|
| 8 | 28.16 | 6.05 | 44.8 |
| 32 | 11.66 | 4.84 | 17.45|
| 64 | 7.16 | 4.33 | 11.07|
| 128 | 4.53 | 4.32 | 6.87 |
Staggered (`serve_bench.py` k=128 n=160 stagger0.25), aggregate decode tok/s and
graph-reuse: baseline (reuse 0%) **757.6**; S1+S3 (reuse 72%) **763.3**; **PAD
(reuse 38%) 558.0**.
**Why it fails (four independent reasons):**
1. **Serving decode is GPU-compute-bound, not host-rebuild-bound (this run).**
Baseline reuse 0% (757.6 agg) is statistically equal to S1+S3 reuse 72% (763.3
agg): `hostproc` is only ~4-8% of the per-step wall, so eliminating the host
graph rebuild buys ~nothing. (This **corrects the host-bound hypothesis** above
for this hardware: the earlier 542->762 host-bound delta did **not** reproduce
- it was GPU-state/contention variance, not a stable reuse effect.)
2. **Padding ADDS dummy-row compute** (full-width decode), costing throughput in
direct proportion to `pad_width - real_load`: catastrophic at low concurrency
(n=8: 28.16 -> 6.05, ~4.6x slower, because 8 real streams pay for a 128-wide
step).
3. **In continuous serving padding can't even hold the width constant**: arrivals
are perpetually mid-prefill, so the idle-slot count varies and reuse DROPS
72% -> 38% (the opposite of the goal). It only stabilises the pure-decode
*tail* of a burst (verified: width pinned at 64 as real decoders fell 49->5),
which is exactly where the dummy compute is most wasteful.
4. **The completion-driven batch shrink that padding prevents is itself a
throughput WIN** in a compute-bound regime (fewer real streams -> cheaper
steps -> survivors finish faster); forcing constant width forfeits it.
**Conclusion.** The residual burst gap (paged 4.53 vs vLLM 6.87 at n=128 ~= 66%)
is a **GPU-compute** gap (vLLM's MoE decode kernel + scheduler are ~1.3x faster on
aggregate), not a host-loop gap. A host-side graph-reuse lever cannot close it.
Do not re-pursue padded/fixed-slot shapes for throughput; if the host loop is ever
re-confirmed dominant on other hardware (re-run reason 1's baseline-vs-S1+S3 A/B
first), revisit - but only with an *adaptive* width matched to live load, never a
fixed pad-to-`--parallel`.
---
Per the
"profile-don't-assume" rule in
[`.agents/vllm-parity-methodology.md`](../../../../.agents/vllm-parity-methodology.md),
**Phase 0 (section 5) is to confirm the bottleneck on GPU before touching any
code.** Everything below the Phase-0 line is a hypothesis ranked by
value/effort/risk, not a measured result.
> **Regime warning (read first).** Every "decode is at the BW floor / ties vLLM"
> and "host scheduling loop is the structural residual" conclusion in
> [`README.md`](../README.md) section 5 was measured with **`llama-batched-bench`**:
> a STATIC serving width (fixed `npl`, all sequences in lockstep, constant
> batch shape every step). That is the **decode KERNEL** regime, and there the
> patch series is at parity (paged ~6.1 tok/s/seq vs vLLM ~5.9 at npl128). This
> document is about a **different regime**: real **continuous SERVING** through
> `llama-server`'s `update_slots()` loop, where requests arrive and complete
> asynchronously, the batch shape churns every step, and paged drops to ~3.7
> tok/s/seq (-39%) while vLLM sustains ~5.9. The gap is the **scheduler / host
> loop**, not the kernel. This is the serving analogue of the prefill-GEMM regime
> split called out in [`PREFILL_GEMM_SCOPE.md`](PREFILL_GEMM_SCOPE.md).
Cross-links: [`README.md`](../README.md) sections 2 (scheduler), 3 (patches
0008/0013/0016/0024/0025/0029), 5 (rejected levers - lever 2 graph coverage was
FLAT *in the static regime*; this doc reopens it for the *serving* regime);
[`.agents/llama-cpp-localai-paged-backend.md`](../../../../.agents/llama-cpp-localai-paged-backend.md)
(bit-exact gate);
[`.agents/vllm-parity-methodology.md`](../../../../.agents/vllm-parity-methodology.md)
(both-engine ground-truth, per-lever A/B, record-rejected-levers).
---
## 1. The two regimes, and why the kernel-parity result does not carry over
`llama-batched-bench` and a real serving workload exercise the **same decode
kernels** but **different host loops**:
| | `llama-batched-bench` (kernel regime) | `llama-server` continuous serving |
|---|---|---|
| batch shape per step | **constant** (fixed `npl`, lockstep) | **churns** (arrivals/completions, interleaved prefill) |
| participating seq-set | **fixed** for the whole run | **changes** as requests start/finish |
| graph reuse (see s.2) | holds after warmup -> 1 capture, replayed | breaks nearly every step -> rebuild + re-capture |
| measured | paged ~6.1 tok/s/seq ~ vLLM ~5.9 | paged ~3.7 vs vLLM ~5.9 (-39%) |
The README's decode parity, BW-floor, and "host loop is the irreducible
residual" findings are all **kernel-regime** findings. They prove the *kernels*
are not the serving gap. They do **not** prove the host loop is irreducible *in
serving* - the static bench holds the batch shape constant, which is exactly the
condition that lets both graph-reuse layers (section 2) stay hot. Serving
violates that condition. So the serving gap is reopened here as a host /
scheduler problem, orthogonal to the kernel.
---
## 2. Root-cause hypothesis (from source, pin `9d5d882d` + the dev tree)
There are **two independent graph-reuse layers**, and continuous batching breaks
**both** on nearly every step. This is the leading hypothesis for the -39%.
### 2a. Layer A - llama-context graph reuse (`can_reuse` / `allow_reuse`)
`llama_context::process_ubatch` (`src/llama-context.cpp` ~L1366) only **reuses
the built ggml graph** when `res->can_reuse(gparams)` holds. `allow_reuse`
(`src/llama-graph.h` ~L631) requires, among others:
```
ubatch.n_tokens == other.ubatch.n_tokens &&
ubatch.n_seqs == other.ubatch.n_seqs &&
ubatch.n_seqs_unq == other.ubatch.n_seqs_unq &&
ubatch.equal_seqs() == other.ubatch.equal_seqs()
// + (when equal_seqs) the participating sequence-id SET must match
```
In serving, `n_tokens` changes whenever the decode load `D` changes or a prefill
chunk is co-batched, and the **sequence-id set** changes whenever a request
starts or finishes. Either makes `can_reuse` return false, so `process_ubatch`
falls into the `else` branch: **rebuild the graph** (`model.build_graph`) +
`ggml_backend_sched_reset` + `ggml_backend_sched_alloc_graph` - full host-side
graph construction + allocation, **every step**. In batched-bench all sequences
are lockstep so `n_tokens`/seq-set are constant and `can_reuse` is true after
warmup (the `graphs reused = N` perf line is ~all steps).
### 2b. Layer B - CUDA graph capture (`ggml_cuda_graph_*`)
Even when layer A reuses, the CUDA backend re-checks
`ggml_cuda_graph_update_required` (`ggml-cuda.cu` ~L3367): it `memcmp`s every
node's `ne`, `nb`, and `src[]->data` pointers against the captured graph. Any
shape change -> `cudaGraphExecUpdate` / re-instantiate. Two serving-specific
triggers:
- **shape churn** (same root cause as layer A): different `n_tokens` -> different
node `ne` -> update required.
- **paged data-pointer churn**: when a co-batched prefill allocates new KV blocks
(or a finished sequence frees them), the per-step KV view tensors' `data`
pointers move, so even a constant-shape decode step can trip the `memcmp`. (The
block-table *contents* live in a fixed device buffer filled by `set_inputs`, so
the table tensor pointer itself is stable - 0029 keeps that cheap - but the K/V
cache views are not.)
Net: under serving, the GPU sits idle between launches while the host rebuilds
the graph (layer A) and re-instantiates the CUDA graph (layer B), then runs an
un-graphed `set_inputs` (H2D input copies) before each launch. vLLM avoids this
with **padded/bucketed decode batch shapes + piecewise CUDA graphs**: it pads the
decode batch to a fixed set of sizes and captures one persistent graph per
bucket, so the steady-state decode step is a single `cudaGraphLaunch` with no
host rebuild. Its scheduler is also a tight C++ loop with chunked-prefill
interleave that keeps the GPU fed.
### 2c. Per-step host work that runs un-graphed regardless (already instrumented)
The dev tree carries a built-in `[L5INSTR]` profiler (`src/paged-attn.cpp`,
hooks in `src/llama-context.cpp` and `src/llama-kv-cache.cpp`) that already
isolates the host buckets we care about, printed at process exit:
```
[L5INSTR] get_block_table n=.. sum=..ms mean=..ms | set_inputs n=.. mean=..ms | hostproc n=.. mean=..ms
```
- `hostproc` = `mctx->apply()` + graph reuse-check/rebuild + `set_inputs`, i.e.
the whole host window **before** `graph_compute` (it does NOT include the GPU
launch). Prior profiles put this near ~1.4 ms/step.
- `set_inputs` = the H2D input fills (positions, masks, block table, idxs).
- `get_block_table` = the paged block-table host build (0029 caches it
within-step; `LLAMA_PAGED_NO_BT_CACHE` A/B-toggles that).
If `hostproc` per step is a large fraction of the serving per-step wall time
(and the `graphs reused` count is low), the gap is host-bound, not kernel-bound.
### 2d. The serial-SSM host loop (named in README s.5, secondary here)
The gated-DeltaNet decode advances recurrent state per step; sampling cannot
start until logits land. The README already names this as a structural floor in
the *kernel* regime. It is the same in serving but is the *smaller* term - the
graph-rebuild/re-capture overhead (2a/2b) is the new, serving-specific cost the
static bench hides, and it is the one to attack first.
---
## 3. What the already-shipped scheduler patches do (and do NOT do)
These exist; understand them before proposing anything. **None of them touch the
two graph-reuse layers** - they target prefill freezing and burst collapse, not
steady-state decode-step host overhead. That is why the serving gap survives them.
| Patch | What it does | What it does NOT do |
|---|---|---|
| 0008 cross-request prefix-share (server loop) | Concurrent shared-prefix requests prefill only the divergent suffix (fewer prefill tokens). | Does not stabilise decode batch shape; does not graph-reuse. |
| 0013 `LLAMA_PREFILL_BUDGET` | Static per-step prefill-token cap (vLLM `--max-num-batched-tokens` analogue); flattens the ITL spike a long prefill inflicts on co-batched decode. | Ignores decode load; per-workload tuning; no effect on decode-step graph reuse. |
| 0016 dynamic decode-first budget | `max(n_ubatch, T-D)` leftover-after-decode budget + per-slot chunk cap; decode claimed first, auto-shrinks as `D` rises. Stops a prefill chunk from inflating the step past `T`. | **Still lets the per-step decode `n_tokens` and seq-set vary**, so it does not make the decode step graph-reusable; it shapes prefill admission, not decode-shape stability. |
| 0024 paged-pool burst-reclaim | Truncate/defrag/release KV blocks; fixes long-server prefill burst collapse (488->44->532 t/s). | Host accounting only; nothing about decode-step graph capture. |
| 0025 `LLAMA_MOE_FORCE_GRAPHS` | Keeps CUDA graphs ON for the grouped-MMQ MoE decode step (lifts the conservative `MUL_MAT_ID` graph-disable). | Helps the CUDA-graph *eligibility* of one op; does **not** make layer-A/B *reuse* hold across churning steps. It is necessary-not-sufficient: a step that rebuilds anyway gets recaptured regardless. |
| 0029 block-table within-step cache | `get_block_table` computed once per step, memcpy'd to other full-attn layers (-87/-91%). | Shrinks one `set_inputs`/`hostproc` sub-term; does not address rebuild/re-capture. |
**README s.5 "lever 2 (graph/stream coverage): FLAT"** was concluded **in the
static batched-bench regime**, where graphs already reuse - so more graph
coverage was correctly a no-op there. That conclusion does **not** apply to the
serving regime, where graphs do **not** reuse. This doc reopens graph coverage
**for serving only**; record it as a regime-scoped reopening, not a contradiction.
---
## 4. Ranked lever plan (hypotheses - gate on Phase 0 first)
Ranked by value/effort with bit-exactness/risk called out. All are **host-side /
scheduler** levers (no decode-kernel changes), so all are *bit-exact-safe by
construction* provided padding tokens are masked-inert and verified against the
per-path md5 gate.
### Lever S1 (TOP) - bucketed/padded decode-step shape for graph reuse
**Value: high (targets the dominant -39% mechanism). Effort: medium-high. Risk:
medium (correctness of padding inertness; seq-set churn is harder than n_tokens).**
Make the steady-state decode step present a **stable, bucketed shape** to both
reuse layers, mirroring vLLM's padded decode batch + piecewise CUDA graphs:
- Pad the per-step decode `n_tokens` (and the stream/seq count the graph sees) up
to the next bucket in a small fixed set (e.g. {power-of-two or fixed grid}), so
`allow_reuse` (layer A) and `update_required` (layer B) hold across steps with
the same bucket. Padding tokens are dummy, masked positions that contribute
nothing to any real sequence's logits.
- Bound the number of distinct live buckets so a handful of persistent CUDA
graphs cover steady decode (vLLM captures ~tens).
- Handle the seq-set component of `allow_reuse`: bucketing `n_tokens` alone is
insufficient because the *participating sequence-id set* must also match. Either
(a) pad to a fixed stream-slot layout so the seq-set is stable across arrivals
/completions, or (b) relax/extend the reuse key so a pure-decode step keyed on
bucket+slot-layout reuses regardless of which slots are occupied. (b) is the
higher-leverage but more invasive option.
Bit-exact gate: greedy md5 per path with padding ON must equal the recorded
references (`5951a5b4` dense, `8cb0ce23` paged-MoE); `test-backend-ops`
unaffected (no op changes). The risk is that masked/padded positions leak into a
real logit (off-by-one in the mask) - the md5 gate catches it.
### Lever S2 - overlap per-step host work with GPU decode (double-buffer inputs)
**Value: medium-high (recovers the `hostproc` window even when S1 partial).
Effort: medium. Risk: low (host-side reordering only, bit-exact-safe).**
Even with graphs reused, `set_inputs` (+ the pre-`set_inputs` sync) runs
un-graphed and serially *before* each launch (`hostproc` ~1.4 ms/step in prior
profiles). Overlap the host scheduling + input build of step N+1 with the GPU
decode of step N: double-buffer the input device tensors so the host can fill
N+1's inputs while N's graph is in flight, and prepare the next ubatch / block
table on the host concurrently. This is the llama.cpp analogue of vLLM keeping
the GPU fed. Strictly host-side, no numeric change -> bit-exact. (0029 already
banks part of this for the block table within a step; S2 extends it across
steps.)
### Lever S3 - graph-shape-stable scheduling (bridge from 0016)
**Value: medium (multiplies S1; low marginal value without S1). Effort: low-medium
(extends the existing 0016 policy). Risk: low (scheduler policy, bit-exact when
the decode result is unchanged).**
Extend the existing decode-first budget (0016) so the scheduler actively *prefers
graph-reusable steps*: keep prefill chunks out of the decode step (run prefill in
its own steps, or at a fixed chunk size) so the decode batch shape stays on a
bucket rather than being perturbed by interleaved prefill tokens every step. This
is the policy half of S1 - S1 makes a bucketed step reusable; S3 makes the
scheduler emit bucketed steps. Pair them.
**Rejected/deferred (record so they are not re-tried):**
- **More CUDA-graph *coverage* alone (the README lever-2 redo): still FLAT
without S1.** Forcing more ops graph-eligible (beyond 0025) does nothing while
layer A rebuilds the graph every step - the recapture dominates. Only valuable
*after* S1 makes reuse hold.
- **`GGML_CUDA_DISABLE_GRAPHS` / disabling graphs in serving: REJECTED a priori
as a fix** (it is an A/B *probe* for Phase 0, not a lever) - it removes capture
cost but also removes replay benefit; expected net-negative.
- **Precision levers (W4A16, bf16-SSM): out of scope** - this gap is host-bound,
not GEMM/BW-bound (see README s.5 rejections; do not reopen).
---
## 5. Phase 0 - confirm it is host-bound BEFORE building (run when the GPU frees)
Do NOT build any lever until this confirms host-bound. The dev tree already has
all the instrumentation; this is a measurement, not a code change. **One GPU
bencher at a time** (GPU-contention rule).
**Workload.** Real continuous serving, not batched-bench: run `llama-server`
(paged build) with the paged config and drive it with a steady concurrent
streaming load (e.g. a K-client async generator hitting `/completion` with
staggered arrivals so requests start/finish asynchronously - the regime
batched-bench cannot produce). Use the same models/flags as README s.4:
`-fa on -ngl 99`, `LLAMA_KV_PAGED=1` (+ `LLAMA_MOE_FORCE_GRAPHS=1` for MoE),
dense Qwen3.6-27B-NVFP4 and MoE Qwen3.6-35B-A3B-NVFP4. Pick K so the *effective
decode width* matches a static `npl` you have a kernel-regime number for (e.g.
~128) - that gives the apples comparison: static 6.1 vs serving 3.7 tok/s/seq.
**Signals to capture (all already exist):**
1. **Graph reuse rate.** The `graphs reused = N` perf line (`llama-context.cpp`
~L4146, from `data.n_reused`) over total decode steps. Hypothesis: ~100% in
batched-bench, near 0% in serving. This is the single most decisive number.
A/B with `LLAMA_GRAPH_REUSE_DISABLE=1` (forces the rebuild path) - if serving
is already near that floor, layer-A reuse is the gap.
2. **`[L5INSTR]` host buckets** (printed at exit): `hostproc`, `set_inputs`,
`get_block_table` mean ms/step. Compare serving vs batched-bench. A/B the
block-table cache with `LLAMA_PAGED_NO_BT_CACHE`.
3. **GPU-busy %** in a steady-state serving window via nsys (sum of kernel
durations / wall) and the **inter-launch host gap** (time between consecutive
`cudaGraphLaunch`/kernel launches). Hypothesis: batched-bench ~96-99% busy
(README/methodology note the early "low util" was a window artifact); serving
materially lower, with the gap ~= `hostproc`/step. *Watch the same window
artifact* the methodology warns about - measure a clean steady-state span.
4. **CUDA-graph re-instantiation count** - confirm layer B is also re-capturing
(nsys shows `cudaGraphInstantiate`/`cudaGraphExecUpdate` per step, or add a
host-side counter print - host-side only, no kernel code).
**Decision rule.** Host-bound (proceed with S1/S2/S3) if: serving `graphs reused`
is low AND `hostproc`/step is a large fraction of serving per-step wall AND
GPU-busy% drops vs batched-bench by ~the observed throughput ratio (~3.7/6.1).
If instead GPU-busy% stays high and per-kernel time grows, the cause is
elsewhere (e.g. serving runs a worse effective batch shape into the kernels) -
re-scope before building.
**Ground-truth vLLM (both-engine rule).** Capture vLLM at the same concurrency:
GPU-busy% / step cadence (nsys) and its scheduler step time. Confirm vLLM stays
GPU-bound (persistent graphs) where paged goes host-bound - that is the
direct evidence the gap is the host loop, and it sizes the achievable win.
---
## 6. Summary
- The serving gap (paged 3.7 vs vLLM 5.9 tok/s/seq, -39%) is a **host/scheduler**
problem, distinct from the decode **kernel** (at parity in batched-bench). The
README's BW-floor/host-loop-residual findings are kernel-regime and do not
bound the serving regime.
- Leading mechanism: continuous batching's **batch-shape + seq-set churn breaks
both graph-reuse layers** (llama-context `can_reuse`, CUDA `update_required`)
every step, so the GPU idles while the host rebuilds + re-captures + runs
un-graphed `set_inputs`. vLLM avoids this with padded/bucketed decode shapes +
piecewise CUDA graphs.
- The shipped scheduler patches (0008/0013/0016/0024/0025/0029) target prefill
freezing + burst collapse, **not** decode-step graph reuse - which is why the
serving gap survives them.
- Top levers (all host-side, bit-exact-safe): **S1** bucketed/padded decode-step
shape for graph reuse, **S2** double-buffer/overlap per-step host work, **S3**
graph-shape-stable scheduling (extend 0016). Gate everything on **Phase 0**:
the `graphs reused` rate + `[L5INSTR]` host buckets + nsys GPU-busy% in real
`llama-server` serving vs batched-bench, with vLLM ground-truthed at the same
concurrency.
</content>
</invoke>

View File

File diff suppressed because it is too large Load Diff

View File

File diff suppressed because it is too large Load Diff

View File

@@ -1,376 +0,0 @@
# GB10 vLLM Parity Reopen Spec
Status: scoped follow-up. This document intentionally challenges the current
`VLLM_PARITY_FINAL.md` conclusion that GB10 parity is closed. The final record is
still useful as a baseline, but the follow-up work must treat it as a hypothesis
to test, not as a proof of impossibility.
## Goal
Determine whether llama.cpp / ggml can close the remaining GB10 parity gap for
Qwen3.6 NVFP4 hybrid gated-DeltaNet models by porting or adapting concrete vLLM
implementation ideas, while preserving LocalAI's hard correctness gates.
Success means one of two outcomes:
1. A measured, source-backed path improves paged llama.cpp materially toward vLLM
parity on GB10.
2. The remaining gap is rejected with clean provenance: clean source, clean DGX
host state, artifact-pinned A/B results, and explicit correctness gates.
## Non-goals
- Do not accept a "closed" conclusion based only on existing docs.
- Do not run long builds or benchmarks without a recorded DGX preflight.
- Do not edit `patches/paged/*.patch` directly. Kernel changes land fork-first in
`mudler/llama.cpp:localai-paged`, then the LocalAI patch series is regenerated.
- Do not treat a standalone PoC as a result. Every performance claim requires an
in-backend A/B.
- Do not ship lossy paths default-on. Non-byte-identical paths require KL gates.
## Required Preflight
Before any DGX build, benchmark, or profile:
1. `docker ps` must show no running containers, especially no `local-ai-worker`.
2. `nvidia-smi --query-compute-apps=pid` must show zero compute apps.
3. `~/gpu_bench_lock/owner` must be absent or `FREE*`.
4. Record hostname, git SHA, dirty status, build arch, binary mtimes, model paths,
benchmark command, and environment variables.
Use `~/_git/llama.cpp` as the local source of truth. DGX source trees are allowed
for builds and artifact inspection, but dirty DGX checkouts must not be treated as
canonical source.
## Evidence From Subagent Audit
Four read-only subagents audited the current state:
- llama.cpp / ggml source audit.
- vLLM source and installed package audit.
- LocalAI patch and docs audit.
- DGX artifact and profile audit.
Their shared conclusion: the final docs are a useful snapshot, but several
claims are broader than the available evidence.
Key findings:
- The strongest unresolved implementation target is W4A16 grouped MoE prefill.
vLLM uses Marlin W4A16 on GB10, and llama.cpp already has a correct but untuned
scaffold in `ggml/src/ggml-cuda/w4a16-gemm.cu`.
- The existing W4A16 rejection is a first-implementation failure, not a proof of
impossibility. The patch header names fixable costs: f32 to bf16 cast pre-pass,
host tile-map setup, small copies, scalar dequant, and ragged tile waste.
- The 924 t/s paged GPU-steady decode figure is artifact-backed, but the vLLM
1078 t/s true GPU-steady figure was not found as a self-contained
ntg16/ntg64 difference-method artifact. Reproduce before relying on the 86%
claim.
- GDN M5 is real, but M5/M8 provenance is muddy because CDEF records a dirty
dev-tree M8 commit while docs describe production M5 defaults.
- S3 fixed-period scheduling and fixed-slot padding were rejected, but adaptive
scheduling remains unproven.
## Candidate Workstreams
### A. Provenance And Baseline Reproduction
Purpose: make later claims defensible.
Tasks:
- Build from clean `~/_git/llama.cpp` `localai-paged` source, or a clean DGX clone
generated from that source.
- Re-run canonical md5 gates for paged MoE and dense:
- paged MoE: `8cb0ce23777bf55f92f63d0292c756b0`
- dense: `5951a5b4d624ce891e22ab5fca9bc439`
- Re-run a short prefill baseline for MoE and dense at `npp=512,2048`.
- Re-run graph-node-traced decode for paged and vLLM using the same
difference-method shape: `ntg=16` and `ntg=64` at N=128 or N=256.
Gate:
- No implementation work starts until the baseline artifact names, source SHAs,
and commands are recorded.
### B. W4A16 Grouped MoE Prefill Attack
Purpose: port the vLLM Marlin W4A16 advantage into ggml's in-backend MoE prefill
path.
Current hooks:
- `ggml/src/ggml-cuda/w4a16-gemm.cu`
- `ggml/src/ggml-cuda/w4a16-gemm.cuh`
- `ggml/src/ggml-cuda/ggml-cuda.cu` around `ggml_cuda_mul_mat_id`
- `ggml/src/ggml-cuda/mmq.cu` around `LLAMA_W4A16_PREFILL_M`
Known current costs:
- Separate f32 to bf16 activation cast pass.
- Host-built tile metadata and H2D copies.
- Scalar in-register FP4 to bf16 dequant.
- 4-byte weight staging.
- Ragged expert tile waste.
- Interaction with the generic token-sorting fallback.
Phased experiments:
1. Reconfirm current 0035 W4A16 performance with clean provenance.
2. Remove or fuse the f32 to bf16 activation cast pre-pass.
3. Move tile metadata generation device-side or cache it across repeated shapes.
4. Improve weight staging width and shared-memory layout.
5. Tune tile shapes for ragged per-expert M distribution.
6. Compare against FP4-MMQ and vLLM Marlin buckets with nsys.
Correctness gate:
- `test-backend-ops MUL_MAT_ID` forced W4A16.
- Greedy md5 for unaffected default-off path.
- KL gate for engaged W4A16 path: `KLD(W4A16||f16) <= KLD(FP4-MMQ||f16)`.
- Decode path unchanged when `LLAMA_W4A16_PREFILL_M=0`.
Benchmark gate:
- Beat default FP4-MMQ on MoE `S_PP` at `npp=512` and `npp=2048`.
- No material peak-memory increase.
- No decode regression in the default path.
### C. Native Ragged Grouped FP4-MMA Prefill
Purpose: test whether patch 0034's native FP4-MMA PoC failed due to integration,
not due to the core kernel idea.
Current hooks:
- `ggml/src/ggml-cuda/fp4-gemm.cu`
- `ggml/src/ggml-cuda/fp4-gemm.cuh`
- `LLAMA_FP4_PREFILL_M`
Experiment:
- Build a graph-safe ragged grouped FP4-MMA MoE prefill kernel that avoids the
per-expert host-sync loop.
Correctness gate:
- Same KL and op gates as W4A16.
- Explicit proof that the per-expert host fallback is not on the hot path.
Benchmark gate:
- Beat current FP4-MMQ or lose decisively enough to close this branch.
### D. GDN Chunked Scan Follow-up
Purpose: compare vLLM's in-tree FLA-derived GDN path against the current M5
implementation without relying on muddy dev-tree artifacts.
Current hooks:
- `ggml/src/ggml-cuda/gated_delta_net.cu`
- `GDN_TC`
- `GDN_CHUNK_MIN`
- existing M5 tensor-core ladder
Phased experiments:
1. Clean A/B: current production M5 against sequential and against recorded M8
dev-tree behavior.
2. C=32 and C=64 variants.
3. dv slab variants.
4. cp.async staging variants.
5. Register-state variant only if the lower-risk variants show headroom.
Correctness gate:
- `test-backend-ops GATED_DELTA_NET`, including multi-chunk, tail-chunk,
multi-seq, and adversarial decay cases.
- Greedy md5 per path.
- KL gate for any non-byte-identical path.
Benchmark gate:
- Beat current M5, not just old sequential.
- Preserve decode behavior by keeping `GDN_CHUNK_MIN > 1`.
### E. MoE Weighted Fan-in Fusion
Purpose: remove generic graph-level MoE reduction overhead that vLLM avoids or
amortizes through fused MoE handling.
Current source:
- `src/llama-graph.cpp`, MoE down projection and expert reduction.
- `ggml/src/ggml-cuda/ggml-cuda.cu`, CUDA fusion and MoE support.
Experiment:
- Add a CUDA-specific fused path for `down_experts * weights -> sum expert_used`
while preserving the current reduction order where required.
Correctness gate:
- Bit-exact for supported shapes, or KL-benign if reduction order changes.
- Handles all `n_expert_used` used by Qwen3.6 MoE.
Benchmark gate:
- Move MoE prefill or decode wall time by more than noise. If it is only a
2-3% dispatch bucket, record and deprioritize.
### F. Adaptive Serving Scheduler
Purpose: keep S3's decode-window benefit without reproducing its TTFT collapse.
Current hooks:
- `tools/server/server-context.cpp`
- `LLAMA_PAGED_DECODE_STABLE`
- `LLAMA_PAGED_PREFILL_PERIOD`
- existing dynamic prefill budget patches.
Experiment:
- Replace fixed-period prefill deferral with adaptive admission based on live
decode width, waiting prefill backlog, and TTFT budget.
Correctness gate:
- Serving output correctness unchanged.
- No starvation of prefill requests.
Benchmark gate:
- Improve aggregate throughput or decode throughput at N=128 or N=256 without
the 2.5x TTFT regression from fixed S3.
### G. Projection And GDN Glue Fusion
Purpose: steal vLLM's `prepare_gdn_attention_core_inputs` idea where ggml still
pays small copy, cat, slice, or unpack kernels.
Current source:
- `src/models/qwen35.cpp`
- `src/models/qwen35moe.cpp`
- `ggml/src/ggml-cuda/ggml-cuda.cu`
Experiment:
- Fuse q/k/v/z unpacking, BA projection preparation, RMSNorm-gated output prep,
and FP4/FP8 quant prep where the graph pattern is stable.
Correctness gate:
- Per-op tests for new fusion.
- Greedy md5 for model paths.
Benchmark gate:
- Only continue if nsys shows this bucket is material after MoE and GDN work.
## Subagent Plan
Use subagents for independent read, implementation, and review slices. Do not use
subagents to edit the same files in parallel.
Recommended roles by phase:
- Phase 0 source/provenance agent: owns command capture and source SHA checks.
- Phase 0 artifact agent: owns parsing existing and new benchmark artifacts.
- W4A16 kernel agent: owns `w4a16-gemm.*`.
- W4A16 integration agent: owns `ggml-cuda.cu` and `mmq.cu` dispatch plumbing.
- GDN kernel agent: owns `gated_delta_net.cu`.
- Scheduler agent: owns server scheduling files only.
- Reviewer agent: reviews gates, provenance, and whether measured claims match
artifacts.
Subagent output requirements:
- File paths and functions inspected or changed.
- Exact commands run.
- Exact artifacts produced.
- Pass/fail result against the phase gate.
- Any uncertainty labeled explicitly.
## Phase Order
### Phase 0 - Reproduce And Correct The Record
Do first.
Deliverables:
- Clean source/build provenance.
- Short prefill baseline.
- Graph-node-traced decode difference-method for paged and vLLM.
- Updated docs if the 86% decode claim or CDEF provenance changes.
Exit criteria:
- Baseline is trustworthy enough to judge optimization deltas.
### Phase 1 - W4A16 MoE Prefill
Do second.
Deliverables:
- Reconfirmed current W4A16 baseline.
- At least one targeted W4A16 overhead removal.
- A/B against default FP4-MMQ.
Exit criteria:
- Either W4A16 beats FP4-MMQ and continues, or it is rejected with direct
artifact-backed evidence.
### Phase 2 - GDN Follow-up
Do after Phase 1 unless Phase 0 proves decode/GDN is the larger immediate gap.
Deliverables:
- Clean M5 vs candidate geometry A/B.
- Correctness gates for all candidate variants.
Exit criteria:
- Keep the best variant or close the branch with measured evidence.
### Phase 3 - MoE Fan-in And Glue Fusions
Do after kernel work identifies remaining non-kernel buckets.
Deliverables:
- nsys-backed bucket selection.
- Fusion implementation only for material buckets.
Exit criteria:
- Keep only fusions that move end-to-end numbers beyond noise.
### Phase 4 - Adaptive Serving
Do after compute kernels are stable.
Deliverables:
- Adaptive scheduling policy.
- Serving A/B at N=8,32,128,256.
Exit criteria:
- Improve serving without TTFT collapse.
## Decision Rules
- Prefer measured in-backend results over source plausibility.
- Prefer small kill-gate experiments over multi-week rewrites.
- Continue a branch only if it beats the current shipped path, not an obsolete
baseline.
- Document rejected branches with artifact paths so they are not rerun.
- Keep the fork branch canonical and regenerate LocalAI patches from it.

View File

@@ -1,172 +0,0 @@
# GDN Shared-A/Ai Cost Model
Phase 12 decides whether the next GDN prefill attempt should implement a
shared-A/Ai global-scratch prototype or stop GDN kernel work on GB10.
## Reference Points
llama.cpp:
- `/home/mudler/_git/llama.cpp/ggml/src/ggml-cuda/gated_delta_net.cu`
- `gated_delta_net_chunked_cuda`
- `launch_gdn_chunked`
- `launch_gated_delta_net`
- `ggml_cuda_op_gated_delta_net`
vLLM/FLA:
- `/home/mudler/_git/vllm/vllm/model_executor/layers/fla/ops/chunk.py`
- `chunk_gated_delta_rule_fwd`
- `/home/mudler/_git/vllm/vllm/model_executor/layers/fla/ops/solve_tril.py`
- `solve_tril`
- `solve_tril_16x16_kernel`
- `merge_16x16_to_32x32_inverse_kernel`
- `merge_16x16_to_64x64_inverse_kernel`
- `/home/mudler/_git/vllm/vllm/model_executor/layers/fla/ops/wy_fast.py`
- `recompute_w_u_fwd`
## Metadata
DGX metadata artifact:
- `/home/mudler/bench/phase12_gdn_shared_ai_cost_model/model_metadata.txt`
GGUF metadata:
| Model | Arch | Blocks | Full-attn interval | GDN layers | SSM inner | SSM state | GDN heads |
|-------|------|--------|--------------------|------------|-----------|-----------|-----------|
| MoE | `qwen35moe` | 41 | 4 | 30 inferred | 4096 | 128 | 32 inferred |
| Dense | `qwen35` | 64 | 4 | 48 inferred | 6144 | 128 | 48 inferred |
Notes:
- `GDN heads = ssm.inner_size / ssm.state_size`.
- MoE has one `nextn` layer; the serving/prefill stack uses the 40 normal
layers, with 30 GDN layers at interval 4.
- Dense has 64 layers, 48 GDN layers at interval 4.
## Dynamic Shared Memory
Formula:
```text
C16 full-width current M5:
floats = S_v*S_v + 2*C*S_v + S_v*C + C*C + 3*C + 2*C*C
C32 full-width:
floats = S_v*S_v + 2*C*S_v + S_v*C + C*C + 3*C + 2*C*C
C32 slab64 with U staging:
floats = S_v*64 + 2*C*S_v + 64*C + C*C + 3*C + 2*C*C + 64*C
```
For `S_v=128`:
| Shape | Bytes | KiB | Fits GB10 dynamic smem? |
|-------|-------|-----|-------------------------|
| C16 full-width | 93,376 | 91.19 | yes |
| C32 full-width | 127,360 | 124.38 | no |
| C32 slab64 + U staging | 94,592 | 92.38 | yes |
Implication:
- C32 full-width cannot be a single current-style CTA on GB10.
- C32 only fits by splitting value columns or by changing state residency.
- Splitting value columns must share A/Ai or it repeats the Phase 10 failure.
## Ai Scratch Size
Formula:
```text
Ai scratch bytes = npl * H * ceil(npp / BT) * BT * BT * sizeof(dtype)
```
Benchmark shape: `npl=32`, `S_v=128`.
| Model | H | npp | BT | Ai dtype | Chunks | Ai scratch MiB | 3x Ai traffic MiB |
|-------|---|-----|----|----------|--------|----------------|-------------------|
| MoE | 32 | 512 | 32 | f32 | 16 | 64.0 | 192.0 |
| MoE | 32 | 512 | 32 | f16 | 16 | 32.0 | 96.0 |
| MoE | 32 | 512 | 64 | f32 | 8 | 128.0 | 384.0 |
| MoE | 32 | 512 | 64 | f16 | 8 | 64.0 | 192.0 |
| MoE | 32 | 2048 | 32 | f32 | 64 | 256.0 | 768.0 |
| MoE | 32 | 2048 | 32 | f16 | 64 | 128.0 | 384.0 |
| MoE | 32 | 2048 | 64 | f32 | 32 | 512.0 | 1536.0 |
| MoE | 32 | 2048 | 64 | f16 | 32 | 256.0 | 768.0 |
| Dense | 48 | 512 | 32 | f32 | 16 | 96.0 | 288.0 |
| Dense | 48 | 512 | 32 | f16 | 16 | 48.0 | 144.0 |
| Dense | 48 | 512 | 64 | f32 | 8 | 192.0 | 576.0 |
| Dense | 48 | 512 | 64 | f16 | 8 | 96.0 | 288.0 |
| Dense | 48 | 2048 | 32 | f32 | 64 | 384.0 | 1152.0 |
| Dense | 48 | 2048 | 32 | f16 | 64 | 192.0 | 576.0 |
| Dense | 48 | 2048 | 64 | f32 | 32 | 768.0 | 2304.0 |
| Dense | 48 | 2048 | 64 | f16 | 32 | 384.0 | 1152.0 |
`3x Ai traffic` means one Ai write plus two Ai reads for two value slabs.
## Interpretation
The f32 `BT=32` scratch path is large but plausible:
- Peak scratch is 256 MiB for MoE and 384 MiB for dense at `npp=2048,npl=32`.
- Ai traffic is 768 MiB for MoE and 1.125 GiB for dense per GDN layer call.
- This is not free on LPDDR5x, but it is not automatically worse than
recomputing A/Ai in every value slab.
The f16/BF16 Ai path halves traffic but should not be first because Phase 10 and
Phase 11 showed correctness must be established before performance. The first
prototype should store Ai in f32, stay default-off, and use md5/KL gates before
trying a lossy Ai dtype.
## Decision
GO: Phase 13 should implement a default-off global-Ai scratch prototype.
Rationale:
- The only remaining C32 path that addresses Phase 10's failure is sharing A/Ai
across value slabs.
- `BT=32` f32 scratch has acceptable peak memory for the existing GB10
benchmark shapes.
- The implementation can be default-off and rejected cleanly if global scratch
traffic or extra launch boundaries dominate.
Phase 13 constraints:
- Prototype only `BT=32`, f32 Ai, two `dv_tile=64` value slabs.
- Keep decode out via `GDN_CHUNK_MIN > 1`.
- Gate with `GATED_DELTA_NET`, canonical MoE/dense md5, and same-session A/B.
- If md5 changes, run KL before benchmarking.
- If the prototype is flat or slower, reject it and stop GDN kernel work on
GB10; do not iterate into f16 Ai until f32 proves the schedule can win.
## Phase 13 Result
Phase 13 implemented the f32 Global-Ai32 prototype and rejected it.
Correctness:
- MoE md5: `8cb0ce23777bf55f92f63d0292c756b0`.
- Dense md5: `5951a5b4d624ce891e22ab5fca9bc439`.
Performance:
| Model | Mode | PP | S_PP t/s |
|-------|------|----|----------|
| MoE | M5 base | 2048 | 2425.10 |
| MoE | Global Ai32 | 2048 | 2097.76 |
| Dense | M5 base | 2048 | 1016.14 |
| Dense | Global Ai32 | 2048 | 918.19 |
Artifacts:
- `/home/mudler/bench/phase13_gdn_global_ai32/gates/`
- `/home/mudler/bench/phase13_gdn_global_ai32/ab/`
- `/home/mudler/bench/phase13_gdn_global_ai32/rejected/global_ai32_rejected.diff`
Final decision:
- Reject Global-Ai32.
- Stop GDN kernel work on GB10. The remaining vLLM GDN advantage is not
reachable through the low-conflict C16/C32 patch shapes tested here.

View File

@@ -1,514 +0,0 @@
# Plan: ship the paged llama.cpp as its OWN backend + NVFP4 Qwen3.6 gallery items
Scoping deliverable only. NOTHING is changed by this document. It is grounded in the
actual repo structure (read 2026-06-26 in worktree feat+paged-attention), not assumptions.
SHIPPED REALITY (update 2026-06-27): the backend ships CUDA-only. The matrix rows and
the index.yaml meta-backend keep ONLY the CUDA/cublas variants (cuda-12, cuda-13, and
the nvidia-l4t arm64 cuda-12/cuda-13 Jetson rows). The cpu / vulkan / sycl / hipblas /
metal-darwin variants discussed below as optional/phase-2 were NOT shipped (and the
darwin row was removed): off-CUDA the patchset's wins gate off, so it is neutral-to-
negative there and non-CUDA users should use the stock llama-cpp backend (README 4c).
================================================================================
0. GROUND TRUTH (what the repo actually does today)
================================================================================
The paged patchset is ALREADY integrated into the stock llama-cpp backend in this
worktree. Two mechanisms, both already present:
(a) BUILD: backend/cpp/llama-cpp/Makefile has `LLAMA_PAGED?=on`. The `llama.cpp:`
target git-applies patches/0*.patch (base series) then, when LLAMA_PAGED != off,
patches/paged/0*.patch (the 0018-0023 paged series + the earlier 0001-0017).
prepare.sh has a fallback `patch`-based apply guarded by a sentinel
(llama.cpp/src/paged-kv-manager.cpp). So a stock `make backends/llama-cpp` TODAY
already ships the paged engine compiled in.
(b) RUNTIME GATING: backend/cpp/llama-cpp/grpc-server.cpp ALREADY carries the option
hooks (lines ~752-842). They only call setenv() before context init:
- option `kv_paged` / `paged_kv` / `paged_attention` -> setenv LLAMA_KV_PAGED=1
- option `kv_paged_debug` / `paged_kv_debug` -> setenv LLAMA_KV_PAGED_DEBUG=1
- option `max_prefill_tokens` / `mpt` / `prefill_budget` -> setenv LLAMA_PREFILL_BUDGET
- option `max_batch_tokens` / `mbt` -> setenv LLAMA_MAX_BATCH_TOKENS
- option `prefill_cap` -> setenv LLAMA_PREFILL_CAP
Against UNPATCHED llama.cpp these setenv() calls are inert (nothing reads the env),
so grpc-server.cpp is byte-safe to share between a clean build and a paged build.
The paged engine itself lives entirely inside the patched llama.cpp lib
(paged-kv-manager.cpp etc.), NOT in grpc-server.cpp.
Conclusion: "stock llama-cpp + paged patchset, runtime-gated" is the CURRENT state of
ONE backend. The task is to SPLIT that into two backends:
- llama-cpp = clean upstream llama.cpp (de-risked: a dep-bump can never break on a
paged hook), grpc-server.cpp keeps the dormant hooks.
- <newname> = stock grpc-server.cpp + paged patch series applied + paged on.
The turboquant backend is the EXACT precedent for "a llama.cpp variant that reuses the
backend/cpp/llama-cpp grpc-server sources via a thin wrapper Makefile + its own Dockerfile
+ its own matrix rows". Copy turboquant's shape, with two simplifications (see section 1).
CPU_ALL_VARIANTS reuse: backend/cpp/llama-cpp/Makefile already has `llama-cpp-cpu-all`
(one grpc-server + dlopen libggml-cpu-*.so via -DGGML_BACKEND_DL/-DGGML_CPU_ALL_VARIANTS,
SHARED_LIBS=ON make-var). turboquant mirrors it with `turboquant-cpu-all`. The new backend
gets the same single-build CPU target for free by reusing the same Makefile machinery.
--------------------------------------------------------------------------------
RECOMMENDED BACKEND NAME: `llama-cpp-paged` (see section 4 for the full rationale)
--------------------------------------------------------------------------------
Everywhere below, NAME = llama-cpp-paged, DOCKERFILE = Dockerfile.llama-cpp-paged,
SRC DIR = backend/cpp/llama-cpp-paged/, MAKE VAR = BACKEND_LLAMA_CPP_PAGED.
DO NOT use the dotted working name `localai-llama.cpp`: a dot in Dockerfile.<suffix> and
in the tag-suffix is unprecedented (every sibling is hyphenated: llama-cpp, ik-llama-cpp,
turboquant, ds4) and complicates the changed-backends.js endsWith() suffix matching.
================================================================================
1. NEW BACKEND - file by file
================================================================================
--------------------------------------------------------------------------------
1.1 backend/cpp/llama-cpp/Makefile (the ONE necessary touch to stock)
--------------------------------------------------------------------------------
Change exactly one default so the STOCK image ships clean against upstream:
-LLAMA_PAGED?=on
+LLAMA_PAGED?=off
Why: this is the entire point of the split - stock llama-cpp must build clean so an
upstream LLAMA_VERSION bump can never fail on a paged hook. The runtime hooks in
grpc-server.cpp stay (inert). The new backend forces LLAMA_PAGED=on explicitly (1.2), so
it does not depend on this default. NOTE this DOES change stock's shipped artifact (it
currently ships paged-compiled-in-but-gated); that is intended de-risking, call it out in
the PR. If the team prefers stock literally untouched, the alternative is to leave
`?=on` and accept that stock keeps carrying the patch series - but then "clean stock" is
not achieved. Recommendation: flip to off.
(No other change to backend/cpp/llama-cpp/ - grpc-server.cpp, CMakeLists.txt, prepare.sh,
patches/, patches/paged/ are all reused as-is by the new backend.)
--------------------------------------------------------------------------------
1.2 backend/cpp/llama-cpp-paged/Makefile (NEW - thin wrapper, model on turboquant)
--------------------------------------------------------------------------------
Mirror backend/cpp/turboquant/Makefile, but SIMPLER (two things turboquant needs that we
do NOT):
- turboquant overrides LLAMA_REPO/LLAMA_VERSION to a fork. We use the SAME upstream pin
as stock (it lives in backend/cpp/llama-cpp/Makefile, already auto-bumped). So we do
NOT set LLAMA_VERSION here -> no bump_deps.yaml entry needed (big simplification vs
turboquant). We only force LLAMA_PAGED=on.
- turboquant runs patch-grpc-server.sh (augments the KV-cache type allow-list) and
apply-patches.sh (fork catch-up). We need NEITHER: grpc-server.cpp already has the
paged hooks, and the paged patch series is applied by the copied llama-cpp Makefile's
own `llama.cpp:` target when LLAMA_PAGED=on.
Shape (one flavor shown; replicate the turboquant flavor set: avx/avx2/avx512/fallback/
cpu-all/grpc/rpc-server):
LLAMA_CPP_DIR := $(CURRENT_MAKEFILE_DIR)/../llama-cpp
define paged-build # $(1)=flavor $(2)=cmake flags $(3)=target
rm -rf $(CURRENT_MAKEFILE_DIR)/../llama-cpp-paged-$(1)-build
cp -rf $(LLAMA_CPP_DIR) $(CURRENT_MAKEFILE_DIR)/../llama-cpp-paged-$(1)-build
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-paged-$(1)-build purge
# clone upstream + apply base AND paged patch series (LLAMA_PAGED=on forces it)
LLAMA_PAGED=on $(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-paged-$(1)-build llama.cpp
CMAKE_ARGS="$(CMAKE_ARGS) $(2)" TARGET="$(3)" LLAMA_PAGED=on \
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-paged-$(1)-build grpc-server
cp -rfv $(CURRENT_MAKEFILE_DIR)/../llama-cpp-paged-$(1)-build/grpc-server llama-cpp-paged-$(1)
endef
llama-cpp-paged-cpu-all:
# identical to turboquant-cpu-all: SHARED_LIBS=ON + GGML_BACKEND_DL + CPU_ALL_VARIANTS
# + --target ggml; then collect ggml-shared-libs/ for package.sh to bundle.
... LLAMA_PAGED=on SHARED_LIBS=ON \
EXTRA_CMAKE_ARGS="-DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON" \
TARGET="--target grpc-server --target ggml" ...
package: ; bash package.sh
purge: ; rm -rf $(CURRENT_MAKEFILE_DIR)/../llama-cpp-paged-*-build; rm -rf llama-cpp-paged-* package
clean: purge
Binaries are named llama-cpp-paged-{cpu-all,fallback,grpc,rpc-server,...} so run.sh and
package.sh glob them.
--------------------------------------------------------------------------------
1.3 backend/cpp/llama-cpp-paged/run.sh (NEW - copy turboquant/run.sh, rename binaries)
--------------------------------------------------------------------------------
s/turboquant/llama-cpp-paged/g. Prefers llama-cpp-paged-cpu-all if present, falls back to
llama-cpp-paged-fallback; llama-cpp-paged-grpc when LLAMACPP_GRPC_SERVERS set; Darwin
DYLD_LIBRARY_PATH branch; lib/ld.so launch. Keep verbatim otherwise.
--------------------------------------------------------------------------------
1.4 backend/cpp/llama-cpp-paged/package.sh (NEW - copy turboquant/package.sh, rename)
--------------------------------------------------------------------------------
s/turboquant/llama-cpp-paged/g. Copies llama-cpp-paged-* into package/, bundles
ggml-shared-libs/*.so* into package/lib (the CPU_ALL_VARIANTS dlopen set), copies run.sh,
and the per-arch libc/ld.so set (unchanged).
--------------------------------------------------------------------------------
1.5 backend/Dockerfile.llama-cpp-paged (NEW - copy Dockerfile.turboquant, swap paths)
--------------------------------------------------------------------------------
Identical 3-stage structure (builder-fromsource / builder-prebuilt / FROM scratch). Edits:
- bind/run .docker/llama-cpp-paged-compile.sh (new, 1.6) instead of turboquant-compile.sh
- ccache id: id=llama-cpp-paged-ccache-${TARGETARCH}-${BUILD_TYPE}
(OPTIONAL OPTIMIZATION: set id=llama-cpp-ccache-${TARGETARCH}-${BUILD_TYPE} to SHARE
stock llama-cpp's ccache - the paged TUs are mostly byte-identical to stock, so a warm
stock cache would give the paged build near-free object reuse. Trade-off: a regression
in one could surface as a cold miss in the other. Recommend sharing; revisit if noisy.)
- both `make -BC /LocalAI/backend/cpp/llama-cpp-paged package`
- final COPY --from=builder /LocalAI/backend/cpp/llama-cpp-paged/package/. ./
--------------------------------------------------------------------------------
1.6 .docker/llama-cpp-paged-compile.sh (NEW - copy llama-cpp-compile.sh, swap make targets)
--------------------------------------------------------------------------------
Identical to .docker/llama-cpp-compile.sh except `cd .../llama-cpp-paged` and call
`make llama-cpp-paged-cpu-all` (BUILD_TYPE empty / CPU) or `make llama-cpp-paged-fallback`
(GPU), then `make llama-cpp-paged-grpc` + `make llama-cpp-paged-rpc-server`. Keep the
arm64 gcc-14 apt step (CPU_ALL_VARIANTS armv9.2 SME needs gcc-14). ccache export unchanged.
--------------------------------------------------------------------------------
1.7 Makefile (top-level) - 6 edits, mirror the turboquant lines
--------------------------------------------------------------------------------
a) .NOTPARALLEL (line 2): append `backends/llama-cpp-paged`
b) Backend def (after BACKEND_TURBOQUANT, line ~1172):
# llama-cpp-paged = stock llama.cpp grpc-server + LocalAI paged-attention patch
# series (LLAMA_PAGED=on). Reuses backend/cpp/llama-cpp sources via a thin wrapper.
BACKEND_LLAMA_CPP_PAGED = llama-cpp-paged|llama-cpp-paged|.|false|false
(lang field `llama-cpp-paged` -> Dockerfile.llama-cpp-paged, matching the
llama-cpp / ik-llama-cpp / turboquant convention where lang==backend name.)
c) generate-docker-build-target eval (after BACKEND_TURBOQUANT, line ~1273):
$(eval $(call generate-docker-build-target,$(BACKEND_LLAMA_CPP_PAGED)))
d) docker-build-backends (line ~1337): append docker-build-llama-cpp-paged
e) test-extra-backend-llama-cpp-paged target (mirror test-extra-backend-turboquant,
line ~673): BACKEND_IMAGE=local-ai-backend:llama-cpp-paged $(MAKE) test-extra-backend
f) (optional) backends/llama-cpp-paged-darwin target if shipping metal (mirror
backends/llama-cpp-darwin at line 1124; see 1.11).
--------------------------------------------------------------------------------
1.8 .github/backend-matrix.yml - add rows (mirror every llama-cpp row, swap names)
--------------------------------------------------------------------------------
For EACH variant you choose to ship (see phased recommendation in section 4), add a row
copied from the corresponding llama-cpp row with:
- backend: "llama-cpp-paged"
- dockerfile: "./backend/Dockerfile.llama-cpp-paged"
- tag-suffix: swap `-llama-cpp` -> `-llama-cpp-paged`
(e.g. -cpu-llama-cpp -> -cpu-llama-cpp-paged;
-gpu-nvidia-cuda-12-llama-cpp -> -gpu-nvidia-cuda-12-llama-cpp-paged; etc.)
- builder-base-image: UNCHANGED - reuse the same base-grpc-* tags as llama-cpp
(this backend compiles the same gRPC + same toolchain; no new base-images.yml variant
is needed, so NO base-images bootstrap step). This is the cheap-variant payoff.
- CPU: TWO per-arch rows (amd64 ubuntu-latest + arm64 ubuntu-24.04-arm) sharing
tag-suffix '-cpu-llama-cpp-paged' so changed-backends.js emits a merge-matrix entry and
backend-merge-jobs assembles the manifest list. Same per-arch native + manifest-merge
pattern as -cpu-llama-cpp.
- Darwin (if shipping): add to includeDarwin:
- backend: "llama-cpp-paged"
tag-suffix: "-metal-darwin-arm64-llama-cpp-paged"
lang: "go"
(omit build-type, exactly like the llama-cpp darwin row at line 4908.)
REMINDER: the CI path filter only builds a backend on a PR when a file under its dir
changes. The PR that adds this backend touches backend/cpp/llama-cpp-paged/* so it self-
triggers. But also add the cross-trigger in 1.9 so future edits to backend/cpp/llama-cpp/
(the shared source) retrigger this backend too.
--------------------------------------------------------------------------------
1.9 scripts/changed-backends.js - two edits (mirror turboquant exactly)
--------------------------------------------------------------------------------
a) inferBackendPath(): add BEFORE the generic `endsWith("llama-cpp")` branch (line 56),
next to the turboquant branch (line 45):
if (item.dockerfile.endsWith("llama-cpp-paged")) {
// reuses backend/cpp/llama-cpp sources via a thin wrapper Makefile
return `backend/cpp/llama-cpp-paged/`;
}
ORDER MATTERS: "Dockerfile.llama-cpp-paged".endsWith("llama-cpp") is false today, but
keep the specific branch first regardless (defensive, and returns the right path).
b) inferBackendPathDarwin(): add a case (next to the llama-cpp one at line 66):
if (item.backend === "llama-cpp-paged") { return `backend/cpp/llama-cpp-paged/`; }
c) Per-backend cross-trigger (line 274-278, mirror the turboquant block):
if (backend === "llama-cpp-paged" && !changed) {
changed = changedFiles.some(file => file.startsWith("backend/cpp/llama-cpp/"));
}
Verify: node -e "... e.dockerfile.endsWith('llama-cpp-paged') ..." per adding-backends.md.
--------------------------------------------------------------------------------
1.10 backend/index.yaml - meta + image entries (META-BACKEND - capabilities map, NO uri)
--------------------------------------------------------------------------------
GOTCHA (project_backend_meta_gotcha): a backend that ships per-platform images MUST be a
meta backend = an anchor with a `capabilities:` map and NO top-level `uri:`; the concrete
per-platform entries carry the uri. Copy the *llamacpp anchor (lines 3-31).
Step a - meta anchor in `## metas` (after *turboquant, ~line 74):
- &llamacpppaged
name: "llama-cpp-paged"
alias: "llama-cpp-paged"
license: mit
icon: <same as llama-cpp>
description: |
LocalAI's paged-attention llama.cpp: on-demand paged KV cache + decode-first
prefill budget. Stock llama.cpp grpc-server + the LocalAI paged patch series.
Tuned for NVFP4 dense/MoE on Blackwell/GB10. Reuses the llama-cpp gRPC server.
urls: [ https://github.com/ggerganov/llama.cpp ]
tags: [ text-to-text, LLM, CPU, GPU, CUDA, Metal, paged-attention, nvfp4 ]
capabilities:
default: "cpu-llama-cpp-paged"
nvidia: "cuda12-llama-cpp-paged"
nvidia-cuda-12: "cuda12-llama-cpp-paged"
nvidia-cuda-13: "cuda13-llama-cpp-paged"
nvidia-l4t: "nvidia-l4t-arm64-llama-cpp-paged"
nvidia-l4t-cuda-12: "nvidia-l4t-arm64-llama-cpp-paged"
nvidia-l4t-cuda-13: "cuda13-nvidia-l4t-arm64-llama-cpp-paged"
metal: "metal-llama-cpp-paged"
# add amd/intel/vulkan keys ONLY for variants you actually build (section 4)
Step b - a `-development` meta (mirror llama-cpp-development, line 1611) with the same
capabilities map pointing at the `*-development` image names.
Step c - concrete image entries at end of file (mirror the llama-cpp block lines
2106-2200), one latest + one development per variant, each as:
- !!merge <<: *llamacpppaged
name: "cpu-llama-cpp-paged"
uri: "quay.io/go-skynet/local-ai-backends:latest-cpu-llama-cpp-paged"
mirrors: [ localai/localai-backends:latest-cpu-llama-cpp-paged ]
- !!merge <<: *llamacpppaged
name: "cpu-llama-cpp-paged-development"
uri: "quay.io/go-skynet/local-ai-backends:master-cpu-llama-cpp-paged"
mirrors: [ localai/localai-backends:master-cpu-llama-cpp-paged ]
...repeat for cuda12 / cuda13 / l4t / metal etc.
The `latest-` / `master-` uri prefix + tag-suffix MUST match the matrix tag-suffix exactly.
--------------------------------------------------------------------------------
1.11 Darwin (only if shipping metal; the NVFP4 target is CUDA, so metal is optional/phase 2)
--------------------------------------------------------------------------------
If metal is shipped, also:
- scripts/build/llama-cpp-paged-darwin.sh (copy scripts/build/llama-cpp-darwin.sh; it
drives the 3 CMake variants + otool dylib bundling). Ensure it forces LLAMA_PAGED=on.
- Makefile `backends/llama-cpp-paged-darwin` target (mirror backends/llama-cpp-darwin).
- backend_build_darwin.yml: add the llama-cpp-paged branch (mirror the llama-cpp-specific
step that calls `make backends/llama-cpp-darwin`).
- index.yaml metal-llama-cpp-paged / -development image entries (already in 1.10).
- C++ proto gotcha already handled (reuses llama-cpp CMakeLists.txt with hw_grpc_proto
linking protobuf/grpc++), so no Homebrew-include failure.
--------------------------------------------------------------------------------
1.12 Importer / /backends/known dropdown (drop-in, NOT a new importer)
--------------------------------------------------------------------------------
This backend consumes GGUF exactly like llama-cpp -> extend the EXISTING importer, do not
add a new one (per adding-backends.md rule 2). Edit core/gallery/importers/llama-cpp.go:
- AdditionalBackends() (line 37): append
{Name: "llama-cpp-paged", Modality: "text",
Description: "Paged-attention llama.cpp (on-demand paged KV + decode-first budget)"}
- Import() backend allow-list (line 133): add "llama-cpp-paged" to the switch case so a
preferences.backend == "llama-cpp-paged" is honored:
case "ik-llama-cpp", "turboquant", "llama-cpp-paged": backend = b
- core/gallery/importers/importers_test.go: add a table case asserting the preference
override emits backend: llama-cpp-paged (Ginkgo/Gomega; reuse an existing public GGUF
HF fixture). Run `go test ./core/gallery/importers/...`.
--------------------------------------------------------------------------------
1.13 Docs
--------------------------------------------------------------------------------
- docs/content/features/backends.md: add llama-cpp-paged to the text-to-text/LLM list,
one line noting paged KV + NVFP4 Blackwell tuning. (Not an in-house from-scratch engine
-> it is a llama.cpp variant -> do NOT add to the README maintained-engines table.)
--------------------------------------------------------------------------------
1.14 Does grpc-server.cpp need the paged hooks? YES - already present, reused unchanged.
--------------------------------------------------------------------------------
The hooks (kv_paged / max_batch_tokens / prefill_budget / prefill_cap) are already in the
SHARED backend/cpp/llama-cpp/grpc-server.cpp. The paged backend reuses that file verbatim
(via the Makefile copy). No patch-grpc-server.sh step is needed (unlike turboquant). The
hooks are what translate the gallery `options:` (1.10 section 2) into the LLAMA_KV_PAGED /
LLAMA_MAX_BATCH_TOKENS env that the paged llama.cpp lib reads.
================================================================================
2. GALLERY ITEMS - NVFP4 Qwen3.6 dense + MoE
================================================================================
Add two entries to gallery/index.yaml. Schema (verified against existing GGUF items and
the LocalAI config structs): backend selection via `overrides.backend`; runtime knobs via
either typed config fields (context_size/f16/flash_attention/gpu_layers/batch) or the
`options:` string list (key:value, parsed by grpc-server.cpp set_option).
--------------------------------------------------------------------------------
2.1 Benchmark llama-server flags -> LocalAI model-config mapping
--------------------------------------------------------------------------------
-c 131072 -> context_size: 131072 (LLMConfig.ContextSize, yaml context_size)
-fa on -> flash_attention: "on" (LLMConfig.FlashAttention, yaml flash_attention; string)
-ngl 99 -> gpu_layers: 99 (LLMConfig.NGPULayers, yaml gpu_layers; or omit -> DefaultNGPULayers offloads all)
-b 2048 -> batch: 2048 (schema.PredictionOptions.Batch, yaml batch) [see caveat]
--parallel 128 -> options: ["parallel:128"] (grpc-server.cpp:629; alias n_parallel)
LLAMA_KV_PAGED=1 -> options: ["paged_kv:true"] (grpc-server.cpp:778)
LLAMA_MAX_BATCH_TOKENS=512 -> options: ["max_batch_tokens:512"] (grpc-server.cpp:821; alias mbt)
f16 KV -> f16: true (LLMConfig.F16, yaml f16)
(recommended for paged) -> options: ["kv_unified:false"] (grpc-server.cpp:746 - the per-slot paged
capacity/memory benefit only materializes with a per-sequence cache;
the patch comment explicitly recommends pairing paged with kv_unified:false)
CAVEAT (-ub 512): LocalAI sets params.n_ubatch = params.n_batch = request->nbatch()
(grpc-server.cpp:528,532). There is NO separate config field for n_ubatch, so the
benchmark's `-b 2048 -ub 512` split is NOT exactly reproducible. Options:
(i) set batch: 512 -> n_batch=n_ubatch=512 (matches -ub; the decode-first
max_batch_tokens=512 budget is the dominant prefill lever anyway, and the
benchmark states decode throughput is budget-independent), OR
(ii) set batch: 2048 -> n_ubatch also 2048 (bigger physical batch, more KV scratch).
RECOMMEND (i) batch: 512 for the shipped gallery config (closest to the measured run +
lighter memory). Flag separately: a tiny grpc-server.cpp option `n_ubatch`/`ubatch` could
be added later to honor -b/-ub independently (not required to ship).
--------------------------------------------------------------------------------
2.2 gallery/index.yaml entry - DENSE q36-27b-nvfp4
--------------------------------------------------------------------------------
- name: "qwen3.6-27b-nvfp4-paged"
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
urls:
- https://huggingface.co/<ORG>/Qwen3.6-27B-NVFP4-GGUF # placeholder, section 3
description: |
Qwen3.6-27B dense, native Blackwell NVFP4 (FP4-MMA) GGUF. Configured for LocalAI's
paged-attention llama.cpp backend: on-demand paged KV + decode-first prefill budget.
Benchmarked on GB10/DGX Spark at 90-117% of vLLM dense decode at 1.5-3x lower memory.
license: "apache-2.0" # confirm vs Qwen license
tags: [ llm, gguf, nvfp4, reasoning ]
icon: https://user-images.githubusercontent.com/1991296/230134379-7181e485-c521-4d23-a0d6-f7b3b61ba524.png
overrides:
backend: llama-cpp-paged
f16: true
flash_attention: "on"
context_size: 131072
gpu_layers: 99
batch: 512 # see -ub caveat 2.1; matches the 512 ubatch floor
known_usecases: [ chat ]
options:
- use_jinja:true
- paged_kv:true # LLAMA_KV_PAGED=1
- max_batch_tokens:512 # LLAMA_MAX_BATCH_TOKENS=512 (decode-first QoS budget)
- kv_unified:false # enables the per-slot paged capacity/memory benefit
- parallel:128 # --parallel 128 serving slots
parameters:
model: llama-cpp/models/Qwen3.6-27B-NVFP4-GGUF/q36-27b-nvfp4.gguf
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/Qwen3.6-27B-NVFP4-GGUF/q36-27b-nvfp4.gguf
sha256: <FILL after publish>
uri: https://huggingface.co/<ORG>/Qwen3.6-27B-NVFP4-GGUF/resolve/main/q36-27b-nvfp4.gguf
--------------------------------------------------------------------------------
2.3 gallery/index.yaml entry - MoE q36-35b-a3b-nvfp4
--------------------------------------------------------------------------------
Same shape; the MoE is lighter on memory (~3B active). parallel:128 + budget 256 was the
MoE decode-throughput sweet spot in the sweep, but 512 is fine as a default; if optimizing
purely for saturated MoE decode use max_batch_tokens:256.
- name: "qwen3.6-35b-a3b-nvfp4-paged"
urls: [ https://huggingface.co/<ORG>/Qwen3.6-35B-A3B-NVFP4-GGUF ]
...
overrides:
backend: llama-cpp-paged
f16: true
flash_attention: "on"
context_size: 131072
batch: 512
options:
- use_jinja:true
- paged_kv:true
- max_batch_tokens:512 # or 256 for max saturated MoE decode (sweep winner)
- kv_unified:false
- parallel:128
parameters:
model: llama-cpp/models/Qwen3.6-35B-A3B-NVFP4-GGUF/q36-35b-a3b-nvfp4.gguf
files:
- filename: llama-cpp/models/Qwen3.6-35B-A3B-NVFP4-GGUF/q36-35b-a3b-nvfp4.gguf
sha256: <FILL after publish>
uri: https://huggingface.co/<ORG>/Qwen3.6-35B-A3B-NVFP4-GGUF/resolve/main/q36-35b-a3b-nvfp4.gguf
Note: these are the BENCHMARK serving configs. For an interactive single-user default you
may want a second lighter gallery variant (context_size 16384, parallel 4, drop the budget)
- optional, not required to ship the benchmark reproduction.
================================================================================
3. GGUF PUBLISHING (so the gallery uri: resolves)
================================================================================
The two GGUFs already exist on the DGX dev box (final_benchmark.csv references
q36-27b-nvfp4.gguf and q36-35b-a3b-nvfp4.gguf; README.md "Models" + "Benchmarks"
document provenance: dense = native Blackwell FP4 unsloth W4A4 lineage; MoE = 241 NVFP4
tensors from nvidia modelopt weights). To publish:
1. HF repos (suggest two, under the org that owns the gallery-referenced weights):
<ORG>/Qwen3.6-27B-NVFP4-GGUF (single q36-27b-nvfp4.gguf)
<ORG>/Qwen3.6-35B-A3B-NVFP4-GGUF (single q36-35b-a3b-nvfp4.gguf)
ORG = localai-org (brand) or mudler (personal); pick per ownership of the conversions.
2. Upload each .gguf; compute sha256 (sha256sum) and paste into the gallery `files:` sha256
(LocalAI verifies it on download). Without sha256 the entry still works but loses the
integrity check - fill it.
3. Model card metadata: base_model Qwen/Qwen3.6-*, library_name gguf, quantization NVFP4,
pipeline_tag text-generation, license (confirm Qwen3.6 license terms - apache-2.0 vs
Qwen community license), a note that it REQUIRES the llama-cpp-paged backend (NVFP4 +
paged), and the GB10 benchmark table (link README.md "Benchmarks" numbers).
4. NVFP4 requires a llama.cpp new enough to read the NVFP4 GGUF type. Confirm the pinned
LLAMA_VERSION in backend/cpp/llama-cpp/Makefile supports NVFP4 tensor types (the dev
tree that produced the GGUFs did). If the current pin predates NVFP4 GGUF support, the
backend pin must be bumped OR the paged patch series must carry the NVFP4 reader. THIS
IS A GATING CHECK before the gallery items are usable - verify on a GPU box.
5. Provenance/licensing: the dense conversion derives from unsloth; the MoE from nvidia
modelopt weights. Ensure redistribution of the converted GGUFs is permitted and
attribute upstream in the card.
================================================================================
4. OPEN DECISIONS / BLOCKERS / BUILD COST
================================================================================
BACKEND NAME - RECOMMEND `llama-cpp-paged`.
- llama-cpp-paged (RECOMMENDED): descriptive (it IS the paged variant), hyphenated like
every sibling (llama-cpp/ik-llama-cpp/turboquant/ds4), collision-free in the
changed-backends.js endsWith() suffix scheme, self-documenting in the /backends/known
importer dropdown. Reads correctly next to "turboquant" and "ik-llama-cpp".
- localai-llama-cpp (branding alternative, ACCEPTABLE): keeps the LocalAI brand without a
dot; hyphenated and safe. Use this if marketing wants "LocalAI's own llama.cpp" framing.
Slightly less self-explanatory about WHAT differs (paged) in the dropdown.
- localai-llama.cpp (the working name; NOT RECOMMENDED): the dot makes Dockerfile.localai-
llama.cpp and tag-suffix -cpu-localai-llama.cpp the only dotted ones in the repo, and
".cpp" looks like a file extension to the suffix matcher. Avoid.
BLOCKERS / GATING CHECKS (cannot be closed read-only, no GPU here):
1. NVFP4 GGUF read support in the pinned LLAMA_VERSION (section 3.4). Must verify on GPU.
If unsupported, bump the pin (which also affects stock llama-cpp) or carry the reader.
2. The two GGUFs are not yet on HF (section 3). Gallery uri + sha256 are placeholders
until upload. Blocks gallery validation only, not the backend build.
3. -ub vs -b split (section 2.1) is not exactly reproducible without a tiny grpc-server
option; shipped config uses batch:512. Minor, not a blocker.
4. Flipping stock LLAMA_PAGED?=off changes stock's shipped artifact (de-risking, intended)
- get explicit sign-off since it alters a heavily-used backend's build.
PLATFORM SHIP MATRIX (RECOMMENDED PHASING - the variant is cheap because it reuses the same
base-grpc-* prebuilt bases and the same compile machinery, so each row is just CI minutes):
Phase 1 (the benchmark target - GB10/Blackwell is CUDA):
- cuda12 amd64, cuda13 amd64, cuda13 arm64 (sbsa), l4t-cuda-12 arm64 (NVFP4/paged win)
- cpu-all amd64 + cpu-all arm64 (the single CPU_ALL_VARIANTS build; baseline coverage)
Phase 2 (parity with stock llama-cpp coverage, only if demand):
- metal-darwin-arm64 (1.11), vulkan amd64/arm64, rocm amd64, intel sycl f16/f32
Defer rocm/sycl/vulkan/metal unless asked - the paged + NVFP4 story is GPU/CUDA-centric
and these add CI cost without a clear consumer.
BUILD-COST ESTIMATE PER PLATFORM (with warm base-grpc-* base + ccache; the paged TUs are
~byte-identical to stock so a SHARED ccache id makes most objects free):
- CPU_ALL_VARIANTS (per arch): ~15-30 min warm / ~35-50 min cold. arm64 adds a gcc-14
apt step. Two arches + a merge job.
- CUDA (per arch): ~25-45 min warm / ~45-75 min cold (nvcc dominates; ccache helps less
across CUDA arch flag changes). amd64 cuda12 + cuda13, arm64 cuda13 + l4t = 4 jobs.
- Metal/Darwin (if Phase 2): native macos-14 runner, ~20-35 min with the ccache cache.
- No base-images.yml change and no bootstrap dispatch (reuses existing base-grpc-* tags),
so the only new CI cost is the per-row build minutes above. PR builds read cache, don't
write; first master build per row pays the cold cost once, then warm.
VERIFICATION (post-implementation, needs a GPU box - out of scope here):
- `make backends/llama-cpp-paged` builds + installs locally (from-source path).
- Confirm stock `make backends/llama-cpp` now builds clean (no paged-kv-manager.cpp in the
checkout) - proves the split.
- Load a published NVFP4 GGUF via the gallery entry, hit /v1/chat/completions, confirm the
server log shows LLAMA_KV_PAGED engaged (LLAMA_KV_PAGED_DEBUG trace) and the configured
max_batch_tokens/parallel took effect.
- go test ./core/gallery/importers/... green (importer drop-in case).
- node scripts/changed-backends.js dry-run: editing backend/cpp/llama-cpp/* retriggers
llama-cpp-paged (cross-trigger), editing backend/cpp/llama-cpp-paged/* triggers it too.
================================================================================
END OF PLAN
================================================================================

View File

@@ -1,75 +0,0 @@
# Paged bit-exactness gate - per path (canonical references)
## TL;DR
The greedy decode of the **paged** path does not byte-match the **non-paged**
path for the MoE model. This is a **benign FP-accumulation-order difference of
the paged attention reduction**, KL-validated against the f16 reference. It is
**not a bug**. The bit-exactness gate is therefore **per path**:
| path | model | canonical md5 |
|------|-------|---------------|
| non-paged | MoE q36-35b-a3b-nvfp4 | `07db32c2bcb78d17a43ed18bc22705cd` |
| paged | MoE q36-35b-a3b-nvfp4 | `8cb0ce23777bf55f92f63d0292c756b0` |
| non-paged | dense q36-27b-nvfp4 | `5951a5b4d624ce891e22ab5fca9bc439` |
| paged | dense q36-27b-nvfp4 | `5951a5b4d624ce891e22ab5fca9bc439` (bit-exact to non-paged) |
Gate command (chat-template / conversation path):
```
llama-completion -m MODEL -ngl 99 -fa on -p "The capital of France is" \
-n 48 --temp 0 --seed 1
# paged: prefix with LLAMA_KV_PAGED=1 LLAMA_MOE_FORCE_GRAPHS=1
```
Note: use the default chat-template path (do **not** pass `-no-cnv`; raw
completion lands in a different md5 namespace).
**Future paged-MoE regressions compare to the PAGED reference `8cb0ce23`, not to
the non-paged `07db32c2`.** Dense is bit-exact across paths, so dense uses the
single reference `5951a5b4`.
## Why dense is bit-exact but MoE is not
Dense paged decode reproduces the non-paged reduction order exactly, so dense
greedy md5 is identical across paths. The MoE path runs additional kernels (the
NVFP4 MoE GEMM + expert routing) whose multi-kernel accumulation order differs
between the paged and non-paged attention layouts. Over a long greedy decode this
flips a small number of near-tied argmaxes, changing the byte stream. The same
divergence is present on the 0028 baseline, with `LLAMA_MOE_FORCE_GRAPHS` on or
off, and with the patch-0029 block-table cache on or off - it is a property of
the paged attention path, not of any one lever.
## KL evidence that the paged path is sound (the load-bearing check)
`llama-perplexity --kl-divergence` on `q36-35b-a3b-nvfp4.gguf`, 16 chunks,
`-c 512 -ngl 99 --seed 1`, base logits from the f16 reference
(`darwin_36b_opus/f16.gguf`, PPL 7.3734):
| comparison | PPL(Q) | KL divergence | Same top p | Cor |
|------------|-------:|--------------:|-----------:|----:|
| f16 reference | 7.3734 | - | - | - |
| **non-paged** vs f16 | 7.3896 | 0.136597 +/- 0.003157 | 84.314% | 97.68% |
| **paged** vs f16 | 7.4009 | 0.136000 +/- 0.003285 | 84.828% | 97.58% |
| paged vs non-paged (direct) | 7.4009 (base 7.3818) | 0.050011 +/- 0.001653 | 89.044% | 99.04% |
Direct paged-vs-non-paged: Mean Delta-p = 0.079% (no bias), RMS Delta-p = 6.187%.
### Verdict: BENIGN
- **Paged does not diverge from the f16 ground truth more than non-paged does.**
KLD(paged||f16) = 0.13600 <= KLD(nonpaged||f16) = 0.13660, and PPL(paged) =
7.4009 ~ PPL(nonpaged) = 7.3896 (difference 0.011, far inside the +/- 0.29
error bars). A real paged-MoE correctness bug would push paged measurably
*further* from f16; it does not (it is marginally closer).
- **Paged and non-paged cluster together.** They agree with each other (KLD 0.050,
89.0% same-top-p) more than either agrees with f16 (KLD ~0.137, ~84% same-top-p),
with essentially zero probability bias. That is the signature of two equivalent
FP-reorderings of the same quantized model, both equally approximating the f16
ground truth - not a quality regression.
- The direct same-top-p of 89.0% is below a naive ">99%" heuristic, but that
heuristic is calibrated for higher-precision models. In a 4-bit (NVFP4) model
logit near-ties are abundant, so a different-but-equivalent reduction order
flips ~11% of argmaxes with no quality cost (proven by the equal KLD-to-f16 and
zero Delta-p bias).
Therefore the canonical gate is per path, and `8cb0ce23` is the validated paged
reference for the MoE deployment path.

View File

File diff suppressed because it is too large Load Diff

View File

@@ -1,156 +0,0 @@
# llama.cpp patch series — paged attention (vLLM-parity engine)
A **stacking** series: each patch is a small, self-contained, independently-buildable step toward an
in-model paged-attention engine. They apply in numeric order on top of the pinned `LLAMA_VERSION`
(`backend/cpp/llama-cpp/Makefile`). The build applies them automatically after checkout (see the
`llama.cpp:` target). Keeping the work as ordered patches — rather than one big diff — is what lets us
**rebase cleanly across llama.cpp bumps and avoid drift**: when a patch stops applying, only that small
patch needs fixing, and the failure points at exactly which step the upstream change touched.
## Base
- `LLAMA_VERSION` pin in `../Makefile`. **All patches are generated against that exact commit.** Bumping
the pin = re-run the regen workflow below and fix only the patches that no longer apply.
## The series (phases → patches)
| # | Patch | What | Verifies |
|---|-------|------|----------|
| 0001 | `0001-vendor-paged-kv-manager.patch` | Add `src/paged-kv-manager.{h,cpp}` (vLLM-parity block manager, CPU foundation) + CMake; no behavior change | builds; unit-tested separately |
| 0002 | `0002-paged-kv-storage.patch` | Shared block-pool KV tensor + `set_rows`-by-slot writes, behind `LLAMA_KV_PAGED` | builds; write/gather round-trip |
| 0003 | `0003-paged-gather-read.patch` | `build_attn_paged` gather-read in `llama-graph.cpp` | **Gate 0**: token-identical greedy gen, single + multi-seq |
| 0004 | `0004-paged-ondemand-alloc.patch` | On-demand block allocation via PagedKVManager | max concurrent seqs before OOM |
| 0005 | `0005-paged-continuous-batching.patch` | Block-granular admit/evict in the server slot path | tok/s vs concurrency, mixed-length |
| 0006 | `0006-paged-prefix-caching.patch` | Block-hash cross-request prefix dedup | TTFT + memory on shared prefixes |
Each row is a separate `git commit` on the dev branch (below), exported 1:1 as a patch. Default off
(`LLAMA_KV_PAGED`) until Gate 0 (0003) is green, so partial series never changes stock behavior.
## Regen workflow (the anti-drift recipe)
```sh
# 1. check out the exact pin into a dev tree
git -C /tmp clone https://github.com/ggml-org/llama.cpp llama-dev && cd /tmp/llama-dev
git checkout <LLAMA_VERSION from ../Makefile>
git checkout -b paged
# 2. apply the current series (each becomes a commit), or develop the next patch
git am /path/to/backend/cpp/llama-cpp-localai-paged/patches/paged/00*.patch # or `git apply` + commit per patch
# 3. iterate a phase as ONE commit, then export the whole series 1:1
git format-patch <LLAMA_VERSION>..paged -o /path/to/backend/cpp/llama-cpp-localai-paged/patches/paged/ --zero-commit -N
# 4. on a pin bump: rebase `paged` onto the new pin; only conflicting patches need edits; re-export.
```
## Build integration
The series is owned by this backend (`backend/cpp/llama-cpp-localai-paged`), not by the stock
`llama-cpp` backend, which is pure upstream. `../Makefile` (the paged wrapper) clones the pinned
`llama.cpp` via the copied stock build infra, then applies this series onto the cloned tree with the
same strict `git apply` the stock build uses for base patches:
```
for p in $(PAGED_PATCHES_DIR)/0*.patch; do git apply --verbose "$p" || exit 1; done
```
All variants (avx/avx2/avx512/cuda/…) clone + apply into their own build copy, so the series ships
everywhere without ever touching the stock `llama-cpp` source tree.
## Latest mirror check
Phase 37 re-verified the mirror invariant after adding patch `0063`:
```text
base=0ed235ea2c17a19fc8238668653946721ed136fd
applied_tree=dedb1182910eafe9f6875588dc8285bfb544cce5
fork_tree=dedb1182910eafe9f6875588dc8285bfb544cce5
```
The check used a fresh worktree at `LLAMA_VERSION`, applied every
`patches/paged/0*.patch` with strict `git apply`, staged the result, and compared
`git write-tree` to canonical fork branch `localai-paged` at
`2d590d770 feat(cuda): trace cublas tensor names`.
Phase 69 re-verified that the committed LocalAI patch series still matches the
Phase37 fork tip, and then dry-ran the additive patch export needed for the
current local fork HEAD. No generated patch files were edited in Phase69 because
the repo policy requires pushing the fork branch before regenerating the LocalAI
series, and pushes still require explicit approval.
Committed-series check:
```text
base=0ed235ea2c17a19fc8238668653946721ed136fd
applied_tree=dedb1182910eafe9f6875588dc8285bfb544cce5
patch_tip_tree=dedb1182910eafe9f6875588dc8285bfb544cce5
fork_head_tree=fcf5720b659c5e1e2b487ccf3c8f7289bb12b9c4
match_patch_tip=yes
match_fork_head=no
patch_count=54
```
Dry-run export from `2d590d770..ea0875d14` produced ten source-only candidate
patches:
```text
0064-feat-server-trace-serving-admission-batches.patch
0065-feat-server-add-admission-trace-histograms.patch
0066-feat-server-add-TTFT-prefill-first-scheduler-mode.patch
0067-feat-server-cap-TTFT-prefill-first-decode-deferral.patch
0068-feat-server-gate-TTFT-defer-by-prompt-backlog.patch
0069-test-cuda-cover-W4A16-direct-activation-policy.patch
0070-feat-cuda-route-W4A16-direct-activation-stub.patch
0071-feat-cuda-trace-layout-tensor-names.patch
0072-feat-cuda-trace-activation-quant-routes.patch
0073-feat-cuda-gate-BF16-cuBLAS-F32-output.patch
```
Projected-series check with current `0001..0063` plus temp `0064..0073`:
```text
base=0ed235ea2c17a19fc8238668653946721ed136fd
applied_plus_missing_tree=fcf5720b659c5e1e2b487ccf3c8f7289bb12b9c4
fork_head_tree=fcf5720b659c5e1e2b487ccf3c8f7289bb12b9c4
match_fork_head=yes
current_patch_count=54
missing_patch_count=10
projected_patch_count=64
```
Next mirror action after explicit push approval:
1. Push `/home/mudler/_git/llama.cpp` branch `localai-paged` to
`fork/localai-paged`.
2. Regenerate or copy the equivalent source-only `0064..0073` patches from the
pushed fork.
3. Repeat the projected-series tree hash check above against fork HEAD before
committing generated patches.
## Status
- **0001 vendor manager — DONE.** Applies clean to the pin; builds into `libllama`.
- **0002 block placement — DONE + VERIFIED.** Built `llama-simple` at the pin; greedy generation is
**token-identical** stock vs `LLAMA_KV_PAGED=1` (Qwen3-0.6B), paged branch confirmed firing.
- **0003 gather-read — DONE + VERIFIED (Gate 0 green).** Implemented in the **additive** form
(see `../README.md`): all logic in new `src/paged-attn.{h,cpp}` (a `llm_graph_input_i` gather-index
subclass + the K/V/mask gather), hooked by **one** line in `build_attn` + **two** thin accessors on
`llama_kv_cache_context` + 1 CMake line (216 insertions; no edit to `llm_graph_input_attn_kv` or
`llama-graph.h`). Greedy generation is **token-identical** stock vs `LLAMA_KV_PAGED=1` (Qwen3-0.6B,
**9/9** across 3 prompts × {32,96,128} tokens), with `n_gather=71 < n_kv=256` confirming real
compaction. Patch: `0003-paged-gather-read-env-LLAMA_KV_PAGED.patch`.
- **Key correctness finding:** `get_gather_idxs` must emit cells **sorted by token position**. The CPU
flash-attn online softmax reduces cells in physical-array order and is FP-order-sensitive, so 0002's
scattered placement *alone* (full-window read, no gather) diverges from stock once a sequence crosses
the first 16-cell block. The position-sorted gather reproduces stock's exact reduction order -> bit-
identical, not merely mathematically equivalent. So 0002 is the placement substrate; **0003 is what
makes paged placement token-identical under flash-attn.**
- 00040006 follow.
### Honest parity note (important)
This series delivers the paged-attention **engine** (capacity + scheduling + prefix sharing). It does **not**
by itself reach vLLM throughput parity, because the measured prefill bottleneck is the **FP4 MoE GEMM kernel**
(Lever 3: `mul_mat_q<MXFP4>` ~22 TFLOP/s, ~27× behind vLLM) — a *per-token compute* gap that paging does not
touch. Paged attention closes the **concurrency/memory** gap (more sequences, prefix reuse); the prefill/throughput
gap additionally needs the tcgen05/CUTLASS grouped-GEMM (deferred, upstream-grade, no shortcut — see
`../README.md`). So full vLLM parity = this series **AND** the
kernel; neither alone suffices.

View File

@@ -1,76 +0,0 @@
# PREFILL_GEMM_RESULTS - option (a) dequant->bf16 cuBLAS, measured on GB10
Companion to `PREFILL_GEMM_SCOPE.md`. This records the GPU A/B for the #1
prefill lever (route large-M NVFP4 dense GEMMs off FP4-MMQ onto dequant->bf16
cuBLAS / nvjet). Shipped as patch `0033`, **default-off** because the measured
result is a regression on this hardware.
Hardware: NVIDIA GB10 (sm_121), CUDA 13.0. Backend pin `9d5d882d`.
Models: `q36-27b-nvfp4.gguf` (dense), `q36-35b-a3b-nvfp4.gguf` (MoE).
Binary: `build-cuda/bin/llama-batched-bench -fa on -ngl 99`, `LLAMA_KV_PAGED=1`.
A/B is a single build toggled by `LLAMA_FP4_PREFILL_M` (0 = MMQ baseline, >0 =
route prefill M>threshold to bf16 cuBLAS), so it isolates exactly this lever.
## 1. Bit-exact / numeric gate (PASS - divergence benign)
| Gate | Result |
|---|---|
| `test-backend-ops -o MUL_MAT` (default, threshold off) | 1146/1146 pass |
| `test-backend-ops -o MUL_MAT_ID` (default) | 806/806 pass (MoE untouched) |
| `test-backend-ops -o MUL_MAT`, path FORCED (`LLAMA_FP4_PREFILL_M=64`) | NVFP4 large-M cases (m=2048/1600/2050, n=128, k=2048) green CUDA-vs-CPU |
| greedy md5, short prefill (< threshold), lever vs base | identical: `5951a5b4d624ce891e22ab5fca9bc439` (== documented dense reference; decode byte-untouched) |
| greedy md5, long prefill (> threshold, exercises bf16 path), lever vs base | identical: `5f3967df5781445feeb25762abb9eae7` (the new FP path flips no greedy argmax) |
The new path (NVFP4->bf16 round, bf16 tensor cores, f32 accumulate) is a
different FP path from fused FP4xQ8_1 MMQ, but it is precision-neutral-to-better:
keeping activations in bf16 instead of Q8_1 is strictly more precise, and the
greedy output is byte-identical. This matches the scope's prediction
(KLD(dequant-bf16 || f16) <= KLD(FP4-MMQ || f16)).
## 2. Performance (REGRESSION - the lever loses on GB10)
S_PP (prefill tokens/s), q36-27b dense, A/B `LLAMA_FP4_PREFILL_M` off vs on:
| prefill ubatch M | npl | base S_PP (MMQ) | lever S_PP (bf16 cuBLAS) | delta |
|---|---|---|---|---|
| 512 | 32 | 958.99 | 486.65 | -49% |
| 1024 | 8 | 1013.65 | 587.27 | -42% |
| 2048 | 8 | 918.46 | 649.42 | -29% |
Default-off control (no env): S_PP 966.98 == base (within noise) -> the patch is
inert by default.
## 3. Why it loses (the scope premise was wrong for GB10)
The scope assumed FP4-MMQ is register-bound to ~3% of FP4 peak at large M, so a
vendor large-M kernel would win. **Measured, FP4-MMQ at M=512..2048 beats
dequant->bf16 cuBLAS by 29-49%.** Two compounding reasons:
1. **bf16 tensor-core peak is ~half FP4 peak on GB10.** Even a perfect bf16 GEMM
caps at ~half the throughput the FP4-MMA path can reach.
2. **The dequant tax is an un-amortized memory pass.** Per prefill step the new
path reads FP4 weights (~0.5 B/elt), writes bf16 (2 B/elt), then the GEMM
reads bf16 (2 B/elt) = ~8x the weight byte traffic of the FP4-MMQ read
(~0.5 B/elt). The dequant write is M-independent, so it only amortizes as M
grows: the gap shrinks 49% -> 42% -> 29% from M=512 -> 2048 but never crosses
even at M=2048 (above the default n_ubatch).
This is also consistent with the README decode finding that the dense path was
already ~96-97% of vLLM - the dense GEMM was never the bottleneck the way the
prefill ground-truth (measured on the MoE decision model) implied.
## 4. Status of the phases
- **Phase 1 (dense): REJECTED on GB10**, landed default-off as a validated,
env-gated scaffold (mechanism + bit-exact gate reusable by option (b) and by
non-GB10 hardware where bf16 may fare differently).
- **Phase 2 (MoE grouped large-M): NOT implemented.** It inherits the same
bf16-peak < FP4-peak ceiling plus a per-expert dequant, so a grouped
bf16-cuBLAS would regress for the same reason; the MoE id-path also has the
graph-safety catch (a false `should_use_mmq` falls to the host-sync sorted
loop, not CUDA-graph-safe). Not worth the multi-day grouped-cuBLAS + graph
work on a path the dense A/B already shows loses.
- **The only route to a real prefill GEMM win is option (b)** - a native
Blackwell FP4-MMA large-M kernel (multi-week), to greenlight only if the
prefill regime is funded. The committed scaffold gives option (b) its
M-threshold routing and its bit-exact gate for free.

View File

@@ -1,264 +0,0 @@
# PREFILL_GEMM_SCOPE - large-M NVFP4 expert/dense GEMM (design only)
**Status: DESIGN + PLAN ONLY. No kernel written, no GPU run in this pass.**
This scopes the #1 prefill lever for `llama-cpp-localai-paged`: the NVFP4 weight
GEMM at large M (prefill), where llama.cpp's `mul_mat_q` (MMQ) NVFP4 path is far
slower than vLLM's `marlin_moe_wna16` (MoE) + cutlass/nvjet (dense). Per the
prefill ground-truth that motivated this scope, the GEMM bucket is ~232 us/tok
(paged) vs ~68 us/tok (vLLM) - 3.4x slower, ~51% of the paged-vs-vLLM prefill
gap (164 us/tok).
> **Regime warning (read first).** Every "GEMM is at the BW floor / ties vLLM"
> conclusion in `README.md` section 5 is a **DECODE** finding (M<=128,
> bandwidth-bound). This document is about **PREFILL** (large M, compute /
> tensor-core-throughput bound) - a different regime, which is exactly why the
> rejected "W4A16-Marlin MoE GEMM" lever is revisited here **for prefill only**.
> The 232/164/68 us/tok prefill bucket came from the prefill ground-truth that
> commissioned this scope and is **not** in a committed in-repo profile (the
> committed profiling - `GAP_PROGRESS.md` etc. - is decode-focused). Per the
> "profile-don't-assume" rule in `.agents/vllm-parity-methodology.md`, **step 0 of
> any build is to re-confirm the prefill GEMM bucket on GPU** (nsys, prefill-only
> window) before touching code.
---
## 1. Why `mul_mat_q` is slow at large M (confirmed from source)
Source: `ggml/src/ggml-cuda/mmq.cu`, `mmq.cuh` at this backend's pin (`9d5d882d`).
MMQ is built for the **M<=128 decode tile**. Three structural facts from the code:
1. **The M (column/token) tile is capped at 128.**
`get_mmq_x_max_host()` / `get_mmq_x_max_device()` (mmq.cuh ~108-140) return
`128` on Blackwell (`turing_mma_available(cc)`), and the host launch loop
(mmq.cuh ~4237) picks `mmq_x_best` only to *minimise the column-tile count for
`ncols_max`, never exceeding `mmq_x_max`*. So a prefill ubatch of M=512 (or
4096) tokens is processed as many `mmq_x<=128` column-tiles. The compile-time
accumulator tile is `mmq_x`-wide; there is no large-M (e.g. 256-wide) tile
variant. The whole tile-selection machinery exists to pick a *small* tile for
*small* batches, not to grow for large ones.
2. **The FP4-MMA kernel is register-bound to 1 CTA/SM.**
`mul_mat_q` for FP4 is `__launch_bounds__(warp_size*nwarps, min_blocks=1)`
(mmq.cuh ~3579-3585), i.e. 256 threads, 1 resident block/SM (~255 regs/thread).
The patch-0017 comment in-tree states this plainly: the kernel is
"REGISTER-bound to 1 CTA/SM ... the under-occupancy that strands the kernel at
~3% of FP4 peak at M=128." At large M the work per tile is bigger, but with one
CTA/SM the tensor cores still stall on LPDDR5x / shared-memory weight loads
with no CTA-level latency hiding - the design has no async multi-stage global->
shared pipeline (cp.async double-buffering) that large-M GEMMs need.
3. **Per-tile fixed overheads amortise poorly only because the tile stays small.**
Each tile re-stages weights into shared memory, runs the `MMQ_ITER_K_FP4=512`
K-loop, and the activations are quantized to Q8_1 (`quantize_mmq_fp4_cuda`,
block_fp4_mmq = FP4 weights x int8 activations). For decode this is the right
trade (FP4 weight traffic is the bottleneck). For large-M prefill the GEMM is
compute-bound, so the right structure is big tensor-core output tiles (e.g.
128x256), a deep async load pipeline, and full SM occupancy - exactly what
cutlass 3.x / nvjet (cuBLAS) and marlin implement and MMQ does not.
Patch 0017 already proved every *cheap* large-tile/occupancy lever inside MMQ
(`GGML_CUDA_FP4_MMQ_Y`, `GGML_CUDA_FP4_MINBLOCKS`) is a no-win on GB10 - because
the limit is the small-tile kernel *structure*, not a tunable. To win at large M
you must leave MMQ for a large-M kernel.
---
## 2. Options (feasibility / bit-exactness / effort)
### Key enabling facts already in the tree
- **NVFP4 -> bf16/f16 dequant kernels already exist.** `convert.cu` defines
`dequantize_row_nvfp4_cuda`; `ggml_get_to_bf16_cuda` / `ggml_get_to_fp16_cuda`
/ `ggml_get_to_fp16_nc_cuda` all return it for `GGML_TYPE_NVFP4`. The
non-Blackwell fallback ("falls back to dequant", README s2) already uses this.
- **cuBLAS on GB10 dispatches to nvjet** (NVIDIA's JIT tensor-core GEMM) - the
committed profiles already show `nvjet lm_head` and `nvjet non-FP4 cublas GEMM`
rows. So a dequant->cuBLAS bf16 GEMM lands on a vendor-tuned large-M kernel for
free.
- **BUT NVFP4 is explicitly excluded from the tensor-core cuBLAS path.** In
`ggml_cuda_op_mul_mat_cublas` (ggml-cuda.cu ~1659) the `use_fp16` predicate
begins `src0->type != GGML_TYPE_NVFP4 && ...`. So if NVFP4 reaches cuBLAS today
it falls to the `else` branch: dequant to **F32** + `cublasSgemm` (**no tensor
cores**) - useless for prefill. Relaxing this one exclusion (route NVFP4 to the
bf16/f16 tensor-core branch, where `to_*_cuda(NVFP4)` already exists) is the
pivot that makes option (a) a few-line change rather than a kernel.
### (a) Dequant -> cuBLAS/cutlass bf16 GEMM for large M -- RECOMMENDED
Dequant the NVFP4 weights to bf16 (transient pool buffer) once per prefill step,
then a large-M tensor-core `cublasGemmEx` (CUBLAS_COMPUTE_32F accumulate, bf16
inputs). Activations stay bf16 (not Q8_1-quantized).
- **Feasibility: HIGH.** All pieces exist (dequant kernels, cuBLAS bf16 path,
pool allocator). The only code change for the dense path is (i) make
`ggml_cuda_should_use_mmq` return false for NVFP4 dense above an M threshold so
the dispatch falls through to `ggml_cuda_op_mul_mat_cublas`, and (ii) relax the
`src0->type != GGML_TYPE_NVFP4` exclusion so it dequants to bf16 and uses
`cublasGemmEx` tensor-core, not f32 Sgemm.
- **Cost model (the crux - why it wins ONLY at large M).** Dequant is one extra
weight-sized memory pass (read ~0.5B/elt FP4 + scales, write 2B/elt bf16). The
bf16 GEMM then reads weights as bf16 = **4x the byte traffic of the FP4-MMQ
read**. At small M (decode) this 4x weight traffic dominates -> bf16-cuBLAS
loses -> keep MMQ (this is why decode stays FP4-MMQ; consistent with the
README decode verdict). At large M the GEMM is compute-bound and weight traffic
is amortised over hundreds of columns, so the 4x is cheap and cuBLAS's mature
large tiles + async pipeline + full occupancy dominate MMQ's 3%-of-peak small
tile. The dequant pass itself is ~one weight-read amortised over the whole
prefill step - negligible at large M.
- **Honest ceiling.** GB10 bf16 tensor-core peak is ~**half** the FP4 tensor-core
peak. A bf16 cuBLAS GEMM at ~70-80% of bf16 peak is ~35-40% of FP4 peak. That
is a huge jump from MMQ's ~3% large-M utilisation, but it is **not** automatic
full vLLM parity (vLLM prefill uses 4-bit weight tiles, staying near FP4-class
throughput). Expect this to recover most, not all, of the 232->68 gap. See s4.
- **Bit-exactness: NEW FP path** (NVFP4->bf16 round, bf16 TC, f32 accumulate) vs
fused FP4xQ8_1 MMQ. **Not byte-identical** - gate per-path via KLD exactly like
the paged-MoE `8cb0ce23` precedent (README s5 / `PAGED_BITEXACT_NOTE.md`). It
should pass *easily and favourably*: keeping activations in bf16 instead of
Q8_1 is strictly more precise than the MMQ path, so KLD(dequant-bf16 || f16)
should be <= KLD(FP4-MMQ || f16). This is a precision-neutral-to-better change,
not a precision regression like the rejected lever 4.
- **Effort: LOW-MEDIUM (a few days).** Dispatch flip + exclusion relax + an M
threshold + the KL gate + a prefill bench. No new kernel. Dense first; MoE is
the harder follow-on (see (c)/plan).
- **Memory note.** Dequant into a *transient* pool scratch per step (do **not**
cache bf16 weights - a persistent bf16 copy is 4x VRAM for those tensors and
would erase the backend's "1.5-3x less memory" property). The per-step dequant
pass is the price of keeping the model FP4-resident.
### (b) Marlin-style fused NVFP4 large-M MoE GEMM (port `marlin_moe_wna16`)
Port vLLM's marlin grouped MoE kernel (4-bit weights, f16 activations, dequant-
in-register, async cp.async pipelines, swizzled layouts).
- **Feasibility: LOW (hardest).** Marlin is a hand-tuned CUTLASS-class kernel and
is **not NVFP4-aware** (it targets wna16 group-quant, not NVFP4's 16-elt blocks
with ue4m3 micro-scales). You would either (i) adapt marlin to dequant NVFP4
in-register and accumulate in f16 (abandoning native Blackwell FP4-MMA), or
(ii) write a brand-new Blackwell sm_121 FP4-MMA large-M kernel - which is
essentially re-implementing what cutlass 3.x / nvjet already give you via (a).
- **Bit-exactness:** new FP path, KL-gate (same as (a)).
- **Effort: HIGH (multi-week, high risk),** kernel + layout + Blackwell MMA
scheduling + graph-safety + the bit-exact gate.
- **Verdict: do NOT start here.** Its only structural advantage over (a) is 4-bit
weight traffic, which matters only when BW-bound = small M = **decode**, the
regime already rejected. At large M (a) reaches the same vendor large-M kernels
for ~1% of the effort. Keep (b) on the shelf as the *only* route to true 68
us/tok parity if (a)'s bf16 ceiling proves insufficient and the win justifies a
multi-week kernel.
### (c) M-threshold routing (the integration mechanism for (a))
Not an alternative to (a) - it is *how* (a) is wired. Keep FP4-MMQ for decode
(M<=threshold), switch to the large-M path for prefill.
- **Cleanest hook:** `ggml_cuda_should_use_mmq(type, cc, ne11_or_ne12, n_experts)`
already receives M (`ne11` dense / `ne12` MoE tokens). Add an NVFP4+Blackwell
branch: return false when M > `LLAMA_FP4_PREFILL_M` (default e.g. 256-512,
env/`-D` tunable, default value chosen so default == today's behaviour until
validated). It is called from both `ggml_cuda_mul_mat` (~2573/2582) and
`ggml_cuda_mul_mat_id` (~2664), so one edit covers dense + MoE routing.
- **Dense fallthrough is clean:** `ggml_cuda_mul_mat` final `else` ->
`ggml_cuda_op_mul_mat(..., ggml_cuda_op_mul_mat_cublas, ...)` -> with the
exclusion relaxed, dequant->bf16->`cublasGemmEx`. Works.
- **MoE fallthrough is NOT clean (the catch):** in `ggml_cuda_mul_mat_id`, a
false `should_use_mmq` falls to `should_use_mmf` (no NVFP4 support) then to the
**host-side sorted per-expert loop** with a `cudaStreamSynchronize` (ggml-cuda.cu
~2700) - slow and **not CUDA-graph-safe** (it would break the MoE re-graph,
patch 0025). So MoE large-M needs a *dedicated graph-safe grouped GEMM* (dequant
the expert-gathered weights to bf16 + `cublasGemmGroupedBatchedEx`, CUDA 12.5+,
over the existing `expert_bounds`/`ids_dst` sorted layout), not a bare
fallthrough. This is why the plan ships **dense first, MoE second**.
---
## 3. Recommended approach + implementation plan
**Recommendation: (a) dequant->bf16 cuBLAS, wired via (c) M-threshold routing,
dense-path first, MoE grouped-cuBLAS second. Reject (b).**
### Phase 0 - confirm the bucket on GPU (no code)
- nsys prefill-only window (`-npp <large> -ntg 0/1`, exclude the graph-capture
step) on q36-27b dense and q36-35b-a3b MoE at the backend pin. Confirm the
NVFP4 `mul_mat_q` / `mul_mat_id` bucket is ~232 us/tok and that it is
compute-bound at prefill M (check tensor-core active % low, not BW-saturated).
If the bucket is not what the ground-truth claims, stop and re-scope.
### Phase 1 - dense large-M NVFP4 -> bf16 cuBLAS (the bankable win)
Files / edits:
1. `ggml/src/ggml-cuda/mmq.cu` - `ggml_cuda_should_use_mmq`: add
`if (type==GGML_TYPE_NVFP4 && blackwell_mma_available(cc) && ne11 > LLAMA_FP4_PREFILL_M && n_experts==0) return false;`
(n_experts==0 = dense only in Phase 1). Default threshold == effectively
disabled until A/B-validated, env/`-D` overridable (mirror the 0017
`GGML_CUDA_FP4_*` knob style + in-tree comment).
2. `ggml/src/ggml-cuda/ggml-cuda.cu` - `ggml_cuda_op_mul_mat_cublas`: relax the
`src0->type != GGML_TYPE_NVFP4` guard in `use_fp16` (prefer a dedicated bf16
branch: NVFP4 -> `ggml_get_to_bf16_cuda` -> `cublasGemmEx` CUDA_R_16BF /
COMPUTE_32F, matching the existing BF16 src0 branch for best accuracy).
3. Transient pool scratch for the dequanted weights (reuse `ggml_cuda_pool_alloc`
as the existing branch does; no persistent allocation).
### Phase 2 - MoE grouped large-M (the harder, higher-value follow-on)
1. New grouped path reached from `ggml_cuda_mul_mat_id` when
`should_use_mmq`==false for NVFP4+large-M+`n_experts>0`: dequant the
expert-gathered weights to bf16 and run `cublasGemmGroupedBatchedEx` over the
existing `expert_bounds` / `ids_dst` sorted layout that `mul_mat_q` already
builds. Reuse the patch-0023 de-dup'd activation gather where applicable.
2. **Must stay CUDA-graph-safe** - no host sync (do not fall into the legacy
sorted loop). Validate the MoE re-graph (patch 0025 / `LLAMA_MOE_FORCE_GRAPHS`)
still captures.
### The bit-exact / KL gate (both phases)
- Greedy md5 on the standard prompt (README s5) to detect *unexpected* divergence
on the non-prefill paths (must stay == the per-path reference: dense
`5951a5b4`, paged-MoE `8cb0ce23`). The large-M path itself will differ -> gate
it by KLD vs the f16 reference, requiring `KLD(new||f16) <= KLD(FP4-MMQ||f16)`
and PPL within the established band, recorded in `PAGED_BITEXACT_NOTE.md`.
- `test-backend-ops` MUL_MAT / MUL_MAT_ID at NVFP4 **prefill shapes** (large M)
CUDA0-vs-CPU, plus the existing decode shapes to prove decode is byte-untouched
(default threshold keeps decode on MMQ).
### The bench
- `llama-batched-bench -fa on -ngl 99` reporting **S_PP** (prefill t/s), swept
over prefill length and `npl`, A/B with `LLAMA_FP4_PREFILL_M` off vs on, dense
and MoE, vs stock and vs the vLLM prefill reference. Per-lever A/B discipline
(`.agents/vllm-parity-methodology.md`): one knob at a time, record the rejected
threshold values too.
---
## 4. Honest risk + expected speedup
- **Phase 1 (dense) is a tractable routing change, not a kernel project** - days,
low risk. It reuses existing dequant kernels and the existing nvjet/cuBLAS
large-M path; the net new code is a threshold + a one-line exclusion relax + a
KL gate.
- **Phase 2 (MoE) is medium risk** - the grouped-batched cuBLAS wiring +
CUDA-graph-safety is real work (the bare fallthrough is a slow, graph-breaking
host loop), but still far short of a from-scratch kernel.
- **Will the GEMM bucket hit 232 -> ~68 us/tok (full vLLM parity)? Honestly, no -
not from bf16-cuBLAS alone.** bf16 tensor-core peak on GB10 is ~half FP4 peak,
so the realistic floor for a dequant->bf16 GEMM is ~**90-130 us/tok** (roughly
35-45% of FP4 peak at ~70-80% of bf16 peak). That recovers ~**60-75%** of the
232->68 bucket gap = a large prefill win (the GEMM is ~51% of the total prefill
gap, so closing ~two-thirds of it is a meaningful S_PP improvement), but it
leaves a residual. **True 68 us/tok parity requires a native FP4-MMA large-M
kernel (option (b)) - the multi-week project** to greenlight only if Phase 1's
measured win proves the prefill regime matters enough to fund it.
- **Recommendation:** build Phase 1, measure, and let the measured dense S_PP
gain decide whether Phase 2 (MoE grouped cuBLAS) and ultimately (b) (native FP4
large-M kernel) are worth funding. Bank the cheap two-thirds before paying for
the kernel.
---
## 5. Summary table
| Option | Feasibility | Bit-exact | Effort | Verdict |
|---|---|---|---|---|
| (a) dequant->bf16 cuBLAS large-M | HIGH (parts exist) | new FP path, KL-gate (likely better PPL) | LOW-MED (days) | **RECOMMENDED** (dense first) |
| (b) Marlin/native FP4 large-M kernel | LOW | new FP path, KL-gate | HIGH (multi-week) | shelf - only route to true 68 us/tok |
| (c) M-threshold routing | HIGH | n/a (mechanism) | LOW | **the wiring for (a)** |
Decode is untouched by all of the above (threshold keeps M<=128 on FP4-MMQ); this
is a **prefill-only** lever.

View File

@@ -1,628 +0,0 @@
# Tensor-Core GDN Build Plan
> Auto-generated from the GDN build-design workflow. Build-ready spec for the full tensor-core chunked-scan kernel (2nd prefill lever).
## 1. Remaining intra-chunk products -> mma mapping
I have everything needed: the exact chunked-scan math from patch 0031, the sm_121a constraints from the scope doc, and the concrete ggml tf32 fragment (`mma.sync.aligned.m16n8k8.row.col.f32.tf32.tf32.f32`, `tile<16,8,float> D, tile<16,8,float> A, tile<8,8,float> B`) at `mma.cuh:976-984`. Here is the design.
---
# Tensor-core mapping of the REMAINING intra-chunk GDN products (patch 0031 steps 3-7)
## 0. Building block + what the PoC already covered
**Grounding.** Math: `backend/cpp/llama-cpp-localai-paged/patches/paged/0031-paged-chunked-gdn-prefill-scan-kernel.patch` (steps reproduced inline below). Scope/constraints: `backend/cpp/llama-cpp-localai-paged/docs/TENSORCORE_GDN_SCOPE.md`. Fragment API: `ggml/src/ggml-cuda/mma.cuh:976-984` (the only f32-accumulate tf32 overload on sm_121a).
The single warp-level primitive on sm_121a is **`m16n8k8` tf32 / f32-accumulate**:
- `A` fragment = `tile<16,8,float>` (M=16, K=8; 4 floats/thread, `Axi[0..3]`)
- `B` fragment = `tile<8,8,float>` (K=8, N=8, `.col` operand; 2 floats/thread, `Bxi[0..1]`)
- `D` accumulator = `tile<16,8,float>` (M=16, N=8; 4 floats/thread)
- A GEMM `[M×K]·[K×N]` tiles to `ceil(M/16) × ceil(N/8) × ceil(K/8)` mma calls, f32-accumulating over the K-subtiles.
- bf16 alternative `m16n8k16` (`mma.cuh:1064`, K=16/mma, 7-bit mantissa) exists but is **only** admissible for the tf32-safe Gram class — never the state/decay-coupled class.
- 3xtf32 ladder = split each f32 operand into 3 tf32 limbs, run 3 limb-products per K-subtile (hi·hi, hi·lo, lo·hi), accumulate high→low. ~3x the mma count, ~f32 accuracy.
**PoC covered products 1 + 2** (the two `C×C` Gram products, both tf32-safe, NMSE ~3e-9): `KK[t,t']=k_t·k_t'``A`, and `QK[t,t']=q_t·k_t'``P`. Both are `(C×dk)·(dk×C)`, M=C N=C K=dk=128, decay+beta applied in f32 after. They already share the `Kc^T` B-fragments.
The remaining families are **steps 3,4,5,6,7**. Notation: `C` = chunk (default 64; PoC 16), `dk=dv=128`, per `(head,seq)` block. Tile counts below are for **C=64**.
---
## 1. Per-product mma mapping table (the deliverable)
| # | Product (0031 step) | Result = matmul | M | N | K | mma tiles `(M/16)·(N/8)·(K/8)` @C=64 | Accumulation order | Precision class | Shares staged operand with |
|---|---|---|---|---|---|---|---|---|---|
| 1 | `KK→A` (PoC) | `Kc · Kcᵀ` | C | C | dk=128 | 4·8·16 = **512** (~½ tri) | over 16 k-subtiles | **tf32-safe** (proven) | `Kcᵀ` B-frag ↔ P2; `Kc` LHS ↔ P3 |
| 2 | `QK→P` (PoC) | `Qc · Kcᵀ` | C | C | dk=128 | 4·8·16 = **512** (~½ tri) | over 16 k-subtiles | **tf32-safe** (proven) | `Kcᵀ` B-frag ↔ P1; `Qc` LHS ↔ P4 |
| 3 | `KS = S0ᵀk_t` | `Kc · S0` | C | dv=128 | dk=128 | 4·16·16 = **1024** | 16 k-subtiles, limbs hi→lo | **3xtf32 / f32** (state-boundary, feeds solve) | `S0` B-frag ↔ P4; `Kc` LHS ↔ P1 |
| 4 | `QS = S0ᵀq_t` | `Qc · S0` | C | dv=128 | dk=128 | 4·16·16 = **1024** | 16 k-subtiles, limbs hi→lo | **3xtf32 → demote-first** (×γ_t≤1 attenuated) | `S0` B-frag ↔ P3; `Qc` LHS ↔ P2 |
| 5 | `O += P·U` | `P · U` | C | dv=128 | C=64 | 4·16·8 = **512** (~½ tri over K) | C/8 k-subtiles, triangular | **tf32-safe** (P decay-masked & bounded in f32 first) | `P`(=Amat) ↔ P2; `U` B-frag ↔ P6 |
| 6 | `S_C += Kᵀ(D·U)` | `Kcᵀ · DU` | dk=128 | dv=128 | C=64 | 8·16·8 = **1024** | scale state by γ_last (f32) **first**, then C/8 k-subtiles, limbs hi→lo | **3xtf32 / f32** (THE cross-chunk carry, compounds over n_tok/C) | `U` B-frag ↔ P5; `Kc` (transposed) ↔ P1/3 |
| 7 | `U = A⁻¹·RHS` off-diag coupling `A_ij·U_j` | `A_ij · U_j` | b=16 | dv=128 | b=16 | 1·16·2 = **32**/pair → **~192** (6 pairs) +~128 diag | forward sweep i=0..C/b; off-diag subtractions before diagonal solve | **tf32-safe off-diag + f32 in-register `16×16` diagonal** | `A`(=Amat) ↔ P1; `U` blocks ↔ P5/P6 |
3xtf32 inflation if the ladder is taken: P3 1024→**3072**, P4→**3072**, P6 1024→**3072**.
---
## 2. Per-product detail (the 5 remaining families)
### Product 3 - `KS = S0ᵀ k_t` (RHS state-boundary term)
0031: `ks = Σ_i Sd[j·dk+i]·Kc[t·dk+i]`; feeds `RHS[t][j] = β_t(v_t[j] γ_t·ks)`.
- **As a GEMM:** `KS[t][j] = Σ_i Kc[t][i]·S0[i][j]``KS = Kc[C×dk] · S0[dk×dv]`. **M=C, N=dv=128, K=dk=128.** Contraction over the state-row index `i`.
- **Schedule:** `Kc` is the LHS (M-major over `t`, K over `i`) — already staged for P1. `S0` is the B operand, K-major over `i`, N over `j`. The patch's `Sd[j·dk+i]` layout (i contiguous for fixed j) **is already a K-major B layout**`ldmatrix`-friendly as `tile<8,8>` B fragments. Accumulate 16 k-subtiles into f32 D.
- **Precision: 3xtf32/f32.** This is a state-boundary product: `S0` carries the full sequence history (wide dynamic range), and the result is *differenced* against `v_t` then fed into the solve, so error here propagates through `U` into both `O` and `S_C`. Default to the 3xtf32 ladder; A/B a plain-tf32 demote only after P4.
### Product 4 - `QS = S0ᵀ q_t` (γ cross-chunk `O` term)
0031: `qs = Σ_i Sd[j·dk+i]·Qc[t·dk+i]`; `o = γ_t·qs + Σ P·U`.
- **As a GEMM:** `QS = Qc[C×dk] · S0[dk×dv]`. **M=C, N=dv=128, K=dk=128** — identical shape to P3.
- **Schedule:** identical to P3 but LHS=`Qc` (shared with P2). **Fuse with P3 on the shared `S0` B-fragments:** stage `S0` once as B, run `Kc·S0` then `Qc·S0` back-to-back — `S0` is the heavy operand (128×128) and is loaded once for both.
- **Precision: 3xtf32 but the demote-first candidate.** The term is scaled by `γ_t ≤ 1` in f32 after the mma, so when the chunk has decayed (`γ_t→0`) the absolute error is attenuated. Least sensitive of the three state-boundary products; it is the first to try at plain tf32 in the precision A/B.
### Product 5 - `O += P · U` (attention-weighted output)
0031: `o += Amat[t·Cc+tp]·Ud[j·C+tp]` for `tp≤t`.
- **As a GEMM:** `O[C×dv] += P[C×C] · U[C×dv]`. **M=C, N=dv=128, K=C=64.** Contraction over the chunk index `t'`.
- **Schedule:** `P` (=Amat scratch from P2, with `d(t',t)` applied in f32) is LHS (M over t, K over t'); `U` (solved, in `Ud`) is the B operand, K-major over t'. `P` is **lower-triangular** ⇒ for M-tile `m` only K-subtiles `≤ m` are non-zero → ~½ the mma. Accumulate `C/8` k-subtiles. Add the `γ_t·QS` term (P4) into the same f32 D accumulator before write-out.
- **Precision: tf32-safe.** `P = d·QK` with `d≤1` is formed and bounded **in f32 first** (strong-decay rows already underflowed to ~0), so down-casting the bounded `P` to tf32 for this mma is benign. The decay is never inside the accumulation — it is pre-baked in f32, preserving the bounded de-gating invariant.
### Product 6 - `S_C += Kᵀ(D·U)` (the state update)
0031: `s = γ_last·Sd[j·dk+i] + Σ_t d(t,last)·Kc[t·dk+i]·Ud[j·C+t]`.
- **As a GEMM:** let `DU[t][j] = d(t,last)·U[t][j]` (D=diag applied in f32). `S_C[i][j] += Σ_t Kc[t][i]·DU[t][j]``S_C[dk×dv] += Kcᵀ[dk×C] · DU[C×dv]`. **M=dk=128, N=dv=128, K=C=64.** Contraction over the chunk index `t`.
- **Schedule:** the accumulator D fragments **are the register-resident state** that persists across the chunk loop. Order is strict: (i) scale the state fragments by `γ_last` in f32 in-register, **then** (ii) mma-accumulate `Kcᵀ·DU` over `C/8` k-subtiles into them. LHS = `Kc` read **transposed** (i as M-row, t as K) — a different fragment view of the same `Kc` smem buffer (use the `ldmatrix` transpose / J-major tile view). B = `DU` = `U` scaled by `d(t,last)` in f32, K-major over t — **same `U` B-layout as P5**.
- **Precision: 3xtf32 / f32 — the strongest ladder candidate.** This is the only product whose error *compounds across all `n_tokens/C` chunk steps*; it defines the state trajectory. Keep at 3xtf32 longest; this is the last product to ever consider demoting, and the place where a full-f32 accumulate (3xtf32) is most justified even if everything else passes plain tf32.
### Product 7 - the A-inverse (blocked forward substitution, FLA UT-transform)
0031 does a serial per-thread fwd-subst. Tensor-core form (block `b=16` = one mma M-tile, `C/b=4` blocks at C=64):
- For block `i`: `U_i = Ainv_ii·(RHS_i Σ_{j<i} A_ij·U_j)`.
- **The A-inverse-adjacent matmul = the off-diagonal coupling `A_ij·U_j`:** **M=b=16, N=dv=128, K=b=16**`1·16·2 = 32` mma/pair; 6 lower pairs at C=64 → **192** mma. Optional materialized-`Ainv_ii` apply is the same shape (~128 more).
- **Schedule:** forward sweep `i=0..3`; for each `i` accumulate all `j<i` couplings into a `b×dv` register tile (subtract from `RHS_i`), then apply the `b×b` diagonal inverse. `A`=Amat (from P1, β·d applied in f32) is the LHS; `U_j` blocks are read from `Ud` and updated in place as the sweep advances.
- **Precision: split.** Off-diagonal coupling = **tf32-safe** (`A_ij`=β·d·kk is bounded, `d≤1`; well-conditioned for the stable de-gating). The `16×16` **diagonal block inverse stays f32/in-register** (Neumann series on the b-nilpotent, ≤b1 terms, or a short serial solve) — exact, sensitive, but tiny. This is exactly the scope's recommended structure.
---
## 3. Staged-operand sharing graph (load amortization)
Five smem/register operands, and which products read them — the fusion that makes the added flops nearly free:
- **`Kc` (C×dk)** — the most-shared buffer. M-major-over-t LHS: P1, P3. K-major-over-i B (`Kcᵀ`): P1, P2. Transposed (i-major, contract t): P6. ⇒ stage once per chunk, feeds 1,2,3,6.
- **`Qc` (C×dk)** — LHS for P2 and P4.
- **`S0` B-fragments (dk×dv, register-resident state)** — P3 and P4. **Stage once as B, run KS then QS** (heaviest operand, amortized 2×).
- **`Amat` (C×C)** — P1 writes `A` → P7 reads `A` → P2 overwrites with `P` → P5 reads `P`. One buffer, lifecycle-reused (as 0031 already does).
- **`Ud` (C×dv)** — P7 writes `U` → P5 reads `U` (B, contract t) → P6 reads `U` scaled to `DU` (B, contract t). **P5 and P6 share the identical `U` B-layout** (both contract the chunk dim) → fully shared B-fragments.
Three concrete fusions worth coding as fused passes:
1. **P1+P2** share `Kcᵀ` B (PoC already does this).
2. **P3+P4** share `S0` B (stage the 128×128 state-as-B once).
3. **P5+P6** share `U` as B (both K=C contractions); compute `P·U` and `Kcᵀ·DU` from one `U` staging, P6 accumulating straight into the persistent state fragments.
---
## 4. tf32-safe vs 3xtf32 ladder - summary + recommended A/B order
**Plain-tf32-safe (well-conditioned, bounded, intra-chunk; bf16 `m16n8k16` is even an option if more throughput is needed):**
- P1 `KK`, P2 `QK` (PoC-proven), P5 `P·U` (P bounded/f32-pre-masked), P7 off-diagonal coupling.
**3xtf32 / f32 ladder (state-boundary, cross-chunk carry, error compounds):**
- P6 `Kᵀ(D·U)` — keep at 3xtf32 longest (compounds over every chunk).
- P3 `KS` — feeds the solve; 3xtf32 by default.
- P4 `QS` — 3xtf32 by default but γ_t-attenuated → **first to demote** to plain tf32 in the precision A/B.
- P7 `16×16` diagonal block inverse — stays **f32/in-register** (not a tensor-core op).
**Recommended precision A/B ladder (drives the KL-gate from `PAGED_BITEXACT_NOTE.md`):** start P3/P4/P6 at 3xtf32 and P1/P2/P5/P7-offdiag at plain tf32. If the KL-gate has margin, demote in order **P4 → P3**, holding **P6 at 3xtf32**. If even all-3xtf32 misses the KL-gate, the residual is the `16×16` diagonal solve precision, not the mma — that already stays f32.
---
## 5. Two honest implementation gotchas (not in the scope doc, surface in the mapping)
1. **Accumulator→B relayout of the state at each chunk boundary.** The register-resident state lives as P6's **D/accumulator** fragments (`tile<16,8>`), but P3/P4 need it as a **B operand** (`tile<8,8>`, K-major over `i`). These fragment layouts differ, so at chunk entry the state must be re-laid-out accumulator→B. Cheapest correct path: bounce the 128×128 state through a transient smem tile (write D fragments, `ldmatrix` back as B fragments) once per chunk — `n_tokens/C` times total, negligible vs the `C`x state-BW saved, but it means the "freed 64KB" needs a *transient* state-shaped smem tile at boundaries (not resident across the intra-chunk work). With dv-slabbing it's per-slab (`dk×dv_tile`).
2. **`Kc` needs two fragment views.** P1/P2/P3 read `Kc` contracting over `i` (dk); P6 reads `Kc` contracting over `t` (transposed, i as M-row). One smem buffer, but P6 must use the `ldmatrix` transpose / J-major `tile` view — budget for the transposed load, don't assume one staging serves both.
Ordering within a chunk is already correct in 0031 and must be preserved by the tensor-core version: P3,P4 read **pre-update** `S0` → P7 solve → P5 → **P6 overwrites** `S0``S_C`. Single accumulator, read-first/write-last, no state double-buffer needed.
## 2. A-inverse solve (form-T then apply, FLA UT transform)
I have everything I need: the chunked GDN math (patch 0031), the scope doc's recommendation, the ggml tf32 `mma.sync` API (`mma(tile<16,8,float>&D, tile<16,8,float>&A, tile<8,8,float>&B)` = m16n8k8), and the proven Gram PoC layout (`g=lane>>2, t=lane&3`; tf32 NMSE ~3e-9). Here is the design.
---
# A-inverse solve on `mma.sync` tensor cores (C=64, sm_121a) - design
Notation: `C=64`, head dim `dk=dv=128`, block size `b=16` (= one `m16n8k8` m-tile), `n_b=C/b=4`. `A = I + N`, `N = tril(beta_t·d(t',t)·(k_t·k_t'), -1)` strictly-lower (nilpotent, `N^C=0`); `RHS[t][j] = beta_t(v_t[j] - gamma_t(S0^T k_t)[j])` is `C×dv`; we want `U = A^{-1}·RHS`.
## 0. Core decision: form `T=A^{-1}` explicitly, then one wide apply (not direct back-subst)
Two routes were on the table. **Form `T = A^{-1}` in the `C×C` domain (FLA "UT transform"), then `U = T·RHS` as a single tf32 GEMM** - rather than blocked forward-substitution applied directly to the `C×dv` RHS. Reasons, all decisive on this part:
1. **Confines the only triangular dependency to the cheap `C×C` domain.** The expensive `dv=128`-wide work (`U=T·RHS`) becomes a dependency-free dense GEMM. The serial part is just the tiny `T`-formation. This is the single most important move for "don't serialize the warps."
2. **Fewer serial passes vs `dv`.** Inverting the `16×16` diagonal block once = a 16-column solve. Direct-solving against `RHS` re-solves against all `dv=128` columns per block. Form-`T`-once + reuse via mma is far cheaper in serial work.
3. **dv-slab reuse (the occupancy lever).** `T` depends only on `K`, not on `dv`. Form once, reuse for every `dv`-slab's `T·RHS_slab` apply. (Improvement over the scope's conservative "recompute per slab": when single-block, `T` lives in 16KB shared and is broadcast; only when dv-slabbing across separate blocks for occupancy do we recompute - which is cheap anyway, ~12% of the apply's mma count.)
4. **Isolates the error amplifier.** All recursion (the part that "amplifies error") lives in the small `T`-formation where 3xtf32 is nearly free; the bulk apply is a single well-conditioned GEMM.
This still **is** the scope's "blocked forward substitution: in-register diagonal solves + mma off-diagonal coupling" - just organized to produce `T` explicitly so the wide apply is dependency-free.
## 1. Solve algorithm
Block-partition `A` into a `4×4` lower-triangular grid of `16×16` blocks. `A_ii = I_b + N_ii` (unit-lower-tri, `N_ii` strictly-lower nilpotent); `A_ij` (i>j) full `16×16`. `T=A^{-1}` is block-lower-tri with:
```
T_ii = A_ii^{-1} (diagonal block inverse)
T_ij = -A_ii^{-1} · ( Σ_{m=j}^{i-1} A_im · T_mj ) for i > j (block fwd subst)
```
Then `U = T·RHS`, with `U_i = Σ_{j≤i} T_ij·RHS_j`.
**Phase D - diagonal inverses (4 blocks, fully parallel).** Each `A_ii` is `16×16` unit-lower-tri. Invert **exactly in f32** via shared-memory column-parallel forward substitution: stage `A_ii` to shared; thread `c` (c=0..15) solves `A_ii x = e_c` (`x_c=1`, `x_r = -Σ_{m=c}^{r-1} A_ii[r][m]·x_m`), writes column `c` of `T_ii`. 16 columns in parallel, ≤16 serial MACs each, all 4 blocks on 4 warps simultaneously. **No tensor cores here, and no reduced precision** - this is where the strongest coupling lives (see §4).
**Phase O - off-diagonal, mma.** For each i>j: accumulate `P_ij = Σ_m A_im·T_mj` (δ block-products), then `T_ij = -T_ii·P_ij`. All on `mma.sync` (`16×16×16` = `2 n-tiles × 2 k-steps` = 4 m16n8k8 per block-product).
**Apply.** `U = T·RHS`: warp `w` owns output rows `16w..16w+15`, sweeps all `dv=128` (16 n-tiles) × `C=64` (8 k-steps) = 128 m16n8k8/warp. This is the bulk and is embarrassingly parallel.
`A`, `P` (the QK term), `RHS`, and `T` are all assembled from tf32 Gram mma's (`KK`,`QK`,`KS`,`QS` - the PoC-proven step-1/2 plus step-3/4) with **all decay/`gamma`/`beta` applied in f32 outside the mma** (preserves bounded de-gating).
## 2. Tile schedule - keeping the triangular dependency off the warps
Block = 128 threads = 4 warps; **"warp == 16-row m-tile" throughout** (same mapping as the PoC's C=64 kernel, `rowbase = warp*16`, `g=lane>>2`, `t=lane&3`). Three layered mechanisms keep the warps busy despite the triangular dependency:
**(a) Wavefront (anti-diagonal) parallelism in `T`-formation.** The 6 off-diagonal blocks have a critical path of only `n_b-1=3`, not 6. Group by distance `δ=i-j`:
| Wave | Blocks (δ) | count | depends on | mapped to |
|---|---|---|---|---|
| D | (0,0)(1,1)(2,2)(3,3) | 4 | - | 4 warps ‖ |
| 1 | (1,0)(2,1)(3,2) | 3 | D | 3 warps ‖ |
| 2 | (2,0)(3,1) | 2 | D,1 | 2 warps ‖ |
| 3 | (3,0) | 1 | D,1,2 | 1 warp |
Within each wave all blocks are independent → one block per warp, no intra-wave serialization. Critical path = 4 dependency levels. Total `T`-formation mma: ~10 accumulation block-products + 6 inverse-applies = ~16 block-products × 4 = **~64 m16n8k8**, vs the apply's **512** (128/warp × 4) - so `T`-formation is ~12% of apply width and carries the only dependency.
**(b) Confinement.** Because we form `T` then apply, the dependency-laden work is the ~64-mma `C×C` formation; the 512-mma `dv`-wide apply has zero triangular dependency. The serial chain never touches the throughput-critical width.
**(c) Latency hiding via RHS overlap.** `T` depends only on `K` (→ `A` ← KK Gram). `RHS` depends on `V` and `S0^T k` (KS Gram, `dv`-wide, the expensive RHS term) and is **independent of the solve**. Schedule the wavefront `T`-formation (cheap, short critical path) concurrently with the `dv`-wide KS/QS Grams that build `RHS` and the `O` cross-term. The Phase-D shared scalar inverse (~16 shared round-trips × 4 warps) hides entirely under the KS Gram (thousands of cycles). By the time `T` is ready, `RHS` is staged and the apply fires immediately.
**Shared/register budget (C=64, state register-resident per the scope):**
| Buffer | bytes | note |
|---|---|---|
| `Kc`,`Qc` (bf16/tf32-staged) | 16KB+16KB | Gram operands |
| `A``T` scratch (`C×C` f32, in place) | 16KB | `A` consumed into `T`; reuses scope's A/P slot |
| `RHS`/`U` (`C×dv`) | 16-32KB | bf16 for the P·U and KᵀU mma's |
| diag-inverse scratch | ~1KB | `16×16` per warp, transient |
| gates `cs/gam/beta` | <1KB | f32 |
| state `S` | 0 (registers) | frees the 64KB that forced 0031's C=16 |
Total ~65-80KB, under the 99KB opt-in - the solve adds **no** net shared pressure (T overwrites A; diag scratch is transient). Per-thread diag-inverse needs ~16 regs (one column of `x`), released before the apply - does not compound the already-heavy state-accumulator register budget.
## 3. Precision risk assessment
**Error model.** `‖ΔU‖/‖U‖ ≲ κ(A)·(‖ΔA‖/‖A‖ + ‖ΔRHS‖/‖RHS‖) + ‖Δ_apply‖/‖U‖`. The inverse is the amplifier; `κ(A)` is data-dependent. For DeltaNet, keys are L2-normalized so `|k_t·k_t'|≤1``|N[t][t']|≤beta_t≤1`; in the decaying regime `‖N‖<1` and `κ` is modest, but in the weak-decay/aligned-keys corner `κ` grows and the `δ=3` column path (`T_30`) compounds 3 multiplies. tf32 input rounding is ~`2⁻¹¹``5e-4` relative (f32-accumulate; PoC measured Gram NMSE ~`3e-9`). 3xtf32 (3-limb split, the CUTLASS fp32-emulation trick) buys ~f32 (~`1e-7`) at ~3× that step's mma cost.
**Where the strong coupling actually sits (the key structural fact):** the *inverse* `T_ii` is computed f32-exact, **but the dominant near-diagonal mixing is applied in the tf32 apply GEMM** (`U_i ⊃ T_ii·RHS_i`), and block-boundary adjacent pairs (e.g. tokens 15↔16) live in the `δ=1` off-diagonal `T_10`. So "f32 protects the strong coupling" is only true for the inverse *computation*; its *application* is tf32 unless promoted. This drives the ladder.
**Precision config + 3xtf32 ladder (mandatory vs optional):**
| Step | Default | Mandatory? | 3xtf32 cost |
|---|---|---|---|
| Diagonal inverse `T_ii` | **f32 (shared scalar)** | **Mandatory-and-free** (it's already f32) | n/a |
| Off-diag coupling `A_im·T_mj`, `T_ii·P_ij` | **3xtf32 (default-on)** | Effectively mandatory; ~3× of ~64 tiny mma = **negligible** | free insurance |
| KK/QK Gram → A,P | tf32 | optional (rung 1) | 3× of C×C Grams (cheap) |
| Apply `U=T·RHS` | tf32 | optional (rungs 2-4) | up to 3× the bulk |
| KS/QS Gram → RHS, O | tf32 | optional (rung 5) | vLLM keeps these bf16 (L4-rejected precedent) |
Decays/`gamma`/`beta` **always f32, outside the mma** - invariant, not a rung.
**Ladder ordering if the default config misses the KL-gate (cheapest → most expensive):**
1. KK Gram (feeds `A`) → 3xtf32 [cheap, C×C].
2. Apply **block-diagonal terms only** `T_ii·RHS_i` → 3xtf32 [≈+0.8× apply; protects within-window strong coupling - mixed-precision-by-distance].
3. + apply `δ=1` off-diagonal terms → 3xtf32 [covers block-boundary adjacent pairs].
4. Full apply → 3xtf32 [≈+2× apply; expensive escape hatch].
5. KS/QS Gram → 3xtf32.
6. Fall back to direct blocked back-substitution against RHS in 3xtf32 (the alternative route, slightly more accurate than form-`T`-then-multiply at the cost of the parallelism), else keep 0031's serial path.
**Adversarial `g∈[-20,-1e-4]`:** strong decay ⇒ `d=exp(big-negative)→0` ⇒ off-diagonal `N→0``A≈I`, `T≈I`, apply≈identity, tf32 error vanishes; bounded de-gating (f32) guarantees underflow-to-zero, never inf. Weak decay (`g→0`) ⇒ `d≈1`, `A` well-conditioned, tf32's 8-bit exponent (vs f16's 5) holds the `gamma` dynamic range. The dangerous middle is the only KL-empirical risk - re-run this op case explicitly per the scope.
**KL impact / gating.** Same gate as the backend's new-FP-path precedents: NMSE is expected to *fail* at reduced precision (this is a new path on a new path) - **the binding gate is KL** (`KLD(tc‖f16) ≤ KLD(seq‖f16)` + PPL band) plus greedy-md5 stability (md5 will not match 0031's serial path - per-path, validated benign). Expectation: the **default config (f32 diagonal + 3xtf32 off-diagonal-coupling + tf32 everything-else)** clears the KL-gate, because (i) the dominant apply matches the PoC Gram's `~3e-9`/tf32-input grade and (ii) the recursion-amplified `C×C` work is f32-grade for free. The expensive apply-3xtf32 rungs are reserved escapes. Worst case all-3xtf32 ≈ 3× the mma cost - still an order of magnitude under 0031's serial-f32 reductions and still net-positive given the `~C×` state-BW cut.
## 4. Integration + validation
- Build on `ggml/src/ggml-cuda/mma.cuh`: the tf32 path is `mma(tile<16,8,float>&D, tile<16,8,float>&A, tile<8,8,float>&B)``mma.sync.aligned.m16n8k8.row.col.f32.tf32.tf32.f32` (line ~1089), gated by `AMPERE_MMA_AVAILABLE` (sm_121-correct). tf32 operands stage to shared and load via `load_generic` (or the PoC's `cvt.rna.tf32.f32` register packing); `ldmatrix` is `.b16`-only so it is **not** usable for tf32 fragments - use `load_generic`. 3xtf32 = 3 limb-split tf32 mma's + f32 recombine, reusing the same tile type.
- Slots into patch 0031's `gated_delta_net_chunked_cuda`: replace the serial `KK`/`QK` Grams (lines 183-198, 222-235) with the PoC tiles; replace the serial forward-subst (lines 211-218) with this Phase-D + Phase-O `T`-formation; replace the serial `O`/`S_C` reductions (steps 5/6) with the apply + `KᵀU` mma. `C=16``C=64`, state register-resident.
- Validation step before the full kernel: extend `~/scratch_tc_gdn_poc/gdn_gram_bench.cu` (DGX) to microbench Phase-D+O `T`-formation and the `T·RHS` apply in isolation, with the f64 host oracle measuring `κ(A)` and the tf32-vs-3xtf32 NMSE per rung - mirrors the Gram PoC that proved 6.7×→9.3×. This tells you, cheaply, whether the default precision config will clear the KL-gate before committing to the multi-week kernel (Phase 2 of the scope's plan).
**Net:** the solve adds zero shared pressure, ~12% mma overhead over the apply, a 4-level (not 64-step) critical path fully hidden under the RHS Grams, an f32-exact inverse, and a cheap-by-default 3xtf32 placement aimed exactly at the recursion-amplified term - so the A-inverse is **not** the bottleneck and **not** the precision wall; the residual risk is the data-dependent `κ(A)` in the weak-decay corner, arbitrated by the KL-gate with the apply-3xtf32 rungs as the escape.
Files referenced (absolute): `/home/mudler/_git/LocalAI/.claude/worktrees/feat+paged-attention/backend/cpp/llama-cpp-localai-paged/docs/TENSORCORE_GDN_SCOPE.md`; `/home/mudler/_git/LocalAI/.claude/worktrees/feat+paged-attention/backend/cpp/llama-cpp-localai-paged/patches/paged/0031-paged-chunked-gdn-prefill-scan-kernel.patch`; ggml tf32 mma at `ggml/src/ggml-cuda/mma.cuh` (line ~1089); DGX PoC `~/scratch_tc_gdn_poc/gdn_gram_bench.cu`.
## 3. Register-resident state + C=64 budget
I have everything I need: the scope doc, the full 0031 kernel (smem layout, the 7 product families, the chunk-loop sequencing), and the sibling FP4-MMA GEMM kernel's concrete patterns (PAD=4 u32 rows, STAGES=2 cp.async, WM=4/WN=2 8-warp grid, ldmatrix, register-resident accumulators). Here is the concrete data-layout design.
---
# Register-resident state + C=64 design (tensor-core chunked GDN, sm_121a / mma.sync)
## 0. Block/warp configuration (the frame everything hangs on)
0031 used `block_dims(128,1,1)` = 128 threads, "thread j owns v-column j" - a **column-parallel scalar** model. The tensor-core kernel must abandon that and adopt a **warp-tiled** model (same as the sibling GEMM kernel):
- **256 threads = 8 warps**, arranged as a **WARPS_M x WARPS_N = 4 x 2** warp grid (the GEMM kernel's proven `WM=4, WN=2`).
- Threads no longer own columns; warp `(wm,wn)` owns a rectangular sub-tile of each matrix and drives `mma.sync` on it.
- Precision: **tf32 m16n8k8** for the S-coupled / decay-coupled products, **bf16 m16n8k16** allowed only for the well-conditioned intra-chunk Gram terms (KK, QK). f32 accumulate throughout. Decays/`gamma`/`beta` stay f32, applied outside the mma (preserve bounded de-gating).
This 4x2 warp grid is the denominator for every ownership calc below.
---
## 1. The one hard problem: S is an *accumulator* for step 6 but an *operand* for steps 3/4
This is the crux the scope hand-waves ("read S as the stationary operand; step 6 accumulates into it"). The register fragment layouts are **not** interchangeable:
| Use | Role | mma shape | S indexing | Fragment layout |
|---|---|---|---|---|
| Step 6 `S += Kᵀ(D·U)` | **accumulator (D/C)** | m=dk, n=dv, k=C | `S[i][j]`, m=i, n=j | `tile<16,8,float>` acc grid |
| Step 3 `KS = K·S` | **B operand** | m=C, n=dv, k=dk | `S[i][j]`, k=i, n=j | `tile<8,8,float>` B frag |
| Step 4 `QS = Q·S` | **B operand** | m=C, n=dv, k=dk | same as step 3 | `tile<8,8,float>` B frag |
An accumulator fragment's thread→element map differs from a B-operand fragment's, so **you cannot feed the persistent S registers directly into the step-3/4 mma.** A bridge is mandatory. The design decision:
> **S lives register-resident in the step-6 ACCUMULATOR layout** (it is written every chunk; that is the hot path). Steps 3/4 reach it via a **once-per-chunk restage to a small smem tile, re-read with `ldmatrix`** as B-operand fragments.
The restage cost is paid `n_tokens/C` times (not per token) - it is *inside* the BW saving the whole lever buys. And critically, the restage smem **time-multiplexes onto the Uc/Amat region**: at chunk entry (when KS/QS are needed) U and A for this chunk are not yet computed, so their buffers are free to hold the S restage. **Net additional persistent smem for the state: 0KB** - the scope's "0KB shared state" holds, with this scheduling caveat made explicit.
KS and QS read the **same** pre-update S0, so restage once → do both → then overwrite with U.
---
## 2. Register allocation map (per thread, 256-thread block)
State `S` is `dk x dv = 128 x 128` f32 = 16384 elems. Distributed over 256 threads = **64 f32/thread** at full dv.
| Register class | Lifetime | Full dv=128 | dv-slab=64 | dv-slab=32 | Layout / ownership |
|---|---|---|---|---|---|
| **Persistent S accumulator** | whole chunk loop | **64 regs** | **32 regs** | **16 regs** | warp `(wm,wn)` owns dk-rows `[wm·32, +32)` x dv-cols `[wn·(dv/2), +dv/2)`; = 2 m-tiles x (dv/2/8) n-tiles of `tile<16,8,float>`, 4 f32 each |
| Transient A-operand frag | per product | 4 regs/tile | same | same | `tile<16,8,float>` (tf32 packs k8) reused across KK/QK/KS/QS/O/Supd |
| Transient B-operand frag | per product | 2 regs/tile | same | same | `tile<8,8,float>` |
| Transient product accumulator (KK/QK/KS/QS/O) | per product, then spilled to smem | ≤8 tiles·4 = 32 regs | ≤16 | ≤8 | these outputs go to smem; acc is transient, reused |
| A⁻¹ diagonal-block solve (16x16, in-registers) | step 7 only | ~8-12 regs | same | same | one `b=16` unit-lower-tri block per warp-row, scalar Neumann/`<b-1` terms |
| loop/index/gate scalars | always | ~12 regs | same | same | c0, Cc, cs/gam/beta locals |
**Per-thread totals (256 threads):**
- Full dv: 64 (S) + ~50 (transients, non-overlapping with S) + ~12 ≈ **~130 regs/thread** → fits **1 block/SM** (256 regs/thread budget at 65536/SM÷256). 2 blocks/SM (128 regs/thread cap) would spill - hence dv-slab for occupancy.
- dv-slab 64: 32 (S) + ~50 + 12 ≈ **~94 regs/thread** → fits **2 blocks/SM** (128-reg cap). ✓
- dv-slab 32: 16 (S) → ~78 regs/thread → headroom; grid x4.
Persistent-state register pressure is the occupancy gate; everything else is transient and reused across the 7 products.
---
## 3. Shared-memory allocation map (PAD-padded, conflict-free)
Apply the GEMM kernel's lesson verbatim: **PAD = 4 (in the row's element width)** so a 128-wide row (a multiple of the 32 banks → 8-way conflict for the 8-row `ldmatrix`) becomes stride `132`, and the 8 rows of an `ldmatrix.m8n8` land in 8 distinct banks: `(r·132 + c) mod 32 = (4r + c) mod 32`, distinct for `r=0..7`. ✓
| Buffer | Logical shape | Row stride (padded) | Element | Notes |
|---|---|---|---|---|
| `Kc` (chunk K) | `[C][dk]` | `dk + 8` bf16 (= +4 u32) | bf16 | A-operand for KK/QK; A/Bᵀ for KS; transposed-A for Supd. bf16 default |
| `Qc` (chunk Q) | `[C][dk]` | `dk + 8` bf16 | bf16 | A-operand for QK/QS/O |
| `Uc` (solved U) | `[dv][C]` | `C + 4` f32 | f32 | f32 for the triangular solve accuracy; B-operand (down-cast tf32) for O & Supd |
| `Amat` (A then P) | `[C][C]` | `C + 4` f32 | f32 | KK→A→solve, then reused for QK→P; decays applied in f32 here |
| `gates` cs/gam/beta | `[3·C]` | (1-D, no pad) | f32 | prefix-sum + `expf`, f32 always |
| **S-restage tile** | `[dk][dv_strip]` | `dv_strip + 4` f32 | f32 | **overlays UcAmat** at chunk entry; not additive at peak |
| `cp.async` stage dup | (STAGES=2 on Kc/Qc) | as Kc/Qc | bf16 | Phase-3 latency hiding only |
PAD widths: f32 tiles +4 elems; bf16 tiles +8 elems (= +4 u32, identical bank offset as the GEMM kernel). `Uc` and `Amat` are padded on the C dimension (their `ldmatrix` access dimension).
---
## 4. C=64 shared budget table (under the 99KB opt-in)
Byte math with PAD included (`KB = bytes/1024`):
| Buffer | CONFIG A — **default**: C=64, dv=128, K/Q **bf16** | CONFIG B: C=64, dv=128, K/Q **tf32** | CONFIG C — **2 blk/SM**: C=32, dv-slab=64, K/Q bf16 |
|---|---|---|---|
| `Kc` | 64·136·2 = **17.0KB** | 64·132·4 = 33.0KB | 32·136·2 = 8.5KB |
| `Qc` | **17.0KB** | 33.0KB | 8.5KB |
| `Uc` (f32) | 128·68·4 = **34.0KB** | 34.0KB | 64·36·4 = 9.0KB |
| `Amat` (f32) | 64·68·4 = **17.0KB** | 17.0KB | 32·36·4 = 4.5KB |
| gates | **0.75KB** | 0.75KB | 0.4KB |
| S-restage | overlay (0 net) | overlay (0 net) | overlay (0 net) |
| **Per-block total** | **≈ 85.8KB** ✅ < 99 | **≈ 117.8KB** ❌ | **≈ 30.9KB** |
| Blocks/SM (≈100KB/SM) | **1** | n/a | **2** (61.8KB) ✅ |
Read-out:
- **CONFIG A is the recommended default**: C=64 (4x the 0031 chunk), full dv, fits at ~86KB with margin, 1 block/SM. Peak is the O/Supd phase (all of Kc+Qc+Uc+Amat live).
- **CONFIG B (tf32 K/Q) is budget-hostile** (117KB) - tf32 K/Q tiles don't shrink with dv-slab (they're `C x dk`), so even dv-slab=64 lands ~101KB. **Conclusion: stage K/Q as bf16; reserve tf32/3xtf32 for the S-coupled and decay-coupled terms** (which arrive via the small streamed S-restage and the f32 gate scaling), exactly per the scope's "bf16 only for well-conditioned Gram terms."
- **CONFIG C is the 2-block/SM lever**: C=32 + dv-slab=64 → 31KB/block, two resident blocks under the ~100KB/SM total, and the grid grows to `H x n_seqs x 2`.
---
## 5. dv-slab strategy (the 2nd block/SM + grid-starvation fix)
Split the `dv=128` value dimension into `n_slabs` blocks; each block computes a `dv_tile`-wide vertical strip of O and of the state.
- **Grid**: `dim3(H, n_seqs, n_slabs)` (was `(H, n_seqs)`). `n_slabs ∈ {1,2,4}` for `dv_tile ∈ {128,64,32}`. This **multiplies the grid by `n_slabs`**, directly attacking 0031's low-`n_seqs` grid starvation.
- **What is dv-independent and must be recomputed per slab**: `A` (KK Gram), the `A⁻¹` solve, the gate prefix - all depend only on K and the gates, *not* dv. Each slab recomputes them. Cheap once they are on tensor cores (the whole point); this is the FLA per-slab pattern.
- **What is dv-sliced**: the S accumulator (`128 x dv_tile`), `Uc` (`dv_tile x C`), KS/QS/O outputs, step-6 update. Halving/quartering dv halves/quarters both the **S register footprint** (64→32→16 regs/thread, §2) and the dv-scaled smem (`Uc`, restage).
- **Restage budget bonus**: at `dv_tile=64` the per-block S is `128 x 64` = 32KB, so the once-per-chunk restage fits the UcAmat overlay window in a single pass (no strip loop). At full dv=128 the restage is done as **2 dv-strips of 32KB** reusing the same overlay (or 16 k-strips of 8x128 if registers are tighter than smem).
`b`-block forward substitution (step 7) is independent of dv too, so the in-register `16x16` diagonal solves are computed once and the off-diagonal mma coupling `Uᵢ -= Aᵢⱼ Uⱼ` runs per slab as a `(16x16)·(16 x dv_tile)` mma.
---
## 6. Bank-conflict-free layout - the GEMM PAD lesson applied
Concretely, per the sibling kernel's `ARS = KBLK·SAW + PAD` with `PAD=4` (which gave +19%):
- Every smem matrix read by `ldmatrix` (or its tf32 equivalent in `ggml/src/ggml-cuda/mma.cuh`) is stored with **row stride = logical_width + PAD**, PAD chosen so `stride mod 32 ≠ 0`: f32 width-128 → 132 (`132 mod 32 = 4`); bf16 width-128 (packed 64 u32) → 68 u32 (`68 mod 32 = 4`).
- This guarantees the 8 rows an `m8n8` `ldmatrix` touches map to 8 distinct banks for any fixed column → no replays on the operand loads, which are the kernel's inner-loop smem traffic.
- `cp.async` (CONFIG, Phase 3): `STAGES=2` double-buffer on `Kc`/`Qc` only (the GEMM kernel found multistage saturates BW past depth 2). 16B `cp.async.cg` copies into the padded rows; `commit_group`/`wait_group` Ampere-style (no TMA on sm_121). The pad keeps the staged writes coalesced and the mma reads conflict-free simultaneously.
---
## 7. Summary of the allocation decisions
| Decision | Value |
|---|---|
| Threads / warp grid | 256 / 4x2 (WM=4, WN=2) |
| **State residency** | register-resident in **step-6 accumulator layout** (`tile<16,8,float>` grid), 64/32/16 f32-regs/thread at dv 128/64/32 |
| **Accumulator↔operand bridge** | once-per-chunk `ldmatrix` restage of S to a small smem tile that **overlays UcAmat** (0 net persistent smem); KS+QS share one restage |
| K/Q precision | **bf16** staged (tf32 K/Q breaks the 99KB budget); tf32/f32 reserved for S-coupled + decay-coupled terms |
| Uc / Amat | f32, padded on C (+4) |
| **PAD** | +4 f32 (+8 bf16 = +4 u32) row stride → `ldmatrix` 8-row conflict-free (GEMM-proven) |
| **C=64 default budget** | **≈86KB**, 1 block/SM (CONFIG A) ✅ |
| 2 blocks/SM | C=32 + dv-slab=64 → ≈31KB/block, grid x2 (CONFIG C) |
| dv-slab | grid `(H, n_seqs, n_slabs)`; A/A⁻¹/gates recomputed per slab; S/Uc/O dv-sliced |
| cp.async | STAGES=2 on Kc/Qc (Phase 3 only) |
One honest caveat surfaced beyond the scope doc: the scope's "~64-80KB / 0KB shared state" budget only holds with **bf16 K/Q staging** and the **overlay restage**; tf32 K/Q (CONFIG B) does not fit even with dv-slab, which is why bf16-Gram + tf32-only-for-S/decay is the forced (and scope-consistent) precision split.
Files referenced: `/home/mudler/_git/LocalAI/.claude/worktrees/feat+paged-attention/backend/cpp/llama-cpp-localai-paged/docs/TENSORCORE_GDN_SCOPE.md`, `/home/mudler/_git/LocalAI/.claude/worktrees/feat+paged-attention/backend/cpp/llama-cpp-localai-paged/patches/paged/0031-paged-chunked-gdn-prefill-scan-kernel.patch`, and the sibling GEMM patterns in `/home/mudler/_git/LocalAI/.claude/worktrees/feat+paged-attention/backend/cpp/llama-cpp-localai-paged/patches/paged/0034-feat-paged-native-NVFP4-W4A4-FP4-MMA-large-M-prefill.patch`.
## 4. Occupancy + launch config
I have everything I need. Key facts confirmed: tf32 tile is `mma.sync.aligned.m16n8k8.f32.tf32.tf32.f32` (mma.cuh:1089, accumulator `tile<16,8,float>` = 4 f32/thread/tile); grid.x = `H` = `n_v_heads` = `ssm_dt_rank` (llama-model.cpp:504; Qwen3-Next family = 32 GDN value heads); the sequential kernel already uses a 3D grid `(H, n_seqs, ceil(S_v/num_warps))` (gated_delta_net.cu:184) — the chunked 0031 collapsed that z-axis to 1, which is exactly what starves the grid.
Here is the occupancy + launch design.
---
# Occupancy + launch design — tensor-core chunked GDN prefill (sm_121a)
## 0. The two independent caps 0031 hit (must relieve BOTH for ≥2 blocks/SM)
0031's -22% is not one wall, it's two stacked walls, and they are relieved by *different* levers:
| Cap | 0031 value | Binding resource | Lever |
|---|---|---|---|
| **Shared-memory cap** | 89 KB (64 KB all-shared state) | 100 KB/SM, 99 KB dyn opt-in | state→registers **+ smaller C** |
| **Register cap** | n/a (was scalar) | 65536 regs/SM | **dv-slab** the register-resident state |
| **Grid cap** | `(H, n_seqs, 1)` = 32·n_seqs blocks | 48 SMs | **dv-slab multiplies grid** by n_slabs |
sm_120/121-class per-SM limits used throughout: **1536 threads/SM, 65536 32-bit regs/SM, 100 KB shared/SM (99 KB dynamic opt-in), 255 regs/thread, ≤24-32 blocks/SM (hw, never the binding limit here).** The binding limits are **shared and registers.**
Critical correction to the scope-doc budget table: it assumes **bf16** K/Q staging (2 B). The precision default is **tf32**, which is a 32-bit container in shared — tf32 K/Q would *double* Kc/Qc and blow C=64 past 99 KB. So the occupancy config **stages K/Q as bf16** (the well-conditioned KK/QK Gram products per the scope's "bf16 only for Gram terms"), keeps gates/decays/beta/the solve in f32. This is a real precision↔occupancy coupling, flagged in §5.
## 1. Grid mapping — three parallel axes, the chunk axis is serial
The inter-chunk recurrence carries state `S` across chunks, so **the chunk axis cannot be a grid axis** (it's the sequential dependency — that's the whole algorithm). The only legitimate grid axes that don't break the recurrence are:
```
dim3 grid(H, n_seqs, n_slabs); // H = n_v_heads = 32 (ssm_dt_rank)
// n_slabs = dv / dv_tile (the new lever)
```
- `blockIdx.x = head` (0..31), `blockIdx.y = seq`, `blockIdx.z = dv-slab`.
- A block owns v-columns `[z·dv_tile, (z+1)·dv_tile)`, walks the chunk loop serially, and keeps **only its `dk × dv_tile` state slab** register-resident.
- This reuses the **same 3D grid shape the sequential kernel already has** (gated_delta_net.cu:184 uses z for S_v-splitting); the chunked kernel repurposes z from S_v-split to dv-slab. The dispatcher change is minimal.
**Saturation math (the core of the task).** Target ≥2 blocks/SM on 48 SMs ⇒ **≥96 concurrent blocks**. With H=32:
| n_seqs | dv_tile=128 (n_slabs=1) | dv_tile=64 (2) | dv_tile=32 (4) |
|---|---|---|---|
| 1 | 32 (starved, 0031) | 64 (48/48 SMs busy, 67% warp-occ) | **128 (100%)** |
| 2 | 64 | **128 (100%)** | 256 |
| 4 | 128 | 256 | 512 |
So **dv-slabbing is simultaneously the register-relief lever and the grid-multiplier** — it's the single most important move. Rejected grid alternatives: split-K over dk (needs cross-block atomic reduction + fights the state carry); batching heads/seqs per block (reduces grid, wrong direction).
## 2. Block dim / warp count — 8 warps / 256 threads
```
constexpr int WARPS = 8;
dim3 block(32 * WARPS, 1, 1); // 256 threads
```
Why 8 warps:
- **Clean mma tile partition at C=32:** KK/QK output is `C×C = 32×32` = (32/16)·(32/8) = **8 m16n8 tiles → exactly 1 tile/warp**, dk=128 = 16 k8-steps. Steps 3/4 (KS/QS) and 5 (P·U) → 2 tiles/warp. Step 6 state update `dk×dv_tile`=128×64 → 64 tiles → **8 tiles/warp** (these are the persistent register-resident accumulators).
- **Register dilution:** the register-resident state accumulator is spread across all 256 threads (see §3) — more warps = fewer state-regs/thread.
- **Threads are not the cap:** 256 threads ⇒ up to 6 blocks/SM by the 1536 thread limit, so registers/shared decide.
Fallback if register-capped (§5): **12 warps (384 threads)** dilutes the state accumulator further (dv_tile=64: 32→21 state-regs/thread) at the cost of thinner per-warp tiles and ≤4 blocks/SM by threads.
## 3. Register-resident state ↔ dv-slab ↔ occupancy interaction
The state slab is held as **tf32 mma accumulator fragments** (`tile<16,8,float>`, 4 f32/thread/tile) persisting across the chunk loop. Per-thread state-register cost = `dk·dv_tile / 256`:
| dv_tile | state f32/block | state regs/thread (256 thr) | + working (est.) | regs/thread | regs/block | reg-allowed blocks/SM |
|---|---|---|---|---|---|---|
| 128 (no slab) | 16384 | 64 | ~50 | ~114 | ~29 K | 2 (tight) |
| 64 | 8192 | 32 | ~50 | ~82 | ~21 K | 3 |
| 32 | 4096 | 16 | ~50 | ~66 | ~17 K | 3 |
So on registers alone, dv_tile≤64 admits ≥2 blocks/SM. **Shared memory is then the binding cap**, and it's governed by **C**, not dv_tile (Kc/Qc/A all scale with C, only U scales with dv_tile):
| Config | Kc+Qc (bf16) | A/P (f32) | U (f32) | single | +cp.async dbl-buf K/Q | blocks/SM (shared) |
|---|---|---|---|---|---|---|
| C=64, dv_tile=128 | 32 KB | 16 KB | 32 KB | 80 KB | 112 KB ✗(no room!) | **1** |
| C=64, dv_tile=64 | 32 KB | 16 KB | 16 KB | 64 KB | 96 KB ✓ | **1** |
| **C=32, dv_tile=64** | 16 KB | 4 KB | 8 KB | **28 KB** | **44 KB ✓** | **2** |
| C=32, dv_tile=32 | 16 KB | 4 KB | 4 KB | 24 KB | 40 KB ✓ | **2** |
**Finding the scope doc missed:** C=64-no-slab is shared-saturated at 80 KB — there is **no room for cp.async double-buffering**, so the 1-block/SM kernel would have *no latency hiding* and likely still lose. C=64 needs dv_tile≤64 *just to make room for cp.async*, and is still 1 block/SM. **Genuine 2 blocks/SM requires C=32** (to drop Kc/Qc/A under the 49.5 KB/block budget).
## 4. cp.async double-buffering (depth 2, no TMA)
At 1 block/SM (C=64 path) cp.async is the *only* latency-hiding mechanism, so it's mandatory, not optional. Plain Ampere `cp.async` (`cp.async.commit_group` / `cp.async.wait_group`) — **no `cp.async.bulk`/TMA on sm_121.** Stage the **next chunk's Kc, Qc** (and Vc if the KL-gate doesn't force V from global) into a second shared buffer while the current chunk's mma runs. Depth **2 only** — the sibling GEMM kernel proved multistage saturates BW past depth 2. The double-buffer cost is already in the "+cp.async" column above (44 KB at C=32 keeps 2 blocks/SM).
## 5. Launch config (concrete) + honest occupancy estimate
**Recommended default (batched-prefill serving regime, n_seqs≥2):**
```
C = 32 ; dv_tile = 64 ; n_slabs = 2 ; WARPS = 8
grid = dim3(H=32, n_seqs, 2)
block = dim3(256, 1, 1)
smem = 44 KB (Kc/Qc bf16 ×2 dbl-buf + A/P f32 + U f32) // cudaFuncSetAttribute return CHECKED (0031 precedent)
→ 2 blocks/SM. n_seqs≥2 ⇒ ≥128 blocks ⇒ 48/48 SMs at full 2-block occupancy (100%), 1.33 waves.
A/Gram/solve recomputed 2× across slabs (state-update per slab is 2× the A work ⇒ ~25% redundant-flop overhead).
```
**Single-stream prefill (n_seqs=1) saturator:**
```
C = 32 ; dv_tile = 32 ; n_slabs = 4 ; WARPS = 8
grid = dim3(32, 1, 4) = 128 blocks ⇒ 2 blocks/SM on all 48 SMs (100%) even at n_seqs=1.
Cost: A recomputed 4×, and at dv_tile=32 the A bucket ≈ the per-slab state bucket ⇒ ~1.5-2× total-flop overhead.
```
**BW-max alternative (1 block/SM, bench against the above):**
```
C = 64 ; dv_tile = 64 ; n_slabs = 2 ; WARPS = 8 ; smem = 96 KB (dbl-buf, fits 99 KB)
→ 1 block/SM, but 4× state-BW cut (vs 2× at C=32) + grid ×2. n_seqs=1 ⇒ 64 blocks ⇒ 48/48 SMs busy (67% warp-occ).
```
**Occupancy summary:**
| Config | blocks/SM | regs/thread | smem/block | SM util @ n_seqs=1 | SM util @ n_seqs≥2 | state-BW cut | redundant-A |
|---|---|---|---|---|---|---|---|
| 0031 | 1 | scalar | 89 KB | 32/48 busy (starved) | 1024 blk, no overlap | 1× (C=16) | none |
| C=32 dv64 (default) | **2** | ~82 | 44 KB | 48 busy, 67% occ | **100%** | 2× | 2× (~25%) |
| C=32 dv32 (1-seq) | **2** | ~66 | 40 KB | **100%** | 100% | 2× | 4× (~1.5-2×) |
| C=64 dv64 (BW-max) | 1 | ~114 | 96 KB | 48 busy, 67% occ | 100%, multi-wave | **4×** | 2× |
The C=32 (occupancy) vs C=64 (BW) choice is the empirical fork the scope doc defers to Phase-3 bench: 2 blocks/SM at half the BW saving, vs 1 block/SM at full BW saving + cp.async. **Wire both behind the existing `GDN_CHUNK_MIN` gate plus a `GDN_CHUNK_C` / `GDN_DV_TILE` selector and A/B them; do not assume.**
## 6. Residual risk — register pressure likely caps it at 1 block/SM (honest)
The ≥2-blocks/SM result rests on the **~50 working-regs/thread estimate**, which is optimistic:
- **The blocked-forward-subst A⁻¹ (step 7) is the swing factor.** The in-register 16×16 unit-lower-triangular diagonal inverse + the off-diagonal mma coupling + mma operand fragments + the **accumulator→operand fragment transpose** for reusing the register-resident S as a step-3/4 operand (a `movmatrix`/shared round-trip, since S lives in C-fragment layout but steps 3/4 need it as an A/B operand) can push working regs to **80-120**. At 256 threads, regs/thread > 128 ⇒ > 32 K regs/block ⇒ **silently drops to 1 block/SM** regardless of the 44 KB shared headroom. The scope doc names this exactly: "blocked-forward-subst register pressure trades against state-register pressure; both compete for the same budget."
- **Mitigation ladder, in order:** (i) 12 warps to dilute (dv_tile=64: state 32→21 regs/thread); (ii) `__launch_bounds__(256, 2)` to force the compiler under 128 regs/thread (risks spills to local → BW back); (iii) smaller dv_tile (more grid, more redundant A). If all fail, accept **1 block/SM and lean on cp.async double-buffering + the 4× BW cut + mma throughput** — which is still very likely a win over 0031's serial-f32/-22%, just not the 2-block target.
- **Grid-starvation at n_seqs=1 is structural** (H=32 < 96): only dv_tile=32 fully saturates a single stream, and it pays ~1.5-2× redundant-A flops. This is an inherent floor of the chunked formulation on a 32-head layer; batched serving (n_seqs≥2) is where the design cleanly saturates.
- **Precision↔occupancy coupling:** the 2-block budget assumes **bf16 K/Q** staging. If the KL-gate demands tf32 for the KS/QS S0-products (decay-coupled), that needs a second 32-bit K/Q copy or 3×tf32 — both inflate shared/registers and can knock C=32 back toward 1 block/SM. The occupancy win is contingent on bf16 Gram clearing the gate.
**Bottom line:** 2 blocks/SM and full 48-SM saturation are *reachable* — via **C=32 + dv-slab (64 for serving, 32 for single-stream) + 8 warps + depth-2 cp.async** — on the **shared-memory** budget. Whether the **register** budget also permits it is the one genuinely open risk, and the A⁻¹ solve's footprint is what decides it; the realistic outcome remains the scope doc's honest **1-2 blocks/SM**, with the win carried by mma + cp.async + the C× BW cut rather than by high occupancy.
Files: design lives in `backend/cpp/llama-cpp-localai-paged/docs/TENSORCORE_GDN_SCOPE.md` (§3 occupancy plan — this analysis sharpens it); kernel/dispatcher to modify `backend/cpp/llama-cpp-paged-dev/ggml/src/ggml-cuda/gated_delta_net.cu` (grid at :184, chunked launch added by patch `patches/paged/0031-paged-chunked-gdn-prefill-scan-kernel.patch`); tf32 tile at `ggml/src/ggml-cuda/mma.cuh:1089`; H source `src/llama-model.cpp:504`.
## 5. Synthesized build plan + milestones + gate
Historical note: this plan predates the shipped f32-only M5 tensor-core GDN
path in patch `0047`. Current code parses `GDN_CHUNK_MIN` and `GDN_TC`; the
older `GDN_CHUNK_OFF` and `GDN_CHUNK_TC` names in this section are obsolete.
Phase71 revalidated the current default against sequential-disabled and
serial-chunked modes on DGX and kept M5 as shipped. Use this document as
background for any larger FLA/CuteDSL-class redesign, not as the active next
patch queue.
All anchors were verified at the time of writing. 0031's kernel body, the
7-step structure, the `GDN_CHUNK_MIN` gating at the `if constexpr (!KDA &&
!keep_rs_t)` site, the `launch_gdn_chunked<128,16>` template, the smem formula,
and the test-backend-ops shapes were confirmed. The scope doc's KL gate,
3xtf32 ladder, risk register, and Phase 0-3 plan were confirmed. Here is the
historical build-ready synthesis.
---
# BUILD-READY PLAN: tensor-core chunked-GDN prefill kernel (sm_121a)
Anchors (absolute): kernel + launch + gate live in `ggml/src/ggml-cuda/gated_delta_net.cu` (grid at :184), added by `/home/mudler/_git/LocalAI/.claude/worktrees/feat+paged-attention/backend/cpp/llama-cpp-localai-paged/patches/paged/0031-paged-chunked-gdn-prefill-scan-kernel.patch`. tf32 tile `mma(tile<16,8,float>&D, tile<16,8,float>&A, tile<8,8,float>&B)` = `mma.sync.aligned.m16n8k8.row.col.f32.tf32.tf32.f32` in `ggml/src/ggml-cuda/mma.cuh` (m16n8k8 overload ~976-984, dispatch ~1089), gated by `AMPERE_MMA_AVAILABLE`. PAD/cp.async patterns from `patches/paged/0034-feat-paged-native-NVFP4-W4A4-FP4-MMA-large-M-prefill.patch`. Gate/precedent docs: `docs/TENSORCORE_GDN_SCOPE.md`, `docs/PAGED_BITEXACT_NOTE.md`, `README.md` s5. Microbench: `~/scratch_tc_gdn_poc/gdn_gram_bench.cu` (DGX). Last patch in series is 0042 → this work is patches 0043+.
The new kernel is `gated_delta_net_chunked_tc_cuda<S_v, C, DV_TILE>`, a sibling to 0031's `gated_delta_net_chunked_cuda`. Symbols below reuse 0031's smem names (`Sd, Kc, Qc, Ud, Amat, csh, gam, bet`).
---
## (1) Phase-by-phase kernel structure
Block = **256 threads / 8 warps** in a **4×2 (WM×WN)** warp grid. State `S` (`dk×dv_tile`) is **register-resident in the step-6 accumulator layout** (`tile<16,8,float>` grid). Grid = `dim3(H, n_seqs, n_slabs)`, `blockIdx.z` = dv-slab. Chunk axis is the serial recurrence (NOT a grid axis). Invariant preserved from 0031: read pre-update `S0` (P3/P4) → solve → output (P5) → **overwrite S last** (P6). Single accumulator, no state double-buffer.
Per chunk `c0` (the loop body):
**Phase A - chunk load + gate prefix (f32, cooperative).** Load `Kc[C][dk]`, `Qc[C][dk]` **as bf16** (tf32 K/Q blows the 99KB budget - see §5 of the state design), load `V` chunk. Compute `csh = cumsum(g)` (≤0), `gam = exp(csh)` (≤1), `bet` - all f32, identical to 0031 lines (the `j==0` prefix scan, kept scalar; it is <1KB and hides under the Grams). cp.async depth-2 prefetch of the *next* chunk's `Kc/Qc` starts here.
**Phase B - state restage (accumulator→B bridge).** The crux. `S0` lives as P6's D/accumulator fragments but P3/P4 need it as a **B operand** (`tile<8,8>`, K-major over `i`). Bounce the `dk×dv_tile` state through a transient smem tile that **overlays the `UdAmat` region** (free at chunk entry - U/A not yet computed) → `load_generic` back as B fragments (NOT `ldmatrix`: it is `.b16`-only, unusable for tf32; use `load_generic`). Paid `n_tokens/C` times, **0KB net persistent smem**. KS and QS share this one restage.
**Phase C - Gram + state-boundary products (the matmuls that read pre-update S0).**
- **P1 `KK→A`** = `Kc·Kcᵀ`, M=C N=C K=dk, lower-tri (~½ tiles). **tf32-safe** (PoC-proven NMSE ~3e-9). Apply `A = I + tril(βₜ·d(t',t)·KK, -1)` in **f32** after the mma.
- **P3 `KS`** = `Kc·S0`, M=C N=dv K=dk. **3xtf32** (state-boundary, feeds the solve). Output → `Ud` region (becomes RHS).
- **P4 `QS`** = `Qc·S0`, M=C N=dv K=dk, **fused with P3 on the shared S0 B-fragments**. **3xtf32** (γ-attenuated → first demote candidate). Seed the **O accumulator fragments register-resident with `γₜ·QS`** immediately (avoids parking QS in smem through to Phase F). Restage overlay is now free; `Ud`/`Amat` reclaim it.
**Phase D - A-inverse (form T = A⁻¹ explicitly, then wide apply).**
- **Phase-D inverses:** 4 diagonal `16×16` unit-lower-tri blocks, **f32 in shared-memory column-parallel forward substitution** (thread `c` solves `A_ii x = e_c`). No tensor cores, no reduced precision (this is the strong-coupling amplifier). 4 blocks on 4 warps in parallel, hides entirely under the Phase-C/RHS Grams.
- **Phase-O off-diagonal:** wavefront (anti-diagonal) schedule, critical path `n_b-1=3` not 6. For each i>j: `P_ij = Σ_m A_im·T_mj` then `T_ij = -T_ii·P_ij`, on `m16n8k8`. **3xtf32 default-on** (~64 tiny mma total, negligible). `T` overwrites the `A` scratch in place.
**Phase E - RHS + apply.** `RHS = βₜ(vₜ - γₜ·KS)` in **f32** (uses P3 result + V) → `Ud`. **`U = T·RHS`** as one dependency-free wide **tf32** GEMM, M=C N=dv K=C (the bulk, 128 mma/warp at full dv), in place → `Ud`.
**Phase F - intra-chunk output.**
- **P2 `QK→P`** = `Qc·Kcᵀ`, reuse `Amat` (now free after T consumed). **tf32-safe**. Apply `P = d(t',t)·QK` in **f32** (bounded, decay pre-baked - preserves the bounded de-gating invariant).
- **P5 `O += P·U`**, M=C N=dv K=C, P lower-tri (~½ tiles). **tf32-safe** (P f32-bounded first). Accumulate into the O fragments already seeded with `γₜ·QS`. Write `O*scale` to `dst`.
**Phase G - state carry (overwrites S0 last).** `DU = d(t,last)·U` in f32. **Scale the persistent S accumulator fragments by `γ_last` in f32 in-register first**, then **P6 `S_C += Kcᵀ·DU`** = `Kcᵀ·DU`, M=dk N=dv K=C, **3xtf32 (the strongest ladder candidate - compounds over every chunk)**, accumulated straight into the persistent registers. `Kc` is read **transposed** here (second fragment view, `load_generic` transpose). No restage-out: S stays resident for the next chunk.
After the loop: final-state write-back (M-layout), identical to 0031's tail.
Buffer lifecycle (single `Amat`, single `Ud`, as 0031): `Amat`: A(P1) → T(Phase-D/O, in place) → consumed by apply → P(P2) → consumed by P5. `Ud`: KS(P3) → RHS(Phase-E) → U(apply, in place) → read by P5 (B) and P6 (B, scaled to DU). Restage tile overlays `UdAmat` only at chunk entry (Phase B), before either is written.
---
## (2) Build sequence - incremental, each independently GPU-verifiable vs 0031
Each milestone is a **separate patch** stacked on 0031, **green on `test-backend-ops GATED_DELTA_NET` + greedy-md5 stable before the next is started**. Reference for every step = the `test_gated_delta_net` op's f64/CPU oracle (already in-tree) and 0031's serial-chunked output. **No milestone integrates on top of an unverified one.**
| M | Scope | Patch | GPU verification gate (vs 0031 / op oracle) |
|---|---|---|---|
| **M0** | Re-confirm regime, NO code (scope Phase 0) | - | Profile 0031 (`GDN_CHUNK_MIN` low): confirm GDN prefill bucket dominates + grid-starved at low n_seqs. If not, kill the lever now. |
| **M1** | **DGX microbench (NO kernel yet)** - extend `gdn_gram_bench.cu` with KS/QS/PU/KᵀU microkernels + Phase-D/O T-formation + T·RHS apply, each with f64 host oracle measuring **κ(A)** and tf32-vs-3xtf32 NMSE per rung, incl. adversarial `g∈[-20,-1e-4]` | - | **The cheap go/no-go before multi-week kernel work.** Pass = default precision config (f32 diag + 3xtf32 off-diag + tf32 bulk) reaches ~PoC `3e-9`-grade on benign data and survives the κ(A) weak-decay corner within the ladder. Mirrors the PoC that proved 6.7×→9.3×. |
| **M2** | In-kernel: replace **only** step-1/2 serial Grams (KK/QK) with tensor-core tiles. **C=16, scalar everything else, same occupancy** (scope Phase 1 / PoC integration) | 0043 | test-backend-ops 128-shapes green via KL gate (NMSE if it passes); greedy-md5 stable. |
| **M3** | Add **P3/P4 (KS/QS)** tensor-core (3xtf32) + S restage bridge. Still C=16, scalar solve + scalar O/state | 0044 | Same gate. Isolates the accumulator→B bridge correctness. |
| **M4** | **A-inverse** Phase-D (f32 diag) + Phase-O (3xtf32 off-diag), form T; replace 0031's serial fwd-subst. Still C=16 | 0045 | Same gate + the adversarial-decay op case (this is the amplifier). |
| **M5** | **Apply `U=T·RHS`** + **P5 `P·U`** tensor-core. Still C=16 | 0046 | Same gate. |
| **M6** | **P6 `Kᵀ(D·U)`** tensor-core + **register-resident state** (step-6 accumulator layout) + accumulator→B restage in steady state. State leaves smem here | 0047 | Same gate. Frees the 64KB that forced C=16. |
| **M7** | **Flip C=16→C=64, full dv (CONFIG A ~86KB, 1 blk/SM)**, 8-warp 4×2 grid, PAD=4 smem | 0048 | Gate + **first A/B bench vs sequential** (S_PP at n_seqs≥2). |
| **M8** | **Occupancy:** C=32 + dv-slab grid `(H,n_seqs,n_slabs)` (CONFIG C, 2 blk/SM) + cp.async depth-2; selectors `GDN_CHUNK_C`/`GDN_DV_TILE` | 0049 | Gate + A/B bench across {C=32/dv64, C=32/dv32, C=64/dv64-BW-max}; pick winner per regime. |
---
## (3) Bit-exact / KL gate plan
**md5 is per-path and will NOT match** 0031-serial or the sequential recurrence (different FP reduction order). This is the established `-paged` precedent (`PAGED_BITEXACT_NOTE.md`): per-path md5, validated benign. So:
- **Binding gate = KL** (not strict NMSE): `KLD(tensorcore ‖ f16) ≤ KLD(sequential ‖ f16)` plus a PPL band, on the README s5 harness. NMSE is *expected to fail* at reduced precision (new path on a new path); NMSE-pass is a bonus, KL-pass is the bar.
- **Stability gate:** greedy-md5 **stable across runs** (deterministic), not equal to the serial path.
- **Adversarial op case mandatory:** `g∈[-20,-1e-4]` (the dangerous middle-decay regime where κ(A) grows); strong-decay underflows to 0 (safe), weak-decay is well-conditioned (tf32's 8-bit exponent holds γ range), the middle is the only empirical risk.
**Precision default config (the bet that clears the gate):** f32 diagonal inverse (mandatory, already f32) · **3xtf32 off-diagonal coupling** (default-on, negligible ~64-mma cost) · **tf32** Grams + apply · **bf16** K/Q staging (well-conditioned KK/QK only) · decays/γ**always f32 outside the mma** (invariant, not a rung). Hold **P6 state carry at 3xtf32 longest** (it compounds over every chunk).
**3xtf32 ladder (cheapest→dearest) if default misses the gate:** (1) KK Gram→3xtf32; (2) apply **block-diagonal `T_ii·RHS_i`**→3xtf32 (within-window strong coupling, mixed-precision-by-distance); (3) +**δ=1 off-diagonal** apply→3xtf32 (block-boundary adjacent pairs e.g. tokens 15↔16); (4) **full apply**→3xtf32 (≈+2× apply, expensive escape); (5) KS/QS→3xtf32; (6) fall back to direct blocked back-substitution in 3xtf32, else keep 0031's serial path. **Demote order if the gate has margin:** P4→P3, holding P6 at 3xtf32. If even all-3xtf32 misses, the residual is the f32 diagonal solve (already f32) → not fixable by more mma precision → fall to (6). Record the final rung in `PAGED_BITEXACT_NOTE.md` + README s5.
---
## (4) Slot into 0031's existing framework (historical, superseded by 0047)
Same dispatch site - the `if constexpr (!KDA && !keep_rs_t)` block inside `launch_gated_delta_net` (0031 patch, after `init_fastdiv_values`). Extend, don't replace:
- Current code keeps `GDN_CHUNK_MIN` as the token threshold and uses `GDN_TC`
as the tensor-core level selector. It does not parse `GDN_CHUNK_OFF` or
`GDN_CHUNK_TC`.
- Historical plan: add **`GDN_CHUNK_TC`** selector: `0` = 0031 serial-solve chunked (fallback, retained), `1` = tensor-core. Add **`GDN_CHUNK_C` ∈ {16,32,64}** and **`GDN_DV_TILE` ∈ {32,64,128}** for A/B; defaults `C=32, DV_TILE=64` (CONFIG C) for serving, `DV_TILE=32` saturator for n_seqs=1.
- New launcher `launch_gdn_chunked_tc<128, C, DV_TILE>` mirrors `launch_gdn_chunked`: `cudaFuncSetAttribute(...MaxDynamicSharedMemorySize...)` **return-checked** (0031 precedent), `grid = dim3(H, n_seqs, n_slabs)`, `block = dim3(256,1,1)`. Per-slab the kernel recomputes A/A⁻¹/gates (dv-independent), dv-slices S/Ud/O.
- **Default OFF** (`gdn_chunk_min=INT_MAX`) exactly as 0031 ships. Flip the default to on **only when** the M8 A/B shows an S_PP win over the tuned sequential recurrence at the serving regime (n_seqs≥2) **and** the KL gate + adversarial op case hold - recorded in README s5 (dev notes / rejected-flat levers) and `PAGED_BITEXACT_NOTE.md`. Until then it ships like 0031: opt-in, regression-free default.
- Extend the test-backend-ops block 0031 added (the `S_v==128` shapes at lines after :9398) so the tc path is exercised at C=64 and C=32 in CI.
- New per-path md5 acknowledged in the dispatch comment (tc-md5 ≠ serial-chunked-md5 ≠ sequential-md5; all benign, KL-validated).
---
## (5) Top 3 risks that could make it NOT beat sequential + kill-criteria
**Risk 1 - Register pressure forces 1 block/SM (the swing factor).** The ~50 working-regs/thread estimate is optimistic; the A⁻¹ blocked solve (in-register `16×16` diag inverse), the accumulator→B restage transpose, and the O+state transient accumulators can push working regs to 80-120. At 256 threads, >128 regs/thread → >32K regs/block → **silently 1 block/SM regardless of the 44KB shared headroom**, and local-memory spills push BW back. *Mitigation ladder:* (i) 12 warps (dilute state 32→21 regs/thread); (ii) `__launch_bounds__(256,2)`; (iii) smaller `DV_TILE`. **Kill criterion:** if after the full ladder the M8 occupancy build still spills to local OR stays 1 block/SM, **and** the CONFIG-A BW-max 1-block path (C=64, dv64, 96KB, cp.async, 4× state-BW cut) **also** fails to beat sequential S_PP at n_seqs≥2 in the A/B bench → the occupancy lever is dead; keep 0031 serial-chunked behind `GDN_CHUNK_TC=0`, record rejected in README s5.
**Risk 2 - Precision: tf32 (and even all-3xtf32) misses the KL gate in the weak-decay/aligned-keys κ(A) corner.** The inverse amplifies error; κ(A) is data-dependent and grows where keys align and decay is weak. **Detected cheaply at M1** (microbench measures κ(A) + per-rung NMSE on the adversarial case *before* the kernel exists). **Kill criterion:** if at M1 the **top of the ladder (all-3xtf32 + f32 diagonal)** cannot reach f32-grade on `g∈[-20,-1e-4]`, OR at M4+ `KLD(tc‖f16) ≤ KLD(seq‖f16)` fails on that op case at the top rung → the tensor-core solve is not numerically viable as a default; fall to ladder rung (6) (direct back-subst 3xtf32); if that also misses, abandon the tc solve and keep 0031 serial. **Fail-fast:** M1 gates this before any multi-week kernel commitment.
**Risk 3 - Grid starvation at n_seqs=1 is structural (H=32 < the ~96 blocks needed for 2 blk/SM × 48 SM).** Only `DV_TILE=32` (4 slabs) fully saturates a single stream, and it pays ~1.5-2× redundant-A flops (A/A⁻¹/gates recomputed per slab) plus the per-chunk restage. **Kill criterion:** if the M8 bench shows single-stream (n_seqs=1) S_PP is slower than sequential even at full saturation (dv32×4) due to redundant-A + restage overhead, **and** the batched regime (n_seqs≥2) gain also fails to materialize → the lever only helps a regime the target workload doesn't hit → keep default-OFF, ship as opt-in experiment only, record. (If n_seqs≥2 *does* win, ship enabled for the serving regime and gate single-stream back to sequential via `GDN_CHUNK_MIN` + an n_seqs check - a partial, honest win.)
**Overarching kill gate:** the disposition is the bench, not the theory. The kernel flips to default-on only when it beats the tuned sequential recurrence at the serving regime AND clears the KL + adversarial gates. Any milestone that regresses test-backend-ops or md5-stability halts the stack until fixed; M1 and M0 are the cheap fail-fast exits before the expensive kernel work.

Some files were not shown because too many files have changed in this diff Show More