mirror of
https://github.com/mudler/LocalAI.git
synced 2026-08-06 05:15:11 -04:00
Compare commits
321 Commits
feat/audio
...
worktree-f
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
3159ed0637 | ||
|
|
a1a3b99960 | ||
|
|
ac2b0211ff | ||
|
|
5b8b33a302 | ||
|
|
7b129a51f1 | ||
|
|
865e77c4ec | ||
|
|
586639d016 | ||
|
|
ccf75d1dcd | ||
|
|
500d653bfa | ||
|
|
b2784ccbca | ||
|
|
bf61db6214 | ||
|
|
b529cc5420 | ||
|
|
1aba41082b | ||
|
|
67d2c4c9d4 | ||
|
|
d091eb30f2 | ||
|
|
bbfaa66f02 | ||
|
|
04ed7fe52f | ||
|
|
a9454b45c8 | ||
|
|
f21b393746 | ||
|
|
26a41fad1a | ||
|
|
5369219729 | ||
|
|
eb82ff138f | ||
|
|
2efb0ec362 | ||
|
|
e5c5746c0a | ||
|
|
6cf8b782d1 | ||
|
|
e573194799 | ||
|
|
2b2b1f0b25 | ||
|
|
e67b329eb1 | ||
|
|
60954d484a | ||
|
|
3fbdfc21c9 | ||
|
|
55df9100dc | ||
|
|
2e19e5c90f | ||
|
|
6a2618b6dc | ||
|
|
f7d76389b0 | ||
|
|
4645935fa5 | ||
|
|
b425d8ce03 | ||
|
|
ef578866c8 | ||
|
|
fc5d5e4ff3 | ||
|
|
ef7dbfa5f7 | ||
|
|
c41d1a5b4f | ||
|
|
9be291e6b0 | ||
|
|
902bcc7717 | ||
|
|
999cf09532 | ||
|
|
3dbf34e739 | ||
|
|
347a5c05bd | ||
|
|
2aa76702df | ||
|
|
b5f65152e2 | ||
|
|
c299dcd231 | ||
|
|
cd59e5d61f | ||
|
|
96825a224e | ||
|
|
440129c98e | ||
|
|
e69ee0e867 | ||
|
|
2a0fc0f4b9 | ||
|
|
ae8284f5fb | ||
|
|
ecaf406c0b | ||
|
|
b9eff5bca3 | ||
|
|
aa848d5afb | ||
|
|
d44e164c96 | ||
|
|
52c11b1ce5 | ||
|
|
5354adcffb | ||
|
|
9f75da01f9 | ||
|
|
fbdc200886 | ||
|
|
49cce0b5a2 | ||
|
|
ba1979a689 | ||
|
|
7665422bfa | ||
|
|
70a4c31f36 | ||
|
|
e189e5a4ca | ||
|
|
b28b448c68 | ||
|
|
2148fa466b | ||
|
|
3b9ec3e1f1 | ||
|
|
3c2cb9f4ab | ||
|
|
ace1ffab28 | ||
|
|
a0194125f5 | ||
|
|
7108b68a70 | ||
|
|
7aa15ce539 | ||
|
|
6c165747a9 | ||
|
|
ff3f0620de | ||
|
|
c99678da42 | ||
|
|
310eb3c866 | ||
|
|
cced07c7fe | ||
|
|
6e35476340 | ||
|
|
ae76d42a96 | ||
|
|
4d171e62bb | ||
|
|
70394364a3 | ||
|
|
e169058e73 | ||
|
|
ede23df333 | ||
|
|
abc70c209e | ||
|
|
2074b4fb5b | ||
|
|
adabd11919 | ||
|
|
1b5ae227eb | ||
|
|
24e778de47 | ||
|
|
3da3b169fb | ||
|
|
ff3ad84191 | ||
|
|
9bbe02c161 | ||
|
|
b862e2c568 | ||
|
|
b009de0ee0 | ||
|
|
89ef3a4020 | ||
|
|
ef14748f06 | ||
|
|
b6885aa446 | ||
|
|
4b6fc0fa1c | ||
|
|
22a93ce1a3 | ||
|
|
3cf7fa1715 | ||
|
|
d0fa463eac | ||
|
|
34c4b5ce8d | ||
|
|
b647460dee | ||
|
|
f9e015d8e2 | ||
|
|
85c88320ef | ||
|
|
8b413d1cbd | ||
|
|
c5f2545cdd | ||
|
|
d8edc615e7 | ||
|
|
1c0709b700 | ||
|
|
337ebb8a37 | ||
|
|
ef5d4af203 | ||
|
|
a9a2efb296 | ||
|
|
b1a1b721bd | ||
|
|
b3cfdfac4a | ||
|
|
6ac06734e9 | ||
|
|
d288a0300f | ||
|
|
f8d7b026cf | ||
|
|
de34cd5954 | ||
|
|
1b9176c2c8 | ||
|
|
2033086f60 | ||
|
|
8bb47e5a8a | ||
|
|
2431090ff3 | ||
|
|
baf1025245 | ||
|
|
6edbb56b06 | ||
|
|
bd100dd20a | ||
|
|
be65438eac | ||
|
|
7b38c6b2a3 | ||
|
|
042deab40e | ||
|
|
c4058eb4da | ||
|
|
f1c98ff0b9 | ||
|
|
b028c81eda | ||
|
|
2fa8ef8fc5 | ||
|
|
d706980c2b | ||
|
|
000705321f | ||
|
|
4bdd26a7f0 | ||
|
|
9a28f23134 | ||
|
|
e610347367 | ||
|
|
11128cb080 | ||
|
|
4cd90bfae9 | ||
|
|
2c59805267 | ||
|
|
c51ff4cec9 | ||
|
|
ea72a56e2c | ||
|
|
1f3e5ba301 | ||
|
|
4da769c1ca | ||
|
|
23b11a5239 | ||
|
|
9bb8994c4e | ||
|
|
0b84fda496 | ||
|
|
1431f72b92 | ||
|
|
266fcc79ad | ||
|
|
3466094c68 | ||
|
|
ed5eb705c7 | ||
|
|
53f66a6f03 | ||
|
|
08b754f910 | ||
|
|
db14006fcd | ||
|
|
a4e730979d | ||
|
|
9115c2c52c | ||
|
|
984c8fcbea | ||
|
|
4a9a1dd247 | ||
|
|
78fac9a28f | ||
|
|
fb2dc33d52 | ||
|
|
a5a5b2ad80 | ||
|
|
7e1832b868 | ||
|
|
2bee7a5ab1 | ||
|
|
e160041f05 | ||
|
|
400930db19 | ||
|
|
202a29f980 | ||
|
|
621a20d2b5 | ||
|
|
2332587fdc | ||
|
|
af6e133759 | ||
|
|
87cfd1fadb | ||
|
|
2a2de1d6c1 | ||
|
|
5667dfe461 | ||
|
|
34abf392fc | ||
|
|
683e22500f | ||
|
|
db6ebc53b2 | ||
|
|
9b0e4e544c | ||
|
|
e3f8149f3b | ||
|
|
9a1be79f04 | ||
|
|
62c407ed55 | ||
|
|
c1f1d1e8ea | ||
|
|
6dd8a3d895 | ||
|
|
79edfd26a3 | ||
|
|
bf9b4fafa8 | ||
|
|
b1667b48ea | ||
|
|
6c6a925213 | ||
|
|
3b59571579 | ||
|
|
b3d3323105 | ||
|
|
9c1c2a6a16 | ||
|
|
1f857f179e | ||
|
|
33dfe7fd41 | ||
|
|
fe5bd3f53d | ||
|
|
6bfca146d6 | ||
|
|
4d3fecd524 | ||
|
|
ec7c1b1f68 | ||
|
|
30a2b590d9 | ||
|
|
167768cac3 | ||
|
|
125d10a782 | ||
|
|
b061e4aef0 | ||
|
|
89e62fc74f | ||
|
|
001d833426 | ||
|
|
00f92659f8 | ||
|
|
7dd3431040 | ||
|
|
ae0042f214 | ||
|
|
aaaa90ae4b | ||
|
|
7c45447c9e | ||
|
|
24833f0966 | ||
|
|
634c0e5a0f | ||
|
|
64766ecc85 | ||
|
|
02cbae5ea9 | ||
|
|
3c1ed67b4b | ||
|
|
8f8777e0f4 | ||
|
|
5cec1a6a21 | ||
|
|
17855735c7 | ||
|
|
2a8103c419 | ||
|
|
fd4332e8f0 | ||
|
|
5825b073a5 | ||
|
|
a72385257a | ||
|
|
2b57997df0 | ||
|
|
e597a8ac78 | ||
|
|
b895f4dff8 | ||
|
|
c0e0ed3865 | ||
|
|
ee13fd18ce | ||
|
|
6f0792c3be | ||
|
|
5ce2f1df51 | ||
|
|
34cadb64af | ||
|
|
2dd5d68e6d | ||
|
|
da67fd87e2 | ||
|
|
40f019e761 | ||
|
|
39e16cc2c4 | ||
|
|
7434d64c75 | ||
|
|
c1d7f336cb | ||
|
|
ea634ee958 | ||
|
|
e4c63179e0 | ||
|
|
f7500df64e | ||
|
|
24ce7d0823 | ||
|
|
fccbb4082d | ||
|
|
5a38dd3f09 | ||
|
|
ed17fc804e | ||
|
|
362eea90ff | ||
|
|
c7075fb796 | ||
|
|
c8b1f16507 | ||
|
|
2975a74fb4 | ||
|
|
ee78ae4a11 | ||
|
|
acb22a66ed | ||
|
|
010067d900 | ||
|
|
8925c009b7 | ||
|
|
a3abd60ae0 | ||
|
|
dd6a4425e0 | ||
|
|
4bc2b4a9b2 | ||
|
|
ba6bd94976 | ||
|
|
e983919516 | ||
|
|
2c5adda28c | ||
|
|
ee13a94a8c | ||
|
|
4dcbcfcf92 | ||
|
|
80e0c1ac6b | ||
|
|
52f0f7b8cf | ||
|
|
f347f7ca1d | ||
|
|
0dd45f0da5 | ||
|
|
9537726649 | ||
|
|
d1ba327843 | ||
|
|
ecffd4b097 | ||
|
|
67c6208b3a | ||
|
|
667a21c119 | ||
|
|
04e3d04ab8 | ||
|
|
4968cd8a94 | ||
|
|
37e0e1ef55 | ||
|
|
d9d846e04b | ||
|
|
84d59e659b | ||
|
|
931793aa24 | ||
|
|
0337505dc8 | ||
|
|
faeb5b457c | ||
|
|
6e0b910210 | ||
|
|
aaf7b4112e | ||
|
|
037ad82b7c | ||
|
|
1887385b79 | ||
|
|
40ee9cdd13 | ||
|
|
d6c91b7d62 | ||
|
|
92e93dfc34 | ||
|
|
fdb7f56bb7 | ||
|
|
07985ba45b | ||
|
|
fc589b3fad | ||
|
|
2b79083b71 | ||
|
|
2f648dc6a0 | ||
|
|
9973fa995a | ||
|
|
4de0c3b1b2 | ||
|
|
9a71e81fc4 | ||
|
|
718b31d063 | ||
|
|
d291e15114 | ||
|
|
dae2679c3b | ||
|
|
13e6ee89c7 | ||
|
|
76cc0b6abc | ||
|
|
122df1c620 | ||
|
|
14e3da25b6 | ||
|
|
f5e9caece1 | ||
|
|
d2651c86d9 | ||
|
|
19742aee64 | ||
|
|
ce60737fc5 | ||
|
|
37cbc089b0 | ||
|
|
b7b2e8291c | ||
|
|
cb28deda6b | ||
|
|
2a500c371f | ||
|
|
48fbb9384f | ||
|
|
145e45b6f2 | ||
|
|
c4b4f3a3e4 | ||
|
|
61ff738177 | ||
|
|
ce48cc0751 | ||
|
|
ba3fa5a633 | ||
|
|
62f0ae17e3 | ||
|
|
b14214620c | ||
|
|
1449b806ab | ||
|
|
9f16a907be | ||
|
|
aba0bfd24f | ||
|
|
7aa61d4c32 | ||
|
|
bbc84a9889 | ||
|
|
3ed3279739 | ||
|
|
ddace5fb6a | ||
|
|
5a5d3df8c8 | ||
|
|
c6698dd4bf | ||
|
|
edb1a11abc |
@@ -34,7 +34,7 @@ The build matrix is data-only YAML at `.github/backend-matrix.yml` (not inside `
|
||||
|
||||
**Without an entry here no image is ever built or pushed, and the gallery entry in `backend/index.yaml` will point at a tag that does not exist.** The `dockerfile:` field must point at `./backend/Dockerfile.<lang>` matching the language bucket from step 1 (e.g. `Dockerfile.python`, `Dockerfile.golang`, `Dockerfile.rust`). The `tag-suffix` must match the `uri:` in the corresponding `backend/index.yaml` image entry exactly.
|
||||
|
||||
**Path-filter registration — REQUIRED for any new dockerfile suffix.** This is the single most common omission, because it has no effect on the PR that adds the backend (when no prior path filter could catch it anyway) — it only breaks the *next* PR that touches your backend's directory, which then gets zero CI jobs and looks broken for unrelated reasons. Edit `scripts/lib/backend-filter.mjs:inferBackendPath` and add a branch BEFORE the more-generic suffixes:
|
||||
**`scripts/changed-backends.js` registration — REQUIRED for any new dockerfile suffix.** This is the single most common omission, because it has no effect on the PR that adds the backend (when no prior path filter could catch it anyway) — it only breaks the *next* PR that touches your backend's directory, which then gets zero CI jobs and looks broken for unrelated reasons. Edit `scripts/changed-backends.js:inferBackendPath` and add a branch BEFORE the more-generic suffixes:
|
||||
|
||||
```js
|
||||
if (item.dockerfile.endsWith("<your-dockerfile-suffix>")) {
|
||||
@@ -54,9 +54,7 @@ for (const e of m.include.filter(e => e.backend === '<your-backend>')) {
|
||||
}"
|
||||
```
|
||||
|
||||
A quick way to find the right insertion point: `grep -n 'item.dockerfile.endsWith' scripts/lib/backend-filter.mjs`.
|
||||
|
||||
If your backend consumes a *shared* build input that lives outside its own directory (a new script under `scripts/build/`, a new file copied into every image), add a rule to `SHARED_BUILD_INPUTS` in the same file — the per-backend prefix match cannot see those, and a miss ships your change to no image at all. See `scripts/lib/backend-filter_test.mjs` for the pattern; `make test-ci-scripts` runs it.
|
||||
A quick way to find the right insertion point: `grep -n 'item.dockerfile.endsWith' scripts/changed-backends.js`.
|
||||
|
||||
**`bump_deps.yaml` registration — REQUIRED for any backend pinning an upstream commit.** If your backend's Makefile has a `*_VERSION?=<sha>` pin to a third-party repo, the daily auto-bump bot at `.github/workflows/bump_deps.yaml` won't notice it unless you register the backend in its matrix. The bot runs `.github/bump_deps.sh` which `grep`s for `^$VAR?=` in the Makefile you list — so the pin MUST live in the Makefile (not in a separate shell script). The bump for ds4 (#9761) had to walk this back because the original landed the pin in `prepare.sh`, which the bot can't see. Pattern (for `antirez/ds4`):
|
||||
|
||||
@@ -117,7 +115,7 @@ Wiring a backend into `includeDarwin:` is more than the matrix entry:
|
||||
|
||||
1. **`includeDarwin:` entry** — `tag-suffix: "-metal-darwin-arm64-<backend>"`, `build-type: "metal"`, `lang: "go"` for go+ggml backends; omit `build-type` for the bespoke C++ ones (llama-cpp / ds4 / privacy-filter). Match an existing entry of the same shape.
|
||||
2. **`backend/index.yaml`** — add `metal:` to the backend's `capabilities` map (main and `-development`) and concrete `metal-<backend>` / `metal-<backend>-development` image entries pointing at the `-metal-darwin-arm64-<backend>` images.
|
||||
3. **C/C++ backends only** — add an `inferBackendPathDarwin` case in `scripts/lib/backend-filter.mjs` returning `backend/cpp/<backend>/` (the generic fallthrough assumes `backend/<lang>/`, which is wrong for a C++ source tree driven with `lang: go`), and give `run.sh` a Darwin branch that exports `DYLD_LIBRARY_PATH` instead of `LD_LIBRARY_PATH`. If the build is bespoke (single `grpc-server` + dylib bundling), model it on `scripts/build/ds4-darwin.sh` and add a `backends/<backend>-darwin` make target plus a gated step in `.github/workflows/backend_build_darwin.yml`.
|
||||
3. **C/C++ backends only** — add an `inferBackendPathDarwin` case in `scripts/changed-backends.js` returning `backend/cpp/<backend>/` (the generic fallthrough assumes `backend/<lang>/`, which is wrong for a C++ source tree driven with `lang: go`), and give `run.sh` a Darwin branch that exports `DYLD_LIBRARY_PATH` instead of `LD_LIBRARY_PATH`. If the build is bespoke (single `grpc-server` + dylib bundling), model it on `scripts/build/ds4-darwin.sh` and add a `backends/<backend>-darwin` make target plus a gated step in `.github/workflows/backend_build_darwin.yml`.
|
||||
4. **C++ proto gotcha** — if the backend compiles the generated gRPC/protobuf in a separate CMake target (e.g. `hw_grpc_proto`), that target must link `protobuf::libprotobuf` + `gRPC::grpc++` so the Homebrew include dirs propagate; otherwise macOS fails with `google/protobuf/runtime_version.h not found` (Linux hides this because apt headers sit in `/usr/include`).
|
||||
|
||||
The CI path filter only builds a backend on a PR when a file under its directory changes, so a darwin-only YAML edit builds nothing — touch a file under `backend/<lang>/<backend>/` (a one-line comment is enough) in the same PR.
|
||||
@@ -218,69 +216,6 @@ docker-build-backends: ... docker-build-<backend-name>
|
||||
- If the backend is in `backend/python/<backend-name>/` but uses `.` as context in the workflow file, use `.` context
|
||||
- Check similar backends to determine the correct context
|
||||
|
||||
## Engine preference for gallery model variants
|
||||
|
||||
A gallery entry can declare `variants`, alternative builds of the same weights,
|
||||
and LocalAI picks one per host: it drops builds whose backend cannot run here or
|
||||
that do not fit memory, then ranks the survivors by **engine preference
|
||||
first, serving feature second, size third** (`SelectVariant` in
|
||||
`core/gallery/resolve_variant.go`).
|
||||
|
||||
Ask whether your backend should outrank another one on some hardware. If it
|
||||
should, add it to `engineNamePreferenceRules` in `pkg/system/capabilities.go`,
|
||||
best engine first for that capability:
|
||||
|
||||
```go
|
||||
{Nvidia, []string{engineVLLM, engineSGLang, engineLlamaCpp}},
|
||||
+ {Nvidia, []string{engineVLLM, engineSGLang, engineMyEngine, engineLlamaCpp}},
|
||||
```
|
||||
|
||||
That is the ENGINE NAME table, matched as a substring of a gallery entry's
|
||||
`backend:` value. Two sibling tables in the same file speak different
|
||||
vocabularies and are matched against different things:
|
||||
|
||||
| Table | Vocabulary | Matched against | Consumer |
|
||||
|-------|-----------|-----------------|----------|
|
||||
| `backendBuildTagPreferenceRules` | build tags (`cuda`, `rocm`, `metal`) | installed build directory names, as a substring | alias resolution in `ListSystemBackends` |
|
||||
| `engineNamePreferenceRules` | engine names (`vllm`, `llama-cpp`, `mlx`) | a gallery entry's `backend:`, as a substring | gallery variant ranking |
|
||||
| `servingFeaturePreferenceTokens` | serving features (`dflash`, `mtp`) | a gallery entry's `tags:`, compared whole and case-insensitively, and nothing else | gallery variant ranking, one rank below the engine |
|
||||
|
||||
**Putting a token in the wrong table matches nothing and does not error**: every
|
||||
candidate scores equal and the next sort key decides, so the preference silently
|
||||
stops existing. The block comment above all three tables spells the contract out.
|
||||
|
||||
The serving feature table is the odd one: it is not keyed by capability, because
|
||||
no hardware prefers a plain build over an equivalent faster build of the same
|
||||
weights. It reads a declared tag and nothing else. The entry name was the
|
||||
original signal and is gone: a naming convention is not a contract, and names
|
||||
are author-supplied free text where a short marker like `mtp` turns up inside
|
||||
unrelated words or on weights whose entry enables nothing.
|
||||
`overrides.options` was rejected for the mirror-image reason: `spec_type:` is
|
||||
llama.cpp's config vocabulary, whereas a cross-backend ranking decision must
|
||||
work the same for `ds4`'s `mtp_path:` and `sglang`'s `speculative_algorithm:`.
|
||||
|
||||
**If your backend can serve the same weights faster** (speculative decoding,
|
||||
multi-token prediction), say so in the docs for its gallery entries so curators
|
||||
tag them: the tagging rule and the per-backend evidence table live in
|
||||
[adding-gallery-models.md](adding-gallery-models.md). A backend never needs to
|
||||
appear in the token table itself; it ranks builds, not engines.
|
||||
|
||||
Leaving your backend out is a valid choice when no ordering can be justified for
|
||||
it. It then ranks below every known engine and selection falls back to size,
|
||||
which is the behaviour that predates preference.
|
||||
|
||||
**Leaving a whole capability out is not.** A missing row gives that host an
|
||||
empty preference list, so size alone decides among everything that survives the
|
||||
filters, and the filter will not save you: `IsBackendCompatible` derives hardware
|
||||
support from the engine NAME, so `vllm` and `sglang` carry no darwin, cuda, rocm
|
||||
or sycl token and are never dropped on a host with no GPU. That is why `default`
|
||||
(no usable accelerator, including a GPU under the 4 GiB VRAM floor) and
|
||||
`darwin-x86` both have rows putting `llama-cpp` first. Every capability
|
||||
`getSystemCapabilities()` can return needs a row unless every engine really is
|
||||
equally at home there. When you add one, enumerate the engines you are demoting
|
||||
rather than relying on them falling through unmatched: unmatched engines all tie
|
||||
with each other, so size decides among them.
|
||||
|
||||
## Documenting the backend (README + docs)
|
||||
|
||||
A backend is not "added" until it is discoverable. Update the user-facing docs:
|
||||
@@ -308,7 +243,7 @@ After adding a new backend, verify:
|
||||
|
||||
- [ ] Backend directory structure is complete with all necessary files
|
||||
- [ ] Build configurations added to `.github/backend-matrix.yml` for all desired platforms (per-arch entries with `platform-tag` for multi-arch; `builder-base-image` for llama-cpp / ik-llama-cpp / turboquant)
|
||||
- [ ] **OS coverage considered**: added to `includeDarwin:` (macOS/Apple Silicon) if the backend can build there — with the `backend/index.yaml` `metal:` capability + `metal-<backend>` image entries, a `run.sh` Darwin/DYLD branch and `inferBackendPathDarwin` case (in `scripts/lib/backend-filter.mjs`) for C++ backends — or the PR explains why an OS is unsupported. Do not ship Linux-only by default.
|
||||
- [ ] **OS coverage considered**: added to `includeDarwin:` (macOS/Apple Silicon) if the backend can build there — with the `backend/index.yaml` `metal:` capability + `metal-<backend>` image entries, a `run.sh` Darwin/DYLD branch and `inferBackendPathDarwin` case for C++ backends — or the PR explains why an OS is unsupported. Do not ship Linux-only by default.
|
||||
- [ ] Meta definition added to `backend/index.yaml` in the `## metas` section
|
||||
- [ ] Image entries added to `backend/index.yaml` for all build variants (latest + development)
|
||||
- [ ] Tag suffixes match between workflow file and index.yaml
|
||||
@@ -316,8 +251,6 @@ After adding a new backend, verify:
|
||||
- [ ] No YAML syntax errors (check with linter)
|
||||
- [ ] No Makefile syntax errors (check with linter)
|
||||
- [ ] Follows the same pattern as similar backends (e.g., if it's a transcription backend, follow `faster-whisper` pattern)
|
||||
- [ ] **`Load` validates its input and refuses models it can't serve.** When a model config has no explicit `backend:`, the model loader greedily probes *every* installed backend with the model's name and binds to the first `Load` that succeeds — an accept-anything `Load` will capture arbitrary LLMs (issue #9287). Backends that load a real artefact get this for free (the load fails); backends with no artefact must gate on the name: `opus` accepts only its own name (or none), `local-store` requires the `store.NamespacePrefix` namespace marker sent by `core/backend/stores.go`.
|
||||
- [ ] **Gallery variant ranking considered**: if this backend should be preferred over another on some hardware, it is listed in `engineNamePreferenceRules` (NOT `backendBuildTagPreferenceRules`, NOT `servingFeaturePreferenceTokens`) in `pkg/system/capabilities.go`. A missing entry silently ranks it last and lets the next sort key decide.
|
||||
- [ ] Documented: added to the category list in `docs/content/features/backends.md` (and any new endpoint/realtime capability documented under `docs/content/`)
|
||||
- [ ] If it is an in-house native C/C++/GGML engine, added to the maintained-engines table in the top-level `README.md`
|
||||
|
||||
|
||||
@@ -91,108 +91,6 @@ To add a variant (e.g., different quantization), use YAML merge:
|
||||
uri: huggingface://<gguf-org>/<gguf-repo>/<filename>-Q8_0.gguf
|
||||
```
|
||||
|
||||
## Offering several builds of one model (`variants`)
|
||||
|
||||
When the same model is published in more than one quantization, or is also
|
||||
servable by another engine, add each build as its own ordinary gallery entry and
|
||||
then point one of them at the others with `variants`:
|
||||
|
||||
```yaml
|
||||
- !!merge <<: *chatml
|
||||
name: "nanbeige4.1-3b-q4"
|
||||
# ... the usual urls / overrides / files for the Q4 build ...
|
||||
variants:
|
||||
- model: nanbeige4.1-3b-q8
|
||||
```
|
||||
|
||||
Rules:
|
||||
|
||||
- The declaring entry is a **complete, normal entry**. It keeps its own
|
||||
`files`/`overrides` and stays installable on every host and by every older
|
||||
LocalAI release, which simply ignore `variants`.
|
||||
- A variant references another gallery entry **by name**. That entry must exist
|
||||
and must not declare `variants` of its own.
|
||||
- **A referenced entry keeps its own gallery row by default.** It is hidden only
|
||||
in the collapsed listing (`collapse_variants=true`, which the web UI requests
|
||||
by default), where the declaring entry stands in for it. Searching there still
|
||||
matches the referenced entry and answers with the entry declaring it, so
|
||||
referencing an entry never makes it unfindable; turning the collapse off
|
||||
returns it under its own name.
|
||||
- **Order carries no meaning.** Do not try to encode a preference; write the
|
||||
list in whatever order reads best.
|
||||
- **A variant may be smaller than the declaring entry.** Offering a downgrade
|
||||
for small hosts is a normal shape: the declaring entry's own build competes
|
||||
like every other candidate, so a large host keeps the large build.
|
||||
- **Do not describe hardware.** At install time LocalAI drops variants whose
|
||||
backend cannot run on the host, then drops those that do not fit available
|
||||
memory. The declaring entry's own build is exempt from both filters, so
|
||||
selection always terminates on something installable. Sizes are measured live
|
||||
from the weights and cached, so nothing has to be written down.
|
||||
- **Engine preference outranks size.** Among the builds that survive the
|
||||
filters, the host's preferred engine wins first and only then does the larger
|
||||
footprint win. On NVIDIA a vLLM build beats a larger llama.cpp one; on Apple
|
||||
silicon an MLX build beats a larger GGUF one; on a host with no preference for
|
||||
either engine the larger build wins, since a bigger footprint is a higher
|
||||
quality quantization of the same weights. Predict what a user gets by asking
|
||||
which engine the host prefers before asking which build is biggest. The
|
||||
per-capability order lives in `engineNamePreferenceRules`
|
||||
(`pkg/system/capabilities.go`); see
|
||||
[adding-backends.md](adding-backends.md) for how a backend gets into it.
|
||||
- **Serving feature preference sits between engine and size.** Among builds on
|
||||
an equally preferred engine, one that speculates or predicts several tokens
|
||||
per step beats the plain build of the same weights, because it answers faster
|
||||
for the same output: a `dflash` build beats an `mtp` one, and either beats a
|
||||
plain build. The order lives in `servingFeaturePreferenceTokens`
|
||||
(`pkg/system/capabilities.go`) and is matched against the entry's `tags:` and
|
||||
**nothing else**: not the entry name, not `overrides.options`. See
|
||||
[the tagging rule](#the-dflash--mtp-tagging-rule) below. Engine deliberately
|
||||
outranks it: a serving feature makes the right engine faster, it does not make
|
||||
a wrong engine right. Fit still outranks both, so a drafter pairing (strictly
|
||||
larger than the plain build, since it ships a drafter alongside it) is dropped
|
||||
on a host too small for it before this order is ever consulted.
|
||||
- A variant is nothing but a name; there is no per-variant memory field. When
|
||||
the measured size for a build is wrong, correct it on the referenced entry by
|
||||
setting that entry's own `size:` (e.g. `size: "20GiB"`). The estimator prefers
|
||||
a declared size over its own guesswork, so the fix applies everywhere the size
|
||||
is shown or compared rather than only to variant selection.
|
||||
|
||||
Users can override the automatic choice with `variant` on `POST /models/apply`,
|
||||
`local-ai models install --variant`, or the `install_model` MCP tool. See
|
||||
`docs/content/features/model-gallery.md`.
|
||||
|
||||
The gallery lint specs live in `core/gallery`, so run that suite after adding a
|
||||
`variants` list.
|
||||
|
||||
### The `dflash` / `mtp` tagging rule
|
||||
|
||||
**Tag an entry `dflash` or `mtp` when the entry actually configures that
|
||||
feature. Variant ranking reads the tag and nothing else.**
|
||||
|
||||
Decide by looking at what the entry configures, in whatever vocabulary its
|
||||
backend uses:
|
||||
|
||||
| Backend | Configures the feature when it declares |
|
||||
|---------|------------------------------------------|
|
||||
| `llama-cpp` | `overrides.options` contains `spec_type:draft-dflash` or `spec_type:draft-mtp` |
|
||||
| `ds4` | `overrides.options` contains `mtp_path:` / `mtp_draft:` |
|
||||
| `sglang` | the referenced `gallery/*.yaml` sets `speculative_algorithm:` |
|
||||
|
||||
That check is curation-time only. `spec_type` is llama.cpp's config vocabulary,
|
||||
and a cross-backend ranking decision must not depend on one backend's option
|
||||
syntax, which is precisely why the ranker reads the tag instead of the options.
|
||||
|
||||
Two mistakes the rule exists to prevent:
|
||||
|
||||
- **Weights that carry the heads are not an entry that enables them.** The
|
||||
NVFP4 GGUF entries ship MTP-bearing weights but set only `use_jinja:true`, so
|
||||
they enable no speculative decoding and must NOT be tagged. Tagging them wins
|
||||
them the feature axis without being any faster.
|
||||
- **A name is not a declaration.** An entry whose name spells `-mtp` while
|
||||
configuring nothing gets no tag, and an entry that configures the feature is
|
||||
tagged even when its name says nothing (`hy3`, `glm-5.2`). Ranking never reads
|
||||
the name, so an untagged build that does enable the feature is simply ranked
|
||||
as plain rather than promoted on a marker nobody meant.
|
||||
|
||||
## Available template configs
|
||||
|
||||
Look at existing `.yaml` files in `gallery/` to find the right prompt template for your model architecture:
|
||||
|
||||
@@ -28,6 +28,7 @@ The core Go suites (`./pkg`, `./core`, plus the in-process integration suite `./
|
||||
- **Build tags (`COVERAGE_TAGS`, passed via `GINKGO_TAGS`):** defaults to `debug auth`. The `auth` tag is required to compile the real (sqlite-backed) auth implementation and its ~150 `//go:build auth` tests — without it those files aren't built, the tests don't run, and the gate scores auth against a stub (~3.7% instead of ~38%). If you add new tag-gated tests, extend `COVERAGE_TAGS` or they won't count (and likely won't run in CI at all).
|
||||
- `make test-coverage-check` — runs `test-coverage`, then `scripts/coverage-check.sh` fails the build if total coverage is **below** the committed baseline in `coverage-baseline.txt`. The Linux job in `.github/workflows/test.yml` runs this instead of `make test`.
|
||||
- `make test-coverage-baseline` — regenerates and overwrites `coverage-baseline.txt` from the current run.
|
||||
- `make install-hooks` — sets `core.hooksPath` to the versioned `.githooks/`, whose `pre-commit` runs checks scoped to what's staged: Go changes → `make lint` + `make test-coverage-check`; `core/http/react-ui/` changes → `make test-ui-coverage-check` (Playwright e2e + UI coverage gate). A commit touching neither is skipped; bypass with `git commit --no-verify`. The hook resolves golangci-lint's new-from base to `upstream/master` → `origin/master` → `master`, so it works from a fork clone where `origin/master` is stale (passed to `make lint` via `LINT_NEW_FROM`).
|
||||
|
||||
### React UI coverage
|
||||
|
||||
@@ -37,11 +38,12 @@ The React UI (`core/http/react-ui/`) has **no component/unit tests** — its onl
|
||||
- **Browser:** the flake dev shell ships `chromium` and exports `PLAYWRIGHT_CHROMIUM_PATH`; `playwright.config.js` uses it via `launchOptions.executablePath`, and the Makefile skips `playwright install` when it's set. This avoids Playwright's downloaded browser, which can't resolve system libs (`libglib-2.0`, …) on NixOS. In CI (no `PLAYWRIGHT_CHROMIUM_PATH`) the Makefile falls back to `playwright install --with-deps chromium`.
|
||||
- The app is a React SPA, so coverage accumulates across in-app navigation within a test; a full `page.goto`/reload resets it.
|
||||
- `.nycrc.json` uses `all: true`, so **every `src/**` file is in the report**, including 0%-coverage ones — that's how you spot features with no test at all (sort the HTML report or `coverage-summary.json` by line% ascending).
|
||||
- **UI coverage gate:** `make test-ui-coverage-check` runs the suite then `scripts/ui-coverage-check.sh`, failing if total line coverage drops more than `UI_COVERAGE_TOLERANCE` below `core/http/react-ui/coverage-baseline.txt`. `make test-ui-coverage-baseline` regenerates the baseline. Runs in CI (`tests-ui-e2e.yml`).
|
||||
- **UI coverage gate:** `make test-ui-coverage-check` runs the suite then `scripts/ui-coverage-check.sh`, failing if total line coverage drops more than `UI_COVERAGE_TOLERANCE` below `core/http/react-ui/coverage-baseline.txt`. `make test-ui-coverage-baseline` regenerates the baseline. Runs in CI (`tests-ui-e2e.yml`) and pre-commit on `core/http/react-ui/` changes.
|
||||
- **Why it has a tolerance (unlike the strict Go gate):** UI e2e coverage is *non-deterministic*. Specs that assert on state and end while async/lazy render work is still in flight collect those lines only when the render beats the coverage teardown — so the total drifts with machine speed/load (a fast local box reads higher than a slow CI runner), diffusely across many specs. The tolerance absorbs that drift, so set the baseline *below* the slow-CI floor, never to a fast-local `make test-ui-coverage-baseline` number, or CI flaps.
|
||||
- **Raising coverage is cheap:** a *render-smoke* spec (navigate to a route, assert its header renders) mounts a lazy page and runs its full render + initial effects, capturing most of its lines in a few lines of test — see `e2e/page-render-smoke.spec.js`. Auth is disabled in the test server (`isAdmin=true`), so `RequireAdmin`/`RequireFeature` routes render without a mock. The most *deterministic* win is removing a race: make a spec `await` a rendered element before ending (see `e2e/agents.spec.js` → AgentCreate) so its lines count every run.
|
||||
|
||||
Rules (both gates):
|
||||
- **Don't weaken the gate:** never hand-lower a baseline or widen a tolerance to turn a red gate green. The ratchet only moves up.
|
||||
- **Install the hooks:** `make install-hooks` once per clone so lint + coverage run pre-commit. Don't lean on CI for what the hook catches.
|
||||
- **Don't work around the gate:** never `git commit --no-verify`, and never hand-lower a baseline or widen a tolerance to turn a red gate green. The ratchet only moves up.
|
||||
- If a change drops coverage, **add tests** (sort `coverage-summary.json` by line% ascending to find untested code) rather than editing the baseline. When coverage legitimately rises, commit the regenerated baseline (`make test-coverage-baseline` / `test-ui-coverage-baseline`).
|
||||
- The Go gate is **strict — no tolerance**; `covermode=atomic` keeps it deterministic. The UI gate keeps a small tolerance only because its e2e coverage isn't.
|
||||
|
||||
@@ -114,24 +114,6 @@ Both `backend.yml` (push) and `backend_pr.yml` (PR) generate their matrix dynami
|
||||
- **Tag pushes**: `FORCE_ALL=true` is set from the workflow side (`startsWith(github.ref, 'refs/tags/')`) — releases rebuild every backend regardless of diff.
|
||||
- **Schedule / `workflow_dispatch`**: no `event.before`, falls through to "run everything" automatically.
|
||||
|
||||
### Shared build inputs
|
||||
|
||||
The per-backend prefix match only sees files under a backend's own directory, so a change to shared build infrastructure would rebuild *nothing* — an empty matrix, every job green, and the change reaching no image. That silently un-shipped PR #10946 (a partial-cuDNN packaging fix in `scripts/build/package-gpu-libs.sh`), which merged 1h48m after the weekly cron and so sat unbuilt for a week.
|
||||
|
||||
`SHARED_BUILD_INPUTS` in `scripts/lib/backend-filter.mjs` closes that hole. Each rule maps a shared path to the narrowest set of matrix entries it can honestly invalidate, since a full matrix is 417 Linux + 56 Darwin builds:
|
||||
|
||||
| Changed path | Rebuilds |
|
||||
|---|---|
|
||||
| `backend/backend.proto` | everything (all languages compile or copy it) |
|
||||
| `backend/Dockerfile.<x>` | the Linux entries whose `dockerfile:` names it |
|
||||
| `backend/python/common/` | Python, Linux + Darwin |
|
||||
| `scripts/build/package-gpu-libs.sh` | Python, Linux only |
|
||||
| `scripts/build/<lang>-darwin.sh` | the Darwin entries that build target routes to |
|
||||
| `.github/workflows/backend_build[_darwin].yml` | everything on that OS |
|
||||
| anything else under `scripts/build/` (except `*_test.sh`) | everything — conservative default for unclassified packaging inputs |
|
||||
|
||||
Deliberately excluded: `backend/index.yaml` (gallery metadata, never enters an image), `.github/backend-matrix.yml` (adding a backend would rebuild all of them), `backend/Dockerfile.base-grpc-builder` (owned by `base-images.yml`), and the root `Makefile` (touched in ~11% of commits, and its backend-relevant edits arrive alongside the backend directory anyway). `make test-ci-scripts` pins all of this.
|
||||
|
||||
The Sunday 06:00 UTC cron on `backend.yml` exists specifically because path filtering can leave Python backends frozen on stale wheels. `DEPS_REFRESH` (below) only fires when the build actually runs, so an untouched Python backend would never re-resolve its unpinned deps. The weekly cron is the safety net.
|
||||
|
||||
## The `DEPS_REFRESH` cache-buster (Python backends)
|
||||
|
||||
@@ -65,7 +65,6 @@ This is enforced by `forbidigo` (see `.golangci.yml`): `http.DefaultClient` and
|
||||
|
||||
The project documentation is located in `docs/content`. When adding new features or changing existing functionality, it is crucial to update the documentation to reflect these changes. This helps users understand how to use the new capabilities and ensures the documentation stays relevant.
|
||||
|
||||
- **Docs-with-code rule**: When you change user-facing behavior (API endpoints, CLI flags, config keys, or features), update the corresponding page under `docs/content/` in the SAME change, not as a follow-up. A user-facing change without a matching docs update is incomplete. The PR template carries a checklist item for this.
|
||||
- **Feature Documentation**: If you add a new feature (like a new backend or API endpoint), create a new markdown file in `docs/content/features/` explaining what it is, how to configure it, and how to use it.
|
||||
- **Configuration**: If you modify configuration options, update the relevant sections in `docs/content/`.
|
||||
- **Examples**: providing concrete examples (like YAML configuration blocks) is highly encouraged to help users get started quickly.
|
||||
|
||||
143
.agents/llama-cpp-localai-paged-backend.md
Normal file
143
.agents/llama-cpp-localai-paged-backend.md
Normal file
@@ -0,0 +1,143 @@
|
||||
# llama-cpp-localai-paged Backend (paged attention + Blackwell NVFP4 decode)
|
||||
|
||||
`llama-cpp-localai-paged` is LocalAI's **CUDA-only** paged-attention variant of the
|
||||
llama.cpp backend. It targets high-concurrency decode for the Qwen3.6 hybrid
|
||||
gated-DeltaNet (SSM) models on Blackwell (GB10 / DGX Spark). It reuses the stock
|
||||
`llama-cpp` backend's sources and applies a vendored patch series on top at build
|
||||
time. It is **not** a fork: a source-only `*.patch` stack plus one canonical doc.
|
||||
|
||||
**Canonical reference:** `backend/cpp/llama-cpp-localai-paged/README.md`
|
||||
(architecture, the patch series 0001-0030, benchmarks, dev notes, generality,
|
||||
pin/canary policy). Read it for any technical detail; this guide is the maintenance
|
||||
how-to.
|
||||
|
||||
## Where things live
|
||||
|
||||
- `backend/cpp/llama-cpp-localai-paged/Makefile` - the thin wrapper. It copies the
|
||||
stock `backend/cpp/llama-cpp/` build infra into a build dir, clones llama.cpp at
|
||||
this backend's **own** pin (`LLAMA_VERSION`), applies the paged series via the
|
||||
`apply-paged-patches` define (strict `git apply`), then builds `grpc-server`.
|
||||
- `backend/cpp/llama-cpp-localai-paged/patches/paged/` - the source-only `.patch`
|
||||
series (0001-0030), nothing else.
|
||||
- `backend/cpp/llama-cpp-localai-paged/README.md` - the canonical doc. The
|
||||
operational docs (`PAGED_BITEXACT_NOTE.md`, `UPSTREAM_LAYER2_SCOPE.md`) and
|
||||
dev artifacts live in
|
||||
`backend/cpp/llama-cpp-localai-paged/docs/`.
|
||||
- `backend/Dockerfile.llama-cpp-localai-paged`, `.docker/llama-cpp-localai-paged-compile.sh`
|
||||
- the CUDA build entry points.
|
||||
- `backend/cpp/llama-cpp/` - the **stock** backend, pure upstream. It carries no
|
||||
paged patches.
|
||||
|
||||
## Invariants (do not break these)
|
||||
|
||||
- **Stock stays pure.** The paged patches live ONLY in this backend. Never add a
|
||||
`patches/paged/` dir or `LLAMA_PAGED` logic to `backend/cpp/llama-cpp/`.
|
||||
- **CUDA-only.** Ship cublas/cuda targets only. Off-CUDA the fusions are gated off
|
||||
(patch 0030) and NVFP4 falls back to dequant, so the backend is neutral-to-
|
||||
slightly-negative there - non-CUDA users use the stock `llama-cpp`. Do not add
|
||||
cpu/vulkan/sycl/metal rows for this backend in `.github/backend-matrix.yml`.
|
||||
(Those builds also fail to link `grpc-server` on darwin/arm64 against upstream
|
||||
`stream_*` server symbols - another reason it is CUDA-only.)
|
||||
- **Source-only patches.** A `.patch` may touch only llama.cpp source - never a
|
||||
dev doc or `*.md`. Strict `git apply` on a clean checkout must reach exit 0. (A
|
||||
stray `SSM_DECODE_FIX_RESULTS.md` hunk in patch 0019 once broke the CI build.)
|
||||
- **Bit-exact by default.** Every shipped patch is byte-identical to the f32
|
||||
baseline. (The one opt-in precision trade, `ssm_bf16_tau` / patch 0026, was
|
||||
DROPPED: it went flat once the decode fusions landed - forcing all gated-DeltaNet
|
||||
heads to bf16 gave 780.6 vs 780.0 t/s, zero benefit - so the series is now
|
||||
bit-exact end to end. Do not reintroduce a per-head SSM-precision lever; see the
|
||||
rejected-levers note in the backend README section 5.)
|
||||
|
||||
## Fork-first workflow (MANDATORY)
|
||||
|
||||
The fork **`mudler/llama.cpp` branch `localai-paged`** is the CANONICAL source
|
||||
of truth for ALL paged-backend kernel and patch work. The vendored
|
||||
`patches/paged/*.patch` series is a **derivative**: the fork is the source, the
|
||||
series is a generated mirror of it.
|
||||
|
||||
**Always update the fork FIRST, in this exact order:**
|
||||
|
||||
1. **Commit the change on the `localai-paged` branch and push it.** Every
|
||||
kernel or patch change lands as a fork commit first.
|
||||
2. **Then regenerate the LocalAI series from the fork** via `git format-patch`
|
||||
(one patch per fork commit, source-only) into
|
||||
`backend/cpp/llama-cpp-localai-paged/patches/paged/`, so the series stays a
|
||||
**1:1, drift-free mirror** of the branch.
|
||||
|
||||
Hard rules, no exceptions:
|
||||
|
||||
- **NEVER edit the `patches/paged/*.patch` files directly.** They are generated
|
||||
output, not source.
|
||||
- **NEVER add a patch to the series that has no corresponding fork-branch
|
||||
commit.** Every `.patch` must be the `git format-patch` of a real commit on
|
||||
`localai-paged`.
|
||||
- The fork branch is **where the build and the per-path bit-exact md5 gate
|
||||
actually run**, so it is the **only** place a change is truly validated. A
|
||||
patch living only in the LocalAI series has never been built or gated.
|
||||
|
||||
Verify the mirror by tree hash: applying the full on-disk series on the pin
|
||||
must reproduce the fork branch tree byte-for-byte. (The patch maintenance
|
||||
detail is in `backend/cpp/llama-cpp-localai-paged/docs/PATCH_MAINTENANCE.md`;
|
||||
the hard-gate is section 2.5 of `docs/PARITY_HANDOFF.md`.)
|
||||
|
||||
## Maintaining the pin against new llama.cpp
|
||||
|
||||
The pin (`LLAMA_VERSION` in the wrapper Makefile) is advanced ONLY by the manual
|
||||
pin-sync. It is deliberately **excluded from the nightly auto-bumper**
|
||||
(`bump_deps.yaml`): a naive bump would shift the tree out from under the patches
|
||||
and break `git apply` at build time.
|
||||
|
||||
1. **The canary tells you when to sync.** `.github/workflows/llama-cpp-paged-canary.yml`
|
||||
runs weekly: it applies + builds the series against the latest upstream tip and
|
||||
goes **red** when upstream drifts past the patches. Canary red -> run a pin-sync.
|
||||
2. **The pin-sync** (recorded in the README section 7 and git history): rebase the series onto the new
|
||||
tip (resolve conflicts; re-export **source-only** with a pathspec like
|
||||
`-- src/ ggml/ common/ include/ tools/ tests/ cmake/`), rebuild on a CUDA box,
|
||||
pass the bit-exact gate on **every** path + `test-backend-ops`, **and confirm
|
||||
the full grpc-server build/link is green on CI**, then bump `LLAMA_VERSION`.
|
||||
|
||||
**Hard constraint: keep the pin == the stock `llama-cpp` pin.** `grpc-server.cpp`
|
||||
is shared with the stock backend and tracks the stock pin. A paged pin that
|
||||
diverges PAST an upstream server-API refactor breaks the grpc-server LINK even
|
||||
when the patches are byte-for-byte bit-exact - the bit-exact gate alone does NOT
|
||||
catch it. The `c299a92c` bump did exactly this (patches applied + greedy-md5
|
||||
bit-exact, but `grpc-server.cpp` failed to link with undefined `stream_*` server
|
||||
helpers the refactor pulled into its headers), so it was reverted to `9d5d882d`.
|
||||
A pin bump is shippable only once the full CI grpc-server build is green, which in
|
||||
practice means moving in lockstep with the stock pin (or vendoring a
|
||||
pin-matched grpc-server.cpp, which we deliberately do not, to keep stock pure).
|
||||
|
||||
## The bit-exact gate (run for every change)
|
||||
|
||||
- greedy md5: `llama-completion -m MODEL -ngl 99 -fa on -p "The capital of France is" -n 48 --temp 0 --seed 1 </dev/null | md5sum`,
|
||||
paged paths prefixed `LLAMA_KV_PAGED=1` (+ `LLAMA_MOE_FORCE_GRAPHS=1` for paged
|
||||
MoE). Must match the recorded baseline. Redirect stdin from `/dev/null` or
|
||||
`llama-completion` hangs in conversation mode.
|
||||
- `test-backend-ops` (CUDA0 vs CPU oracle) for every touched op (`SSM_CONV*`,
|
||||
`GATED_DELTA_NET`, `MUL_MAT`, `MUL_MAT_ID`).
|
||||
- **The gate is per-path.** The paged-MoE md5 differs from the non-paged md5 - a
|
||||
benign, KL-validated FP-accumulation-order difference (see `docs/PAGED_BITEXACT_NOTE.md`).
|
||||
Compare a paged-MoE change to the **paged** reference, not the non-paged one.
|
||||
|
||||
## Encapsulating your work
|
||||
|
||||
- When you change a kernel, follow the **Fork-first workflow** above: commit and
|
||||
push on the `localai-paged` branch first, then regenerate the `.patch`
|
||||
(source-only) from the fork so this worktree mirrors the branch byte-for-byte.
|
||||
Commit with sign-off.
|
||||
- New optimization -> next patch number (gaps 0005/0027 are intentional). Update
|
||||
the README's patch table and dev notes - keep the README the single doc; do not
|
||||
scatter `*_RESULTS.md` files.
|
||||
- Record rejected/flat levers in the README too (they stop the next person from
|
||||
re-running dead ends).
|
||||
|
||||
## Follow-ups (Metal / SYCL / Vulkan)
|
||||
|
||||
The decode fusions are implemented for **CUDA + CPU only**. The base
|
||||
gated-DeltaNet + SSM_CONV ops already exist upstream on Metal, SYCL, and Vulkan,
|
||||
so the models **run** there via the non-fused path - what is missing is the
|
||||
fusion speedup. Porting it (strictly mirroring the CUDA kernels, since we have no
|
||||
Metal/SYCL/Vulkan hardware to test on here) is scoped in `docs/UPSTREAM_LAYER2_SCOPE.md`
|
||||
(recommended order: Metal, then SYCL, then Vulkan; ops-first upstream PR, then one
|
||||
PR per backend, each gated by `test-backend-ops` on the target hardware). The
|
||||
methodology for that work is in [.agents/vllm-parity-methodology.md](vllm-parity-methodology.md).
|
||||
101
.agents/vllm-parity-methodology.md
Normal file
101
.agents/vllm-parity-methodology.md
Normal file
@@ -0,0 +1,101 @@
|
||||
# Methodology: Closing the vLLM Decode-Throughput Gap in llama.cpp
|
||||
|
||||
This is the playbook that took the paged backend
|
||||
([.agents/llama-cpp-localai-paged-backend.md](llama-cpp-localai-paged-backend.md))
|
||||
from ~38% of vLLM decode to **parity-to-ahead on dense** (and a proven, honest
|
||||
ceiling on MoE) on GB10. Use it for any "make llama.cpp match or beat engine X on
|
||||
accelerator Y" effort. The *levers* are model- and hardware-specific; the
|
||||
*discipline* below is not. The worked example, with all numbers, is the paged
|
||||
backend README.
|
||||
|
||||
## The core loop
|
||||
|
||||
1. **Establish a bit-exact baseline and gate FIRST.** Record the greedy md5 (per
|
||||
path) and an f32 reference. Every optimization must stay byte-identical to it -
|
||||
or ship as an explicit, default-off precision opt-in. This is what lets you
|
||||
optimize aggressively without silently regressing quality. Gate two ways:
|
||||
greedy md5, and `test-backend-ops` against the CPU oracle.
|
||||
|
||||
2. **Profile - do not assume.** nsys the steady-state decode step, broken down per
|
||||
*kernel* AND per *memcpy*. Find the dominant cost. "It's the GEMM" was wrong
|
||||
here: on hybrid gated-DeltaNet models the bottleneck was the recurrent-state
|
||||
**plumbing** (state memcpy + gathers, ~67% of the step), not the weight GEMM.
|
||||
Also sanity-check GPU-busy %: an early "low utilization" reading was a profiling
|
||||
window artifact (decode was 96-99% GPU-busy), not real idle.
|
||||
|
||||
3. **Ground-truth BOTH engines.** Decompose *your* decode step AND the
|
||||
competitor's, side by side, per bucket, and compute the per-bucket delta. This
|
||||
tells you WHERE the gap actually is - not where you would guess. It overturned
|
||||
premises here: e.g. vLLM does NOT run the GDN/attn projections as NVFP4 (it
|
||||
keeps them bf16, same as us); the MoE expert GEMM was a llama *win*, not the gap.
|
||||
|
||||
4. **Per-lever discipline.** For each candidate: implement -> bit-exact gate ->
|
||||
same-harness A/B bench. Use a runtime env-toggle (flag off vs on) ONLY for
|
||||
levers that are actually runtime-gated; a lever **compiled into** the binary
|
||||
(e.g. the SSM decode fusions here) is NOT isolated by a runtime flag, so measure
|
||||
it build-vs-build. The full-patchset "stock" baseline likewise needs a
|
||||
**separately-built unpatched binary at the same pin** - toggling the runtime
|
||||
flag on the patched binary does not reproduce stock (it measures only the gated
|
||||
part; here that was ~neutral, which is exactly how this gotcha hides). Bank only
|
||||
what lifts AND gates. **Record every rejected or flat lever with the reason** -
|
||||
over time this is the most valuable part: it stops the next person re-running
|
||||
dead ends.
|
||||
|
||||
5. **Name the structural floor.** Prove the bit-exact ceiling exhaustively (every
|
||||
lever measured, not assumed). What remains is physical - the memory-bandwidth
|
||||
floor, the irreducible serial-SSM host loop (sampling can't start until logits
|
||||
land). Name it; do not claim more than you measured.
|
||||
|
||||
## Hard rules learned
|
||||
|
||||
- **Apples-to-apples, or label it.** Stock-vs-patched on the SAME harness
|
||||
(`llama-batched-bench`) is exact - lead with it. But "stock" must be a
|
||||
separately-built unpatched binary at the SAME pin, NOT the patched binary with
|
||||
the runtime flag off (compiled-in wins survive the toggle). Cross-engine "% of vLLM"
|
||||
(batched-bench vs vLLM server+client) is *indicative*; always caveat the harness
|
||||
and config (context length alone shifted the MoE figure 76% <-> 86%).
|
||||
- **Re-measure a "win" after later levers land - it may evaporate.** bf16 SSM
|
||||
state (the `ssm_bf16_tau` lever) benched +12% early and failed the f32 KL gate
|
||||
(vLLM keeps f32 too), so it was kept default-off opt-in. Once the decode fusions
|
||||
(recurrent-state gather-fusion + block-table cache) landed, a clean re-measure
|
||||
forcing ALL gated-DeltaNet heads to bf16 (`tau=100000`) went **flat** - 780.6 vs
|
||||
780.0 t/s. The "+12%" was subsumed by the fusions: the lever bought nothing, so
|
||||
it was **dropped** (precision trade + bug surface + extra CUDA template-instantiation
|
||||
compile cost, zero benefit). A win measured before the rest of the series is not a
|
||||
win after it.
|
||||
- **Reject the obvious-but-wrong, with evidence.** A faster kernel that is off the
|
||||
critical path benches FLAT (the freed time becomes idle). Quantizing the bf16
|
||||
projections to NVFP4 cost ~6% PPL - and vLLM keeps them bf16 for the same reason.
|
||||
Always measure before believing; a plausible mechanism is not a result.
|
||||
- **The gate can be per-path.** Paged vs non-paged attention legitimately produces
|
||||
different (equivalent) FP-reduction orders; validate the difference is benign
|
||||
(KLD to f32) and then gate each path against its own reference.
|
||||
|
||||
## Orchestration (multi-agent)
|
||||
|
||||
- **One GPU profiler/bencher at a time** (the GPU-contention rule). Parallel
|
||||
design/analysis/read agents are fine; concurrent GPU benches pollute each other's
|
||||
numbers.
|
||||
- **Adversarial verify.** Before banking a finding, spawn skeptics prompted to
|
||||
*refute* it; majority-refute kills it. Prevents plausible-but-wrong results.
|
||||
- **Anti-punt.** Use foreground, blocking ssh loops with short benches and a
|
||||
progress-file checkpoint. Agents that background work and "wait for the monitor
|
||||
event" stall - forbid that pattern.
|
||||
- **GPU coexistence.** On a shared host, stop the user's deployments for a clean
|
||||
benchmark window (with their OK) and ALWAYS restore them (wrap the bench so a
|
||||
failure cannot strand them).
|
||||
|
||||
## What generalizes (and what doesn't)
|
||||
|
||||
The *speedups* may be hardware-specific (here: CUDA/Blackwell - the SSM fusions,
|
||||
NVFP4 FP4-MMA, the occupancy tune), which is why other accelerators did not
|
||||
benefit. But the *findings* often generalize and are worth upstreaming: the
|
||||
"decode is plumbing-bound, not GEMM-bound" insight and the bit-exact, CPU-mirrored
|
||||
fusion ops help any backend running these models. Separate "ship our tuned backend"
|
||||
from "upstream the portable op" - they are different deliverables.
|
||||
|
||||
## The closing record
|
||||
|
||||
Write up the result HONESTLY: the shipped wins, the rejected levers (with reasons),
|
||||
the structural ceiling, and the cross-backend / cross-quant generality. Negative
|
||||
results are as valuable as wins. The paged backend README is the template.
|
||||
@@ -28,10 +28,6 @@ if [ -z "${BUILD_TYPE:-}" ]; then
|
||||
# variants with it (the host never *selects* SME unless it has it, but every variant must
|
||||
# still compile).
|
||||
if [ "${TARGETARCH}" = "arm64" ]; then
|
||||
# The prebuilt base inherits default ports.ubuntu.com sources; honor the
|
||||
# APT_*_MIRROR build args here like the from-source path does, so this
|
||||
# apt step survives a mirror outage.
|
||||
sh /LocalAI/.docker/apt-mirror.sh || true
|
||||
apt-get update -qq && apt-get install -y -qq gcc-14 g++-14
|
||||
export CC=gcc-14 CXX=g++-14
|
||||
fi
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
#!/usr/bin/env bash
|
||||
# Shared compile logic for backend/Dockerfile.bonsai.
|
||||
# Shared compile logic for backend/Dockerfile.llama-cpp-localai-paged.
|
||||
# Sourced (via bind mount) from both builder-fromsource and builder-prebuilt stages.
|
||||
|
||||
set -euxo pipefail
|
||||
@@ -14,10 +14,10 @@ if [[ -n "${CUDA_DOCKER_ARCH:-}" ]]; then
|
||||
CUDA_ARCH_ESC="${CUDA_DOCKER_ARCH//;/\\;}"
|
||||
export CMAKE_ARGS="${CMAKE_ARGS} -DCMAKE_CUDA_ARCHITECTURES=${CUDA_ARCH_ESC}"
|
||||
echo "CMAKE_ARGS(env) = ${CMAKE_ARGS}"
|
||||
rm -rf /LocalAI/backend/cpp/bonsai-*-build
|
||||
rm -rf /LocalAI/backend/cpp/llama-cpp-localai-paged-*-build
|
||||
fi
|
||||
|
||||
cd /LocalAI/backend/cpp/bonsai
|
||||
cd /LocalAI/backend/cpp/llama-cpp-localai-paged
|
||||
|
||||
if [ -z "${BUILD_TYPE:-}" ]; then
|
||||
# Pure CPU image: one ggml CPU_ALL_VARIANTS build replaces the per-microarch binaries.
|
||||
@@ -26,14 +26,14 @@ if [ -z "${BUILD_TYPE:-}" ]; then
|
||||
apt-get update -qq && apt-get install -y -qq gcc-14 g++-14
|
||||
export CC=gcc-14 CXX=g++-14
|
||||
fi
|
||||
make bonsai-cpu-all
|
||||
make llama-cpp-localai-paged-cpu-all
|
||||
else
|
||||
# GPU build (cublas/hipblas/sycl/vulkan/...): single fallback CPU build, the accelerator
|
||||
# does the compute. Keeps the GPU compile from also building the CPU variant matrix and
|
||||
# avoids the gcc-14 apt step on GPU base images such as nvidia l4t.
|
||||
make bonsai-fallback
|
||||
make llama-cpp-localai-paged-fallback
|
||||
fi
|
||||
make bonsai-grpc
|
||||
make bonsai-rpc-server
|
||||
make llama-cpp-localai-paged-grpc
|
||||
make llama-cpp-localai-paged-rpc-server
|
||||
|
||||
ccache -s || true
|
||||
60
.githooks/pre-commit
Executable file
60
.githooks/pre-commit
Executable file
@@ -0,0 +1,60 @@
|
||||
#!/usr/bin/env sh
|
||||
#
|
||||
# LocalAI pre-commit hook. Install it (once per clone) with:
|
||||
#
|
||||
# make install-hooks
|
||||
#
|
||||
# Runs only the checks relevant to what's staged:
|
||||
# - Go files -> make lint + make test-coverage-check
|
||||
# - core/http/react-ui -> make test-ui-coverage-check (Playwright e2e + gate)
|
||||
# A commit touching neither is skipped entirely (docs/YAML/etc. can't change
|
||||
# lint findings, Go coverage, or the UI).
|
||||
#
|
||||
# To bypass for a single commit (e.g. a WIP checkpoint): git commit --no-verify
|
||||
set -eu
|
||||
|
||||
repo_root="$(git rev-parse --show-toplevel)"
|
||||
cd "$repo_root"
|
||||
|
||||
staged="$(git diff --cached --name-only --diff-filter=ACMRD)"
|
||||
|
||||
go_changed=0
|
||||
ui_changed=0
|
||||
if echo "$staged" | grep -qE '\.go$'; then go_changed=1; fi
|
||||
if echo "$staged" | grep -qE '^core/http/react-ui/'; then ui_changed=1; fi
|
||||
|
||||
if [ "$go_changed" -eq 0 ] && [ "$ui_changed" -eq 0 ]; then
|
||||
echo "pre-commit: no Go or React UI changes staged — skipping."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
if [ "$go_changed" -eq 1 ]; then
|
||||
# Resolve the ref golangci-lint's new-from-merge-base should compare
|
||||
# against. .golangci.yml pins origin/master, which is correct in CI
|
||||
# (origin == the canonical repo) but wrong from a fork clone, where
|
||||
# origin/master lags behind and lint would report the whole upstream
|
||||
# backlog. Prefer upstream/master, then origin/master, then master.
|
||||
lint_base=""
|
||||
for ref in upstream/master origin/master master; do
|
||||
if git rev-parse --verify --quiet "${ref}^{commit}" >/dev/null 2>&1; then
|
||||
lint_base="$ref"
|
||||
break
|
||||
fi
|
||||
done
|
||||
|
||||
echo "pre-commit ▶ golangci-lint (make lint${lint_base:+, new-from $lint_base})"
|
||||
make lint LINT_NEW_FROM="$lint_base"
|
||||
|
||||
echo "pre-commit ▶ coverage gate (make test-coverage-check) — builds and runs the"
|
||||
echo " pkg/core suites plus tests/e2e; can take a few minutes."
|
||||
make test-coverage-check
|
||||
fi
|
||||
|
||||
if [ "$ui_changed" -eq 1 ]; then
|
||||
echo "pre-commit ▶ React UI e2e + coverage gate (make test-ui-coverage-check) —"
|
||||
echo " rebuilds the UI + ui-test-server, runs the Playwright specs, and"
|
||||
echo " fails if line coverage regressed; can take a couple of minutes."
|
||||
make test-ui-coverage-check
|
||||
fi
|
||||
|
||||
echo "pre-commit ✓ all relevant checks passed"
|
||||
1
.github/PULL_REQUEST_TEMPLATE.md
vendored
1
.github/PULL_REQUEST_TEMPLATE.md
vendored
@@ -7,7 +7,6 @@ This PR fixes #
|
||||
|
||||
**[Signed commits](../CONTRIBUTING.md#signing-off-on-commits-developer-certificate-of-origin)**
|
||||
- [ ] Yes, I signed my commits.
|
||||
- [ ] Documentation updated (docs/content/) for user-facing changes, or not applicable
|
||||
|
||||
<!--
|
||||
Thank you for contributing to LocalAI!
|
||||
|
||||
826
.github/backend-matrix.yml
vendored
826
.github/backend-matrix.yml
vendored
File diff suppressed because it is too large
Load Diff
18
.github/bump_deps.sh
vendored
18
.github/bump_deps.sh
vendored
@@ -1,8 +1,5 @@
|
||||
#!/bin/bash
|
||||
set -xe
|
||||
|
||||
source "$(dirname "${BASH_SOURCE[0]}")/gh_curl.sh"
|
||||
|
||||
REPO=$1
|
||||
BRANCH=$2
|
||||
VAR=$3
|
||||
@@ -12,20 +9,7 @@ if [ -z "$FILE" ]; then
|
||||
FILE="Makefile"
|
||||
fi
|
||||
|
||||
# gh_curl follows redirects so a renamed/transferred upstream repo (GitHub
|
||||
# answers 301) still resolves, and fails on HTTP errors rather than letting an
|
||||
# error page reach sed below. `|| true` keeps a failed lookup from aborting the
|
||||
# script at exit 22 with no context — the SHA guard below reports it instead.
|
||||
LAST_COMMIT=$(gh_curl -H "Accept: application/vnd.github.VERSION.sha" "https://api.github.com/repos/$REPO/commits/$BRANCH" || true)
|
||||
|
||||
# Guard the sed input: anything that is not a bare 40-hex SHA (an API error
|
||||
# body, an empty response) would otherwise be spliced into the Makefile pin —
|
||||
# either corrupting it silently or blowing up sed with an unterminated
|
||||
# expression, which is how this job failed for a renamed repo.
|
||||
if ! [[ "$LAST_COMMIT" =~ ^[0-9a-f]{40}$ ]]; then
|
||||
echo "Refusing to bump $VAR: expected a 40-char commit SHA for $REPO@$BRANCH, got: $LAST_COMMIT" >&2
|
||||
exit 1
|
||||
fi
|
||||
LAST_COMMIT=$(curl -s -H "Accept: application/vnd.github.VERSION.sha" "https://api.github.com/repos/$REPO/commits/$BRANCH")
|
||||
|
||||
# Read $VAR from Makefile (only first match)
|
||||
set +e
|
||||
|
||||
13
.github/bump_docs.sh
vendored
13
.github/bump_docs.sh
vendored
@@ -1,18 +1,7 @@
|
||||
#!/bin/bash
|
||||
set -xe
|
||||
|
||||
source "$(dirname "${BASH_SOURCE[0]}")/gh_curl.sh"
|
||||
|
||||
REPO=$1
|
||||
|
||||
LATEST_TAG=$(gh_curl -H "Accept: application/vnd.github+json" \
|
||||
"https://api.github.com/repos/$REPO/releases/latest" | jq -r '.tag_name')
|
||||
|
||||
# jq prints the string "null" for a missing key, so a throttled or otherwise
|
||||
# unexpected API response would otherwise be published as the docs version.
|
||||
if [ -z "$LATEST_TAG" ] || [ "$LATEST_TAG" = "null" ]; then
|
||||
echo "Refusing to bump docs version: could not resolve the latest release tag for $REPO." >&2
|
||||
exit 1
|
||||
fi
|
||||
LATEST_TAG=$(curl -s "https://api.github.com/repos/$REPO/releases/latest" | jq -r '.tag_name')
|
||||
|
||||
cat <<< $(jq ".version = \"$LATEST_TAG\"" docs/data/version.json) > docs/data/version.json
|
||||
|
||||
7
.github/bump_vllm_metal.sh
vendored
7
.github/bump_vllm_metal.sh
vendored
@@ -11,9 +11,6 @@
|
||||
# darwin build can only use the exact vLLM version vllm-metal supports, so it may
|
||||
# lag the Linux pin (requirements-cublas13-after.txt) until vllm-metal catches up.
|
||||
set -xe
|
||||
|
||||
source "$(dirname "${BASH_SOURCE[0]}")/gh_curl.sh"
|
||||
|
||||
REPO=$1 # vllm-project/vllm-metal
|
||||
FILE=$2 # backend/python/vllm/install.sh
|
||||
VAR=$3 # VLLM_METAL_VERSION (used for the workflow's output file names)
|
||||
@@ -25,12 +22,12 @@ fi
|
||||
|
||||
# vllm-metal ships frequent dev releases, all flagged as non-prerelease, so
|
||||
# /releases/latest returns the newest one (with its cp312 wheel asset).
|
||||
LATEST_TAG=$(gh_curl -H "Accept: application/vnd.github+json" \
|
||||
LATEST_TAG=$(curl -sS -H "Accept: application/vnd.github+json" \
|
||||
"https://api.github.com/repos/$REPO/releases/latest" \
|
||||
| python3 -c "import json,sys; print(json.load(sys.stdin)['tag_name'])")
|
||||
|
||||
# The coupled vLLM source version lives in vllm-metal's installer at that tag.
|
||||
NEW_VLLM_VERSION=$(gh_curl \
|
||||
NEW_VLLM_VERSION=$(curl -fsSL \
|
||||
"https://raw.githubusercontent.com/$REPO/$LATEST_TAG/install.sh" \
|
||||
| grep -oE 'vllm_v="[0-9]+\.[0-9]+\.[0-9]+"' | head -1 | cut -d'"' -f2)
|
||||
|
||||
|
||||
5
.github/bump_vllm_wheel.sh
vendored
5
.github/bump_vllm_wheel.sh
vendored
@@ -9,9 +9,6 @@
|
||||
# vars in Makefiles; this script handles the two-value rewrite specific to the
|
||||
# vLLM requirements file.
|
||||
set -xe
|
||||
|
||||
source "$(dirname "${BASH_SOURCE[0]}")/gh_curl.sh"
|
||||
|
||||
REPO=$1 # vllm-project/vllm
|
||||
FILE=$2 # backend/python/vllm/requirements-cublas13-after.txt
|
||||
VAR=$3 # VLLM_VERSION (used for output file names so the workflow can read them)
|
||||
@@ -22,7 +19,7 @@ if [ -z "$FILE" ] || [ -z "$REPO" ] || [ -z "$VAR" ]; then
|
||||
fi
|
||||
|
||||
# /releases/latest returns the most recent non-prerelease tag.
|
||||
LATEST_TAG=$(gh_curl -H "Accept: application/vnd.github+json" \
|
||||
LATEST_TAG=$(curl -sS -H "Accept: application/vnd.github+json" \
|
||||
"https://api.github.com/repos/$REPO/releases/latest" \
|
||||
| python3 -c "import json,sys; print(json.load(sys.stdin)['tag_name'])")
|
||||
|
||||
|
||||
194
.github/ci/apexentries/README.md
vendored
194
.github/ci/apexentries/README.md
vendored
@@ -1,194 +0,0 @@
|
||||
# apexentries
|
||||
|
||||
Generates gallery entries for the `mudler/*-APEX-GGUF` HuggingFace repositories.
|
||||
|
||||
Each APEX repo becomes one **family**: one entry per quality rung the repo
|
||||
publishes and one per quantization rung its unsloth counterpart publishes, all
|
||||
gathered under the **base model's** entry. LocalAI's variant selector then picks
|
||||
the build that fits the hardware in front of it.
|
||||
|
||||
## The hub is the base model entry, never a generated `*-apex` parent
|
||||
|
||||
Somebody looking for `qwen3.6-35b-a3b` must find every build of those weights
|
||||
under that one name: the APEX imatrix rungs, the unsloth quant rungs and any
|
||||
speculative build. A separate `qwen3.6-35b-a3b-apex` hub competing with the base
|
||||
entry would split the family in two and leave whichever half the user did not
|
||||
search for effectively invisible.
|
||||
|
||||
So the generator resolves the hub by stripping the `-APEX`, `-MTP` and `-TQ`
|
||||
markers and looking the result up in the index, trying both the repo-derived and
|
||||
the stem-derived candidate the same way `CounterpartCandidates` does. Then:
|
||||
|
||||
- **The hub exists** (14 of the 45 repos, resolving to 10 distinct entries).
|
||||
Nothing new is emitted for the family root. A `variants:` block is spliced into
|
||||
the entry that is already there, textually, leaving its description, icon,
|
||||
tags, overrides and files untouched. The line editing is shared with the
|
||||
`variantproposals` job via `.github/ci/galleryedit`.
|
||||
- **The hub is absent** (the other 31). A new hub is emitted, named for the base
|
||||
model and never for the APEX repo. It carries one of the discovered builds as
|
||||
its own payload so it is a complete installable entry rather than a bare index,
|
||||
and that payload is what gives it an `overrides.backend`. Without a declared
|
||||
backend the verifier would skip it, so a hub carrying feature tags would escape
|
||||
the tagging check in silence.
|
||||
|
||||
Several APEX repos routinely resolve to one base model, so both paths accumulate
|
||||
by hub name rather than assuming one family per hub.
|
||||
|
||||
Two references are always filtered out of a hub's list: anything the entry
|
||||
already declares, and the hub's own name. The self reference is not merely
|
||||
redundant. An unsloth rung whose weights the gallery already ships under the base
|
||||
name resolves, through the merge, straight back to the hub, and the verifier
|
||||
reads a self reference as a variant that declares variants of its own.
|
||||
|
||||
The four hand-written `*-apex` entries (`qwen3.6-35b-a3b-apex`,
|
||||
`gemma-4-26b-a4b-it-apex`, `qwen3.5-35b-a3b-apex`,
|
||||
`nemotron-3-nano-omni-30b-a3b-reasoning-apex`) are **ordinary builds**, not hubs.
|
||||
They are referenced from their hub's variants list like any other rung, and are
|
||||
never deleted or renamed.
|
||||
|
||||
## Flags
|
||||
|
||||
| Flag | Default | Meaning |
|
||||
|------|---------|---------|
|
||||
| `-index <path>` | `gallery/index.yaml` | Gallery index to dedup against. Read only, unless `-apply` is passed. |
|
||||
| `-only <a,b,c>` | (all) | Comma-separated full repo names (`mudler/Foo-APEX-GGUF`) to restrict generation to. A name that matches nothing is reported as a warning, since it is a typo rather than an empty result. |
|
||||
| `-out <path>` | (none) | Write the entries to add to this file. Nothing is written to the gallery. |
|
||||
| `-apply` | `false` | Splice the variants into `-index` and append the new entries to it. |
|
||||
| `-verify <path>` | (none) | Verify a gallery index and exit. Ignores every other flag. |
|
||||
|
||||
Either `-out` or `-apply` is required, otherwise the run has nothing to do.
|
||||
|
||||
`-apply` splices variant lines into existing entries and **appends** new ones. It
|
||||
never re-serialises the index: it is roughly 40,000 lines, and a YAML round trip
|
||||
would reflow the whole file, drop the anchors and merge keys the gallery relies
|
||||
on, and produce a diff nobody can review. On the three-family sample the splice
|
||||
is 24 added lines across 3 hunks with zero deletions.
|
||||
|
||||
## Discovery is by filename suffix, never by repo name
|
||||
|
||||
Builds come from the files a repo actually publishes. A filename is never
|
||||
constructed from a repo name, because the two disagree:
|
||||
`mudler/gemma-4-26B-A4B-it-APEX-GGUF` ships `gemma-4-26B-A4B-APEX-*.gguf`, and
|
||||
five other repos likewise drop a suffix (`-it`, `-2603`) or a vendor prefix
|
||||
(`NVIDIA-`) that the repo name carries. Composing a URL from the repo name would
|
||||
produce a 404 for every one of them, and the 404 would only surface after the
|
||||
entry shipped.
|
||||
|
||||
The quality ladder is matched on the trailing tier marker, `-(I-)?(Quality|
|
||||
Balanced|Compact|Mini|Nano).gguf`. The `I-` prefix marks the imatrix ladder. The
|
||||
imatrix ladder is emitted when it is non-empty and the plain ladder is used only
|
||||
as a fallback, because two of the 45 repos publish no imatrix tiers at all and
|
||||
must still contribute. Eleven repos carry a fifth `I-Nano` rung, so nothing
|
||||
assumes a fixed number of rungs.
|
||||
|
||||
Every run prints, per repo, the counts that discovery accounted for. If the
|
||||
number of classified files is short of the number of `.gguf` files the repo
|
||||
publishes, the shortfall is printed as `UNCLASSIFIED`. That check is a set
|
||||
difference on counts rather than a second pass over filenames: a second matcher
|
||||
would duplicate the tier regex and the two copies would drift. The failure it
|
||||
catches is quiet. A publishing-script typo that breaks every imatrix filename in
|
||||
a repo does not produce a short ladder; it makes the imatrix ladder empty, and
|
||||
the fallback then downgrades the whole family to the plain ladder with nothing
|
||||
said. A downstream HTTP check cannot catch it either, because it validates the
|
||||
URLs that were emitted, and an undiscovered tier emits none.
|
||||
|
||||
The same reasoning applies to `UNACCOUNTED QUANT`, printed when the unsloth
|
||||
counterpart demonstrably publishes a wanted quant that produced no build. It is
|
||||
reported at discovery time because a dropped quant leaves no trace at all in the
|
||||
finished gallery file.
|
||||
|
||||
## sha256 always comes from the API
|
||||
|
||||
Every file stanza takes its `sha256` from the HuggingFace models API
|
||||
(`lfs.sha256`). A GGUF the API describes without one is a fatal error for that
|
||||
family: the repo is reported by name and the run ends non-zero. It is never
|
||||
substituted from another field, because that is exactly how a Xet hash ends up
|
||||
masquerading as a content hash.
|
||||
|
||||
## The dflash / mtp tagging rule
|
||||
|
||||
An entry is tagged `dflash` or `mtp` **if and only if** it configures the
|
||||
matching `spec_type:draft-<feature>`. Variant ranking reads tags and nothing
|
||||
else, so a tag that does not match the configuration either promotes a build
|
||||
that is no faster or hides one that genuinely is.
|
||||
|
||||
A repo name is not configuration. `mudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF` ships
|
||||
weights that carry MTP heads; an entry that does not enable them is not an MTP
|
||||
entry and is not tagged as one.
|
||||
|
||||
A generated hub inherits the tags of the build it carries as its payload, rather
|
||||
than rebuilding them from the base set, so a hub whose payload configures a
|
||||
`spec_type` stays tagged consistently with the overrides copied alongside it.
|
||||
|
||||
## Reuse reporting: two categories, not one
|
||||
|
||||
Generated entries are deduped against the gallery and against the batch itself.
|
||||
The run prints the result under two separate headings, because the two cases are
|
||||
not equivalent:
|
||||
|
||||
- **URI MATCHES** mean the gallery, or an earlier entry in this batch, already
|
||||
ships exactly these weights. Pointing the hub at the existing entry is correct
|
||||
and needs no thought.
|
||||
- **NAME COLLISIONS** mean an entry already owns the name but holds different
|
||||
weights. Referencing it would point the hub at a build other than the one
|
||||
generated. Every one of these must be inspected by hand.
|
||||
|
||||
The run then prints `HUBS SPLICED`, listing every reference that will be added to
|
||||
an entry the gallery already ships along with the line it will be added at, and
|
||||
`HUBS CREATED` for the families that get a new hub. The splices are the part a
|
||||
review has to read closely, because they modify entries somebody else wrote.
|
||||
|
||||
Hubs are deliberately kept out of the merge. A new hub carries the family's top
|
||||
rung as its own payload, so URI dedup would fold the hub into that rung and the
|
||||
family would lose the very entry point this command exists to create.
|
||||
|
||||
## Workflow: sample first, then the full set
|
||||
|
||||
Never run the full generation straight into the gallery. Generate a small,
|
||||
deliberately awkward sample, have it reviewed, then run the rest.
|
||||
|
||||
```bash
|
||||
# 1. Sample three families that between them cover the awkward shapes:
|
||||
# a standard four-rung repo, one with the extra I-Nano rung AND a file stem
|
||||
# that differs from its repo name, and one whose unsloth counterpart shards
|
||||
# its quants across subdirectories.
|
||||
go run ./.github/ci/apexentries \
|
||||
-index gallery/index.yaml \
|
||||
-only mudler/Qwen3.6-35B-A3B-APEX-GGUF,mudler/gemma-4-26B-A4B-it-APEX-GGUF,mudler/Step-3.7-Flash-APEX-GGUF \
|
||||
-out /tmp/sample.yaml
|
||||
|
||||
# 2. Verify the sample against the gallery it would join, splices included. Apply
|
||||
# to a COPY, never to the real index, and check that the diff is only the
|
||||
# intended variant lines. Compare the verifier output to the gallery's own
|
||||
# baseline: what matters is that the sample adds no new problem, not that the
|
||||
# total is zero.
|
||||
cp gallery/index.yaml /tmp/index-copy.yaml
|
||||
go run ./.github/ci/apexentries -index /tmp/index-copy.yaml -only <same list> -apply
|
||||
diff -u gallery/index.yaml /tmp/index-copy.yaml # expect zero deletions
|
||||
|
||||
go run ./.github/ci/apexentries -verify gallery/index.yaml > /tmp/baseline.log 2>&1
|
||||
go run ./.github/ci/apexentries -verify /tmp/index-copy.yaml > /tmp/spliced.log 2>&1
|
||||
diff /tmp/baseline.log /tmp/spliced.log
|
||||
|
||||
# 3. Have a human review /tmp/sample.yaml and every reported name collision.
|
||||
|
||||
# 4. Only then, the full set.
|
||||
go run ./.github/ci/apexentries -index gallery/index.yaml -apply
|
||||
```
|
||||
|
||||
## Tests
|
||||
|
||||
```bash
|
||||
go test ./.github/ci/apexentries/
|
||||
```
|
||||
|
||||
The shared line editor has its own package:
|
||||
|
||||
```bash
|
||||
go test ./.github/ci/galleryedit/
|
||||
```
|
||||
|
||||
`.github/ci/` is invisible to `go list ./...`, so these specs are not covered by
|
||||
`make lint` or the repository test run. `.github/workflows/ci-tools-tests.yaml`
|
||||
names the package explicitly; keep that workflow in step with any package added
|
||||
under `.github/ci/`.
|
||||
70
.github/ci/apexentries/discover.go
vendored
70
.github/ci/apexentries/discover.go
vendored
@@ -1,70 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"regexp"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// tierRE matches the tier marker APEX repos put at the end of a weight
|
||||
// filename. Discovery is by suffix because the stem is not predictable from
|
||||
// the repo name: six of the 45 repos drop a suffix ("-it", "-2603") or a
|
||||
// vendor prefix ("NVIDIA-") that the repo name carries.
|
||||
var tierRE = regexp.MustCompile(`-(I-)?(Quality|Balanced|Compact|Mini|Nano)\.gguf$`)
|
||||
|
||||
// fullPrecisionRE matches the unquantized source weights an APEX repo publishes
|
||||
// alongside its ladder, flat (-F16.gguf) or sharded across a numbered set
|
||||
// (-F16-00001-of-00010.gguf). bf16 is accepted because some repos publish that
|
||||
// instead, and the match is case-insensitive because the casing varies between
|
||||
// publishing scripts.
|
||||
//
|
||||
// These are deliberately not tiers: they are the weights the ladder is quantized
|
||||
// FROM, and generation is scoped to the ladder itself.
|
||||
var fullPrecisionRE = regexp.MustCompile(`(?i)-b?f16(-\d{5}-of-\d{5})?\.gguf$`)
|
||||
|
||||
// IsFullPrecision reports whether a weight filename is an unquantized source.
|
||||
func IsFullPrecision(name string) bool {
|
||||
return fullPrecisionRE.MatchString(name)
|
||||
}
|
||||
|
||||
// Tier is one discovered build of an APEX repo.
|
||||
type Tier struct {
|
||||
Label string
|
||||
File GGUFFile
|
||||
}
|
||||
|
||||
// DiscoverAPEXTiers splits a repo's weight files into the imatrix ladder and
|
||||
// the plain ladder. mmproj files are never tiers.
|
||||
func DiscoverAPEXTiers(files []GGUFFile) (imatrix, plain []Tier) {
|
||||
for _, f := range files {
|
||||
if strings.HasPrefix(f.Name, "mmproj") {
|
||||
continue
|
||||
}
|
||||
m := tierRE.FindStringSubmatch(f.Name)
|
||||
if m == nil {
|
||||
continue
|
||||
}
|
||||
if m[1] != "" {
|
||||
imatrix = append(imatrix, Tier{Label: "I-" + m[2], File: f})
|
||||
continue
|
||||
}
|
||||
plain = append(plain, Tier{Label: m[2], File: f})
|
||||
}
|
||||
return imatrix, plain
|
||||
}
|
||||
|
||||
// DiscoverMMProj returns the repo's projector file, if it publishes one. The
|
||||
// name varies across repos (mmproj.gguf, mmproj-F16.gguf,
|
||||
// mmproj-step3.7-flash-f16.gguf), so match the prefix rather than a fixed name.
|
||||
func DiscoverMMProj(files []GGUFFile) (GGUFFile, bool) {
|
||||
for _, f := range files {
|
||||
if strings.HasPrefix(f.Name, "mmproj") {
|
||||
return f, true
|
||||
}
|
||||
}
|
||||
return GGUFFile{}, false
|
||||
}
|
||||
|
||||
// FileStem returns a tier's filename with its tier suffix removed.
|
||||
func FileStem(t Tier) string {
|
||||
return tierRE.ReplaceAllString(t.File.Name, "")
|
||||
}
|
||||
68
.github/ci/apexentries/discover_test.go
vendored
68
.github/ci/apexentries/discover_test.go
vendored
@@ -1,68 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
)
|
||||
|
||||
var _ = Describe("DiscoverAPEXTiers", func() {
|
||||
It("finds tiers regardless of how the stem relates to the repo name", func() {
|
||||
// This repo is mudler/gemma-4-26B-A4B-it-APEX-GGUF but its files drop "-it".
|
||||
files := []GGUFFile{
|
||||
{Name: "gemma-4-26B-A4B-APEX-I-Quality.gguf", SHA256: "a"},
|
||||
{Name: "gemma-4-26B-A4B-APEX-I-Nano.gguf", SHA256: "b"},
|
||||
{Name: "gemma-4-26B-A4B-APEX-Quality.gguf", SHA256: "c"},
|
||||
{Name: "mmproj-F16.gguf", SHA256: "d"},
|
||||
}
|
||||
|
||||
imatrix, plain := DiscoverAPEXTiers(files)
|
||||
|
||||
Expect(labels(imatrix)).To(ConsistOf("I-Quality", "I-Nano"))
|
||||
Expect(labels(plain)).To(ConsistOf("Quality"))
|
||||
})
|
||||
|
||||
It("excludes mmproj from the tier list", func() {
|
||||
files := []GGUFFile{{Name: "mmproj.gguf", SHA256: "d"}}
|
||||
|
||||
imatrix, plain := DiscoverAPEXTiers(files)
|
||||
|
||||
Expect(imatrix).To(BeEmpty())
|
||||
Expect(plain).To(BeEmpty())
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("DiscoverMMProj", func() {
|
||||
It("finds an mmproj whatever its suffix", func() {
|
||||
files := []GGUFFile{
|
||||
{Name: "Model-APEX-I-Mini.gguf", SHA256: "a"},
|
||||
{Name: "mmproj-step3.7-flash-f16.gguf", SHA256: "b"},
|
||||
}
|
||||
|
||||
got, ok := DiscoverMMProj(files)
|
||||
|
||||
Expect(ok).To(BeTrue())
|
||||
Expect(got.Name).To(Equal("mmproj-step3.7-flash-f16.gguf"))
|
||||
})
|
||||
|
||||
It("reports absence when the repo ships none", func() {
|
||||
_, ok := DiscoverMMProj([]GGUFFile{{Name: "Model-APEX-Quality.gguf", SHA256: "a"}})
|
||||
|
||||
Expect(ok).To(BeFalse())
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("FileStem", func() {
|
||||
It("strips the tier suffix", func() {
|
||||
t := Tier{Label: "I-Quality", File: GGUFFile{Name: "gemma-4-26B-A4B-APEX-I-Quality.gguf"}}
|
||||
|
||||
Expect(FileStem(t)).To(Equal("gemma-4-26B-A4B-APEX"))
|
||||
})
|
||||
})
|
||||
|
||||
func labels(ts []Tier) []string {
|
||||
out := make([]string, 0, len(ts))
|
||||
for _, t := range ts {
|
||||
out = append(out, t.Label)
|
||||
}
|
||||
return out
|
||||
}
|
||||
130
.github/ci/apexentries/hf.go
vendored
130
.github/ci/apexentries/hf.go
vendored
@@ -1,130 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"net/http"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"github.com/mudler/LocalAI/pkg/httpclient"
|
||||
)
|
||||
|
||||
// ErrNoSHA256 marks a GGUF the HuggingFace API describes without an
|
||||
// lfs.sha256. Emitting an entry without a hash would ship an unverifiable
|
||||
// download, and guessing one from another field is how a Xet hash ends up
|
||||
// masquerading as a content hash, so this is fatal rather than skippable.
|
||||
var ErrNoSHA256 = errors.New("gguf file has no lfs.sha256")
|
||||
|
||||
// GGUFFile is one .gguf sibling of a HuggingFace repo.
|
||||
type GGUFFile struct {
|
||||
Name string
|
||||
Size int64
|
||||
SHA256 string
|
||||
}
|
||||
|
||||
type apiSibling struct {
|
||||
RFilename string `json:"rfilename"`
|
||||
Size int64 `json:"size"`
|
||||
LFS *struct {
|
||||
SHA256 string `json:"sha256"`
|
||||
} `json:"lfs"`
|
||||
}
|
||||
|
||||
type apiModel struct {
|
||||
Siblings []apiSibling `json:"siblings"`
|
||||
}
|
||||
|
||||
// ParseRepoFiles returns every .gguf sibling described by a models API body.
|
||||
func ParseRepoFiles(body []byte) ([]GGUFFile, error) {
|
||||
var m apiModel
|
||||
if err := json.Unmarshal(body, &m); err != nil {
|
||||
return nil, fmt.Errorf("decoding model response: %w", err)
|
||||
}
|
||||
|
||||
var out []GGUFFile
|
||||
for _, s := range m.Siblings {
|
||||
if !strings.HasSuffix(s.RFilename, ".gguf") {
|
||||
continue
|
||||
}
|
||||
if s.LFS == nil || s.LFS.SHA256 == "" {
|
||||
return nil, fmt.Errorf("%s: %w", s.RFilename, ErrNoSHA256)
|
||||
}
|
||||
out = append(out, GGUFFile{Name: s.RFilename, Size: s.Size, SHA256: s.LFS.SHA256})
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// FetchOptionalRepoFiles asks the models API for a repo the caller can do
|
||||
// without, and reports separately whether the repo was merely unreadable.
|
||||
//
|
||||
// HuggingFace answers 401 Unauthorized, not 404, for a repo that does not exist
|
||||
// when the request carries no credentials. Without a token there is therefore no
|
||||
// way to tell "this repo was never published" from "this repo is private", so an
|
||||
// optional probe has to treat 401 and 403 exactly like 404: whatever the reason,
|
||||
// there is nothing here for us to read, so there is no counterpart.
|
||||
//
|
||||
// The second return value exists because that collapse is lossy in one
|
||||
// direction: 401/403 can also mean a real, gated repo whose quants we would
|
||||
// genuinely want. The caller reports those repos so a silently dropped
|
||||
// counterpart is visible to a human rather than invisible.
|
||||
func FetchOptionalRepoFiles(client *http.Client, repo string) ([]GGUFFile, bool, error) {
|
||||
files, status, err := fetchRepoFiles(client, repo)
|
||||
if err != nil && (status == http.StatusUnauthorized || status == http.StatusForbidden) {
|
||||
return nil, true, nil
|
||||
}
|
||||
return files, false, err
|
||||
}
|
||||
|
||||
// FetchRepoFiles asks the models API for one repo. A 404 yields (nil, nil) so
|
||||
// that probing for an optional counterpart repo is not an error. Every other
|
||||
// non-200, 401 and 403 included, is an error: for a repo the run REQUIRES there
|
||||
// is no benign reading of "we cannot see it".
|
||||
func FetchRepoFiles(client *http.Client, repo string) ([]GGUFFile, error) {
|
||||
files, _, err := fetchRepoFiles(client, repo)
|
||||
return files, err
|
||||
}
|
||||
|
||||
// fetchRepoFiles does the request and returns the HTTP status alongside the
|
||||
// result, so the optional and required callers can apply different policies to
|
||||
// the same response without duplicating the request.
|
||||
func fetchRepoFiles(client *http.Client, repo string) ([]GGUFFile, int, error) {
|
||||
url := fmt.Sprintf("https://huggingface.co/api/models/%s?blobs=true", repo)
|
||||
req, err := http.NewRequest(http.MethodGet, url, nil)
|
||||
if err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
req.Header.Set("User-Agent", "localai-apexentries/1.0")
|
||||
|
||||
resp, err := client.Do(req)
|
||||
if err != nil {
|
||||
return nil, 0, err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
|
||||
if resp.StatusCode == http.StatusNotFound {
|
||||
return nil, resp.StatusCode, nil
|
||||
}
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
return nil, resp.StatusCode, fmt.Errorf("%s: unexpected status %d", repo, resp.StatusCode)
|
||||
}
|
||||
|
||||
body, err := io.ReadAll(resp.Body)
|
||||
if err != nil {
|
||||
return nil, resp.StatusCode, err
|
||||
}
|
||||
files, err := ParseRepoFiles(body)
|
||||
return files, resp.StatusCode, err
|
||||
}
|
||||
|
||||
// newHTTPClient builds the client used against the HuggingFace API. It goes
|
||||
// through pkg/httpclient rather than a bare &http.Client{} because the std
|
||||
// client follows redirects and forwards custom credential headers to the
|
||||
// redirect target on a cross-host hop (GHSA-3mj3-57v2-4636). This caller sends
|
||||
// only a User-Agent today, but it talks to an external API that could start
|
||||
// redirecting, and an HF_TOKEN header here later would then leak.
|
||||
func newHTTPClient() *http.Client {
|
||||
return httpclient.NewWithTimeout(60 * time.Second)
|
||||
}
|
||||
142
.github/ci/apexentries/hf_test.go
vendored
142
.github/ci/apexentries/hf_test.go
vendored
@@ -1,142 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"io"
|
||||
"net/http"
|
||||
"testing"
|
||||
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
)
|
||||
|
||||
func TestApexEntries(t *testing.T) {
|
||||
RegisterFailHandler(Fail)
|
||||
RunSpecs(t, "apexentries")
|
||||
}
|
||||
|
||||
// stubTransport answers every request with one canned status and body, so the
|
||||
// status handling of the fetchers can be exercised without reaching the real
|
||||
// HuggingFace API.
|
||||
type stubTransport struct {
|
||||
status int
|
||||
body string
|
||||
}
|
||||
|
||||
func (t stubTransport) RoundTrip(req *http.Request) (*http.Response, error) {
|
||||
return &http.Response{
|
||||
StatusCode: t.status,
|
||||
Body: io.NopCloser(bytes.NewBufferString(t.body)),
|
||||
Header: make(http.Header),
|
||||
Request: req,
|
||||
}, nil
|
||||
}
|
||||
|
||||
func stubClient(status int, body string) *http.Client {
|
||||
return &http.Client{Transport: stubTransport{status: status, body: body}}
|
||||
}
|
||||
|
||||
const oneGGUFBody = `{"siblings":[{"rfilename":"Model-APEX-I-Quality.gguf","size":10,"lfs":{"sha256":"aa","size":10}}]}`
|
||||
|
||||
var _ = Describe("FetchOptionalRepoFiles", func() {
|
||||
// HuggingFace answers 401 rather than 404 for a repo that does not exist
|
||||
// when the client carries no credentials, so an optional probe cannot tell
|
||||
// "absent" from "unauthorized" and must treat both as "no counterpart".
|
||||
It("treats a 401 as an absent repo and flags it as unavailable", func() {
|
||||
files, unavailable, err := FetchOptionalRepoFiles(stubClient(http.StatusUnauthorized, ""), "unsloth/Nope-GGUF")
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(files).To(BeEmpty())
|
||||
Expect(unavailable).To(BeTrue())
|
||||
})
|
||||
|
||||
It("treats a 403 as an absent repo and flags it as unavailable", func() {
|
||||
files, unavailable, err := FetchOptionalRepoFiles(stubClient(http.StatusForbidden, ""), "unsloth/Gated-GGUF")
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(files).To(BeEmpty())
|
||||
Expect(unavailable).To(BeTrue())
|
||||
})
|
||||
|
||||
// A clean 404 is an unambiguous absence, so it must NOT be reported as
|
||||
// unavailable: the whole point of the flag is to separate the ambiguous
|
||||
// case a human may need to look at from the settled one.
|
||||
It("treats a 404 as an absent repo without flagging it as unavailable", func() {
|
||||
files, unavailable, err := FetchOptionalRepoFiles(stubClient(http.StatusNotFound, ""), "unsloth/Nope-GGUF")
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(files).To(BeEmpty())
|
||||
Expect(unavailable).To(BeFalse())
|
||||
})
|
||||
|
||||
It("parses a 200 body as usual", func() {
|
||||
files, unavailable, err := FetchOptionalRepoFiles(stubClient(http.StatusOK, oneGGUFBody), "unsloth/Real-GGUF")
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(unavailable).To(BeFalse())
|
||||
Expect(files).To(HaveLen(1))
|
||||
Expect(files[0].Name).To(Equal("Model-APEX-I-Quality.gguf"))
|
||||
Expect(files[0].SHA256).To(Equal("aa"))
|
||||
})
|
||||
|
||||
// Tolerating 401/403 must not widen into tolerating everything: a 500 is a
|
||||
// broken API, not evidence about whether the repo exists.
|
||||
It("still errors on a 500", func() {
|
||||
_, _, err := FetchOptionalRepoFiles(stubClient(http.StatusInternalServerError, ""), "unsloth/Real-GGUF")
|
||||
|
||||
Expect(err).To(HaveOccurred())
|
||||
Expect(err.Error()).To(ContainSubstring("unexpected status 500"))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("FetchRepoFiles", func() {
|
||||
// The APEX repo itself is not optional. A 401 there means the repo the run
|
||||
// was asked to publish cannot be read, which is a real failure and must not
|
||||
// be quietly downgraded to "no files".
|
||||
It("errors on a 401 for a required repo", func() {
|
||||
_, err := FetchRepoFiles(stubClient(http.StatusUnauthorized, ""), "mudler/Model-APEX-GGUF")
|
||||
|
||||
Expect(err).To(HaveOccurred())
|
||||
Expect(err.Error()).To(ContainSubstring("unexpected status 401"))
|
||||
})
|
||||
|
||||
It("errors on a 403 for a required repo", func() {
|
||||
_, err := FetchRepoFiles(stubClient(http.StatusForbidden, ""), "mudler/Model-APEX-GGUF")
|
||||
|
||||
Expect(err).To(HaveOccurred())
|
||||
Expect(err.Error()).To(ContainSubstring("unexpected status 403"))
|
||||
})
|
||||
|
||||
It("still treats a 404 as an absent repo", func() {
|
||||
files, err := FetchRepoFiles(stubClient(http.StatusNotFound, ""), "mudler/Model-APEX-GGUF")
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(files).To(BeEmpty())
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("ParseRepoFiles", func() {
|
||||
It("returns gguf siblings with their lfs sha256", func() {
|
||||
body := []byte(`{"siblings":[
|
||||
{"rfilename":"Model-APEX-I-Quality.gguf","size":10,"lfs":{"sha256":"aa","size":10}},
|
||||
{"rfilename":"README.md"},
|
||||
{"rfilename":"mmproj.gguf","size":5,"lfs":{"sha256":"bb","size":5}}
|
||||
]}`)
|
||||
|
||||
files, err := ParseRepoFiles(body)
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(files).To(HaveLen(2))
|
||||
Expect(files[0].Name).To(Equal("Model-APEX-I-Quality.gguf"))
|
||||
Expect(files[0].SHA256).To(Equal("aa"))
|
||||
Expect(files[1].Name).To(Equal("mmproj.gguf"))
|
||||
})
|
||||
|
||||
It("reports a gguf that carries no lfs sha256", func() {
|
||||
body := []byte(`{"siblings":[{"rfilename":"mmproj.gguf","size":5}]}`)
|
||||
|
||||
_, err := ParseRepoFiles(body)
|
||||
|
||||
Expect(err).To(MatchError(ErrNoSHA256))
|
||||
})
|
||||
})
|
||||
143
.github/ci/apexentries/hub.go
vendored
143
.github/ci/apexentries/hub.go
vendored
@@ -1,143 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"strings"
|
||||
|
||||
"gopkg.in/yaml.v3"
|
||||
|
||||
"github.com/mudler/LocalAI/.github/ci/galleryedit"
|
||||
)
|
||||
|
||||
// IndexText is the gallery index seen as text: the entries it declares plus the
|
||||
// exact lines each one occupies, which is what splicing a variants block into an
|
||||
// entry the gallery already ships requires.
|
||||
//
|
||||
// It is a second, narrower read of the same file LoadExisting parses. The two
|
||||
// answer different questions: LoadExisting answers "do these weights already
|
||||
// exist anywhere", this one answers "where in the file does this entry live".
|
||||
type IndexText struct {
|
||||
Lines []string
|
||||
Entries []*indexEntry
|
||||
|
||||
byName map[string]*indexEntry
|
||||
}
|
||||
|
||||
// indexEntry is one entry of the index: its name, the variants it already
|
||||
// declares, and its coordinates in the file.
|
||||
type indexEntry struct {
|
||||
Name string `yaml:"name"`
|
||||
Variants []VariantRef `yaml:"variants"`
|
||||
|
||||
Pos galleryedit.Entry `yaml:"-"`
|
||||
}
|
||||
|
||||
// LoadIndexText reads the gallery index for editing.
|
||||
func LoadIndexText(path string) (*IndexText, error) {
|
||||
raw, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return ParseIndexText(string(raw))
|
||||
}
|
||||
|
||||
// ParseIndexText pairs the decoded entries with the top level list items the
|
||||
// text actually contains.
|
||||
//
|
||||
// If the two views disagree on how many entries there are then every line number
|
||||
// a splice would compute is suspect, and the failure mode is writing a variants
|
||||
// block into the wrong model. The parse refuses instead.
|
||||
func ParseIndexText(text string) (*IndexText, error) {
|
||||
var entries []*indexEntry
|
||||
if err := yaml.Unmarshal([]byte(text), &entries); err != nil {
|
||||
return nil, fmt.Errorf("decoding gallery index: %w", err)
|
||||
}
|
||||
|
||||
lines, starts := galleryedit.Scan(text)
|
||||
if len(starts) != len(entries) {
|
||||
return nil, fmt.Errorf("gallery index has %d decoded entries but %d top level list items; refusing to edit by line number",
|
||||
len(entries), len(starts))
|
||||
}
|
||||
|
||||
ix := &IndexText{Lines: lines, Entries: entries, byName: map[string]*indexEntry{}}
|
||||
for i, e := range entries {
|
||||
if e == nil {
|
||||
return nil, fmt.Errorf("gallery index list item %d is empty; refusing to edit by line number", i)
|
||||
}
|
||||
end := len(lines)
|
||||
if i+1 < len(starts) {
|
||||
end = starts[i+1]
|
||||
}
|
||||
e.Pos = galleryedit.Entry{Name: e.Name, StartLine: starts[i], EndLine: end}
|
||||
|
||||
// First occurrence wins, matching the gallery's own resolution.
|
||||
key := strings.ToLower(e.Name)
|
||||
if _, seen := ix.byName[key]; !seen {
|
||||
ix.byName[key] = e
|
||||
}
|
||||
}
|
||||
return ix, nil
|
||||
}
|
||||
|
||||
// Find looks an entry up by name, case insensitively.
|
||||
func (ix *IndexText) Find(name string) *indexEntry {
|
||||
return ix.byName[strings.ToLower(name)]
|
||||
}
|
||||
|
||||
// ResolveHub returns the gallery name of a family's hub and whether the gallery
|
||||
// already ships an entry under it.
|
||||
//
|
||||
// The hub is the BASE model entry, never a generated *-apex parent. Somebody
|
||||
// looking for qwen3.6-35b-a3b has to find every build of those weights under
|
||||
// that one name: the APEX imatrix rungs, the unsloth quant rungs and any
|
||||
// speculative build. A separate qwen3.6-35b-a3b-apex hub competing with the base
|
||||
// entry would split the family in two and leave whichever half the user did not
|
||||
// search for invisible.
|
||||
//
|
||||
// Both candidates are tried for the same reason CounterpartCandidates tries
|
||||
// both. The repo name and the published file stem disagree for several of these
|
||||
// repos, and either one may be what the base entry was named after.
|
||||
func ResolveHub(ix *IndexText, repoBase, stem string) (name string, exists bool) {
|
||||
candidates := CounterpartCandidates(repoBase, stem)
|
||||
for _, c := range candidates {
|
||||
if n := slug(c); ix.Find(n) != nil {
|
||||
return n, true
|
||||
}
|
||||
}
|
||||
// Nothing matched, so the family needs a hub of its own under the repo
|
||||
// derived name, which is the more reliable of the two.
|
||||
return slug(candidates[0]), false
|
||||
}
|
||||
|
||||
// HubLabel is the human-cased base model name, for prose rather than lookup.
|
||||
func HubLabel(repoBase, stem string) string {
|
||||
return CounterpartCandidates(repoBase, stem)[0]
|
||||
}
|
||||
|
||||
// filterVariants drops the references a hub must not carry: itself, and anything
|
||||
// it already lists.
|
||||
//
|
||||
// The self reference is not merely redundant. A hub that names itself makes the
|
||||
// verifier resolve the reference back to the hub, see that the hub declares
|
||||
// variants, and report a variant that declares variants of its own. It arises
|
||||
// for real rather than in theory: an unsloth rung whose weights the gallery
|
||||
// already ships under the base model name resolves, through Merge, straight back
|
||||
// to the hub that is about to reference it.
|
||||
func filterVariants(hub string, already []VariantRef, want []string) []string {
|
||||
seen := map[string]bool{strings.ToLower(hub): true}
|
||||
for _, v := range already {
|
||||
seen[strings.ToLower(v.Model)] = true
|
||||
}
|
||||
|
||||
var out []string
|
||||
for _, w := range want {
|
||||
key := strings.ToLower(w)
|
||||
if seen[key] {
|
||||
continue
|
||||
}
|
||||
seen[key] = true
|
||||
out = append(out, w)
|
||||
}
|
||||
return out
|
||||
}
|
||||
795
.github/ci/apexentries/main.go
vendored
795
.github/ci/apexentries/main.go
vendored
@@ -1,795 +0,0 @@
|
||||
// Command apexentries generates gallery entries for the mudler APEX GGUF
|
||||
// repositories: one entry per imatrix tier and per unsloth quant rung, all
|
||||
// gathered under the BASE model's entry. Builds off a *-APEX-MTP-GGUF repo turn
|
||||
// speculative decoding on, because those weights retain the model's MTP heads
|
||||
// and are only worth their extra size with the heads in use.
|
||||
//
|
||||
// The base model entry is the hub. Somebody looking for qwen3.6-35b-a3b must
|
||||
// find every build of those weights under that one name, so when the gallery
|
||||
// already ships the base entry this command splices a variants block into it
|
||||
// rather than emitting a competing *-apex parent beside it. Only a family whose
|
||||
// base model the gallery does not ship at all gets a new hub entry, and that one
|
||||
// is still named for the base model.
|
||||
//
|
||||
// Builds are discovered by inspecting the filenames a repo actually publishes.
|
||||
// Repo names do not reliably predict them: mudler/gemma-4-26B-A4B-it-APEX-GGUF
|
||||
// ships gemma-4-26B-A4B-APEX-*.gguf, and six of the 45 repos drop a suffix or a
|
||||
// vendor prefix in the same way.
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"net/http"
|
||||
"os"
|
||||
"path"
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"gopkg.in/yaml.v3"
|
||||
|
||||
"github.com/mudler/LocalAI/.github/ci/galleryedit"
|
||||
)
|
||||
|
||||
const (
|
||||
// entryTemplate carries no backend and no parameters of its own, which is
|
||||
// why RenderChild states everything inline.
|
||||
entryTemplate = "virtual.yaml"
|
||||
unslothOwner = "unsloth"
|
||||
authorListURL = "https://huggingface.co/api/models?author=mudler&limit=300"
|
||||
)
|
||||
|
||||
// rungRank orders the quality ladder from best to smallest. The HuggingFace API
|
||||
// returns siblings alphabetically and DiscoverAPEXTiers preserves that order, so
|
||||
// an unsorted variants list reads I-Balanced, I-Compact, I-Mini, I-Nano,
|
||||
// I-Quality. Selection ignores authored order, so this is purely so the file a
|
||||
// human reviews scans in a meaningful sequence.
|
||||
var rungRank = map[string]int{
|
||||
"I-Quality": 0, "I-Balanced": 1, "I-Compact": 2, "I-Mini": 3, "I-Nano": 4,
|
||||
"Quality": 5, "Balanced": 6, "Compact": 7, "Mini": 8, "Nano": 9,
|
||||
}
|
||||
|
||||
// baseTags are the tags every generated entry carries. dflash and mtp are never
|
||||
// among them: RenderChild adds those if and only if the entry configures the
|
||||
// matching spec_type.
|
||||
var baseTags = []string{"llm", "gguf", "cpu", "gpu"}
|
||||
|
||||
// childBuild pairs a rendered entry with its position on the quality ladder, so
|
||||
// the parent's variants list can be sorted without re-parsing entry names.
|
||||
type childBuild struct {
|
||||
entry GalleryEntry
|
||||
rank int
|
||||
}
|
||||
|
||||
// family is one APEX repo's full generated output.
|
||||
type family struct {
|
||||
repo string
|
||||
repoBase string
|
||||
stem string
|
||||
hasMMProj bool
|
||||
children []childBuild
|
||||
|
||||
// skippedRepos are counterpart candidates HuggingFace would not describe.
|
||||
// Carried on the family rather than printed and forgotten so the run can
|
||||
// summarize them next to everything else a reviewer has to eyeball.
|
||||
skippedRepos []string
|
||||
census fileCensus
|
||||
unaccounted int
|
||||
}
|
||||
|
||||
// fileCensus splits the files discovery emitted nothing for into the ones a
|
||||
// reviewer must chase and the ones that are deliberately out of scope.
|
||||
//
|
||||
// Full-precision sources are the second kind: they are the unquantized weights
|
||||
// the ladder is derived FROM, not a rung of it. Folding them into the
|
||||
// unclassified total would leave a permanent benign baseline, and a permanent
|
||||
// baseline is exactly what hides the one file that ever genuinely matters.
|
||||
type fileCensus struct {
|
||||
unclassified int
|
||||
fullPrecision int
|
||||
}
|
||||
|
||||
// add accumulates one repo's census into a running total.
|
||||
func (c *fileCensus) add(o fileCensus) {
|
||||
c.unclassified += o.unclassified
|
||||
c.fullPrecision += o.fullPrecision
|
||||
}
|
||||
|
||||
// sortedChildren returns the family's builds in ladder order, best first.
|
||||
func (f *family) sortedChildren() []childBuild {
|
||||
sorted := append([]childBuild{}, f.children...)
|
||||
sort.SliceStable(sorted, func(i, j int) bool { return sorted[i].rank < sorted[j].rank })
|
||||
return sorted
|
||||
}
|
||||
|
||||
func main() {
|
||||
verify := flag.String("verify", "", "verify a gallery index and exit")
|
||||
index := flag.String("index", "gallery/index.yaml", "gallery index to dedup against")
|
||||
only := flag.String("only", "", "comma-separated repo names to restrict generation to")
|
||||
out := flag.String("out", "", "write the entries to add to this file")
|
||||
apply := flag.Bool("apply", false, "append the entries to add to -index")
|
||||
flag.Parse()
|
||||
|
||||
if *verify != "" {
|
||||
problems := Verify(*verify)
|
||||
for _, p := range problems {
|
||||
fmt.Fprintln(os.Stderr, p)
|
||||
}
|
||||
if len(problems) > 0 {
|
||||
fmt.Fprintf(os.Stderr, "%d problem(s)\n", len(problems))
|
||||
os.Exit(1)
|
||||
}
|
||||
fmt.Println("index is sound")
|
||||
return
|
||||
}
|
||||
|
||||
if err := generate(*index, *only, *out, *apply); err != nil {
|
||||
fmt.Fprintln(os.Stderr, "error:", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
func generate(indexPath, only, outPath string, apply bool) error {
|
||||
if outPath == "" && !apply {
|
||||
return fmt.Errorf("nothing to do: pass -out <file> or -apply")
|
||||
}
|
||||
|
||||
client := newHTTPClient()
|
||||
|
||||
repos, err := listAPEXRepos(client)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if only != "" {
|
||||
repos = restrict(repos, only)
|
||||
}
|
||||
if len(repos) == 0 {
|
||||
return fmt.Errorf("no APEX repos selected")
|
||||
}
|
||||
fmt.Printf("repos selected: %d\n", len(repos))
|
||||
|
||||
var families []family
|
||||
var failed []string
|
||||
|
||||
for _, repo := range repos {
|
||||
f, err := buildFamily(client, repo)
|
||||
if err != nil {
|
||||
// A missing sha256 is fatal for the family rather than skippable: an
|
||||
// entry without one ships an unverifiable download. Report which repo
|
||||
// and keep going, so one bad repo does not hide the state of the rest.
|
||||
fmt.Fprintf(os.Stderr, "FAILED %s: %v\n", repo, err)
|
||||
failed = append(failed, repo)
|
||||
continue
|
||||
}
|
||||
families = append(families, *f)
|
||||
}
|
||||
|
||||
existing, err := LoadExisting(indexPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
ixText, err := LoadIndexText(indexPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
fmt.Printf("existing index: %d names, %d weight URIs, %d lines\n",
|
||||
len(existing.ByName), len(existing.ByURI), len(ixText.Lines))
|
||||
|
||||
// Only the builds go through Merge. A hub is deliberately kept out of it: a
|
||||
// new hub carries the family's top rung as its own payload, so Merge's URI
|
||||
// dedup would fold the hub into that rung and the family would lose the very
|
||||
// entry point this command exists to create. Hub names are checked against
|
||||
// the index directly, by ResolveHub.
|
||||
var generated []GalleryEntry
|
||||
for _, f := range families {
|
||||
for _, c := range f.children {
|
||||
generated = append(generated, c.entry)
|
||||
}
|
||||
}
|
||||
|
||||
add, reused := Merge(existing, generated)
|
||||
reportReuse(existing, generated, reused)
|
||||
|
||||
// Variant references are resolved from `reused`, never used to decide what to
|
||||
// emit: on a within-batch name collision Merge records reused[name] = name
|
||||
// while the first entry of that name is still in `add`, so treating presence
|
||||
// in `reused` as "dropped" would silently emit nothing for it.
|
||||
added := map[string]bool{}
|
||||
for _, e := range add {
|
||||
added[e.Name] = true
|
||||
}
|
||||
|
||||
inserts, newHubs, err := planHubs(families, ixText, reused, added)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
reportHubs(ixText, inserts, newHubs)
|
||||
|
||||
skipped, census, fullPrecisionRepos, unaccounted := reportSkipped(families)
|
||||
|
||||
add = append(add, newHubs...)
|
||||
|
||||
fmt.Printf("\nentries generated: %d\nentries to add: %d\nentries reused: %d\nhubs spliced: %d\nhubs created: %d\nrepos skipped: %d\nexcluded (full precision): %d files across %d repos\nunclassified: %d\nunaccounted: %d\n",
|
||||
len(generated), len(add), len(reused), len(inserts), len(newHubs), len(skipped),
|
||||
census.fullPrecision, fullPrecisionRepos, census.unclassified, unaccounted)
|
||||
|
||||
lines, err := galleryedit.Apply(ixText.Lines, inserts)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
if err := writeEntries(add, lines, outPath, apply, indexPath); err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
if len(failed) > 0 {
|
||||
return fmt.Errorf("%d repo(s) failed: %s", len(failed), strings.Join(failed, ", "))
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// resolveVariant maps a generated child name onto whatever entry actually stands
|
||||
// for it after the merge. `added` is consulted first because a within-batch name
|
||||
// collision puts a name in BOTH add and reused, and the entry that was emitted
|
||||
// is the one the parent must reference.
|
||||
func resolveVariant(name string, reused map[string]string, added map[string]bool) string {
|
||||
if added[name] {
|
||||
return name
|
||||
}
|
||||
if target, ok := reused[name]; ok {
|
||||
return target
|
||||
}
|
||||
return name
|
||||
}
|
||||
|
||||
// SpecTypeForRepo reports the speculative decoding mechanism a repo's builds can
|
||||
// turn on with no extra download.
|
||||
//
|
||||
// The *-APEX-MTP-GGUF repos republish the base weights with the model's own MTP
|
||||
// heads retained, so those builds are only worth their extra size if the heads
|
||||
// are actually used. Every other APEX repo drops them, and switching MTP on
|
||||
// there would name a mechanism the weights cannot serve.
|
||||
//
|
||||
// The suffix is read off the repo the FILES come from, so nothing downstream has
|
||||
// to infer a capability from an entry name.
|
||||
func SpecTypeForRepo(repo string) string {
|
||||
if strings.HasSuffix(path.Base(repo), "-APEX-MTP-GGUF") {
|
||||
return "draft-mtp"
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// buildFamily discovers everything one APEX repo and its unsloth counterpart
|
||||
// publish, and renders it.
|
||||
func buildFamily(client *http.Client, repo string) (*family, error) {
|
||||
files, err := FetchRepoFiles(client, repo)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if len(files) == 0 {
|
||||
return nil, fmt.Errorf("no gguf files")
|
||||
}
|
||||
|
||||
imatrix, plain := DiscoverAPEXTiers(files)
|
||||
mmproj, hasMMProj := DiscoverMMProj(files)
|
||||
|
||||
census := reportUnclassified(repo, files, imatrix, plain)
|
||||
|
||||
// The imatrix ladder is preferred, but two of the 45 repos publish no
|
||||
// imatrix tiers at all and must still contribute their plain ladder.
|
||||
ladder := imatrix
|
||||
ladderKind := "imatrix"
|
||||
if len(ladder) == 0 {
|
||||
ladder = plain
|
||||
ladderKind = "plain"
|
||||
}
|
||||
if len(ladder) == 0 {
|
||||
return nil, fmt.Errorf("no tiers discovered")
|
||||
}
|
||||
|
||||
sortTiers(ladder)
|
||||
|
||||
var mm *GGUFFile
|
||||
if hasMMProj {
|
||||
mm = &mmproj
|
||||
}
|
||||
|
||||
repoBase := strings.TrimSuffix(path.Base(repo), "-GGUF")
|
||||
f := &family{repo: repo, repoBase: repoBase, hasMMProj: hasMMProj, census: census}
|
||||
|
||||
// Only the APEX ladder can carry MTP heads; the unsloth counterpart quantizes
|
||||
// the plain weights and gets nothing from this.
|
||||
specType := SpecTypeForRepo(repo)
|
||||
|
||||
for _, t := range ladder {
|
||||
f.children = append(f.children, childBuild{
|
||||
rank: rungRank[t.Label],
|
||||
entry: RenderChild(ChildInput{
|
||||
Name: slug(repoBase) + "-" + slug(t.Label),
|
||||
Repo: repo,
|
||||
Template: entryTemplate,
|
||||
SpecType: specType,
|
||||
Weights: []GGUFFile{t.File},
|
||||
MMProj: mm,
|
||||
BaseTags: baseTags,
|
||||
}),
|
||||
})
|
||||
}
|
||||
|
||||
stem := FileStem(ladder[0])
|
||||
f.stem = stem
|
||||
fmt.Printf("%s: %d %s tier(s) [%s], stem %s, mmproj %v\n",
|
||||
repo, len(ladder), ladderKind, tierLabels(ladder), stem, hasMMProj)
|
||||
|
||||
counterpart, cpFiles, skipped, err := resolveCounterpart(client, repoBase, stem)
|
||||
f.skippedRepos = skipped
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if counterpart != "" {
|
||||
builds := DiscoverUnslothQuants(cpFiles)
|
||||
|
||||
// Called here rather than inside Verify: a quant dropped at discovery
|
||||
// leaves no trace at all in the finished gallery file, so the only place
|
||||
// the shortfall is still visible is the moment of discovery.
|
||||
unaccounted := UnaccountedQuants(cpFiles, builds)
|
||||
f.unaccounted = len(unaccounted)
|
||||
for _, p := range unaccounted {
|
||||
fmt.Fprintf(os.Stderr, "UNACCOUNTED QUANT %s: %s\n", counterpart, p)
|
||||
}
|
||||
|
||||
cpMMProj, hasCPMMProj := DiscoverMMProj(cpFiles)
|
||||
var cpMM *GGUFFile
|
||||
if hasCPMMProj {
|
||||
cpMM = &cpMMProj
|
||||
}
|
||||
cpBase := strings.TrimSuffix(path.Base(counterpart), "-GGUF")
|
||||
|
||||
for i, b := range builds {
|
||||
f.children = append(f.children, childBuild{
|
||||
rank: 100 + i,
|
||||
entry: RenderChild(ChildInput{
|
||||
Name: slug(cpBase) + "-" + slug(b.Quant),
|
||||
Repo: counterpart,
|
||||
Template: entryTemplate,
|
||||
Weights: b.Files,
|
||||
MMProj: cpMM,
|
||||
BaseTags: baseTags,
|
||||
}),
|
||||
})
|
||||
}
|
||||
fmt.Printf("%s: counterpart %s, %d quant build(s) %s\n", repo, counterpart, len(builds), quantLabels(builds))
|
||||
} else {
|
||||
fmt.Printf("%s: no unsloth counterpart\n", repo)
|
||||
}
|
||||
|
||||
return f, nil
|
||||
}
|
||||
|
||||
// planHubs decides, per family, whether the family's builds are spliced into a
|
||||
// base model entry the gallery already ships or gathered under a new hub.
|
||||
//
|
||||
// Splicing is strongly preferred and is the measured majority-adjacent case. The
|
||||
// existing entry keeps its description, icon, tags, overrides and files
|
||||
// untouched; only variant lines are added to it.
|
||||
func planHubs(families []family, ix *IndexText, reused map[string]string, added map[string]bool) ([]galleryedit.Insert, []GalleryEntry, error) {
|
||||
// Several APEX repos can resolve to one base model, so both paths accumulate
|
||||
// by hub name rather than assuming one family per hub.
|
||||
wantByHub := map[string][]string{}
|
||||
var spliceOrder []string
|
||||
|
||||
var newHubs []GalleryEntry
|
||||
hubAt := map[string]int{}
|
||||
|
||||
for i := range families {
|
||||
f := &families[i]
|
||||
|
||||
hubName, exists := ResolveHub(ix, f.repoBase, f.stem)
|
||||
want := hubVariants(f, ix, reused, added)
|
||||
|
||||
if exists {
|
||||
if _, seen := wantByHub[hubName]; !seen {
|
||||
spliceOrder = append(spliceOrder, hubName)
|
||||
}
|
||||
wantByHub[hubName] = append(wantByHub[hubName], want...)
|
||||
continue
|
||||
}
|
||||
|
||||
if at, dup := hubAt[hubName]; dup {
|
||||
for _, v := range filterVariants(hubName, newHubs[at].Variants, want) {
|
||||
newHubs[at].Variants = append(newHubs[at].Variants, VariantRef{Model: v})
|
||||
}
|
||||
continue
|
||||
}
|
||||
|
||||
builds := f.sortedChildren()
|
||||
if len(builds) == 0 {
|
||||
return nil, nil, fmt.Errorf("%s: no builds to hang a hub on", f.repo)
|
||||
}
|
||||
hubAt[hubName] = len(newHubs)
|
||||
newHubs = append(newHubs, renderHub(hubName, f, builds[0], filterVariants(hubName, nil, want)))
|
||||
}
|
||||
|
||||
var inserts []galleryedit.Insert
|
||||
for _, name := range spliceOrder {
|
||||
e := ix.Find(name)
|
||||
items := filterVariants(name, e.Variants, wantByHub[name])
|
||||
if len(items) == 0 {
|
||||
continue
|
||||
}
|
||||
inserts = append(inserts, galleryedit.Insert{Entry: e.Pos, Variants: items})
|
||||
}
|
||||
return inserts, newHubs, nil
|
||||
}
|
||||
|
||||
// hubVariants is a family's full build list, in ladder order, named as the hub
|
||||
// must reference them after the merge.
|
||||
func hubVariants(f *family, ix *IndexText, reused map[string]string, added map[string]bool) []string {
|
||||
var out []string
|
||||
|
||||
// A hand-written *-apex entry is an ordinary build of these weights. It is
|
||||
// never deleted, never renamed and never treated as a hub; it is simply
|
||||
// referenced like any other rung.
|
||||
if apex := slug(f.repoBase); ix.Find(apex) != nil {
|
||||
out = append(out, apex)
|
||||
}
|
||||
for _, c := range f.sortedChildren() {
|
||||
out = append(out, resolveVariant(c.entry.Name, reused, added))
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// renderHub builds the hub for a family whose base model the gallery does not
|
||||
// ship at all. It is named for the BASE model, never for the APEX repo.
|
||||
//
|
||||
// It carries one of the discovered builds as its own payload so it is a complete
|
||||
// installable entry rather than a bare index pointing at other entries. That
|
||||
// payload is what supplies overrides.backend, which matters beyond installation:
|
||||
// the verifier can only judge the tagging rule for a backend it can read, so a
|
||||
// hub carrying feature tags and no backend would escape the check in silence.
|
||||
//
|
||||
// The payload's own tags are kept rather than rebuilt from baseTags, so a hub
|
||||
// whose payload configures a spec_type stays tagged for it and consistent with
|
||||
// the overrides copied alongside.
|
||||
func renderHub(name string, f *family, payload childBuild, variants []string) GalleryEntry {
|
||||
e := payload.entry
|
||||
e.Name = name
|
||||
e.Description = fmt.Sprintf(
|
||||
"%s. Quality ladder and quantization rungs published by %s and its unsloth counterpart; LocalAI picks the build that fits the hardware.",
|
||||
HubLabel(f.repoBase, f.stem), f.repo)
|
||||
|
||||
e.Tags = append([]string{}, payload.entry.Tags...)
|
||||
if f.hasMMProj && !hasTag(e.Tags, "vision") {
|
||||
e.Tags = append(e.Tags, "vision")
|
||||
}
|
||||
|
||||
e.Variants = nil
|
||||
for _, v := range variants {
|
||||
e.Variants = append(e.Variants, VariantRef{Model: v})
|
||||
}
|
||||
return e
|
||||
}
|
||||
|
||||
func hasTag(tags []string, want string) bool {
|
||||
for _, t := range tags {
|
||||
if t == want {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// resolveCounterpart probes the unsloth candidates in order and returns the
|
||||
// first that publishes files.
|
||||
//
|
||||
// CounterpartCandidates is handed a BARE repo name: its cleaner does not strip
|
||||
// an owner prefix, so passing "mudler/Foo-APEX-GGUF" would yield "mudler/Foo"
|
||||
// and compose into the nonsense probe "unsloth/mudler/Foo".
|
||||
//
|
||||
// It also returns the candidates HuggingFace refused to describe. Those are
|
||||
// indistinguishable from absent without credentials, so they are skipped, but
|
||||
// they are named rather than dropped: one of them could be a real gated repo
|
||||
// whose quants belong in the gallery.
|
||||
func resolveCounterpart(client *http.Client, repoBase, stem string) (string, []GGUFFile, []string, error) {
|
||||
var unavailable []string
|
||||
for _, cand := range CounterpartCandidates(repoBase, stem) {
|
||||
repo := unslothOwner + "/" + cand + "-GGUF"
|
||||
files, unreadable, err := FetchOptionalRepoFiles(client, repo)
|
||||
if err != nil {
|
||||
return "", nil, unavailable, fmt.Errorf("probing %s: %w", repo, err)
|
||||
}
|
||||
if unreadable {
|
||||
unavailable = append(unavailable, repo)
|
||||
continue
|
||||
}
|
||||
if len(files) > 0 {
|
||||
return repo, files, unavailable, nil
|
||||
}
|
||||
}
|
||||
return "", nil, unavailable, nil
|
||||
}
|
||||
|
||||
// reportUnclassified prints the files discovery turned into nothing.
|
||||
//
|
||||
// It is a set difference on COUNTS, not a re-match of filenames: re-matching
|
||||
// would duplicate the tier regex from discover.go and the two copies would
|
||||
// drift. The likeliest trigger is a typo or case change from a publishing script
|
||||
// rather than a genuine sixth tier, and because generation falls back to the
|
||||
// plain ladder when the imatrix one is empty, a repo whose imatrix files all
|
||||
// fail to match silently downgrades the whole family instead of erroring. The
|
||||
// downstream HTTP check cannot catch that: it validates URLs that were emitted,
|
||||
// and an undiscovered tier emits none.
|
||||
// It returns the census so the run can total it.
|
||||
func reportUnclassified(repo string, files []GGUFFile, imatrix, plain []Tier) fileCensus {
|
||||
mmprojCount, fullPrecision := 0, 0
|
||||
for _, f := range files {
|
||||
// The mmproj test comes first because projectors are themselves often
|
||||
// published at f16 (mmproj-F16.gguf), and counting such a file in both
|
||||
// buckets would understate the unclassified remainder.
|
||||
if strings.HasPrefix(f.Name, "mmproj") {
|
||||
mmprojCount++
|
||||
continue
|
||||
}
|
||||
if IsFullPrecision(f.Name) {
|
||||
fullPrecision++
|
||||
}
|
||||
}
|
||||
|
||||
classified := len(imatrix) + len(plain) + mmprojCount + fullPrecision
|
||||
if classified >= len(files) {
|
||||
return fileCensus{fullPrecision: fullPrecision}
|
||||
}
|
||||
fmt.Fprintf(os.Stderr, "UNCLASSIFIED %s: %d of %d .gguf files classified, %d unaccounted for\n",
|
||||
repo, classified, len(files), len(files)-classified)
|
||||
return fileCensus{unclassified: len(files) - classified, fullPrecision: fullPrecision}
|
||||
}
|
||||
|
||||
// reportReuse splits Merge's single reused map into the two cases it conflates.
|
||||
//
|
||||
// A URI match means the gallery already ships exactly these weights, and
|
||||
// pointing the parent at the existing entry is correct. A NAME match with a
|
||||
// different URI means an unrelated entry happens to own the name, and
|
||||
// referencing it would point the parent at different weights than were
|
||||
// generated, substituting a build without saying so. Only the first is safe to
|
||||
// wave through.
|
||||
func reportReuse(existing *ExistingIndex, generated []GalleryEntry, reused map[string]string) {
|
||||
byName := map[string]GalleryEntry{}
|
||||
for _, e := range generated {
|
||||
if _, seen := byName[e.Name]; !seen {
|
||||
byName[e.Name] = e
|
||||
}
|
||||
}
|
||||
|
||||
var nameCollisions, uriMatches []string
|
||||
for name, target := range reused {
|
||||
gen := byName[name]
|
||||
uri := ""
|
||||
if len(gen.Files) > 0 {
|
||||
uri = gen.Files[0].URI
|
||||
}
|
||||
|
||||
switch {
|
||||
case hasName(existing, name):
|
||||
nameCollisions = append(nameCollisions,
|
||||
fmt.Sprintf(" %s -> gallery entry of the same name (generated uri: %s)", name, orNone(uri)))
|
||||
case target == name:
|
||||
nameCollisions = append(nameCollisions,
|
||||
fmt.Sprintf(" %s -> earlier entry of the same name in this batch (generated uri: %s)", name, orNone(uri)))
|
||||
default:
|
||||
uriMatches = append(uriMatches, fmt.Sprintf(" %s -> %s (same weights: %s)", name, target, orNone(uri)))
|
||||
}
|
||||
}
|
||||
sort.Strings(nameCollisions)
|
||||
sort.Strings(uriMatches)
|
||||
|
||||
fmt.Printf("\nNAME COLLISIONS (%d) - inspect each by hand, the target may hold different weights\n", len(nameCollisions))
|
||||
for _, l := range nameCollisions {
|
||||
fmt.Println(l)
|
||||
}
|
||||
fmt.Printf("\nURI MATCHES (%d) - the gallery or this batch already ships these exact weights\n", len(uriMatches))
|
||||
for _, l := range uriMatches {
|
||||
fmt.Println(l)
|
||||
}
|
||||
}
|
||||
|
||||
// reportHubs prints exactly what will be written where. The splices are the part
|
||||
// a human has to read: they modify entries the gallery already ships, so the
|
||||
// review needs the target, the line, and every added reference spelled out.
|
||||
func reportHubs(ix *IndexText, inserts []galleryedit.Insert, newHubs []GalleryEntry) {
|
||||
fmt.Printf("\nHUBS SPLICED (%d) - variants added to the EXISTING base model entry, nothing else touched\n", len(inserts))
|
||||
for _, in := range inserts {
|
||||
e := ix.Find(in.Entry.Name)
|
||||
fmt.Printf(" %s (line %d, %d variant(s) already declared):\n", in.Entry.Name, in.Entry.StartLine+1, len(e.Variants))
|
||||
for _, v := range in.Variants {
|
||||
fmt.Printf(" + - model: %s\n", galleryedit.QuoteName(v))
|
||||
}
|
||||
}
|
||||
|
||||
fmt.Printf("\nHUBS CREATED (%d) - the gallery ships no base model entry, so one is emitted for it\n", len(newHubs))
|
||||
for _, h := range newHubs {
|
||||
fmt.Printf(" %s:\n", h.Name)
|
||||
for _, v := range h.Variants {
|
||||
fmt.Printf(" - model: %s\n", v.Model)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// reportSkipped names the counterpart repos HuggingFace would not describe, and
|
||||
// totals the other two silent-shortfall counters alongside them.
|
||||
//
|
||||
// A skipped repo is not the same as a clean 404. HuggingFace answers 401 for a
|
||||
// nonexistent repo to an unauthenticated client, so the overwhelmingly likely
|
||||
// reading is "there is no such counterpart", which is the normal case for the
|
||||
// community merges. But a private or gated repo answers 401 too, and that one
|
||||
// WOULD have quants worth shipping. Printing the list is what keeps that
|
||||
// possibility auditable instead of silently discarded.
|
||||
func reportSkipped(families []family) ([]string, fileCensus, int, int) {
|
||||
var skipped []string
|
||||
var census fileCensus
|
||||
fullPrecisionRepos, unaccounted := 0, 0
|
||||
for _, f := range families {
|
||||
skipped = append(skipped, f.skippedRepos...)
|
||||
census.add(f.census)
|
||||
if f.census.fullPrecision > 0 {
|
||||
fullPrecisionRepos++
|
||||
}
|
||||
unaccounted += f.unaccounted
|
||||
}
|
||||
sort.Strings(skipped)
|
||||
|
||||
fmt.Printf("\nREPOS SKIPPED AS UNAVAILABLE (%d) - HuggingFace answered 401/403, which is indistinguishable from absent without a token; check none of these is a real gated repo\n", len(skipped))
|
||||
for _, r := range skipped {
|
||||
fmt.Printf(" %s\n", r)
|
||||
}
|
||||
return skipped, census, fullPrecisionRepos, unaccounted
|
||||
}
|
||||
|
||||
func hasName(ix *ExistingIndex, name string) bool {
|
||||
_, ok := ix.ByName[name]
|
||||
return ok
|
||||
}
|
||||
|
||||
func orNone(s string) string {
|
||||
if s == "" {
|
||||
return "(no files)"
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// writeEntries emits the additions.
|
||||
//
|
||||
// -apply does two things in one pass: it writes back the spliced lines, which
|
||||
// differ from the original only by the variant lines galleryedit inserted, and
|
||||
// then appends the new entries. New entries are APPENDED rather than merged into
|
||||
// the structure, for the same reason the splice is textual: a YAML round trip
|
||||
// over 40,000 lines would reflow the whole file into an unreviewable diff.
|
||||
func writeEntries(add []GalleryEntry, lines []string, outPath string, apply bool, indexPath string) error {
|
||||
if apply {
|
||||
if err := os.WriteFile(indexPath, []byte(strings.Join(lines, "\n")), 0o644); err != nil {
|
||||
return err
|
||||
}
|
||||
fmt.Printf("spliced %s\n", indexPath)
|
||||
}
|
||||
|
||||
if len(add) == 0 {
|
||||
fmt.Println("nothing to append")
|
||||
return nil
|
||||
}
|
||||
|
||||
blob, err := yaml.Marshal(add)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
if outPath != "" {
|
||||
if err := os.WriteFile(outPath, blob, 0o644); err != nil {
|
||||
return err
|
||||
}
|
||||
fmt.Printf("wrote %d entries to %s\n", len(add), outPath)
|
||||
}
|
||||
|
||||
if apply {
|
||||
f, err := os.OpenFile(indexPath, os.O_APPEND|os.O_WRONLY, 0o644)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer f.Close()
|
||||
if _, err := f.Write(blob); err != nil {
|
||||
return err
|
||||
}
|
||||
fmt.Printf("appended %d entries to %s\n", len(add), indexPath)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// listAPEXRepos returns the mudler repos whose name marks them as APEX builds.
|
||||
func listAPEXRepos(client *http.Client) ([]string, error) {
|
||||
req, err := http.NewRequest(http.MethodGet, authorListURL, nil)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
req.Header.Set("User-Agent", "localai-apexentries/1.0")
|
||||
|
||||
resp, err := client.Do(req)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
if resp.StatusCode != http.StatusOK {
|
||||
return nil, fmt.Errorf("listing models: unexpected status %d", resp.StatusCode)
|
||||
}
|
||||
|
||||
body, err := io.ReadAll(resp.Body)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
var models []struct {
|
||||
ID string `json:"id"`
|
||||
}
|
||||
if err := json.Unmarshal(body, &models); err != nil {
|
||||
return nil, fmt.Errorf("decoding model list: %w", err)
|
||||
}
|
||||
|
||||
var out []string
|
||||
for _, m := range models {
|
||||
if strings.Contains(m.ID, "APEX") {
|
||||
out = append(out, m.ID)
|
||||
}
|
||||
}
|
||||
sort.Strings(out)
|
||||
return out, nil
|
||||
}
|
||||
|
||||
func restrict(repos []string, only string) []string {
|
||||
want := map[string]bool{}
|
||||
for _, r := range strings.Split(only, ",") {
|
||||
if r = strings.TrimSpace(r); r != "" {
|
||||
want[r] = true
|
||||
}
|
||||
}
|
||||
|
||||
var out []string
|
||||
for _, r := range repos {
|
||||
if want[r] {
|
||||
out = append(out, r)
|
||||
delete(want, r)
|
||||
}
|
||||
}
|
||||
// A name in -only that matched nothing is a typo, not an empty result.
|
||||
for r := range want {
|
||||
fmt.Fprintf(os.Stderr, "WARNING: -only names %s, which is not an APEX repo of this author\n", r)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func sortTiers(tiers []Tier) {
|
||||
sort.SliceStable(tiers, func(i, j int) bool { return rungRank[tiers[i].Label] < rungRank[tiers[j].Label] })
|
||||
}
|
||||
|
||||
func tierLabels(tiers []Tier) string {
|
||||
var out []string
|
||||
for _, t := range tiers {
|
||||
out = append(out, t.Label)
|
||||
}
|
||||
return strings.Join(out, ",")
|
||||
}
|
||||
|
||||
func quantLabels(builds []QuantBuild) string {
|
||||
var out []string
|
||||
for _, b := range builds {
|
||||
l := b.Quant
|
||||
if b.Sharded {
|
||||
l += fmt.Sprintf("(%d shards)", len(b.Files))
|
||||
}
|
||||
out = append(out, l)
|
||||
}
|
||||
return strings.Join(out, ",")
|
||||
}
|
||||
|
||||
// slug turns a repo, tier or quant label into a gallery entry name component.
|
||||
func slug(s string) string {
|
||||
return strings.ReplaceAll(strings.ToLower(s), "_", "-")
|
||||
}
|
||||
344
.github/ci/apexentries/main_test.go
vendored
344
.github/ci/apexentries/main_test.go
vendored
@@ -1,344 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"strings"
|
||||
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
|
||||
"github.com/mudler/LocalAI/.github/ci/galleryedit"
|
||||
)
|
||||
|
||||
func mustIndex(text string) *IndexText {
|
||||
ix, err := ParseIndexText(text)
|
||||
ExpectWithOffset(1, err).ToNot(HaveOccurred())
|
||||
return ix
|
||||
}
|
||||
|
||||
// buildOf renders a realistic child so the specs exercise the payload a hub
|
||||
// actually inherits rather than a bare name.
|
||||
func buildOf(name, repo, file string, rank int) childBuild {
|
||||
return childBuild{
|
||||
rank: rank,
|
||||
entry: RenderChild(ChildInput{
|
||||
Name: name,
|
||||
Repo: repo,
|
||||
Template: entryTemplate,
|
||||
Weights: []GGUFFile{{Name: file, SHA256: "aa"}},
|
||||
BaseTags: baseTags,
|
||||
}),
|
||||
}
|
||||
}
|
||||
|
||||
var _ = Describe("ResolveHub", func() {
|
||||
It("picks the base model name over the APEX name, even when both are in the gallery", func() {
|
||||
// The hub is the entry a user searches for. If the *-apex entry were
|
||||
// chosen the family would be gathered under a name nobody looks up, and
|
||||
// the base entry would go on advertising only its own build.
|
||||
ix := mustIndex("- name: qwen3.6-35b-a3b\n url: u\n- name: qwen3.6-35b-a3b-apex\n url: u\n")
|
||||
|
||||
name, exists := ResolveHub(ix, "Qwen3.6-35B-A3B-APEX", "Qwen3.6-35B-A3B-APEX")
|
||||
|
||||
Expect(name).To(Equal("qwen3.6-35b-a3b"))
|
||||
Expect(exists).To(BeTrue())
|
||||
})
|
||||
|
||||
It("falls back to the stem-derived candidate when the repo-derived one is absent", func() {
|
||||
// gemma's repo says "-it" and its published files do not, so only one of
|
||||
// the two candidates can match whatever the base entry was named after.
|
||||
ix := mustIndex("- name: gemma-4-26b-a4b\n url: u\n")
|
||||
|
||||
name, exists := ResolveHub(ix, "gemma-4-26B-A4B-it-APEX", "gemma-4-26B-A4B-APEX")
|
||||
|
||||
Expect(name).To(Equal("gemma-4-26b-a4b"))
|
||||
Expect(exists).To(BeTrue())
|
||||
})
|
||||
|
||||
It("reports the base name as absent rather than settling for the APEX entry", func() {
|
||||
ix := mustIndex("- name: qwen3.5-35b-a3b-apex\n url: u\n")
|
||||
|
||||
name, exists := ResolveHub(ix, "Qwen3.5-35B-A3B-APEX", "Qwen3.5-35B-A3B-APEX")
|
||||
|
||||
Expect(name).To(Equal("qwen3.5-35b-a3b"))
|
||||
Expect(exists).To(BeFalse())
|
||||
})
|
||||
|
||||
It("strips the MTP and TQ markers as well as APEX", func() {
|
||||
ix := mustIndex("- name: qwen3.6-35b-a3b\n url: u\n")
|
||||
|
||||
name, exists := ResolveHub(ix, "Qwen3.6-35B-A3B-APEX-MTP", "Qwen3.6-35B-A3B-APEX-MTP")
|
||||
|
||||
Expect(name).To(Equal("qwen3.6-35b-a3b"))
|
||||
Expect(exists).To(BeTrue())
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("planHubs", func() {
|
||||
noReuse := map[string]string{}
|
||||
allAdded := func(names ...string) map[string]bool {
|
||||
out := map[string]bool{}
|
||||
for _, n := range names {
|
||||
out[n] = true
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
It("splices into the existing base entry instead of emitting an *-apex parent", func() {
|
||||
ix := mustIndex("- name: step-3.7-flash\n url: u\n- name: other\n url: u\n")
|
||||
fams := []family{{
|
||||
repo: "mudler/Step-3.7-Flash-APEX-GGUF",
|
||||
repoBase: "Step-3.7-Flash-APEX",
|
||||
stem: "Step-3.7-Flash-APEX",
|
||||
children: []childBuild{buildOf("step-3.7-flash-apex-i-quality", "mudler/Step-3.7-Flash-APEX-GGUF", "a.gguf", 0)},
|
||||
}}
|
||||
|
||||
inserts, newHubs, err := planHubs(fams, ix, noReuse, allAdded("step-3.7-flash-apex-i-quality"))
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(newHubs).To(BeEmpty())
|
||||
Expect(inserts).To(HaveLen(1))
|
||||
Expect(inserts[0].Entry.Name).To(Equal("step-3.7-flash"))
|
||||
Expect(inserts[0].Variants).To(Equal([]string{"step-3.7-flash-apex-i-quality"}))
|
||||
})
|
||||
|
||||
It("merges into an entry that already declares variants, without repeating one", func() {
|
||||
// The gallery's qwen3.6-35b-a3b already lists its APEX build. Re-adding it
|
||||
// would put a duplicate key's worth of noise in the diff and a duplicate
|
||||
// reference in the entry.
|
||||
ix := mustIndex("- name: qwen3.6-35b-a3b\n variants:\n - model: qwen3.6-35b-a3b-apex\n url: u\n" +
|
||||
"- name: qwen3.6-35b-a3b-apex\n url: u\n")
|
||||
fams := []family{{
|
||||
repo: "mudler/Qwen3.6-35B-A3B-APEX-GGUF",
|
||||
repoBase: "Qwen3.6-35B-A3B-APEX",
|
||||
stem: "Qwen3.6-35B-A3B-APEX",
|
||||
children: []childBuild{buildOf("qwen3.6-35b-a3b-apex-i-quality", "mudler/Qwen3.6-35B-A3B-APEX-GGUF", "a.gguf", 0)},
|
||||
}}
|
||||
|
||||
inserts, newHubs, err := planHubs(fams, ix, noReuse, allAdded("qwen3.6-35b-a3b-apex-i-quality"))
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(newHubs).To(BeEmpty())
|
||||
Expect(inserts[0].Variants).To(Equal([]string{"qwen3.6-35b-a3b-apex-i-quality"}))
|
||||
|
||||
out, err := galleryedit.Apply(ix.Lines, inserts)
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(strings.Count(strings.Join(out, "\n"), "variants:")).To(Equal(1))
|
||||
Expect(out).To(HaveLen(len(ix.Lines) + 1))
|
||||
})
|
||||
|
||||
It("never lets the hub reference itself", func() {
|
||||
// An unsloth rung whose weights the gallery already ships under the base
|
||||
// name resolves, through Merge, straight back to the hub. The verifier
|
||||
// reads a self reference as a variant that declares variants of its own.
|
||||
ix := mustIndex("- name: step-3.7-flash\n url: u\n")
|
||||
fams := []family{{
|
||||
repo: "mudler/Step-3.7-Flash-APEX-GGUF",
|
||||
repoBase: "Step-3.7-Flash-APEX",
|
||||
stem: "Step-3.7-Flash-APEX",
|
||||
children: []childBuild{buildOf("step-3.7-flash-ud-q4-k-m", "unsloth/Step-3.7-Flash-GGUF", "a.gguf", 100)},
|
||||
}}
|
||||
|
||||
inserts, _, err := planHubs(fams, ix, map[string]string{"step-3.7-flash-ud-q4-k-m": "step-3.7-flash"}, map[string]bool{})
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(inserts).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("emits a hub named for the base model when the gallery has none", func() {
|
||||
ix := mustIndex("- name: qwen3.5-35b-a3b-apex\n url: u\n")
|
||||
fams := []family{{
|
||||
repo: "mudler/Qwen3.5-35B-A3B-APEX-GGUF",
|
||||
repoBase: "Qwen3.5-35B-A3B-APEX",
|
||||
stem: "Qwen3.5-35B-A3B-APEX",
|
||||
hasMMProj: true,
|
||||
children: []childBuild{
|
||||
buildOf("qwen3.5-35b-a3b-apex-i-quality", "mudler/Qwen3.5-35B-A3B-APEX-GGUF", "a.gguf", 0),
|
||||
buildOf("qwen3.5-35b-a3b-ud-q6-k", "unsloth/Qwen3.5-35B-A3B-GGUF", "b.gguf", 102),
|
||||
},
|
||||
}}
|
||||
|
||||
inserts, newHubs, err := planHubs(fams, ix, noReuse,
|
||||
allAdded("qwen3.5-35b-a3b-apex-i-quality", "qwen3.5-35b-a3b-ud-q6-k"))
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(inserts).To(BeEmpty())
|
||||
Expect(newHubs).To(HaveLen(1))
|
||||
|
||||
hub := newHubs[0]
|
||||
Expect(hub.Name).To(Equal("qwen3.5-35b-a3b"))
|
||||
Expect(hub.Name).ToNot(HaveSuffix("-apex"))
|
||||
|
||||
// A hand-written *-apex entry is an ordinary build, referenced like any
|
||||
// other rung and never deleted or renamed.
|
||||
Expect(hub.Variants).To(Equal([]VariantRef{
|
||||
{Model: "qwen3.5-35b-a3b-apex"},
|
||||
{Model: "qwen3.5-35b-a3b-apex-i-quality"},
|
||||
{Model: "qwen3.5-35b-a3b-ud-q6-k"},
|
||||
}))
|
||||
|
||||
// The verifier skips entries with no declared backend, so a hub without
|
||||
// one would escape the tagging check in silence.
|
||||
Expect(hub.Overrides).To(HaveKeyWithValue("backend", "llama-cpp"))
|
||||
Expect(hub.Files).ToNot(BeEmpty())
|
||||
Expect(hub.Tags).To(ContainElement("vision"))
|
||||
})
|
||||
|
||||
It("gathers two APEX repos that share one base model under a single hub", func() {
|
||||
ix := mustIndex("- name: unrelated\n url: u\n")
|
||||
fams := []family{
|
||||
{
|
||||
repo: "mudler/Solo-APEX-GGUF",
|
||||
repoBase: "Solo-APEX",
|
||||
stem: "Solo-APEX",
|
||||
children: []childBuild{buildOf("solo-apex-i-quality", "mudler/Solo-APEX-GGUF", "a.gguf", 0)},
|
||||
},
|
||||
{
|
||||
repo: "mudler/Solo-APEX-MTP-GGUF",
|
||||
repoBase: "Solo-APEX-MTP",
|
||||
stem: "Solo-APEX-MTP",
|
||||
children: []childBuild{buildOf("solo-apex-mtp-i-quality", "mudler/Solo-APEX-MTP-GGUF", "b.gguf", 0)},
|
||||
},
|
||||
}
|
||||
|
||||
_, newHubs, err := planHubs(fams, ix, noReuse, allAdded("solo-apex-i-quality", "solo-apex-mtp-i-quality"))
|
||||
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(newHubs).To(HaveLen(1))
|
||||
Expect(newHubs[0].Name).To(Equal("solo"))
|
||||
Expect(newHubs[0].Variants).To(Equal([]VariantRef{
|
||||
{Model: "solo-apex-i-quality"},
|
||||
{Model: "solo-apex-mtp-i-quality"},
|
||||
}))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("hubVariants", func() {
|
||||
It("orders builds by quality rung rather than discovery order", func() {
|
||||
// DiscoverAPEXTiers preserves input order and the HF API returns siblings
|
||||
// alphabetically, so an unsorted list reads I-Balanced, I-Compact, I-Mini,
|
||||
// I-Nano, I-Quality. Selection ignores authored order; this is for the
|
||||
// human reading the file.
|
||||
f := family{repoBase: "X-APEX", stem: "X-APEX", children: []childBuild{
|
||||
{rank: rungRank["I-Nano"], entry: GalleryEntry{Name: "x-i-nano"}},
|
||||
{rank: 100, entry: GalleryEntry{Name: "x-ud-q4-k-m"}},
|
||||
{rank: rungRank["I-Quality"], entry: GalleryEntry{Name: "x-i-quality"}},
|
||||
{rank: rungRank["I-Compact"], entry: GalleryEntry{Name: "x-i-compact"}},
|
||||
}}
|
||||
|
||||
got := hubVariants(&f, mustIndex("- name: x\n url: u\n"), map[string]string{}, map[string]bool{})
|
||||
|
||||
Expect(got).To(Equal([]string{"x-i-quality", "x-i-compact", "x-i-nano", "x-ud-q4-k-m"}))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("ParseIndexText", func() {
|
||||
It("refuses to edit by line number when the two views of the file disagree", func() {
|
||||
_, err := ParseIndexText("- name: one\n url: u\n-\n")
|
||||
Expect(err).To(MatchError(ContainSubstring("empty")))
|
||||
})
|
||||
|
||||
It("records the line range of each entry", func() {
|
||||
ix := mustIndex("- name: first\n url: u\n- name: second\n url: u\n")
|
||||
|
||||
Expect(ix.Find("FIRST").Pos.StartLine).To(Equal(0))
|
||||
Expect(ix.Find("first").Pos.EndLine).To(Equal(2))
|
||||
Expect(ix.Find("second").Pos.StartLine).To(Equal(2))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("resolveVariant", func() {
|
||||
It("keeps an entry that was emitted even when it is also in reused", func() {
|
||||
// A within-batch name collision records reused[name] = name while the
|
||||
// FIRST entry of that name is still in add. Treating presence in reused as
|
||||
// "dropped" would emit nothing for it.
|
||||
added := map[string]bool{"dup": true}
|
||||
reused := map[string]string{"dup": "dup"}
|
||||
|
||||
Expect(resolveVariant("dup", reused, added)).To(Equal("dup"))
|
||||
})
|
||||
|
||||
It("redirects a reused name at the entry that stands in for it", func() {
|
||||
added := map[string]bool{}
|
||||
reused := map[string]string{"generated": "already-in-gallery"}
|
||||
|
||||
Expect(resolveVariant("generated", reused, added)).To(Equal("already-in-gallery"))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("slug", func() {
|
||||
It("lowercases and turns quant underscores into hyphens", func() {
|
||||
Expect(slug("UD-Q4_K_M")).To(Equal("ud-q4-k-m"))
|
||||
Expect(slug("gemma-4-26B-A4B-it-APEX")).To(Equal("gemma-4-26b-a4b-it-apex"))
|
||||
Expect(slug("I-Nano")).To(Equal("i-nano"))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("sortTiers", func() {
|
||||
It("puts the imatrix ladder in descending quality order", func() {
|
||||
tiers := []Tier{
|
||||
{Label: "I-Balanced"}, {Label: "I-Compact"}, {Label: "I-Mini"},
|
||||
{Label: "I-Nano"}, {Label: "I-Quality"},
|
||||
}
|
||||
sortTiers(tiers)
|
||||
Expect(tierLabels(tiers)).To(Equal("I-Quality,I-Balanced,I-Compact,I-Mini,I-Nano"))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("restrict", func() {
|
||||
It("keeps only the named repos", func() {
|
||||
got := restrict([]string{"mudler/A-APEX-GGUF", "mudler/B-APEX-GGUF"}, "mudler/B-APEX-GGUF")
|
||||
Expect(got).To(Equal([]string{"mudler/B-APEX-GGUF"}))
|
||||
})
|
||||
|
||||
It("returns nothing when the filter matches nothing", func() {
|
||||
Expect(restrict([]string{"mudler/A-APEX-GGUF"}, "mudler/typo")).To(BeEmpty())
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("reportUnclassified", func() {
|
||||
// One real imatrix rung is always present so the specs measure how the
|
||||
// remaining files are bucketed, not an empty-repo edge case.
|
||||
tier := Tier{Label: "I-Quality", File: GGUFFile{Name: "Model-APEX-I-Quality.gguf"}}
|
||||
|
||||
censusOf := func(names ...string) fileCensus {
|
||||
files := []GGUFFile{tier.File}
|
||||
for _, n := range names {
|
||||
files = append(files, GGUFFile{Name: n})
|
||||
}
|
||||
return reportUnclassified("mudler/Model-APEX-GGUF", files, []Tier{tier}, nil)
|
||||
}
|
||||
|
||||
It("counts a flat full-precision source as excluded, not unclassified", func() {
|
||||
got := censusOf("Carnice-MoE-35B-A3B-F16.gguf")
|
||||
Expect(got.fullPrecision).To(Equal(1))
|
||||
Expect(got.unclassified).To(Equal(0))
|
||||
})
|
||||
|
||||
It("counts every shard of a sharded full-precision source as excluded", func() {
|
||||
got := censusOf(
|
||||
"MiniMax-M2.7-APEX-F16-00001-of-00003.gguf",
|
||||
"MiniMax-M2.7-APEX-F16-00002-of-00003.gguf",
|
||||
"MiniMax-M2.7-APEX-F16-00003-of-00003.gguf",
|
||||
)
|
||||
Expect(got.fullPrecision).To(Equal(3))
|
||||
Expect(got.unclassified).To(Equal(0))
|
||||
})
|
||||
|
||||
It("treats bf16 the same as f16, in either case", func() {
|
||||
got := censusOf("Model-APEX-BF16.gguf", "Model-APEX-bf16-00001-of-00002.gguf", "Model-APEX-f16.gguf")
|
||||
Expect(got.fullPrecision).To(Equal(3))
|
||||
Expect(got.unclassified).To(Equal(0))
|
||||
})
|
||||
|
||||
It("still reports a genuinely unknown filename as unclassified", func() {
|
||||
got := censusOf("Model-APEX-Turbo.gguf")
|
||||
Expect(got.unclassified).To(Equal(1))
|
||||
Expect(got.fullPrecision).To(Equal(0))
|
||||
})
|
||||
|
||||
It("separates the two kinds when a repo publishes both", func() {
|
||||
got := censusOf("Model-APEX-F16.gguf", "Model-APEX-Turbo.gguf")
|
||||
Expect(got.fullPrecision).To(Equal(1))
|
||||
Expect(got.unclassified).To(Equal(1))
|
||||
})
|
||||
})
|
||||
143
.github/ci/apexentries/merge.go
vendored
143
.github/ci/apexentries/merge.go
vendored
@@ -1,143 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"strings"
|
||||
|
||||
"gopkg.in/yaml.v3"
|
||||
)
|
||||
|
||||
const (
|
||||
hfShorthandPrefix = "huggingface://"
|
||||
hfResolvePrefix = "https://huggingface.co/"
|
||||
hfResolveInfix = "/resolve/main/"
|
||||
)
|
||||
|
||||
// canonicalURI reduces the two interchangeable spellings of a HuggingFace file
|
||||
// to one key, so a generated resolve/main URI dedups against the shorthand the
|
||||
// gallery uses for the majority of its entries.
|
||||
//
|
||||
// The repo is exactly the first two path segments; everything after is the file
|
||||
// path, which may itself contain slashes because sharded quants live in a
|
||||
// subdirectory. Anything that is not recognisably one of the two forms is
|
||||
// returned unchanged rather than guessed at, so mirrors and other hosts still
|
||||
// dedup on their literal string.
|
||||
func canonicalURI(uri string) string {
|
||||
switch {
|
||||
case strings.HasPrefix(uri, hfShorthandPrefix):
|
||||
rest := strings.TrimPrefix(uri, hfShorthandPrefix)
|
||||
owner, after, ok := strings.Cut(rest, "/")
|
||||
if !ok {
|
||||
return uri
|
||||
}
|
||||
name, file, ok := strings.Cut(after, "/")
|
||||
if !ok || owner == "" || name == "" || file == "" {
|
||||
return uri
|
||||
}
|
||||
return hfShorthandPrefix + owner + "/" + name + "/" + file
|
||||
|
||||
case strings.HasPrefix(uri, hfResolvePrefix):
|
||||
rest := strings.TrimPrefix(uri, hfResolvePrefix)
|
||||
repo, file, ok := strings.Cut(rest, hfResolveInfix)
|
||||
if !ok || file == "" {
|
||||
return uri
|
||||
}
|
||||
// A repo is owner/name and nothing more; a longer prefix means this is
|
||||
// some other huggingface.co URL that must not be rewritten.
|
||||
owner, name, ok := strings.Cut(repo, "/")
|
||||
if !ok || owner == "" || name == "" || strings.Contains(name, "/") {
|
||||
return uri
|
||||
}
|
||||
return hfShorthandPrefix + repo + "/" + file
|
||||
|
||||
default:
|
||||
return uri
|
||||
}
|
||||
}
|
||||
|
||||
// ExistingIndex is the lookup built from the current gallery: entry names, and
|
||||
// which entry claims each weight URI.
|
||||
type ExistingIndex struct {
|
||||
ByName map[string]int
|
||||
ByURI map[string]string
|
||||
}
|
||||
|
||||
// LoadExisting reads the gallery index for dedup purposes only. It is
|
||||
// deliberately not used to rewrite the file: the index is 40,000 lines, and a
|
||||
// YAML round trip would reflow the whole thing into an unreviewable diff.
|
||||
func LoadExisting(path string) (*ExistingIndex, error) {
|
||||
raw, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
var entries []struct {
|
||||
Name string `yaml:"name"`
|
||||
Files []struct {
|
||||
URI string `yaml:"uri"`
|
||||
} `yaml:"files"`
|
||||
}
|
||||
if err := yaml.Unmarshal(raw, &entries); err != nil {
|
||||
return nil, fmt.Errorf("parsing %s: %w", path, err)
|
||||
}
|
||||
|
||||
ix := &ExistingIndex{ByName: map[string]int{}, ByURI: map[string]string{}}
|
||||
for i, e := range entries {
|
||||
ix.ByName[e.Name] = i
|
||||
for _, f := range e.Files {
|
||||
if f.URI != "" {
|
||||
ix.ByURI[canonicalURI(f.URI)] = e.Name
|
||||
}
|
||||
}
|
||||
}
|
||||
return ix, nil
|
||||
}
|
||||
|
||||
// Merge splits generated entries into those to add and those already covered.
|
||||
// reused maps a generated name to the existing entry that stands in for it, so
|
||||
// a parent can reference what is already there instead of duplicating weights.
|
||||
// Several APEX repos share one base model, so the same counterpart rungs are
|
||||
// generated more than once in a batch. The batch has to dedup against itself as
|
||||
// well as against the gallery, tracked locally because the caller may reuse the
|
||||
// ExistingIndex it passed in.
|
||||
func Merge(existing *ExistingIndex, generated []GalleryEntry) (add []GalleryEntry, reused map[string]string) {
|
||||
reused = map[string]string{}
|
||||
batchNames := map[string]string{}
|
||||
batchURIs := map[string]string{}
|
||||
|
||||
// Canonicalized into a local copy rather than in place: an ExistingIndex may
|
||||
// be hand-built or reused by the caller, so Merge must not rewrite it.
|
||||
existingURIs := make(map[string]string, len(existing.ByURI))
|
||||
for uri, owner := range existing.ByURI {
|
||||
existingURIs[canonicalURI(uri)] = owner
|
||||
}
|
||||
|
||||
for _, e := range generated {
|
||||
// Name is checked before URI: a name collision must block the add
|
||||
// whatever the weights say, since duplicate names corrupt the index.
|
||||
if _, clash := existing.ByName[e.Name]; clash {
|
||||
reused[e.Name] = e.Name
|
||||
continue
|
||||
}
|
||||
if claimant, clash := batchNames[e.Name]; clash {
|
||||
reused[e.Name] = claimant
|
||||
continue
|
||||
}
|
||||
if len(e.Files) > 0 {
|
||||
uri := canonicalURI(e.Files[0].URI)
|
||||
if owner, ok := existingURIs[uri]; ok {
|
||||
reused[e.Name] = owner
|
||||
continue
|
||||
}
|
||||
if claimant, ok := batchURIs[uri]; ok {
|
||||
reused[e.Name] = claimant
|
||||
continue
|
||||
}
|
||||
batchURIs[uri] = e.Name
|
||||
}
|
||||
batchNames[e.Name] = e.Name
|
||||
add = append(add, e)
|
||||
}
|
||||
return add, reused
|
||||
}
|
||||
183
.github/ci/apexentries/merge_test.go
vendored
183
.github/ci/apexentries/merge_test.go
vendored
@@ -1,183 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
)
|
||||
|
||||
var _ = Describe("Merge", func() {
|
||||
It("drops a generated entry whose weight URI already exists and reports the existing name", func() {
|
||||
existing := &ExistingIndex{
|
||||
ByName: map[string]int{"qwen3.6-35b-a3b-apex": 0},
|
||||
ByURI: map[string]string{
|
||||
"https://huggingface.co/mudler/X-APEX-GGUF/resolve/main/X-APEX-I-Quality.gguf": "qwen3.6-35b-a3b-apex",
|
||||
},
|
||||
}
|
||||
gen := []GalleryEntry{{
|
||||
Name: "x-apex-i-quality",
|
||||
Files: []EntryFile{{URI: "https://huggingface.co/mudler/X-APEX-GGUF/resolve/main/X-APEX-I-Quality.gguf"}},
|
||||
}}
|
||||
|
||||
add, reused := Merge(existing, gen)
|
||||
|
||||
Expect(add).To(BeEmpty())
|
||||
Expect(reused).To(HaveKeyWithValue("x-apex-i-quality", "qwen3.6-35b-a3b-apex"))
|
||||
})
|
||||
|
||||
It("keeps a generated entry whose weights are new", func() {
|
||||
existing := &ExistingIndex{ByName: map[string]int{}, ByURI: map[string]string{}}
|
||||
gen := []GalleryEntry{{
|
||||
Name: "x-apex-i-mini",
|
||||
Files: []EntryFile{{URI: "https://huggingface.co/mudler/X-APEX-GGUF/resolve/main/X-APEX-I-Mini.gguf"}},
|
||||
}}
|
||||
|
||||
add, reused := Merge(existing, gen)
|
||||
|
||||
Expect(add).To(HaveLen(1))
|
||||
Expect(reused).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("refuses to add an entry whose name collides with an existing one", func() {
|
||||
existing := &ExistingIndex{
|
||||
ByName: map[string]int{"x-apex-i-mini": 0},
|
||||
ByURI: map[string]string{},
|
||||
}
|
||||
gen := []GalleryEntry{{
|
||||
Name: "x-apex-i-mini",
|
||||
Files: []EntryFile{{URI: "https://huggingface.co/mudler/X-APEX-GGUF/resolve/main/other.gguf"}},
|
||||
}}
|
||||
|
||||
add, reused := Merge(existing, gen)
|
||||
|
||||
Expect(add).To(BeEmpty())
|
||||
Expect(reused).To(HaveKeyWithValue("x-apex-i-mini", "x-apex-i-mini"))
|
||||
})
|
||||
|
||||
// The gallery records most of its URIs in huggingface:// shorthand while
|
||||
// render.go only ever emits the resolve/main form, so without
|
||||
// canonicalization the majority of the file is invisible to the dedup.
|
||||
It("matches a generated https URI against the shorthand form recorded in the gallery", func() {
|
||||
existing := &ExistingIndex{
|
||||
ByName: map[string]int{"foo-gguf-q8-0": 0},
|
||||
ByURI: map[string]string{
|
||||
"huggingface://unsloth/Foo-GGUF/Foo-Q8_0.gguf": "foo-gguf-q8-0",
|
||||
},
|
||||
}
|
||||
gen := []GalleryEntry{{
|
||||
Name: "foo-apex-q8-0",
|
||||
Files: []EntryFile{{URI: "https://huggingface.co/unsloth/Foo-GGUF/resolve/main/Foo-Q8_0.gguf"}},
|
||||
}}
|
||||
|
||||
add, reused := Merge(existing, gen)
|
||||
|
||||
Expect(add).To(BeEmpty())
|
||||
Expect(reused).To(HaveKeyWithValue("foo-apex-q8-0", "foo-gguf-q8-0"))
|
||||
})
|
||||
|
||||
It("matches a generated shorthand URI against the https form recorded in the gallery", func() {
|
||||
existing := &ExistingIndex{
|
||||
ByName: map[string]int{"foo-gguf-q8-0": 0},
|
||||
ByURI: map[string]string{
|
||||
"https://huggingface.co/unsloth/Foo-GGUF/resolve/main/Foo-Q8_0.gguf": "foo-gguf-q8-0",
|
||||
},
|
||||
}
|
||||
gen := []GalleryEntry{{
|
||||
Name: "foo-apex-q8-0",
|
||||
Files: []EntryFile{{URI: "huggingface://unsloth/Foo-GGUF/Foo-Q8_0.gguf"}},
|
||||
}}
|
||||
|
||||
add, reused := Merge(existing, gen)
|
||||
|
||||
Expect(add).To(BeEmpty())
|
||||
Expect(reused).To(HaveKeyWithValue("foo-apex-q8-0", "foo-gguf-q8-0"))
|
||||
})
|
||||
|
||||
// Sharded quants live under a subdirectory, so the file path carries slashes
|
||||
// of its own and only the first two segments are the repo.
|
||||
It("matches across both forms when the file path has a subdirectory", func() {
|
||||
existing := &ExistingIndex{
|
||||
ByName: map[string]int{"model-ud-q4-k-m": 0},
|
||||
ByURI: map[string]string{
|
||||
"huggingface://unsloth/Model-GGUF/UD-Q4_K_M/Model-UD-Q4_K_M-00001-of-00002.gguf": "model-ud-q4-k-m",
|
||||
},
|
||||
}
|
||||
gen := []GalleryEntry{{
|
||||
Name: "model-apex-ud-q4-k-m",
|
||||
Files: []EntryFile{{URI: "https://huggingface.co/unsloth/Model-GGUF/resolve/main/UD-Q4_K_M/Model-UD-Q4_K_M-00001-of-00002.gguf"}},
|
||||
}}
|
||||
|
||||
add, reused := Merge(existing, gen)
|
||||
|
||||
Expect(add).To(BeEmpty())
|
||||
Expect(reused).To(HaveKeyWithValue("model-apex-ud-q4-k-m", "model-ud-q4-k-m"))
|
||||
})
|
||||
|
||||
// Several APEX repos share one base model, so the same unsloth rungs are
|
||||
// generated more than once in a single batch.
|
||||
It("adds only the first of two generated entries sharing a name", func() {
|
||||
existing := &ExistingIndex{ByName: map[string]int{}, ByURI: map[string]string{}}
|
||||
gen := []GalleryEntry{
|
||||
{
|
||||
Name: "shared-rung-q8-0",
|
||||
Files: []EntryFile{{URI: "https://huggingface.co/unsloth/Shared-GGUF/resolve/main/Shared-Q8_0.gguf"}},
|
||||
},
|
||||
{
|
||||
Name: "shared-rung-q8-0",
|
||||
Files: []EntryFile{{URI: "https://huggingface.co/unsloth/Other-GGUF/resolve/main/Other-Q8_0.gguf"}},
|
||||
},
|
||||
}
|
||||
|
||||
add, reused := Merge(existing, gen)
|
||||
|
||||
Expect(add).To(HaveLen(1))
|
||||
Expect(add[0].Files[0].URI).To(Equal("https://huggingface.co/unsloth/Shared-GGUF/resolve/main/Shared-Q8_0.gguf"))
|
||||
Expect(reused).To(HaveKeyWithValue("shared-rung-q8-0", "shared-rung-q8-0"))
|
||||
})
|
||||
|
||||
It("adds only the first of two generated entries sharing a primary URI", func() {
|
||||
existing := &ExistingIndex{ByName: map[string]int{}, ByURI: map[string]string{}}
|
||||
gen := []GalleryEntry{
|
||||
{
|
||||
Name: "shared-rung-from-apex",
|
||||
Files: []EntryFile{{URI: "https://huggingface.co/unsloth/Shared-GGUF/resolve/main/Shared-Q8_0.gguf"}},
|
||||
},
|
||||
{
|
||||
Name: "shared-rung-from-apex-mtp",
|
||||
Files: []EntryFile{{URI: "huggingface://unsloth/Shared-GGUF/Shared-Q8_0.gguf"}},
|
||||
},
|
||||
}
|
||||
|
||||
add, reused := Merge(existing, gen)
|
||||
|
||||
Expect(add).To(HaveLen(1))
|
||||
Expect(add[0].Name).To(Equal("shared-rung-from-apex"))
|
||||
Expect(reused).To(HaveKeyWithValue("shared-rung-from-apex-mtp", "shared-rung-from-apex"))
|
||||
})
|
||||
|
||||
// Anything that is not a HuggingFace URI must survive untouched, so an
|
||||
// unrecognised scheme still dedups against the very same string.
|
||||
It("leaves a URI in neither recognised form alone and still dedups it exactly", func() {
|
||||
existing := &ExistingIndex{
|
||||
ByName: map[string]int{"mirrored-model": 0},
|
||||
ByURI: map[string]string{
|
||||
"https://mirror.example.com/weights/Model-Q8_0.gguf": "mirrored-model",
|
||||
},
|
||||
}
|
||||
gen := []GalleryEntry{
|
||||
{
|
||||
Name: "mirrored-apex",
|
||||
Files: []EntryFile{{URI: "https://mirror.example.com/weights/Model-Q8_0.gguf"}},
|
||||
},
|
||||
{
|
||||
Name: "elsewhere-apex",
|
||||
Files: []EntryFile{{URI: "https://mirror.example.com/weights/Other-Q8_0.gguf"}},
|
||||
},
|
||||
}
|
||||
|
||||
add, reused := Merge(existing, gen)
|
||||
|
||||
Expect(add).To(HaveLen(1))
|
||||
Expect(add[0].Name).To(Equal("elsewhere-apex"))
|
||||
Expect(reused).To(HaveKeyWithValue("mirrored-apex", "mirrored-model"))
|
||||
})
|
||||
})
|
||||
175
.github/ci/apexentries/render.go
vendored
175
.github/ci/apexentries/render.go
vendored
@@ -1,175 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"path"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// EntryFile is one downloadable file of a gallery entry.
|
||||
type EntryFile struct {
|
||||
Filename string `yaml:"filename"`
|
||||
SHA256 string `yaml:"sha256"`
|
||||
URI string `yaml:"uri"`
|
||||
}
|
||||
|
||||
// GalleryEntry is the subset of a gallery entry this generator writes.
|
||||
//
|
||||
// Named GalleryEntry rather than Entry because the test files dot-import
|
||||
// Ginkgo, whose table DSL exports an Entry that a package-level Entry would
|
||||
// collide with. The yaml tags are what the gallery index sees, so the Go
|
||||
// identifier is free to differ.
|
||||
type GalleryEntry struct {
|
||||
Name string `yaml:"name"`
|
||||
URL string `yaml:"url"`
|
||||
Description string `yaml:"description,omitempty"`
|
||||
Tags []string `yaml:"tags,omitempty"`
|
||||
Overrides map[string]any `yaml:"overrides,omitempty"`
|
||||
Files []EntryFile `yaml:"files,omitempty"`
|
||||
Variants []VariantRef `yaml:"variants,omitempty"`
|
||||
}
|
||||
|
||||
// VariantRef mirrors the gallery's variant reference: a name and nothing else.
|
||||
type VariantRef struct {
|
||||
Model string `yaml:"model"`
|
||||
}
|
||||
|
||||
// ChildInput is everything needed to render one non-parent entry.
|
||||
type ChildInput struct {
|
||||
Name string
|
||||
Repo string
|
||||
// DraftRepo is the repo publishing the drafter, when it is not the repo
|
||||
// publishing the weights. Speculative pairings routinely cross repos, so
|
||||
// the drafter cannot be assumed to sit next to the weights. Empty means
|
||||
// same-repo, which is how the *-APEX-MTP-GGUF repos ship.
|
||||
DraftRepo string
|
||||
Template string
|
||||
Weights []GGUFFile
|
||||
MMProj *GGUFFile
|
||||
SpecType string
|
||||
DraftFile *GGUFFile
|
||||
BaseTags []string
|
||||
}
|
||||
|
||||
// specTuning is the acceptance-window tuning each spec type ships with, copied
|
||||
// from the hand-written entries that already run these two mechanisms rather
|
||||
// than invented here. The two differ because the drafters differ: self-drafted
|
||||
// MTP heads produce a short, high-confidence proposal (15+ hand-written entries
|
||||
// use 6 with a 0.75 floor), while a separate DFlash drafter is cheap enough to
|
||||
// run far ahead unconditionally (the five hand-written dflash entries use 15 and
|
||||
// set no floor).
|
||||
var specTuning = map[string][]string{
|
||||
"draft-mtp": {"spec_n_max:6", "spec_p_min:0.75"},
|
||||
"draft-dflash": {"spec_n_max:15"},
|
||||
}
|
||||
|
||||
func hfURI(repo, file string) string {
|
||||
return fmt.Sprintf("https://huggingface.co/%s/resolve/main/%s", repo, file)
|
||||
}
|
||||
|
||||
// localPath is where a downloaded file lands.
|
||||
//
|
||||
// The hand-written entries namespace by the repo's BARE name
|
||||
// (llama-cpp/models/<repo>/<file>), which is not unique. LiquidAI/LFM2.5-8B-A1B-GGUF
|
||||
// and unsloth/LFM2.5-8B-A1B-GGUF share a basename, so both claim
|
||||
// llama-cpp/models/LFM2.5-8B-A1B-GGUF/, and installing the second after the first
|
||||
// either overwrites weights whose recorded sha256 belongs to the other file or is
|
||||
// skipped as already present. Two owners publishing the same model name is the
|
||||
// normal case for quantizers, not an edge case, so the owner has to be in the path.
|
||||
//
|
||||
// The owner becomes its own path segment rather than being folded into the
|
||||
// directory name: owner/repo is unique on HuggingFace and "/" cannot occur inside
|
||||
// either half, so this is the only form that is collision-proof by construction.
|
||||
// It still reads as the hand-written convention with the owner restored, and the
|
||||
// extra depth is already present in the index for sharded builds.
|
||||
func localPath(kind, repo, file string) string {
|
||||
// path.Dir yields "." for a repo named without an owner, which path.Join
|
||||
// drops, so such a caller keeps the historical two-segment layout.
|
||||
return path.Join("llama-cpp", kind, path.Dir(repo), path.Base(repo), file)
|
||||
}
|
||||
|
||||
// RenderChild builds one child entry.
|
||||
//
|
||||
// The dflash/mtp tag is added if and only if this entry sets a spec_type,
|
||||
// because variant ranking reads tags and nothing else, and a tag that does not
|
||||
// match what the entry configures either promotes a build that is no faster or
|
||||
// hides one that is.
|
||||
func RenderChild(in ChildInput) GalleryEntry {
|
||||
e := GalleryEntry{
|
||||
Name: in.Name,
|
||||
URL: fmt.Sprintf("github:mudler/LocalAI/gallery/%s@master", in.Template),
|
||||
Tags: append([]string{}, in.BaseTags...),
|
||||
Overrides: map[string]any{},
|
||||
}
|
||||
|
||||
// gallery/virtual.yaml carries no backend, so nothing else would name an
|
||||
// engine for these entries. Matching the hand-written entries on
|
||||
// known_usecases too: LocalAI would fall back to the backend defaults, but
|
||||
// generated entries should not read differently from their neighbours.
|
||||
e.Overrides["backend"] = "llama-cpp"
|
||||
e.Overrides["known_usecases"] = []string{"chat"}
|
||||
|
||||
options := []string{"use_jinja:true"}
|
||||
|
||||
for _, w := range in.Weights {
|
||||
e.Files = append(e.Files, EntryFile{
|
||||
Filename: localPath("models", in.Repo, w.Name),
|
||||
SHA256: w.SHA256,
|
||||
URI: hfURI(in.Repo, w.Name),
|
||||
})
|
||||
}
|
||||
e.Overrides["parameters"] = map[string]any{
|
||||
"model": localPath("models", in.Repo, in.Weights[0].Name),
|
||||
}
|
||||
|
||||
if in.MMProj != nil {
|
||||
// An explicit known_usecases SUPPRESSES the backend-default fallback in
|
||||
// core/gallery/models_types.go, so a multimodal entry left at chat-only
|
||||
// never matches FilterGalleryModelsByUsecase(FLAG_VISION) or
|
||||
// FilterGalleryModelsByMultimodal and vanishes from the UI's vision and
|
||||
// multimodal filters. 19 of the 45 APEX repos ship an mmproj.
|
||||
e.Overrides["known_usecases"] = []string{"chat", "vision"}
|
||||
e.Overrides["mmproj"] = localPath("mmproj", in.Repo, in.MMProj.Name)
|
||||
e.Files = append(e.Files, EntryFile{
|
||||
Filename: localPath("mmproj", in.Repo, in.MMProj.Name),
|
||||
SHA256: in.MMProj.SHA256,
|
||||
URI: hfURI(in.Repo, in.MMProj.Name),
|
||||
})
|
||||
}
|
||||
|
||||
// A spec type is configured independently of a drafter FILE. Weights that
|
||||
// carry their own MTP heads need no second download, and requiring one left
|
||||
// the *-APEX-MTP-GGUF builds shipping the larger heads-bearing weights with
|
||||
// the heads switched off: a strictly bigger download at the same speed,
|
||||
// ranked identically to the plain rung at the same tier.
|
||||
if in.SpecType != "" {
|
||||
options = append(options, "spec_type:"+in.SpecType)
|
||||
options = append(options, specTuning[in.SpecType]...)
|
||||
// The tag is derived from the spec type this entry sets and from nothing
|
||||
// else. Variant ranking reads tags only, so a tag taken from a repo or
|
||||
// entry NAME would promote a build that is no faster whenever the name
|
||||
// and the configuration disagree.
|
||||
e.Tags = append(e.Tags, strings.TrimPrefix(in.SpecType, "draft-"))
|
||||
}
|
||||
|
||||
if in.SpecType != "" && in.DraftFile != nil {
|
||||
// Fall back to the weights repo so pairings that publish the drafter
|
||||
// alongside the weights keep working without restating the repo.
|
||||
draftRepo := in.DraftRepo
|
||||
if draftRepo == "" {
|
||||
draftRepo = in.Repo
|
||||
}
|
||||
draftPath := localPath("models", draftRepo, in.DraftFile.Name)
|
||||
|
||||
e.Overrides["draft_model"] = draftPath
|
||||
e.Overrides["flash_attention"] = "on"
|
||||
e.Files = append(e.Files, EntryFile{
|
||||
Filename: draftPath,
|
||||
SHA256: in.DraftFile.SHA256,
|
||||
URI: hfURI(draftRepo, in.DraftFile.Name),
|
||||
})
|
||||
}
|
||||
|
||||
e.Overrides["options"] = options
|
||||
return e
|
||||
}
|
||||
249
.github/ci/apexentries/render_test.go
vendored
249
.github/ci/apexentries/render_test.go
vendored
@@ -1,249 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
)
|
||||
|
||||
var _ = Describe("RenderChild", func() {
|
||||
It("tags an entry that configures draft-dflash", func() {
|
||||
e := RenderChild(ChildInput{
|
||||
Name: "qwen3.5-9b-dflash",
|
||||
Repo: "mudler/Example-APEX-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "Example-APEX-I-Quality.gguf", SHA256: "a"}},
|
||||
SpecType: "draft-dflash",
|
||||
DraftFile: &GGUFFile{Name: "Example-DFlash.Q8_0.gguf", SHA256: "b"},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(e.Tags).To(ContainElement("dflash"))
|
||||
Expect(e.Tags).ToNot(ContainElement("mtp"))
|
||||
Expect(e.Overrides["options"]).To(ContainElement("spec_type:draft-dflash"))
|
||||
Expect(e.Overrides["draft_model"]).ToNot(BeNil())
|
||||
})
|
||||
|
||||
It("does not tag an MTP-named repo that configures no speculation", func() {
|
||||
// mudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF ships MTP-bearing weights. Weights
|
||||
// that carry the heads are not an entry that enables them, and tagging it
|
||||
// would win the feature axis without being any faster.
|
||||
e := RenderChild(ChildInput{
|
||||
Name: "qwen3.6-35b-a3b-apex-mtp-i-quality",
|
||||
Repo: "mudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "Qwen3.6-35B-A3B-APEX-MTP-I-Quality.gguf", SHA256: "a"}},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(e.Tags).ToNot(ContainElement("mtp"))
|
||||
Expect(e.Tags).ToNot(ContainElement("dflash"))
|
||||
Expect(e.Overrides).ToNot(HaveKey("draft_model"))
|
||||
})
|
||||
|
||||
It("lists every shard of a sharded build and points the model at the first", func() {
|
||||
e := RenderChild(ChildInput{
|
||||
Name: "step-3.7-flash-ud-q4-k-m",
|
||||
Repo: "unsloth/Step-3.7-Flash-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{
|
||||
{Name: "UD-Q4_K_M/Step-3.7-Flash-UD-Q4_K_M-00001-of-00002.gguf", SHA256: "a"},
|
||||
{Name: "UD-Q4_K_M/Step-3.7-Flash-UD-Q4_K_M-00002-of-00002.gguf", SHA256: "b"},
|
||||
},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(e.Files).To(HaveLen(2))
|
||||
params, ok := e.Overrides["parameters"].(map[string]any)
|
||||
Expect(ok).To(BeTrue())
|
||||
Expect(params["model"]).To(HaveSuffix("00001-of-00002.gguf"))
|
||||
Expect(e.Files[0].URI).To(Equal(
|
||||
"https://huggingface.co/unsloth/Step-3.7-Flash-GGUF/resolve/main/UD-Q4_K_M/Step-3.7-Flash-UD-Q4_K_M-00001-of-00002.gguf"))
|
||||
})
|
||||
|
||||
It("wires mmproj when the repo publishes one", func() {
|
||||
e := RenderChild(ChildInput{
|
||||
Name: "example-i-mini",
|
||||
Repo: "mudler/Example-APEX-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "Example-APEX-I-Mini.gguf", SHA256: "a"}},
|
||||
MMProj: &GGUFFile{Name: "mmproj-F16.gguf", SHA256: "c"},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(e.Overrides["mmproj"]).ToNot(BeNil())
|
||||
Expect(e.Files).To(HaveLen(2))
|
||||
})
|
||||
|
||||
It("names the engine and the usecases the hand-written entries name", func() {
|
||||
// gallery/virtual.yaml supplies no backend, so an entry that omits one
|
||||
// names no engine at all and cannot load.
|
||||
e := RenderChild(ChildInput{
|
||||
Name: "example-i-mini",
|
||||
Repo: "mudler/Example-APEX-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "Example-APEX-I-Mini.gguf", SHA256: "a"}},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(e.Overrides["backend"]).To(Equal("llama-cpp"))
|
||||
Expect(e.Overrides["known_usecases"]).To(ContainElement("chat"))
|
||||
})
|
||||
|
||||
It("draws the drafter from DraftRepo when the pairing spans two repos", func() {
|
||||
// unsloth/Qwen3-4B-GGUF pairs with a drafter published separately by
|
||||
// AtomicChat, so a drafter URI built from the weights repo 404s.
|
||||
e := RenderChild(ChildInput{
|
||||
Name: "qwen3-4b-dflash",
|
||||
Repo: "unsloth/Qwen3-4B-GGUF",
|
||||
DraftRepo: "AtomicChat/Qwen3-4B-DFlash-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "Qwen3-4B-Q4_K_M.gguf", SHA256: "a"}},
|
||||
SpecType: "draft-dflash",
|
||||
DraftFile: &GGUFFile{Name: "Qwen3-4B-DFlash.Q8_0.gguf", SHA256: "b"},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(e.Files[0].URI).To(Equal(
|
||||
"https://huggingface.co/unsloth/Qwen3-4B-GGUF/resolve/main/Qwen3-4B-Q4_K_M.gguf"))
|
||||
Expect(e.Files[1].URI).To(Equal(
|
||||
"https://huggingface.co/AtomicChat/Qwen3-4B-DFlash-GGUF/resolve/main/Qwen3-4B-DFlash.Q8_0.gguf"))
|
||||
Expect(e.Files[1].Filename).To(Equal(
|
||||
"llama-cpp/models/AtomicChat/Qwen3-4B-DFlash-GGUF/Qwen3-4B-DFlash.Q8_0.gguf"))
|
||||
Expect(e.Overrides["draft_model"]).To(Equal(
|
||||
"llama-cpp/models/AtomicChat/Qwen3-4B-DFlash-GGUF/Qwen3-4B-DFlash.Q8_0.gguf"))
|
||||
})
|
||||
|
||||
It("falls back to the weights repo for the drafter when DraftRepo is empty", func() {
|
||||
// The *-APEX-MTP-GGUF repos ship the drafter alongside the weights.
|
||||
e := RenderChild(ChildInput{
|
||||
Name: "example-apex-dflash",
|
||||
Repo: "mudler/Example-APEX-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "Example-APEX-I-Quality.gguf", SHA256: "a"}},
|
||||
SpecType: "draft-dflash",
|
||||
DraftFile: &GGUFFile{Name: "Example-DFlash.Q8_0.gguf", SHA256: "b"},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(e.Files[1].URI).To(Equal(
|
||||
"https://huggingface.co/mudler/Example-APEX-GGUF/resolve/main/Example-DFlash.Q8_0.gguf"))
|
||||
Expect(e.Files[1].Filename).To(Equal(
|
||||
"llama-cpp/models/mudler/Example-APEX-GGUF/Example-DFlash.Q8_0.gguf"))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("RenderChild known_usecases", func() {
|
||||
It("declares vision alongside chat when the entry carries an mmproj", func() {
|
||||
// An explicit known_usecases suppresses the backend-default fallback, so a
|
||||
// chat-only multimodal entry disappears from the UI's vision filter.
|
||||
e := RenderChild(ChildInput{
|
||||
Name: "example-i-quality",
|
||||
Repo: "mudler/Example-APEX-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "Example-APEX-I-Quality.gguf", SHA256: "a"}},
|
||||
MMProj: &GGUFFile{Name: "mmproj-F16.gguf", SHA256: "c"},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(e.Overrides["known_usecases"]).To(ConsistOf("chat", "vision"))
|
||||
})
|
||||
|
||||
It("leaves a text-only entry at chat", func() {
|
||||
e := RenderChild(ChildInput{
|
||||
Name: "example-i-quality",
|
||||
Repo: "mudler/Example-APEX-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "Example-APEX-I-Quality.gguf", SHA256: "a"}},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(e.Overrides["known_usecases"]).To(ConsistOf("chat"))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("localPath", func() {
|
||||
It("keeps two repos with the same basename but different owners apart", func() {
|
||||
// LiquidAI and unsloth both publish LFM2.5-8B-A1B-GGUF. A path built from
|
||||
// the bare repo name gives both the same local file, so installing the
|
||||
// second overwrites or skips the first and one of them then serves bytes
|
||||
// that do not match its recorded sha256.
|
||||
liquid := RenderChild(ChildInput{
|
||||
Name: "lfm2.5-8b-a1b-i-quality",
|
||||
Repo: "LiquidAI/LFM2.5-8B-A1B-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "LFM2.5-8B-A1B-Q8_0.gguf", SHA256: "33ab3b8c"}},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
unsloth := RenderChild(ChildInput{
|
||||
Name: "lfm2.5-8b-a1b-q8-0",
|
||||
Repo: "unsloth/LFM2.5-8B-A1B-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "LFM2.5-8B-A1B-Q8_0.gguf", SHA256: "ec11666b"}},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(liquid.Files[0].Filename).ToNot(Equal(unsloth.Files[0].Filename))
|
||||
Expect(unsloth.Files[0].Filename).To(Equal(
|
||||
"llama-cpp/models/unsloth/LFM2.5-8B-A1B-GGUF/LFM2.5-8B-A1B-Q8_0.gguf"))
|
||||
})
|
||||
|
||||
It("namespaces the mmproj by owner too", func() {
|
||||
e := RenderChild(ChildInput{
|
||||
Name: "example-i-quality",
|
||||
Repo: "mudler/Example-APEX-GGUF",
|
||||
Template: "virtual.yaml",
|
||||
Weights: []GGUFFile{{Name: "Example-APEX-I-Quality.gguf", SHA256: "a"}},
|
||||
MMProj: &GGUFFile{Name: "mmproj-F16.gguf", SHA256: "c"},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
|
||||
Expect(e.Overrides["mmproj"]).To(Equal(
|
||||
"llama-cpp/mmproj/mudler/Example-APEX-GGUF/mmproj-F16.gguf"))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("MTP builds", func() {
|
||||
renderTier := func(repo string) GalleryEntry {
|
||||
return RenderChild(ChildInput{
|
||||
Name: "example-i-quality",
|
||||
Repo: repo,
|
||||
Template: "virtual.yaml",
|
||||
SpecType: SpecTypeForRepo(repo),
|
||||
Weights: []GGUFFile{{Name: "Example-I-Quality.gguf", SHA256: "a"}},
|
||||
BaseTags: []string{"llm", "gguf"},
|
||||
})
|
||||
}
|
||||
|
||||
It("turns MTP on for a build off an APEX-MTP repo", func() {
|
||||
// These weights retain the model's own MTP heads, so shipping them with
|
||||
// speculation off is a strictly larger download at the same speed,
|
||||
// ranked identically to the plain rung at the same tier.
|
||||
e := renderTier("mudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF")
|
||||
|
||||
Expect(e.Overrides["options"]).To(ContainElements(
|
||||
"spec_type:draft-mtp", "spec_n_max:6", "spec_p_min:0.75"))
|
||||
Expect(e.Tags).To(ContainElement("mtp"))
|
||||
})
|
||||
|
||||
It("needs no drafter file, because the heads travel with the weights", func() {
|
||||
e := renderTier("mudler/Qwen3.6-35B-A3B-APEX-MTP-GGUF")
|
||||
|
||||
Expect(e.Overrides).ToNot(HaveKey("draft_model"))
|
||||
Expect(e.Files).To(HaveLen(1))
|
||||
})
|
||||
|
||||
It("leaves a build off a plain APEX repo alone", func() {
|
||||
e := renderTier("mudler/Qwen3.6-35B-A3B-APEX-GGUF")
|
||||
|
||||
Expect(e.Tags).ToNot(ContainElement("mtp"))
|
||||
Expect(e.Overrides["options"]).To(ConsistOf("use_jinja:true"))
|
||||
})
|
||||
|
||||
It("leaves an unsloth counterpart rung alone", func() {
|
||||
// The counterpart quantizes the plain weights; nothing there carries heads.
|
||||
e := renderTier("unsloth/Qwen3.6-35B-A3B-GGUF")
|
||||
|
||||
Expect(e.Tags).ToNot(ContainElement("mtp"))
|
||||
Expect(e.Overrides["options"]).To(ConsistOf("use_jinja:true"))
|
||||
})
|
||||
})
|
||||
71
.github/ci/apexentries/unsloth.go
vendored
71
.github/ci/apexentries/unsloth.go
vendored
@@ -1,71 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"regexp"
|
||||
"sort"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// WantedQuants is the fixed unsloth subset this generator emits. It is a
|
||||
// deliberate subset: unsloth publishes north of 20 quants per repo, and the
|
||||
// selector needs useful fitness points rather than every rung.
|
||||
var WantedQuants = []string{"UD-Q4_K_M", "UD-Q5_K_M", "UD-Q6_K", "Q8_0"}
|
||||
|
||||
var shardRE = regexp.MustCompile(`-(\d{5})-of-(\d{5})\.gguf$`)
|
||||
|
||||
// QuantBuild is one unsloth quantization, which may be a single file or an
|
||||
// ordered set of shards.
|
||||
type QuantBuild struct {
|
||||
Quant string
|
||||
Files []GGUFFile
|
||||
Sharded bool
|
||||
}
|
||||
|
||||
// CounterpartCandidates returns the unsloth repo base names worth probing, most
|
||||
// likely first. Both derivations are needed: the repo name finds
|
||||
// unsloth/gemma-4-26B-A4B-it-GGUF, while the file stem is what matches for
|
||||
// repos whose stem is the canonical model name.
|
||||
func CounterpartCandidates(repoName, fileStem string) []string {
|
||||
clean := func(s string) string {
|
||||
s = strings.TrimSuffix(s, "-GGUF")
|
||||
s = regexp.MustCompile(`-(MTP|TQ)$`).ReplaceAllString(s, "")
|
||||
s = strings.TrimSuffix(s, "-APEX")
|
||||
return regexp.MustCompile(`-(MTP|TQ)$`).ReplaceAllString(s, "")
|
||||
}
|
||||
|
||||
out := []string{clean(repoName)}
|
||||
if stem := clean(fileStem); stem != out[0] {
|
||||
out = append(out, stem)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// DiscoverUnslothQuants returns the wanted quants a repo publishes, handling
|
||||
// both the flat single-file layout and the sharded layout where a quant lives
|
||||
// in its own subdirectory.
|
||||
func DiscoverUnslothQuants(files []GGUFFile) []QuantBuild {
|
||||
var out []QuantBuild
|
||||
|
||||
for _, q := range WantedQuants {
|
||||
var flat []GGUFFile
|
||||
var shards []GGUFFile
|
||||
|
||||
for _, f := range files {
|
||||
switch {
|
||||
case !strings.Contains(f.Name, "/") && strings.HasSuffix(f.Name, "-"+q+".gguf"):
|
||||
flat = append(flat, f)
|
||||
case strings.HasPrefix(f.Name, q+"/") && shardRE.MatchString(f.Name):
|
||||
shards = append(shards, f)
|
||||
}
|
||||
}
|
||||
|
||||
switch {
|
||||
case len(flat) > 0:
|
||||
out = append(out, QuantBuild{Quant: q, Files: flat})
|
||||
case len(shards) > 0:
|
||||
sort.Slice(shards, func(i, j int) bool { return shards[i].Name < shards[j].Name })
|
||||
out = append(out, QuantBuild{Quant: q, Files: shards, Sharded: true})
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
75
.github/ci/apexentries/unsloth_test.go
vendored
75
.github/ci/apexentries/unsloth_test.go
vendored
@@ -1,75 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
)
|
||||
|
||||
var _ = Describe("CounterpartCandidates", func() {
|
||||
It("offers both the repo-derived and stem-derived names", func() {
|
||||
// mudler/gemma-4-26B-A4B-it-APEX-GGUF ships gemma-4-26B-A4B-APEX-*.gguf,
|
||||
// and only the repo-derived name finds unsloth/gemma-4-26B-A4B-it-GGUF.
|
||||
got := CounterpartCandidates("gemma-4-26B-A4B-it-APEX-GGUF", "gemma-4-26B-A4B-APEX")
|
||||
|
||||
Expect(got).To(Equal([]string{"gemma-4-26B-A4B-it", "gemma-4-26B-A4B"}))
|
||||
})
|
||||
|
||||
It("strips the MTP marker", func() {
|
||||
got := CounterpartCandidates("Qwopus3.6-35B-A3B-v1-APEX-MTP-GGUF", "Qwopus3.6-35B-A3B-v1-APEX-MTP")
|
||||
|
||||
Expect(got[0]).To(Equal("Qwopus3.6-35B-A3B-v1"))
|
||||
})
|
||||
|
||||
It("strips the TQ marker", func() {
|
||||
// This is the branch that folds mudler/Qwen3.5-35B-A3B-APEX-TQ-GGUF into
|
||||
// the qwen3.5-35b-a3b hub. Without it the probe is
|
||||
// unsloth/Qwen3.5-35B-A3B-TQ-GGUF, which does not exist, so the family
|
||||
// silently loses every unsloth rung.
|
||||
got := CounterpartCandidates("Qwen3.5-35B-A3B-APEX-TQ-GGUF", "Qwen3.5-35B-A3B-APEX-TQ")
|
||||
|
||||
Expect(got).To(Equal([]string{"Qwen3.5-35B-A3B"}))
|
||||
})
|
||||
|
||||
It("does not repeat a candidate when both derivations agree", func() {
|
||||
got := CounterpartCandidates("Qwen3.6-35B-A3B-APEX-GGUF", "Qwen3.6-35B-A3B-APEX")
|
||||
|
||||
Expect(got).To(Equal([]string{"Qwen3.6-35B-A3B"}))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("DiscoverUnslothQuants", func() {
|
||||
It("finds flat single-file quants", func() {
|
||||
files := []GGUFFile{
|
||||
{Name: "Qwen3.6-35B-A3B-UD-Q4_K_M.gguf", SHA256: "a"},
|
||||
{Name: "Qwen3.6-35B-A3B-UD-IQ1_M.gguf", SHA256: "b"},
|
||||
}
|
||||
|
||||
got := DiscoverUnslothQuants(files)
|
||||
|
||||
Expect(got).To(HaveLen(1))
|
||||
Expect(got[0].Quant).To(Equal("UD-Q4_K_M"))
|
||||
Expect(got[0].Sharded).To(BeFalse())
|
||||
Expect(got[0].Files).To(HaveLen(1))
|
||||
})
|
||||
|
||||
It("collects a sharded quant from its subdirectory in shard order", func() {
|
||||
files := []GGUFFile{
|
||||
{Name: "UD-Q4_K_M/Step-3.7-Flash-UD-Q4_K_M-00002-of-00002.gguf", SHA256: "b"},
|
||||
{Name: "UD-Q4_K_M/Step-3.7-Flash-UD-Q4_K_M-00001-of-00002.gguf", SHA256: "a"},
|
||||
}
|
||||
|
||||
got := DiscoverUnslothQuants(files)
|
||||
|
||||
Expect(got).To(HaveLen(1))
|
||||
Expect(got[0].Quant).To(Equal("UD-Q4_K_M"))
|
||||
Expect(got[0].Sharded).To(BeTrue())
|
||||
Expect(got[0].Files).To(HaveLen(2))
|
||||
Expect(got[0].Files[0].Name).To(HaveSuffix("00001-of-00002.gguf"))
|
||||
})
|
||||
|
||||
It("ignores quants outside the wanted subset", func() {
|
||||
files := []GGUFFile{{Name: "Model-UD-IQ2_XXS.gguf", SHA256: "a"}}
|
||||
|
||||
Expect(DiscoverUnslothQuants(files)).To(BeEmpty())
|
||||
})
|
||||
})
|
||||
312
.github/ci/apexentries/verify.go
vendored
312
.github/ci/apexentries/verify.go
vendored
@@ -1,312 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"strings"
|
||||
|
||||
"gopkg.in/yaml.v3"
|
||||
)
|
||||
|
||||
type verifyEntry struct {
|
||||
Name string `yaml:"name"`
|
||||
Tags []string `yaml:"tags"`
|
||||
Variants []VariantRef `yaml:"variants"`
|
||||
Overrides struct {
|
||||
// Backend scopes the checks that only hold for one engine. An entry that
|
||||
// declares none takes its configuration from the referenced url: template,
|
||||
// which this verifier never reads, so it cannot be judged either way.
|
||||
Backend string `yaml:"backend"`
|
||||
Options []string `yaml:"options"`
|
||||
// MMProj and DraftModel name the files that are not weights. They are
|
||||
// the only signal for it: a drafter lands in the same models/ prefix as
|
||||
// the weights, so the path alone cannot tell them apart.
|
||||
MMProj string `yaml:"mmproj"`
|
||||
DraftModel string `yaml:"draft_model"`
|
||||
} `yaml:"overrides"`
|
||||
Files []struct {
|
||||
Filename string `yaml:"filename"`
|
||||
SHA256 string `yaml:"sha256"`
|
||||
URI string `yaml:"uri"`
|
||||
} `yaml:"files"`
|
||||
}
|
||||
|
||||
// Verify checks the invariants the variants schema and the tagging rule
|
||||
// require. It returns every problem rather than the first, so one run tells the
|
||||
// author everything that needs fixing.
|
||||
func Verify(path string) []string {
|
||||
raw, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return []string{fmt.Sprintf("reading %s: %v", path, err)}
|
||||
}
|
||||
|
||||
var entries []verifyEntry
|
||||
if err := yaml.Unmarshal(raw, &entries); err != nil {
|
||||
return []string{fmt.Sprintf("parsing %s: %v", path, err)}
|
||||
}
|
||||
|
||||
var problems []string
|
||||
|
||||
byName := map[string]verifyEntry{}
|
||||
for _, e := range entries {
|
||||
if _, seen := byName[e.Name]; seen {
|
||||
problems = append(problems, fmt.Sprintf("duplicate entry name: %s", e.Name))
|
||||
continue
|
||||
}
|
||||
byName[e.Name] = e
|
||||
}
|
||||
|
||||
for _, e := range entries {
|
||||
for _, v := range e.Variants {
|
||||
target, ok := byName[v.Model]
|
||||
if !ok {
|
||||
problems = append(problems, fmt.Sprintf("%s: variant %q does not exist", e.Name, v.Model))
|
||||
continue
|
||||
}
|
||||
if len(target.Variants) > 0 {
|
||||
problems = append(problems, fmt.Sprintf("%s: variant %q declares variants of its own", e.Name, v.Model))
|
||||
}
|
||||
}
|
||||
|
||||
for _, f := range e.Files {
|
||||
if requiresSHA256(f.Filename) && f.SHA256 == "" {
|
||||
problems = append(problems, fmt.Sprintf("%s: file %s has no sha256", e.Name, f.Filename))
|
||||
}
|
||||
}
|
||||
|
||||
problems = append(problems, checkWeightCount(e)...)
|
||||
problems = append(problems, checkFeatureTag(e, "dflash")...)
|
||||
problems = append(problems, checkFeatureTag(e, "mtp")...)
|
||||
}
|
||||
|
||||
problems = append(problems, checkPathCollisions(entries)...)
|
||||
|
||||
return problems
|
||||
}
|
||||
|
||||
// checkPathCollisions catches two different upstream files claiming one local
|
||||
// path. The install layer keys on the local filename, so whichever entry is
|
||||
// installed second either overwrites weights the first entry recorded a
|
||||
// different sha256 for or is skipped as already present. Either way some entry
|
||||
// afterwards serves bytes that do not match its own checksum, and nothing at
|
||||
// install time says so.
|
||||
//
|
||||
// This is an index-wide invariant rather than a per-entry one: neither entry is
|
||||
// wrong on its own and the collision exists only in their pairing. The usual
|
||||
// source is a path scheme built from the repo's BARE name, because two owners
|
||||
// publishing the same model name is routine for quantizers.
|
||||
//
|
||||
// Sharing a path is fine when the uri is the same, which is how several entries
|
||||
// legitimately reuse one projector. Files with no uri are skipped: there is
|
||||
// nothing to compare.
|
||||
func checkPathCollisions(entries []verifyEntry) []string {
|
||||
type source struct{ uri, entry string }
|
||||
|
||||
first := map[string]source{}
|
||||
reported := map[string]bool{}
|
||||
|
||||
var problems []string
|
||||
for _, e := range entries {
|
||||
for _, f := range e.Files {
|
||||
if f.Filename == "" || f.URI == "" {
|
||||
continue
|
||||
}
|
||||
prev, seen := first[f.Filename]
|
||||
if !seen {
|
||||
first[f.Filename] = source{uri: f.URI, entry: e.Name}
|
||||
continue
|
||||
}
|
||||
if prev.uri == f.URI || reported[f.Filename] {
|
||||
continue
|
||||
}
|
||||
// Reported once per path however many entries pile onto it, so one
|
||||
// heavily reused filename cannot bury the rest of the report.
|
||||
reported[f.Filename] = true
|
||||
problems = append(problems, fmt.Sprintf(
|
||||
"local path %s is claimed by two different uris: %s (%s) and %s (%s)",
|
||||
f.Filename, prev.uri, prev.entry, f.URI, e.Name))
|
||||
}
|
||||
}
|
||||
return problems
|
||||
}
|
||||
|
||||
// auxiliaryExtensions are the metadata formats an entry ships beside its
|
||||
// weights, where an unverified download is a nuisance rather than a hole.
|
||||
//
|
||||
// The exclusion is stated as a list of metadata formats on purpose. Requiring
|
||||
// the checksum only on a blessed list of weight formats would silently exempt
|
||||
// every format nobody has shipped yet, and it already exempted safetensors
|
||||
// weights, which are downloaded and loaded exactly like GGUF ones.
|
||||
var auxiliaryExtensions = []string{".json", ".txt", ".md"}
|
||||
|
||||
// requiresSHA256 reports whether an unverified download of this file would be
|
||||
// a supply-chain hole rather than a cosmetic gap.
|
||||
func requiresSHA256(filename string) bool {
|
||||
for _, ext := range auxiliaryExtensions {
|
||||
if strings.HasSuffix(filename, ext) {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// checkWeightCount catches an entry carrying two whole models. The flat-match
|
||||
// branch in DiscoverUnslothQuants appends every match, so a quant label that is
|
||||
// a suffix of another one (Q8_0 of UD-Q8_0) collects both files into one build
|
||||
// while the rendered model: points at only the first. The result downloads
|
||||
// twice the bytes and serves whichever file sorted first, silently.
|
||||
//
|
||||
// Shards are exempt because a sharded build is legitimately many files.
|
||||
//
|
||||
// The collision is a property of llama-cpp quant discovery, so the check is
|
||||
// scoped to that backend. Multi-component TTS, ASR and diffusion engines ship an
|
||||
// encoder, a decoder and a vocoder as one model, and there the second GGUF is
|
||||
// the design rather than a bug.
|
||||
func checkWeightCount(e verifyEntry) []string {
|
||||
if e.Overrides.Backend != "llama-cpp" {
|
||||
return nil
|
||||
}
|
||||
|
||||
var weights []string
|
||||
for _, f := range e.Files {
|
||||
switch {
|
||||
case !strings.HasSuffix(f.Filename, ".gguf"):
|
||||
case shardRE.MatchString(f.Filename):
|
||||
case f.Filename == e.Overrides.MMProj:
|
||||
case f.Filename == e.Overrides.DraftModel:
|
||||
default:
|
||||
weights = append(weights, f.Filename)
|
||||
}
|
||||
}
|
||||
|
||||
if len(weights) > 1 {
|
||||
return []string{fmt.Sprintf("%s: more than one weight file: %s", e.Name, strings.Join(weights, ", "))}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// checkFeatureTag enforces the rule in both directions. A tag without the
|
||||
// configuration promotes a build that is no faster; configuration without the
|
||||
// tag leaves a genuinely faster build ranked as plain.
|
||||
//
|
||||
// It only speaks about backends whose declaration it can actually read, because
|
||||
// a rule applied where the evidence is invisible reports noise rather than bugs.
|
||||
func checkFeatureTag(e verifyEntry, feature string) []string {
|
||||
decl, configured, judgeable := featureDeclaration(e, feature)
|
||||
if !judgeable {
|
||||
return nil
|
||||
}
|
||||
|
||||
tagged := false
|
||||
for _, t := range e.Tags {
|
||||
if t == feature {
|
||||
tagged = true
|
||||
break
|
||||
}
|
||||
}
|
||||
|
||||
switch {
|
||||
case tagged && !configured:
|
||||
return []string{fmt.Sprintf("%s: tagged %s but sets no %s", e.Name, feature, decl)}
|
||||
case configured && !tagged:
|
||||
return []string{fmt.Sprintf("%s: sets %s but is not tagged %s", e.Name, decl, feature)}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// featureDeclaration implements the per-backend table in
|
||||
// .agents/adding-gallery-models.md. It returns the declaration the backend uses
|
||||
// to configure the feature, whether the entry carries it, and whether this
|
||||
// verifier is in a position to answer at all.
|
||||
func featureDeclaration(e verifyEntry, feature string) (decl string, configured, judgeable bool) {
|
||||
switch e.Overrides.Backend {
|
||||
case "llama-cpp":
|
||||
decl = "spec_type:draft-" + feature
|
||||
for _, o := range e.Overrides.Options {
|
||||
if strings.TrimSpace(o) == decl {
|
||||
return decl, true, true
|
||||
}
|
||||
}
|
||||
return decl, false, true
|
||||
|
||||
case "ds4":
|
||||
// ds4 carries the MTP heads in the weights and turns them on with
|
||||
// mtp_path / mtp_draft. It has no dflash counterpart, so dflash is not a
|
||||
// question that can be asked of a ds4 entry.
|
||||
if feature != "mtp" {
|
||||
return "", false, false
|
||||
}
|
||||
decl = "mtp_path:"
|
||||
for _, o := range e.Overrides.Options {
|
||||
o = strings.TrimSpace(o)
|
||||
if strings.HasPrefix(o, "mtp_path:") || strings.HasPrefix(o, "mtp_draft:") {
|
||||
return decl, true, true
|
||||
}
|
||||
}
|
||||
return decl, false, true
|
||||
|
||||
default:
|
||||
// sglang configures the feature with speculative_algorithm: in the
|
||||
// referenced gallery/*.yaml, and an entry that declares no backend takes
|
||||
// its whole configuration from its url: template. Verify reads one index
|
||||
// file and follows neither, so it must not judge these in either
|
||||
// direction.
|
||||
return "", false, false
|
||||
}
|
||||
}
|
||||
|
||||
// UnaccountedQuants reports a wanted quant the repo demonstrably publishes but
|
||||
// that discovery produced no build for. The layout that triggers it today is
|
||||
// root-level shards, which match neither branch of DiscoverUnslothQuants; no
|
||||
// counterpart ships that way yet, but a batch generator must not drop a build
|
||||
// with nothing said about it.
|
||||
func UnaccountedQuants(files []GGUFFile, builds []QuantBuild) []string {
|
||||
built := map[string]bool{}
|
||||
for _, b := range builds {
|
||||
built[b.Quant] = true
|
||||
}
|
||||
|
||||
var problems []string
|
||||
for _, q := range WantedQuants {
|
||||
if built[q] {
|
||||
continue
|
||||
}
|
||||
for _, f := range files {
|
||||
if filePublishesQuant(f.Name, q) {
|
||||
problems = append(problems, fmt.Sprintf("quant %s is published upstream (%s) but produced no build", q, f.Name))
|
||||
break
|
||||
}
|
||||
}
|
||||
}
|
||||
return problems
|
||||
}
|
||||
|
||||
// filePublishesQuant reports whether an upstream file is a publication of
|
||||
// quant q. It anchors on the quant label the way DiscoverUnslothQuants does,
|
||||
// as the trailing token of the base name or as the sharding subdirectory, so
|
||||
// the diagnostic and the discovery it audits cannot disagree about what a file
|
||||
// is.
|
||||
//
|
||||
// An unanchored match would reproduce the very collision this diagnostic warns
|
||||
// about: Q8_0 is a substring of UD-Q8_0, so a repo publishing only UD-Q8_0
|
||||
// would be reported as publishing an unbuilt Q8_0, which it does not, and
|
||||
// UD-Q8_0 is not a wanted quant at all.
|
||||
func filePublishesQuant(name, q string) bool {
|
||||
if strings.HasPrefix(name, q+"/") {
|
||||
return true
|
||||
}
|
||||
|
||||
base := name[strings.LastIndex(name, "/")+1:]
|
||||
// Shard numbering sits between the quant label and the extension, so it has
|
||||
// to come off before the label can be read as the trailing token. Root-level
|
||||
// shards are the layout that matches neither branch of
|
||||
// DiscoverUnslothQuants, and so the layout this diagnostic mainly catches.
|
||||
base = shardRE.ReplaceAllString(base, ".gguf")
|
||||
|
||||
if !strings.HasSuffix(base, "-"+q+".gguf") {
|
||||
return false
|
||||
}
|
||||
// UD- is unsloth's dynamic-quant modifier, and UD-<q> is a distinct quant
|
||||
// label rather than a publication of <q>.
|
||||
return !strings.HasSuffix(base, "-UD-"+q+".gguf")
|
||||
}
|
||||
480
.github/ci/apexentries/verify_test.go
vendored
480
.github/ci/apexentries/verify_test.go
vendored
@@ -1,480 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
)
|
||||
|
||||
var _ = Describe("Verify", func() {
|
||||
write := func(body string) string {
|
||||
dir := GinkgoT().TempDir()
|
||||
p := filepath.Join(dir, "index.yaml")
|
||||
Expect(os.WriteFile(p, []byte(body), 0o600)).To(Succeed())
|
||||
return p
|
||||
}
|
||||
|
||||
It("passes a sound index", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: parent
|
||||
variants:
|
||||
- model: child
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
- name: child
|
||||
files:
|
||||
- filename: b.gguf
|
||||
sha256: bb
|
||||
uri: https://example.com/b.gguf
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("reports a variant pointing at a missing entry", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: parent
|
||||
variants:
|
||||
- model: ghost
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
`))).To(ContainElement(ContainSubstring("ghost")))
|
||||
})
|
||||
|
||||
It("reports a variant that itself declares variants", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: parent
|
||||
variants:
|
||||
- model: child
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
- name: child
|
||||
variants:
|
||||
- model: grandchild
|
||||
files:
|
||||
- filename: b.gguf
|
||||
sha256: bb
|
||||
uri: https://example.com/b.gguf
|
||||
- name: grandchild
|
||||
files:
|
||||
- filename: c.gguf
|
||||
sha256: cc
|
||||
uri: https://example.com/c.gguf
|
||||
`))).To(ContainElement(ContainSubstring("declares variants of its own")))
|
||||
})
|
||||
|
||||
It("reports duplicate entry names", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: dup
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
- name: dup
|
||||
files:
|
||||
- filename: b.gguf
|
||||
sha256: bb
|
||||
uri: https://example.com/b.gguf
|
||||
`))).To(ContainElement(ContainSubstring("duplicate entry name")))
|
||||
})
|
||||
|
||||
It("reports a file with no sha256", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: one
|
||||
files:
|
||||
- filename: a.gguf
|
||||
uri: https://example.com/a.gguf
|
||||
`))).To(ContainElement(ContainSubstring("no sha256")))
|
||||
})
|
||||
|
||||
It("reports an entry tagged dflash without a matching spec_type", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: liar
|
||||
tags:
|
||||
- dflash
|
||||
overrides:
|
||||
backend: llama-cpp
|
||||
options:
|
||||
- use_jinja:true
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
`))).To(ContainElement(ContainSubstring("tagged dflash")))
|
||||
})
|
||||
|
||||
It("reports an entry configuring spec_type without the tag", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: shy
|
||||
overrides:
|
||||
backend: llama-cpp
|
||||
options:
|
||||
- spec_type:draft-mtp
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
`))).To(ContainElement(ContainSubstring("not tagged mtp")))
|
||||
})
|
||||
|
||||
// ds4 carries the MTP heads in the weights and names them with mtp_path, so
|
||||
// the rule holds there in a different vocabulary rather than not at all.
|
||||
It("reports a ds4 entry configuring mtp_path without the tag", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: ds4-shy
|
||||
overrides:
|
||||
backend: ds4
|
||||
options:
|
||||
- mtp_path:model-mtp.gguf
|
||||
- mtp_draft:2
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
`))).To(ContainElement(ContainSubstring("not tagged mtp")))
|
||||
})
|
||||
|
||||
It("reports a ds4 entry tagged mtp that configures no mtp_path", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: ds4-liar
|
||||
tags:
|
||||
- mtp
|
||||
overrides:
|
||||
backend: ds4
|
||||
options:
|
||||
- context_size:4096
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
`))).To(ContainElement(ContainSubstring("tagged mtp")))
|
||||
})
|
||||
|
||||
It("accepts a ds4 entry that both configures mtp_path and carries the tag", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: ds4-honest
|
||||
tags:
|
||||
- mtp
|
||||
overrides:
|
||||
backend: ds4
|
||||
options:
|
||||
- mtp_path:model-mtp.gguf
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
|
||||
// sglang declares speculative_algorithm in the referenced gallery/*.yaml,
|
||||
// which Verify never reads, so it may not judge such an entry either way.
|
||||
It("says nothing about an sglang entry tagged mtp", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: sglang-mtp
|
||||
tags:
|
||||
- mtp
|
||||
overrides:
|
||||
backend: sglang
|
||||
files: []
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("says nothing about the tag on an entry with no declared backend", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: templated
|
||||
tags:
|
||||
- mtp
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
|
||||
// The flat-match branch in unsloth.go appends every match, so a repo
|
||||
// publishing both a plain and a UD Q8_0 renders one entry holding two full
|
||||
// models while model: points at only the first.
|
||||
It("reports an entry holding more than one non-shard weight file", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: greedy
|
||||
overrides:
|
||||
backend: llama-cpp
|
||||
options:
|
||||
- use_jinja:true
|
||||
parameters:
|
||||
model: llama-cpp/models/repo/Model-Q8_0.gguf
|
||||
files:
|
||||
- filename: llama-cpp/models/repo/Model-Q8_0.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
- filename: llama-cpp/models/repo/Model-UD-Q8_0.gguf
|
||||
sha256: bb
|
||||
uri: https://example.com/b.gguf
|
||||
`))).To(ContainElement(ContainSubstring("more than one weight file")))
|
||||
})
|
||||
|
||||
It("accepts many shards alongside an mmproj and a drafter", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: sharded
|
||||
tags:
|
||||
- mtp
|
||||
overrides:
|
||||
backend: llama-cpp
|
||||
options:
|
||||
- spec_type:draft-mtp
|
||||
mmproj: llama-cpp/mmproj/repo/mm.gguf
|
||||
draft_model: llama-cpp/models/repo/Model-draft.gguf
|
||||
files:
|
||||
- filename: llama-cpp/models/repo/Model-00001-of-00002.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
- filename: llama-cpp/models/repo/Model-00002-of-00002.gguf
|
||||
sha256: bb
|
||||
uri: https://example.com/b.gguf
|
||||
- filename: llama-cpp/mmproj/repo/mm.gguf
|
||||
sha256: cc
|
||||
uri: https://example.com/c.gguf
|
||||
- filename: llama-cpp/models/repo/Model-draft.gguf
|
||||
sha256: dd
|
||||
uri: https://example.com/d.gguf
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
|
||||
// Multi-component TTS and ASR engines legitimately ship an encoder, a
|
||||
// tokenizer, a vocoder and so on as one model, so the collision the weight
|
||||
// count catches does not exist for them.
|
||||
It("accepts a multi-component non-llama-cpp entry declaring five weights", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: multi
|
||||
overrides:
|
||||
backend: qwen3-tts-cpp
|
||||
files:
|
||||
- filename: talker.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
- filename: tokenizer.gguf
|
||||
sha256: bb
|
||||
uri: https://example.com/b.gguf
|
||||
- filename: vocoder.gguf
|
||||
sha256: cc
|
||||
uri: https://example.com/c.gguf
|
||||
- filename: encoder.gguf
|
||||
sha256: dd
|
||||
uri: https://example.com/d.gguf
|
||||
- filename: vae.gguf
|
||||
sha256: ee
|
||||
uri: https://example.com/e.gguf
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("says nothing about the weight count of an entry with no declared backend", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: templated-weights
|
||||
files:
|
||||
- filename: model-Q4_K_M.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
- filename: model-mmproj-f16.gguf
|
||||
sha256: bb
|
||||
uri: https://example.com/b.gguf
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("says nothing about an auxiliary metadata file carrying no sha256", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: aux
|
||||
files:
|
||||
- filename: a.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
- filename: params.json
|
||||
sha256: ""
|
||||
uri: https://example.com/params.json
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
|
||||
// safetensors weights are downloaded and loaded exactly like GGUF weights,
|
||||
// so an unverified one is the same supply-chain hole.
|
||||
It("reports a safetensors weight carrying no sha256", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: vae
|
||||
files:
|
||||
- filename: wan_2.1_vae.safetensors
|
||||
sha256: ""
|
||||
uri: https://example.com/vae.safetensors
|
||||
`))).To(ContainElement(ContainSubstring("no sha256")))
|
||||
})
|
||||
|
||||
It("says nothing about a txt or md file carrying no sha256", func() {
|
||||
Expect(Verify(write(`
|
||||
- name: docs
|
||||
files:
|
||||
- filename: notes.txt
|
||||
sha256: ""
|
||||
uri: https://example.com/notes.txt
|
||||
- filename: README.md
|
||||
sha256: ""
|
||||
uri: https://example.com/README.md
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("UnaccountedQuants", func() {
|
||||
// A quant published only as root-level shards matches neither branch in
|
||||
// DiscoverUnslothQuants, so without this diagnostic the build would vanish
|
||||
// from a batch run with nothing said about it.
|
||||
It("reports a wanted quant upstream publishes but discovery dropped", func() {
|
||||
files := []GGUFFile{
|
||||
{Name: "Model-UD-Q4_K_M-00001-of-00003.gguf", SHA256: "aa"},
|
||||
{Name: "Model-UD-Q4_K_M-00002-of-00003.gguf", SHA256: "bb"},
|
||||
{Name: "Model-UD-Q4_K_M-00003-of-00003.gguf", SHA256: "cc"},
|
||||
}
|
||||
|
||||
Expect(UnaccountedQuants(files, DiscoverUnslothQuants(files))).
|
||||
To(ContainElement(ContainSubstring("UD-Q4_K_M")))
|
||||
})
|
||||
|
||||
It("says nothing when every published wanted quant produced a build", func() {
|
||||
files := []GGUFFile{
|
||||
{Name: "Model-UD-Q4_K_M.gguf", SHA256: "aa"},
|
||||
{Name: "UD-Q6_K/Model-UD-Q6_K-00001-of-00002.gguf", SHA256: "bb"},
|
||||
{Name: "UD-Q6_K/Model-UD-Q6_K-00002-of-00002.gguf", SHA256: "cc"},
|
||||
}
|
||||
|
||||
Expect(UnaccountedQuants(files, DiscoverUnslothQuants(files))).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("says nothing about a wanted quant the repo does not publish at all", func() {
|
||||
files := []GGUFFile{{Name: "Model-UD-Q4_K_M.gguf", SHA256: "aa"}}
|
||||
|
||||
Expect(UnaccountedQuants(files, DiscoverUnslothQuants(files))).To(BeEmpty())
|
||||
})
|
||||
|
||||
// UD-Q8_0 is its own quant label and is not a wanted one. Reading it as a
|
||||
// publication of Q8_0 is the substring collision this diagnostic exists to
|
||||
// warn about, and subdirectory-sharded UD quants are the normal unsloth
|
||||
// layout for large repos, so the false positive would fire on every batch.
|
||||
It("does not read a subdirectory-sharded UD-Q8_0 as a published Q8_0", func() {
|
||||
files := []GGUFFile{
|
||||
{Name: "UD-Q8_0/Model-UD-Q8_0-00001-of-00002.gguf", SHA256: "aa"},
|
||||
{Name: "UD-Q8_0/Model-UD-Q8_0-00002-of-00002.gguf", SHA256: "bb"},
|
||||
}
|
||||
|
||||
Expect(UnaccountedQuants(files, DiscoverUnslothQuants(files))).To(BeEmpty())
|
||||
})
|
||||
|
||||
// A quant in its own subdirectory but not shard-numbered matches neither
|
||||
// branch of DiscoverUnslothQuants, so it is genuinely published and
|
||||
// genuinely undiscovered.
|
||||
It("reports a wanted quant published in its own subdirectory without shard numbering", func() {
|
||||
files := []GGUFFile{{Name: "Q8_0/Model-Q8_0.gguf", SHA256: "aa"}}
|
||||
|
||||
Expect(UnaccountedQuants(files, DiscoverUnslothQuants(files))).
|
||||
To(ContainElement(ContainSubstring("quant Q8_0 is published upstream")))
|
||||
})
|
||||
|
||||
// builds is empty on purpose: it isolates the file-to-quant match from
|
||||
// whatever DiscoverUnslothQuants would have made of the same file.
|
||||
It("matches the flat single-file layout", func() {
|
||||
files := []GGUFFile{{Name: "Model-Q8_0.gguf", SHA256: "aa"}}
|
||||
|
||||
Expect(UnaccountedQuants(files, nil)).
|
||||
To(ConsistOf(ContainSubstring("quant Q8_0 is published upstream")))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("Verify local path collisions", func() {
|
||||
write := func(body string) string {
|
||||
dir := GinkgoT().TempDir()
|
||||
p := filepath.Join(dir, "index.yaml")
|
||||
Expect(os.WriteFile(p, []byte(body), 0o600)).To(Succeed())
|
||||
return p
|
||||
}
|
||||
|
||||
It("reports one local path claimed by two different uris", func() {
|
||||
// The shape that shipped: LiquidAI and unsloth both publish
|
||||
// LFM2.5-8B-A1B-GGUF, so a path built from the bare repo name gives both
|
||||
// entries the same local file under two different checksums.
|
||||
Expect(Verify(write(`
|
||||
- name: lfm2.5-8b-a1b
|
||||
files:
|
||||
- filename: llama-cpp/models/LFM2.5-8B-A1B-GGUF/LFM2.5-8B-A1B-Q8_0.gguf
|
||||
sha256: 33ab3b8c
|
||||
uri: https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-GGUF/resolve/main/LFM2.5-8B-A1B-Q8_0.gguf
|
||||
- name: lfm2.5-8b-a1b-q8-0
|
||||
files:
|
||||
- filename: llama-cpp/models/LFM2.5-8B-A1B-GGUF/LFM2.5-8B-A1B-Q8_0.gguf
|
||||
sha256: ec11666b
|
||||
uri: https://huggingface.co/unsloth/LFM2.5-8B-A1B-GGUF/resolve/main/LFM2.5-8B-A1B-Q8_0.gguf
|
||||
`))).To(ContainElement(SatisfyAll(
|
||||
ContainSubstring("claimed by two different uris"),
|
||||
ContainSubstring("lfm2.5-8b-a1b-q8-0"),
|
||||
)))
|
||||
})
|
||||
|
||||
It("accepts two entries reusing one file from the same uri", func() {
|
||||
// Sibling builds of one repo legitimately share a projector.
|
||||
Expect(Verify(write(`
|
||||
- name: a
|
||||
files:
|
||||
- filename: llama-cpp/mmproj/mudler/Example-GGUF/mmproj-F16.gguf
|
||||
sha256: cc
|
||||
uri: https://huggingface.co/mudler/Example-GGUF/resolve/main/mmproj-F16.gguf
|
||||
- name: b
|
||||
files:
|
||||
- filename: llama-cpp/mmproj/mudler/Example-GGUF/mmproj-F16.gguf
|
||||
sha256: cc
|
||||
uri: https://huggingface.co/mudler/Example-GGUF/resolve/main/mmproj-F16.gguf
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("reports a collision once however many entries pile onto the path", func() {
|
||||
problems := Verify(write(`
|
||||
- name: a
|
||||
files:
|
||||
- filename: shared.gguf
|
||||
sha256: aa
|
||||
uri: https://example.com/a.gguf
|
||||
- name: b
|
||||
files:
|
||||
- filename: shared.gguf
|
||||
sha256: bb
|
||||
uri: https://example.com/b.gguf
|
||||
- name: c
|
||||
files:
|
||||
- filename: shared.gguf
|
||||
sha256: cc
|
||||
uri: https://example.com/c.gguf
|
||||
`))
|
||||
|
||||
var collisions int
|
||||
for _, p := range problems {
|
||||
if strings.Contains(p, "claimed by two different uris") {
|
||||
collisions++
|
||||
}
|
||||
}
|
||||
Expect(collisions).To(Equal(1))
|
||||
})
|
||||
|
||||
It("says nothing about files that carry no uri", func() {
|
||||
// A hand-written entry may record only a checksum. There is no upstream
|
||||
// to compare, so the check cannot conclude anything either way.
|
||||
Expect(Verify(write(`
|
||||
- name: a
|
||||
files:
|
||||
- filename: shared.gguf
|
||||
sha256: aa
|
||||
- name: b
|
||||
files:
|
||||
- filename: shared.gguf
|
||||
sha256: bb
|
||||
`))).To(BeEmpty())
|
||||
})
|
||||
})
|
||||
152
.github/ci/galleryedit/edit.go
vendored
152
.github/ci/galleryedit/edit.go
vendored
@@ -1,152 +0,0 @@
|
||||
// Package galleryedit splices variant references into the LocalAI gallery index
|
||||
// as TEXT.
|
||||
//
|
||||
// Re-serialising the index through a YAML marshaller would reflow 40,000 lines,
|
||||
// drop the anchors and merge keys the gallery relies on, and produce a diff no
|
||||
// reviewer could read, which makes a pull request worthless even when the
|
||||
// content inside it is right. Every generator that adds variants to an entry the
|
||||
// gallery already ships therefore edits lines, and they share this package so
|
||||
// that two of them cannot drift apart on where a variants block belongs.
|
||||
package galleryedit
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strings"
|
||||
)
|
||||
|
||||
var (
|
||||
entryStart = regexp.MustCompile(`^-(?: |$)`)
|
||||
inlineName = regexp.MustCompile(`^- (?:&\S+ )?name:`)
|
||||
keyName = regexp.MustCompile(`^ name:`)
|
||||
keyVariants = regexp.MustCompile(`^ variants:\s*(.*)$`)
|
||||
variantItem = regexp.MustCompile(`^ - `)
|
||||
unsafeInName = regexp.MustCompile(`[:#{}\[\],&*?|>'"%@` + "`" + `]|^\s|\s$`)
|
||||
)
|
||||
|
||||
// Entry is the positional view of one gallery entry: what it is called and
|
||||
// which lines it occupies. Nothing about what the entry MEANS belongs here, so
|
||||
// each caller keeps its own semantic decode and only hands over the coordinates.
|
||||
type Entry struct {
|
||||
Name string
|
||||
// StartLine and EndLine bound the entry, zero based and half open.
|
||||
StartLine int
|
||||
EndLine int
|
||||
}
|
||||
|
||||
// Insert is one entry's pending variants addition. The caller owns the contents
|
||||
// of Variants: this package neither orders nor deduplicates them, because the
|
||||
// right order and the right dedup rule differ between generators.
|
||||
type Insert struct {
|
||||
Entry Entry
|
||||
Variants []string
|
||||
}
|
||||
|
||||
// Scan splits index text into lines and reports the line each top level list
|
||||
// item begins on.
|
||||
func Scan(text string) (lines []string, starts []int) {
|
||||
lines = strings.Split(text, "\n")
|
||||
for i, line := range lines {
|
||||
if entryStart.MatchString(line) {
|
||||
starts = append(starts, i)
|
||||
}
|
||||
}
|
||||
return lines, starts
|
||||
}
|
||||
|
||||
// Apply splices every insert into the index lines and returns the new text.
|
||||
func Apply(lines []string, inserts []Insert) ([]string, error) {
|
||||
type edit struct {
|
||||
at int
|
||||
remove int
|
||||
insert []string
|
||||
}
|
||||
var edits []edit
|
||||
|
||||
for _, in := range inserts {
|
||||
if len(in.Variants) == 0 {
|
||||
continue
|
||||
}
|
||||
|
||||
items := make([]string, 0, len(in.Variants))
|
||||
for _, v := range in.Variants {
|
||||
items = append(items, " - model: "+QuoteName(v))
|
||||
}
|
||||
|
||||
at, remove, err := insertionPoint(lines, in.Entry)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
block := items
|
||||
if remove > 0 || !hasVariantsKey(lines, in.Entry) {
|
||||
block = append([]string{" variants:"}, items...)
|
||||
}
|
||||
edits = append(edits, edit{at: at, remove: remove, insert: block})
|
||||
}
|
||||
|
||||
// Applying from the bottom up keeps every line number computed against the
|
||||
// original text valid while earlier edits are still pending.
|
||||
sort.Slice(edits, func(i, j int) bool { return edits[i].at > edits[j].at })
|
||||
|
||||
out := append([]string(nil), lines...)
|
||||
for _, e := range edits {
|
||||
tail := append([]string(nil), out[e.at+e.remove:]...)
|
||||
out = append(out[:e.at], append(append([]string(nil), e.insert...), tail...)...)
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
func hasVariantsKey(lines []string, e Entry) bool {
|
||||
for i := e.StartLine; i < e.EndLine; i++ {
|
||||
if keyVariants.MatchString(lines[i]) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// insertionPoint reports where new variant items belong, and how many existing
|
||||
// lines the insertion replaces.
|
||||
//
|
||||
// An entry with no variants key gets one right after its name, which is where
|
||||
// the hand-written families put it. An entry with an empty "variants: []" has
|
||||
// that line replaced by a block. An entry with a block gets its items appended.
|
||||
func insertionPoint(lines []string, e Entry) (at int, remove int, err error) {
|
||||
for i := e.StartLine; i < e.EndLine; i++ {
|
||||
m := keyVariants.FindStringSubmatch(lines[i])
|
||||
if m == nil {
|
||||
continue
|
||||
}
|
||||
if strings.TrimSpace(m[1]) == "[]" {
|
||||
return i, 1, nil
|
||||
}
|
||||
if strings.TrimSpace(m[1]) != "" {
|
||||
return 0, 0, fmt.Errorf("entry %q writes its variants inline (%q); this job only edits block lists", e.Name, strings.TrimSpace(m[1]))
|
||||
}
|
||||
last := i
|
||||
for j := i + 1; j < e.EndLine && variantItem.MatchString(lines[j]); j++ {
|
||||
last = j
|
||||
}
|
||||
return last + 1, 0, nil
|
||||
}
|
||||
|
||||
if inlineName.MatchString(lines[e.StartLine]) {
|
||||
return e.StartLine + 1, 0, nil
|
||||
}
|
||||
for i := e.StartLine; i < e.EndLine; i++ {
|
||||
if keyName.MatchString(lines[i]) {
|
||||
return i + 1, 0, nil
|
||||
}
|
||||
}
|
||||
return 0, 0, fmt.Errorf("entry %q has no name line to anchor the insertion to", e.Name)
|
||||
}
|
||||
|
||||
// QuoteName quotes a variant reference when the name would otherwise change
|
||||
// meaning as bare YAML. Config-suffixed names carry a ":" and always need it.
|
||||
func QuoteName(name string) string {
|
||||
if unsafeInName.MatchString(name) {
|
||||
return `"` + strings.ReplaceAll(name, `"`, `\"`) + `"`
|
||||
}
|
||||
return name
|
||||
}
|
||||
133
.github/ci/variantproposals/body.go
vendored
133
.github/ci/variantproposals/body.go
vendored
@@ -1,133 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// RenderBody writes the pull request body.
|
||||
//
|
||||
// The body is the product of this job, not the diff. Grouping is a judgement
|
||||
// call that has gone wrong in both directions before, so a reviewer has to be
|
||||
// able to accept or reject each family from the body alone, without opening
|
||||
// HuggingFace to work out whether two entries hold the same weights.
|
||||
func RenderBody(r *Result, ledgerPath string) string {
|
||||
var b strings.Builder
|
||||
|
||||
b.WriteString("## Proposed gallery variant groupings\n\n")
|
||||
b.WriteString("This is a proposal, not a decision. The gallery agent adds one build per model and never joins an existing family, so entries that are alternative builds of the same weights drift apart as the gallery grows. This job re-applies the grouping heuristics from the manual sweeps and asks a human to confirm.\n\n")
|
||||
b.WriteString("Each family below lists the parent, the variants, and the evidence that they are the same weights. **Reject anything whose evidence you do not believe.**\n\n")
|
||||
b.WriteString(fmt.Sprintf("To decline a family permanently, add one line to `%s` in this pull request and close it:\n\n", ledgerPath))
|
||||
b.WriteString("```yaml\npairs:\n - {parent: some-model, variant: some-model-thing, reason: \"different finetune\"}\n```\n\n")
|
||||
|
||||
b.WriteString(fmt.Sprintf("### Proposed families (%d)\n\n", len(r.Families)))
|
||||
if len(r.Families) == 0 {
|
||||
b.WriteString("None.\n\n")
|
||||
}
|
||||
for _, f := range r.Families {
|
||||
b.WriteString(fmt.Sprintf("#### `%s`\n\n", f.Parent))
|
||||
b.WriteString("| variant | signals | evidence |\n|---|---|---|\n")
|
||||
for _, p := range f.Proposals {
|
||||
b.WriteString(fmt.Sprintf("| `%s` | %s | %s |\n", p.Variant, joinSignals(p.Evidence.Signals), describeEvidence(p.Evidence)))
|
||||
}
|
||||
b.WriteString("\n")
|
||||
}
|
||||
|
||||
b.WriteString(fmt.Sprintf("### Declined by the ledger (%d)\n\n", len(r.Suppressed)))
|
||||
if len(r.Suppressed) == 0 {
|
||||
b.WriteString("Nothing the heuristics found was already on the ledger.\n\n")
|
||||
} else {
|
||||
b.WriteString("Candidates the heuristics found and the ledger has already settled. They are listed so the ledger's effect stays visible rather than silently shrinking the job's output.\n\n")
|
||||
for _, s := range r.Suppressed {
|
||||
b.WriteString(fmt.Sprintf("- `%s` + `%s`: %s\n", s.A, s.B, s.Reason))
|
||||
}
|
||||
b.WriteString("\n")
|
||||
}
|
||||
|
||||
if len(r.AliasSkipped) > 0 {
|
||||
b.WriteString(fmt.Sprintf("### Aliases, not variants (%d)\n\n", len(r.AliasSkipped)))
|
||||
b.WriteString("These entries install byte for byte the same payload. An alias exists so clients can send a particular name; folding it under another entry would hide that name.\n\n")
|
||||
for _, s := range r.AliasSkipped {
|
||||
b.WriteString(fmt.Sprintf("- `%s` + `%s`: %s\n", s.A, s.B, s.Reason))
|
||||
}
|
||||
b.WriteString("\n")
|
||||
}
|
||||
|
||||
if len(r.Refusals) > 0 {
|
||||
b.WriteString(fmt.Sprintf("### Found but refused (%d)\n\n", len(r.Refusals)))
|
||||
b.WriteString("Candidates the heuristics found but the authoring rules would not let this job write. They need a human edit or a rule change.\n\n")
|
||||
for _, ref := range r.Refusals {
|
||||
b.WriteString(fmt.Sprintf("- %s: %s\n", codeList(ref.Members), ref.Reason))
|
||||
}
|
||||
b.WriteString("\n")
|
||||
}
|
||||
|
||||
b.WriteString("---\n\nOpened by `.github/ci/variantproposals`. Heuristics and the rejection ledger live there and in the ledger file; a wrong proposal is a bug in one of the two.\n")
|
||||
return b.String()
|
||||
}
|
||||
|
||||
func joinSignals(signals []Signal) string {
|
||||
if len(signals) == 0 {
|
||||
return "inferred through another member of the family"
|
||||
}
|
||||
out := make([]string, 0, len(signals))
|
||||
for _, s := range signals {
|
||||
out = append(out, "`"+string(s)+"`")
|
||||
}
|
||||
return strings.Join(out, ", ")
|
||||
}
|
||||
|
||||
func describeEvidence(e Evidence) string {
|
||||
var parts []string
|
||||
if e.SharedStem != "" {
|
||||
parts = append(parts, fmt.Sprintf("same name once quantization markers are stripped: `%s`", e.SharedStem))
|
||||
}
|
||||
if e.SharedFile != "" {
|
||||
parts = append(parts, fmt.Sprintf("same primary weight filename once quantization markers are stripped: `%s`", e.SharedFile))
|
||||
}
|
||||
if e.SharedRepo != "" {
|
||||
parts = append(parts, fmt.Sprintf("same upstream repo `%s`", e.SharedRepo))
|
||||
}
|
||||
if len(e.QuantTokens) > 0 {
|
||||
parts = append(parts, "differing quantization tokens: `"+strings.Join(e.QuantTokens, "`, `")+"`")
|
||||
}
|
||||
if len(parts) == 0 {
|
||||
return "reached this family through another member"
|
||||
}
|
||||
return strings.Join(parts, "; ")
|
||||
}
|
||||
|
||||
func codeList(names []string) string {
|
||||
out := make([]string, 0, len(names))
|
||||
for _, n := range names {
|
||||
out = append(out, "`"+n+"`")
|
||||
}
|
||||
return strings.Join(out, " + ")
|
||||
}
|
||||
|
||||
// RenderSummary is the terminal-facing digest of a run, so the workflow log
|
||||
// says what happened without anyone opening the pull request.
|
||||
func RenderSummary(r *Result) string {
|
||||
var b strings.Builder
|
||||
fmt.Fprintf(&b, "families proposed: %d\n", len(r.Families))
|
||||
for _, f := range r.Families {
|
||||
names := make([]string, 0, len(f.Proposals))
|
||||
for _, p := range f.Proposals {
|
||||
names = append(names, p.Variant)
|
||||
}
|
||||
fmt.Fprintf(&b, " %s <- %s\n", f.Parent, strings.Join(names, ", "))
|
||||
}
|
||||
fmt.Fprintf(&b, "declined by ledger: %d\n", len(r.Suppressed))
|
||||
for _, s := range r.Suppressed {
|
||||
fmt.Fprintf(&b, " %s\n", s)
|
||||
}
|
||||
fmt.Fprintf(&b, "aliases skipped: %d\n", len(r.AliasSkipped))
|
||||
for _, s := range r.AliasSkipped {
|
||||
fmt.Fprintf(&b, " %s\n", s)
|
||||
}
|
||||
fmt.Fprintf(&b, "refused: %d\n", len(r.Refusals))
|
||||
for _, ref := range r.Refusals {
|
||||
fmt.Fprintf(&b, " %s: %s\n", strings.Join(ref.Members, " + "), ref.Reason)
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
42
.github/ci/variantproposals/edit.go
vendored
42
.github/ci/variantproposals/edit.go
vendored
@@ -1,42 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"github.com/mudler/LocalAI/.github/ci/galleryedit"
|
||||
)
|
||||
|
||||
// ApplyFamilies writes the proposed variant lists into the index text.
|
||||
//
|
||||
// The line editing itself lives in galleryedit, shared with the apexentries
|
||||
// generator. Both jobs add variants to entries the gallery already ships, and a
|
||||
// second answer to "where does a variants block go" would drift from this one;
|
||||
// see that package for why the edit is textual rather than a YAML round trip.
|
||||
func ApplyFamilies(ix *Index, families []Family) ([]string, error) {
|
||||
byName, _ := ix.ByName()
|
||||
|
||||
var inserts []galleryedit.Insert
|
||||
for _, f := range families {
|
||||
entry, ok := byName[strings.ToLower(f.Parent)]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("parent %q is not in the index", f.Parent)
|
||||
}
|
||||
|
||||
variants := make([]string, 0, len(f.Proposals))
|
||||
for _, p := range f.Proposals {
|
||||
variants = append(variants, p.Variant)
|
||||
}
|
||||
|
||||
inserts = append(inserts, galleryedit.Insert{
|
||||
Entry: galleryedit.Entry{
|
||||
Name: entry.Name,
|
||||
StartLine: entry.StartLine,
|
||||
EndLine: entry.EndLine,
|
||||
},
|
||||
Variants: variants,
|
||||
})
|
||||
}
|
||||
|
||||
return galleryedit.Apply(ix.Lines, inserts)
|
||||
}
|
||||
153
.github/ci/variantproposals/edit_test.go
vendored
153
.github/ci/variantproposals/edit_test.go
vendored
@@ -1,153 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"strings"
|
||||
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
)
|
||||
|
||||
var _ = Describe("ApplyFamilies", func() {
|
||||
apply := func(ix *Index, families []Family) []string {
|
||||
lines, err := ApplyFamilies(ix, families)
|
||||
ExpectWithOffset(1, err).ToNot(HaveOccurred())
|
||||
return lines
|
||||
}
|
||||
|
||||
// insertedLines is what a reviewer would see in the diff. A textual editor
|
||||
// that reflowed the file would show thousands here, which is the failure
|
||||
// this whole approach exists to avoid.
|
||||
insertedLines := func(before, after []string) int {
|
||||
remaining := map[string]int{}
|
||||
for _, l := range before {
|
||||
remaining[l]++
|
||||
}
|
||||
n := 0
|
||||
for _, l := range after {
|
||||
if remaining[l] > 0 {
|
||||
remaining[l]--
|
||||
continue
|
||||
}
|
||||
n++
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
It("adds a variants block right after the entry's name and touches nothing else", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("foo-model", "acme/repo", "foo-model-Q4_K_M.gguf", "aa"),
|
||||
entryYAML("foo-model-q8_0", "acme/repo", "foo-model-Q8_0.gguf", "bb"),
|
||||
)
|
||||
out := apply(ix, []Family{{Parent: "foo-model", Proposals: []Proposal{{Variant: "foo-model-q8_0"}}}})
|
||||
|
||||
Expect(out[0]).To(Equal("- name: foo-model"))
|
||||
Expect(out[1]).To(Equal(" variants:"))
|
||||
Expect(out[2]).To(Equal(" - model: foo-model-q8_0"))
|
||||
Expect(len(out)).To(Equal(len(ix.Lines) + 2))
|
||||
Expect(insertedLines(ix.Lines, out)).To(Equal(2))
|
||||
})
|
||||
|
||||
It("appends to a variants block that already exists", func() {
|
||||
ix := indexOf(`- name: partial
|
||||
variants:
|
||||
- model: partial-q8_0
|
||||
url: u
|
||||
overrides:
|
||||
parameters:
|
||||
model: partial-Q4_K_M.gguf
|
||||
`, entryYAML("partial-f16", "acme/repo", "partial-f16.gguf", "cc"))
|
||||
out := apply(ix, []Family{{Parent: "partial", Proposals: []Proposal{{Variant: "partial-f16"}}}})
|
||||
|
||||
Expect(out[1]).To(Equal(" variants:"))
|
||||
Expect(out[2]).To(Equal(" - model: partial-q8_0"))
|
||||
Expect(out[3]).To(Equal(" - model: partial-f16"))
|
||||
Expect(out[4]).To(Equal(" url: u"))
|
||||
})
|
||||
|
||||
It("replaces an explicit empty list rather than leaving two variants keys", func() {
|
||||
ix := indexOf(`- name: emptied
|
||||
variants: []
|
||||
url: u
|
||||
`, entryYAML("emptied-q8_0", "acme/repo", "emptied-Q8_0.gguf", "cc"))
|
||||
out := apply(ix, []Family{{Parent: "emptied", Proposals: []Proposal{{Variant: "emptied-q8_0"}}}})
|
||||
|
||||
Expect(strings.Join(out[:4], "\n")).To(Equal("- name: emptied\n variants:\n - model: emptied-q8_0\n url: u"))
|
||||
Expect(strings.Count(strings.Join(out, "\n"), "variants:")).To(Equal(1))
|
||||
})
|
||||
|
||||
It("quotes a config-suffixed name so the reference stays a string", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("phi-2-chat", "acme/repo", "phi-2-chat-Q4_K_M.gguf", "aa"),
|
||||
entryYAML("phi-2-chat:Q8_0", "acme/repo", "phi-2-chat-Q8_0.gguf", "bb"),
|
||||
)
|
||||
out := apply(ix, []Family{{Parent: "phi-2-chat", Proposals: []Proposal{{Variant: "phi-2-chat:Q8_0"}}}})
|
||||
Expect(out[2]).To(Equal(` - model: "phi-2-chat:Q8_0"`))
|
||||
|
||||
// The result has to still be a gallery, and the reference has to
|
||||
// resolve to the entry it names.
|
||||
reparsed, err := ParseIndex(strings.Join(out, "\n"))
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(reparsed.Entries[0].Variants).To(ConsistOf(VariantRef{Model: "phi-2-chat:Q8_0"}))
|
||||
})
|
||||
|
||||
It("keeps line numbers correct when several entries are edited at once", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("alpha", "acme/repo", "alpha-Q4_K_M.gguf", "aa"),
|
||||
entryYAML("alpha-q8_0", "acme/repo", "alpha-Q8_0.gguf", "bb"),
|
||||
entryYAML("beta", "acme/repo", "beta-Q4_K_M.gguf", "cc"),
|
||||
entryYAML("beta-q8_0", "acme/repo", "beta-Q8_0.gguf", "dd"),
|
||||
)
|
||||
out := apply(ix, []Family{
|
||||
{Parent: "alpha", Proposals: []Proposal{{Variant: "alpha-q8_0"}}},
|
||||
{Parent: "beta", Proposals: []Proposal{{Variant: "beta-q8_0"}}},
|
||||
})
|
||||
|
||||
reparsed, err := ParseIndex(strings.Join(out, "\n"))
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
Expect(reparsed.Entries).To(HaveLen(4))
|
||||
Expect(reparsed.Entries[0].Variants).To(ConsistOf(VariantRef{Model: "alpha-q8_0"}))
|
||||
Expect(reparsed.Entries[2].Variants).To(ConsistOf(VariantRef{Model: "beta-q8_0"}))
|
||||
Expect(reparsed.Entries[1].Variants).To(BeEmpty())
|
||||
Expect(reparsed.Entries[3].Variants).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("fails loudly rather than editing an entry it cannot find", func() {
|
||||
ix := indexOf(entryYAML("only", "acme/repo", "only-Q4_K_M.gguf", "aa"))
|
||||
_, err := ApplyFamilies(ix, []Family{{Parent: "missing", Proposals: []Proposal{{Variant: "x"}}}})
|
||||
Expect(err).To(MatchError(ContainSubstring("not in the index")))
|
||||
})
|
||||
})
|
||||
|
||||
var _ = Describe("ParseIndex", func() {
|
||||
It("records the anchor an entry defines and the anchor an entry merges", func() {
|
||||
ix := indexOf(`- &anc
|
||||
name: anchored
|
||||
url: u
|
||||
`, `- !!merge <<: *anc
|
||||
name: child
|
||||
`)
|
||||
Expect(ix.Entries[0].AnchorName).To(Equal("anc"))
|
||||
Expect(ix.Entries[1].MergesFrom).To(Equal("anc"))
|
||||
Expect(ix.MergeChildren("anc")).To(HaveLen(1))
|
||||
})
|
||||
|
||||
It("carries merged values into the child, so an inherited variants key is visible", func() {
|
||||
ix := indexOf(`- &anc
|
||||
name: anchored
|
||||
url: u
|
||||
variants:
|
||||
- model: something
|
||||
`, `- !!merge <<: *anc
|
||||
name: child
|
||||
`)
|
||||
Expect(ix.Entries[1].HasVariants()).To(BeTrue())
|
||||
})
|
||||
|
||||
It("refuses a list item that decodes to nothing", func() {
|
||||
// Every line number the editor works from comes from pairing decoded
|
||||
// entries with top level list items. If those two views can disagree,
|
||||
// the editor writes into the wrong entry, so the parse refuses instead.
|
||||
_, err := ParseIndex("- name: one\n url: u\n-\n")
|
||||
Expect(err).To(MatchError(ContainSubstring("empty")))
|
||||
})
|
||||
})
|
||||
281
.github/ci/variantproposals/index.go
vendored
281
.github/ci/variantproposals/index.go
vendored
@@ -1,281 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"gopkg.in/yaml.v3"
|
||||
|
||||
"github.com/mudler/LocalAI/.github/ci/galleryedit"
|
||||
)
|
||||
|
||||
// File is the subset of a gallery file entry the proposer reads.
|
||||
type File struct {
|
||||
Filename string `yaml:"filename"`
|
||||
URI string `yaml:"uri"`
|
||||
SHA256 string `yaml:"sha256"`
|
||||
}
|
||||
|
||||
// VariantRef mirrors the gallery's variant reference.
|
||||
type VariantRef struct {
|
||||
Model string `yaml:"model"`
|
||||
}
|
||||
|
||||
// GalleryEntry is one gallery entry, carrying both the semantics the heuristics need
|
||||
// and the text range the editor needs.
|
||||
//
|
||||
// The two views are kept together deliberately. The editor must not round-trip
|
||||
// the index through a YAML marshaller: the gallery is 40,000 lines and a
|
||||
// reflowed diff cannot be reviewed, which defeats the entire point of a job
|
||||
// whose output is a human decision.
|
||||
type GalleryEntry struct {
|
||||
Name string `yaml:"name"`
|
||||
URL string `yaml:"url"`
|
||||
ConfigFile map[string]any `yaml:"config_file"`
|
||||
Overrides map[string]any `yaml:"overrides"`
|
||||
Files []File `yaml:"files"`
|
||||
Variants []VariantRef `yaml:"variants"`
|
||||
|
||||
// Index is the entry's position in gallery order.
|
||||
Index int `yaml:"-"`
|
||||
// StartLine and EndLine bound the entry's lines, zero based and half open.
|
||||
StartLine int `yaml:"-"`
|
||||
EndLine int `yaml:"-"`
|
||||
// AnchorName is set when the entry defines a YAML anchor. Adding a variants
|
||||
// key to such an entry is inherited by everything that merges it, which is
|
||||
// why proposals involving anchors get special treatment.
|
||||
AnchorName string `yaml:"-"`
|
||||
// MergesFrom is the anchor this entry pulls in with "!!merge <<:".
|
||||
MergesFrom string `yaml:"-"`
|
||||
}
|
||||
|
||||
// Index is a parsed gallery index: entries plus the exact lines they came from.
|
||||
type Index struct {
|
||||
Lines []string
|
||||
Entries []*GalleryEntry
|
||||
}
|
||||
|
||||
var (
|
||||
anchorStart = regexp.MustCompile(`^- &(\S+)`)
|
||||
mergeStart = regexp.MustCompile(`^- !!merge <<: \*(\S+)`)
|
||||
)
|
||||
|
||||
// LoadIndex reads and parses a gallery index file.
|
||||
func LoadIndex(path string) (*Index, error) {
|
||||
data, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return ParseIndex(string(data))
|
||||
}
|
||||
|
||||
// ParseIndex builds an Index from the raw text of a gallery index.
|
||||
//
|
||||
// The YAML decode and the textual scan are cross checked against each other: if
|
||||
// they disagree on how many entries there are, every line number the editor
|
||||
// would use is suspect, so the run fails rather than editing the wrong entry.
|
||||
func ParseIndex(text string) (*Index, error) {
|
||||
var entries []*GalleryEntry
|
||||
if err := yaml.Unmarshal([]byte(text), &entries); err != nil {
|
||||
return nil, fmt.Errorf("decoding gallery index: %w", err)
|
||||
}
|
||||
|
||||
lines, starts := galleryedit.Scan(text)
|
||||
if len(starts) != len(entries) {
|
||||
return nil, fmt.Errorf("gallery index has %d decoded entries but %d top level list items; refusing to edit by line number", len(entries), len(starts))
|
||||
}
|
||||
|
||||
for i, e := range entries {
|
||||
if e == nil {
|
||||
return nil, fmt.Errorf("gallery index list item %d is empty; refusing to edit by line number", i)
|
||||
}
|
||||
e.Index = i
|
||||
e.StartLine = starts[i]
|
||||
if i+1 < len(starts) {
|
||||
e.EndLine = starts[i+1]
|
||||
} else {
|
||||
e.EndLine = len(lines)
|
||||
}
|
||||
if m := anchorStart.FindStringSubmatch(lines[e.StartLine]); m != nil {
|
||||
e.AnchorName = m[1]
|
||||
}
|
||||
if m := mergeStart.FindStringSubmatch(lines[e.StartLine]); m != nil {
|
||||
e.MergesFrom = m[1]
|
||||
}
|
||||
}
|
||||
|
||||
return &Index{Lines: lines, Entries: entries}, nil
|
||||
}
|
||||
|
||||
// MergeChildren lists the entries that pull in the given anchor.
|
||||
func (ix *Index) MergeChildren(anchor string) []*GalleryEntry {
|
||||
var out []*GalleryEntry
|
||||
for _, e := range ix.Entries {
|
||||
if e.MergesFrom == anchor {
|
||||
out = append(out, e)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// ByName indexes entries by lowercased name. A name appearing twice keeps the
|
||||
// first occurrence, matching the gallery's own first-match-wins resolution, and
|
||||
// the duplicates are returned so the caller can refuse to touch them: a
|
||||
// proposal naming an ambiguous entry cannot be reviewed.
|
||||
func (ix *Index) ByName() (map[string]*GalleryEntry, map[string]int) {
|
||||
byName := make(map[string]*GalleryEntry, len(ix.Entries))
|
||||
counts := make(map[string]int, len(ix.Entries))
|
||||
for _, e := range ix.Entries {
|
||||
key := strings.ToLower(e.Name)
|
||||
counts[key]++
|
||||
if _, seen := byName[key]; !seen {
|
||||
byName[key] = e
|
||||
}
|
||||
}
|
||||
dupes := map[string]int{}
|
||||
for name, n := range counts {
|
||||
if n > 1 {
|
||||
dupes[name] = n
|
||||
}
|
||||
}
|
||||
return byName, dupes
|
||||
}
|
||||
|
||||
// Installable reports whether installing this entry would put anything on disk.
|
||||
// A variant target that installs nothing is a dead end for the selector, so it
|
||||
// is never proposed as one.
|
||||
func (e *GalleryEntry) Installable() bool {
|
||||
return e.URL != "" || len(e.ConfigFile) > 0 || len(e.Overrides) > 0 || len(e.Files) > 0
|
||||
}
|
||||
|
||||
// HasVariants reports whether the entry already offers builds of its own. Such
|
||||
// an entry cannot be a variant target: nesting is what the gallery's own
|
||||
// resolution refuses.
|
||||
func (e *GalleryEntry) HasVariants() bool {
|
||||
return len(e.Variants) > 0
|
||||
}
|
||||
|
||||
// auxiliaryFile matches the shared side files that several unrelated models
|
||||
// legitimately hand out the same copy of. Grouping on one of these is how an
|
||||
// earlier sweep linked four wan-2.1 entries to each other and Z-Image-Turbo to
|
||||
// qwen3-4b: they shared a text encoder, not weights.
|
||||
var auxiliaryFile = regexp.MustCompile(`(?i)(mmproj|vae|clip|t5|umt5|text_?encoder|tokenizer|\bae\b|^ae\.|scheduler|config)`)
|
||||
|
||||
// IsAuxiliaryFile reports whether a filename is a side file rather than the
|
||||
// model's own weights.
|
||||
func IsAuxiliaryFile(filename string) bool {
|
||||
base := filename
|
||||
if i := strings.LastIndex(base, "/"); i >= 0 {
|
||||
base = base[i+1:]
|
||||
}
|
||||
return auxiliaryFile.MatchString(base)
|
||||
}
|
||||
|
||||
// PrimaryWeightFile returns the filename of the entry's own weights, and
|
||||
// whether one could be identified unambiguously.
|
||||
//
|
||||
// The declared overrides.parameters.model wins because that is the file the
|
||||
// backend is actually pointed at. Falling back to the file list only works when
|
||||
// exactly one non-auxiliary file is present; anything else is ambiguous, and
|
||||
// guessing is precisely the failure mode this heuristic has already had.
|
||||
func (e *GalleryEntry) PrimaryWeightFile() (string, bool) {
|
||||
if params, ok := e.Overrides["parameters"].(map[string]any); ok {
|
||||
if model, ok := params["model"].(string); ok && model != "" && !IsAuxiliaryFile(model) {
|
||||
return model, true
|
||||
}
|
||||
}
|
||||
var candidates []string
|
||||
for _, f := range e.Files {
|
||||
if f.Filename == "" || IsAuxiliaryFile(f.Filename) {
|
||||
continue
|
||||
}
|
||||
candidates = append(candidates, f.Filename)
|
||||
}
|
||||
if len(candidates) == 1 {
|
||||
return candidates[0], true
|
||||
}
|
||||
return "", false
|
||||
}
|
||||
|
||||
// SourceRepo returns the upstream repository the entry's files come from, as a
|
||||
// coarse "host + owner + repo" key.
|
||||
func (e *GalleryEntry) SourceRepo() string {
|
||||
for _, f := range e.Files {
|
||||
if f.URI == "" {
|
||||
continue
|
||||
}
|
||||
return repoKey(f.URI)
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
func repoKey(uri string) string {
|
||||
u := strings.ToLower(uri)
|
||||
u = strings.TrimPrefix(u, "huggingface://")
|
||||
u = strings.TrimPrefix(u, "https://huggingface.co/")
|
||||
u = strings.TrimPrefix(u, "http://huggingface.co/")
|
||||
parts := strings.Split(u, "/")
|
||||
if len(parts) >= 2 {
|
||||
return parts[0] + "/" + parts[1]
|
||||
}
|
||||
return u
|
||||
}
|
||||
|
||||
// SameInstallPayload reports whether two entries install byte for byte the same
|
||||
// thing.
|
||||
//
|
||||
// Entries like this are aliases, not variants. whisper-1 exists so a client
|
||||
// speaking the OpenAI API can send that name and get whisper-base; folding it
|
||||
// under whisper-base as a variant would hide the very name clients send.
|
||||
func SameInstallPayload(a, b *GalleryEntry) bool {
|
||||
if a.URL != b.URL {
|
||||
return false
|
||||
}
|
||||
if !sameYAML(a.Overrides, b.Overrides) || !sameYAML(a.ConfigFile, b.ConfigFile) {
|
||||
return false
|
||||
}
|
||||
return sameChecksums(a.Files, b.Files)
|
||||
}
|
||||
|
||||
func sameChecksums(a, b []File) bool {
|
||||
if len(a) != len(b) || len(a) == 0 {
|
||||
return false
|
||||
}
|
||||
ha := make([]string, 0, len(a))
|
||||
hb := make([]string, 0, len(b))
|
||||
for _, f := range a {
|
||||
if f.SHA256 == "" {
|
||||
return false
|
||||
}
|
||||
ha = append(ha, f.SHA256)
|
||||
}
|
||||
for _, f := range b {
|
||||
if f.SHA256 == "" {
|
||||
return false
|
||||
}
|
||||
hb = append(hb, f.SHA256)
|
||||
}
|
||||
sort.Strings(ha)
|
||||
sort.Strings(hb)
|
||||
for i := range ha {
|
||||
if ha[i] != hb[i] {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
func sameYAML(a, b any) bool {
|
||||
ba, err := yaml.Marshal(a)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
bb, err := yaml.Marshal(b)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
return string(ba) == string(bb)
|
||||
}
|
||||
180
.github/ci/variantproposals/ledger.go
vendored
180
.github/ci/variantproposals/ledger.go
vendored
@@ -1,180 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"gopkg.in/yaml.v3"
|
||||
)
|
||||
|
||||
// Ledger records the grouping decisions a human has already made against the
|
||||
// proposer, so a declined candidate stays declined instead of coming back every
|
||||
// night until reviewers stop reading the job's pull requests.
|
||||
//
|
||||
// It is checked in next to the gallery and is meant to be edited inside the
|
||||
// proposal pull request itself: declining a family is adding one flow-mapping
|
||||
// line under pairs or groups and closing the PR.
|
||||
type Ledger struct {
|
||||
// Tokens are name segments that mark a distinct model rather than another
|
||||
// build of the same one: finetune names, language codes, product suffixes.
|
||||
// A candidate whose two names differ by any of these is never proposed.
|
||||
Tokens []LedgerToken `yaml:"tokens"`
|
||||
// Pairs are individual candidates a human considered and declined. Order
|
||||
// does not matter: the pair is matched both ways round.
|
||||
Pairs []LedgerPair `yaml:"pairs"`
|
||||
// Groups decline every pair drawn from a set at once, for families like a
|
||||
// per-language release where listing each pair would be unreadable.
|
||||
Groups []LedgerGroup `yaml:"groups"`
|
||||
}
|
||||
|
||||
type LedgerToken struct {
|
||||
Token string `yaml:"token"`
|
||||
Reason string `yaml:"reason"`
|
||||
}
|
||||
|
||||
type LedgerPair struct {
|
||||
Parent string `yaml:"parent"`
|
||||
Variant string `yaml:"variant"`
|
||||
Reason string `yaml:"reason"`
|
||||
}
|
||||
|
||||
type LedgerGroup struct {
|
||||
Members []string `yaml:"members"`
|
||||
Reason string `yaml:"reason"`
|
||||
}
|
||||
|
||||
// LoadLedger reads a ledger file. A missing file is not an error: a gallery
|
||||
// that has declined nothing yet is a legitimate state, and failing the job over
|
||||
// it would only teach people to keep an empty file around.
|
||||
func LoadLedger(path string) (*Ledger, error) {
|
||||
data, err := os.ReadFile(path)
|
||||
if os.IsNotExist(err) {
|
||||
return &Ledger{}, nil
|
||||
}
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return ParseLedger(data)
|
||||
}
|
||||
|
||||
func ParseLedger(data []byte) (*Ledger, error) {
|
||||
l := &Ledger{}
|
||||
if err := yaml.Unmarshal(data, l); err != nil {
|
||||
return nil, fmt.Errorf("parsing ledger: %w", err)
|
||||
}
|
||||
for i, t := range l.Tokens {
|
||||
if strings.TrimSpace(t.Token) == "" {
|
||||
return nil, fmt.Errorf("ledger tokens[%d] has an empty token", i)
|
||||
}
|
||||
}
|
||||
for i, p := range l.Pairs {
|
||||
if strings.TrimSpace(p.Parent) == "" || strings.TrimSpace(p.Variant) == "" {
|
||||
return nil, fmt.Errorf("ledger pairs[%d] needs both parent and variant", i)
|
||||
}
|
||||
}
|
||||
return l, nil
|
||||
}
|
||||
|
||||
// Suppression is a ledger hit: why a candidate was not proposed, in words a
|
||||
// reviewer can check against the ledger file.
|
||||
type Suppression struct {
|
||||
A string
|
||||
B string
|
||||
Reason string
|
||||
}
|
||||
|
||||
func (s Suppression) String() string {
|
||||
return fmt.Sprintf("%s + %s: %s", s.A, s.B, s.Reason)
|
||||
}
|
||||
|
||||
// Suppresses reports whether the ledger has already declined pairing these two
|
||||
// entries, and why.
|
||||
//
|
||||
// The token rule is applied to the segments the two names do not share. Two
|
||||
// builds of the same weights differ only in quantization markers, so any
|
||||
// ledgered token showing up in that difference is by construction a claim that
|
||||
// the entries are different models.
|
||||
func (l *Ledger) Suppresses(a, b string) (Suppression, bool) {
|
||||
la, lb := strings.ToLower(a), strings.ToLower(b)
|
||||
for _, p := range l.Pairs {
|
||||
lp, lv := strings.ToLower(p.Parent), strings.ToLower(p.Variant)
|
||||
if (lp == la && lv == lb) || (lp == lb && lv == la) {
|
||||
return Suppression{A: a, B: b, Reason: p.Reason}, true
|
||||
}
|
||||
}
|
||||
for _, g := range l.Groups {
|
||||
var seenA, seenB bool
|
||||
for _, m := range g.Members {
|
||||
lm := strings.ToLower(m)
|
||||
if lm == la {
|
||||
seenA = true
|
||||
}
|
||||
if lm == lb {
|
||||
seenB = true
|
||||
}
|
||||
}
|
||||
if seenA && seenB {
|
||||
return Suppression{A: a, B: b, Reason: g.Reason}, true
|
||||
}
|
||||
}
|
||||
diff := differingSegments(la, lb)
|
||||
for _, t := range l.Tokens {
|
||||
token := strings.ToLower(strings.TrimSpace(t.Token))
|
||||
if _, ok := diff[token]; ok {
|
||||
reason := t.Reason
|
||||
if reason == "" {
|
||||
reason = fmt.Sprintf("names differ by %q", token)
|
||||
}
|
||||
return Suppression{A: a, B: b, Reason: fmt.Sprintf("%s (token %q)", reason, token)}, true
|
||||
}
|
||||
}
|
||||
return Suppression{}, false
|
||||
}
|
||||
|
||||
// segments splits a name into the atoms the token rules are written against.
|
||||
func segments(name string) []string {
|
||||
fields := strings.FieldsFunc(strings.ToLower(name), func(r rune) bool {
|
||||
return r == '-' || r == '_' || r == '.' || r == ':' || r == '/'
|
||||
})
|
||||
return fields
|
||||
}
|
||||
|
||||
// differingSegments returns the set of segments present in exactly one of the
|
||||
// two names.
|
||||
func differingSegments(a, b string) map[string]struct{} {
|
||||
setA := map[string]int{}
|
||||
for _, s := range segments(a) {
|
||||
setA[s]++
|
||||
}
|
||||
setB := map[string]int{}
|
||||
for _, s := range segments(b) {
|
||||
setB[s]++
|
||||
}
|
||||
diff := map[string]struct{}{}
|
||||
for s := range setA {
|
||||
if setB[s] == 0 {
|
||||
diff[s] = struct{}{}
|
||||
}
|
||||
}
|
||||
for s := range setB {
|
||||
if setA[s] == 0 {
|
||||
diff[s] = struct{}{}
|
||||
}
|
||||
}
|
||||
return diff
|
||||
}
|
||||
|
||||
// SortedSuppressions gives the ledger's effect on one run in a stable order, so
|
||||
// the pull request body reads the same way for the same gallery.
|
||||
func SortedSuppressions(in []Suppression) []Suppression {
|
||||
out := append([]Suppression(nil), in...)
|
||||
sort.Slice(out, func(i, j int) bool {
|
||||
if out[i].A != out[j].A {
|
||||
return out[i].A < out[j].A
|
||||
}
|
||||
return out[i].B < out[j].B
|
||||
})
|
||||
return out
|
||||
}
|
||||
65
.github/ci/variantproposals/main.go
vendored
65
.github/ci/variantproposals/main.go
vendored
@@ -1,65 +0,0 @@
|
||||
// Command variant-proposals looks for gallery entries that are alternative
|
||||
// builds of the same weights but are not grouped under one another, and writes
|
||||
// a proposal for a human to accept or reject.
|
||||
//
|
||||
// It never decides. Grouping has gone wrong repeatedly in both directions, so
|
||||
// the job's value is catching drift and surfacing candidates with their
|
||||
// evidence, not automating the call. The scheduled workflow feeds its output to
|
||||
// a pull request in the same shape as .github/checksum_checker.sh.
|
||||
package main
|
||||
|
||||
import (
|
||||
"flag"
|
||||
"fmt"
|
||||
"os"
|
||||
"strings"
|
||||
)
|
||||
|
||||
func main() {
|
||||
index := flag.String("index", "gallery/index.yaml", "path to the gallery index")
|
||||
ledger := flag.String("ledger", "gallery/variant-exclusions.yaml", "path to the rejection ledger")
|
||||
bodyOut := flag.String("body-out", "", "write the pull request body here")
|
||||
apply := flag.Bool("apply", false, "write the proposed groupings back into the index")
|
||||
flag.Parse()
|
||||
|
||||
if err := run(*index, *ledger, *bodyOut, *apply); err != nil {
|
||||
fmt.Fprintln(os.Stderr, "variant-proposals:", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
func run(indexPath, ledgerPath, bodyOut string, apply bool) error {
|
||||
ix, err := LoadIndex(indexPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
ledger, err := LoadLedger(ledgerPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
result := Propose(ix, ledger)
|
||||
fmt.Print(RenderSummary(result))
|
||||
|
||||
if !result.HasProposals() {
|
||||
// An empty pull request every night is how a proposal job gets muted.
|
||||
fmt.Println("nothing to propose")
|
||||
return nil
|
||||
}
|
||||
|
||||
if bodyOut != "" {
|
||||
if err := os.WriteFile(bodyOut, []byte(RenderBody(result, ledgerPath)), 0o644); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
|
||||
if !apply {
|
||||
return nil
|
||||
}
|
||||
|
||||
lines, err := ApplyFamilies(ix, result.Families)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
return os.WriteFile(indexPath, []byte(strings.Join(lines, "\n")), 0o644)
|
||||
}
|
||||
615
.github/ci/variantproposals/propose.go
vendored
615
.github/ci/variantproposals/propose.go
vendored
@@ -1,615 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// Signal names the grouping heuristic that linked two entries.
|
||||
type Signal string
|
||||
|
||||
const (
|
||||
// SignalName is "same name once quantization markers are stripped".
|
||||
SignalName Signal = "name-modulo-quant"
|
||||
// SignalConfigSuffix is the ":" convention, foo:q8_0 as a build of foo.
|
||||
SignalConfigSuffix Signal = "config-suffix"
|
||||
// SignalWeightFile is "same primary weight filename once quantization
|
||||
// markers are stripped", auxiliary files excluded.
|
||||
SignalWeightFile Signal = "weight-filename"
|
||||
)
|
||||
|
||||
// Evidence is what a reviewer needs in order to agree or disagree without
|
||||
// opening HuggingFace: what the two entries share, and what differs.
|
||||
type Evidence struct {
|
||||
Signals []Signal
|
||||
SharedStem string
|
||||
SharedFile string
|
||||
SharedRepo string
|
||||
QuantTokens []string
|
||||
}
|
||||
|
||||
// Proposal is one variant target offered to one parent.
|
||||
type Proposal struct {
|
||||
Variant string
|
||||
Evidence Evidence
|
||||
}
|
||||
|
||||
// Family is a complete proposal: one parent gaining one or more variants.
|
||||
type Family struct {
|
||||
Parent string
|
||||
Proposals []Proposal
|
||||
}
|
||||
|
||||
// Refusal is a family the heuristics found but the rules would not let through.
|
||||
// Refusals are reported rather than dropped: a candidate the job keeps refusing
|
||||
// is either a rule worth revisiting or a gallery bug worth fixing.
|
||||
type Refusal struct {
|
||||
Members []string
|
||||
Reason string
|
||||
}
|
||||
|
||||
// Result is one run of the proposer.
|
||||
type Result struct {
|
||||
Families []Family
|
||||
Refusals []Refusal
|
||||
Suppressed []Suppression
|
||||
AliasSkipped []Suppression
|
||||
}
|
||||
|
||||
// HasProposals reports whether the run found anything to open a pull request
|
||||
// about. A job that opens an empty pull request every night is a job people
|
||||
// filter out of their inbox.
|
||||
func (r *Result) HasProposals() bool {
|
||||
return len(r.Families) > 0
|
||||
}
|
||||
|
||||
// sizeToken matches a parameter-count marker: 8b, 1.7b, a3b for an active
|
||||
// expert count, e2b for the Gemma effective sizes, 8x7b for a mixture.
|
||||
//
|
||||
// This is a structural rule rather than a ledger entry because it is about the
|
||||
// shape of the token, not about any one model. Different parameter sizes were
|
||||
// mis-grouped by an earlier sweep and the failure is systematic.
|
||||
var sizeToken = regexp.MustCompile(`^(?:[0-9]+(?:\.[0-9]+)?[bm]|[ae][0-9]+(?:\.[0-9]+)?b|[0-9]+x[0-9]+(?:\.[0-9]+)?b)$`)
|
||||
|
||||
func differsByParameterSize(a, b string) (string, bool) {
|
||||
for seg := range differingSegments(a, b) {
|
||||
if sizeToken.MatchString(seg) {
|
||||
return seg, true
|
||||
}
|
||||
}
|
||||
return "", false
|
||||
}
|
||||
|
||||
// genericFileStem lists weight filenames too generic to be evidence of
|
||||
// anything. Two entries both shipping "model.safetensors" share a convention,
|
||||
// not a set of weights.
|
||||
var genericFileStem = map[string]struct{}{
|
||||
"model": {}, "weights": {}, "pytorch_model": {}, "diffusion_pytorch_model": {},
|
||||
"consolidated": {}, "ggml-model": {}, "model-00001-of-00002": {},
|
||||
}
|
||||
|
||||
// minFileStemLength keeps short, collision-prone filename stems from linking
|
||||
// unrelated entries.
|
||||
const minFileStemLength = 6
|
||||
|
||||
type pair struct {
|
||||
a, b int
|
||||
evidence Evidence
|
||||
}
|
||||
|
||||
// Propose runs the grouping heuristics over a gallery index and returns what it
|
||||
// would offer a human, what it refused, and what the ledger silenced.
|
||||
//
|
||||
// Nothing here touches the network or git, and the index is not modified.
|
||||
func Propose(ix *Index, ledger *Ledger) *Result {
|
||||
if ledger == nil {
|
||||
ledger = &Ledger{}
|
||||
}
|
||||
result := &Result{}
|
||||
|
||||
byName, dupes := ix.ByName()
|
||||
|
||||
// Existing relationships. A target already claimed must not be claimed
|
||||
// again, and two entries already in one family need no proposal.
|
||||
claimedBy := map[string]string{}
|
||||
familyOf := map[string]string{}
|
||||
for _, e := range ix.Entries {
|
||||
if !e.HasVariants() {
|
||||
continue
|
||||
}
|
||||
familyOf[strings.ToLower(e.Name)] = strings.ToLower(e.Name)
|
||||
for _, v := range e.Variants {
|
||||
target := strings.ToLower(v.Model)
|
||||
if _, taken := claimedBy[target]; !taken {
|
||||
claimedBy[target] = strings.ToLower(e.Name)
|
||||
}
|
||||
familyOf[target] = strings.ToLower(e.Name)
|
||||
}
|
||||
}
|
||||
|
||||
candidates := map[[2]int]*Evidence{}
|
||||
|
||||
addPair := func(i, j int, sig Signal, apply func(*Evidence)) {
|
||||
if i == j {
|
||||
return
|
||||
}
|
||||
if i > j {
|
||||
i, j = j, i
|
||||
}
|
||||
key := [2]int{i, j}
|
||||
ev, ok := candidates[key]
|
||||
if !ok {
|
||||
ev = &Evidence{}
|
||||
candidates[key] = ev
|
||||
}
|
||||
for _, s := range ev.Signals {
|
||||
if s == sig {
|
||||
apply(ev)
|
||||
return
|
||||
}
|
||||
}
|
||||
ev.Signals = append(ev.Signals, sig)
|
||||
apply(ev)
|
||||
}
|
||||
|
||||
// Signal 1 and 2: entries sharing a name stem.
|
||||
byStem := map[string][]int{}
|
||||
for _, e := range ix.Entries {
|
||||
if e.Name == "" {
|
||||
continue
|
||||
}
|
||||
byStem[NameStem(e.Name)] = append(byStem[NameStem(e.Name)], e.Index)
|
||||
}
|
||||
for stem, members := range byStem {
|
||||
if len(members) < 2 {
|
||||
continue
|
||||
}
|
||||
for i := 0; i < len(members); i++ {
|
||||
for j := i + 1; j < len(members); j++ {
|
||||
a, b := ix.Entries[members[i]], ix.Entries[members[j]]
|
||||
sig := SignalName
|
||||
if HasConfigSuffix(a.Name) || HasConfigSuffix(b.Name) {
|
||||
sig = SignalConfigSuffix
|
||||
}
|
||||
// The bare parent carries no marker in its name, so the
|
||||
// evidence would read "differs by q8_0" and say nothing about
|
||||
// what the parent is. The weight filenames fill that in.
|
||||
fa, _ := a.PrimaryWeightFile()
|
||||
fb, _ := b.PrimaryWeightFile()
|
||||
addPair(members[i], members[j], sig, func(ev *Evidence) {
|
||||
ev.SharedStem = stem
|
||||
ev.QuantTokens = quantDifference(a.Name, b.Name, fa, fb)
|
||||
})
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Signal 3: entries whose own weight file is the same file at a different
|
||||
// quantization. Auxiliary files never take part.
|
||||
byFile := map[string][]int{}
|
||||
for _, e := range ix.Entries {
|
||||
primary, ok := e.PrimaryWeightFile()
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
stem := FileStem(primary)
|
||||
if len(stem) < minFileStemLength {
|
||||
continue
|
||||
}
|
||||
if _, generic := genericFileStem[stem]; generic {
|
||||
continue
|
||||
}
|
||||
byFile[stem] = append(byFile[stem], e.Index)
|
||||
}
|
||||
for stem, members := range byFile {
|
||||
if len(members) < 2 {
|
||||
continue
|
||||
}
|
||||
for i := 0; i < len(members); i++ {
|
||||
for j := i + 1; j < len(members); j++ {
|
||||
a, b := ix.Entries[members[i]], ix.Entries[members[j]]
|
||||
// The filename alone is not evidence. Publishers reuse the
|
||||
// upstream filename for finetunes and for models that merely
|
||||
// embed the base weights: bert-embeddings, an ultravox audio
|
||||
// model and a roleplay finetune all ship a file called
|
||||
// llama-3.2-1b-instruct-q4_k_m.gguf. Requiring the same
|
||||
// upstream repository turns the signal back into what it
|
||||
// claims to be, one repo publishing one file at two
|
||||
// quantizations. Two repos holding the same weights is a fact
|
||||
// no filename proves, so it stays a human call.
|
||||
repo := a.SourceRepo()
|
||||
if repo == "" || repo != b.SourceRepo() {
|
||||
continue
|
||||
}
|
||||
fa, _ := a.PrimaryWeightFile()
|
||||
fb, _ := b.PrimaryWeightFile()
|
||||
addPair(members[i], members[j], SignalWeightFile, func(ev *Evidence) {
|
||||
ev.SharedFile = stem
|
||||
ev.SharedRepo = repo
|
||||
if len(ev.QuantTokens) == 0 {
|
||||
ev.QuantTokens = quantDifference(fa, fb)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Filter candidates. Everything dropped here is dropped for a reason a
|
||||
// reviewer can read back off the ledger or the rules.
|
||||
var kept []pair
|
||||
for key, ev := range candidates {
|
||||
a, b := ix.Entries[key[0]], ix.Entries[key[1]]
|
||||
la, lb := strings.ToLower(a.Name), strings.ToLower(b.Name)
|
||||
if la == lb {
|
||||
continue
|
||||
}
|
||||
if dupes[la] > 0 || dupes[lb] > 0 {
|
||||
result.Refusals = append(result.Refusals, Refusal{
|
||||
Members: []string{a.Name, b.Name},
|
||||
Reason: "one of these names appears more than once in the gallery, so a variant reference to it is ambiguous",
|
||||
})
|
||||
continue
|
||||
}
|
||||
if fa, fb := familyOf[la], familyOf[lb]; fa != "" && fa == fb {
|
||||
continue
|
||||
}
|
||||
if seg, differs := differsByParameterSize(la, lb); differs {
|
||||
result.Suppressed = append(result.Suppressed, Suppression{
|
||||
A: a.Name, B: b.Name, Reason: fmt.Sprintf("different parameter sizes (segment %q)", seg),
|
||||
})
|
||||
continue
|
||||
}
|
||||
if s, ok := ledger.Suppresses(a.Name, b.Name); ok {
|
||||
result.Suppressed = append(result.Suppressed, s)
|
||||
continue
|
||||
}
|
||||
if SameInstallPayload(a, b) {
|
||||
result.AliasSkipped = append(result.AliasSkipped, Suppression{
|
||||
A: a.Name, B: b.Name,
|
||||
Reason: "identical install payload; these are aliases of one build, not alternative builds",
|
||||
})
|
||||
continue
|
||||
}
|
||||
kept = append(kept, pair{a: key[0], b: key[1], evidence: *ev})
|
||||
}
|
||||
|
||||
sort.Slice(kept, func(i, j int) bool {
|
||||
if kept[i].a != kept[j].a {
|
||||
return kept[i].a < kept[j].a
|
||||
}
|
||||
return kept[i].b < kept[j].b
|
||||
})
|
||||
|
||||
// Components. A pair from either signal joins the same family, so a chain
|
||||
// of alternative builds discovered by different signals stays one family
|
||||
// rather than two overlapping ones that would double claim a target.
|
||||
parent := map[int]int{}
|
||||
var find func(int) int
|
||||
find = func(x int) int {
|
||||
if p, ok := parent[x]; ok && p != x {
|
||||
parent[x] = find(p)
|
||||
return parent[x]
|
||||
}
|
||||
if _, ok := parent[x]; !ok {
|
||||
parent[x] = x
|
||||
}
|
||||
return parent[x]
|
||||
}
|
||||
union := func(x, y int) {
|
||||
rx, ry := find(x), find(y)
|
||||
if rx != ry {
|
||||
parent[ry] = rx
|
||||
}
|
||||
}
|
||||
evidenceFor := map[[2]int]Evidence{}
|
||||
for _, p := range kept {
|
||||
union(p.a, p.b)
|
||||
evidenceFor[[2]int{p.a, p.b}] = p.evidence
|
||||
}
|
||||
|
||||
components := map[int][]int{}
|
||||
for _, p := range kept {
|
||||
for _, m := range []int{p.a, p.b} {
|
||||
root := find(m)
|
||||
if !contains(components[root], m) {
|
||||
components[root] = append(components[root], m)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
roots := make([]int, 0, len(components))
|
||||
for r := range components {
|
||||
roots = append(roots, r)
|
||||
}
|
||||
sort.Ints(roots)
|
||||
|
||||
proposedTargets := map[string]string{}
|
||||
for _, root := range roots {
|
||||
members := components[root]
|
||||
sort.Ints(members)
|
||||
family, refusal := buildFamily(ix, members, evidenceFor, claimedBy, proposedTargets, byName)
|
||||
if refusal != nil {
|
||||
result.Refusals = append(result.Refusals, *refusal)
|
||||
continue
|
||||
}
|
||||
if family == nil {
|
||||
continue
|
||||
}
|
||||
for _, p := range family.Proposals {
|
||||
proposedTargets[strings.ToLower(p.Variant)] = family.Parent
|
||||
}
|
||||
result.Families = append(result.Families, *family)
|
||||
}
|
||||
|
||||
sort.Slice(result.Families, func(i, j int) bool { return result.Families[i].Parent < result.Families[j].Parent })
|
||||
result.Suppressed = SortedSuppressions(result.Suppressed)
|
||||
result.AliasSkipped = SortedSuppressions(result.AliasSkipped)
|
||||
result.Refusals = dedupeRefusals(result.Refusals)
|
||||
return result
|
||||
}
|
||||
|
||||
// dedupeRefusals collapses the same refusal reached from both orderings of a
|
||||
// pair, and sorts what is left. A reviewer reading the same complaint twice
|
||||
// learns to skim the section.
|
||||
func dedupeRefusals(in []Refusal) []Refusal {
|
||||
seen := map[string]struct{}{}
|
||||
var out []Refusal
|
||||
for _, r := range in {
|
||||
members := append([]string(nil), r.Members...)
|
||||
sort.Strings(members)
|
||||
key := strings.Join(members, "\x00") + "\x00" + r.Reason
|
||||
if _, dup := seen[key]; dup {
|
||||
continue
|
||||
}
|
||||
seen[key] = struct{}{}
|
||||
out = append(out, r)
|
||||
}
|
||||
sort.Slice(out, func(i, j int) bool {
|
||||
if a, b := strings.Join(out[i].Members, ","), strings.Join(out[j].Members, ","); a != b {
|
||||
return a < b
|
||||
}
|
||||
return out[i].Reason < out[j].Reason
|
||||
})
|
||||
return out
|
||||
}
|
||||
|
||||
func contains(xs []int, x int) bool {
|
||||
for _, v := range xs {
|
||||
if v == x {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// buildFamily turns a connected component into a proposal, or refuses it.
|
||||
func buildFamily(ix *Index, members []int, evidenceFor map[[2]int]Evidence, claimedBy map[string]string, proposedTargets map[string]string, byName map[string]*GalleryEntry) (*Family, *Refusal) {
|
||||
names := make([]string, 0, len(members))
|
||||
for _, m := range members {
|
||||
names = append(names, ix.Entries[m].Name)
|
||||
}
|
||||
|
||||
parentIdx, err := selectParent(ix, members)
|
||||
if err != nil {
|
||||
return nil, &Refusal{Members: names, Reason: err.Error()}
|
||||
}
|
||||
parentEntry := ix.Entries[parentIdx]
|
||||
parentName := strings.ToLower(parentEntry.Name)
|
||||
|
||||
// A parent that is itself somebody's variant would create a chain, which
|
||||
// the gallery's own resolution refuses to install.
|
||||
if owner, claimed := claimedBy[parentName]; claimed {
|
||||
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("the natural parent %q is already a variant of %q; proposing it as a parent would nest variants", parentEntry.Name, owner)}
|
||||
}
|
||||
if owner, claimed := proposedTargets[parentName]; claimed {
|
||||
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("the natural parent %q is already proposed as a variant of %q; proposing it as a parent would nest variants", parentEntry.Name, owner)}
|
||||
}
|
||||
|
||||
// Adding a variants key to an anchor is inherited by every entry that
|
||||
// merges it, silently grouping models nobody proposed. Handling that means
|
||||
// editing each merging child too, which is a larger change than this job
|
||||
// should make unsupervised, so it refuses and hands the reviewer the list.
|
||||
if parentEntry.AnchorName != "" {
|
||||
children := ix.MergeChildren(parentEntry.AnchorName)
|
||||
if len(children) > 0 {
|
||||
childNames := make([]string, 0, len(children))
|
||||
for _, c := range children {
|
||||
childNames = append(childNames, c.Name)
|
||||
}
|
||||
return nil, &Refusal{
|
||||
Members: names,
|
||||
Reason: fmt.Sprintf("the parent %q defines YAML anchor &%s, and a variants key added there is inherited by the %d entries that merge it (%s). Grouping this family by hand also means adding an explicit `variants: []` to each of those entries",
|
||||
parentEntry.Name, parentEntry.AnchorName, len(children), strings.Join(childNames, ", ")),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
existing := map[string]struct{}{}
|
||||
for _, v := range parentEntry.Variants {
|
||||
existing[strings.ToLower(v.Model)] = struct{}{}
|
||||
}
|
||||
|
||||
family := &Family{Parent: parentEntry.Name}
|
||||
for _, m := range members {
|
||||
if m == parentIdx {
|
||||
continue
|
||||
}
|
||||
target := ix.Entries[m]
|
||||
lower := strings.ToLower(target.Name)
|
||||
if _, already := existing[lower]; already {
|
||||
continue
|
||||
}
|
||||
if target.HasVariants() {
|
||||
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("%q already offers variants of its own, so it cannot itself be a variant target", target.Name)}
|
||||
}
|
||||
if !target.Installable() {
|
||||
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("%q has no url, config_file, overrides or files, so it is not independently installable", target.Name)}
|
||||
}
|
||||
if owner, claimed := claimedBy[lower]; claimed && owner != parentName {
|
||||
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("%q is already a variant of %q; a target claimed by two parents is not something the gallery resolves predictably", target.Name, owner)}
|
||||
}
|
||||
if owner, claimed := proposedTargets[lower]; claimed && owner != parentEntry.Name {
|
||||
return nil, &Refusal{Members: names, Reason: fmt.Sprintf("%q is already proposed as a variant of %q in this same run", target.Name, owner)}
|
||||
}
|
||||
family.Proposals = append(family.Proposals, Proposal{
|
||||
Variant: target.Name,
|
||||
Evidence: lookupEvidence(evidenceFor, parentIdx, m),
|
||||
})
|
||||
}
|
||||
|
||||
if len(family.Proposals) == 0 {
|
||||
return nil, nil
|
||||
}
|
||||
sort.Slice(family.Proposals, func(i, j int) bool { return family.Proposals[i].Variant < family.Proposals[j].Variant })
|
||||
return family, nil
|
||||
}
|
||||
|
||||
func lookupEvidence(evidenceFor map[[2]int]Evidence, a, b int) Evidence {
|
||||
if a > b {
|
||||
a, b = b, a
|
||||
}
|
||||
if ev, ok := evidenceFor[[2]int{a, b}]; ok {
|
||||
return ev
|
||||
}
|
||||
// The two entries reached the same family through a third one. Say so
|
||||
// rather than inventing evidence that was never observed for this pair.
|
||||
return Evidence{Signals: []Signal{SignalName}}
|
||||
}
|
||||
|
||||
// selectParent picks the entry the others should hang off.
|
||||
//
|
||||
// The bare name wins when there is one: it is the name a user types and the one
|
||||
// documentation links to. Otherwise the smallest build wins, judged by the
|
||||
// quantization token in the entry's own weight filename, so the default install
|
||||
// is the one most hosts can actually run.
|
||||
func selectParent(ix *Index, members []int) (int, error) {
|
||||
// The family's own stem: the one the most members reduce to, shortest name
|
||||
// breaking a tie. An entry named exactly that is the bare entry.
|
||||
stemCount := map[string]int{}
|
||||
for _, m := range members {
|
||||
stemCount[NameStem(ix.Entries[m].Name)]++
|
||||
}
|
||||
// Only a stem two or more members reduce to is the family's own stem. A
|
||||
// stem reached by exactly one member is just that member's name, and
|
||||
// treating it as the family stem would crown whichever name happens to be
|
||||
// shortest rather than whichever build is the base one.
|
||||
familyStem := ""
|
||||
for stem, n := range stemCount {
|
||||
if n < 2 {
|
||||
continue
|
||||
}
|
||||
if familyStem == "" || n > stemCount[familyStem] ||
|
||||
(n == stemCount[familyStem] && len(stem) < len(familyStem)) ||
|
||||
(n == stemCount[familyStem] && len(stem) == len(familyStem) && stem < familyStem) {
|
||||
familyStem = stem
|
||||
}
|
||||
}
|
||||
|
||||
var bare []int
|
||||
for _, m := range members {
|
||||
e := ix.Entries[m]
|
||||
if HasConfigSuffix(e.Name) {
|
||||
continue
|
||||
}
|
||||
if strings.ToLower(e.Name) == familyStem {
|
||||
bare = append(bare, m)
|
||||
}
|
||||
}
|
||||
if len(bare) == 1 {
|
||||
return bare[0], nil
|
||||
}
|
||||
if len(bare) > 1 {
|
||||
names := make([]string, 0, len(bare))
|
||||
for _, m := range bare {
|
||||
names = append(names, ix.Entries[m].Name)
|
||||
}
|
||||
return 0, fmt.Errorf("more than one entry is named exactly %q (%s), so which one is the base build is a judgement this job will not make", familyStem, strings.Join(names, ", "))
|
||||
}
|
||||
|
||||
// No shared stem to be named after. An entry whose name every other member
|
||||
// extends is still recognisably the base one, and this is the only handle
|
||||
// left for families whose weights carry no readable quantization token at
|
||||
// all, such as the ONNX builds.
|
||||
if prefix, ok := uniquePrefixMember(ix, members); ok {
|
||||
return prefix, nil
|
||||
}
|
||||
|
||||
best := -1
|
||||
bestWidth := 1 << 20
|
||||
for _, m := range members {
|
||||
e := ix.Entries[m]
|
||||
width := unknownWidth
|
||||
if primary, ok := e.PrimaryWeightFile(); ok {
|
||||
width = BuildWidth(primary)
|
||||
}
|
||||
// Members are visited in gallery order, so a strict comparison leaves
|
||||
// the earliest entry holding a tie and the choice is deterministic.
|
||||
if width < bestWidth {
|
||||
best, bestWidth = m, width
|
||||
}
|
||||
}
|
||||
if best < 0 {
|
||||
return 0, fmt.Errorf("no member could be identified as the smallest build")
|
||||
}
|
||||
if bestWidth == unknownWidth {
|
||||
names := make([]string, 0, len(members))
|
||||
for _, m := range members {
|
||||
names = append(names, ix.Entries[m].Name)
|
||||
}
|
||||
return 0, fmt.Errorf("no member declares a weight file whose quantization can be read (%s), so the smallest build cannot be identified", strings.Join(names, ", "))
|
||||
}
|
||||
return best, nil
|
||||
}
|
||||
|
||||
// uniquePrefixMember reports the single member whose name every other member's
|
||||
// name starts with, if there is exactly one.
|
||||
func uniquePrefixMember(ix *Index, members []int) (int, bool) {
|
||||
found := -1
|
||||
for _, m := range members {
|
||||
name := strings.ToLower(ix.Entries[m].Name)
|
||||
isPrefix := true
|
||||
for _, other := range members {
|
||||
if other == m {
|
||||
continue
|
||||
}
|
||||
if !strings.HasPrefix(strings.ToLower(ix.Entries[other].Name), name) {
|
||||
isPrefix = false
|
||||
break
|
||||
}
|
||||
}
|
||||
if !isPrefix {
|
||||
continue
|
||||
}
|
||||
if found >= 0 {
|
||||
return 0, false
|
||||
}
|
||||
found = m
|
||||
}
|
||||
return found, found >= 0
|
||||
}
|
||||
|
||||
// quantDifference lists the quantization tokens that tell two names apart. It
|
||||
// is the compact form of the evidence: "these differ only by q4_k_m vs q8_0".
|
||||
func quantDifference(names ...string) []string {
|
||||
var out []string
|
||||
seen := map[string]struct{}{}
|
||||
for _, name := range names {
|
||||
// Filenames arrive here too, so the extension goes first and "/" counts
|
||||
// as a separator. "_" deliberately does not: it holds "q4_k_m" together.
|
||||
trimmed := weightExtension.ReplaceAllString(name, "")
|
||||
for _, seg := range strings.FieldsFunc(strings.ToLower(trimmed), func(r rune) bool { return r == '-' || r == '/' }) {
|
||||
if !IsQuantToken(seg) {
|
||||
continue
|
||||
}
|
||||
if _, ok := seen[seg]; ok {
|
||||
continue
|
||||
}
|
||||
seen[seg] = struct{}{}
|
||||
out = append(out, seg)
|
||||
}
|
||||
}
|
||||
sort.Strings(out)
|
||||
return out
|
||||
}
|
||||
429
.github/ci/variantproposals/propose_test.go
vendored
429
.github/ci/variantproposals/propose_test.go
vendored
@@ -1,429 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
)
|
||||
|
||||
// entryYAML writes one gallery entry with a single weight file, which is the
|
||||
// shape almost every real entry has. Specs that need something else write the
|
||||
// YAML out by hand.
|
||||
func entryYAML(name, repo, filename, sha string) string {
|
||||
return fmt.Sprintf(`- name: %s
|
||||
url: github:mudler/LocalAI/gallery/virtual.yaml@master
|
||||
overrides:
|
||||
parameters:
|
||||
model: %s
|
||||
files:
|
||||
- filename: %s
|
||||
uri: huggingface://%s/%s
|
||||
sha256: %s
|
||||
`, name, filename, filename, repo, filename, sha)
|
||||
}
|
||||
|
||||
func indexOf(entries ...string) *Index {
|
||||
ix, err := ParseIndex(strings.Join(entries, ""))
|
||||
ExpectWithOffset(1, err).ToNot(HaveOccurred())
|
||||
return ix
|
||||
}
|
||||
|
||||
// familyNames flattens a result into "parent <- variant, variant" strings, the
|
||||
// form the specs assert against.
|
||||
func familyNames(r *Result) []string {
|
||||
out := make([]string, 0, len(r.Families))
|
||||
for _, f := range r.Families {
|
||||
names := make([]string, 0, len(f.Proposals))
|
||||
for _, p := range f.Proposals {
|
||||
names = append(names, p.Variant)
|
||||
}
|
||||
out = append(out, f.Parent+" <- "+strings.Join(names, ", "))
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func refusalReasons(r *Result) string {
|
||||
var b strings.Builder
|
||||
for _, ref := range r.Refusals {
|
||||
b.WriteString(strings.Join(ref.Members, " + ") + ": " + ref.Reason + "\n")
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
func suppressionReasons(r *Result) string {
|
||||
var b strings.Builder
|
||||
for _, s := range r.Suppressed {
|
||||
b.WriteString(s.String() + "\n")
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
var _ = Describe("Propose", func() {
|
||||
Describe("the grouping signals", func() {
|
||||
It("groups entries whose names differ only by a quantization marker", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("foo-model", "acme/foo-GGUF", "foo-model-Q4_K_M.gguf", "aa"),
|
||||
entryYAML("foo-model-q8_0", "acme/foo-GGUF", "foo-model-Q8_0.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(ConsistOf("foo-model <- foo-model-q8_0"))
|
||||
Expect(r.Families[0].Proposals[0].Evidence.Signals).To(ContainElement(SignalName))
|
||||
Expect(r.Families[0].Proposals[0].Evidence.SharedStem).To(Equal("foo-model"))
|
||||
Expect(r.Families[0].Proposals[0].Evidence.QuantTokens).To(ContainElements("q4_k_m", "q8_0"))
|
||||
})
|
||||
|
||||
It("groups entries that use the colon config-suffix convention", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("bar-model", "acme/bar-GGUF", "bar-model-Q4_K_M.gguf", "aa"),
|
||||
entryYAML("bar-model:grammar-functioncall", "acme/bar-GGUF", "bar-model-Q4_K_M-grammar.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(ConsistOf("bar-model <- bar-model:grammar-functioncall"))
|
||||
Expect(r.Families[0].Proposals[0].Evidence.Signals).To(ContainElement(SignalConfigSuffix))
|
||||
})
|
||||
|
||||
It("groups entries whose own weight file is the same file at another quantization", func() {
|
||||
// The names share no stem, so only the filename signal can link
|
||||
// these two.
|
||||
ix := indexOf(
|
||||
entryYAML("omni-cpp", "Serveurperso/Omni-GGUF", "omnivoice-base-Q8_0.gguf", "aa"),
|
||||
entryYAML("omni-cpp-hq", "Serveurperso/Omni-GGUF", "omnivoice-base-BF16.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(ConsistOf("omni-cpp <- omni-cpp-hq"))
|
||||
ev := r.Families[0].Proposals[0].Evidence
|
||||
Expect(ev.Signals).To(ConsistOf(SignalWeightFile))
|
||||
Expect(ev.SharedFile).To(Equal("omnivoice-base"))
|
||||
Expect(ev.SharedRepo).To(Equal("serveurperso/omni-gguf"))
|
||||
})
|
||||
|
||||
It("does not let a shared auxiliary file link unrelated models", func() {
|
||||
// Both entries ship the same text encoder. That is a packaging
|
||||
// convention, not evidence of shared weights: this is how an
|
||||
// earlier sweep linked four wan-2.1 entries to each other.
|
||||
ix := indexOf(`- name: wan-2.1-t2v
|
||||
url: u
|
||||
files:
|
||||
- filename: wan-2.1-t2v-Q4_K_M.gguf
|
||||
uri: huggingface://acme/wan/wan-2.1-t2v-Q4_K_M.gguf
|
||||
sha256: aa
|
||||
- filename: umt5-xxl-encoder-Q8_0.gguf
|
||||
uri: huggingface://acme/wan/umt5-xxl-encoder-Q8_0.gguf
|
||||
sha256: cc
|
||||
`, `- name: z-image-turbo
|
||||
url: u
|
||||
files:
|
||||
- filename: z-image-turbo-Q4_K_M.gguf
|
||||
uri: huggingface://acme/wan/z-image-turbo-Q4_K_M.gguf
|
||||
sha256: bb
|
||||
- filename: umt5-xxl-encoder-Q8_0.gguf
|
||||
uri: huggingface://acme/wan/umt5-xxl-encoder-Q8_0.gguf
|
||||
sha256: cc
|
||||
`)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("does not treat a shared filename in two different repos as evidence", func() {
|
||||
// A finetune republished under the base model's filename is the
|
||||
// most common way this signal misfires.
|
||||
ix := indexOf(
|
||||
entryYAML("llama-3.2-3b-instruct", "hugging-quants/Llama-3.2-3B-Instruct-GGUF", "llama-3.2-3b-instruct-q4_k_m.gguf", "aa"),
|
||||
entryYAML("llama-3.2-3b-shiro-roleplay", "someone/Shiro-GGUF", "Llama-3.2-3B-Instruct.Q8_0.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
})
|
||||
})
|
||||
|
||||
Describe("what must never be proposed", func() {
|
||||
It("does not group different parameter sizes that share a prefix", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("qwen3-tts-cpp-0.6b-base", "Serveurperso/Qwen3-TTS-GGUF", "qwen3-tts-talker-Q4_K_M.gguf", "aa"),
|
||||
entryYAML("qwen3-tts-cpp-1.7b-base", "Serveurperso/Qwen3-TTS-GGUF", "qwen3-tts-talker-Q8_0.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
Expect(suppressionReasons(r)).To(ContainSubstring("different parameter sizes"))
|
||||
})
|
||||
|
||||
It("does not group the Gemma effective sizes", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("gemma-4-e2b-it", "google/gemma-GGUF", "gemma-4-it-Q4_K_M.gguf", "aa"),
|
||||
entryYAML("gemma-4-e4b-it", "google/gemma-GGUF", "gemma-4-it-Q8_0.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
Expect(suppressionReasons(r)).To(ContainSubstring("different parameter sizes"))
|
||||
})
|
||||
|
||||
It("does not group entries with a byte-identical install payload", func() {
|
||||
// whisper-1 exists so OpenAI-compatible clients can send that name.
|
||||
// Folding it under whisper-base would hide the name they send.
|
||||
payload := ` url: github:mudler/LocalAI/gallery/whisper-base.yaml@master
|
||||
overrides:
|
||||
parameters:
|
||||
model: ggml-whisper-base.bin
|
||||
files:
|
||||
- filename: ggml-whisper-base.bin
|
||||
uri: huggingface://ggerganov/whisper.cpp/ggml-base.bin
|
||||
sha256: aa
|
||||
`
|
||||
ix := indexOf("- name: whisper-base\n"+payload, "- name: whisper-1\n"+payload)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
Expect(r.AliasSkipped).To(HaveLen(1))
|
||||
Expect(r.AliasSkipped[0].Reason).To(ContainSubstring("aliases"))
|
||||
})
|
||||
|
||||
DescribeTable("declines the categories the ledger records",
|
||||
func(nameA, nameB string, ledgerYAML string) {
|
||||
ix := indexOf(
|
||||
entryYAML(nameA, "acme/repo", "shared-weights-Q4_K_M.gguf", "aa"),
|
||||
entryYAML(nameB, "acme/repo", "shared-weights-Q8_0.gguf", "bb"),
|
||||
)
|
||||
ledger, err := ParseLedger([]byte(ledgerYAML))
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
|
||||
// Without the ledger these would be proposed, which is what
|
||||
// makes the ledger load bearing rather than decorative.
|
||||
Expect(familyNames(Propose(ix, nil))).ToNot(BeEmpty())
|
||||
|
||||
r := Propose(ix, ledger)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
Expect(r.Suppressed).To(HaveLen(1))
|
||||
},
|
||||
Entry("a finetune", "base-model", "base-model-abliterated",
|
||||
"tokens:\n - {token: abliterated, reason: finetune}\n"),
|
||||
Entry("a distill", "base-model", "base-model-distilled",
|
||||
"tokens:\n - {token: distilled, reason: distilled}\n"),
|
||||
Entry("English-only versus multilingual ASR", "whisper-small", "whisper-small-en",
|
||||
"pairs:\n - {parent: whisper-small, variant: whisper-small-en, reason: English-only versus multilingual}\n"),
|
||||
Entry("two products sharing a prefix", "vibevoice-cpp", "vibevoice-cpp-asr",
|
||||
"pairs:\n - {parent: vibevoice-cpp, variant: vibevoice-cpp-asr, reason: different products}\n"),
|
||||
Entry("a per-language release", "kokoros-de", "kokoros-ja",
|
||||
"groups:\n - {members: [kokoros, kokoros-de, kokoros-ja], reason: different languages}\n"),
|
||||
)
|
||||
|
||||
It("reports the ledger's reason so its effect stays visible", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("base-model", "acme/repo", "shared-weights-Q4_K_M.gguf", "aa"),
|
||||
entryYAML("base-model-heretic", "acme/repo", "shared-weights-Q8_0.gguf", "bb"),
|
||||
)
|
||||
ledger, err := ParseLedger([]byte("tokens:\n - {token: heretic, reason: \"finetune, not a re-quantization\"}\n"))
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
r := Propose(ix, ledger)
|
||||
Expect(suppressionReasons(r)).To(ContainSubstring("finetune, not a re-quantization"))
|
||||
Expect(suppressionReasons(r)).To(ContainSubstring(`token "heretic"`))
|
||||
})
|
||||
})
|
||||
|
||||
Describe("parent selection", func() {
|
||||
It("picks the bare-named entry when one exists", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("base-model-q8_0", "acme/repo", "base-model-Q8_0.gguf", "aa"),
|
||||
entryYAML("base-model", "acme/repo", "base-model-Q4_K_M.gguf", "bb"),
|
||||
entryYAML("base-model-f16", "acme/repo", "base-model-f16.gguf", "cc"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(ConsistOf("base-model <- base-model-f16, base-model-q8_0"))
|
||||
})
|
||||
|
||||
It("picks the smallest build when no entry is bare-named", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("ced-base-f16", "acme/repo", "ced-base-f16.gguf", "aa"),
|
||||
entryYAML("ced-base-q8", "acme/repo", "ced-base-Q8_0.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(ConsistOf("ced-base-q8 <- ced-base-f16"))
|
||||
})
|
||||
|
||||
It("judges the smallest build by the quantization in the model filename, not the name", func() {
|
||||
// The names carry no marker at all; only the filenames say which
|
||||
// build is which.
|
||||
ix := indexOf(
|
||||
entryYAML("thing-hq", "acme/repo", "thing-weights-BF16.gguf", "aa"),
|
||||
entryYAML("thing-lite", "acme/repo", "thing-weights-Q4_K_M.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(ConsistOf("thing-lite <- thing-hq"))
|
||||
})
|
||||
})
|
||||
|
||||
Describe("the rules a proposal has to respect", func() {
|
||||
It("refuses to nest: a target that already offers variants of its own", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("nest-model", "acme/repo", "nest-model-Q4_K_M.gguf", "aa"),
|
||||
`- name: nest-model-q8_0
|
||||
url: u
|
||||
variants:
|
||||
- model: nest-model-q8_0-mtp
|
||||
overrides:
|
||||
parameters:
|
||||
model: nest-model-Q8_0.gguf
|
||||
files:
|
||||
- filename: nest-model-Q8_0.gguf
|
||||
uri: huggingface://acme/repo/nest-model-Q8_0.gguf
|
||||
sha256: bb
|
||||
`,
|
||||
entryYAML("nest-model-q8_0-mtp", "other/repo", "nest-model-mtp.gguf", "cc"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
Expect(refusalReasons(r)).To(ContainSubstring("already offers variants of its own"))
|
||||
})
|
||||
|
||||
It("refuses to nest: a parent that is already somebody else's variant", func() {
|
||||
ix := indexOf(
|
||||
`- name: outer
|
||||
url: u
|
||||
variants:
|
||||
- model: middle
|
||||
overrides:
|
||||
parameters:
|
||||
model: outer-Q4_K_M.gguf
|
||||
files:
|
||||
- filename: outer-Q4_K_M.gguf
|
||||
uri: huggingface://acme/repo/outer-Q4_K_M.gguf
|
||||
sha256: aa
|
||||
`,
|
||||
entryYAML("middle", "acme/other", "middle-Q4_K_M.gguf", "bb"),
|
||||
entryYAML("middle-q8_0", "acme/other", "middle-Q8_0.gguf", "cc"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
Expect(refusalReasons(r)).To(ContainSubstring("would nest variants"))
|
||||
})
|
||||
|
||||
It("refuses to let two parents claim one target", func() {
|
||||
ix := indexOf(
|
||||
`- name: claimant
|
||||
url: u
|
||||
variants:
|
||||
- model: contested-q8_0
|
||||
overrides:
|
||||
parameters:
|
||||
model: claimant-Q4_K_M.gguf
|
||||
files:
|
||||
- filename: claimant-Q4_K_M.gguf
|
||||
uri: huggingface://acme/repo/claimant-Q4_K_M.gguf
|
||||
sha256: aa
|
||||
`,
|
||||
entryYAML("contested", "acme/other", "contested-Q4_K_M.gguf", "bb"),
|
||||
entryYAML("contested-q8_0", "acme/other", "contested-Q8_0.gguf", "cc"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
Expect(refusalReasons(r)).To(ContainSubstring("already a variant of"))
|
||||
})
|
||||
|
||||
It("refuses a target that is not independently installable", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("stub-model", "acme/repo", "stub-model-Q4_K_M.gguf", "aa"),
|
||||
"- name: stub-model-q8_0\n description: a stanza nobody finished\n",
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
Expect(refusalReasons(r)).To(ContainSubstring("not independently installable"))
|
||||
})
|
||||
|
||||
It("refuses a family whose parent defines a merge anchor, naming the entries that would inherit", func() {
|
||||
ix := indexOf(
|
||||
`- &anchored
|
||||
name: anchored-model
|
||||
url: u
|
||||
overrides:
|
||||
parameters:
|
||||
model: anchored-Q4_K_M.gguf
|
||||
files:
|
||||
- filename: anchored-Q4_K_M.gguf
|
||||
uri: huggingface://acme/repo/anchored-Q4_K_M.gguf
|
||||
sha256: aa
|
||||
`,
|
||||
`- !!merge <<: *anchored
|
||||
name: anchored-child
|
||||
variants: []
|
||||
overrides:
|
||||
parameters:
|
||||
model: unrelated-child-Q4_K_M.gguf
|
||||
files:
|
||||
- filename: unrelated-child-Q4_K_M.gguf
|
||||
uri: huggingface://other/repo/unrelated-child-Q4_K_M.gguf
|
||||
sha256: cc
|
||||
`,
|
||||
entryYAML("anchored-model-q8_0", "acme/repo", "anchored-Q8_0.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
Expect(refusalReasons(r)).To(ContainSubstring("defines YAML anchor &anchored"))
|
||||
Expect(refusalReasons(r)).To(ContainSubstring("anchored-child"))
|
||||
Expect(refusalReasons(r)).To(ContainSubstring("variants: []"))
|
||||
})
|
||||
|
||||
It("refuses an entry whose name is not unique in the gallery", func() {
|
||||
ix := indexOf(
|
||||
entryYAML("twin", "acme/repo", "twin-Q4_K_M.gguf", "aa"),
|
||||
entryYAML("twin", "acme/repo", "twin-Q4_K_M.gguf", "aa"),
|
||||
entryYAML("twin-q8_0", "acme/repo", "twin-Q8_0.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(BeEmpty())
|
||||
Expect(refusalReasons(r)).To(ContainSubstring("appears more than once"))
|
||||
})
|
||||
|
||||
It("says nothing about a pair that is already grouped", func() {
|
||||
ix := indexOf(
|
||||
`- name: settled
|
||||
url: u
|
||||
variants:
|
||||
- model: settled-q8_0
|
||||
overrides:
|
||||
parameters:
|
||||
model: settled-Q4_K_M.gguf
|
||||
files:
|
||||
- filename: settled-Q4_K_M.gguf
|
||||
uri: huggingface://acme/repo/settled-Q4_K_M.gguf
|
||||
sha256: aa
|
||||
`,
|
||||
entryYAML("settled-q8_0", "acme/repo", "settled-Q8_0.gguf", "bb"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(r.HasProposals()).To(BeFalse())
|
||||
Expect(r.Refusals).To(BeEmpty())
|
||||
Expect(r.Suppressed).To(BeEmpty())
|
||||
})
|
||||
|
||||
It("adds only the missing members to a family that already exists", func() {
|
||||
ix := indexOf(
|
||||
`- name: partial
|
||||
url: u
|
||||
variants:
|
||||
- model: partial-q8_0
|
||||
overrides:
|
||||
parameters:
|
||||
model: partial-Q4_K_M.gguf
|
||||
files:
|
||||
- filename: partial-Q4_K_M.gguf
|
||||
uri: huggingface://acme/repo/partial-Q4_K_M.gguf
|
||||
sha256: aa
|
||||
`,
|
||||
entryYAML("partial-q8_0", "acme/repo", "partial-Q8_0.gguf", "bb"),
|
||||
entryYAML("partial-f16", "acme/repo", "partial-f16.gguf", "cc"),
|
||||
)
|
||||
r := Propose(ix, nil)
|
||||
Expect(familyNames(r)).To(ConsistOf("partial <- partial-f16"))
|
||||
})
|
||||
})
|
||||
|
||||
It("does not modify the index it was given", func() {
|
||||
text := entryYAML("foo-model", "acme/foo-GGUF", "foo-model-Q4_K_M.gguf", "aa") +
|
||||
entryYAML("foo-model-q8_0", "acme/foo-GGUF", "foo-model-Q8_0.gguf", "bb")
|
||||
ix, err := ParseIndex(text)
|
||||
Expect(err).ToNot(HaveOccurred())
|
||||
before := strings.Join(ix.Lines, "\n")
|
||||
Propose(ix, nil)
|
||||
Expect(strings.Join(ix.Lines, "\n")).To(Equal(before))
|
||||
})
|
||||
})
|
||||
151
.github/ci/variantproposals/quant.go
vendored
151
.github/ci/variantproposals/quant.go
vendored
@@ -1,151 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// Quantization and precision markers that distinguish one build of a set of
|
||||
// weights from another build of the same weights. Stripping them from a name
|
||||
// is what lets the proposer notice that two entries are the same model.
|
||||
//
|
||||
// qat and apex are in this list on a maintainer ruling: they are quantization
|
||||
// techniques applied to published weights, not separate weights. Names that use
|
||||
// "apex" to mean a finetune are handled by the rejection ledger instead, because
|
||||
// no amount of pattern matching can tell the two uses apart.
|
||||
const quantAlternation = `q[2-8](?:_[0-9a-z]+)*|pq[2-8](?:_[0-9a-z]+)*|iq[1-9][0-9a-z]*(?:_[0-9a-z]+)*|i1|` +
|
||||
`f16|f32|bf16|fp16|fp32|fp8|fp4|nvfp4|mxfp4(?:_moe)*|awq|gptq|qat|apex|gguf|ggml|[0-9]+bit|g[0-9]+`
|
||||
|
||||
// quantSegment matches a whole hyphen-delimited segment of an entry name.
|
||||
// Names separate their parts with "-" and keep quantization tokens internally
|
||||
// joined with "_", so a segment is the right unit here: "q4_k_m" arrives whole.
|
||||
var quantSegment = regexp.MustCompile(`^(?:` + quantAlternation + `)$`)
|
||||
|
||||
// quantFileSuffix matches a trailing quantization token in a weight filename.
|
||||
// Filenames mix "-", "_" and "." as separators, so unlike entry names they
|
||||
// cannot be split into segments up front without tearing "Q4_K_M" apart.
|
||||
var quantFileSuffix = regexp.MustCompile(`(?i)[-_.](?:` + quantAlternation + `)$`)
|
||||
|
||||
var weightExtension = regexp.MustCompile(`(?i)\.(gguf|ggml|safetensors|bin|pt|pth|onnx)$`)
|
||||
|
||||
// IsQuantToken reports whether a single name segment is a quantization or
|
||||
// precision marker rather than part of the model's identity.
|
||||
func IsQuantToken(segment string) bool {
|
||||
return quantSegment.MatchString(strings.ToLower(segment))
|
||||
}
|
||||
|
||||
// NameStem reduces an entry name to the identity it shares with its alternative
|
||||
// builds: the config suffix after ":" is dropped, then trailing quantization
|
||||
// segments are stripped.
|
||||
//
|
||||
// It implements the first two grouping signals together because they answer the
|
||||
// same question. "foo:q8_0" and "foo-q8_0" are both alternative builds of "foo",
|
||||
// and the caller that needs to report which convention was used can compare the
|
||||
// name against the stem itself.
|
||||
//
|
||||
// At least one segment always survives, so a name made entirely of quantization
|
||||
// tokens does not collapse to the empty stem and swallow every other such name.
|
||||
func NameStem(name string) string {
|
||||
base := strings.ToLower(strings.TrimSpace(name))
|
||||
if i := strings.Index(base, ":"); i >= 0 {
|
||||
base = base[:i]
|
||||
}
|
||||
segments := strings.Split(base, "-")
|
||||
for len(segments) > 1 && quantSegment.MatchString(segments[len(segments)-1]) {
|
||||
segments = segments[:len(segments)-1]
|
||||
}
|
||||
return strings.Join(segments, "-")
|
||||
}
|
||||
|
||||
// HasConfigSuffix reports whether a name uses the ":" convention for naming a
|
||||
// config variant of another entry.
|
||||
func HasConfigSuffix(name string) bool {
|
||||
return strings.Contains(name, ":")
|
||||
}
|
||||
|
||||
// FileStem reduces a weight filename to the identity shared by its other
|
||||
// quantizations: directories, extension and trailing quantization tokens go.
|
||||
//
|
||||
// This is the third grouping signal. It is the one that has misfired before, so
|
||||
// callers must filter auxiliary files out before handing a filename here: a
|
||||
// shared text encoder is not evidence of shared weights.
|
||||
func FileStem(filename string) string {
|
||||
base := filename
|
||||
if i := strings.LastIndex(base, "/"); i >= 0 {
|
||||
base = base[i+1:]
|
||||
}
|
||||
base = weightExtension.ReplaceAllString(base, "")
|
||||
for {
|
||||
stripped := quantFileSuffix.ReplaceAllString(base, "")
|
||||
if stripped == base {
|
||||
break
|
||||
}
|
||||
base = stripped
|
||||
}
|
||||
return strings.ToLower(base)
|
||||
}
|
||||
|
||||
// bitsPerWeight ranks quantization tokens so the smallest build of a family can
|
||||
// be identified when no bare-named entry exists to be the parent.
|
||||
//
|
||||
// The figures are nominal bits per weight, not measured file sizes. Ranking is
|
||||
// all that is asked of them, and a nominal figure is available from the name
|
||||
// alone without downloading anything.
|
||||
func bitsPerWeight(token string) (int, bool) {
|
||||
t := strings.ToLower(token)
|
||||
switch {
|
||||
case t == "i1":
|
||||
return 1, true
|
||||
case strings.HasPrefix(t, "nvfp4"), strings.HasPrefix(t, "mxfp4"), t == "fp4":
|
||||
return 4, true
|
||||
case t == "fp8":
|
||||
return 8, true
|
||||
case t == "f16", t == "bf16", t == "fp16":
|
||||
return 16, true
|
||||
case t == "f32", t == "fp32":
|
||||
return 32, true
|
||||
case t == "awq", t == "gptq":
|
||||
return 4, true
|
||||
}
|
||||
if m := regexp.MustCompile(`^p?q([1-9])`).FindStringSubmatch(t); m != nil {
|
||||
n, _ := strconv.Atoi(m[1])
|
||||
return n, true
|
||||
}
|
||||
if m := regexp.MustCompile(`^iq([1-9])`).FindStringSubmatch(t); m != nil {
|
||||
n, _ := strconv.Atoi(m[1])
|
||||
return n, true
|
||||
}
|
||||
if m := regexp.MustCompile(`^([0-9]+)bit$`).FindStringSubmatch(t); m != nil {
|
||||
n, _ := strconv.Atoi(m[1])
|
||||
return n, true
|
||||
}
|
||||
return 0, false
|
||||
}
|
||||
|
||||
// unknownWidth sorts after every recognised quantization so an entry whose
|
||||
// build cannot be read from its filename never wins the "smallest build" tie
|
||||
// break by accident.
|
||||
const unknownWidth = 1 << 10
|
||||
|
||||
// BuildWidth reports the nominal bits per weight of the build a filename holds.
|
||||
// An unreadable filename gets unknownWidth.
|
||||
func BuildWidth(filename string) int {
|
||||
base := filename
|
||||
if i := strings.LastIndex(base, "/"); i >= 0 {
|
||||
base = base[i+1:]
|
||||
}
|
||||
base = weightExtension.ReplaceAllString(base, "")
|
||||
best := unknownWidth
|
||||
for {
|
||||
m := quantFileSuffix.FindString(base)
|
||||
if m == "" {
|
||||
break
|
||||
}
|
||||
if bits, ok := bitsPerWeight(m[1:]); ok && bits < best {
|
||||
best = bits
|
||||
}
|
||||
base = base[:len(base)-len(m)]
|
||||
}
|
||||
return best
|
||||
}
|
||||
92
.github/ci/variantproposals/quant_test.go
vendored
92
.github/ci/variantproposals/quant_test.go
vendored
@@ -1,92 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
)
|
||||
|
||||
var _ = Describe("quantization markers", func() {
|
||||
DescribeTable("NameStem strips the markers that distinguish builds, not models",
|
||||
func(name, expected string) {
|
||||
Expect(NameStem(name)).To(Equal(expected))
|
||||
},
|
||||
Entry("plain q4", "foo-model-q4_k_m", "foo-model"),
|
||||
Entry("q8_0", "foo-model-q8_0", "foo-model"),
|
||||
Entry("q5_1", "foo-model-q5_1", "foo-model"),
|
||||
Entry("q2 with group size", "ternary-bonsai-8b-q2-g64", "ternary-bonsai-8b"),
|
||||
Entry("iq variant", "ideogram-4-iq4nl-ggml", "ideogram-4"),
|
||||
Entry("i1 imatrix", "orca-agent-v0.1-i1", "orca-agent-v0.1"),
|
||||
Entry("f16", "ced-base-f16", "ced-base"),
|
||||
Entry("bf16", "some-model-bf16", "some-model"),
|
||||
Entry("fp8", "some-model-fp8", "some-model"),
|
||||
Entry("nvfp4", "qwen3.6-27b-nvfp4", "qwen3.6-27b"),
|
||||
Entry("mxfp4_moe", "huihui-qwen3-vl-30b-a3b-instruct-abliterated-mxfp4_moe", "huihui-qwen3-vl-30b-a3b-instruct-abliterated"),
|
||||
Entry("pq2", "ternary-bonsai-8b-pq2", "ternary-bonsai-8b"),
|
||||
Entry("awq", "some-model-awq", "some-model"),
|
||||
Entry("gptq", "some-model-gptq", "some-model"),
|
||||
Entry("Nbit", "qwen3-8b-mlx-4bit", "qwen3-8b-mlx"),
|
||||
Entry("gguf", "some-model-gguf", "some-model"),
|
||||
Entry("ggml", "flux.1-dev-ggml", "flux.1-dev"),
|
||||
Entry("qat is a quantization technique", "gemma-3-27b-it-qat", "gemma-3-27b-it"),
|
||||
Entry("apex is a quantization technique", "qwen3.6-35b-a3b-apex", "qwen3.6-35b-a3b"),
|
||||
Entry("stacked markers", "gemma-4-e2b-it-qat-q4_0", "gemma-4-e2b-it"),
|
||||
Entry("the config suffix is dropped", "phi-2-chat:Q8_0", "phi-2-chat"),
|
||||
Entry("a non-quant config suffix is dropped too", "meta-llama-3.1-8b-instruct:grammar-functioncall", "meta-llama-3.1-8b-instruct"),
|
||||
)
|
||||
|
||||
DescribeTable("NameStem leaves alone what identifies a different model",
|
||||
func(name, expected string) {
|
||||
Expect(NameStem(name)).To(Equal(expected))
|
||||
},
|
||||
Entry("parameter size", "qwen3-tts-cpp-0.6b-base", "qwen3-tts-cpp-0.6b-base"),
|
||||
Entry("language suffix", "kokoros-de", "kokoros-de"),
|
||||
Entry("English-only ASR", "whisper-small-en", "whisper-small-en"),
|
||||
Entry("finetune", "qwen3-30b-a3b-abliterated", "qwen3-30b-a3b-abliterated"),
|
||||
Entry("product suffix", "vibevoice-cpp-asr", "vibevoice-cpp-asr"),
|
||||
)
|
||||
|
||||
It("never strips a name down to nothing", func() {
|
||||
Expect(NameStem("q4_k_m")).To(Equal("q4_k_m"))
|
||||
Expect(NameStem("f16-q8_0")).To(Equal("f16"))
|
||||
})
|
||||
|
||||
DescribeTable("FileStem reduces a weight filename to the weights it holds",
|
||||
func(filename, expected string) {
|
||||
Expect(FileStem(filename)).To(Equal(expected))
|
||||
},
|
||||
Entry("directory and extension go", "bonsai/models/Ternary-Bonsai-8B-gguf/Ternary-Bonsai-8B-Q2_0.gguf", "ternary-bonsai-8b"),
|
||||
Entry("underscored quant token stays whole", "Llama-3.2-1B-Instruct-Q4_K_M.gguf", "llama-3.2-1b-instruct"),
|
||||
Entry("dot separated quant token", "Llama-3.2-3B-Instruct.Q4_K_M.gguf", "llama-3.2-3b-instruct"),
|
||||
Entry("group size suffix", "Ternary-Bonsai-8B-Q2_0_g64.gguf", "ternary-bonsai-8b"),
|
||||
Entry("bf16", "omnivoice-cpp-hq/omnivoice-base-BF16.gguf", "omnivoice-base"),
|
||||
Entry("safetensors", "some/dir/Model-Name-fp8.safetensors", "model-name"),
|
||||
)
|
||||
|
||||
DescribeTable("BuildWidth reads the nominal width out of a filename",
|
||||
func(filename string, expected int) {
|
||||
Expect(BuildWidth(filename)).To(Equal(expected))
|
||||
},
|
||||
Entry("q4", "foo-Q4_K_M.gguf", 4),
|
||||
Entry("q8", "foo-Q8_0.gguf", 8),
|
||||
Entry("q2", "foo-Q2_0.gguf", 2),
|
||||
Entry("f16", "foo-f16.gguf", 16),
|
||||
Entry("bf16", "foo-BF16.gguf", 16),
|
||||
Entry("iq3", "foo-iq3_xxs.gguf", 3),
|
||||
Entry("nothing readable sorts last", "foo.gguf", unknownWidth),
|
||||
)
|
||||
|
||||
It("treats an auxiliary file as never being the model's own weights", func() {
|
||||
for _, f := range []string{
|
||||
"mmproj-model-f16.gguf",
|
||||
"dir/vae-BF16.gguf",
|
||||
"clip_l.safetensors",
|
||||
"umt5-xxl-encoder-Q8_0.gguf",
|
||||
"t5xxl_fp16.safetensors",
|
||||
"ae.safetensors",
|
||||
"omnivoice-tokenizer-Q8_0.gguf",
|
||||
} {
|
||||
Expect(IsAuxiliaryFile(f)).To(BeTrue(), "expected %q to be auxiliary", f)
|
||||
}
|
||||
Expect(IsAuxiliaryFile("gemma-3-27b-it-Q4_K_M.gguf")).To(BeFalse())
|
||||
})
|
||||
})
|
||||
@@ -1,13 +0,0 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
. "github.com/onsi/ginkgo/v2"
|
||||
. "github.com/onsi/gomega"
|
||||
)
|
||||
|
||||
func TestVariantProposals(t *testing.T) {
|
||||
RegisterFailHandler(Fail)
|
||||
RunSpecs(t, "gallery variant proposals")
|
||||
}
|
||||
10
.github/dependabot.yml
vendored
10
.github/dependabot.yml
vendored
@@ -45,16 +45,6 @@ updates:
|
||||
directory: "/backend/python/diffusers"
|
||||
schedule:
|
||||
interval: "weekly"
|
||||
# torch and transformers are deliberately pinned in this backend (see
|
||||
# backend/python/diffusers/requirements-*.txt and issue #9979), and the
|
||||
# l4t12 variant resolves them from the Jetson pip index
|
||||
# (https://pypi.jetson-ai-lab.io/jp6/cu129/). dependabot cannot authenticate
|
||||
# against that index and fails the whole weekly update with a
|
||||
# private_source_authentication_failure. Ignore the two pinned deps we don't
|
||||
# want bumped anyway so the job stays green.
|
||||
ignore:
|
||||
- dependency-name: "torch"
|
||||
- dependency-name: "transformers"
|
||||
- package-ecosystem: "pip"
|
||||
directory: "/backend/python/exllama"
|
||||
schedule:
|
||||
|
||||
39
.github/gh_curl.sh
vendored
39
.github/gh_curl.sh
vendored
@@ -1,39 +0,0 @@
|
||||
#!/bin/bash
|
||||
# Shared curl wrapper for the nightly dependency-bump scripts.
|
||||
#
|
||||
# The bump workflow fans out to ~25 parallel matrix jobs, each querying
|
||||
# api.github.com. Anonymous API calls are capped at 60/hour per source IP and
|
||||
# GitHub-hosted runners egress through shared NAT addresses, so a random handful
|
||||
# of jobs were getting rate-limited (HTTP 403 -> curl exit 22, empty response)
|
||||
# every single night. Authenticating with GITHUB_TOKEN lifts the ceiling to
|
||||
# 1000/hour; the retries absorb whatever transient blips remain.
|
||||
|
||||
# Wraps curl with GitHub auth (when a token is present) plus retry/timeout
|
||||
# hardening. Callers pass their own headers and the URL.
|
||||
gh_curl() {
|
||||
# The bump scripts run under `set -x`; without this the Authorization header
|
||||
# would be echoed into the job log on every call.
|
||||
local had_xtrace=0
|
||||
case "$-" in
|
||||
*x*) had_xtrace=1; set +x ;;
|
||||
esac
|
||||
|
||||
local args=(
|
||||
--silent --show-error --location --fail
|
||||
# --retry-all-errors so 403 rate-limit responses are retried too; plain
|
||||
# --retry only covers 408/429/5xx. curl honours Retry-After when sent.
|
||||
--retry 5 --retry-delay 3 --retry-all-errors
|
||||
--connect-timeout 15 --max-time 60
|
||||
)
|
||||
if [ -n "${GITHUB_TOKEN:-}" ]; then
|
||||
args+=(--header "Authorization: Bearer ${GITHUB_TOKEN}")
|
||||
fi
|
||||
|
||||
curl "${args[@]}" "$@"
|
||||
local rc=$?
|
||||
|
||||
if [ "$had_xtrace" -eq 1 ]; then
|
||||
set -x
|
||||
fi
|
||||
return $rc
|
||||
}
|
||||
77
.github/scripts/paged-canary-apply.sh
vendored
Executable file
77
.github/scripts/paged-canary-apply.sh
vendored
Executable file
@@ -0,0 +1,77 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# paged-canary-apply.sh - apply the vendored paged-attention patch series
|
||||
# (backend/cpp/llama-cpp-localai-paged/patches/paged/0001-0030) to a llama.cpp checkout, the
|
||||
# same way the build does, but tolerating the ONE known-benign pre-existing
|
||||
# quirk in the series. Used by the early-warning canary
|
||||
# (.github/workflows/llama-cpp-paged-canary.yml) so it only goes red on a REAL
|
||||
# upstream break, never on that quirk.
|
||||
#
|
||||
# Usage: paged-canary-apply.sh <llama.cpp-checkout-dir> <patches-dir>
|
||||
# <patches-dir> is normally backend/cpp/llama-cpp-localai-paged/patches (it holds the
|
||||
# top-level base series 0*.patch, currently empty, and the paged/ subseries).
|
||||
#
|
||||
# Exit 0 = the whole series applied -> patches still fit upstream.
|
||||
# Exit !=0 = a patch failed to apply = the red signal: an upstream change moved
|
||||
# the tree out from under the patches, so it is time to run a PIN_SYNC.
|
||||
#
|
||||
# Apply method MIRRORS backend/cpp/llama-cpp/Makefile's `llama.cpp` target:
|
||||
# plain `git apply --verbose`, which natively tolerates @@ line-number offsets
|
||||
# but NOT context-line changes. Matching the build's method is the point - the
|
||||
# canary's apply result is exactly what the real build's apply would do.
|
||||
#
|
||||
# The ONLY tolerance, and it is path-scoped (not a blanket `|| true`): patch
|
||||
# 0019 carries a stray *modify* hunk against the dev-only doc
|
||||
# SSM_DECODE_FIX_RESULTS.md, a file that exists only on the DGX dev tree and is
|
||||
# absent from any clean upstream checkout. `git apply` is atomic, so that single
|
||||
# missing-file hunk rejects the whole patch - and because 0021/0022/0026/0028
|
||||
# build on 0019's code, the rejection cascades to them too. This is a
|
||||
# PRE-EXISTING shipped-series defect, present identically on every pin, NOT an
|
||||
# upstream break (see backend/cpp/llama-cpp-localai-paged/README.md section 7,
|
||||
# "Pin + maintenance policy"). We exclude ONLY that dev-doc path and still
|
||||
# apply 0019's real code hunks atomically, so a genuine code-hunk break in 0019
|
||||
# still fails the canary. prepare.sh tolerates the same hunk via
|
||||
# `patch ... || true`; this mirrors that tolerance precisely.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
CHECKOUT="${1:?usage: paged-canary-apply.sh <llama.cpp-checkout> <patches-dir>}"
|
||||
PATCHES="${2:?usage: paged-canary-apply.sh <llama.cpp-checkout> <patches-dir>}"
|
||||
|
||||
# The lone tolerated dev-doc, and the only patch allowed to carry it.
|
||||
DEVDOC_GLOB='*SSM_DECODE_FIX_RESULTS.md'
|
||||
DEVDOC_PATCH='0019-qwen35-ssm-decode-fused-gather.patch'
|
||||
|
||||
# Resolve to absolute paths so the apply works after we cd into the checkout.
|
||||
PATCHES="$(cd "$PATCHES" && pwd)"
|
||||
cd "$CHECKOUT"
|
||||
|
||||
shopt -s nullglob
|
||||
|
||||
apply_one() {
|
||||
local p="$1"; shift
|
||||
echo "paged-canary: applying $(basename "$p")"
|
||||
if ! git apply --verbose "$@" "$p"; then
|
||||
echo "::error::paged patch no longer applies to the upstream llama.cpp tip: $(basename "$p")"
|
||||
echo "::error::upstream drifted past the vendored paged series - run a PIN_SYNC (see backend/cpp/llama-cpp-localai-paged/README.md section 7, Pin + maintenance policy), do NOT bump the pin blindly"
|
||||
exit 1
|
||||
fi
|
||||
}
|
||||
|
||||
# Base series first (parity with the build: patches/0*.patch before
|
||||
# patches/paged/0*.patch). Currently empty; nullglob makes this a no-op.
|
||||
for p in "$PATCHES"/0*.patch; do
|
||||
apply_one "$p"
|
||||
done
|
||||
|
||||
# Paged series, in order.
|
||||
for p in "$PATCHES"/paged/0*.patch; do
|
||||
if [ "$(basename "$p")" = "$DEVDOC_PATCH" ]; then
|
||||
# Apply 0019's real code hunks; exclude ONLY the benign dev-doc hunk.
|
||||
apply_one "$p" --exclude="$DEVDOC_GLOB"
|
||||
else
|
||||
apply_one "$p"
|
||||
fi
|
||||
done
|
||||
|
||||
echo "paged-canary: the full paged patch series applied cleanly to the upstream tip"
|
||||
181
.github/workflows/backend.yml
vendored
181
.github/workflows/backend.yml
vendored
@@ -32,30 +32,16 @@ jobs:
|
||||
if: github.repository == 'mudler/LocalAI'
|
||||
runs-on: ubuntu-latest
|
||||
outputs:
|
||||
matrix-singlearch: ${{ steps.set-matrix.outputs['matrix-singlearch'] }}
|
||||
matrix-multiarch: ${{ steps.set-matrix.outputs['matrix-multiarch'] }}
|
||||
matrix-darwin: ${{ steps.set-matrix.outputs['matrix-darwin'] }}
|
||||
merge-matrix-multiarch: ${{ steps.set-matrix.outputs['merge-matrix-multiarch'] }}
|
||||
merge-matrix-singlearch: ${{ steps.set-matrix.outputs['merge-matrix-singlearch'] }}
|
||||
has-backends-singlearch: ${{ steps.set-matrix.outputs['has-backends-singlearch'] }}
|
||||
has-backends-multiarch: ${{ steps.set-matrix.outputs['has-backends-multiarch'] }}
|
||||
has-backends-darwin: ${{ steps.set-matrix.outputs['has-backends-darwin'] }}
|
||||
has-merges-multiarch: ${{ steps.set-matrix.outputs['has-merges-multiarch'] }}
|
||||
# Single-arch backends are sharded across SINGLEARCH_SHARDS matrix jobs to
|
||||
# stay under GitHub's 256-jobs-per-matrix limit (see changed-backends.js).
|
||||
matrix-singlearch-1: ${{ steps.set-matrix.outputs['matrix-singlearch-1'] }}
|
||||
merge-matrix-singlearch-1: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-1'] }}
|
||||
has-backends-singlearch-1: ${{ steps.set-matrix.outputs['has-backends-singlearch-1'] }}
|
||||
has-merges-singlearch-1: ${{ steps.set-matrix.outputs['has-merges-singlearch-1'] }}
|
||||
matrix-singlearch-2: ${{ steps.set-matrix.outputs['matrix-singlearch-2'] }}
|
||||
merge-matrix-singlearch-2: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-2'] }}
|
||||
has-backends-singlearch-2: ${{ steps.set-matrix.outputs['has-backends-singlearch-2'] }}
|
||||
has-merges-singlearch-2: ${{ steps.set-matrix.outputs['has-merges-singlearch-2'] }}
|
||||
matrix-singlearch-3: ${{ steps.set-matrix.outputs['matrix-singlearch-3'] }}
|
||||
merge-matrix-singlearch-3: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-3'] }}
|
||||
has-backends-singlearch-3: ${{ steps.set-matrix.outputs['has-backends-singlearch-3'] }}
|
||||
has-merges-singlearch-3: ${{ steps.set-matrix.outputs['has-merges-singlearch-3'] }}
|
||||
matrix-singlearch-4: ${{ steps.set-matrix.outputs['matrix-singlearch-4'] }}
|
||||
merge-matrix-singlearch-4: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-4'] }}
|
||||
has-backends-singlearch-4: ${{ steps.set-matrix.outputs['has-backends-singlearch-4'] }}
|
||||
has-merges-singlearch-4: ${{ steps.set-matrix.outputs['has-merges-singlearch-4'] }}
|
||||
has-merges-singlearch: ${{ steps.set-matrix.outputs['has-merges-singlearch'] }}
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@v7
|
||||
@@ -123,9 +109,9 @@ jobs:
|
||||
# take their full ~6h cold without blocking manifest assembly for the
|
||||
# multi-arch backends whose per-arch digests would otherwise sit untagged
|
||||
# on quay long enough to be GC'd.
|
||||
backend-jobs-singlearch-1:
|
||||
backend-jobs-singlearch:
|
||||
needs: generate-matrix
|
||||
if: needs.generate-matrix.outputs['has-backends-singlearch-1'] == 'true'
|
||||
if: needs.generate-matrix.outputs['has-backends-singlearch'] == 'true'
|
||||
uses: ./.github/workflows/backend_build.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
@@ -152,100 +138,7 @@ jobs:
|
||||
strategy:
|
||||
fail-fast: false
|
||||
max-parallel: 8
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-1']) }}
|
||||
|
||||
backend-jobs-singlearch-2:
|
||||
needs: generate-matrix
|
||||
if: needs.generate-matrix.outputs['has-backends-singlearch-2'] == 'true'
|
||||
uses: ./.github/workflows/backend_build.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
build-type: ${{ matrix.build-type }}
|
||||
cuda-major-version: ${{ matrix.cuda-major-version }}
|
||||
cuda-minor-version: ${{ matrix.cuda-minor-version }}
|
||||
platforms: ${{ matrix.platforms }}
|
||||
platform-tag: ${{ matrix.platform-tag || '' }}
|
||||
runs-on: ${{ matrix.runs-on }}
|
||||
builder-base-image: ${{ matrix.builder-base-image || '' }}
|
||||
base-image: ${{ matrix.base-image }}
|
||||
backend: ${{ matrix.backend }}
|
||||
dockerfile: ${{ matrix.dockerfile }}
|
||||
skip-drivers: ${{ matrix.skip-drivers }}
|
||||
context: ${{ matrix.context }}
|
||||
ubuntu-version: ${{ matrix.ubuntu-version }}
|
||||
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
|
||||
secrets:
|
||||
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
|
||||
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
max-parallel: 8
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-2']) }}
|
||||
|
||||
backend-jobs-singlearch-3:
|
||||
needs: generate-matrix
|
||||
if: needs.generate-matrix.outputs['has-backends-singlearch-3'] == 'true'
|
||||
uses: ./.github/workflows/backend_build.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
build-type: ${{ matrix.build-type }}
|
||||
cuda-major-version: ${{ matrix.cuda-major-version }}
|
||||
cuda-minor-version: ${{ matrix.cuda-minor-version }}
|
||||
platforms: ${{ matrix.platforms }}
|
||||
platform-tag: ${{ matrix.platform-tag || '' }}
|
||||
runs-on: ${{ matrix.runs-on }}
|
||||
builder-base-image: ${{ matrix.builder-base-image || '' }}
|
||||
base-image: ${{ matrix.base-image }}
|
||||
backend: ${{ matrix.backend }}
|
||||
dockerfile: ${{ matrix.dockerfile }}
|
||||
skip-drivers: ${{ matrix.skip-drivers }}
|
||||
context: ${{ matrix.context }}
|
||||
ubuntu-version: ${{ matrix.ubuntu-version }}
|
||||
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
|
||||
secrets:
|
||||
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
|
||||
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
max-parallel: 8
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-3']) }}
|
||||
|
||||
backend-jobs-singlearch-4:
|
||||
needs: generate-matrix
|
||||
if: needs.generate-matrix.outputs['has-backends-singlearch-4'] == 'true'
|
||||
uses: ./.github/workflows/backend_build.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
build-type: ${{ matrix.build-type }}
|
||||
cuda-major-version: ${{ matrix.cuda-major-version }}
|
||||
cuda-minor-version: ${{ matrix.cuda-minor-version }}
|
||||
platforms: ${{ matrix.platforms }}
|
||||
platform-tag: ${{ matrix.platform-tag || '' }}
|
||||
runs-on: ${{ matrix.runs-on }}
|
||||
builder-base-image: ${{ matrix.builder-base-image || '' }}
|
||||
base-image: ${{ matrix.base-image }}
|
||||
backend: ${{ matrix.backend }}
|
||||
dockerfile: ${{ matrix.dockerfile }}
|
||||
skip-drivers: ${{ matrix.skip-drivers }}
|
||||
context: ${{ matrix.context }}
|
||||
ubuntu-version: ${{ matrix.ubuntu-version }}
|
||||
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
|
||||
secrets:
|
||||
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
|
||||
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
max-parallel: 8
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-4']) }}
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch']) }}
|
||||
|
||||
# Apply tags to per-arch digests via `imagetools create`. Split into two
|
||||
# jobs that mirror the build split so each merge waits ONLY on its
|
||||
@@ -281,12 +174,10 @@ jobs:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-multiarch']) }}
|
||||
|
||||
# One merge shard per build shard: backend-merge-jobs-singlearch-<n> needs only
|
||||
# backend-jobs-singlearch-<n>, preserving the "merge waits only on its own
|
||||
# build" property while staying under the 256-jobs-per-matrix limit.
|
||||
backend-merge-jobs-singlearch-1:
|
||||
needs: [generate-matrix, backend-jobs-singlearch-1]
|
||||
if: ${{ !cancelled() && needs.generate-matrix.outputs['has-merges-singlearch-1'] == 'true' }}
|
||||
backend-merge-jobs-singlearch:
|
||||
needs: [generate-matrix, backend-jobs-singlearch]
|
||||
# See note on backend-merge-jobs-multiarch above for !cancelled().
|
||||
if: ${{ !cancelled() && needs.generate-matrix.outputs['has-merges-singlearch'] == 'true' }}
|
||||
uses: ./.github/workflows/backend_merge.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
@@ -298,55 +189,7 @@ jobs:
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-1']) }}
|
||||
|
||||
backend-merge-jobs-singlearch-2:
|
||||
needs: [generate-matrix, backend-jobs-singlearch-2]
|
||||
if: ${{ !cancelled() && needs.generate-matrix.outputs['has-merges-singlearch-2'] == 'true' }}
|
||||
uses: ./.github/workflows/backend_merge.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
secrets:
|
||||
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
|
||||
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-2']) }}
|
||||
|
||||
backend-merge-jobs-singlearch-3:
|
||||
needs: [generate-matrix, backend-jobs-singlearch-3]
|
||||
if: ${{ !cancelled() && needs.generate-matrix.outputs['has-merges-singlearch-3'] == 'true' }}
|
||||
uses: ./.github/workflows/backend_merge.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
secrets:
|
||||
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
|
||||
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-3']) }}
|
||||
|
||||
backend-merge-jobs-singlearch-4:
|
||||
needs: [generate-matrix, backend-jobs-singlearch-4]
|
||||
if: ${{ !cancelled() && needs.generate-matrix.outputs['has-merges-singlearch-4'] == 'true' }}
|
||||
uses: ./.github/workflows/backend_merge.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
secrets:
|
||||
dockerUsername: ${{ secrets.DOCKERHUB_USERNAME }}
|
||||
dockerPassword: ${{ secrets.DOCKERHUB_PASSWORD }}
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-4']) }}
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch']) }}
|
||||
|
||||
backend-jobs-darwin:
|
||||
needs: generate-matrix
|
||||
|
||||
25
.github/workflows/backend_build_darwin.yml
vendored
25
.github/workflows/backend_build_darwin.yml
vendored
@@ -82,7 +82,7 @@ jobs:
|
||||
# as the Linux registry cache.
|
||||
- name: Restore Homebrew cache
|
||||
id: brew-cache
|
||||
uses: actions/cache/restore@v6
|
||||
uses: actions/cache/restore@v4
|
||||
with:
|
||||
path: |
|
||||
~/Library/Caches/Homebrew/downloads
|
||||
@@ -142,7 +142,7 @@ jobs:
|
||||
|
||||
- name: Save Homebrew cache
|
||||
if: github.event_name != 'pull_request' && steps.brew-cache.outputs.cache-hit != 'true'
|
||||
uses: actions/cache/save@v6
|
||||
uses: actions/cache/save@v4
|
||||
with:
|
||||
path: |
|
||||
~/Library/Caches/Homebrew/downloads
|
||||
@@ -169,16 +169,16 @@ jobs:
|
||||
# invalidates cleanly; restore-keys fall back to the latest entry for the
|
||||
# same pin so unchanged TUs stay warm even when the cache is fresh.
|
||||
- name: Compute llama.cpp version
|
||||
if: inputs.backend == 'llama-cpp'
|
||||
if: inputs.backend == 'llama-cpp' || inputs.backend == 'llama-cpp-localai-paged'
|
||||
id: llama-version
|
||||
run: |
|
||||
version=$(grep '^LLAMA_VERSION' backend/cpp/llama-cpp/Makefile | head -1 | cut -d= -f2 | cut -d'?' -f1 | tr -d ' ')
|
||||
echo "version=${version}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Restore ccache
|
||||
if: inputs.backend == 'llama-cpp'
|
||||
if: inputs.backend == 'llama-cpp' || inputs.backend == 'llama-cpp-localai-paged'
|
||||
id: ccache-cache
|
||||
uses: actions/cache/restore@v6
|
||||
uses: actions/cache/restore@v4
|
||||
with:
|
||||
path: ~/Library/Caches/ccache
|
||||
key: ccache-llama-${{ runner.arch }}-${{ steps.llama-version.outputs.version }}-${{ github.run_id }}
|
||||
@@ -186,7 +186,7 @@ jobs:
|
||||
ccache-llama-${{ runner.arch }}-${{ steps.llama-version.outputs.version }}-
|
||||
|
||||
- name: Configure ccache
|
||||
if: inputs.backend == 'llama-cpp'
|
||||
if: inputs.backend == 'llama-cpp' || inputs.backend == 'llama-cpp-localai-paged'
|
||||
run: |
|
||||
mkdir -p "$HOME/Library/Caches/ccache"
|
||||
ccache -M 2G
|
||||
@@ -211,7 +211,7 @@ jobs:
|
||||
- name: Restore Python wheel cache
|
||||
if: inputs.lang == 'python'
|
||||
id: pyenv-cache
|
||||
uses: actions/cache/restore@v6
|
||||
uses: actions/cache/restore@v4
|
||||
with:
|
||||
path: |
|
||||
~/Library/Caches/pip
|
||||
@@ -251,19 +251,24 @@ jobs:
|
||||
BACKEND=${{ inputs.backend }} BUILD_TYPE=${{ inputs.build-type }} USE_PIP=${{ inputs.use-pip }} make build-darwin-${{ inputs.lang }}-backend
|
||||
|
||||
- name: ccache stats
|
||||
if: inputs.backend == 'llama-cpp'
|
||||
if: inputs.backend == 'llama-cpp' || inputs.backend == 'llama-cpp-localai-paged'
|
||||
run: ccache -s
|
||||
|
||||
# Only stock llama-cpp persists the ccache: both backends share the same
|
||||
# ccache-llama-<arch>-<version>-<run_id> key, so the paged job restores from
|
||||
# the shared prefix (warm) but must NOT also save under the identical key in
|
||||
# the same run (it would collide). The shared upstream TUs stay warm via the
|
||||
# stock save; the paged-only patched TUs are a small recompile.
|
||||
- name: Save ccache
|
||||
if: inputs.backend == 'llama-cpp' && github.event_name != 'pull_request'
|
||||
uses: actions/cache/save@v6
|
||||
uses: actions/cache/save@v4
|
||||
with:
|
||||
path: ~/Library/Caches/ccache
|
||||
key: ccache-llama-${{ runner.arch }}-${{ steps.llama-version.outputs.version }}-${{ github.run_id }}
|
||||
|
||||
- name: Save Python wheel cache
|
||||
if: inputs.lang == 'python' && github.event_name != 'pull_request' && steps.pyenv-cache.outputs.cache-hit != 'true'
|
||||
uses: actions/cache/save@v6
|
||||
uses: actions/cache/save@v4
|
||||
with:
|
||||
path: |
|
||||
~/Library/Caches/pip
|
||||
|
||||
165
.github/workflows/backend_pr.yml
vendored
165
.github/workflows/backend_pr.yml
vendored
@@ -11,30 +11,16 @@ jobs:
|
||||
generate-matrix:
|
||||
runs-on: ubuntu-latest
|
||||
outputs:
|
||||
matrix-singlearch: ${{ steps.set-matrix.outputs['matrix-singlearch'] }}
|
||||
matrix-multiarch: ${{ steps.set-matrix.outputs['matrix-multiarch'] }}
|
||||
matrix-darwin: ${{ steps.set-matrix.outputs['matrix-darwin'] }}
|
||||
merge-matrix-multiarch: ${{ steps.set-matrix.outputs['merge-matrix-multiarch'] }}
|
||||
merge-matrix-singlearch: ${{ steps.set-matrix.outputs['merge-matrix-singlearch'] }}
|
||||
has-backends-singlearch: ${{ steps.set-matrix.outputs['has-backends-singlearch'] }}
|
||||
has-backends-multiarch: ${{ steps.set-matrix.outputs['has-backends-multiarch'] }}
|
||||
has-backends-darwin: ${{ steps.set-matrix.outputs['has-backends-darwin'] }}
|
||||
has-merges-multiarch: ${{ steps.set-matrix.outputs['has-merges-multiarch'] }}
|
||||
# Single-arch backends are sharded across SINGLEARCH_SHARDS matrix jobs to
|
||||
# stay under GitHub's 256-jobs-per-matrix limit (see changed-backends.js).
|
||||
matrix-singlearch-1: ${{ steps.set-matrix.outputs['matrix-singlearch-1'] }}
|
||||
merge-matrix-singlearch-1: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-1'] }}
|
||||
has-backends-singlearch-1: ${{ steps.set-matrix.outputs['has-backends-singlearch-1'] }}
|
||||
has-merges-singlearch-1: ${{ steps.set-matrix.outputs['has-merges-singlearch-1'] }}
|
||||
matrix-singlearch-2: ${{ steps.set-matrix.outputs['matrix-singlearch-2'] }}
|
||||
merge-matrix-singlearch-2: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-2'] }}
|
||||
has-backends-singlearch-2: ${{ steps.set-matrix.outputs['has-backends-singlearch-2'] }}
|
||||
has-merges-singlearch-2: ${{ steps.set-matrix.outputs['has-merges-singlearch-2'] }}
|
||||
matrix-singlearch-3: ${{ steps.set-matrix.outputs['matrix-singlearch-3'] }}
|
||||
merge-matrix-singlearch-3: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-3'] }}
|
||||
has-backends-singlearch-3: ${{ steps.set-matrix.outputs['has-backends-singlearch-3'] }}
|
||||
has-merges-singlearch-3: ${{ steps.set-matrix.outputs['has-merges-singlearch-3'] }}
|
||||
matrix-singlearch-4: ${{ steps.set-matrix.outputs['matrix-singlearch-4'] }}
|
||||
merge-matrix-singlearch-4: ${{ steps.set-matrix.outputs['merge-matrix-singlearch-4'] }}
|
||||
has-backends-singlearch-4: ${{ steps.set-matrix.outputs['has-backends-singlearch-4'] }}
|
||||
has-merges-singlearch-4: ${{ steps.set-matrix.outputs['has-merges-singlearch-4'] }}
|
||||
has-merges-singlearch: ${{ steps.set-matrix.outputs['has-merges-singlearch'] }}
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@v7
|
||||
@@ -85,10 +71,10 @@ jobs:
|
||||
fail-fast: true
|
||||
max-parallel: 8
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-multiarch']) }}
|
||||
backend-jobs-singlearch-1:
|
||||
backend-jobs-singlearch:
|
||||
needs: generate-matrix
|
||||
if: needs.generate-matrix.outputs['has-backends-singlearch-1'] == 'true'
|
||||
uses: ./.github/workflows/backend_build.yml
|
||||
if: needs.generate-matrix.outputs['has-backends-singlearch'] == 'true'
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
@@ -112,94 +98,7 @@ jobs:
|
||||
strategy:
|
||||
fail-fast: true
|
||||
max-parallel: 8
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-1']) }}
|
||||
|
||||
backend-jobs-singlearch-2:
|
||||
needs: generate-matrix
|
||||
if: needs.generate-matrix.outputs['has-backends-singlearch-2'] == 'true'
|
||||
uses: ./.github/workflows/backend_build.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
build-type: ${{ matrix.build-type }}
|
||||
cuda-major-version: ${{ matrix.cuda-major-version }}
|
||||
cuda-minor-version: ${{ matrix.cuda-minor-version }}
|
||||
platforms: ${{ matrix.platforms }}
|
||||
platform-tag: ${{ matrix.platform-tag || '' }}
|
||||
runs-on: ${{ matrix.runs-on }}
|
||||
builder-base-image: ${{ matrix.builder-base-image || '' }}
|
||||
base-image: ${{ matrix.base-image }}
|
||||
backend: ${{ matrix.backend }}
|
||||
dockerfile: ${{ matrix.dockerfile }}
|
||||
skip-drivers: ${{ matrix.skip-drivers }}
|
||||
context: ${{ matrix.context }}
|
||||
ubuntu-version: ${{ matrix.ubuntu-version }}
|
||||
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
|
||||
secrets:
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: true
|
||||
max-parallel: 8
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-2']) }}
|
||||
|
||||
backend-jobs-singlearch-3:
|
||||
needs: generate-matrix
|
||||
if: needs.generate-matrix.outputs['has-backends-singlearch-3'] == 'true'
|
||||
uses: ./.github/workflows/backend_build.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
build-type: ${{ matrix.build-type }}
|
||||
cuda-major-version: ${{ matrix.cuda-major-version }}
|
||||
cuda-minor-version: ${{ matrix.cuda-minor-version }}
|
||||
platforms: ${{ matrix.platforms }}
|
||||
platform-tag: ${{ matrix.platform-tag || '' }}
|
||||
runs-on: ${{ matrix.runs-on }}
|
||||
builder-base-image: ${{ matrix.builder-base-image || '' }}
|
||||
base-image: ${{ matrix.base-image }}
|
||||
backend: ${{ matrix.backend }}
|
||||
dockerfile: ${{ matrix.dockerfile }}
|
||||
skip-drivers: ${{ matrix.skip-drivers }}
|
||||
context: ${{ matrix.context }}
|
||||
ubuntu-version: ${{ matrix.ubuntu-version }}
|
||||
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
|
||||
secrets:
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: true
|
||||
max-parallel: 8
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-3']) }}
|
||||
|
||||
backend-jobs-singlearch-4:
|
||||
needs: generate-matrix
|
||||
if: needs.generate-matrix.outputs['has-backends-singlearch-4'] == 'true'
|
||||
uses: ./.github/workflows/backend_build.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
build-type: ${{ matrix.build-type }}
|
||||
cuda-major-version: ${{ matrix.cuda-major-version }}
|
||||
cuda-minor-version: ${{ matrix.cuda-minor-version }}
|
||||
platforms: ${{ matrix.platforms }}
|
||||
platform-tag: ${{ matrix.platform-tag || '' }}
|
||||
runs-on: ${{ matrix.runs-on }}
|
||||
builder-base-image: ${{ matrix.builder-base-image || '' }}
|
||||
base-image: ${{ matrix.base-image }}
|
||||
backend: ${{ matrix.backend }}
|
||||
dockerfile: ${{ matrix.dockerfile }}
|
||||
skip-drivers: ${{ matrix.skip-drivers }}
|
||||
context: ${{ matrix.context }}
|
||||
ubuntu-version: ${{ matrix.ubuntu-version }}
|
||||
amdgpu-targets: ${{ matrix.amdgpu-targets || 'gfx908,gfx90a,gfx942,gfx950,gfx1030,gfx1100,gfx1101,gfx1102,gfx1151,gfx1200,gfx1201' }}
|
||||
secrets:
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: true
|
||||
max-parallel: 8
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch-4']) }}
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['matrix-singlearch']) }}
|
||||
backend-merge-jobs-multiarch:
|
||||
needs: [generate-matrix, backend-jobs-multiarch]
|
||||
# backend_merge.yml's push-side steps are all gated on
|
||||
@@ -219,9 +118,9 @@ jobs:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-multiarch']) }}
|
||||
|
||||
backend-merge-jobs-singlearch-1:
|
||||
needs: [generate-matrix, backend-jobs-singlearch-1]
|
||||
if: ${{ !cancelled() && github.event_name != 'pull_request' && needs.generate-matrix.outputs['has-merges-singlearch-1'] == 'true' }}
|
||||
backend-merge-jobs-singlearch:
|
||||
needs: [generate-matrix, backend-jobs-singlearch]
|
||||
if: ${{ !cancelled() && github.event_name != 'pull_request' && needs.generate-matrix.outputs['has-merges-singlearch'] == 'true' }}
|
||||
uses: ./.github/workflows/backend_merge.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
@@ -231,49 +130,7 @@ jobs:
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-1']) }}
|
||||
|
||||
backend-merge-jobs-singlearch-2:
|
||||
needs: [generate-matrix, backend-jobs-singlearch-2]
|
||||
if: ${{ !cancelled() && github.event_name != 'pull_request' && needs.generate-matrix.outputs['has-merges-singlearch-2'] == 'true' }}
|
||||
uses: ./.github/workflows/backend_merge.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
secrets:
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-2']) }}
|
||||
|
||||
backend-merge-jobs-singlearch-3:
|
||||
needs: [generate-matrix, backend-jobs-singlearch-3]
|
||||
if: ${{ !cancelled() && github.event_name != 'pull_request' && needs.generate-matrix.outputs['has-merges-singlearch-3'] == 'true' }}
|
||||
uses: ./.github/workflows/backend_merge.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
secrets:
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-3']) }}
|
||||
|
||||
backend-merge-jobs-singlearch-4:
|
||||
needs: [generate-matrix, backend-jobs-singlearch-4]
|
||||
if: ${{ !cancelled() && github.event_name != 'pull_request' && needs.generate-matrix.outputs['has-merges-singlearch-4'] == 'true' }}
|
||||
uses: ./.github/workflows/backend_merge.yml
|
||||
with:
|
||||
tag-latest: ${{ matrix.tag-latest }}
|
||||
tag-suffix: ${{ matrix.tag-suffix }}
|
||||
secrets:
|
||||
quayUsername: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
quayPassword: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch-4']) }}
|
||||
matrix: ${{ fromJson(needs.generate-matrix.outputs['merge-matrix-singlearch']) }}
|
||||
backend-jobs-darwin:
|
||||
needs: generate-matrix
|
||||
uses: ./.github/workflows/backend_build_darwin.yml
|
||||
|
||||
8
.github/workflows/build-test.yaml
vendored
8
.github/workflows/build-test.yaml
vendored
@@ -6,14 +6,6 @@ on:
|
||||
- master
|
||||
pull_request:
|
||||
|
||||
# Supersede an in-flight run when a PR gets a new push. Keyed on the PR number
|
||||
# so every push to the same PR shares a group; on a master push the key falls
|
||||
# back to github.sha (unique per commit) and cancel-in-progress is false, so
|
||||
# master runs never cancel each other -- each commit is built on its own.
|
||||
concurrency:
|
||||
group: ci-build-test-${{ github.event.pull_request.number || github.sha }}-${{ github.repository }}
|
||||
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
|
||||
|
||||
jobs:
|
||||
build-test:
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
55
.github/workflows/bump_deps.yaml
vendored
55
.github/workflows/bump_deps.yaml
vendored
@@ -9,6 +9,23 @@ jobs:
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
# NOTE: there is intentionally NO entry for the llama-cpp-localai-paged
|
||||
# backend. It carries a vendored paged-attention patch series
|
||||
# (backend/cpp/llama-cpp-localai-paged/patches/paged/) hand-verified bit-exact against
|
||||
# ONE specific llama.cpp tip; a naive nightly bump would move the tip out
|
||||
# from under the patches and break `git apply` at build time. Its pin is
|
||||
# therefore decoupled (its own LLAMA_VERSION in
|
||||
# backend/cpp/llama-cpp-localai-paged/Makefile) and advanced ONLY by the
|
||||
# manual PIN_SYNC process. Do not add it here. (turboquant CAN be
|
||||
# auto-bumped below because its fork branch carries the patches.)
|
||||
#
|
||||
# Excluding it from the auto-bumper removed the early warning of upstream
|
||||
# drift; that signal is restored separately by the dedicated canary
|
||||
# .github/workflows/llama-cpp-paged-canary.yml, which weekly applies +
|
||||
# compiles the paged series against the latest llama.cpp tip and goes red
|
||||
# when upstream breaks it (prompting a PIN_SYNC). The canary is
|
||||
# signal-only - it never opens a bump PR and never moves the pin - so
|
||||
# this dep-bump workflow and its PRs stay green regardless.
|
||||
include:
|
||||
- repository: "ggml-org/llama.cpp"
|
||||
variable: "LLAMA_VERSION"
|
||||
@@ -22,18 +39,10 @@ jobs:
|
||||
variable: "TURBOQUANT_VERSION"
|
||||
branch: "feature/turboquant-kv-cache"
|
||||
file: "backend/cpp/turboquant/Makefile"
|
||||
- repository: "PrismML-Eng/llama.cpp"
|
||||
variable: "BONSAI_VERSION"
|
||||
branch: "prism"
|
||||
file: "backend/cpp/bonsai/Makefile"
|
||||
- repository: "antirez/ds4"
|
||||
variable: "DS4_VERSION"
|
||||
branch: "main"
|
||||
file: "backend/cpp/ds4/Makefile"
|
||||
- repository: "meituan-longcat/LongCat-Video"
|
||||
variable: "LONGCAT_VIDEO_VERSION"
|
||||
branch: "main"
|
||||
file: "backend/python/longcat-video/Makefile"
|
||||
- repository: "localai-org/privacy-filter.cpp"
|
||||
variable: "PRIVACY_FILTER_VERSION"
|
||||
branch: "master"
|
||||
@@ -50,19 +59,11 @@ jobs:
|
||||
variable: "PARAKEET_VERSION"
|
||||
branch: "master"
|
||||
file: "backend/go/parakeet-cpp/Makefile"
|
||||
- repository: "mudler/vllm.cpp"
|
||||
variable: "VLLM_CPP_VERSION"
|
||||
branch: "main"
|
||||
file: "backend/go/vllm-cpp/Makefile"
|
||||
- repository: "localai-org/moss-transcribe.cpp"
|
||||
variable: "MOSS_VERSION"
|
||||
branch: "master"
|
||||
file: "backend/go/moss-transcribe-cpp/Makefile"
|
||||
- repository: "localai-org/ced.cpp"
|
||||
- repository: "mudler/ced.cpp"
|
||||
variable: "CED_VERSION"
|
||||
branch: "main"
|
||||
branch: "master"
|
||||
file: "backend/go/ced/Makefile"
|
||||
- repository: "localai-org/voice-detect.cpp"
|
||||
- repository: "mudler/voice-detect.cpp"
|
||||
variable: "VOICEDETECT_VERSION"
|
||||
branch: "master"
|
||||
file: "backend/go/voice-detect/Makefile"
|
||||
@@ -94,7 +95,7 @@ jobs:
|
||||
variable: "SAM3_VERSION"
|
||||
branch: "main"
|
||||
file: "backend/go/sam3-cpp/Makefile"
|
||||
- repository: "localai-org/rf-detr.cpp"
|
||||
- repository: "mudler/rf-detr.cpp"
|
||||
variable: "RFDETR_VERSION"
|
||||
branch: "main"
|
||||
file: "backend/go/rfdetr-cpp/Makefile"
|
||||
@@ -114,21 +115,11 @@ jobs:
|
||||
variable: "VIBEVOICE_CPP_VERSION"
|
||||
branch: "master"
|
||||
file: "backend/go/vibevoice-cpp/Makefile"
|
||||
- repository: "mudler/magpie-tts.cpp"
|
||||
variable: "MAGPIETTS_CPP_VERSION"
|
||||
branch: "main"
|
||||
file: "backend/go/magpie-tts-cpp/Makefile"
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v7
|
||||
- name: Bump dependencies 🔧
|
||||
id: bump
|
||||
env:
|
||||
# This job fans out to ~25 parallel matrix entries, all querying
|
||||
# api.github.com from runner IPs that share the 60/hour anonymous
|
||||
# rate limit. Authenticating raises it to 1000/hour, which is what
|
||||
# kept a random handful of these red every night.
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
bash .github/bump_deps.sh ${{ matrix.repository }} ${{ matrix.branch }} ${{ matrix.variable }} ${{ matrix.file }}
|
||||
{
|
||||
@@ -165,8 +156,6 @@ jobs:
|
||||
- uses: actions/checkout@v7
|
||||
- name: Bump vLLM cu130 wheel pin 🔧
|
||||
id: bump
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
bash .github/bump_vllm_wheel.sh vllm-project/vllm backend/python/vllm/requirements-cublas13-after.txt VLLM_VERSION
|
||||
{
|
||||
@@ -203,8 +192,6 @@ jobs:
|
||||
- uses: actions/checkout@v7
|
||||
- name: Bump vllm-metal pin 🔧
|
||||
id: bump
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
bash .github/bump_vllm_metal.sh vllm-project/vllm-metal backend/python/vllm/install.sh VLLM_METAL_VERSION
|
||||
{
|
||||
|
||||
4
.github/workflows/bump_docs.yaml
vendored
4
.github/workflows/bump_docs.yaml
vendored
@@ -15,10 +15,6 @@ jobs:
|
||||
steps:
|
||||
- uses: actions/checkout@v7
|
||||
- name: Bump dependencies 🔧
|
||||
env:
|
||||
# Authenticated API calls get 1000 req/hour instead of the 60/hour
|
||||
# anonymous cap that is shared across every job on the runner's IP.
|
||||
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
bash .github/bump_docs.sh ${{ matrix.repository }}
|
||||
- name: Create Pull Request
|
||||
|
||||
39
.github/workflows/ci-tools-tests.yaml
vendored
39
.github/workflows/ci-tools-tests.yaml
vendored
@@ -1,39 +0,0 @@
|
||||
---
|
||||
# The packages under .github/ci/ are invisible to `go list ./...`, so neither
|
||||
# `make lint` nor the repository test run ever touches them. Their specs are
|
||||
# dead weight until a workflow names each package explicitly.
|
||||
name: 'CI tool tests'
|
||||
on:
|
||||
pull_request:
|
||||
paths:
|
||||
- '.github/ci/**'
|
||||
- '.github/workflows/ci-tools-tests.yaml'
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
paths:
|
||||
- '.github/ci/**'
|
||||
jobs:
|
||||
ci-tools:
|
||||
name: 'Test the .github/ci generators'
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v7
|
||||
- uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version-file: go.mod
|
||||
cache: false
|
||||
|
||||
# The discovery heuristics are the risky part of these tools. A regression
|
||||
# produces confident, wrong gallery entries, which is worse than no tool.
|
||||
- name: 'Test the APEX entry generator'
|
||||
run: go test ./.github/ci/apexentries/
|
||||
|
||||
- name: 'Test the variant proposer'
|
||||
run: go test ./.github/ci/variantproposals/
|
||||
|
||||
# Shared by both generators above. Its behaviour is exercised through their
|
||||
# specs; this step exists so a break in the shared package fails under its
|
||||
# own name rather than as a puzzling failure in whichever caller ran first.
|
||||
- name: 'Test the shared gallery editor'
|
||||
run: go test ./.github/ci/galleryedit/
|
||||
54
.github/workflows/gallery_variant_proposals.yaml
vendored
54
.github/workflows/gallery_variant_proposals.yaml
vendored
@@ -1,54 +0,0 @@
|
||||
name: Propose gallery variant groupings
|
||||
on:
|
||||
schedule:
|
||||
- cron: 0 4 * * 1
|
||||
workflow_dispatch:
|
||||
jobs:
|
||||
variant_proposals:
|
||||
if: github.repository == 'mudler/LocalAI'
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v7
|
||||
- uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version-file: go.mod
|
||||
cache: false
|
||||
|
||||
# The heuristics are the risky part of this job. A regression in them
|
||||
# produces confident, wrong proposals, which is worse than no job at all.
|
||||
- name: Test the proposer
|
||||
run: go test ./.github/ci/variantproposals/
|
||||
|
||||
- name: Propose groupings 🔧
|
||||
id: propose
|
||||
run: |
|
||||
rm -f /tmp/variant-proposals-body.md
|
||||
go run ./.github/ci/variantproposals \
|
||||
-index gallery/index.yaml \
|
||||
-ledger gallery/variant-exclusions.yaml \
|
||||
-body-out /tmp/variant-proposals-body.md \
|
||||
-apply
|
||||
if [ -s /tmp/variant-proposals-body.md ]; then
|
||||
echo "have_proposals=true" >> "$GITHUB_OUTPUT"
|
||||
{
|
||||
echo 'body<<VARIANT_PROPOSAL_BODY_EOF'
|
||||
cat /tmp/variant-proposals-body.md
|
||||
echo VARIANT_PROPOSAL_BODY_EOF
|
||||
} >> "$GITHUB_OUTPUT"
|
||||
else
|
||||
echo "have_proposals=false" >> "$GITHUB_OUTPUT"
|
||||
fi
|
||||
|
||||
# No body file means the proposer found nothing. Opening an empty pull
|
||||
# request every run is how a proposal job gets muted by its reviewers.
|
||||
- name: Create Pull Request
|
||||
if: steps.propose.outputs.have_proposals == 'true'
|
||||
uses: peter-evans/create-pull-request@v8
|
||||
with:
|
||||
token: ${{ secrets.UPDATE_BOT_TOKEN }}
|
||||
push-to-fork: ci-forks/LocalAI
|
||||
commit-message: 'chore(model-gallery): propose variant groupings'
|
||||
title: 'chore(model-gallery): propose variant groupings for review'
|
||||
branch: "propose/variant-groupings"
|
||||
body: ${{ steps.propose.outputs.body }}
|
||||
signoff: true
|
||||
2
.github/workflows/image-pr.yml
vendored
2
.github/workflows/image-pr.yml
vendored
@@ -52,7 +52,7 @@
|
||||
tag-latest: 'false'
|
||||
tag-suffix: '-gpu-nvidia-cuda-13'
|
||||
runs-on: 'ubuntu-latest'
|
||||
base-image: "ubuntu:24.04"
|
||||
base-image: "ubuntu:22.04"
|
||||
makeflags: "--jobs=3 --output-sync=target"
|
||||
ubuntu-version: '2404'
|
||||
- build-type: 'hipblas'
|
||||
|
||||
2
.github/workflows/image.yml
vendored
2
.github/workflows/image.yml
vendored
@@ -113,7 +113,7 @@
|
||||
tag-latest: 'auto'
|
||||
tag-suffix: '-gpu-nvidia-cuda-13'
|
||||
runs-on: 'ubuntu-latest'
|
||||
base-image: "ubuntu:24.04"
|
||||
base-image: "ubuntu:22.04"
|
||||
skip-drivers: 'false'
|
||||
makeflags: "--jobs=4 --output-sync=target"
|
||||
ubuntu-version: '2404'
|
||||
|
||||
20
.github/workflows/lint.yml
vendored
20
.github/workflows/lint.yml
vendored
@@ -46,23 +46,3 @@ jobs:
|
||||
touch core/http/react-ui/dist/index.html
|
||||
- name: lint
|
||||
run: make lint
|
||||
|
||||
build-scripts:
|
||||
# The image packaging scripts encode invariants that only surface inside a
|
||||
# container build (a missing transitive dep, a partial cuDNN family). Their
|
||||
# shell tests need nothing but bash + gcc + ldd, so run them on every PR
|
||||
# rather than waiting on a multi-GB cross-arch backend image build.
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v7
|
||||
- name: run packaging script tests
|
||||
run: make test-build-scripts
|
||||
|
||||
# The backend matrix path filter fails silently: a miss emits an empty
|
||||
# matrix, every job goes green, and the change reaches no image (#10946).
|
||||
# Its tests need only node, so they ride along with this job.
|
||||
- uses: actions/setup-node@v7
|
||||
with:
|
||||
node-version: '20'
|
||||
- name: run CI script tests
|
||||
run: make test-ci-scripts
|
||||
|
||||
179
.github/workflows/llama-cpp-paged-canary.yml
vendored
Normal file
179
.github/workflows/llama-cpp-paged-canary.yml
vendored
Normal file
@@ -0,0 +1,179 @@
|
||||
name: 'llama.cpp paged patches: upstream canary'
|
||||
|
||||
# EARLY-WARNING CANARY for the vendored paged-attention patch series
|
||||
# (backend/cpp/llama-cpp-localai-paged/patches/paged/0001-0030).
|
||||
#
|
||||
# WHY THIS EXISTS
|
||||
# The paged backend (backend/cpp/llama-cpp-localai-paged) pins its OWN verified
|
||||
# llama.cpp tip (LLAMA_VERSION in backend/cpp/llama-cpp-localai-paged/Makefile)
|
||||
# and is intentionally EXCLUDED from the nightly auto-bumper
|
||||
# (.github/workflows/bump_deps.yaml), so a naive upstream bump can never silently
|
||||
# break the shipped build. The cost of that safety: nobody finds out when
|
||||
# upstream DRIFTS past the patches. This canary restores that signal WITHOUT
|
||||
# touching the shipped pin - weekly it tries the patch series + a real compile
|
||||
# against the LATEST llama.cpp master tip and goes red the moment upstream breaks
|
||||
# the patches.
|
||||
#
|
||||
# RED HERE means: time to run a PIN_SYNC (rebase the patches onto the new tip,
|
||||
# pass the bit-exact gate on the GPU, re-export the .patch files, THEN advance
|
||||
# the pin in backend/cpp/llama-cpp-localai-paged/Makefile). See the backend README
|
||||
# section 7 (Pin + maintenance policy):
|
||||
# backend/cpp/llama-cpp-localai-paged/README.md.
|
||||
#
|
||||
# SIGNAL-ONLY: this workflow moves no pinned version, ships nothing, and is fully
|
||||
# decoupled from bump_deps - so the main dep-bump PR stays green regardless. A
|
||||
# green run means "the paged series still applies and compiles on upstream HEAD";
|
||||
# a red run means "upstream moved - schedule a pin-sync".
|
||||
|
||||
on:
|
||||
schedule:
|
||||
# Weekly (Mondays 06:00 UTC), mirroring the weekly DEPS_REFRESH / bump_deps
|
||||
# cadence. Offset from bump_deps' nightly 20:00 so the two never pile up.
|
||||
- cron: '0 6 * * 1'
|
||||
workflow_dispatch:
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
concurrency:
|
||||
group: llama-cpp-paged-canary
|
||||
cancel-in-progress: false
|
||||
|
||||
env:
|
||||
# Upstream source of truth - the same repo/branch bump_deps tracks for the
|
||||
# stock llama-cpp pin.
|
||||
LLAMA_UPSTREAM: 'https://github.com/ggml-org/llama.cpp'
|
||||
|
||||
jobs:
|
||||
apply-check:
|
||||
# Cheap, fast, toolchain-free early warning: does the series still APPLY to
|
||||
# the latest upstream tip? A patch no longer applying is by far the most
|
||||
# common way upstream breaks a vendored series, so this runs first, is
|
||||
# reliable on a free runner, and feeds the resolved tip to the compile job.
|
||||
if: github.repository == 'mudler/LocalAI'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 20
|
||||
outputs:
|
||||
tip: ${{ steps.resolve.outputs.tip }}
|
||||
steps:
|
||||
- name: Checkout LocalAI
|
||||
uses: actions/checkout@v7
|
||||
|
||||
- name: Resolve latest llama.cpp master tip
|
||||
id: resolve
|
||||
run: |
|
||||
tip="$(git ls-remote "$LLAMA_UPSTREAM" refs/heads/master | cut -f1)"
|
||||
if [ -z "$tip" ]; then
|
||||
echo "::error::could not resolve llama.cpp master tip from $LLAMA_UPSTREAM"
|
||||
exit 1
|
||||
fi
|
||||
pin="$(grep -m1 'LLAMA_VERSION?=' backend/cpp/llama-cpp-localai-paged/Makefile | cut -d= -f2)"
|
||||
echo "latest llama.cpp master tip: $tip"
|
||||
echo "shipped paged pin: $pin"
|
||||
echo "tip=$tip" >> "$GITHUB_OUTPUT"
|
||||
{
|
||||
echo "## llama.cpp paged canary"
|
||||
echo ""
|
||||
echo "- upstream master tip: \`$tip\`"
|
||||
echo "- shipped paged pin: \`$pin\`"
|
||||
} >> "$GITHUB_STEP_SUMMARY"
|
||||
|
||||
- name: Checkout llama.cpp at latest tip (shallow)
|
||||
run: |
|
||||
mkdir -p /tmp/llama.cpp
|
||||
cd /tmp/llama.cpp
|
||||
git init -q
|
||||
git remote add origin "$LLAMA_UPSTREAM"
|
||||
git fetch -q --depth 1 origin "${{ steps.resolve.outputs.tip }}"
|
||||
git checkout -q FETCH_HEAD
|
||||
git log --oneline -1
|
||||
|
||||
- name: Apply paged patch series (build's git-apply method)
|
||||
run: |
|
||||
bash .github/scripts/paged-canary-apply.sh \
|
||||
/tmp/llama.cpp \
|
||||
"$PWD/backend/cpp/llama-cpp-localai-paged/patches"
|
||||
echo "- apply: full paged series applies to the upstream tip :white_check_mark:" >> "$GITHUB_STEP_SUMMARY"
|
||||
|
||||
compile:
|
||||
# Proves the patches still COMPILE against the latest tip, using the SAME
|
||||
# toolchain + build target the shipped paged backend uses (the
|
||||
# base-grpc-cuda-12 builder base + the Makefile `grpc-server` cublas target),
|
||||
# so a failure means upstream drift, not toolchain noise. CUDA is compiled
|
||||
# (nvcc; no GPU required) because most of the paged series is CUDA kernels.
|
||||
# Runs only if the apply check passed, on the exact tip it validated.
|
||||
#
|
||||
# If a full CUDA compile on the hosted runner ever proves too heavy/flaky,
|
||||
# switch `runs-on` to 'bigger-runner' (the runner class the real paged CUDA
|
||||
# build uses), or drop to a CPU build (BUILD_TYPE='') which still compiles
|
||||
# all host + CPU paged code, leaving CUDA-kernel coverage to the apply check
|
||||
# plus the manual PIN_SYNC GPU gate.
|
||||
needs: apply-check
|
||||
if: github.repository == 'mudler/LocalAI'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 180
|
||||
steps:
|
||||
- name: Checkout LocalAI
|
||||
uses: actions/checkout@v7
|
||||
|
||||
- name: Free disk space
|
||||
uses: ./.github/actions/free-disk-space
|
||||
with:
|
||||
mode: hosted
|
||||
|
||||
- name: Login to Quay.io
|
||||
uses: docker/login-action@v4
|
||||
with:
|
||||
registry: quay.io
|
||||
username: ${{ secrets.LOCALAI_REGISTRY_USERNAME }}
|
||||
password: ${{ secrets.LOCALAI_REGISTRY_PASSWORD }}
|
||||
|
||||
- name: Compile paged backend against latest tip (cublas)
|
||||
env:
|
||||
TIP: ${{ needs.apply-check.outputs.tip }}
|
||||
BUILDER_BASE_IMAGE: 'quay.io/go-skynet/ci-cache:base-grpc-cuda-12-amd64'
|
||||
run: |
|
||||
docker run --rm \
|
||||
-v "$PWD":/LocalAI -w /LocalAI \
|
||||
-e TIP -e LLAMA_UPSTREAM \
|
||||
"$BUILDER_BASE_IMAGE" bash -euxo pipefail -c '
|
||||
# Mirror the Dockerfile: gRPC lives at /opt/grpc in the base image;
|
||||
# copy it to the prefix CMake find_package expects.
|
||||
cp -a /opt/grpc/. /usr/local/
|
||||
|
||||
# Pre-populate the llama.cpp checkout at the latest tip with the
|
||||
# paged series applied via the tolerant canary apply. Because
|
||||
# backend/cpp/llama-cpp/llama.cpp now exists, the stock Makefile's
|
||||
# llama.cpp target (clone + base-patch apply) is skipped and the
|
||||
# now patch-free prepare.sh only copies the grpc-server sources -
|
||||
# so we drive the REAL grpc-server build path on top of our paged
|
||||
# apply. The stock llama-cpp backend no longer carries the paged
|
||||
# series (it lives in backend/cpp/llama-cpp-localai-paged/patches/
|
||||
# paged); we build it here in the stock dir only because that is
|
||||
# where the shared build infra (Makefile / grpc-server.cpp /
|
||||
# CMakeLists.txt / prepare.sh) lives.
|
||||
cd backend/cpp/llama-cpp/
|
||||
mkdir -p llama.cpp
|
||||
cd llama.cpp
|
||||
git init -q
|
||||
git remote add origin "$LLAMA_UPSTREAM"
|
||||
git fetch -q --depth 1 origin "$TIP"
|
||||
git checkout -q FETCH_HEAD
|
||||
cd /LocalAI
|
||||
bash .github/scripts/paged-canary-apply.sh \
|
||||
backend/cpp/llama-cpp/llama.cpp \
|
||||
"$PWD/backend/cpp/llama-cpp-localai-paged/patches"
|
||||
|
||||
# Cheapest real CUDA build that proves the patches compile: one
|
||||
# CUDA arch, cublas. CMAKE_ARGS is passed via the environment (not
|
||||
# as a make arg) so the Makefile += flags are still appended,
|
||||
# exactly like .docker/llama-cpp-localai-paged-compile.sh. The paged
|
||||
# series is already applied to the checkout above, so the stock
|
||||
# build just compiles the patched tree.
|
||||
cd backend/cpp/llama-cpp/
|
||||
BUILD_TYPE=cublas \
|
||||
CMAKE_ARGS="-DCMAKE_CUDA_ARCHITECTURES=80" \
|
||||
make grpc-server
|
||||
test -x grpc-server
|
||||
'
|
||||
echo "- compile: paged series builds (cublas) against the upstream tip :white_check_mark:" >> "$GITHUB_STEP_SUMMARY"
|
||||
69
.github/workflows/realtime-conformance.yml
vendored
69
.github/workflows/realtime-conformance.yml
vendored
@@ -1,69 +0,0 @@
|
||||
---
|
||||
name: 'realtime-conformance'
|
||||
|
||||
# Verifies the realtime state-machine implementations conform to their formal
|
||||
# designs (docs/design/realtime-state-machines.md, formal-verification/). BOTH
|
||||
# layers are enforced and the gate is fail-closed: the Go conformance layer
|
||||
# (respcoord + turncoord transition/rapid tests under -race) AND the FizzBee model check of
|
||||
# the authoritative specs. FizzBee is pinned + checksum-verified
|
||||
# (formal-verification/fizzbee.sha256), so a failed install fails the job rather
|
||||
# than silently skipping verification.
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
paths:
|
||||
- 'core/http/endpoints/openai/coordinator/**'
|
||||
- 'core/http/endpoints/openai/respcoord/**'
|
||||
- 'core/http/endpoints/openai/turncoord/**'
|
||||
- 'core/http/endpoints/openai/conncoord/**'
|
||||
- 'core/http/endpoints/openai/compactcoord/**'
|
||||
- 'core/http/endpoints/openai/ttscoord/**'
|
||||
- 'formal-verification/**'
|
||||
- 'scripts/realtime-conformance.sh'
|
||||
- 'scripts/install-fizzbee.sh'
|
||||
- '.github/workflows/realtime-conformance.yml'
|
||||
push:
|
||||
branches:
|
||||
- master
|
||||
paths:
|
||||
- 'core/http/endpoints/openai/coordinator/**'
|
||||
- 'core/http/endpoints/openai/respcoord/**'
|
||||
- 'core/http/endpoints/openai/turncoord/**'
|
||||
- 'core/http/endpoints/openai/conncoord/**'
|
||||
- 'core/http/endpoints/openai/compactcoord/**'
|
||||
- 'core/http/endpoints/openai/ttscoord/**'
|
||||
- 'formal-verification/**'
|
||||
- 'scripts/realtime-conformance.sh'
|
||||
|
||||
concurrency:
|
||||
group: realtime-conformance-${{ github.event.pull_request.number || github.sha }}-${{ github.repository }}
|
||||
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
|
||||
|
||||
jobs:
|
||||
conformance:
|
||||
runs-on: ubuntu-latest
|
||||
strategy:
|
||||
matrix:
|
||||
go-version: ['1.26.x']
|
||||
steps:
|
||||
- name: Clone
|
||||
uses: actions/checkout@v7
|
||||
- name: Setup Go ${{ matrix.go-version }}
|
||||
uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version: ${{ matrix.go-version }}
|
||||
cache: false
|
||||
- name: Cache FizzBee
|
||||
uses: actions/cache@v6
|
||||
with:
|
||||
path: .tools/fizzbee
|
||||
key: fizzbee-v0.5.2-${{ runner.os }}-${{ hashFiles('formal-verification/fizzbee.sha256') }}
|
||||
- name: Install FizzBee (pinned, checksum-verified)
|
||||
# No `|| true`: a failed/forged download must fail the job, not silently
|
||||
# drop the design verification. install-fizzbee.sh is a no-op if the
|
||||
# cached binary is already present and valid.
|
||||
run: ./scripts/install-fizzbee.sh
|
||||
- name: Run conformance gate (fail-closed)
|
||||
# No skip env: both the Go conformance and the FizzBee model check are
|
||||
# required. The gate auto-detects .tools/fizzbee/fizz.
|
||||
run: make test-realtime-conformance
|
||||
13
.github/workflows/secscan.yaml
vendored
13
.github/workflows/secscan.yaml
vendored
@@ -7,19 +7,6 @@ on:
|
||||
schedule:
|
||||
- cron: '0 0 * * 0'
|
||||
|
||||
# `push:` is deliberately unfiltered, so this fires on every push to every
|
||||
# branch and there is no pull_request event to key on -- the usual
|
||||
# `github.event.pull_request.number || github.sha` idiom used elsewhere would
|
||||
# key on the unique-per-commit sha and dedup nothing. Group on the ref instead
|
||||
# so successive pushes to the same feature branch supersede one another.
|
||||
#
|
||||
# Cancelling is safe here: the only output is a SARIF upload, and code scanning
|
||||
# tracks the latest result per ref, so a superseded scan has nothing to lose.
|
||||
# master is excluded anyway -- every commit on master gets its own scan.
|
||||
concurrency:
|
||||
group: ci-secscan-${{ github.ref }}-${{ github.repository }}
|
||||
cancel-in-progress: ${{ github.ref != 'refs/heads/master' }}
|
||||
|
||||
jobs:
|
||||
tests:
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
2
.github/workflows/stalebot.yml
vendored
2
.github/workflows/stalebot.yml
vendored
@@ -11,7 +11,7 @@ jobs:
|
||||
if: github.repository == 'mudler/LocalAI'
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/stale@1e223db275d687790206a7acac4d1a11bd6fe629 # v9
|
||||
- uses: actions/stale@eb5cf3af3ac0a1aa4c9c45633dd1ae542a27a899 # v9
|
||||
with:
|
||||
stale-issue-message: 'This issue is stale because it has been open 90 days with no activity. Remove stale label or comment or this will be closed in 5 days.'
|
||||
stale-pr-message: 'This PR is stale because it has been open 90 days with no activity. Remove stale label or comment or this will be closed in 10 days.'
|
||||
|
||||
35
.github/workflows/test-extra.yml
vendored
35
.github/workflows/test-extra.yml
vendored
@@ -37,7 +37,6 @@ jobs:
|
||||
sglang: ${{ steps.detect.outputs.sglang }}
|
||||
acestep-cpp: ${{ steps.detect.outputs.acestep-cpp }}
|
||||
qwen3-tts-cpp: ${{ steps.detect.outputs.qwen3-tts-cpp }}
|
||||
magpie-tts-cpp: ${{ steps.detect.outputs.magpie-tts-cpp }}
|
||||
rfdetr-cpp: ${{ steps.detect.outputs.rfdetr-cpp }}
|
||||
locate-anything-cpp: ${{ steps.detect.outputs.locate-anything-cpp }}
|
||||
vibevoice-cpp: ${{ steps.detect.outputs.vibevoice-cpp }}
|
||||
@@ -588,7 +587,7 @@ jobs:
|
||||
with:
|
||||
go-version: '1.25.4'
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v7
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: '22'
|
||||
- name: Build sherpa-onnx backend image and run realtime e2e tests
|
||||
@@ -867,38 +866,6 @@ jobs:
|
||||
- name: Test qwen3-tts-cpp
|
||||
run: |
|
||||
make --jobs=5 --output-sync=target -C backend/go/qwen3-tts-cpp test
|
||||
tests-magpie-tts-cpp:
|
||||
needs: detect-changes
|
||||
if: needs.detect-changes.outputs.magpie-tts-cpp == 'true' || needs.detect-changes.outputs.run-all == 'true'
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Clone
|
||||
uses: actions/checkout@v7
|
||||
with:
|
||||
submodules: true
|
||||
- name: Dependencies
|
||||
run: |
|
||||
sudo apt-get update
|
||||
sudo apt-get install -y build-essential cmake curl libopenblas-dev ffmpeg
|
||||
- name: Setup Go
|
||||
uses: actions/setup-go@v5
|
||||
- name: Display Go version
|
||||
run: go version
|
||||
- name: Proto Dependencies
|
||||
run: |
|
||||
# Install protoc
|
||||
curl -L -s https://github.com/protocolbuffers/protobuf/releases/download/v26.1/protoc-26.1-linux-x86_64.zip -o protoc.zip && \
|
||||
unzip -j -d /usr/local/bin protoc.zip bin/protoc && \
|
||||
rm protoc.zip
|
||||
go install google.golang.org/protobuf/cmd/protoc-gen-go@v1.34.2
|
||||
go install google.golang.org/grpc/cmd/protoc-gen-go-grpc@1958fcbe2ca8bd93af633f11e97d44e567e945af
|
||||
PATH="$PATH:$HOME/go/bin" make protogen-go
|
||||
- name: Build magpie-tts-cpp
|
||||
run: |
|
||||
make --jobs=5 --output-sync=target -C backend/go/magpie-tts-cpp
|
||||
- name: Test magpie-tts-cpp
|
||||
run: |
|
||||
make --jobs=5 --output-sync=target -C backend/go/magpie-tts-cpp test
|
||||
# Per-backend smoke for rfdetr-cpp: builds the .so + Go binary and runs
|
||||
# `make -C backend/go/rfdetr-cpp test`. test.sh fetches the small (~20 MB)
|
||||
# rfdetr-nano-q8_0 GGUF from the published mudler/rfdetr-cpp-nano HF repo
|
||||
|
||||
4
.github/workflows/test.yml
vendored
4
.github/workflows/test.yml
vendored
@@ -48,7 +48,7 @@ jobs:
|
||||
sudo apt-get update
|
||||
sudo apt-get install curl ffmpeg libopus-dev
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v7
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: '22'
|
||||
- name: Build React UI
|
||||
@@ -100,7 +100,7 @@ jobs:
|
||||
brew install protobuf grpc make protoc-gen-go protoc-gen-go-grpc libomp llvm opus ffmpeg
|
||||
pip install --user --no-cache-dir grpcio-tools grpcio
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v7
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: '22'
|
||||
- name: Build React UI
|
||||
|
||||
2
.github/workflows/tests-e2e.yml
vendored
2
.github/workflows/tests-e2e.yml
vendored
@@ -47,7 +47,7 @@ jobs:
|
||||
sudo apt-get update
|
||||
sudo apt-get install -y build-essential libopus-dev
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v7
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: '22'
|
||||
- name: Build React UI
|
||||
|
||||
2
.github/workflows/tests-ui-e2e.yml
vendored
2
.github/workflows/tests-ui-e2e.yml
vendored
@@ -34,7 +34,7 @@ jobs:
|
||||
go-version: ${{ matrix.go-version }}
|
||||
cache: false
|
||||
- name: Setup Node.js
|
||||
uses: actions/setup-node@v7
|
||||
uses: actions/setup-node@v6
|
||||
with:
|
||||
node-version: '22'
|
||||
- name: Setup Bun
|
||||
|
||||
5
.github/workflows/yaml-check.yml
vendored
5
.github/workflows/yaml-check.yml
vendored
@@ -1,11 +1,6 @@
|
||||
name: 'Yamllint GitHub Actions'
|
||||
on:
|
||||
- pull_request
|
||||
|
||||
concurrency:
|
||||
group: ci-yamllint-${{ github.event.pull_request.number || github.sha }}-${{ github.repository }}
|
||||
cancel-in-progress: ${{ github.event_name == 'pull_request' }}
|
||||
|
||||
jobs:
|
||||
yamllint:
|
||||
name: 'Yamllint'
|
||||
|
||||
32
.gitignore
vendored
32
.gitignore
vendored
@@ -9,6 +9,15 @@ prepare-sources
|
||||
/backend/cpp/llama-cpp/llama.cpp
|
||||
/backend/cpp/llama-*
|
||||
!backend/cpp/llama-cpp
|
||||
# llama-cpp-localai-paged is a tracked source dir (a thin wrapper Makefile over
|
||||
# backend/cpp/llama-cpp). Re-include it like llama-cpp above; its sibling
|
||||
# *-build dirs are still ignored by the /backend/cpp/llama-* rule, and its
|
||||
# in-dir build artifacts (binaries, package output, collected ggml .so set) are
|
||||
# re-ignored just below.
|
||||
!backend/cpp/llama-cpp-localai-paged
|
||||
/backend/cpp/llama-cpp-localai-paged/llama-cpp-localai-paged-*
|
||||
/backend/cpp/llama-cpp-localai-paged/package
|
||||
/backend/cpp/llama-cpp-localai-paged/ggml-shared-libs
|
||||
/backends
|
||||
/backend-images
|
||||
/result.yaml
|
||||
@@ -41,12 +50,7 @@ models/*
|
||||
test-models/
|
||||
test-dir/
|
||||
tests/e2e-aio/backends
|
||||
# The mock backend binary built by `make build-mock-backend`. Anchored to its
|
||||
# full path: a bare `mock-backend` also matched the *directory* holding the
|
||||
# source, so git would not descend into it and adding a file there needed -f.
|
||||
# tests/e2e/mock-backend/.gitignore covers the same binary; kept here too so
|
||||
# the artifact stays ignored if that scoped file is ever removed.
|
||||
/tests/e2e/mock-backend/mock-backend
|
||||
mock-backend
|
||||
|
||||
release/
|
||||
|
||||
@@ -102,19 +106,3 @@ core/http/react-ui/test-results/
|
||||
|
||||
# Local Apple signing material (never commit)
|
||||
.certs/
|
||||
|
||||
# Pinned dev tools (e.g. FizzBee for the realtime-conformance gate)
|
||||
.tools/
|
||||
|
||||
# FizzBee model-check artifacts: the parser emits <spec>.json next to each
|
||||
# .fizz and the checker writes run dirs under out/. Both are regenerated by
|
||||
# the realtime-conformance gate; only the .fizz sources are authoritative.
|
||||
formal-verification/*.json
|
||||
formal-verification/out/
|
||||
|
||||
# `go build ./.github/ci/apexentries` drops a binary of the package name into
|
||||
# whatever directory it runs in, one `git add -A` away from being committed.
|
||||
# Both paths are anchored: an unanchored `apexentries` would also match the
|
||||
# package directory itself and untrack the source.
|
||||
/apexentries
|
||||
/.github/ci/apexentries/apexentries
|
||||
|
||||
@@ -1,6 +0,0 @@
|
||||
{
|
||||
"files": ["core/http/react-ui/index.html"],
|
||||
"insertBefore": "</body>",
|
||||
"commentSyntax": "html",
|
||||
"cspChecked": true
|
||||
}
|
||||
@@ -23,6 +23,8 @@ LocalAI follows the Linux kernel project's [guidelines for AI coding assistants]
|
||||
| [.agents/adding-backends.md](.agents/adding-backends.md) | Adding a new backend (Python, Go, or C++) — full step-by-step checklist, including importer integration (the `/import-model` dropdown is server-driven from `GET /backends/known`) |
|
||||
| [.agents/coding-style.md](.agents/coding-style.md) | Code style, editorconfig, logging, documentation conventions |
|
||||
| [.agents/llama-cpp-backend.md](.agents/llama-cpp-backend.md) | Working on the llama.cpp backend — architecture, updating, tool call parsing |
|
||||
| [.agents/llama-cpp-localai-paged-backend.md](.agents/llama-cpp-localai-paged-backend.md) | Working on the CUDA-only paged-attention llama.cpp variant (Qwen3.6 hybrid-SSM / Blackwell NVFP4 decode) - patchset scope, the bit-exact gate, the manual pin-sync + weekly canary, CUDA-only invariants, stock-stays-pure, Metal/SYCL/Vulkan follow-up scope |
|
||||
| [.agents/vllm-parity-methodology.md](.agents/vllm-parity-methodology.md) | The methodology for closing the vLLM decode-throughput gap in llama.cpp - bit-exact gating, profile-don't-assume, both-engine ground-truth, per-lever A/B discipline, recording rejected levers, multi-agent GPU orchestration |
|
||||
| [.agents/vllm-backend.md](.agents/vllm-backend.md) | Working on the vLLM / vLLM-omni backends — native parsers, ChatDelta, CPU build, libnuma packaging, backend hooks |
|
||||
| [.agents/sglang-backend.md](.agents/sglang-backend.md) | Working on the SGLang backend — `engine_args` validation against ServerArgs, speculative-decoding (EAGLE/EAGLE3/DFLASH/MTP) recipes, parser handling |
|
||||
| [.agents/ds4-backend.md](.agents/ds4-backend.md) | Working on the ds4 backend - DSML state machine, thinking modes, KV cache, Metal+CUDA matrix |
|
||||
@@ -35,14 +37,14 @@ LocalAI follows the Linux kernel project's [guidelines for AI coding assistants]
|
||||
|
||||
## Quick Reference
|
||||
|
||||
- **Coverage gates**: Never lower a coverage baseline or widen a gate's tolerance to turn a red gate green — the coverage ratchet only moves up. If a change drops coverage, add tests to raise it (e.g. render-smoke specs). See [.agents/building-and-testing.md](.agents/building-and-testing.md).
|
||||
- **Git hooks & coverage gates**: Run `make install-hooks` once per clone so the pre-commit lint + coverage gates run. **Never bypass them with `git commit --no-verify`, and never lower a coverage baseline or widen a gate's tolerance to turn a red gate green** — the coverage ratchet only moves up. If a change drops coverage, add tests to raise it (e.g. render-smoke specs). See [.agents/building-and-testing.md](.agents/building-and-testing.md).
|
||||
- **Logging**: Use `github.com/mudler/xlog` (same API as slog)
|
||||
- **Paged llama.cpp backend**: `llama-cpp-localai-paged` is a CUDA-only variant that owns its own patch series + its own pinned llama.cpp (manual pin-sync, weekly canary); the stock `llama-cpp` backend stays patch-free. Read [.agents/llama-cpp-localai-paged-backend.md](.agents/llama-cpp-localai-paged-backend.md) before touching either, and [.agents/vllm-parity-methodology.md](.agents/vllm-parity-methodology.md) for the decode-parity methodology behind it.
|
||||
- **Go style**: Prefer `any` over `interface{}`
|
||||
- **Comments**: Explain *why*, not *what*
|
||||
- **Docs (docs-with-code rule)**: When you change user-facing behavior (API endpoints, CLI flags, config keys, or features), update the corresponding page under `docs/content/` in the SAME change, not as a follow-up. A user-facing change without a matching docs update is incomplete. See also the documentation conventions in [.agents/coding-style.md](.agents/coding-style.md).
|
||||
- **Docs**: Update `docs/content/` when adding features or changing config
|
||||
- **New API endpoints**: LocalAI advertises its capability surface in several independent places — swagger `@Tags`, `/api/instructions` registry, auth `RouteFeatureRegistry`, React UI `capabilities.js`, docs. Read [.agents/api-endpoints-and-auth.md](.agents/api-endpoints-and-auth.md) and follow its checklist — missing any surface means clients, admins, and the UI won't know the endpoint exists.
|
||||
- **Admin endpoints → MCP tool**: every admin endpoint that an admin would manage conversationally (install/list/edit/toggle/upgrade) MUST also be exposed as an MCP tool in `pkg/mcp/localaitools/`. The LocalAI Assistant chat modality and the standalone `local-ai mcp-server` consume that package; drift between REST and MCP is a real risk. Read [.agents/localai-assistant-mcp.md](.agents/localai-assistant-mcp.md) — the `TestToolHTTPRouteMappingComplete` test fails until you wire the new tool and update the route map.
|
||||
- **Build**: Inspect `Makefile` and `.github/workflows/` — ask the user before running long builds
|
||||
- **Backend OS coverage**: a new backend must target every OS it can build for, not just Linux. `.github/backend-matrix.yml` has two matrices — `include:` (Linux) and `includeDarwin:` (macOS / Apple Silicon). Most C/C++/GGML and many Python backends build on Darwin too — wire the `includeDarwin` entry + `backend/index.yaml` `metal:` entries, or say in the PR why an OS is unsupported. See the darwin checklist in [.agents/adding-backends.md](.agents/adding-backends.md).
|
||||
- **Gallery variant ranking**: a gallery entry can declare `variants` (alternative builds of the same weights), and LocalAI ranks the ones a host can run by engine preference first, size second. A new backend that should be preferred on some hardware must be listed in `engineNamePreferenceRules` in `pkg/system/capabilities.go`; the sibling `backendBuildTagPreferenceRules` speaks build tags rather than engine names, and using the wrong table matches nothing without erroring. See [.agents/adding-backends.md](.agents/adding-backends.md).
|
||||
- **UI**: The active UI is the React app in `core/http/react-ui/`. The older Alpine.js/HTML UI in `core/http/static/` is pending deprecation — all new UI work goes in the React UI
|
||||
|
||||
@@ -198,6 +198,7 @@ For AI-assisted development, see [`AGENTS.md`](AGENTS.md) (or the equivalent [`C
|
||||
|
||||
- Prefer modern Go idioms — for example, use `any` instead of `interface{}`.
|
||||
- Use [`golangci-lint`](https://golangci-lint.run) to catch common issues before submitting a PR.
|
||||
- Run `make install-hooks` once per clone to enable the pre-commit hook: Go changes run `make lint` + the coverage gate (`make test-coverage-check`); `core/http/react-ui/` changes run the Playwright e2e suite (`make test-ui`). Bypass a single commit with `git commit --no-verify`.
|
||||
- Use [`github.com/mudler/xlog`](https://github.com/mudler/xlog) for logging (same API as `slog`). Do not use `fmt.Println` or the standard `log` package for operational logging.
|
||||
- Use tab indentation for Go files (as defined in `.editorconfig`).
|
||||
|
||||
@@ -267,7 +268,7 @@ make test-e2e
|
||||
|
||||
### React UI tests and coverage
|
||||
|
||||
The React UI (`core/http/react-ui/`) is covered by Playwright e2e specs, gated by a **monotonic line-coverage ratchet** (`make test-ui-coverage-check`, run in CI). The metric is non-deterministic — a fast local box reads higher than a slow CI runner for the same code — so a small tolerance is unavoidable.
|
||||
The React UI (`core/http/react-ui/`) is covered by Playwright e2e specs, gated by a **monotonic line-coverage ratchet** (`make test-ui-coverage-check`, run in CI and pre-commit). The metric is non-deterministic — a fast local box reads higher than a slow CI runner for the same code — so a small tolerance is unavoidable.
|
||||
|
||||
**If your change lowers UI coverage, raise it back by adding specs — do not widen the tolerance or hand-lower the baseline.** A *render-smoke* spec (navigate to a page, assert its header is visible) cheaply covers an entire lazy page. See `core/http/react-ui/e2e/page-render-smoke.spec.js` and the full policy in [.agents/building-and-testing.md](.agents/building-and-testing.md#react-ui-coverage).
|
||||
|
||||
|
||||
44
Dockerfile
44
Dockerfile
@@ -12,16 +12,12 @@ ARG APT_MIRROR
|
||||
ARG APT_PORTS_MIRROR
|
||||
ENV DEBIAN_FRONTEND=noninteractive
|
||||
|
||||
# hwdata ships /usr/share/hwdata/pci.ids. Without it, the ghw library we use
|
||||
# for hardware detection cannot resolve PCI vendor IDs and fails to enumerate
|
||||
# GPUs at all, so the image reports "No GPU detected" (see issue #10941).
|
||||
RUN --mount=type=bind,source=.docker/apt-mirror.sh,target=/usr/local/sbin/apt-mirror \
|
||||
APT_MIRROR="${APT_MIRROR}" APT_PORTS_MIRROR="${APT_PORTS_MIRROR}" sh /usr/local/sbin/apt-mirror && \
|
||||
apt-get update && \
|
||||
apt-get install -y --no-install-recommends \
|
||||
ca-certificates curl wget espeak-ng libgomp1 \
|
||||
ffmpeg libopenblas0 libopenblas-dev libopus0 sox \
|
||||
hwdata && \
|
||||
ffmpeg libopenblas0 libopenblas-dev libopus0 sox && \
|
||||
apt-get clean && \
|
||||
rm -rf /var/lib/apt/lists/*
|
||||
|
||||
@@ -175,17 +171,6 @@ RUN if [ "${BUILD_TYPE}" = "hipblas" ]; then \
|
||||
ln -s /opt/rocm-**/lib/llvm/lib/libomp.so /usr/lib/libomp.so \
|
||||
; fi
|
||||
|
||||
# ROCm's bundled libdrm_amdgpu is built with a hardcoded fallback lookup path
|
||||
# for the ASIC ID table (/opt/amdgpu/share/libdrm/amdgpu.ids), which only exists
|
||||
# if AMD's full amdgpu graphics/DKMS stack is installed. This compute-only image
|
||||
# doesn't have it, so hipblas/rocBLAS log "No such file or directory" on every
|
||||
# model load and can fail to identify the GPU. Point it at the equivalent file
|
||||
# Ubuntu's libdrm-common package already ships.
|
||||
RUN if [ "${BUILD_TYPE}" = "hipblas" ] && [ -f /usr/share/libdrm/amdgpu.ids ] && [ ! -e /opt/amdgpu/share/libdrm/amdgpu.ids ]; then \
|
||||
mkdir -p /opt/amdgpu/share/libdrm && \
|
||||
ln -s /usr/share/libdrm/amdgpu.ids /opt/amdgpu/share/libdrm/amdgpu.ids \
|
||||
; fi
|
||||
|
||||
RUN expr "${BUILD_TYPE}" = intel && echo "intel" > /run/localai/capability || echo "not intel"
|
||||
|
||||
# Cuda
|
||||
@@ -393,12 +378,7 @@ RUN go install github.com/mikefarah/yq/v4@latest
|
||||
# If you cannot find a more suitable place for an addition, this layer is a suitable place for it.
|
||||
FROM requirements-drivers
|
||||
|
||||
# Optional override for the HEALTHCHECK target. Left empty so healthcheck.sh
|
||||
# derives the endpoint from the mode the container is actually running — the
|
||||
# same image runs `local-ai run` (HTTP on 8080) and `local-ai worker` (HTTP on
|
||||
# the gRPC base port minus one), and a hardcoded default marked every worker
|
||||
# permanently unhealthy (#10987). Set it to pin an explicit URL.
|
||||
ENV HEALTHCHECK_ENDPOINT=""
|
||||
ENV HEALTHCHECK_ENDPOINT=http://localhost:8080/readyz
|
||||
|
||||
ARG CUDA_MAJOR_VERSION=12
|
||||
ENV NVIDIA_DRIVER_CAPABILITIES=compute,utility
|
||||
@@ -408,7 +388,6 @@ ENV NVIDIA_VISIBLE_DEVICES=all
|
||||
WORKDIR /
|
||||
|
||||
COPY ./entrypoint.sh .
|
||||
COPY ./scripts/build/healthcheck.sh .
|
||||
|
||||
# Copy the binary
|
||||
COPY --from=builder /build/local-ai ./
|
||||
@@ -419,22 +398,9 @@ RUN --mount=from=builder,src=/build/,dst=/mnt/build \
|
||||
# Make sure the models directory exists
|
||||
RUN mkdir -p /models /backends /data
|
||||
|
||||
# Define the health check command.
|
||||
#
|
||||
# --start-period is the knob for slow starts, not --timeout/--retries. Since
|
||||
# #10949 a frontend's startup preload materializes HuggingFace artifacts before
|
||||
# the HTTP server binds (31 GB observed on a live cluster), so a healthy replica
|
||||
# can legitimately fail probes for a long time. Failures inside the start period
|
||||
# leave the container `starting` instead of burning retries, and the period ends
|
||||
# early on the first success — so a generous value costs a fast-starting
|
||||
# container nothing. A process that actually died is handled by the restart
|
||||
# policy, not by health.
|
||||
#
|
||||
# --timeout is a per-probe deadline: 10m meant a wedged probe could hang for ten
|
||||
# minutes and stretch detection without bound. A localhost curl that has not
|
||||
# answered in 10s is itself the fault being detected.
|
||||
HEALTHCHECK --start-period=60m --interval=1m --timeout=10s --retries=3 \
|
||||
CMD /healthcheck.sh
|
||||
# Define the health check command
|
||||
HEALTHCHECK --interval=1m --timeout=10m --retries=10 \
|
||||
CMD curl -f ${HEALTHCHECK_ENDPOINT} || exit 1
|
||||
|
||||
VOLUME /models /backends /configuration /data
|
||||
EXPOSE 8080
|
||||
|
||||
124
Makefile
124
Makefile
@@ -1,5 +1,5 @@
|
||||
# Disable parallel execution for backend builds
|
||||
.NOTPARALLEL: backends/diffusers backends/llama-cpp backends/turboquant backends/bonsai backends/outetts backends/piper backends/stablediffusion-ggml backends/whisper backends/crispasr backends/parakeet-cpp backends/moss-transcribe-cpp backends/faster-whisper backends/silero-vad backends/local-store backends/cloud-proxy backends/huggingface backends/rfdetr backends/rfdetr-cpp backends/insightface backends/speaker-recognition backends/kitten-tts backends/kokoro backends/chatterbox backends/llama-cpp-darwin backends/neutts build-darwin-python-backend build-darwin-go-backend backends/mlx backends/diffuser-darwin backends/mlx-vlm backends/mlx-audio backends/mlx-distributed backends/stablediffusion-ggml-darwin backends/vllm backends/vllm-omni backends/longcat-video backends/sglang backends/moonshine backends/pocket-tts backends/qwen-tts backends/faster-qwen3-tts backends/qwen-asr backends/nemo backends/voxcpm backends/whisperx backends/ace-step backends/acestep-cpp backends/fish-speech backends/voxtral backends/opus backends/trl backends/llama-cpp-quantization backends/kokoros backends/sam3-cpp backends/qwen3-tts-cpp backends/moss-tts-cpp backends/magpie-tts-cpp backends/vllm-cpp backends/omnivoice-cpp backends/vibevoice-cpp backends/localvqe backends/tinygrad backends/sherpa-onnx backends/ds4 backends/ds4-darwin backends/liquid-audio backends/supertonic backends/depth-anything-cpp backends/privacy-filter backends/privacy-filter-darwin
|
||||
.NOTPARALLEL: backends/diffusers backends/llama-cpp backends/turboquant backends/outetts backends/piper backends/stablediffusion-ggml backends/whisper backends/crispasr backends/parakeet-cpp backends/faster-whisper backends/silero-vad backends/local-store backends/huggingface backends/rfdetr backends/rfdetr-cpp backends/insightface backends/speaker-recognition backends/kitten-tts backends/kokoro backends/chatterbox backends/llama-cpp-darwin backends/neutts build-darwin-python-backend build-darwin-go-backend backends/mlx backends/diffuser-darwin backends/mlx-vlm backends/mlx-audio backends/mlx-distributed backends/stablediffusion-ggml-darwin backends/vllm backends/vllm-omni backends/sglang backends/moonshine backends/pocket-tts backends/qwen-tts backends/faster-qwen3-tts backends/qwen-asr backends/nemo backends/voxcpm backends/whisperx backends/ace-step backends/acestep-cpp backends/fish-speech backends/voxtral backends/opus backends/trl backends/llama-cpp-quantization backends/kokoros backends/sam3-cpp backends/qwen3-tts-cpp backends/omnivoice-cpp backends/vibevoice-cpp backends/localvqe backends/tinygrad backends/sherpa-onnx backends/ds4 backends/ds4-darwin backends/liquid-audio backends/supertonic backends/depth-anything-cpp backends/privacy-filter backends/privacy-filter-darwin backends/llama-cpp-localai-paged
|
||||
|
||||
GOCMD=go
|
||||
GOTEST=$(GOCMD) test
|
||||
@@ -103,7 +103,7 @@ COVERAGE_E2E_LABELS?=!real-models
|
||||
COVERAGE_EXCLUDE_RE?=grpc/proto/.*[.]pb[.]go
|
||||
|
||||
|
||||
.PHONY: all test test-coverage test-coverage-baseline test-coverage-check test-backend-cpp test-build-scripts test-ui test-ui-coverage-baseline test-ui-coverage-check build vendor lint lint-all
|
||||
.PHONY: all test test-coverage test-coverage-baseline test-coverage-check test-backend-cpp test-ui test-ui-coverage-baseline test-ui-coverage-check install-hooks build vendor lint lint-all
|
||||
|
||||
all: help
|
||||
|
||||
@@ -208,20 +208,6 @@ test: prepare-test
|
||||
test-backend-cpp:
|
||||
bash backend/cpp/run-unit-tests.sh
|
||||
|
||||
## Runs the shell-level regression tests for the image packaging scripts
|
||||
## (scripts/build/*_test.sh). These guard invariants that only ever break
|
||||
## inside a container build - a missing transitive dep, a partial cuDNN
|
||||
## family - and that no Go test can observe. Needs only bash + gcc + ldd.
|
||||
test-build-scripts:
|
||||
@set -e; for t in scripts/build/*_test.sh; do echo "== $$t"; bash "$$t"; done
|
||||
|
||||
## Runs the unit tests for the CI helper scripts under scripts/lib/. Currently
|
||||
## the backend matrix path filter, whose failure mode is invisible in CI: it
|
||||
## emits an empty matrix, every job goes green, and the change ships to no
|
||||
## image at all (see PR #10946). Plain `node --test`, no dependencies.
|
||||
test-ci-scripts:
|
||||
@set -e; for t in scripts/lib/*_test.mjs; do echo "== $$t"; node --test "$$t"; done
|
||||
|
||||
## Runs the core suite ($(TEST_PATHS)) with statement-coverage instrumentation
|
||||
## and writes a merged profile to $(COVERAGE_PROFILE). Deliberately omits
|
||||
## --fail-fast so a single failure doesn't truncate the coverage number, and
|
||||
@@ -269,7 +255,8 @@ LINT_EXCLUDE_DIRS_RE=/(backend/go/(piper|silero-vad|llm)|cmd/launcher)(/|$$)
|
||||
|
||||
## Set LINT_NEW_FROM to a git ref to override .golangci.yml's
|
||||
## new-from-merge-base (origin/master). Useful from a fork clone where
|
||||
## origin/master is stale relative to the canonical repo.
|
||||
## origin/master is stale relative to the canonical repo — the pre-commit
|
||||
## hook passes the resolved upstream ref here so local lint matches CI.
|
||||
LINT_NEW_FROM?=
|
||||
lint:
|
||||
@command -v golangci-lint >/dev/null 2>&1 || { \
|
||||
@@ -288,6 +275,17 @@ lint-all:
|
||||
}
|
||||
golangci-lint run --new=false --new-from-merge-base= --new-from-rev= $$(go list -e -f '{{.Dir}}' ./... | grep -vE '$(LINT_EXCLUDE_DIRS_RE)')
|
||||
|
||||
########################################################
|
||||
## Git hooks
|
||||
########################################################
|
||||
## Points git at the versioned .githooks/ directory so the pre-commit hook
|
||||
## (lint + coverage gate) runs locally. Run once per clone. Undo with:
|
||||
## `git config --unset core.hooksPath`. Skip a single commit with
|
||||
## `git commit --no-verify`.
|
||||
install-hooks:
|
||||
git config core.hooksPath .githooks
|
||||
@echo 'Installed git hooks: core.hooksPath -> .githooks (pre-commit runs lint + test-coverage-check on Go changes)'
|
||||
|
||||
########################################################
|
||||
## E2E AIO tests (uses standard image with pre-configured models)
|
||||
########################################################
|
||||
@@ -407,23 +405,6 @@ test-realtime: build-mock-backend
|
||||
@echo 'Running realtime e2e tests (mock backend)'
|
||||
$(GOCMD) run github.com/onsi/ginkgo/v2/ginkgo --label-filter="Realtime && !real-models" --flake-attempts $(TEST_FLAKES) -v -r ./tests/e2e
|
||||
|
||||
# Verify the realtime state-machine implementations conform to their formal
|
||||
# designs (Go transition/rapid tests under -race + FizzBee model check of the
|
||||
# authoritative specs). See docs/design/realtime-state-machines.md (Part 6) and
|
||||
# docs/design/specs/README.md.
|
||||
test-realtime-conformance:
|
||||
GOCMD=$(GOCMD) ./scripts/realtime-conformance.sh
|
||||
|
||||
# Verify the shared model-loader shutdown behavior independently of any API
|
||||
# modality (focused loader/gRPC/distributed/worker tests under -race + FizzBee).
|
||||
test-model-lifecycle-conformance:
|
||||
GOCMD=$(GOCMD) ./scripts/model-lifecycle-conformance.sh
|
||||
|
||||
# Install the pinned, checksum-verified FizzBee model checker (into .tools/,
|
||||
# gitignored) used by the conformance targets. Idempotent; no-op if present.
|
||||
install-fizzbee:
|
||||
./scripts/install-fizzbee.sh
|
||||
|
||||
# Container-based real-model realtime testing. Build env vars / pipeline
|
||||
# definition kept here so test-realtime-models-docker can drive a fully wired
|
||||
# pipeline (VAD + STT + LLM + TTS) from inside a containerised runner.
|
||||
@@ -572,7 +553,6 @@ prepare-test-extra: protogen-python
|
||||
$(MAKE) -C backend/python/chatterbox
|
||||
$(MAKE) -C backend/python/vllm
|
||||
$(MAKE) -C backend/python/vllm-omni
|
||||
$(MAKE) -C backend/python/longcat-video
|
||||
$(MAKE) -C backend/python/sglang
|
||||
$(MAKE) -C backend/python/vibevoice
|
||||
$(MAKE) -C backend/python/liquid-audio
|
||||
@@ -602,7 +582,6 @@ test-extra: prepare-test-extra
|
||||
$(MAKE) -C backend/python/chatterbox test
|
||||
$(MAKE) -C backend/python/vllm test
|
||||
$(MAKE) -C backend/python/vllm-omni test
|
||||
$(MAKE) -C backend/python/longcat-video test
|
||||
$(MAKE) -C backend/python/vibevoice test
|
||||
$(MAKE) -C backend/python/liquid-audio test
|
||||
$(MAKE) -C backend/python/moonshine test
|
||||
@@ -625,7 +604,6 @@ test-extra: prepare-test-extra
|
||||
$(MAKE) -C backend/go/locate-anything-cpp test
|
||||
$(MAKE) -C backend/go/depth-anything-cpp test
|
||||
$(MAKE) -C backend/go/supertonic test
|
||||
$(MAKE) -C backend/go/vllm-cpp test
|
||||
|
||||
##
|
||||
## End-to-end gRPC tests that exercise a built backend container image.
|
||||
@@ -655,9 +633,6 @@ test-extra: prepare-test-extra
|
||||
## suite against it.
|
||||
##
|
||||
BACKEND_TEST_MODEL_URL?=https://huggingface.co/Qwen/Qwen3-0.6B-GGUF/resolve/main/Qwen3-0.6B-Q8_0.gguf
|
||||
## Suite timeout for `go test`. Wrappers whose model download alone can eat
|
||||
## most of the default (multi-GB models on a slow HF CDN day) override this.
|
||||
BACKEND_TEST_TIMEOUT?=30m
|
||||
|
||||
## Generic target — runs the suite against whatever BACKEND_IMAGE points at.
|
||||
## Depends on protogen-go so pkg/grpc/proto is generated before `go test`.
|
||||
@@ -685,7 +660,7 @@ test-extra-backend: protogen-go
|
||||
BACKEND_TEST_FACE_IMAGE_3_URL="$$BACKEND_TEST_FACE_IMAGE_3_URL" \
|
||||
BACKEND_TEST_FACE_IMAGE_3_FILE="$$BACKEND_TEST_FACE_IMAGE_3_FILE" \
|
||||
BACKEND_TEST_VERIFY_DISTANCE_CEILING="$$BACKEND_TEST_VERIFY_DISTANCE_CEILING" \
|
||||
go test -v -timeout $(BACKEND_TEST_TIMEOUT) ./tests/e2e-backends/...
|
||||
go test -v -timeout 30m ./tests/e2e-backends/...
|
||||
|
||||
## Convenience wrappers: build the image, then exercise it.
|
||||
test-extra-backend-llama-cpp: docker-build-llama-cpp
|
||||
@@ -696,6 +671,15 @@ test-extra-backend-llama-cpp: docker-build-llama-cpp
|
||||
test-extra-backend-ik-llama-cpp: docker-build-ik-llama-cpp
|
||||
BACKEND_IMAGE=local-ai-backend:ik-llama-cpp $(MAKE) test-extra-backend
|
||||
|
||||
## llama-cpp-localai-paged: the LocalAI paged-attention llama.cpp variant. Same
|
||||
## GGUF surface as stock llama-cpp (the paged engine is runtime-gated by the
|
||||
## LLAMA_KV_PAGED env the grpc-server option hooks set), so the standard
|
||||
## llama-cpp capability set is what we exercise here.
|
||||
test-extra-backend-llama-cpp-localai-paged: docker-build-llama-cpp-localai-paged
|
||||
BACKEND_IMAGE=local-ai-backend:llama-cpp-localai-paged \
|
||||
BACKEND_TEST_CAPS=health,load,predict,stream,logprobs,logit_bias \
|
||||
$(MAKE) test-extra-backend
|
||||
|
||||
## turboquant: exercises the llama.cpp-fork backend with the fork's
|
||||
## *TurboQuant-specific* KV-cache types (turbo3 for both K and V). turbo3
|
||||
## is what makes this backend distinct from stock llama-cpp — picking q8_0
|
||||
@@ -708,16 +692,6 @@ test-extra-backend-turboquant: docker-build-turboquant
|
||||
BACKEND_TEST_CACHE_TYPE_V=turbo3 \
|
||||
$(MAKE) test-extra-backend
|
||||
|
||||
## bonsai: exercises the llama.cpp-fork backend with a real Q1_0 (1-bit) model —
|
||||
## the PrismML Bonsai-8B GGUF, whose weight quant is *only* decodable by the fork's
|
||||
## Q1_0 kernels. Loading it is what makes this backend distinct from stock llama-cpp;
|
||||
## a standard-quant model would only test the upstream code path the llama-cpp backend
|
||||
## already covers.
|
||||
test-extra-backend-bonsai: docker-build-bonsai
|
||||
BACKEND_IMAGE=local-ai-backend:bonsai \
|
||||
BACKEND_TEST_MODEL_URL=https://huggingface.co/prism-ml/Bonsai-8B-gguf/resolve/main/Bonsai-8B-Q1_0.gguf \
|
||||
$(MAKE) test-extra-backend
|
||||
|
||||
## Audio transcription wrapper for the llama-cpp backend.
|
||||
## Drives the new AudioTranscription / AudioTranscriptionStream RPCs against
|
||||
## ggml-org/Qwen3-ASR-0.6B-GGUF (a small ASR model that requires its mmproj
|
||||
@@ -1035,7 +1009,6 @@ test-extra-backend-vibevoice-cpp-tts: docker-build-vibevoice-cpp
|
||||
## post-image disk budget.
|
||||
test-extra-backend-vibevoice-cpp-transcription: docker-build-vibevoice-cpp
|
||||
BACKEND_IMAGE=local-ai-backend:vibevoice-cpp \
|
||||
BACKEND_TEST_TIMEOUT=120m \
|
||||
BACKEND_TEST_MODEL_URL='https://huggingface.co/mudler/vibevoice.cpp-models/resolve/main/vibevoice-asr-q4_k.gguf#vibevoice-asr-q4_k.gguf' \
|
||||
BACKEND_TEST_EXTRA_FILES='https://huggingface.co/mudler/vibevoice.cpp-models/resolve/main/tokenizer.gguf#tokenizer.gguf' \
|
||||
BACKEND_TEST_AUDIO_URL=https://github.com/ggml-org/whisper.cpp/raw/master/samples/jfk.wav \
|
||||
@@ -1063,19 +1036,7 @@ test-extra-backend-whisper-transcription: docker-build-whisper
|
||||
## is reachable.
|
||||
test-extra-backend-parakeet-cpp-transcription: docker-build-parakeet-cpp
|
||||
BACKEND_IMAGE=local-ai-backend:parakeet-cpp \
|
||||
BACKEND_TEST_MODEL_URL=https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/realtime_eou_120m-v1-f16.gguf \
|
||||
BACKEND_TEST_AUDIO_URL=https://github.com/ggml-org/whisper.cpp/raw/master/samples/jfk.wav \
|
||||
BACKEND_TEST_CAPS=health,load,transcription \
|
||||
$(MAKE) test-extra-backend
|
||||
|
||||
## Audio transcription wrapper for the moss-transcribe-cpp (moss-transcribe.cpp
|
||||
## ggml port) backend. Mirrors test-extra-backend-parakeet-cpp-transcription:
|
||||
## drives the AudioTranscription RPC against a published MOSS GGUF using the JFK
|
||||
## 11s clip from whisper.cpp's CI samples. Not part of the default test suite -
|
||||
## run explicitly once the pinned model URL is reachable.
|
||||
test-extra-backend-moss-transcribe-cpp-transcription: docker-build-moss-transcribe-cpp
|
||||
BACKEND_IMAGE=local-ai-backend:moss-transcribe-cpp \
|
||||
BACKEND_TEST_MODEL_URL=https://huggingface.co/mudler/moss-transcribe.cpp-gguf/resolve/main/moss-transcribe-q5_k.gguf \
|
||||
BACKEND_TEST_MODEL_URL=https://huggingface.co/mudler/parakeet-cpp-gguf/resolve/main/tdt_ctc-110m-f16.gguf \
|
||||
BACKEND_TEST_AUDIO_URL=https://github.com/ggml-org/whisper.cpp/raw/master/samples/jfk.wav \
|
||||
BACKEND_TEST_CAPS=health,load,transcription \
|
||||
$(MAKE) test-extra-backend
|
||||
@@ -1229,10 +1190,10 @@ BACKEND_IK_LLAMA_CPP = ik-llama-cpp|ik-llama-cpp|.|false|false
|
||||
# turboquant is a llama.cpp fork with TurboQuant KV-cache quantization.
|
||||
# Reuses backend/cpp/llama-cpp grpc-server sources via a thin wrapper Makefile.
|
||||
BACKEND_TURBOQUANT = turboquant|turboquant|.|false|false
|
||||
# bonsai is a llama.cpp fork (PrismML) adding the Q1_0 (1-bit) and Q2_0 (ternary)
|
||||
# weight-quant kernels the Bonsai / Ternary-Bonsai models ship in. Reuses
|
||||
# backend/cpp/llama-cpp grpc-server sources via a thin wrapper Makefile.
|
||||
BACKEND_BONSAI = bonsai|bonsai|.|false|false
|
||||
# llama-cpp-localai-paged = stock llama.cpp grpc-server + the LocalAI paged-attention
|
||||
# patch series (vendored in this wrapper backend). Reuses backend/cpp/llama-cpp sources via a thin
|
||||
# wrapper Makefile (same upstream pin as stock llama-cpp; no fork, no patch-grpc-server).
|
||||
BACKEND_LLAMA_CPP_LOCALAI_PAGED = llama-cpp-localai-paged|llama-cpp-localai-paged|.|false|false
|
||||
# ds4 is antirez/ds4, a DeepSeek V4 Flash-specific inference engine.
|
||||
# Single-model; hardware-only validation lives at tests/e2e-backends/
|
||||
# (BACKEND_BINARY mode); see docs/superpowers/plans/2026-05-11-ds4-backend.md.
|
||||
@@ -1252,14 +1213,10 @@ BACKEND_STABLEDIFFUSION_GGML = stablediffusion-ggml|golang|.|--progress=plain|tr
|
||||
BACKEND_WHISPER = whisper|golang|.|false|true
|
||||
BACKEND_CRISPASR = crispasr|golang|.|false|true
|
||||
BACKEND_PARAKEET_CPP = parakeet-cpp|golang|.|false|true
|
||||
BACKEND_MOSS_TRANSCRIBE_CPP = moss-transcribe-cpp|golang|.|false|true
|
||||
BACKEND_DEPTH_ANYTHING_CPP = depth-anything-cpp|golang|.|false|true
|
||||
BACKEND_VOXTRAL = voxtral|golang|.|false|true
|
||||
BACKEND_ACESTEP_CPP = acestep-cpp|golang|.|false|true
|
||||
BACKEND_QWEN3_TTS_CPP = qwen3-tts-cpp|golang|.|false|true
|
||||
BACKEND_MOSS_TTS_CPP = moss-tts-cpp|golang|.|false|true
|
||||
BACKEND_MAGPIE_TTS_CPP = magpie-tts-cpp|golang|.|false|true
|
||||
BACKEND_VLLM_CPP = vllm-cpp|golang|.|false|true
|
||||
BACKEND_OMNIVOICE_CPP = omnivoice-cpp|golang|.|false|true
|
||||
BACKEND_VIBEVOICE_CPP = vibevoice-cpp|golang|.|false|true
|
||||
BACKEND_LOCALVQE = localvqe|golang|.|false|true
|
||||
@@ -1281,7 +1238,6 @@ BACKEND_NEUTTS = neutts|python|.|false|true
|
||||
BACKEND_KOKORO = kokoro|python|.|false|true
|
||||
BACKEND_VLLM = vllm|python|.|false|true
|
||||
BACKEND_VLLM_OMNI = vllm-omni|python|.|false|true
|
||||
BACKEND_LONGCAT_VIDEO = longcat-video|python|.|--progress=plain|true
|
||||
BACKEND_SGLANG = sglang|python|.|false|true
|
||||
BACKEND_DIFFUSERS = diffusers|python|.|--progress=plain|true
|
||||
BACKEND_CHATTERBOX = chatterbox|python|.|false|true
|
||||
@@ -1339,7 +1295,7 @@ endef
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_LLAMA_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_IK_LLAMA_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_TURBOQUANT)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_BONSAI)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_LLAMA_CPP_LOCALAI_PAGED)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_DS4)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_PRIVACY_FILTER)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_PIPER)))
|
||||
@@ -1351,7 +1307,6 @@ $(eval $(call generate-docker-build-target,$(BACKEND_STABLEDIFFUSION_GGML)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_WHISPER)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_CRISPASR)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_PARAKEET_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_MOSS_TRANSCRIBE_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_DEPTH_ANYTHING_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_VOXTRAL)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_OPUS)))
|
||||
@@ -1368,7 +1323,6 @@ $(eval $(call generate-docker-build-target,$(BACKEND_NEUTTS)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_KOKORO)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_VLLM)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_VLLM_OMNI)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_LONGCAT_VIDEO)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_SGLANG)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_DIFFUSERS)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_CHATTERBOX)))
|
||||
@@ -1386,9 +1340,6 @@ $(eval $(call generate-docker-build-target,$(BACKEND_WHISPERX)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_ACE_STEP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_ACESTEP_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_QWEN3_TTS_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_MOSS_TTS_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_MAGPIE_TTS_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_VLLM_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_OMNIVOICE_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_VIBEVOICE_CPP)))
|
||||
$(eval $(call generate-docker-build-target,$(BACKEND_LOCALVQE)))
|
||||
@@ -1408,7 +1359,7 @@ $(eval $(call generate-docker-build-target,$(BACKEND_SUPERTONIC)))
|
||||
docker-save-%: backend-images
|
||||
docker save local-ai-backend:$* -o backend-images/$*.tar
|
||||
|
||||
docker-build-backends: docker-build-llama-cpp docker-build-ik-llama-cpp docker-build-turboquant docker-build-bonsai docker-build-ds4 docker-build-rerankers docker-build-vllm docker-build-vllm-omni docker-build-longcat-video docker-build-sglang docker-build-transformers docker-build-outetts docker-build-diffusers docker-build-kokoro docker-build-faster-whisper docker-build-crispasr docker-build-coqui docker-build-chatterbox docker-build-vibevoice docker-build-liquid-audio docker-build-moonshine docker-build-pocket-tts docker-build-qwen-tts docker-build-fish-speech docker-build-faster-qwen3-tts docker-build-qwen-asr docker-build-nemo docker-build-voxcpm docker-build-whisperx docker-build-ace-step docker-build-acestep-cpp docker-build-voxtral docker-build-mlx-distributed docker-build-trl docker-build-llama-cpp-quantization docker-build-tinygrad docker-build-kokoros docker-build-sam3-cpp docker-build-rfdetr-cpp docker-build-qwen3-tts-cpp docker-build-moss-tts-cpp docker-build-magpie-tts-cpp docker-build-vllm-cpp docker-build-omnivoice-cpp docker-build-vibevoice-cpp docker-build-localvqe docker-build-insightface docker-build-speaker-recognition docker-build-sherpa-onnx docker-build-cloud-proxy docker-build-supertonic docker-build-depth-anything-cpp docker-build-moss-transcribe-cpp docker-build-privacy-filter
|
||||
docker-build-backends: docker-build-llama-cpp docker-build-ik-llama-cpp docker-build-turboquant docker-build-llama-cpp-localai-paged docker-build-ds4 docker-build-rerankers docker-build-vllm docker-build-vllm-omni docker-build-sglang docker-build-transformers docker-build-outetts docker-build-diffusers docker-build-kokoro docker-build-faster-whisper docker-build-crispasr docker-build-coqui docker-build-chatterbox docker-build-vibevoice docker-build-liquid-audio docker-build-moonshine docker-build-pocket-tts docker-build-qwen-tts docker-build-fish-speech docker-build-faster-qwen3-tts docker-build-qwen-asr docker-build-nemo docker-build-voxcpm docker-build-whisperx docker-build-ace-step docker-build-acestep-cpp docker-build-voxtral docker-build-mlx-distributed docker-build-trl docker-build-llama-cpp-quantization docker-build-tinygrad docker-build-kokoros docker-build-sam3-cpp docker-build-rfdetr-cpp docker-build-qwen3-tts-cpp docker-build-omnivoice-cpp docker-build-vibevoice-cpp docker-build-localvqe docker-build-insightface docker-build-speaker-recognition docker-build-sherpa-onnx docker-build-cloud-proxy docker-build-supertonic docker-build-depth-anything-cpp docker-build-privacy-filter
|
||||
|
||||
########################################################
|
||||
### Mock Backend for E2E Tests
|
||||
@@ -1443,7 +1394,7 @@ test-ui-e2e: build-ui-test-server
|
||||
UI_TEST_WORKERS ?=
|
||||
PLAYWRIGHT_WORKERS_FLAG = $(if $(UI_TEST_WORKERS),--workers=$(UI_TEST_WORKERS),)
|
||||
|
||||
## Fast Playwright e2e run for local React UI validation.
|
||||
## Fast Playwright e2e run used by the pre-commit hook on React UI changes.
|
||||
## Force-rebuilds the (non-instrumented) dist so the suite tests the working
|
||||
## tree — not a stale dist the `react-ui` skip-guard would leave — re-embeds
|
||||
## it into ui-test-server, and runs the specs. Uses the nix-provided browser
|
||||
@@ -1533,13 +1484,8 @@ build-launcher-darwin:
|
||||
mv cmd/launcher/LocalAI.app dist/LocalAI.app
|
||||
bash contrib/macos/sign-and-notarize.sh sign dist/LocalAI.app
|
||||
|
||||
# Notarize + staple the .app itself, then wrap it into a drag-to-Applications
|
||||
# DMG via hdiutil and sign the DMG. The app is stapled BEFORE packaging so the
|
||||
# bundle carries its own ticket and verifies offline (a dmg-only staple leaves
|
||||
# the app relying on an online Gatekeeper check, which fails offline / once the
|
||||
# app is copied out of the dmg). No-op without notary secrets.
|
||||
# Wrap the (signed) app into a drag-to-Applications DMG via hdiutil, then sign the DMG.
|
||||
dmg-launcher-darwin: build-launcher-darwin
|
||||
bash contrib/macos/sign-and-notarize.sh notarize-app dist/LocalAI.app
|
||||
rm -rf dist/dmg dist/LocalAI.dmg
|
||||
mkdir -p dist/dmg
|
||||
cp -R dist/LocalAI.app dist/dmg/LocalAI.app
|
||||
@@ -1551,7 +1497,7 @@ dmg-launcher-darwin: build-launcher-darwin
|
||||
notarize-launcher-darwin: dmg-launcher-darwin
|
||||
bash contrib/macos/sign-and-notarize.sh notarize dist/LocalAI.dmg
|
||||
|
||||
# Single entrypoint for CI: build -> sign app -> notarize+staple app -> dmg -> sign dmg -> notarize+staple dmg.
|
||||
# Single entrypoint for CI: build -> sign app -> dmg -> sign dmg -> notarize -> staple.
|
||||
release-launcher-darwin: notarize-launcher-darwin
|
||||
@echo "dist/LocalAI.dmg is ready"
|
||||
|
||||
|
||||
15
README.md
15
README.md
@@ -177,7 +177,7 @@ For more details, see the [Getting Started guide](https://localai.io/basics/gett
|
||||
|
||||
## Latest News
|
||||
|
||||
- **June 2026**: New native biometric backends from the LocalAI team: [voice-detect.cpp](https://github.com/localai-org/voice-detect.cpp) for speaker recognition and voice analysis (ECAPA-TDNN, WeSpeaker, ERes2Net, CAM++, wav2vec2 age/gender/emotion) and [face-detect.cpp](https://github.com/mudler/face-detect.cpp) for face detection, recognition, demographics and anti-spoofing (SCRFD/ArcFace, YuNet/SFace). Both are from-scratch C++/ggml engines with no Python or onnxruntime at inference, self-contained GGUF weights, bit-exact parity with the reference, and GPU cuDNN parity, replacing the heavier Python `insightface` and `speaker-recognition` backends ([PR #10441](https://github.com/mudler/LocalAI/pull/10441)).
|
||||
- **June 2026**: New native biometric backends from the LocalAI team: [voice-detect.cpp](https://github.com/mudler/voice-detect.cpp) for speaker recognition and voice analysis (ECAPA-TDNN, WeSpeaker, ERes2Net, CAM++, wav2vec2 age/gender/emotion) and [face-detect.cpp](https://github.com/mudler/face-detect.cpp) for face detection, recognition, demographics and anti-spoofing (SCRFD/ArcFace, YuNet/SFace). Both are from-scratch C++/ggml engines with no Python or onnxruntime at inference, self-contained GGUF weights, bit-exact parity with the reference, and GPU cuDNN parity, replacing the heavier Python `insightface` and `speaker-recognition` backends ([PR #10441](https://github.com/mudler/LocalAI/pull/10441)).
|
||||
- **June 2026**: New [realtime voice assistant demo](https://github.com/localai-org/localai-realtime-demo) (a tiny Go client for the Realtime API with a full talk-back voice loop and tool calling), plus [streaming of the realtime LLM / TTS / transcription pipeline stages](https://github.com/mudler/LocalAI/pull/10176) and [configurable WebRTC ICE candidates](https://github.com/mudler/LocalAI/pull/10231).
|
||||
- **June 2026**: Big speech push: the [parakeet.cpp](https://github.com/mudler/parakeet.cpp) ASR engine gains [NeMo-faithful segment timestamps](https://github.com/mudler/LocalAI/pull/10207), a [multilingual streaming Nemotron-3.5 model](https://github.com/mudler/LocalAI/pull/10199), [dynamic batching for concurrent transcription](https://github.com/mudler/LocalAI/pull/10112) and [CUDA graphs](https://github.com/mudler/LocalAI/pull/10273); the new [CrispASR backend](https://github.com/mudler/LocalAI/pull/10099) adds multi-architecture ASR + TTS, and [60 Piper TTS voices across 42 languages](https://github.com/mudler/LocalAI/pull/10296) land in the gallery (plus [per-request TTS instructions and params](https://github.com/mudler/LocalAI/pull/10172)).
|
||||
- **June 2026**: New backends and models: [locate-anything.cpp](https://github.com/mudler/LocalAI/pull/10264) for open-vocabulary object detection via ggml, [Ideogram4 image generation](https://github.com/mudler/LocalAI/pull/10201) in stablediffusion-ggml, [llama.cpp video input](https://github.com/mudler/LocalAI/pull/10216), and the [Gemma 4 QAT family with MTP speculative-decoding pairs](https://github.com/mudler/LocalAI/pull/10215). Plus an [interactive CLI chat mode](https://github.com/mudler/LocalAI/pull/10226) and [RAG source citations in agent responses](https://github.com/mudler/LocalAI/pull/10228).
|
||||
@@ -231,20 +231,13 @@ Most backends wrap a best-in-class upstream engine. A handful of them are native
|
||||
|
||||
| Backend | What it does |
|
||||
|---------|-------------|
|
||||
| [vllm.cpp](https://github.com/mudler/vllm.cpp) | From-scratch C++20 port of vLLM for text generation: paged KV cache, continuous batching, prefix caching, safetensors + GGUF loading, engine-enforced structured output, on CPU, CUDA, Metal and Vulkan |
|
||||
| [parakeet.cpp](https://github.com/mudler/parakeet.cpp) | C++/GGML port of NVIDIA NeMo Parakeet ASR (tdt/ctc/rnnt/hybrid), with cache-aware streaming transcription |
|
||||
| [moss-transcribe.cpp](https://github.com/localai-org/moss-transcribe.cpp) | C++/GGML port of OpenMOSS MOSS-Transcribe-Diarize: joint long-form transcription, speaker diarization and timestamping in a single pass |
|
||||
| [moss-tts.cpp](https://github.com/mudler/moss-tts.cpp) | C++/GGML port of the OpenMOSS MOSS-TTS family: text-to-speech (MOSS-TTS-Local v1.5, 48 kHz stereo) with reference-audio voice cloning, through the MOSS-Audio-Tokenizer neural codec |
|
||||
| [magpie-tts.cpp](https://github.com/mudler/magpie-tts.cpp) | C++/GGML port of NVIDIA's Magpie TTS Multilingual 357M: 22.05 kHz mono text-to-speech in 5 voices and 9+ languages, with the NanoCodec neural codec and tokenizer/G2P embedded in a single GGUF |
|
||||
| [ced.cpp](https://github.com/localai-org/ced.cpp) | C++/GGML port of the CED audio-tagging models: sound-event classification (527-class AudioSet) over REST and the realtime API for live recognition |
|
||||
| [voice-detect.cpp](https://github.com/localai-org/voice-detect.cpp) | Speaker recognition and voice analysis (ECAPA-TDNN, WeSpeaker, ERes2Net, CAM++, wav2vec2 age/gender/emotion), replacing the Python speaker-recognition backend |
|
||||
| [voxtral-tts.c](https://github.com/mudler/voxtral-tts.c) | Voxtral Realtime 4B speech-to-text in pure C |
|
||||
| [ced.cpp](https://github.com/mudler/ced.cpp) | C++/GGML port of the CED audio-tagging models: sound-event classification (527-class AudioSet) over REST and the realtime API for live recognition |
|
||||
| [voxtral.c](https://github.com/mudler/voxtral.c) | Voxtral Realtime 4B speech-to-text in pure C |
|
||||
| [vibevoice.cpp](https://github.com/mudler/vibevoice.cpp) | Native port of Microsoft VibeVoice for TTS (voice cloning) and long-form ASR with speaker diarization |
|
||||
| [rf-detr.cpp](https://github.com/localai-org/rf-detr.cpp) | Native RF-DETR object detection and instance segmentation |
|
||||
| [rf-detr.cpp](https://github.com/mudler/rf-detr.cpp) | Native RF-DETR object detection and instance segmentation |
|
||||
| [locate-anything.cpp](https://github.com/mudler/locate-anything.cpp) | Open-vocabulary object detection and visual grounding (LocateAnything-3B) |
|
||||
| [depth-anything.cpp](https://github.com/mudler/depth-anything.cpp) | Depth Anything 3 monocular metric depth + camera pose estimation |
|
||||
| [face-detect.cpp](https://github.com/mudler/face-detect.cpp) | Face detection, recognition, demographics and anti-spoofing (SCRFD/ArcFace, YuNet/SFace), replacing the Python insightface backend |
|
||||
| [free-splatter.cpp](https://github.com/localai-org/free-splatter.cpp) | Pose-free 3D reconstruction (FreeSplatter): turns a handful of plain photos into 3D Gaussians, no camera poses or GPU required |
|
||||
| [privacy-filter.cpp](https://github.com/localai-org/privacy-filter.cpp) | Standalone GGML PII/NER token-classification engine powering LocalAI's PII redaction tier |
|
||||
| [LocalVQE](https://github.com/localai-org/LocalVQE) | Joint acoustic echo cancellation, noise suppression, and dereverberation |
|
||||
| [local-store](https://github.com/mudler/LocalAI) | Local-first vector database for embeddings (shipped in-tree) |
|
||||
|
||||
@@ -221,33 +221,6 @@ RUN if [ "${BACKEND}" = "crispasr" ]; then \
|
||||
apt-get clean && rm -rf /var/lib/apt/lists/*; \
|
||||
fi
|
||||
|
||||
# sherpa-onnx links onnxruntime's CUDA execution provider, and
|
||||
# libonnxruntime_providers_cuda.so has cuDNN as a hard DT_NEEDED. The
|
||||
# onnxruntime GPU tarball does not ship cuDNN itself, so without this the
|
||||
# builder has none (the arm64 + CUDA 13 branch above is the only other place
|
||||
# that installs it) and package-gpu-libs.sh correctly refuses to produce a
|
||||
# package that references cuDNN with no cuDNN available to it.
|
||||
#
|
||||
# Installed per-backend rather than for every cublas build: the auto-detection
|
||||
# in package-gpu-libs.sh bundles only what a package actually references, so
|
||||
# the ggml backends would not grow either way, but they would all pay ~1.1 GB
|
||||
# of builder layer and registry cache for a library they never call.
|
||||
#
|
||||
# Runtime package only, no -dev: sherpa-onnx consumes onnxruntime's prebuilt
|
||||
# CUDA provider and never compiles against cuDNN headers. libcudnn9-cuda-N
|
||||
# carries the dispatcher plus all seven dlopen()ed sublibraries, which is what
|
||||
# complete_cudnn_family needs to assemble a whole bundle.
|
||||
RUN <<EOT bash
|
||||
if [ "${BACKEND}" = "sherpa-onnx" ] && [ "${BUILD_TYPE}" = "cublas" ] && [ "${SKIP_DRIVERS}" = "false" ]; then
|
||||
apt-get update && \
|
||||
apt-get install -y --no-install-recommends \
|
||||
libcudnn9-cuda-${CUDA_MAJOR_VERSION} && \
|
||||
ldconfig && \
|
||||
apt-get clean && \
|
||||
rm -rf /var/lib/apt/lists/*
|
||||
fi
|
||||
EOT
|
||||
|
||||
COPY . /LocalAI
|
||||
|
||||
RUN git config --global --add safe.directory /LocalAI
|
||||
|
||||
@@ -111,10 +111,6 @@ RUN make -BC /LocalAI/backend/cpp/llama-cpp package
|
||||
# ============================================================================
|
||||
FROM ${BUILDER_BASE_IMAGE} AS builder-prebuilt
|
||||
|
||||
ARG APT_MIRROR
|
||||
ENV APT_MIRROR=${APT_MIRROR}
|
||||
ARG APT_PORTS_MIRROR
|
||||
ENV APT_PORTS_MIRROR=${APT_PORTS_MIRROR}
|
||||
ARG BUILD_TYPE
|
||||
ENV BUILD_TYPE=${BUILD_TYPE}
|
||||
ARG CUDA_DOCKER_ARCH
|
||||
|
||||
@@ -7,7 +7,7 @@ ARG BUILDER_BASE_IMAGE=${BASE_IMAGE}
|
||||
# BUILDER_TARGET selects which builder stage the final scratch image copies
|
||||
# package output from. Declared at global scope (before any FROM) so it's
|
||||
# usable in `FROM ${BUILDER_TARGET}` below. Default keeps local
|
||||
# `make backends/bonsai` on the from-source path.
|
||||
# `make backends/llama-cpp-localai-paged` on the from-source path.
|
||||
ARG BUILDER_TARGET=builder-fromsource
|
||||
ARG APT_MIRROR=""
|
||||
ARG APT_PORTS_MIRROR=""
|
||||
@@ -18,7 +18,7 @@ ARG APT_PORTS_MIRROR=""
|
||||
# Runs .docker/install-base-deps.sh (apt deps + cmake + protoc + gRPC +
|
||||
# conditional CUDA/ROCm/Vulkan), copies /opt/grpc to /usr/local, then
|
||||
# compiles the variant. Used when BUILDER_TARGET=builder-fromsource (the
|
||||
# default; local `make backends/bonsai`).
|
||||
# default; local `make backends/llama-cpp-localai-paged`).
|
||||
#
|
||||
# The install script is the same one that backend/Dockerfile.base-grpc-builder
|
||||
# runs, so the result is bit-equivalent to the prebuilt-base path
|
||||
@@ -84,21 +84,22 @@ RUN cp -a /opt/grpc/. /usr/local/
|
||||
COPY . /LocalAI
|
||||
|
||||
# BuildKit cache mount for ccache. See Dockerfile.llama-cpp (commit 9228e5b4)
|
||||
# for rationale. bonsai is a llama.cpp fork that reuses
|
||||
# backend/cpp/llama-cpp source via a thin wrapper Makefile, so MOST TUs
|
||||
# are content-identical to the upstream llama-cpp build. Sharing a cache
|
||||
# id with llama-cpp could give cross-fork hits — but for now keep them
|
||||
# separate so a regression in one doesn't poison the other. Revisit
|
||||
# sharing after measuring the actual hit rate.
|
||||
# for rationale. llama-cpp-localai-paged is the SAME upstream llama.cpp with
|
||||
# the LocalAI paged patch series applied; it reuses backend/cpp/llama-cpp
|
||||
# source via a thin wrapper Makefile, so MOST TUs are content-identical to the
|
||||
# stock llama-cpp build. Sharing a cache id with llama-cpp could give
|
||||
# cross-variant hits — but for now keep them separate (mirroring turboquant) so
|
||||
# a regression in one doesn't poison the other. Revisit sharing after measuring
|
||||
# the actual hit rate.
|
||||
#
|
||||
# The compile body is shared with builder-prebuilt via .docker/bonsai-compile.sh.
|
||||
RUN --mount=type=bind,source=.docker/bonsai-compile.sh,target=/usr/local/sbin/compile.sh \
|
||||
--mount=type=cache,target=/root/.ccache,id=bonsai-ccache-${TARGETARCH}-${BUILD_TYPE},sharing=locked \
|
||||
# The compile body is shared with builder-prebuilt via .docker/llama-cpp-localai-paged-compile.sh.
|
||||
RUN --mount=type=bind,source=.docker/llama-cpp-localai-paged-compile.sh,target=/usr/local/sbin/compile.sh \
|
||||
--mount=type=cache,target=/root/.ccache,id=llama-cpp-localai-paged-ccache-${TARGETARCH}-${BUILD_TYPE},sharing=locked \
|
||||
bash /usr/local/sbin/compile.sh
|
||||
|
||||
|
||||
# Copy libraries using a script to handle architecture differences
|
||||
RUN make -BC /LocalAI/backend/cpp/bonsai package
|
||||
RUN make -BC /LocalAI/backend/cpp/llama-cpp-localai-paged package
|
||||
|
||||
|
||||
# ============================================================================
|
||||
@@ -107,7 +108,9 @@ RUN make -BC /LocalAI/backend/cpp/bonsai package
|
||||
# That image already has gRPC at /opt/grpc + apt deps + CUDA/ROCm/Vulkan
|
||||
# pre-installed, so we just copy gRPC to /usr/local and compile. Used when
|
||||
# BUILDER_TARGET=builder-prebuilt (CI when the matrix entry sets
|
||||
# builder-base-image).
|
||||
# builder-base-image). llama-cpp-localai-paged reuses the SAME base-grpc-* tags
|
||||
# as the stock llama-cpp backend (same gRPC + same toolchain), so no new
|
||||
# base-images.yml variant is required.
|
||||
# ============================================================================
|
||||
FROM ${BUILDER_BASE_IMAGE} AS builder-prebuilt
|
||||
|
||||
@@ -118,9 +121,9 @@ ENV CUDA_DOCKER_ARCH=${CUDA_DOCKER_ARCH}
|
||||
ARG CMAKE_ARGS
|
||||
ENV CMAKE_ARGS=${CMAKE_ARGS}
|
||||
# AMDGPU_TARGETS must be forwarded into the env here too — backend/cpp/llama-cpp/Makefile
|
||||
# (which the bonsai Makefile reuses via a sibling build dir) errors out when the var
|
||||
# is empty on a hipblas build, and the prebuilt path is what CI exercises most of the
|
||||
# time. The builder-fromsource stage above already does this; mirror it here.
|
||||
# (which the llama-cpp-localai-paged Makefile reuses via a sibling build dir) errors out
|
||||
# when the var is empty on a hipblas build, and the prebuilt path is what CI exercises most
|
||||
# of the time. The builder-fromsource stage above already does this; mirror it here.
|
||||
ARG AMDGPU_TARGETS
|
||||
ENV AMDGPU_TARGETS=${AMDGPU_TARGETS}
|
||||
ARG TARGETARCH
|
||||
@@ -133,11 +136,11 @@ RUN cp -a /opt/grpc/. /usr/local/
|
||||
|
||||
COPY . /LocalAI
|
||||
|
||||
RUN --mount=type=bind,source=.docker/bonsai-compile.sh,target=/usr/local/sbin/compile.sh \
|
||||
--mount=type=cache,target=/root/.ccache,id=bonsai-ccache-${TARGETARCH}-${BUILD_TYPE},sharing=locked \
|
||||
RUN --mount=type=bind,source=.docker/llama-cpp-localai-paged-compile.sh,target=/usr/local/sbin/compile.sh \
|
||||
--mount=type=cache,target=/root/.ccache,id=llama-cpp-localai-paged-ccache-${TARGETARCH}-${BUILD_TYPE},sharing=locked \
|
||||
bash /usr/local/sbin/compile.sh
|
||||
|
||||
RUN make -BC /LocalAI/backend/cpp/bonsai package
|
||||
RUN make -BC /LocalAI/backend/cpp/llama-cpp-localai-paged package
|
||||
|
||||
|
||||
# ============================================================================
|
||||
@@ -157,4 +160,4 @@ FROM scratch
|
||||
|
||||
|
||||
# Copy all available binaries (the build process only creates the appropriate ones for the target architecture)
|
||||
COPY --from=builder /LocalAI/backend/cpp/bonsai/package/. ./
|
||||
COPY --from=builder /LocalAI/backend/cpp/llama-cpp-localai-paged/package/. ./
|
||||
@@ -224,11 +224,7 @@ ARG DEPS_REFRESH=initial
|
||||
|
||||
RUN cd /${BACKEND} && PORTABLE_PYTHON=true make
|
||||
|
||||
# Package GPU libraries into the backend's lib directory.
|
||||
#
|
||||
# Must stay after the venv is built above: package-gpu-libs.sh inspects
|
||||
# /${BACKEND}/venv to decide whether this backend already carries a complete
|
||||
# cuDNN from pip, and bundles one only when it does not (issue #10905).
|
||||
# Package GPU libraries into the backend's lib directory
|
||||
RUN mkdir -p /${BACKEND}/lib && \
|
||||
TARGET_LIB_DIR="/${BACKEND}/lib" BUILD_TYPE="${BUILD_TYPE}" CUDA_MAJOR_VERSION="${CUDA_MAJOR_VERSION}" \
|
||||
bash /package-gpu-libs.sh "/${BACKEND}/lib"
|
||||
|
||||
@@ -46,7 +46,6 @@ The backend system provides language-specific Dockerfiles that handle the build
|
||||
- **vllm**: High-performance LLM inference
|
||||
- **mlx**: Apple Silicon optimization
|
||||
- **diffusers**: Stable Diffusion models
|
||||
- **longcat-video**: CUDA text/image-to-video and speech-driven avatar generation
|
||||
- **Audio**: coqui, faster-whisper, kitten-tts
|
||||
- **Vision**: mlx-vlm, rfdetr
|
||||
- **Specialized**: rerankers, chatterbox, kokoro
|
||||
|
||||
@@ -18,18 +18,6 @@ service Backend {
|
||||
rpc GenerateVideo(GenerateVideoRequest) returns (Result) {}
|
||||
rpc AudioTranscription(TranscriptRequest) returns (TranscriptResult) {}
|
||||
rpc AudioTranscriptionStream(TranscriptRequest) returns (stream TranscriptStreamResponse) {}
|
||||
// AudioTranscriptionLive is the bidirectional live-microphone ASR RPC. The
|
||||
// first message MUST carry a Config; subsequent messages carry Audio frames
|
||||
// (mono float PCM at config.sample_rate, 16 kHz default). After a
|
||||
// successful open the backend replies with a single ready ack
|
||||
// (TranscriptLiveResponse{ready:true}); backends or models without
|
||||
// cache-aware streaming support return UNIMPLEMENTED instead. Newly
|
||||
// finalized text streams back as deltas; eou=true marks the model's
|
||||
// end-of-utterance token. One stream spans many utterances (the decoder
|
||||
// resets itself after each EOU). Closing the send side finalizes: the
|
||||
// backend flushes the decoder tail and emits a terminal message carrying
|
||||
// final_result. A second Config mid-stream resets the decode session.
|
||||
rpc AudioTranscriptionLive(stream TranscriptLiveRequest) returns (stream TranscriptLiveResponse) {}
|
||||
rpc TTS(TTSRequest) returns (Result) {}
|
||||
rpc TTSStream(TTSRequest) returns (stream Reply) {}
|
||||
rpc SoundGeneration(SoundGenerationRequest) returns (Result) {}
|
||||
@@ -136,10 +124,6 @@ message MetricsResponse {
|
||||
message TokenClassifyRequest {
|
||||
string text = 1;
|
||||
float threshold = 2;
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 3;
|
||||
}
|
||||
|
||||
// TokenClassifyEntity is one detected entity span. Byte offsets are
|
||||
@@ -177,17 +161,6 @@ message ScoreRequest {
|
||||
// candidates differ in length and the consumer wants a per-token
|
||||
// measure comparable across them (PMI-style scoring).
|
||||
bool length_normalize = 4;
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 5;
|
||||
// Byte length of the prompt prefix that stays identical across
|
||||
// repeated scoring calls (e.g. a classifier's option-list system
|
||||
// prompt — everything before the per-turn probe text). Backends that
|
||||
// snapshot state (hybrid/recurrent models cannot rewind otherwise)
|
||||
// use it to place a reuse point exactly at the boundary, so the next
|
||||
// call re-processes only the tokens after it. 0 means unknown.
|
||||
int32 stable_prefix_len = 6;
|
||||
}
|
||||
|
||||
// CandidateScore is one row in the ScoreResponse, matching by index
|
||||
@@ -219,10 +192,6 @@ message RerankRequest {
|
||||
string query = 1;
|
||||
repeated string documents = 2;
|
||||
int32 top_n = 3;
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 4;
|
||||
}
|
||||
|
||||
message RerankResult {
|
||||
@@ -334,39 +303,6 @@ message PredictOptions {
|
||||
int32 TopLogprobs = 51; // Number of top logprobs to return per token (maps to OpenAI top_logprobs parameter)
|
||||
map<string, string> Metadata = 52; // Generic per-request metadata (e.g., enable_thinking)
|
||||
float MinP = 53; // Minimum probability sampling threshold (0.0 = disabled)
|
||||
|
||||
// ModelIdentity names the model this request is for, so a backend can reject
|
||||
// a request that reached it by mistake instead of answering from whatever
|
||||
// model it happens to hold. In distributed mode a worker can recycle a
|
||||
// stopped backend's gRPC port for a different model's backend, and a
|
||||
// liveness-only health probe cannot tell that apart from a valid cached
|
||||
// route (#10952).
|
||||
//
|
||||
// The value is the controller's ModelConfig.Model, the SAME expression that
|
||||
// produces ModelOptions.Model at LoadModel time, so the two are equal by
|
||||
// construction rather than by convention.
|
||||
//
|
||||
// Empty means "no identity supplied": backends MUST skip the check. That
|
||||
// keeps an old controller talking to a new backend working, and covers
|
||||
// callers that legitimately synthesize a PredictOptions internally.
|
||||
//
|
||||
// Do NOT reuse TTSRequest.model or SoundGenerationRequest.model for this
|
||||
// purpose. FileStagingClient already rewrites those to worker-local absolute
|
||||
// paths (core/services/nodes/file_staging_client.go), so in distributed mode
|
||||
// they already differ from the load-time value and comparing them would
|
||||
// reject valid requests. Extending identity to those RPCs needs a separate
|
||||
// field carrying the untranslated value - which is exactly what
|
||||
// TTSRequest.ModelIdentity and SoundGenerationRequest.ModelIdentity are.
|
||||
//
|
||||
// Every other request message that reaches a backend through the distributed
|
||||
// router now carries the same ModelIdentity field, populated from the same
|
||||
// ModelConfig.Model. FileStagingClient rewrites Src/Dst/Voice/Model/
|
||||
// StartImage/EndImage/Audio and never ModelIdentity, so what the backend
|
||||
// compares is always what the controller sent.
|
||||
string ModelIdentity = 54;
|
||||
|
||||
// 24 was never assigned; reserve it so it is not silently reused.
|
||||
reserved 24;
|
||||
}
|
||||
|
||||
// ToolCallDelta represents an incremental tool call update from the C++ parser.
|
||||
@@ -500,11 +436,6 @@ message ModelOptions {
|
||||
// Proxy carries the cloud-proxy backend's per-model configuration.
|
||||
// Empty for non-proxy backends.
|
||||
ProxyOptions Proxy = 74;
|
||||
|
||||
// EnableScore reserves backend resources for the Score RPC. It is derived
|
||||
// from the model's explicit `known_usecases: [score]` declaration so models
|
||||
// that never score retain their ordinary serving footprint.
|
||||
bool EnableScore = 75;
|
||||
}
|
||||
|
||||
// ProxyOptions configures the cloud-proxy backend. UpstreamURL and
|
||||
@@ -520,12 +451,6 @@ message ProxyOptions {
|
||||
string api_key_file = 5;
|
||||
string upstream_model = 6;
|
||||
int32 request_timeout_seconds = 7;
|
||||
// cache_prompt enables automatic Anthropic prompt-cache breakpoints
|
||||
// (cache_control: ephemeral) on the stable prefix — system, tools, and
|
||||
// the last message block — when translating to the Anthropic provider.
|
||||
// Cuts input cost on repeated/agentic calls (cache read = 0.1x). Only
|
||||
// meaningful for mode=translate + provider=anthropic; ignored otherwise.
|
||||
bool cache_prompt = 8;
|
||||
}
|
||||
|
||||
message Result {
|
||||
@@ -547,10 +472,6 @@ message TranscriptRequest {
|
||||
float temperature = 8;
|
||||
repeated string timestamp_granularities = 9;
|
||||
bool stream = 10;
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 11;
|
||||
}
|
||||
|
||||
message TranscriptResult {
|
||||
@@ -558,10 +479,6 @@ message TranscriptResult {
|
||||
string text = 2;
|
||||
string language = 3;
|
||||
float duration = 4;
|
||||
// True when the decode ended on the model's end-of-utterance special token
|
||||
// (<EOU>/<EOB>, emitted by cache-aware streaming models such as
|
||||
// parakeet_realtime_eou_120m-v1). The marker itself is stripped from text.
|
||||
bool eou = 5;
|
||||
}
|
||||
|
||||
message TranscriptStreamResponse {
|
||||
@@ -569,34 +486,6 @@ message TranscriptStreamResponse {
|
||||
TranscriptResult final_result = 2;
|
||||
}
|
||||
|
||||
// === AudioTranscriptionLive messages =====================================
|
||||
|
||||
message TranscriptLiveRequest {
|
||||
oneof payload {
|
||||
TranscriptLiveConfig config = 1;
|
||||
TranscriptLiveAudio audio = 2;
|
||||
}
|
||||
}
|
||||
|
||||
message TranscriptLiveConfig {
|
||||
string language = 1; // "" => model default
|
||||
int32 sample_rate = 2; // 0 => 16000; backends may reject others
|
||||
map<string, string> params = 3; // backend-specific tuning
|
||||
}
|
||||
|
||||
message TranscriptLiveAudio {
|
||||
repeated float pcm = 1; // mono PCM in [-1,1] at config.sample_rate
|
||||
}
|
||||
|
||||
message TranscriptLiveResponse {
|
||||
bool ready = 1; // open ack: sent once, before any delta
|
||||
string delta = 2; // newly-finalized text since previous response
|
||||
bool eou = 3; // <EOU> fired during this feed (the user yielded the turn)
|
||||
repeated TranscriptWord words = 4; // words finalized by this feed (stream-relative ns)
|
||||
TranscriptResult final_result = 5; // terminal message only, after the send side closes
|
||||
bool eob = 6; // <EOB> fired: a backchannel ("uh-huh") ended — NOT a turn boundary
|
||||
}
|
||||
|
||||
message TranscriptWord {
|
||||
int64 start = 1;
|
||||
int64 end = 2;
|
||||
@@ -629,10 +518,6 @@ message GenerateImageRequest {
|
||||
|
||||
// Reference images for models that support them (e.g., Flux Kontext)
|
||||
repeated string ref_images = 12;
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 13;
|
||||
}
|
||||
|
||||
message GenerateVideoRequest {
|
||||
@@ -648,14 +533,6 @@ message GenerateVideoRequest {
|
||||
float cfg_scale = 10; // Classifier-free guidance scale
|
||||
int32 step = 11; // Number of inference steps
|
||||
string dst = 12; // Output path for the generated video
|
||||
string audio = 13; // Path to staged audio for audio-conditioned video
|
||||
// Backend-specific per-request generation parameters. Values are strings
|
||||
// and are validated/coerced by the selected backend.
|
||||
map<string, string> params = 14;
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 15;
|
||||
}
|
||||
|
||||
message TTSRequest {
|
||||
@@ -673,26 +550,10 @@ message TTSRequest {
|
||||
// (e.g. Chatterbox exaggeration/cfg_weight/temperature). Values are strings and
|
||||
// coerced by the backend; unset leaves the backend's configured defaults.
|
||||
map<string, string> params = 7;
|
||||
// ModelIdentity is a SEPARATE field from `model` above and carries the
|
||||
// UNTRANSLATED controller-side ModelConfig.Model, so a backend can reject a
|
||||
// request that reached it through a stale distributed route (#10952).
|
||||
//
|
||||
// `model` cannot be reused for this: FileStagingClient.TTS/.TTSStream and the
|
||||
// SoundGeneration path rewrite it into a worker-local absolute path
|
||||
// (core/services/nodes/file_staging_client.go), while the load-time value is
|
||||
// untranslated. In distributed mode - exactly the configuration this guards -
|
||||
// the two already differ, so comparing them would reject valid requests.
|
||||
//
|
||||
// Empty means "no identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 8;
|
||||
}
|
||||
|
||||
message VADRequest {
|
||||
repeated float audio = 1;
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 2;
|
||||
}
|
||||
|
||||
message VADSegment {
|
||||
@@ -724,10 +585,6 @@ message DiarizeRequest {
|
||||
float min_duration_on = 8; // discard segments shorter than this (seconds); 0 = backend default
|
||||
float min_duration_off = 9; // merge gaps shorter than this (seconds); 0 = backend default
|
||||
bool include_text = 10; // when the backend can emit per-segment transcript for free, ask it to populate `text`
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 11;
|
||||
}
|
||||
|
||||
message DiarizeSegment {
|
||||
@@ -762,18 +619,6 @@ message SoundGenerationRequest {
|
||||
optional string language = 14;
|
||||
optional string timesignature = 15;
|
||||
optional bool instrumental = 17;
|
||||
// ModelIdentity is a SEPARATE field from `model` above and carries the
|
||||
// UNTRANSLATED controller-side ModelConfig.Model, so a backend can reject a
|
||||
// request that reached it through a stale distributed route (#10952).
|
||||
//
|
||||
// `model` cannot be reused for this: FileStagingClient.TTS/.TTSStream and the
|
||||
// SoundGeneration path rewrite it into a worker-local absolute path
|
||||
// (core/services/nodes/file_staging_client.go), while the load-time value is
|
||||
// untranslated. In distributed mode - exactly the configuration this guards -
|
||||
// the two already differ, so comparing them would reject valid requests.
|
||||
//
|
||||
// Empty means "no identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 18;
|
||||
}
|
||||
|
||||
message TokenizationResponse {
|
||||
@@ -813,10 +658,6 @@ message DetectOptions {
|
||||
repeated float points = 3; // Point coordinates as [x1, y1, label1, x2, y2, label2, ...] (label: 1=pos, 0=neg)
|
||||
repeated float boxes = 4; // Box coordinates as [x1, y1, x2, y2, ...]
|
||||
float threshold = 5; // Detection confidence threshold
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 6;
|
||||
}
|
||||
|
||||
message Detection {
|
||||
@@ -839,10 +680,6 @@ message SoundDetectionRequest {
|
||||
string src = 1; // audio file path (LocalAI writes the upload to disk)
|
||||
int32 top_k = 2; // number of top tags to return (0 = all classes)
|
||||
float threshold = 3; // optional: drop tags scoring below this
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 4;
|
||||
}
|
||||
|
||||
message SoundClass {
|
||||
@@ -867,10 +704,6 @@ message DepthRequest {
|
||||
bool include_points = 7; // back-project to a 3D point cloud (DualDPT)
|
||||
float points_conf_thresh = 8; // keep points with confidence >= this threshold
|
||||
repeated string exports = 9; // requested exports: "glb", "colmap"
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 10;
|
||||
}
|
||||
|
||||
message DepthResponse {
|
||||
@@ -902,10 +735,6 @@ message FaceVerifyRequest {
|
||||
string img2 = 2; // base64-encoded image
|
||||
float threshold = 3; // cosine-distance threshold; 0 = use backend default
|
||||
bool anti_spoofing = 4; // run MiniFASNet liveness on each image; failed liveness forces verified=false
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 5;
|
||||
}
|
||||
|
||||
message FaceVerifyResponse {
|
||||
@@ -927,10 +756,6 @@ message FaceAnalyzeRequest {
|
||||
string img = 1; // base64-encoded image
|
||||
repeated string actions = 2; // subset of ["age","gender","emotion","race"]; empty = all-supported
|
||||
bool anti_spoofing = 3;
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 4;
|
||||
}
|
||||
|
||||
message FaceAnalysis {
|
||||
@@ -963,10 +788,6 @@ message VoiceVerifyRequest {
|
||||
string audio2 = 2; // path to second audio clip
|
||||
float threshold = 3; // cosine-distance threshold; 0 = use backend default
|
||||
bool anti_spoofing = 4; // reserved for future AASIST bolt-on
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 5;
|
||||
}
|
||||
|
||||
message VoiceVerifyResponse {
|
||||
@@ -981,10 +802,6 @@ message VoiceVerifyResponse {
|
||||
message VoiceAnalyzeRequest {
|
||||
string audio = 1; // path to audio clip
|
||||
repeated string actions = 2; // subset of ["age","gender","emotion"]; empty = all-supported
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 3;
|
||||
}
|
||||
|
||||
message VoiceAnalysis {
|
||||
@@ -1003,10 +820,6 @@ message VoiceAnalyzeResponse {
|
||||
|
||||
message VoiceEmbedRequest {
|
||||
string audio = 1; // path to audio clip
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 2;
|
||||
}
|
||||
|
||||
message VoiceEmbedResponse {
|
||||
@@ -1101,10 +914,6 @@ message AudioTransformRequest {
|
||||
string reference_path = 2; // optional auxiliary; empty => zero-fill
|
||||
string dst = 3; // required, output file path
|
||||
map<string, string> params = 4; // backend-specific tuning
|
||||
// ModelIdentity names the model this request is for; see
|
||||
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
||||
// identity supplied" and backends MUST skip the check.
|
||||
string ModelIdentity = 5;
|
||||
}
|
||||
|
||||
message AudioTransformResult {
|
||||
@@ -1403,3 +1212,4 @@ message ForwardReply {
|
||||
repeated ForwardHeader headers = 2;
|
||||
bytes body_chunk = 3;
|
||||
}
|
||||
|
||||
|
||||
@@ -1,95 +0,0 @@
|
||||
# SPDX-License-Identifier: MIT
|
||||
|
||||
cmake_minimum_required(VERSION 3.20)
|
||||
project(audio-cpp-grpc-server LANGUAGES CXX)
|
||||
|
||||
set(CMAKE_CXX_STANDARD 17)
|
||||
set(CMAKE_CXX_STANDARD_REQUIRED ON)
|
||||
set(CMAKE_CXX_EXTENSIONS OFF)
|
||||
|
||||
set(AUDIO_CPP_DIR "${CMAKE_CURRENT_SOURCE_DIR}/audio.cpp"
|
||||
CACHE PATH "Path to the audio.cpp source tree")
|
||||
set(LOCALAI_BACKEND_PROTO "${CMAKE_CURRENT_SOURCE_DIR}/../../backend.proto"
|
||||
CACHE FILEPATH "Path to the LocalAI backend protocol")
|
||||
|
||||
option(ENGINE_ENABLE_CUDA "Build audio.cpp with CUDA support" OFF)
|
||||
option(ENGINE_ENABLE_VULKAN "Build audio.cpp with Vulkan support" OFF)
|
||||
option(ENGINE_ENABLE_METAL "Build audio.cpp with Metal support" OFF)
|
||||
option(AUDIO_CPP_BUILD_TESTS "Build LocalAI audio.cpp unit tests" OFF)
|
||||
option(AUDIO_CPP_BUILD_GRPC "Build the LocalAI gRPC server" ON)
|
||||
|
||||
find_package(Threads REQUIRED)
|
||||
|
||||
if(NOT EXISTS "${AUDIO_CPP_DIR}/CMakeLists.txt")
|
||||
message(FATAL_ERROR
|
||||
"AUDIO_CPP_DIR does not point to an audio.cpp source tree: ${AUDIO_CPP_DIR}")
|
||||
endif()
|
||||
add_subdirectory("${AUDIO_CPP_DIR}" "${CMAKE_CURRENT_BINARY_DIR}/audio.cpp")
|
||||
|
||||
add_library(localai_audio_cpp_runtime STATIC
|
||||
audio_cpp_runtime.cpp
|
||||
model_config.cpp)
|
||||
target_include_directories(localai_audio_cpp_runtime
|
||||
PUBLIC "${CMAKE_CURRENT_SOURCE_DIR}")
|
||||
target_link_libraries(localai_audio_cpp_runtime PUBLIC engine_runtime)
|
||||
|
||||
if(AUDIO_CPP_BUILD_GRPC)
|
||||
find_package(Protobuf CONFIG REQUIRED)
|
||||
find_package(gRPC CONFIG REQUIRED)
|
||||
find_program(PROTOC_EXECUTABLE NAMES protoc REQUIRED)
|
||||
find_program(GRPC_CPP_PLUGIN_EXECUTABLE NAMES grpc_cpp_plugin REQUIRED)
|
||||
|
||||
get_filename_component(LOCALAI_BACKEND_PROTO_DIR
|
||||
"${LOCALAI_BACKEND_PROTO}" DIRECTORY)
|
||||
set(LOCALAI_PROTO_SOURCES
|
||||
"${CMAKE_CURRENT_BINARY_DIR}/backend.pb.cc"
|
||||
"${CMAKE_CURRENT_BINARY_DIR}/backend.grpc.pb.cc")
|
||||
set(LOCALAI_PROTO_HEADERS
|
||||
"${CMAKE_CURRENT_BINARY_DIR}/backend.pb.h"
|
||||
"${CMAKE_CURRENT_BINARY_DIR}/backend.grpc.pb.h")
|
||||
|
||||
add_custom_command(
|
||||
OUTPUT ${LOCALAI_PROTO_SOURCES} ${LOCALAI_PROTO_HEADERS}
|
||||
COMMAND "${PROTOC_EXECUTABLE}"
|
||||
ARGS
|
||||
--cpp_out "${CMAKE_CURRENT_BINARY_DIR}"
|
||||
--grpc_out "${CMAKE_CURRENT_BINARY_DIR}"
|
||||
-I "${LOCALAI_BACKEND_PROTO_DIR}"
|
||||
--plugin=protoc-gen-grpc="${GRPC_CPP_PLUGIN_EXECUTABLE}"
|
||||
"${LOCALAI_BACKEND_PROTO}"
|
||||
DEPENDS "${LOCALAI_BACKEND_PROTO}"
|
||||
VERBATIM)
|
||||
|
||||
add_library(localai_backend_proto STATIC
|
||||
${LOCALAI_PROTO_SOURCES}
|
||||
${LOCALAI_PROTO_HEADERS})
|
||||
target_include_directories(localai_backend_proto
|
||||
PUBLIC "${CMAKE_CURRENT_BINARY_DIR}")
|
||||
target_link_libraries(localai_backend_proto
|
||||
PUBLIC protobuf::libprotobuf gRPC::grpc++)
|
||||
|
||||
# Task 2 replaces this generated entry point with the LocalAI service.
|
||||
set(AUDIO_CPP_SERVER_PLACEHOLDER
|
||||
"${CMAKE_CURRENT_BINARY_DIR}/audio-cpp-grpc-server-placeholder.cpp")
|
||||
file(GENERATE OUTPUT "${AUDIO_CPP_SERVER_PLACEHOLDER}"
|
||||
CONTENT "int main() { return 0; }\n")
|
||||
|
||||
add_executable(audio-cpp-grpc-server "${AUDIO_CPP_SERVER_PLACEHOLDER}")
|
||||
target_link_libraries(audio-cpp-grpc-server PRIVATE
|
||||
localai_audio_cpp_runtime
|
||||
engine_runtime
|
||||
localai_backend_proto
|
||||
gRPC::grpc++
|
||||
gRPC::grpc++_reflection)
|
||||
endif()
|
||||
|
||||
if(AUDIO_CPP_BUILD_TESTS)
|
||||
enable_testing()
|
||||
add_executable(audio-cpp-runtime-test tests/runtime_tests.cpp)
|
||||
target_link_libraries(audio-cpp-runtime-test PRIVATE
|
||||
localai_audio_cpp_runtime
|
||||
Threads::Threads)
|
||||
add_test(
|
||||
NAME audio-cpp-runtime-test
|
||||
COMMAND audio-cpp-runtime-test "${AUDIO_CPP_DIR}")
|
||||
endif()
|
||||
@@ -1,67 +0,0 @@
|
||||
# SPDX-License-Identifier: MIT
|
||||
|
||||
AUDIO_CPP_VERSION?=f8fb0c19739193adfad0d9e58da99f25eda65256
|
||||
AUDIO_CPP_REPO?=https://github.com/0xShug0/audio.cpp
|
||||
AUDIO_CPP_SRC?=
|
||||
|
||||
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))
|
||||
BUILD_DIR := build
|
||||
|
||||
BUILD_TYPE ?=
|
||||
JOBS ?= $(shell nproc 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 4)
|
||||
UNAME_S := $(shell uname -s)
|
||||
|
||||
CMAKE_ARGS ?= -DCMAKE_BUILD_TYPE=Release
|
||||
CMAKE_ARGS += -DENGINE_ENABLE_CUDA=OFF
|
||||
CMAKE_ARGS += -DENGINE_ENABLE_VULKAN=OFF
|
||||
CMAKE_ARGS += -DENGINE_ENABLE_METAL=OFF
|
||||
|
||||
ifeq ($(BUILD_TYPE),cublas)
|
||||
CMAKE_ARGS += -DENGINE_ENABLE_CUDA=ON
|
||||
else ifeq ($(BUILD_TYPE),vulkan)
|
||||
CMAKE_ARGS += -DENGINE_ENABLE_VULKAN=ON
|
||||
else ifeq ($(UNAME_S),Darwin)
|
||||
CMAKE_ARGS += -DENGINE_ENABLE_METAL=ON
|
||||
endif
|
||||
|
||||
.PHONY: all grpc-server test test-unit clean purge
|
||||
|
||||
all: grpc-server
|
||||
|
||||
audio.cpp:
|
||||
ifneq ($(AUDIO_CPP_SRC),)
|
||||
ln -sfn $(abspath $(AUDIO_CPP_SRC)) audio.cpp
|
||||
else
|
||||
mkdir -p audio.cpp
|
||||
cd audio.cpp && \
|
||||
git init -q && \
|
||||
git remote add origin $(AUDIO_CPP_REPO) && \
|
||||
git fetch --depth 1 origin $(AUDIO_CPP_VERSION) && \
|
||||
git checkout --detach FETCH_HEAD && \
|
||||
git submodule update --init --recursive --depth 1
|
||||
endif
|
||||
|
||||
grpc-server: audio.cpp
|
||||
mkdir -p $(BUILD_DIR)
|
||||
cd $(BUILD_DIR) && cmake $(CMAKE_ARGS) $(CURRENT_MAKEFILE_DIR)
|
||||
cmake --build $(BUILD_DIR) --config Release \
|
||||
--target audio-cpp-grpc-server -j $(JOBS)
|
||||
cp $(BUILD_DIR)/audio-cpp-grpc-server grpc-server
|
||||
|
||||
test:
|
||||
bash tests/build_contract_test.sh
|
||||
|
||||
test-unit: audio.cpp
|
||||
mkdir -p $(BUILD_DIR)-unit
|
||||
cd $(BUILD_DIR)-unit && cmake $(CMAKE_ARGS) \
|
||||
-DAUDIO_CPP_BUILD_TESTS=ON -DAUDIO_CPP_BUILD_GRPC=OFF \
|
||||
$(CURRENT_MAKEFILE_DIR)
|
||||
cmake --build $(BUILD_DIR)-unit --config Release \
|
||||
--target audio-cpp-runtime-test -j $(JOBS)
|
||||
ctest --test-dir $(BUILD_DIR)-unit --output-on-failure
|
||||
|
||||
clean:
|
||||
rm -rf $(BUILD_DIR) $(BUILD_DIR)-unit grpc-server
|
||||
|
||||
purge: clean
|
||||
rm -rf audio.cpp
|
||||
@@ -1,178 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
|
||||
#include "audio_cpp_runtime.h"
|
||||
|
||||
#include <algorithm>
|
||||
#include <stdexcept>
|
||||
#include <string>
|
||||
#include <utility>
|
||||
|
||||
namespace audio_cpp {
|
||||
namespace {
|
||||
|
||||
using engine::runtime::CapabilitySet;
|
||||
using engine::runtime::RunMode;
|
||||
using engine::runtime::TaskSpec;
|
||||
|
||||
const engine::runtime::TaskCapability * find_task_capability(
|
||||
const CapabilitySet & capabilities,
|
||||
const TaskSpec & task) {
|
||||
const auto it = std::find_if(
|
||||
capabilities.supported_tasks.begin(),
|
||||
capabilities.supported_tasks.end(),
|
||||
[&](const engine::runtime::TaskCapability & capability) {
|
||||
return capability.task == task.task;
|
||||
});
|
||||
return it == capabilities.supported_tasks.end() ? nullptr : &*it;
|
||||
}
|
||||
|
||||
void validate_capability(
|
||||
const engine::runtime::ILoadedVoiceModel & model,
|
||||
const AudioCppModelConfig & config) {
|
||||
if (model.metadata().family != config.family) {
|
||||
throw std::runtime_error(
|
||||
"loaded audio.cpp model family '" + model.metadata().family +
|
||||
"' does not match requested family '" + config.family + "'");
|
||||
}
|
||||
|
||||
const auto * capability = find_task_capability(
|
||||
model.capabilities(),
|
||||
config.task);
|
||||
if (capability == nullptr) {
|
||||
throw std::runtime_error(
|
||||
"loaded audio.cpp model does not support requested task '" +
|
||||
std::string(engine::runtime::to_string(config.task.task)) + "'");
|
||||
}
|
||||
if (std::find(
|
||||
capability->modes.begin(),
|
||||
capability->modes.end(),
|
||||
config.task.mode) == capability->modes.end()) {
|
||||
throw std::runtime_error(
|
||||
"loaded audio.cpp model does not support requested mode '" +
|
||||
std::string(engine::runtime::to_string(config.task.mode)) +
|
||||
"' for task '" +
|
||||
std::string(engine::runtime::to_string(config.task.task)) + "'");
|
||||
}
|
||||
}
|
||||
|
||||
void validate_session(
|
||||
const engine::runtime::IVoiceTaskSession & session,
|
||||
const AudioCppModelConfig & config) {
|
||||
if (session.family() != config.family) {
|
||||
throw std::runtime_error("audio.cpp session returned the wrong family");
|
||||
}
|
||||
if (session.task_kind() != config.task.task) {
|
||||
throw std::runtime_error("audio.cpp session returned the wrong task");
|
||||
}
|
||||
if (session.run_mode() != config.task.mode) {
|
||||
throw std::runtime_error("audio.cpp session returned the wrong mode");
|
||||
}
|
||||
if (config.task.mode == RunMode::Offline &&
|
||||
dynamic_cast<const engine::runtime::IOfflineVoiceTaskSession *>(&session) == nullptr) {
|
||||
throw std::runtime_error("audio.cpp session does not implement offline execution");
|
||||
}
|
||||
if (config.task.mode == RunMode::Streaming &&
|
||||
dynamic_cast<const engine::runtime::IStreamingVoiceTaskSession *>(&session) == nullptr) {
|
||||
throw std::runtime_error("audio.cpp session does not implement streaming execution");
|
||||
}
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
AudioCppRuntime::AudioCppRuntime()
|
||||
: registry_(engine::runtime::make_default_registry()) {}
|
||||
|
||||
AudioCppRuntime::AudioCppRuntime(engine::runtime::ModelRegistry registry)
|
||||
: registry_(std::move(registry)) {}
|
||||
|
||||
AudioCppRuntime::~AudioCppRuntime() {
|
||||
free();
|
||||
}
|
||||
|
||||
void AudioCppRuntime::load(const AudioCppModelConfig & config) {
|
||||
std::lock_guard<std::mutex> lock(mutex_);
|
||||
|
||||
auto candidate_model = registry_.load(config.load);
|
||||
if (candidate_model == nullptr) {
|
||||
throw std::runtime_error("audio.cpp registry returned a null model");
|
||||
}
|
||||
validate_capability(*candidate_model, config);
|
||||
|
||||
auto candidate_session = candidate_model->create_task_session(
|
||||
config.task,
|
||||
config.session);
|
||||
if (candidate_session == nullptr) {
|
||||
throw std::runtime_error("audio.cpp model returned a null session");
|
||||
}
|
||||
validate_session(*candidate_session, config);
|
||||
|
||||
free_locked();
|
||||
model_ = std::move(candidate_model);
|
||||
session_ = std::move(candidate_session);
|
||||
}
|
||||
|
||||
void AudioCppRuntime::free() {
|
||||
std::lock_guard<std::mutex> lock(mutex_);
|
||||
free_locked();
|
||||
}
|
||||
|
||||
engine::runtime::TaskResult AudioCppRuntime::run(
|
||||
const engine::runtime::TaskRequest & request) {
|
||||
std::lock_guard<std::mutex> lock(mutex_);
|
||||
auto & session = require_session_locked();
|
||||
session.prepare(engine::runtime::build_preparation_request(request));
|
||||
return require_offline_locked().run(request);
|
||||
}
|
||||
|
||||
void AudioCppRuntime::start_stream(
|
||||
const engine::runtime::TaskRequest & request) {
|
||||
std::lock_guard<std::mutex> lock(mutex_);
|
||||
auto & session = require_session_locked();
|
||||
session.prepare(engine::runtime::build_preparation_request(request));
|
||||
require_streaming_locked().start_stream(request);
|
||||
}
|
||||
|
||||
engine::runtime::StreamEvent AudioCppRuntime::process_audio_chunk(
|
||||
const engine::runtime::AudioChunk & chunk) {
|
||||
std::lock_guard<std::mutex> lock(mutex_);
|
||||
return require_streaming_locked().process_audio_chunk(chunk);
|
||||
}
|
||||
|
||||
engine::runtime::TaskResult AudioCppRuntime::finish_stream() {
|
||||
std::lock_guard<std::mutex> lock(mutex_);
|
||||
return require_streaming_locked().finish_stream();
|
||||
}
|
||||
|
||||
engine::runtime::IVoiceTaskSession & AudioCppRuntime::require_session_locked() {
|
||||
if (session_ == nullptr) {
|
||||
throw std::runtime_error("audio.cpp runtime has no loaded session");
|
||||
}
|
||||
return *session_;
|
||||
}
|
||||
|
||||
engine::runtime::IOfflineVoiceTaskSession &
|
||||
AudioCppRuntime::require_offline_locked() {
|
||||
auto * offline = dynamic_cast<engine::runtime::IOfflineVoiceTaskSession *>(
|
||||
&require_session_locked());
|
||||
if (offline == nullptr) {
|
||||
throw std::runtime_error("loaded audio.cpp session is not offline");
|
||||
}
|
||||
return *offline;
|
||||
}
|
||||
|
||||
engine::runtime::IStreamingVoiceTaskSession &
|
||||
AudioCppRuntime::require_streaming_locked() {
|
||||
auto * streaming = dynamic_cast<engine::runtime::IStreamingVoiceTaskSession *>(
|
||||
&require_session_locked());
|
||||
if (streaming == nullptr) {
|
||||
throw std::runtime_error("loaded audio.cpp session is not streaming");
|
||||
}
|
||||
return *streaming;
|
||||
}
|
||||
|
||||
void AudioCppRuntime::free_locked() {
|
||||
session_.reset();
|
||||
model_.reset();
|
||||
}
|
||||
|
||||
} // namespace audio_cpp
|
||||
@@ -1,46 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
|
||||
#pragma once
|
||||
|
||||
#include "model_config.h"
|
||||
|
||||
#include "engine/framework/runtime/registry.h"
|
||||
#include "engine/framework/runtime/session.h"
|
||||
|
||||
#include <memory>
|
||||
#include <mutex>
|
||||
|
||||
namespace audio_cpp {
|
||||
|
||||
class AudioCppRuntime {
|
||||
public:
|
||||
AudioCppRuntime();
|
||||
explicit AudioCppRuntime(engine::runtime::ModelRegistry registry);
|
||||
~AudioCppRuntime();
|
||||
|
||||
AudioCppRuntime(const AudioCppRuntime &) = delete;
|
||||
AudioCppRuntime & operator=(const AudioCppRuntime &) = delete;
|
||||
|
||||
void load(const AudioCppModelConfig & config);
|
||||
void free();
|
||||
|
||||
engine::runtime::TaskResult run(
|
||||
const engine::runtime::TaskRequest & request);
|
||||
void start_stream(const engine::runtime::TaskRequest & request);
|
||||
engine::runtime::StreamEvent process_audio_chunk(
|
||||
const engine::runtime::AudioChunk & chunk);
|
||||
engine::runtime::TaskResult finish_stream();
|
||||
|
||||
private:
|
||||
engine::runtime::IVoiceTaskSession & require_session_locked();
|
||||
engine::runtime::IOfflineVoiceTaskSession & require_offline_locked();
|
||||
engine::runtime::IStreamingVoiceTaskSession & require_streaming_locked();
|
||||
void free_locked();
|
||||
|
||||
std::mutex mutex_;
|
||||
engine::runtime::ModelRegistry registry_;
|
||||
std::unique_ptr<engine::runtime::ILoadedVoiceModel> model_;
|
||||
std::unique_ptr<engine::runtime::IVoiceTaskSession> session_;
|
||||
};
|
||||
|
||||
} // namespace audio_cpp
|
||||
@@ -1,156 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
|
||||
#include "model_config.h"
|
||||
|
||||
#include <limits>
|
||||
#include <stdexcept>
|
||||
#include <string>
|
||||
|
||||
namespace audio_cpp {
|
||||
namespace {
|
||||
|
||||
using engine::core::BackendType;
|
||||
using engine::runtime::RunMode;
|
||||
using engine::runtime::VoiceTaskKind;
|
||||
|
||||
std::string require_option(
|
||||
const std::unordered_map<std::string, std::string> & options,
|
||||
const std::string & name) {
|
||||
const auto it = options.find(name);
|
||||
if (it == options.end() || it->second.empty()) {
|
||||
throw std::invalid_argument("audio.cpp model config requires " + name);
|
||||
}
|
||||
return it->second;
|
||||
}
|
||||
|
||||
VoiceTaskKind parse_task(const std::string & value) {
|
||||
static const std::unordered_map<std::string, VoiceTaskKind> tasks = {
|
||||
{"vad", VoiceTaskKind::Vad},
|
||||
{"asr", VoiceTaskKind::Asr},
|
||||
{"diarization", VoiceTaskKind::Diarization},
|
||||
{"source-separation", VoiceTaskKind::SourceSeparation},
|
||||
{"audio-generation", VoiceTaskKind::AudioGeneration},
|
||||
{"tts", VoiceTaskKind::Tts},
|
||||
{"voice-cloning", VoiceTaskKind::VoiceCloning},
|
||||
{"voice-conversion", VoiceTaskKind::VoiceConversion},
|
||||
{"speech-to-speech", VoiceTaskKind::SpeechToSpeech},
|
||||
{"alignment", VoiceTaskKind::Alignment},
|
||||
{"voice-design", VoiceTaskKind::VoiceDesign},
|
||||
{"speaker-recognition", VoiceTaskKind::SpeakerRecognition},
|
||||
{"svc", VoiceTaskKind::Svc},
|
||||
};
|
||||
const auto it = tasks.find(value);
|
||||
if (it == tasks.end()) {
|
||||
throw std::invalid_argument("unsupported audio.cpp task: " + value);
|
||||
}
|
||||
return it->second;
|
||||
}
|
||||
|
||||
RunMode parse_mode(const std::string & value) {
|
||||
if (value == "offline") {
|
||||
return RunMode::Offline;
|
||||
}
|
||||
if (value == "streaming") {
|
||||
return RunMode::Streaming;
|
||||
}
|
||||
throw std::invalid_argument("unsupported audio.cpp mode: " + value);
|
||||
}
|
||||
|
||||
BackendType parse_backend(const std::string & value) {
|
||||
if (value == "cpu") {
|
||||
return BackendType::Cpu;
|
||||
}
|
||||
if (value == "cuda") {
|
||||
return BackendType::Cuda;
|
||||
}
|
||||
if (value == "vulkan") {
|
||||
return BackendType::Vulkan;
|
||||
}
|
||||
if (value == "metal") {
|
||||
return BackendType::Metal;
|
||||
}
|
||||
if (value == "best") {
|
||||
return BackendType::BestAvailable;
|
||||
}
|
||||
throw std::invalid_argument("unsupported audio.cpp backend: " + value);
|
||||
}
|
||||
|
||||
int parse_integer(
|
||||
const std::string & name,
|
||||
const std::string & value,
|
||||
int minimum) {
|
||||
size_t parsed = 0;
|
||||
long result = 0;
|
||||
try {
|
||||
result = std::stol(value, &parsed);
|
||||
} catch (const std::exception &) {
|
||||
throw std::invalid_argument("invalid audio.cpp " + name + ": " + value);
|
||||
}
|
||||
if (parsed != value.size() ||
|
||||
result < minimum ||
|
||||
result > std::numeric_limits<int>::max()) {
|
||||
throw std::invalid_argument("invalid audio.cpp " + name + ": " + value);
|
||||
}
|
||||
return static_cast<int>(result);
|
||||
}
|
||||
|
||||
void copy_namespaced_option(
|
||||
const std::string & key,
|
||||
const std::string & prefix,
|
||||
const std::string & value,
|
||||
std::unordered_map<std::string, std::string> & destination) {
|
||||
const std::string name = key.substr(prefix.size());
|
||||
if (name.empty()) {
|
||||
throw std::invalid_argument("audio.cpp option namespace requires a name: " + key);
|
||||
}
|
||||
destination[name] = value;
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
AudioCppModelConfig parse_model_config(
|
||||
const std::filesystem::path & model_path,
|
||||
const std::unordered_map<std::string, std::string> & options) {
|
||||
AudioCppModelConfig config;
|
||||
config.model_path = model_path;
|
||||
config.family = require_option(options, "family");
|
||||
config.task.task = parse_task(require_option(options, "task"));
|
||||
config.task.mode = RunMode::Offline;
|
||||
config.load.model_path = model_path;
|
||||
config.load.family_hint = config.family;
|
||||
config.session.backend.type = BackendType::Cpu;
|
||||
|
||||
if (const auto it = options.find("mode"); it != options.end()) {
|
||||
config.task.mode = parse_mode(it->second);
|
||||
}
|
||||
if (const auto it = options.find("backend"); it != options.end()) {
|
||||
config.session.backend.type = parse_backend(it->second);
|
||||
}
|
||||
if (const auto it = options.find("device"); it != options.end()) {
|
||||
config.session.backend.device = parse_integer("device", it->second, 0);
|
||||
}
|
||||
if (const auto it = options.find("threads"); it != options.end()) {
|
||||
config.session.backend.threads = parse_integer("threads", it->second, 1);
|
||||
}
|
||||
if (const auto it = options.find("model_spec"); it != options.end()) {
|
||||
config.load.model_spec_override = std::filesystem::path(it->second);
|
||||
}
|
||||
if (const auto it = options.find("config_id"); it != options.end()) {
|
||||
config.load.config_id = it->second;
|
||||
}
|
||||
if (const auto it = options.find("weight_id"); it != options.end()) {
|
||||
config.load.weight_id = it->second;
|
||||
}
|
||||
|
||||
for (const auto & [key, value] : options) {
|
||||
if (key.rfind("load.", 0) == 0) {
|
||||
copy_namespaced_option(key, "load.", value, config.load.options);
|
||||
} else if (key.rfind("session.", 0) == 0) {
|
||||
copy_namespaced_option(key, "session.", value, config.session.options);
|
||||
}
|
||||
}
|
||||
|
||||
return config;
|
||||
}
|
||||
|
||||
} // namespace audio_cpp
|
||||
@@ -1,26 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
|
||||
#pragma once
|
||||
|
||||
#include "engine/framework/runtime/model.h"
|
||||
#include "engine/framework/runtime/session.h"
|
||||
|
||||
#include <filesystem>
|
||||
#include <string>
|
||||
#include <unordered_map>
|
||||
|
||||
namespace audio_cpp {
|
||||
|
||||
struct AudioCppModelConfig {
|
||||
std::filesystem::path model_path;
|
||||
std::string family;
|
||||
engine::runtime::TaskSpec task;
|
||||
engine::runtime::ModelLoadRequest load;
|
||||
engine::runtime::SessionOptions session;
|
||||
};
|
||||
|
||||
AudioCppModelConfig parse_model_config(
|
||||
const std::filesystem::path & model_path,
|
||||
const std::unordered_map<std::string, std::string> & options);
|
||||
|
||||
} // namespace audio_cpp
|
||||
@@ -1,164 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# SPDX-License-Identifier: MIT
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
backend_dir="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
tmp_dir="$(mktemp -d)"
|
||||
trap 'rm -rf "${tmp_dir}"' EXIT
|
||||
|
||||
fixture_dir="${tmp_dir}/audio.cpp"
|
||||
prefix_dir="${tmp_dir}/prefix"
|
||||
tools_dir="${tmp_dir}/tools"
|
||||
mkdir -p "${fixture_dir}" "${prefix_dir}/lib/cmake/Protobuf" \
|
||||
"${prefix_dir}/lib/cmake/gRPC" "${tools_dir}"
|
||||
|
||||
cat >"${fixture_dir}/engine_runtime.cpp" <<'EOF'
|
||||
void audio_cpp_build_contract_fixture() {}
|
||||
EOF
|
||||
|
||||
cat >"${fixture_dir}/CMakeLists.txt" <<'EOF'
|
||||
cmake_minimum_required(VERSION 3.20)
|
||||
project(AudioCppBuildContractFixture LANGUAGES CXX)
|
||||
|
||||
option(ENGINE_ENABLE_CUDA "Build with CUDA" OFF)
|
||||
option(ENGINE_ENABLE_VULKAN "Build with Vulkan" OFF)
|
||||
option(ENGINE_ENABLE_METAL "Build with Metal" OFF)
|
||||
|
||||
foreach(wrong_option IN ITEMS
|
||||
AUDIO_CPP_ENABLE_CUDA
|
||||
AUDIO_CPP_ENABLE_VULKAN
|
||||
AUDIO_CPP_ENABLE_METAL
|
||||
AUDIOCPP_ENABLE_CUDA
|
||||
AUDIOCPP_ENABLE_VULKAN
|
||||
AUDIOCPP_ENABLE_METAL
|
||||
ENGINE_CUDA
|
||||
ENGINE_VULKAN
|
||||
ENGINE_METAL
|
||||
GGML_CUDA
|
||||
GGML_VULKAN
|
||||
GGML_METAL)
|
||||
if(DEFINED ${wrong_option})
|
||||
message(FATAL_ERROR "legacy or unsupported audio.cpp option: ${wrong_option}")
|
||||
endif()
|
||||
endforeach()
|
||||
|
||||
add_library(engine_runtime STATIC engine_runtime.cpp)
|
||||
EOF
|
||||
|
||||
cat >"${prefix_dir}/lib/cmake/Protobuf/ProtobufConfig.cmake" <<'EOF'
|
||||
set(Protobuf_FOUND TRUE)
|
||||
set(Protobuf_VERSION 0.0.0)
|
||||
if(NOT TARGET protobuf::libprotobuf)
|
||||
add_library(protobuf::libprotobuf INTERFACE IMPORTED)
|
||||
endif()
|
||||
EOF
|
||||
|
||||
cat >"${prefix_dir}/lib/cmake/gRPC/gRPCConfig.cmake" <<'EOF'
|
||||
set(gRPC_FOUND TRUE)
|
||||
if(NOT TARGET gRPC::grpc++)
|
||||
add_library(gRPC::grpc++ INTERFACE IMPORTED)
|
||||
endif()
|
||||
if(NOT TARGET gRPC::grpc++_reflection)
|
||||
add_library(gRPC::grpc++_reflection INTERFACE IMPORTED)
|
||||
endif()
|
||||
EOF
|
||||
|
||||
cat >"${tools_dir}/protoc" <<'EOF'
|
||||
#!/usr/bin/env sh
|
||||
exit 0
|
||||
EOF
|
||||
cat >"${tools_dir}/grpc_cpp_plugin" <<'EOF'
|
||||
#!/usr/bin/env sh
|
||||
exit 0
|
||||
EOF
|
||||
chmod +x "${tools_dir}/protoc" "${tools_dir}/grpc_cpp_plugin"
|
||||
|
||||
assert_cache_bool() {
|
||||
local cache_file="$1"
|
||||
local name="$2"
|
||||
local expected="$3"
|
||||
grep -q "^${name}:BOOL=${expected}$" "${cache_file}" || {
|
||||
echo "expected ${name}:BOOL=${expected} in ${cache_file}" >&2
|
||||
return 1
|
||||
}
|
||||
}
|
||||
|
||||
configure_case() {
|
||||
local name="$1"
|
||||
local cuda="$2"
|
||||
local vulkan="$3"
|
||||
local metal="$4"
|
||||
local build_dir="${tmp_dir}/build-${name}"
|
||||
|
||||
PATH="${tools_dir}:${PATH}" cmake \
|
||||
-S "${backend_dir}" \
|
||||
-B "${build_dir}" \
|
||||
-DCMAKE_PREFIX_PATH="${prefix_dir}" \
|
||||
-DAUDIO_CPP_DIR="${fixture_dir}" \
|
||||
-DENGINE_ENABLE_CUDA="${cuda}" \
|
||||
-DENGINE_ENABLE_VULKAN="${vulkan}" \
|
||||
-DENGINE_ENABLE_METAL="${metal}" \
|
||||
>/dev/null
|
||||
|
||||
assert_cache_bool "${build_dir}/CMakeCache.txt" ENGINE_ENABLE_CUDA "${cuda}"
|
||||
assert_cache_bool "${build_dir}/CMakeCache.txt" ENGINE_ENABLE_VULKAN "${vulkan}"
|
||||
assert_cache_bool "${build_dir}/CMakeCache.txt" ENGINE_ENABLE_METAL "${metal}"
|
||||
|
||||
grep -q 'engine_runtime' \
|
||||
"${build_dir}/CMakeFiles/audio-cpp-grpc-server.dir/link.txt" || {
|
||||
echo "audio-cpp-grpc-server does not link engine_runtime" >&2
|
||||
return 1
|
||||
}
|
||||
}
|
||||
|
||||
configure_case cpu OFF OFF OFF
|
||||
configure_case cuda ON OFF OFF
|
||||
configure_case vulkan OFF ON OFF
|
||||
configure_case metal OFF OFF ON
|
||||
|
||||
if PATH="${tools_dir}:${PATH}" cmake \
|
||||
-S "${backend_dir}" \
|
||||
-B "${tmp_dir}/build-wrong-option" \
|
||||
-DCMAKE_PREFIX_PATH="${prefix_dir}" \
|
||||
-DAUDIO_CPP_DIR="${fixture_dir}" \
|
||||
-DGGML_CUDA=ON \
|
||||
>/dev/null 2>&1; then
|
||||
echo "strict audio.cpp fixture accepted legacy GGML_CUDA option" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
make_database="${tmp_dir}/make-database"
|
||||
make -C "${backend_dir}" -pn >"${make_database}"
|
||||
audio_cpp_version="$(
|
||||
sed -n 's/^AUDIO_CPP_VERSION = //p' "${make_database}" | head -n 1
|
||||
)"
|
||||
[[ "${audio_cpp_version}" =~ ^[0-9a-f]{40}$ ]] || {
|
||||
echo "AUDIO_CPP_VERSION must be a pinned 40-character commit" >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
fetch_plan="${tmp_dir}/fetch-plan"
|
||||
make -C "${backend_dir}" -Bn audio.cpp >"${fetch_plan}"
|
||||
grep -q 'github.com/0xShug0/audio.cpp' "${fetch_plan}"
|
||||
grep -q "${audio_cpp_version}" "${fetch_plan}"
|
||||
|
||||
cat >"${tools_dir}/uname" <<'EOF'
|
||||
#!/usr/bin/env sh
|
||||
if [ "$#" -eq 1 ] && [ "$1" = "-s" ]; then
|
||||
echo Darwin
|
||||
exit 0
|
||||
fi
|
||||
echo "build contract requires uname -s" >&2
|
||||
exit 64
|
||||
EOF
|
||||
chmod +x "${tools_dir}/uname"
|
||||
|
||||
darwin_plan="${tmp_dir}/darwin-plan"
|
||||
PATH="${tools_dir}:${PATH}" make -C "${backend_dir}" -n \
|
||||
AUDIO_CPP_SRC="${fixture_dir}" grpc-server >"${darwin_plan}"
|
||||
grep -q -- '-DENGINE_ENABLE_CUDA=OFF' "${darwin_plan}"
|
||||
grep -q -- '-DENGINE_ENABLE_VULKAN=OFF' "${darwin_plan}"
|
||||
grep -q -- '-DENGINE_ENABLE_METAL=ON' "${darwin_plan}"
|
||||
|
||||
echo "audio.cpp build contract: PASS"
|
||||
@@ -1,490 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
|
||||
#include "audio_cpp_runtime.h"
|
||||
#include "model_config.h"
|
||||
|
||||
#include "engine/framework/runtime/model.h"
|
||||
#include "engine/framework/runtime/registry.h"
|
||||
#include "engine/framework/runtime/session.h"
|
||||
|
||||
#include <atomic>
|
||||
#include <chrono>
|
||||
#include <condition_variable>
|
||||
#include <exception>
|
||||
#include <filesystem>
|
||||
#include <future>
|
||||
#include <iostream>
|
||||
#include <memory>
|
||||
#include <mutex>
|
||||
#include <stdexcept>
|
||||
#include <string>
|
||||
#include <unordered_map>
|
||||
#include <utility>
|
||||
#include <vector>
|
||||
|
||||
namespace {
|
||||
|
||||
using engine::runtime::AudioChunk;
|
||||
using engine::runtime::CapabilitySet;
|
||||
using engine::runtime::ILoadedVoiceModel;
|
||||
using engine::runtime::IOfflineVoiceTaskSession;
|
||||
using engine::runtime::IStreamingVoiceTaskSession;
|
||||
using engine::runtime::IVoiceModelLoader;
|
||||
using engine::runtime::IVoiceTaskSession;
|
||||
using engine::runtime::ModelInspection;
|
||||
using engine::runtime::ModelLoadRequest;
|
||||
using engine::runtime::ModelMetadata;
|
||||
using engine::runtime::RunMode;
|
||||
using engine::runtime::SessionOptions;
|
||||
using engine::runtime::SessionPreparationRequest;
|
||||
using engine::runtime::StreamEvent;
|
||||
using engine::runtime::TaskCapability;
|
||||
using engine::runtime::TaskRequest;
|
||||
using engine::runtime::TaskResult;
|
||||
using engine::runtime::TaskSpec;
|
||||
using engine::runtime::VoiceTaskKind;
|
||||
|
||||
void require(bool condition, const std::string & message) {
|
||||
if (!condition) {
|
||||
throw std::runtime_error(message);
|
||||
}
|
||||
}
|
||||
|
||||
template <typename Function>
|
||||
void require_throws(Function && function, const std::string & expected) {
|
||||
try {
|
||||
function();
|
||||
} catch (const std::exception & error) {
|
||||
require(
|
||||
std::string(error.what()).find(expected) != std::string::npos,
|
||||
"expected error containing '" + expected + "', got '" + error.what() + "'");
|
||||
return;
|
||||
}
|
||||
throw std::runtime_error("expected exception containing '" + expected + "'");
|
||||
}
|
||||
|
||||
struct SessionGate {
|
||||
std::mutex mutex;
|
||||
std::condition_variable condition;
|
||||
bool first_entered = false;
|
||||
bool release_first = false;
|
||||
std::atomic<int> entries{0};
|
||||
};
|
||||
|
||||
struct FakeState {
|
||||
std::mutex mutex;
|
||||
std::vector<std::string> events;
|
||||
CapabilitySet capabilities;
|
||||
bool fail_load = false;
|
||||
bool fail_session = false;
|
||||
int generation = 0;
|
||||
std::shared_ptr<SessionGate> gate;
|
||||
|
||||
void record(std::string event) {
|
||||
std::lock_guard<std::mutex> lock(mutex);
|
||||
events.push_back(std::move(event));
|
||||
}
|
||||
};
|
||||
|
||||
class FakeSession final
|
||||
: public IOfflineVoiceTaskSession,
|
||||
public IStreamingVoiceTaskSession {
|
||||
public:
|
||||
FakeSession(
|
||||
std::shared_ptr<FakeState> state,
|
||||
int generation,
|
||||
TaskSpec task,
|
||||
SessionOptions options)
|
||||
: state_(std::move(state)),
|
||||
generation_(generation),
|
||||
task_(task),
|
||||
options_(std::move(options)) {}
|
||||
|
||||
~FakeSession() override {
|
||||
state_->record("session-" + std::to_string(generation_) + "-destroyed");
|
||||
}
|
||||
|
||||
std::string family() const override { return "fake-family"; }
|
||||
VoiceTaskKind task_kind() const override { return task_.task; }
|
||||
RunMode run_mode() const override { return task_.mode; }
|
||||
|
||||
void prepare(const SessionPreparationRequest & request) override {
|
||||
prepared_ = request;
|
||||
}
|
||||
|
||||
TaskResult run(const TaskRequest &) override {
|
||||
if (state_->gate != nullptr) {
|
||||
const int entry = ++state_->gate->entries;
|
||||
if (entry == 1) {
|
||||
std::unique_lock<std::mutex> lock(state_->gate->mutex);
|
||||
state_->gate->first_entered = true;
|
||||
state_->gate->condition.notify_all();
|
||||
state_->gate->condition.wait(
|
||||
lock,
|
||||
[&] { return state_->gate->release_first; });
|
||||
}
|
||||
}
|
||||
|
||||
TaskResult result;
|
||||
result.text_output = engine::runtime::Transcript{
|
||||
"generation-" + std::to_string(generation_),
|
||||
"en",
|
||||
};
|
||||
return result;
|
||||
}
|
||||
|
||||
engine::runtime::StreamingPolicy streaming_policy() const override {
|
||||
engine::runtime::StreamingPolicy policy;
|
||||
policy.input = engine::runtime::StreamingInputKind::AudioChunks;
|
||||
policy.output = engine::runtime::StreamingOutputKind::PullEvents;
|
||||
policy.preferred_audio_chunk_samples = 160;
|
||||
return policy;
|
||||
}
|
||||
|
||||
void start_stream(const TaskRequest &) override {
|
||||
streaming_ = true;
|
||||
}
|
||||
|
||||
std::optional<StreamEvent> next_stream_event() override {
|
||||
return std::nullopt;
|
||||
}
|
||||
|
||||
void set_stream_event_sink(engine::runtime::StreamEventCallback sink) override {
|
||||
sink_ = std::move(sink);
|
||||
}
|
||||
|
||||
TaskResult finish_stream() override {
|
||||
streaming_ = false;
|
||||
return run({});
|
||||
}
|
||||
|
||||
void reset() override {
|
||||
streaming_ = false;
|
||||
}
|
||||
|
||||
StreamEvent process_audio_chunk(const AudioChunk & chunk) override {
|
||||
require(streaming_, "stream was not started");
|
||||
StreamEvent event;
|
||||
event.audio_output = engine::runtime::AudioBuffer{
|
||||
chunk.sample_rate,
|
||||
chunk.channels,
|
||||
chunk.samples,
|
||||
};
|
||||
if (sink_) {
|
||||
sink_(event);
|
||||
}
|
||||
return event;
|
||||
}
|
||||
|
||||
TaskResult finalize() override {
|
||||
streaming_ = false;
|
||||
return run({});
|
||||
}
|
||||
|
||||
private:
|
||||
std::shared_ptr<FakeState> state_;
|
||||
int generation_;
|
||||
TaskSpec task_;
|
||||
SessionOptions options_;
|
||||
SessionPreparationRequest prepared_;
|
||||
engine::runtime::StreamEventCallback sink_;
|
||||
bool streaming_ = false;
|
||||
};
|
||||
|
||||
class FakeLoadedModel final : public ILoadedVoiceModel {
|
||||
public:
|
||||
FakeLoadedModel(std::shared_ptr<FakeState> state, int generation)
|
||||
: state_(std::move(state)),
|
||||
generation_(generation) {
|
||||
metadata_.family = "fake-family";
|
||||
metadata_.variant = "complete-fake";
|
||||
metadata_.description = "complete test implementation";
|
||||
metadata_.config_candidates = {"config.json"};
|
||||
metadata_.weight_candidates = {"weights.gguf"};
|
||||
}
|
||||
|
||||
~FakeLoadedModel() override {
|
||||
state_->record("model-" + std::to_string(generation_) + "-destroyed");
|
||||
}
|
||||
|
||||
const ModelMetadata & metadata() const noexcept override {
|
||||
return metadata_;
|
||||
}
|
||||
|
||||
const CapabilitySet & capabilities() const noexcept override {
|
||||
return state_->capabilities;
|
||||
}
|
||||
|
||||
std::unique_ptr<IVoiceTaskSession> create_task_session(
|
||||
const TaskSpec & task,
|
||||
const SessionOptions & options) const override {
|
||||
if (state_->fail_session) {
|
||||
throw std::runtime_error("session creation failed");
|
||||
}
|
||||
return std::make_unique<FakeSession>(state_, generation_, task, options);
|
||||
}
|
||||
|
||||
private:
|
||||
std::shared_ptr<FakeState> state_;
|
||||
int generation_;
|
||||
ModelMetadata metadata_;
|
||||
};
|
||||
|
||||
class FakeLoader final : public IVoiceModelLoader {
|
||||
public:
|
||||
explicit FakeLoader(std::shared_ptr<FakeState> state)
|
||||
: state_(std::move(state)) {}
|
||||
|
||||
std::string family() const override { return "fake-family"; }
|
||||
|
||||
bool can_load(const ModelLoadRequest & request) const override {
|
||||
return request.family_hint == family();
|
||||
}
|
||||
|
||||
ModelInspection inspect(const ModelLoadRequest & request) const override {
|
||||
ModelInspection inspection;
|
||||
inspection.metadata.family = family();
|
||||
inspection.metadata.variant = "complete-fake";
|
||||
inspection.metadata.description = "complete test loader";
|
||||
inspection.metadata.config_candidates = {"config.json"};
|
||||
inspection.metadata.weight_candidates = {"weights.gguf"};
|
||||
inspection.capabilities = state_->capabilities;
|
||||
inspection.model_root = request.model_path;
|
||||
return inspection;
|
||||
}
|
||||
|
||||
std::unique_ptr<ILoadedVoiceModel> load(
|
||||
const ModelLoadRequest &) const override {
|
||||
if (state_->fail_load) {
|
||||
throw std::runtime_error("model load failed");
|
||||
}
|
||||
const int generation = ++state_->generation;
|
||||
return std::make_unique<FakeLoadedModel>(state_, generation);
|
||||
}
|
||||
|
||||
CapabilitySet advertised_capabilities() const override {
|
||||
return state_->capabilities;
|
||||
}
|
||||
|
||||
std::string advertised_instructions_policy() const override {
|
||||
return "explicit";
|
||||
}
|
||||
|
||||
std::vector<std::string> advertised_api_endpoints() const override {
|
||||
return {"/v1/audio/transcriptions"};
|
||||
}
|
||||
|
||||
private:
|
||||
std::shared_ptr<FakeState> state_;
|
||||
};
|
||||
|
||||
audio_cpp::AudioCppModelConfig offline_asr_config(
|
||||
const std::filesystem::path & model_path) {
|
||||
return audio_cpp::parse_model_config(
|
||||
model_path,
|
||||
{
|
||||
{"family", "fake-family"},
|
||||
{"task", "asr"},
|
||||
{"mode", "offline"},
|
||||
{"backend", "cpu"},
|
||||
{"device", "2"},
|
||||
{"threads", "3"},
|
||||
{"load.cache", "memory"},
|
||||
{"session.language", "en"},
|
||||
});
|
||||
}
|
||||
|
||||
std::unique_ptr<audio_cpp::AudioCppRuntime> make_runtime(
|
||||
const std::shared_ptr<FakeState> & state) {
|
||||
engine::runtime::ModelRegistry registry;
|
||||
registry.register_loader(std::make_shared<FakeLoader>(state));
|
||||
return std::make_unique<audio_cpp::AudioCppRuntime>(std::move(registry));
|
||||
}
|
||||
|
||||
void test_model_config(const std::filesystem::path & model_path) {
|
||||
const auto config = offline_asr_config(model_path);
|
||||
require(config.model_path == model_path, "model path was not preserved");
|
||||
require(config.family == "fake-family", "family was not parsed");
|
||||
require(config.task.task == VoiceTaskKind::Asr, "task was not parsed");
|
||||
require(config.task.mode == RunMode::Offline, "mode was not parsed");
|
||||
require(
|
||||
config.session.backend.type == engine::core::BackendType::Cpu,
|
||||
"backend was not parsed");
|
||||
require(config.session.backend.device == 2, "device was not parsed");
|
||||
require(config.session.backend.threads == 3, "threads were not parsed");
|
||||
require(config.load.options.at("cache") == "memory", "load option prefix was not stripped");
|
||||
require(
|
||||
config.session.options.at("language") == "en",
|
||||
"session option prefix was not stripped");
|
||||
|
||||
require_throws(
|
||||
[&] { audio_cpp::parse_model_config(model_path, {{"task", "asr"}}); },
|
||||
"family");
|
||||
require_throws(
|
||||
[&] { audio_cpp::parse_model_config(model_path, {{"family", "fake-family"}}); },
|
||||
"task");
|
||||
require_throws(
|
||||
[&] {
|
||||
audio_cpp::parse_model_config(
|
||||
model_path,
|
||||
{{"family", "fake-family"}, {"task", "asr"}, {"mode", "batch"}});
|
||||
},
|
||||
"mode");
|
||||
require_throws(
|
||||
[&] {
|
||||
audio_cpp::parse_model_config(
|
||||
model_path,
|
||||
{{"family", "fake-family"}, {"task", "asr"}, {"backend", "tpu"}});
|
||||
},
|
||||
"backend");
|
||||
}
|
||||
|
||||
void test_capability_validation(const std::filesystem::path & model_path) {
|
||||
auto state = std::make_shared<FakeState>();
|
||||
state->capabilities.supported_tasks = {
|
||||
{VoiceTaskKind::Tts, {RunMode::Offline}},
|
||||
};
|
||||
auto runtime = make_runtime(state);
|
||||
|
||||
require_throws(
|
||||
[&] { runtime->load(offline_asr_config(model_path)); },
|
||||
"task");
|
||||
|
||||
state->capabilities.supported_tasks = {
|
||||
{VoiceTaskKind::Asr, {RunMode::Streaming}},
|
||||
};
|
||||
require_throws(
|
||||
[&] { runtime->load(offline_asr_config(model_path)); },
|
||||
"mode");
|
||||
}
|
||||
|
||||
void test_atomic_replacement(const std::filesystem::path & model_path) {
|
||||
auto state = std::make_shared<FakeState>();
|
||||
state->capabilities.supported_tasks = {
|
||||
{VoiceTaskKind::Asr, {RunMode::Offline}},
|
||||
};
|
||||
auto runtime = make_runtime(state);
|
||||
runtime->load(offline_asr_config(model_path));
|
||||
|
||||
state->fail_load = true;
|
||||
require_throws(
|
||||
[&] { runtime->load(offline_asr_config(model_path)); },
|
||||
"model load failed");
|
||||
require(
|
||||
runtime->run({}).text_output->text == "generation-1",
|
||||
"old model was not retained after load failure");
|
||||
|
||||
state->fail_load = false;
|
||||
state->fail_session = true;
|
||||
require_throws(
|
||||
[&] { runtime->load(offline_asr_config(model_path)); },
|
||||
"session creation failed");
|
||||
require(
|
||||
runtime->run({}).text_output->text == "generation-1",
|
||||
"old model was not retained after session creation failure");
|
||||
}
|
||||
|
||||
void test_teardown_order(const std::filesystem::path & model_path) {
|
||||
auto state = std::make_shared<FakeState>();
|
||||
state->capabilities.supported_tasks = {
|
||||
{VoiceTaskKind::Asr, {RunMode::Offline}},
|
||||
};
|
||||
auto runtime = make_runtime(state);
|
||||
runtime->load(offline_asr_config(model_path));
|
||||
runtime->free();
|
||||
|
||||
std::lock_guard<std::mutex> lock(state->mutex);
|
||||
require(state->events.size() == 2, "expected one session and one model teardown");
|
||||
require(
|
||||
state->events[0] == "session-1-destroyed",
|
||||
"session was not destroyed before model");
|
||||
require(
|
||||
state->events[1] == "model-1-destroyed",
|
||||
"model teardown event was not second");
|
||||
}
|
||||
|
||||
void test_runtime_serializes_calls(const std::filesystem::path & model_path) {
|
||||
auto state = std::make_shared<FakeState>();
|
||||
state->capabilities.supported_tasks = {
|
||||
{VoiceTaskKind::Asr, {RunMode::Offline}},
|
||||
};
|
||||
state->gate = std::make_shared<SessionGate>();
|
||||
auto runtime = make_runtime(state);
|
||||
runtime->load(offline_asr_config(model_path));
|
||||
|
||||
auto first = std::async(std::launch::async, [&] { return runtime->run({}); });
|
||||
{
|
||||
std::unique_lock<std::mutex> lock(state->gate->mutex);
|
||||
state->gate->condition.wait(
|
||||
lock,
|
||||
[&] { return state->gate->first_entered; });
|
||||
}
|
||||
|
||||
std::promise<void> release_second;
|
||||
std::shared_future<void> second_barrier = release_second.get_future().share();
|
||||
std::promise<void> second_attempted_promise;
|
||||
auto second_attempted = second_attempted_promise.get_future();
|
||||
auto second = std::async(std::launch::async, [&] {
|
||||
second_barrier.wait();
|
||||
second_attempted_promise.set_value();
|
||||
return runtime->run({});
|
||||
});
|
||||
|
||||
release_second.set_value();
|
||||
second_attempted.wait();
|
||||
require(
|
||||
second.wait_for(std::chrono::milliseconds(50)) == std::future_status::timeout,
|
||||
"second call completed while first call held the runtime");
|
||||
require(
|
||||
state->gate->entries.load() == 1,
|
||||
"second call entered the upstream session concurrently");
|
||||
|
||||
{
|
||||
std::lock_guard<std::mutex> lock(state->gate->mutex);
|
||||
state->gate->release_first = true;
|
||||
}
|
||||
state->gate->condition.notify_all();
|
||||
first.get();
|
||||
second.get();
|
||||
require(state->gate->entries.load() == 2, "second call never reached the session");
|
||||
}
|
||||
|
||||
void test_streaming_surface(const std::filesystem::path & model_path) {
|
||||
auto state = std::make_shared<FakeState>();
|
||||
state->capabilities.supported_tasks = {
|
||||
{VoiceTaskKind::Asr, {RunMode::Streaming}},
|
||||
};
|
||||
auto runtime = make_runtime(state);
|
||||
auto config = offline_asr_config(model_path);
|
||||
config.task.mode = RunMode::Streaming;
|
||||
runtime->load(config);
|
||||
runtime->start_stream({});
|
||||
const auto event = runtime->process_audio_chunk({16000, 1, 0, {0.25f}});
|
||||
require(event.audio_output.has_value(), "streaming chunk result was lost");
|
||||
require(
|
||||
event.audio_output->samples == std::vector<float>{0.25f},
|
||||
"streaming chunk samples changed");
|
||||
require(
|
||||
runtime->finish_stream().text_output->text == "generation-1",
|
||||
"streaming final result was lost");
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
int main(int argc, char ** argv) {
|
||||
try {
|
||||
require(argc == 2, "runtime_test requires an existing model path argument");
|
||||
const std::filesystem::path model_path(argv[1]);
|
||||
test_model_config(model_path);
|
||||
test_capability_validation(model_path);
|
||||
test_atomic_replacement(model_path);
|
||||
test_teardown_order(model_path);
|
||||
test_runtime_serializes_calls(model_path);
|
||||
test_streaming_surface(model_path);
|
||||
std::cout << "audio.cpp runtime unit tests: PASS\n";
|
||||
return 0;
|
||||
} catch (const std::exception & error) {
|
||||
std::cerr << "audio.cpp runtime unit tests: FAIL: " << error.what() << '\n';
|
||||
return 1;
|
||||
}
|
||||
}
|
||||
@@ -1,107 +0,0 @@
|
||||
|
||||
# Pinned to the HEAD of the `prism` branch on https://github.com/PrismML-Eng/llama.cpp.
|
||||
# Auto-bumped nightly by .github/workflows/bump_deps.yaml.
|
||||
BONSAI_VERSION?=7529fdaaf99ffdc5ca71ace9c7409a56b27ad92f
|
||||
LLAMA_REPO?=https://github.com/PrismML-Eng/llama.cpp
|
||||
|
||||
CMAKE_ARGS?=
|
||||
BUILD_TYPE?=
|
||||
NATIVE?=false
|
||||
ONEAPI_VARS?=/opt/intel/oneapi/setvars.sh
|
||||
TARGET?=--target grpc-server
|
||||
JOBS?=$(shell nproc 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 1)
|
||||
ARCH?=$(shell uname -m)
|
||||
|
||||
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))
|
||||
LLAMA_CPP_DIR := $(CURRENT_MAKEFILE_DIR)/../llama-cpp
|
||||
|
||||
GREEN := \033[0;32m
|
||||
RESET := \033[0m
|
||||
|
||||
# bonsai is a llama.cpp fork (PrismML) adding the Q1_0 (1-bit) and Q2_0 (ternary)
|
||||
# weight-quantization kernels that the Bonsai / Ternary-Bonsai models ship in. Rather
|
||||
# than duplicating grpc-server.cpp / CMakeLists.txt / prepare.sh we reuse the ones in
|
||||
# backend/cpp/llama-cpp, and only swap which repo+sha the fetch step pulls. Each flavor
|
||||
# target copies ../llama-cpp into a sibling ../bonsai-<flavor>-build directory, then
|
||||
# invokes llama-cpp's own build with LLAMA_REPO/LLAMA_VERSION overridden to point at the
|
||||
# fork.
|
||||
#
|
||||
# The Q1_0/Q2_0 additions are model *weight* types decoded inside libllama, transparent
|
||||
# to the reused gRPC server, so (unlike turboquant's KV-cache types) no grpc-server.cpp
|
||||
# allow-list patch is needed. The fork branched from upstream before a few API changes
|
||||
# the shared grpc-server.cpp depends on; those are carried as patch files under
|
||||
# backend/cpp/bonsai/patches/ and applied to the cloned fork by apply-patches.sh.
|
||||
PATCHES_DIR := $(CURRENT_MAKEFILE_DIR)/patches
|
||||
|
||||
define bonsai-build
|
||||
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build
|
||||
cp -rf $(LLAMA_CPP_DIR) $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build
|
||||
# Drop patches vendored for upstream llama.cpp: the fork tree diverges, so
|
||||
# they reject there. Fork-specific patches live in backend/cpp/bonsai/patches/
|
||||
# and are applied by apply-patches.sh below.
|
||||
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build/patches
|
||||
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build purge
|
||||
bash $(LLAMA_CPP_DIR)/disable-score-task.sh $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build/grpc-server.cpp
|
||||
$(info $(GREEN)I bonsai build info:$(1)$(RESET))
|
||||
LLAMA_REPO=$(LLAMA_REPO) LLAMA_VERSION=$(BONSAI_VERSION) \
|
||||
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build llama.cpp
|
||||
bash $(CURRENT_MAKEFILE_DIR)/apply-patches.sh $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build/llama.cpp $(PATCHES_DIR)
|
||||
CMAKE_ARGS="$(CMAKE_ARGS) $(2)" TARGET="$(3)" \
|
||||
LLAMA_REPO=$(LLAMA_REPO) LLAMA_VERSION=$(BONSAI_VERSION) \
|
||||
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build grpc-server
|
||||
cp -rfv $(CURRENT_MAKEFILE_DIR)/../bonsai-$(1)-build/grpc-server bonsai-$(1)
|
||||
endef
|
||||
|
||||
bonsai-avx2:
|
||||
$(call bonsai-build,avx2,-DGGML_AVX=on -DGGML_AVX2=on -DGGML_AVX512=off -DGGML_FMA=on -DGGML_F16C=on,--target grpc-server)
|
||||
|
||||
bonsai-avx512:
|
||||
$(call bonsai-build,avx512,-DGGML_AVX=on -DGGML_AVX2=off -DGGML_AVX512=on -DGGML_FMA=on -DGGML_F16C=on,--target grpc-server)
|
||||
|
||||
bonsai-avx:
|
||||
$(call bonsai-build,avx,-DGGML_AVX=on -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server)
|
||||
|
||||
bonsai-fallback:
|
||||
$(call bonsai-build,fallback,-DGGML_AVX=off -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server)
|
||||
|
||||
# Single-build CPU backend via ggml CPU_ALL_VARIANTS (mirrors llama-cpp-cpu-all).
|
||||
# bonsai reuses backend/cpp/llama-cpp's CMakeLists.txt (hw_grpc_proto STATIC) and
|
||||
# Makefile (SHARED_LIBS make-var + EXTRA_CMAKE_ARGS), so this passes the same overrides
|
||||
# through to the copied build: SHARED_LIBS=ON, the DL flags, and --target ggml (which
|
||||
# pulls in the per-microarch libggml-cpu-*.so via ggml's add_dependencies). The .so set
|
||||
# is collected for package.sh to bundle into package/lib.
|
||||
bonsai-cpu-all:
|
||||
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build
|
||||
cp -rf $(LLAMA_CPP_DIR) $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build
|
||||
# Drop patches vendored for upstream llama.cpp: the fork tree diverges, so
|
||||
# they reject there. Fork-specific patches live in backend/cpp/bonsai/patches/
|
||||
# and are applied by apply-patches.sh below.
|
||||
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/patches
|
||||
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build purge
|
||||
bash $(LLAMA_CPP_DIR)/disable-score-task.sh $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/grpc-server.cpp
|
||||
$(info $(GREEN)I bonsai build info:cpu-all-variants$(RESET))
|
||||
LLAMA_REPO=$(LLAMA_REPO) LLAMA_VERSION=$(BONSAI_VERSION) \
|
||||
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build llama.cpp
|
||||
bash $(CURRENT_MAKEFILE_DIR)/apply-patches.sh $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/llama.cpp $(PATCHES_DIR)
|
||||
SHARED_LIBS=ON EXTRA_CMAKE_ARGS="-DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON" TARGET="--target grpc-server --target ggml" \
|
||||
LLAMA_REPO=$(LLAMA_REPO) LLAMA_VERSION=$(BONSAI_VERSION) \
|
||||
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build grpc-server
|
||||
cp -rfv $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/grpc-server bonsai-cpu-all
|
||||
rm -rf ggml-shared-libs && mkdir -p ggml-shared-libs
|
||||
find $(CURRENT_MAKEFILE_DIR)/../bonsai-cpu-all-build/llama.cpp/build \( -name '*.so*' -o -name '*.dylib' \) -exec cp -av {} ggml-shared-libs/ \;
|
||||
@echo "Collected ggml shared backends:" && ls -la ggml-shared-libs/
|
||||
|
||||
bonsai-grpc:
|
||||
$(call bonsai-build,grpc,-DGGML_RPC=ON -DGGML_AVX=off -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server --target rpc-server)
|
||||
|
||||
bonsai-rpc-server: bonsai-grpc
|
||||
cp -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-grpc-build/llama.cpp/build/bin/rpc-server bonsai-rpc-server
|
||||
|
||||
package:
|
||||
bash package.sh
|
||||
|
||||
purge:
|
||||
rm -rf $(CURRENT_MAKEFILE_DIR)/../bonsai-*-build
|
||||
rm -rf bonsai-* package
|
||||
|
||||
clean: purge
|
||||
@@ -1,48 +0,0 @@
|
||||
#!/bin/bash
|
||||
# Apply the bonsai patch series to a cloned PrismML llama.cpp (prism branch) checkout.
|
||||
#
|
||||
# The prism fork branched from upstream llama.cpp before a number of API changes that the
|
||||
# shared backend/cpp/llama-cpp/grpc-server.cpp depends on. We carry those upstream commits
|
||||
# as patch files under backend/cpp/bonsai/patches/ and apply them here so the reused
|
||||
# grpc-server source compiles against the fork unmodified.
|
||||
#
|
||||
# Drop the corresponding patch from patches/ whenever the fork catches up with upstream —
|
||||
# the build will fail fast if a patch stops applying, which is the signal to retire it.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
if [[ $# -ne 2 ]]; then
|
||||
echo "usage: $0 <llama.cpp-src-dir> <patches-dir>" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
SRC_DIR=$1
|
||||
PATCHES_DIR=$2
|
||||
|
||||
if [[ ! -d "$SRC_DIR" ]]; then
|
||||
echo "source dir does not exist: $SRC_DIR" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
if [[ ! -d "$PATCHES_DIR" ]]; then
|
||||
echo "no patches dir at $PATCHES_DIR, nothing to apply"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
shopt -s nullglob
|
||||
patches=("$PATCHES_DIR"/*.patch)
|
||||
shopt -u nullglob
|
||||
|
||||
if [[ ${#patches[@]} -eq 0 ]]; then
|
||||
echo "no .patch files in $PATCHES_DIR, nothing to apply"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
cd "$SRC_DIR"
|
||||
|
||||
for patch in "${patches[@]}"; do
|
||||
echo "==> applying $patch"
|
||||
git apply --verbose "$patch"
|
||||
done
|
||||
|
||||
echo "all bonsai patches applied successfully"
|
||||
@@ -1,19 +0,0 @@
|
||||
# bonsai fork skew patches
|
||||
|
||||
The `bonsai` backend reuses `backend/cpp/llama-cpp/grpc-server.cpp` (written against
|
||||
LocalAI's pinned *upstream* llama.cpp) but compiles it against the PrismML `prism` fork,
|
||||
which branched from upstream some commits earlier. Any upstream API change that the shared
|
||||
gRPC server depends on, but that the fork does not yet carry, is back-ported here as a
|
||||
`*.patch` file and applied to the cloned fork checkout by `../apply-patches.sh`.
|
||||
|
||||
CI treats both this directory and `backend/cpp/llama-cpp/` as Bonsai inputs, since
|
||||
the wrapper copies and builds the shared llama.cpp backend sources.
|
||||
|
||||
Rules:
|
||||
|
||||
- One upstream commit (or minimal hunk) per patch, named `NNNN-short-description.patch`.
|
||||
- Patches are applied with `git apply` from the fork's checkout root.
|
||||
- `apply-patches.sh` fails fast if a patch stops applying cleanly — that is the signal the
|
||||
fork has caught up (or diverged), so re-cut or drop the patch.
|
||||
- Keep this set as small as possible; the long-term fix is the fork rebasing onto a newer
|
||||
upstream (or Q1_0/Q2_0 landing in mainline llama.cpp, retiring this backend entirely).
|
||||
@@ -1,56 +0,0 @@
|
||||
#!/bin/bash
|
||||
set -ex
|
||||
|
||||
# Get the absolute current dir where the script is located
|
||||
CURDIR=$(dirname "$(realpath "$0")")
|
||||
|
||||
cd /
|
||||
|
||||
echo "CPU info:"
|
||||
grep -e "model\sname" /proc/cpuinfo | head -1
|
||||
grep -e "flags" /proc/cpuinfo | head -1
|
||||
|
||||
BINARY=bonsai-fallback
|
||||
|
||||
# x86/arm64 ship a single bonsai-cpu-all built with ggml CPU_ALL_VARIANTS: ggml's
|
||||
# backend registry dlopens the best libggml-cpu-*.so for this host, so no shell-side
|
||||
# probing. ROCm ships only bonsai-fallback, so fall back to it when cpu-all is absent.
|
||||
if [ -e "$CURDIR"/bonsai-cpu-all ]; then
|
||||
BINARY=bonsai-cpu-all
|
||||
fi
|
||||
|
||||
if [ -n "$LLAMACPP_GRPC_SERVERS" ]; then
|
||||
if [ -e "$CURDIR"/bonsai-grpc ]; then
|
||||
BINARY=bonsai-grpc
|
||||
fi
|
||||
fi
|
||||
|
||||
# Extend ld library path with the dir where this script is located/lib
|
||||
if [ "$(uname)" == "Darwin" ]; then
|
||||
export DYLD_LIBRARY_PATH="$CURDIR"/lib:$DYLD_LIBRARY_PATH
|
||||
else
|
||||
export LD_LIBRARY_PATH="$CURDIR"/lib:$LD_LIBRARY_PATH
|
||||
# Tell rocBLAS where to find TensileLibrary data (GPU kernel tuning files)
|
||||
if [ -d "$CURDIR/lib/rocblas/library" ]; then
|
||||
export ROCBLAS_TENSILE_LIBPATH="$CURDIR"/lib/rocblas/library
|
||||
fi
|
||||
# Same for hipBLASLt (rocblaslt): the bundled libhipblaslt.so resolves its
|
||||
# TensileLibrary_lazy_gfx*.dat kernel data relative to itself, so point it at
|
||||
# the bundled data or it falls back to slow generic kernels (issue #10660).
|
||||
if [ -d "$CURDIR/lib/hipblaslt/library" ]; then
|
||||
export HIPBLASLT_TENSILE_LIBPATH="$CURDIR"/lib/hipblaslt/library
|
||||
fi
|
||||
fi
|
||||
|
||||
# If there is a lib/ld.so, use it
|
||||
if [ -f "$CURDIR"/lib/ld.so ]; then
|
||||
echo "Using lib/ld.so"
|
||||
echo "Using binary: $BINARY"
|
||||
exec "$CURDIR"/lib/ld.so "$CURDIR"/$BINARY "$@"
|
||||
fi
|
||||
|
||||
echo "Using binary: $BINARY"
|
||||
exec "$CURDIR"/$BINARY "$@"
|
||||
|
||||
# We should never reach this point, however just in case we do, run fallback
|
||||
exec "$CURDIR"/bonsai-fallback "$@"
|
||||
@@ -76,13 +76,12 @@ elseif(DS4_GPU STREQUAL "cpu")
|
||||
set(DS4_OBJS "${DS4_DIR}/ds4_cpu.o")
|
||||
endif()
|
||||
|
||||
# Upstream splits distributed inference, tensor-parallel transport, the SSD
|
||||
# expert cache, and layer placement into GPU-agnostic translation units. Link
|
||||
# them regardless of DS4_GPU.
|
||||
# ds4.c now references ds4_distributed.c (distributed inference) and ds4_ssd.c
|
||||
# (SSD expert-cache), each split into its own translation unit upstream. Both
|
||||
# are GPU-agnostic objects shared by every GPU mode, so link them in regardless
|
||||
# of DS4_GPU.
|
||||
list(APPEND DS4_OBJS "${DS4_DIR}/ds4_distributed.o")
|
||||
list(APPEND DS4_OBJS "${DS4_DIR}/ds4_tp.o")
|
||||
list(APPEND DS4_OBJS "${DS4_DIR}/ds4_ssd.o")
|
||||
list(APPEND DS4_OBJS "${DS4_DIR}/ds4_layer_pack.o")
|
||||
|
||||
add_executable(${TARGET}
|
||||
grpc-server.cpp
|
||||
|
||||
@@ -1,10 +1,10 @@
|
||||
# ds4 backend Makefile.
|
||||
#
|
||||
# Upstream pin lives below as DS4_VERSION?=54b36ed9ba42da31b24f2d1a5feb075c2475dbb1
|
||||
# Upstream pin lives below as DS4_VERSION?=80ebbc396aee40eedc1d829222f3362d10fa4c6c
|
||||
# (.github/bump_deps.sh) can find and update it - matches the
|
||||
# llama-cpp / ik-llama-cpp / turboquant convention.
|
||||
|
||||
DS4_VERSION?=54b36ed9ba42da31b24f2d1a5feb075c2475dbb1
|
||||
DS4_VERSION?=80ebbc396aee40eedc1d829222f3362d10fa4c6c
|
||||
DS4_REPO?=https://github.com/antirez/ds4
|
||||
|
||||
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))
|
||||
@@ -18,19 +18,20 @@ UNAME_S := $(shell uname -s)
|
||||
|
||||
CMAKE_ARGS ?= -DCMAKE_BUILD_TYPE=Release
|
||||
|
||||
# Upstream splits distributed inference, tensor-parallel transport, the SSD
|
||||
# expert cache, and layer placement into GPU-agnostic translation units. They
|
||||
# are shared by every GPU mode, so append them unconditionally below.
|
||||
# ds4_distributed.o and ds4_ssd.o are GPU-agnostic translation units that
|
||||
# ds4.c/ds4_cpu.o now reference (upstream split distributed inference and the
|
||||
# SSD expert-cache into their own .c files). Both objects are shared by every
|
||||
# GPU mode, so they are appended unconditionally below.
|
||||
ifeq ($(BUILD_TYPE),cublas)
|
||||
CMAKE_ARGS += -DDS4_GPU=cuda
|
||||
DS4_OBJ_TARGET := ds4.o ds4_cuda.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
|
||||
DS4_OBJ_TARGET := ds4.o ds4_cuda.o ds4_distributed.o ds4_ssd.o
|
||||
else ifeq ($(UNAME_S),Darwin)
|
||||
CMAKE_ARGS += -DDS4_GPU=metal
|
||||
DS4_OBJ_TARGET := ds4.o ds4_metal.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
|
||||
DS4_OBJ_TARGET := ds4.o ds4_metal.o ds4_distributed.o ds4_ssd.o
|
||||
else
|
||||
# CPU reference path (Linux only - macOS CPU path is broken by VM bug per ds4 README).
|
||||
CMAKE_ARGS += -DDS4_GPU=cpu
|
||||
DS4_OBJ_TARGET := ds4_cpu.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
|
||||
DS4_OBJ_TARGET := ds4_cpu.o ds4_distributed.o ds4_ssd.o
|
||||
endif
|
||||
|
||||
ifneq ($(NATIVE),true)
|
||||
@@ -55,11 +56,11 @@ ds4:
|
||||
# the right per-platform compile flags (Objective-C/Metal on Darwin, nvcc on Linux+CUDA).
|
||||
ds4/ds4.o: ds4
|
||||
ifeq ($(BUILD_TYPE),cublas)
|
||||
+$(MAKE) -C ds4 ds4.o ds4_cuda.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
|
||||
+$(MAKE) -C ds4 ds4.o ds4_cuda.o ds4_distributed.o ds4_ssd.o
|
||||
else ifeq ($(UNAME_S),Darwin)
|
||||
+$(MAKE) -C ds4 ds4.o ds4_metal.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
|
||||
+$(MAKE) -C ds4 ds4.o ds4_metal.o ds4_distributed.o ds4_ssd.o
|
||||
else
|
||||
+$(MAKE) -C ds4 ds4_cpu.o ds4_distributed.o ds4_tp.o ds4_ssd.o ds4_layer_pack.o
|
||||
+$(MAKE) -C ds4 ds4_cpu.o ds4_distributed.o ds4_ssd.o
|
||||
endif
|
||||
|
||||
grpc-server: ds4/ds4.o
|
||||
|
||||
@@ -51,11 +51,6 @@ namespace {
|
||||
|
||||
// Global state - ds4 is single-engine-per-process by design.
|
||||
std::mutex g_engine_mu;
|
||||
// The ModelOptions.Model this process loaded, compared against
|
||||
// PredictOptions.ModelIdentity so a request that arrived through a stale
|
||||
// distributed route is rejected rather than answered from the wrong model
|
||||
// (#10952). Guarded by g_engine_mu like the rest of the engine state.
|
||||
std::string g_loaded_model_identity;
|
||||
ds4_engine *g_engine = nullptr;
|
||||
ds4_session *g_session = nullptr;
|
||||
int g_ctx_size = 32768;
|
||||
@@ -567,24 +562,6 @@ static void build_prompt(ds4_engine *engine, const backend::PredictOptions *requ
|
||||
ds4_chat_append_assistant_prefix(engine, out, think);
|
||||
}
|
||||
|
||||
// check_model_identity mirrors pkg/grpc/server.go and
|
||||
// backend/python/common/model_identity.py. Either side empty means "skip": the
|
||||
// request side is empty for a controller that predates the field, the loaded
|
||||
// side when such a controller performed the load. A false rejection is worse
|
||||
// than the miss it prevents. Callers must already hold g_engine_mu.
|
||||
static GStatus check_model_identity(const backend::PredictOptions *request) {
|
||||
if (request == nullptr || request->modelidentity().empty()) return GStatus::OK;
|
||||
if (g_loaded_model_identity.empty() ||
|
||||
g_loaded_model_identity == request->modelidentity()) {
|
||||
return GStatus::OK;
|
||||
}
|
||||
// NOT_FOUND plus this exact sentinel is the cross-language contract the
|
||||
// router matches on (grpcerrors.ModelMismatchSentinel).
|
||||
return GStatus(StatusCode::NOT_FOUND,
|
||||
"ds4: model identity mismatch: loaded \"" + g_loaded_model_identity +
|
||||
"\", requested \"" + request->modelidentity() + "\"");
|
||||
}
|
||||
|
||||
class DS4Backend final : public backend::Backend::Service {
|
||||
public:
|
||||
GStatus Health(ServerContext *, const backend::HealthMessage *,
|
||||
@@ -739,7 +716,6 @@ public:
|
||||
}
|
||||
|
||||
result->set_success(true);
|
||||
g_loaded_model_identity = request->model();
|
||||
result->set_message("loaded " + model_path);
|
||||
return GStatus::OK;
|
||||
}
|
||||
@@ -748,7 +724,6 @@ public:
|
||||
backend::TokenizationResponse *response) override {
|
||||
std::lock_guard<std::mutex> lock(g_engine_mu);
|
||||
if (!g_engine) return GStatus(StatusCode::FAILED_PRECONDITION, "ds4: model not loaded");
|
||||
if (GStatus id = check_model_identity(request); !id.ok()) return id;
|
||||
ds4_tokens out = {};
|
||||
ds4_tokenize_text(g_engine, request->prompt().c_str(), &out);
|
||||
for (int i = 0; i < out.len; ++i) response->add_tokens(out.v[i]);
|
||||
@@ -763,7 +738,6 @@ public:
|
||||
if (!g_engine || !g_session) {
|
||||
return GStatus(StatusCode::FAILED_PRECONDITION, "ds4: model not loaded");
|
||||
}
|
||||
if (GStatus id = check_model_identity(request); !id.ok()) return id;
|
||||
if (std::string route_err = wait_route_ready(lock); !route_err.empty()) {
|
||||
return GStatus(StatusCode::UNAVAILABLE, route_err);
|
||||
}
|
||||
@@ -863,7 +837,6 @@ public:
|
||||
if (!g_engine || !g_session) {
|
||||
return GStatus(StatusCode::FAILED_PRECONDITION, "ds4: model not loaded");
|
||||
}
|
||||
if (GStatus id = check_model_identity(request); !id.ok()) return id;
|
||||
if (std::string route_err = wait_route_ready(lock); !route_err.empty()) {
|
||||
return GStatus(StatusCode::UNAVAILABLE, route_err);
|
||||
}
|
||||
|
||||
@@ -1,14 +1,12 @@
|
||||
#!/bin/bash
|
||||
set -euo pipefail
|
||||
set -e
|
||||
CURDIR=$(dirname "$(realpath "$0")")
|
||||
REPO_ROOT="${CURDIR}/../../.."
|
||||
PACKAGE_DIR="$CURDIR/package"
|
||||
|
||||
rm -rf "$PACKAGE_DIR"
|
||||
mkdir -p "$PACKAGE_DIR/lib"
|
||||
cp -avf "$CURDIR/grpc-server" "$PACKAGE_DIR/"
|
||||
cp -avf "$CURDIR/ds4-worker" "$PACKAGE_DIR/"
|
||||
cp -rfv "$CURDIR/run.sh" "$PACKAGE_DIR/"
|
||||
mkdir -p "$CURDIR/package/lib"
|
||||
cp -avf "$CURDIR/grpc-server" "$CURDIR/package/"
|
||||
cp -avf "$CURDIR/ds4-worker" "$CURDIR/package/"
|
||||
cp -rfv "$CURDIR/run.sh" "$CURDIR/package/"
|
||||
|
||||
UNAME_S=$(uname -s)
|
||||
if [ "$UNAME_S" = "Darwin" ]; then
|
||||
@@ -18,54 +16,25 @@ if [ "$UNAME_S" = "Darwin" ]; then
|
||||
fi
|
||||
|
||||
if [ -f "/lib64/ld-linux-x86-64.so.2" ]; then
|
||||
cp -arfLv /lib64/ld-linux-x86-64.so.2 "$PACKAGE_DIR/lib/ld.so"
|
||||
cp -arfLv /lib64/ld-linux-x86-64.so.2 "$CURDIR/package/lib/ld.so"
|
||||
LIBDIR=/lib/x86_64-linux-gnu
|
||||
elif [ -f "/lib/ld-linux-aarch64.so.1" ]; then
|
||||
cp -arfLv /lib/ld-linux-aarch64.so.1 "$PACKAGE_DIR/lib/ld.so"
|
||||
cp -arfLv /lib/ld-linux-aarch64.so.1 "$CURDIR/package/lib/ld.so"
|
||||
LIBDIR=/lib/aarch64-linux-gnu
|
||||
else
|
||||
echo "package.sh: unknown architecture" >&2; exit 1
|
||||
fi
|
||||
|
||||
# Bundle the complete dependency closure for both executables. In particular,
|
||||
# grpc-server links the distro gRPC/protobuf/absl stack; copying only the core
|
||||
# C/C++ runtime libraries leaves the scratch image unable to start.
|
||||
{
|
||||
ldd "$CURDIR/grpc-server"
|
||||
ldd "$CURDIR/ds4-worker"
|
||||
} | awk '$2 == "=>" && $3 ~ /^\// { print $3 }' | sort -u | \
|
||||
while read -r so; do
|
||||
cp -arfLv "$so" "$PACKAGE_DIR/lib/"
|
||||
for lib in libc.so.6 libgcc_s.so.1 libstdc++.so.6 libm.so.6 libgomp.so.1 \
|
||||
libdl.so.2 librt.so.1 libpthread.so.0; do
|
||||
cp -arfLv "$LIBDIR/$lib" "$CURDIR/package/lib/$lib"
|
||||
done
|
||||
|
||||
GPU_LIB_SCRIPT="${REPO_ROOT}/scripts/build/package-gpu-libs.sh"
|
||||
if [ -f "$GPU_LIB_SCRIPT" ]; then
|
||||
# shellcheck source=/dev/null
|
||||
source "$GPU_LIB_SCRIPT" "$PACKAGE_DIR/lib"
|
||||
source "$GPU_LIB_SCRIPT" "$CURDIR/package/lib"
|
||||
package_gpu_libs
|
||||
fi
|
||||
|
||||
# Resolve every dependency through the same loader and library path used by
|
||||
# the from-scratch image. The loader can still search host defaults, so reject
|
||||
# any absolute dependency path that escapes the package instead of accepting a
|
||||
# false-positive validation against a library that scratch will not contain.
|
||||
validate_packaged_binary() {
|
||||
local binary="$1"
|
||||
local resolution
|
||||
resolution=$("$PACKAGE_DIR/lib/ld.so" \
|
||||
--library-path "$PACKAGE_DIR/lib" \
|
||||
--list "$PACKAGE_DIR/$binary")
|
||||
|
||||
printf '%s\n' "$resolution" | awk -v prefix="$PACKAGE_DIR/lib/" '
|
||||
$2 == "=>" && $3 ~ /^\// && index($3, prefix) != 1 {
|
||||
print "package.sh: dependency resolved outside package: " $0 > "/dev/stderr"
|
||||
invalid = 1
|
||||
}
|
||||
END { exit invalid }
|
||||
'
|
||||
}
|
||||
|
||||
for binary in grpc-server ds4-worker; do
|
||||
validate_packaged_binary "$binary"
|
||||
done
|
||||
|
||||
echo "ds4 package contents:"
|
||||
ls -lah "$PACKAGE_DIR/" "$PACKAGE_DIR/lib/"
|
||||
ls -lah "$CURDIR/package/" "$CURDIR/package/lib/"
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
|
||||
IK_LLAMA_VERSION?=b054a8b983827c01aec59d4dc273a27c492c51c4
|
||||
IK_LLAMA_VERSION?=f96eaddba8bed6a9a5e628bbf6a566775c70b49c
|
||||
LLAMA_REPO?=https://github.com/ikawrakow/ik_llama.cpp
|
||||
|
||||
CMAKE_ARGS?=
|
||||
|
||||
@@ -2412,33 +2412,7 @@ static void params_parse(const backend::ModelOptions* request,
|
||||
|
||||
// GRPC Server start
|
||||
class BackendServiceImpl final : public backend::Backend::Service {
|
||||
private:
|
||||
// The ModelOptions.Model this process was loaded with. Compared against
|
||||
// PredictOptions.ModelIdentity so a request that reached us through a stale
|
||||
// distributed route is rejected instead of answered from the wrong model
|
||||
// (#10952).
|
||||
std::string loaded_model_identity;
|
||||
|
||||
public:
|
||||
// checkModelIdentity mirrors pkg/grpc/server.go and
|
||||
// backend/python/common/model_identity.py. Either side being empty means
|
||||
// "skip": the request side is empty for a controller that predates the field,
|
||||
// and the loaded side is empty when such a controller performed the load. A
|
||||
// false rejection is worse than the miss it prevents.
|
||||
grpc::Status checkModelIdentity(const backend::PredictOptions* request) {
|
||||
if (request == nullptr || request->modelidentity().empty()) {
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
if (loaded_model_identity.empty() || loaded_model_identity == request->modelidentity()) {
|
||||
return grpc::Status::OK;
|
||||
}
|
||||
// NOT_FOUND plus this exact sentinel is the cross-language contract the
|
||||
// router matches on (grpcerrors.ModelMismatchSentinel).
|
||||
return grpc::Status(grpc::StatusCode::NOT_FOUND,
|
||||
"ik-llama-cpp: model identity mismatch: loaded \"" + loaded_model_identity +
|
||||
"\", requested \"" + request->modelidentity() + "\"");
|
||||
}
|
||||
|
||||
grpc::Status Health(ServerContext* context, const backend::HealthMessage* request, backend::Reply* reply) {
|
||||
// Implement Health RPC
|
||||
reply->set_message("OK");
|
||||
@@ -2464,12 +2438,9 @@ public:
|
||||
result->set_message("Loading succeeded");
|
||||
result->set_success(true);
|
||||
loaded_model = true;
|
||||
loaded_model_identity = request->model();
|
||||
return Status::OK;
|
||||
}
|
||||
grpc::Status PredictStream(grpc::ServerContext* context, const backend::PredictOptions* request, grpc::ServerWriter<backend::Reply>* writer) override {
|
||||
auto identity = checkModelIdentity(request);
|
||||
if (!identity.ok()) return identity;
|
||||
json data = parse_options(true, request, llama);
|
||||
const int task_id = llama.queue_tasks.get_new_id();
|
||||
llama.queue_results.add_waiting_task_id(task_id);
|
||||
@@ -2524,8 +2495,6 @@ public:
|
||||
|
||||
|
||||
grpc::Status Predict(ServerContext* context, const backend::PredictOptions* request, backend::Reply* reply) {
|
||||
auto identity = checkModelIdentity(request);
|
||||
if (!identity.ok()) return identity;
|
||||
json data = parse_options(false, request, llama);
|
||||
const int task_id = llama.queue_tasks.get_new_id();
|
||||
llama.queue_results.add_waiting_task_id(task_id);
|
||||
@@ -2563,8 +2532,6 @@ public:
|
||||
|
||||
/// https://github.com/ggerganov/llama.cpp/blob/aa2341298924ac89778252015efcb792f2df1e20/examples/server/server.cpp#L2969
|
||||
grpc::Status Embedding(ServerContext* context, const backend::PredictOptions* request, backend::EmbeddingResult* embeddingResult) {
|
||||
auto identity = checkModelIdentity(request);
|
||||
if (!identity.ok()) return identity;
|
||||
json data = parse_options(false, request, llama);
|
||||
const int task_id = llama.queue_tasks.get_new_id();
|
||||
llama.queue_results.add_waiting_task_id(task_id);
|
||||
@@ -2589,8 +2556,6 @@ public:
|
||||
}
|
||||
|
||||
grpc::Status TokenizeString(ServerContext* context, const backend::PredictOptions* request, backend::TokenizationResponse* response){
|
||||
auto identity = checkModelIdentity(request);
|
||||
if (!identity.ok()) return identity;
|
||||
json data = parse_options(false, request, llama);
|
||||
|
||||
std::vector<llama_token> tokens = llama.tokenize(data["prompt"],false);
|
||||
|
||||
157
backend/cpp/llama-cpp-localai-paged/Makefile
Normal file
157
backend/cpp/llama-cpp-localai-paged/Makefile
Normal file
@@ -0,0 +1,157 @@
|
||||
|
||||
# llama-cpp-localai-paged is LocalAI's paged-attention llama.cpp variant. It
|
||||
# builds upstream llama.cpp with the LocalAI paged-attention patch series
|
||||
# (patches/paged/, vendored in THIS backend) applied on top. It reuses
|
||||
# backend/cpp/llama-cpp's grpc-server.cpp / CMakeLists.txt / prepare.sh / Makefile
|
||||
# sources verbatim via a thin wrapper - the stock llama-cpp backend is pure
|
||||
# upstream and carries NONE of the paged patches; this backend OWNS them.
|
||||
#
|
||||
# Pin handling (mirrors the turboquant wrapper, the precedent this is modelled
|
||||
# on): the paged patch series is hand-verified bit-exact against ONE specific
|
||||
# llama.cpp tip and re-exported by the manual PIN_SYNC process
|
||||
# (README section 7 + .agents/llama-cpp-localai-paged-backend.md). A naive
|
||||
# pin bump would move the tip out from
|
||||
# under the patches and break `git apply` at build time, so this backend OWNS
|
||||
# its pin (LLAMA_VERSION below) instead of inheriting the auto-bumped stock pin
|
||||
# from backend/cpp/llama-cpp/Makefile. The override is forced into every copied
|
||||
# build via `LLAMA_VERSION=$(LLAMA_VERSION)`. There is deliberately NO
|
||||
# bump_deps.yaml entry for it: it is advanced ONLY by PIN_SYNC, never nightly.
|
||||
# (turboquant CAN auto-bump because its fork branch carries the patches; the
|
||||
# paged series is vendored as .patch files here, so it cannot.)
|
||||
#
|
||||
# - NO patch-grpc-server.sh and NO apply-patches.sh: the shared grpc-server.cpp
|
||||
# already carries the (runtime-gated) paged option hooks, and the paged patch
|
||||
# series (patches/paged/) is applied by THIS Makefile's own apply step onto
|
||||
# the freshly cloned tree, using the same strict `git apply` method the stock
|
||||
# build uses for base patches. The stock llama-cpp Makefile applies only its
|
||||
# own (currently empty) base patches/ series, never the paged one.
|
||||
|
||||
# Manually pin-synced llama.cpp tip the paged patch series is verified against.
|
||||
# Decoupled from the auto-bumped stock pin in backend/cpp/llama-cpp/Makefile so
|
||||
# the nightly llama.cpp bump cannot silently break the vendored paged patches.
|
||||
# Advance ONLY via the PIN_SYNC process (rebase patches + bit-exact gate +
|
||||
# re-export), then update this value. See:
|
||||
# README section 7 + .agents/llama-cpp-localai-paged-backend.md
|
||||
#
|
||||
# This pin = the manual, verified sync. The signal telling you WHEN to do the
|
||||
# next sync is the early-warning canary
|
||||
# (.github/workflows/llama-cpp-paged-canary.yml): weekly it applies + compiles
|
||||
# this patch series against the latest upstream llama.cpp tip and goes red the
|
||||
# moment upstream drifts past the patches. Canary red -> run a PIN_SYNC, then
|
||||
# bump this value. The canary never touches this pin; it is signal-only.
|
||||
#
|
||||
# HARD CONSTRAINT: keep this == the stock llama-cpp pin (backend/cpp/llama-cpp/
|
||||
# Makefile). grpc-server.cpp is SHARED with the stock backend and tracks the
|
||||
# stock pin; a paged pin that diverges PAST an upstream server-API refactor
|
||||
# breaks the grpc-server LINK even when the patches are byte-for-byte bit-exact.
|
||||
# The c299a92c bump did exactly this: patches applied + greedy-md5 bit-exact, but
|
||||
# grpc-server.cpp failed to link with undefined references to stream_* server
|
||||
# helpers that the refactor pulled into the headers grpc-server.cpp includes.
|
||||
# Therefore a PIN_SYNC must pass the FULL grpc-server build/link on CI, not only
|
||||
# the bit-exact gate. See README section 7 + .agents/llama-cpp-localai-paged-backend.md.
|
||||
LLAMA_VERSION?=0ed235ea2c17a19fc8238668653946721ed136fd
|
||||
|
||||
CMAKE_ARGS?=
|
||||
BUILD_TYPE?=
|
||||
NATIVE?=false
|
||||
ONEAPI_VARS?=/opt/intel/oneapi/setvars.sh
|
||||
TARGET?=--target grpc-server
|
||||
JOBS?=$(shell nproc 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 1)
|
||||
ARCH?=$(shell uname -m)
|
||||
|
||||
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))
|
||||
LLAMA_CPP_DIR := $(CURRENT_MAKEFILE_DIR)/../llama-cpp
|
||||
# OUR vendored paged-attention patch series. Owned by this backend; the stock
|
||||
# llama-cpp backend no longer carries it. Applied onto each freshly cloned
|
||||
# llama.cpp tree by apply-paged-patches below (strict git apply).
|
||||
PAGED_PATCHES_DIR := $(CURRENT_MAKEFILE_DIR)/patches/paged
|
||||
|
||||
GREEN := \033[0;32m
|
||||
RESET := \033[0m
|
||||
|
||||
# Apply OUR vendored paged-attention patch series (patches/paged/0*.patch) onto a
|
||||
# freshly cloned llama.cpp tree ($(1)) using the SAME strict git-apply method the
|
||||
# stock build uses for its base patches (backend/cpp/llama-cpp/Makefile `llama.cpp`
|
||||
# target). Strict: any patch that no longer applies aborts the build (exit 1) -
|
||||
# that is the signal to run a PIN_SYNC, never to bump the pin blindly. The series
|
||||
# is owned by THIS backend, not by the now-pure stock llama-cpp backend.
|
||||
define apply-paged-patches
|
||||
cd $(1) && \
|
||||
for p in $(PAGED_PATCHES_DIR)/0*.patch; do \
|
||||
[ -e "$$p" ] || continue; \
|
||||
echo "applying llama.cpp PAGED patch: $$p"; \
|
||||
git apply --verbose "$$p" || { echo "paged patch failed: $$p"; exit 1; }; \
|
||||
done
|
||||
endef
|
||||
|
||||
# Each flavor target:
|
||||
# 1. copies backend/cpp/llama-cpp/ (grpc-server.cpp + prepare.sh +
|
||||
# CMakeLists.txt + Makefile) into a sibling
|
||||
# llama-cpp-localai-paged-<flavor>-build directory;
|
||||
# 2. clones OUR pinned upstream llama.cpp into that copy via the copy's own
|
||||
# `llama.cpp` target (which applies the stock base patches/ series, normally
|
||||
# empty), then applies THIS backend's paged patch series (patches/paged/)
|
||||
# onto the cloned tree with strict `git apply` (apply-paged-patches);
|
||||
# 3. runs the copy's `grpc-server` target and copies the produced binary up as
|
||||
# llama-cpp-localai-paged-<flavor>.
|
||||
# We clone+patch only the *copy*, never the original under backend/cpp/llama-cpp/,
|
||||
# so the stock llama-cpp build stays untouched and patch-free.
|
||||
define paged-build
|
||||
rm -rf $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build
|
||||
cp -rf $(LLAMA_CPP_DIR) $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build
|
||||
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build purge
|
||||
$(info $(GREEN)I llama-cpp-localai-paged build info:$(1)$(RESET))
|
||||
LLAMA_VERSION=$(LLAMA_VERSION) $(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build llama.cpp
|
||||
$(call apply-paged-patches,$(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build/llama.cpp)
|
||||
CMAKE_ARGS="$(CMAKE_ARGS) $(2)" TARGET="$(3)" LLAMA_VERSION=$(LLAMA_VERSION) \
|
||||
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build grpc-server
|
||||
cp -rfv $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-$(1)-build/grpc-server llama-cpp-localai-paged-$(1)
|
||||
endef
|
||||
|
||||
llama-cpp-localai-paged-avx2:
|
||||
$(call paged-build,avx2,-DGGML_AVX=on -DGGML_AVX2=on -DGGML_AVX512=off -DGGML_FMA=on -DGGML_F16C=on,--target grpc-server)
|
||||
|
||||
llama-cpp-localai-paged-avx512:
|
||||
$(call paged-build,avx512,-DGGML_AVX=on -DGGML_AVX2=off -DGGML_AVX512=on -DGGML_FMA=on -DGGML_F16C=on,--target grpc-server)
|
||||
|
||||
llama-cpp-localai-paged-avx:
|
||||
$(call paged-build,avx,-DGGML_AVX=on -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server)
|
||||
|
||||
llama-cpp-localai-paged-fallback:
|
||||
$(call paged-build,fallback,-DGGML_AVX=off -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server)
|
||||
|
||||
# Single-build CPU backend via ggml CPU_ALL_VARIANTS (mirrors llama-cpp-cpu-all).
|
||||
# Reuses backend/cpp/llama-cpp's CMakeLists.txt (hw_grpc_proto STATIC) and
|
||||
# Makefile (SHARED_LIBS make-var + EXTRA_CMAKE_ARGS), so this passes the same
|
||||
# overrides through to the copied build: SHARED_LIBS=ON, the DL flags, and
|
||||
# --target ggml (which pulls in the per-microarch libggml-cpu-*.so via ggml's
|
||||
# add_dependencies). The .so set is collected for package.sh to bundle into
|
||||
# package/lib.
|
||||
llama-cpp-localai-paged-cpu-all:
|
||||
rm -rf $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build
|
||||
cp -rf $(LLAMA_CPP_DIR) $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build
|
||||
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build purge
|
||||
$(info $(GREEN)I llama-cpp-localai-paged build info:cpu-all-variants$(RESET))
|
||||
LLAMA_VERSION=$(LLAMA_VERSION) $(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build llama.cpp
|
||||
$(call apply-paged-patches,$(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build/llama.cpp)
|
||||
SHARED_LIBS=ON EXTRA_CMAKE_ARGS="-DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON" TARGET="--target grpc-server --target ggml" LLAMA_VERSION=$(LLAMA_VERSION) \
|
||||
$(MAKE) -C $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build grpc-server
|
||||
cp -rfv $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build/grpc-server llama-cpp-localai-paged-cpu-all
|
||||
rm -rf ggml-shared-libs && mkdir -p ggml-shared-libs
|
||||
find $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-cpu-all-build/llama.cpp/build \( -name '*.so*' -o -name '*.dylib' \) -exec cp -av {} ggml-shared-libs/ \;
|
||||
@echo "Collected ggml shared backends:" && ls -la ggml-shared-libs/
|
||||
|
||||
llama-cpp-localai-paged-grpc:
|
||||
$(call paged-build,grpc,-DGGML_RPC=ON -DGGML_AVX=off -DGGML_AVX2=off -DGGML_AVX512=off -DGGML_FMA=off -DGGML_F16C=off -DGGML_BMI2=off,--target grpc-server --target ggml-rpc-server)
|
||||
|
||||
llama-cpp-localai-paged-rpc-server: llama-cpp-localai-paged-grpc
|
||||
cp -rf $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-grpc-build/llama.cpp/build/bin/ggml-rpc-server llama-cpp-localai-paged-rpc-server
|
||||
|
||||
package:
|
||||
bash package.sh
|
||||
|
||||
purge:
|
||||
rm -rf $(CURRENT_MAKEFILE_DIR)/../llama-cpp-localai-paged-*-build
|
||||
rm -rf llama-cpp-localai-paged-* package
|
||||
|
||||
clean: purge
|
||||
699
backend/cpp/llama-cpp-localai-paged/README.md
Normal file
699
backend/cpp/llama-cpp-localai-paged/README.md
Normal file
@@ -0,0 +1,699 @@
|
||||
# LocalAI paged-attention llama.cpp patch series
|
||||
|
||||
This backend vendors the patch series (in `patches/paged/`) that turns stock
|
||||
llama.cpp into LocalAI's paged-attention variant (`llama-cpp-localai-paged`). The
|
||||
patches are applied on top of a pinned upstream llama.cpp at build time; nothing
|
||||
here is a fork - it is a source-only `*.patch` stack plus this canonical doc.
|
||||
|
||||
> One-file rule: this README is the canonical reference for the patch series. The
|
||||
> only other docs are operational, kept in `docs/`, and linked below:
|
||||
> - [`PAGED_BITEXACT_NOTE.md`](docs/PAGED_BITEXACT_NOTE.md) - the per-path bit-exactness gate (the canonical paged-MoE md5 reference).
|
||||
> - [`LOCALAI_LLAMACPP_BACKEND_PLAN.md`](docs/LOCALAI_LLAMACPP_BACKEND_PLAN.md) - the design-of-record for shipping this as its own backend + the NVFP4 gallery items.
|
||||
> - [`VLLM_PARITY_FINAL.md`](docs/VLLM_PARITY_FINAL.md) - the definitive, closed record of the GB10 vLLM-parity investigation: full benchmark, every lever + verdict, the structural floors, and the parity verdict (summarized in section 9 below). Read this before reopening any parity work.
|
||||
> - [`EXECUTION_REARCH_SCOPE.md`](docs/EXECUTION_REARCH_SCOPE.md) - the reopened scope: ports vLLM's execution *architecture* (bf16-resident stream, expert-major fused MoE region, persistent-CTA GEMM, token-budget scheduler, blocked-solve GDN) into the fork additively, on the thesis that same-silicon 2-3x is software-architecture-conditional, not a hardware floor. Phased (P1-P6), each with a falsifiable P0 kill-gate. Read this to pick up parity work after `VLLM_PARITY_FINAL.md`.
|
||||
|
||||
---
|
||||
|
||||
## 1. What it is
|
||||
|
||||
`llama-cpp-localai-paged` is the LocalAI paged-attention llama.cpp backend: a
|
||||
vendored patch series over upstream llama.cpp that adds
|
||||
|
||||
- a **paged KV cache** (vLLM-style block manager: on-demand fixed-size blocks,
|
||||
free pool, ref-counted blocks) with a **block-table flash-attention** read so
|
||||
the attention kernels index physical cells instead of a contiguous buffer;
|
||||
- **cross-request prefix sharing** - concurrent requests that share a long
|
||||
prefix physically reuse one committed copy of the prefix blocks and prefill
|
||||
only their divergent suffix;
|
||||
- a **decode-first prefill scheduler** - a dynamic per-step prefill-token budget
|
||||
decoupled from `n_batch`, so a long prefill never freezes co-batched decode;
|
||||
- **GB10 / Blackwell NVFP4 decode optimizations** for the Qwen3.6 hybrid
|
||||
gated-DeltaNet (SSM) models, where the recurrent-state plumbing - not the FP4
|
||||
GEMM - dominates the decode step.
|
||||
|
||||
It is **pinned to llama.cpp `0ed235ea2c17a19fc8238668653946721ed136fd`** (kept == the stock `llama-cpp` backend's
|
||||
pin) and advanced only by a manual, bit-exact-gated pin-sync process (see
|
||||
section 7, "Pin + maintenance policy"), decoupled from the nightly auto-bumper. The pin must stay aligned with the stock pin because
|
||||
`grpc-server.cpp` is shared; an earlier bump to `c299a92c` was bit-exact but broke
|
||||
the grpc-server link and was reverted to the then-current stock pin.
|
||||
|
||||
The build gate is `LLAMA_PAGED` (default on in this tree); the paged engine is
|
||||
enabled per-model at runtime via the gallery `options:` knobs (`paged_kv:true`,
|
||||
`max_batch_tokens:`, `kv_unified:false`, ...). Against unpatched llama.cpp the
|
||||
runtime hooks are inert, so a single `grpc-server.cpp` is shared between the
|
||||
clean and the paged build.
|
||||
|
||||
---
|
||||
|
||||
## 2. Architecture
|
||||
|
||||
The decode step on these models breaks into three cost centers; the patch series
|
||||
attacks each one.
|
||||
|
||||
**Paged KV manager + block-table flash-attn.** A host-side `PagedKVManager`
|
||||
(`FreeBlockQueue` / `BlockPool` / chained-hash content cache) hands out
|
||||
fixed-size KV blocks on demand and reclaims them per-sequence (ref-counted, with
|
||||
copy-on-write for shared prefixes). The attention path reads through a **block
|
||||
table** - an `I32 [n_view, n_stream]` position-ordered physical-cell index passed
|
||||
as `src[5]` of `ggml_flash_attn_ext` - so the CUDA fattn vec/tile kernels and the
|
||||
CPU reference map logical KV index `j` to physical cell `block_table[seq*ne11+j]`
|
||||
and read K/V in place. Token-position ordering keeps the flash-attn online-softmax
|
||||
reduction order identical to stock. A null block table is the stock contiguous
|
||||
read, byte-identical.
|
||||
|
||||
**The gated-DeltaNet (GDN / SSM) decode path.** The Qwen3.6 hybrid models are 48
|
||||
gated-DeltaNet (linear-attention / SSM) layers + 16 full-attention layers. On
|
||||
GB10 the recurrent-state plumbing, not the weight GEMM, is the dominant decode
|
||||
cost. The series fuses that plumbing to mirror vLLM's
|
||||
`fused_recurrent_gated_delta_rule`: the recurrent state is read from and written
|
||||
to its cache slot in place (no copy-back, no `get_rows` materialization), the
|
||||
conv state is updated in place, the output projection is reshaped to route to the
|
||||
tensor-core MMQ GEMM, and the recurrence kernel is occupancy-retuned - all
|
||||
bit-exact (md5-gateable) against the f32 baseline.
|
||||
|
||||
**NVFP4 native FP4-MMA on Blackwell.** The NVFP4 dense/expert weight GEMM uses
|
||||
Blackwell's native FP4-MMA. The series removes a redundant activation-requantize
|
||||
in the MoE broadcast projections (bit-exact byte copy of identical blocks) and
|
||||
keeps CUDA graphs on for the grouped-MMQ MoE decode step. These are the only
|
||||
NVFP4-specific optimizations; on non-Blackwell hardware the FP4 path falls back
|
||||
to dequant.
|
||||
|
||||
**The prefill/decode scheduler.** `update_slots()` already emits one unified
|
||||
mixed prefill+decode batch per step. The scheduler patches change only the *count*
|
||||
of prefill tokens admitted per step: decode tokens are claimed first
|
||||
(decode-first), then a dynamic budget `max(n_ubatch, T - D)` (where `D` is the
|
||||
live decode load and `T` is `LLAMA_MAX_BATCH_TOKENS`) admits prefill, auto-
|
||||
shrinking as decode load rises. Pure scheduler policy, byte-identical when off,
|
||||
orthogonal to the paged allocator.
|
||||
|
||||
---
|
||||
|
||||
## 3. Patch series (0001-0063)
|
||||
|
||||
Source-only patches, with intentional numbering gaps (e.g. 0005, 0027). The
|
||||
decode-serving graph-reuse levers are 0040-0041. "Bit-exact" = greedy md5 /
|
||||
`test-backend-ops` byte-identical to the relevant baseline; the gate methodology
|
||||
is in section 5.
|
||||
|
||||
### Paged-KV core (0001-0012)
|
||||
|
||||
| # | What it does | Bit-exact |
|
||||
|---|---|---|
|
||||
| 0001 | Vendor the host-side paged KV block manager (`FreeBlockQueue`, `BlockPool`, `PagedKVManager`, chained-hash prefix cache). Pure C++17, nothing uses it yet. | n/a (no behavior) |
|
||||
| 0002 | Place each sequence at permuted, non-contiguous block positions in `find_slot` (proves attention is invariant to physical KV placement). | yes (token-identical) |
|
||||
| 0003 | Gather K/V/mask down to each stream's non-empty cells before `build_attn_mha`, position-sorted so the FA reduction order matches stock. | yes |
|
||||
| 0004 | Drive paged placement through the vendored manager: blocks popped on demand, returned on seq end. Core kv-cache struct untouched. | yes (stock path byte-identical) |
|
||||
| 0006 | Host-side cross-request prefix caching: hash prefix blocks, reuse matching physical blocks (ref-count++), COW-privatise before a divergent write. | yes (default off) |
|
||||
| 0007 | Wire the prefix cache into the engine so a new sequence physically shares cached prefix blocks and skips recomputing the shared prefix. | yes (verified byte-identical) |
|
||||
| 0008 | Wire cross-request prefix share into the llama-server continuous-batch loop so concurrent shared-prefix requests prefill only the suffix (36x fewer prefill tokens at K=32). | within CUDA batch-shape non-determinism band |
|
||||
| 0009 | Replace the per-step gather with an **in-kernel paged read** (block table as `src[5]`); the K/V `get_rows` is gone. Decode step at batch32 691->696ms (was 1279ms gathered). | yes on CPU/batch1; GPU batch>1 within vec-vs-mma band |
|
||||
| 0010 | Graft the block-table read into the tile kernel; add a dispatch guard so a present block table routes ONLY to vec/tile (never the mma/wmma kernels that ignore it). | yes (CPU byte-identical; vec route) |
|
||||
| 0011 | Route the GQA-grouped F16 decode to the **tile kernel** (native head-group reuse) by default; vec for everything else. Paged decode to within 1.8% of stock. | vs stock-mma: different-kernel rounding; bit-exact vs vec |
|
||||
| 0012 | Defensive `GGML_ASSERT(n_view % 64 == 0)` so a future pad/tile change can't silently reintroduce a past-end KV leak on the tile route. | yes (additive assert) |
|
||||
|
||||
### Decode-first scheduler (0013, 0016)
|
||||
|
||||
| # | What it does | Bit-exact |
|
||||
|---|---|---|
|
||||
| 0013 | `LLAMA_PREFILL_BUDGET`: a static per-step prefill-token budget decoupled from `n_batch` (vLLM `--max-num-batched-tokens` analogue). Flattens the decode ITL spike a long prefill inflicts (8.5x smaller worst freeze). | yes (off/short = byte-identical; == `-b` chunking) |
|
||||
| 0016 | Supersede 0013 with a **dynamic decode-first** budget: `max(n_ubatch, T-D)`, auto-shrinking as decode load `D` rises. Policy-only inside `update_slots()`, zero libllama changes. | yes (default-off byte-identical) |
|
||||
|
||||
(0014/0015 are the MoE token-tile levers: 0014 adds `LLAMA_MOE_MMQ_X` (opt-in
|
||||
high-batch decode micro-opt, +4.8% on Qwen3-Coder-30B), 0015 makes it a
|
||||
default-on, density-aware auto-select that is prefill-safe by construction. Both
|
||||
bit-exact. 0017 is the dense FP4-GEMM occupancy-tune track: bit-exact gate green,
|
||||
but every cheap occupancy lever regressed on GB10, so nothing is enabled - it
|
||||
ships as the parity gate + default-off instrumentation only.)
|
||||
|
||||
### Decode-serving graph reuse (0040, 0041)
|
||||
|
||||
These two close the **continuous-serving** decode gap (distinct from the static
|
||||
batched-bench decode kernel, which is already at vLLM parity - see
|
||||
[`docs/DECODE_SERVING_SCOPE.md`](docs/DECODE_SERVING_SCOPE.md)). In serving the
|
||||
host rebuilt the ggml graph on **every** decode step (layer-A graph reuse was 0%),
|
||||
so the GPU idled while the host rebuilt - the host-bound -39% the static bench
|
||||
hides.
|
||||
|
||||
| # | What it does | Bit-exact |
|
||||
|---|---|---|
|
||||
| 0040 | **S1 paged decode-graph reuse** - the paged decode inputs (`input_block_table` / `input_gather_idxs`) never overrode `can_reuse` (defaults to false), so any graph carrying a paged input could never be reused. Add a correct `can_reuse` keyed on the (256-bucketed) block-table dims + a live-mctx refresh from the owning attn input. `LLAMA_PAGED_NO_GRAPH_REUSE=1` forces the pre-S1 path. | yes (md5 byte-identical reuse on/off; dense `5951a5b4`, paged-MoE `8cb0ce23`) |
|
||||
| 0041 | **S3 decode-shape-stable scheduling** - keep co-batched prefill OUT of decode steps so the pure-decode batch shape stays reuse-stable (S1 makes a pure-decode step reusable; S3 makes the scheduler emit them). Pure `update_slots()` policy on top of 0016; prefill admitted on a bounded cadence (`LLAMA_PAGED_PREFILL_PERIOD`, default 8). **Default OFF** (opt-in via `LLAMA_PAGED_DECODE_STABLE=1`): a measured end-to-end A/B proved default-on is a serving mistake - deferring prefill admission on the period-8 cadence gives **2.5x worse TTFT** (60s vs 24s at N=256) and **20-29% lower end-to-end throughput**, with no end-to-end win at any concurrency; its apparent `decode_agg` gain was a metric artifact (faster per-step decode bought by starving prefill). Default prefers prompt prefill admission for good TTFT; opt in only for decode-dominated, low-arrival traffic where TTFT is not a concern. | yes (byte-identical on/off; per-stream independent in serving) |
|
||||
|
||||
Measured (GB10, MoE Qwen3.6-35B-A3B-NVFP4, 128-client staggered streaming load):
|
||||
graph reuse **0% -> 72.2%**, host window `hostproc` **15.98 -> 6.31 ms/step**,
|
||||
decode **4.05 -> 5.52 tok/s/seq median (4.24 -> 5.96 mean, at vLLM's ~5.9
|
||||
sustained)**. S1 is necessary but **not** sufficient alone (13.8% reuse - prefill
|
||||
co-batching churns the shape nearly every step); S3 is the multiplier of that
|
||||
per-step decode metric. **But those are per-step decode numbers, not an end-to-end
|
||||
serving win**: a later end-to-end A/B showed S3-default-on regresses real serving
|
||||
(2.5x worse TTFT, 20-29% lower end-to-end throughput, no win at any concurrency),
|
||||
because the period-8 cadence defers prefill admission. So **only S1 (0040) ships
|
||||
default-on; S3 (0041) now defaults OFF and is opt-in** (`LLAMA_PAGED_DECODE_STABLE=1`,
|
||||
for decode-dominated low-arrival traffic). The static batched-bench A/B isolates the S1
|
||||
mechanism: paged decode reuse 0% -> 95.5% (throughput flat there, since the static
|
||||
regime is GPU-bound). **S2 (double-buffer `set_inputs`) was dropped**: the Phase-0
|
||||
profile put `set_inputs` at ~0.05 ms/step (the cost is the rebuild, not the input
|
||||
copy), so it has nothing to recover. The remaining ~28% serving rebuilds are
|
||||
request-boundary D/seq-set churn + the prefill-cadence steps. A **padded/fixed-slot
|
||||
decode shape** to capture them was then implemented and GPU-tested (2026-06-28) and
|
||||
**REJECTED** - it is bit-exact/inert but regresses serving throughput at every
|
||||
concurrency, because this serving decode is GPU-compute-bound (baseline reuse 0% ~=
|
||||
S1+S3 reuse 72% on aggregate tok/s), so the dummy-row compute it adds costs more
|
||||
than the reuse it recovers. Full record + numbers in `docs/DECODE_SERVING_SCOPE.md`
|
||||
("Padded-shape lever - rejected").
|
||||
|
||||
### Prefill fusions (0042, 0044)
|
||||
|
||||
CUDA-family graph fusions of the pre-norm residual chain and the gated-DeltaNet
|
||||
output norm: separate `rms_norm` / `mul` / `add` / `silu` launches collapse into
|
||||
one kernel so the intermediate never round-trips to HBM. Bit-exact (the fused
|
||||
kernel reproduces the unfused FP order; float multiply is commutative). Each is
|
||||
env-gated default-ON (`LLAMA_FUSE_*=0` for a clean single-build A/B that reverts
|
||||
to the byte- and kernel-identical unfused path).
|
||||
|
||||
| # | What it does | Bit-exact / effect |
|
||||
|---|---|---|
|
||||
| 0042 | **Fused residual-add + RMS norm + weight multiply** (`rms_norm_pre_add_mul_f32`) - the pre-norm residual `h = x + sub_out; n = rms_norm(h) * w` ran as a `k_bin_bcast` ADD feeding the fused rms_norm+mul; the residual ADD has a second consumer (the skip add) so it can't pass the single-use `ggml_can_fuse`. Recognized via `ggml_can_fuse_subgraph` (ADD + final MUL both outputs), folded into one launch that publishes `h` and emits `scale * h * w`. Gate `LLAMA_FUSE_ADD_RMSNORM`. | yes (dense `5951a5b4`, MoE `8cb0ce23`); dense S_PP +0.5% |
|
||||
| 0044 | **Fused gated RMSNorm + SiLU gate multiply** (`rms_norm_gate_mul_f32`) - the gated-DeltaNet output norm `(rms_norm(x) * w) * silu(z)` (qwen35 / qwen35moe `build_norm_gated`) ran as rms_norm_mul + silu_mul, two launches with the normalized intermediate crossing HBM. The gate z-projection (a MUL_MAT) is scheduled between the weight MUL and the SILU, so the chain is not naturally consecutive; `build_norm_gated` emits the gate multiply as `mul(silu(z), normalized)` (commutative, bit-exact) so the graph lays out the consecutive subgraph `{ SILU, RMS_NORM, MUL, MUL }` that `ggml_cuda_can_fuse` folds into one `scale * x * w * silu(z)` launch. Gate `LLAMA_FUSE_GATE_RMSNORM`. Profile (dense npp512): 672 (rms_norm_mul + silu_mul) -> 336 fused launches. | yes (dense `5951a5b4`, MoE `8cb0ce23`, paged + non-paged; `test-backend-ops` 12979/12979); S_PP dense +1.1% (~+10 us/tok), MoE +0.9% |
|
||||
|
||||
### SSM (gated-DeltaNet) decode levers (0018-0022, 0028)
|
||||
|
||||
These are the dominant decode levers on the Qwen3.6 hybrid models. All bit-exact.
|
||||
|
||||
| # | What it does | Effect (dense q36-27b / MoE q36-35b-a3b @npl128) |
|
||||
|---|---|---|
|
||||
| 0018 | **In-place SSM state write-back** - the recurrence writes its final state directly into the cache slot, removing the ~225MB/copy D2D memcpy (18.9% of decode time). | dense +23.5% / MoE +18.9% |
|
||||
| 0019 | **Fused recurrent-state gather** - the op reads each sequence's prior state directly from `cache[ids[seq]]` (no `get_rows` materialization); race-free in-place + ids read. | dense +37.8% / MoE +35.3% |
|
||||
| 0020 | **o_proj MMVQ->MMQ reshape** - collapse the GDN output to 2D so the output projection routes to the M=128 tensor-core MMQ GEMM (was a batch<=8 MMVQ GEMV). The single biggest decode-parity lever. | dense +31.7% (->85.9% of vLLM) / MoE +23.3% |
|
||||
| 0021 | **Conv-state in-place fusion** - one `ggml_ssm_conv_update_inplace` op replaces the 4-op conv chain (transpose+concat+conv+silu+ring-cpy), writing the shifted ring state in place. | dense +3.2% / MoE +3.5% |
|
||||
| 0022 | **GDN recurrence occupancy/coalescing retune** - column-folding (NUM_WARPS/COLS_PER_WARP) raises memory-level parallelism on the bandwidth-bound B=128 recurrence kernel; per-column f32 FMA order unchanged. 73.4%->84.6% of GB10 peak BW. | dense +11.1% / MoE +8.3% |
|
||||
| 0028 | **Recurrent conv-tap gather fusion** - the last `k_get_rows` in the GDN decode path (the conv-state tap gather) becomes an indexed in-kernel read. | dense ~377 t/s / MoE ~784 t/s |
|
||||
|
||||
### MoE NVFP4 quant (0023, 0025, 0043)
|
||||
|
||||
| # | What it does | Bit-exact |
|
||||
|---|---|---|
|
||||
| 0023 | **NVFP4 activation-quantize de-dup** - the broadcast up/gate projections re-quantize the same token activation once per expert; quantize the unique token activations once and byte-copy them into the expert-gathered layout. The only NVFP4-specific patch. | yes (byte-identical) |
|
||||
| 0025 | **MoE decode re-graph** - keep CUDA graphs on for the grouped-MMQ MoE decode step (the upstream guard disables graphs conservatively; the grouped path has no host sync). Was env-gated `LLAMA_MOE_FORCE_GRAPHS`; now ON by default via 0043. | yes (graph replay re-issues identical kernels) |
|
||||
| 0043 | **MoE decode graph default-on (D1)** - flip 0025 to ON by default: capture/replay the full-step decode CUDA graph (incl. the grouped-MMQ MoE dispatch) instead of re-issuing every kernel each step. Guard is `should_use_mmq()` (FALSE for the large-M NVFP4 prefill of 0034, so prefill keeps graphs disabled - its per-expert host-loop genuinely syncs). `LLAMA_MOE_NO_FORCE_GRAPHS=1` forces the conservative pre-0025 disable for A/B. D1 profiling: the per-expert host-loop (the only device->host MoE-routing readback) is never hit on the NVFP4 grouped path (sync count identical graphs on/off); steady decode is ~99% GPU-busy, so the cost removed is per-step host kernel RE-ISSUE, not a sync. | yes (md5 byte-identical default/off/forced; paged-MoE `8cb0ce23`, dense `5951a5b4`) |
|
||||
|
||||
### Pool reclaim, block-table cache, backend gate
|
||||
|
||||
| # | What it does | Bit-exact |
|
||||
|---|---|---|
|
||||
| 0024 | **Paged-pool burst-reclaim** - truncate trailing blocks on partial-tail `seq_rm`, defrag the free queue when idle, release blocks on slot completion. Fixes the long-server burst-degradation bug (post-burst prefill collapse 488->44 t/s, restored to 532). Host-side accounting only. | yes |
|
||||
| 0029 | **Block-table within-step host cache** - the block table is fixed for the whole step; cache it on first build and memcpy it for the other full-attention layers (get_block_table -87%/-91%). | yes, per path (paged-MoE ref `8cb0ce23`) |
|
||||
| 0030 | **Fused-op backend gate** - the fused GDN / discriminated SSM_CONV ops are CUDA-family + CPU only; force them off on any non-CUDA compute backend so a Vulkan/SYCL/Metal build can't silently run the wrong plain-conv kernel. | yes on CUDA (byte-identical pre-0030); safety gate elsewhere |
|
||||
| 0031 | **Chunked parallel-scan GDN prefill kernel** (upstream TODO) - FLA-style chunked gated-delta-rule for prefill (non-KDA / f32 / final-state): intra-chunk delta rule solved in parallel (UT-transform + forward subst), inter-chunk recurrence over n_tokens/C steps. The scalar-serial form (`GDN_TC=0`) was bit-exact-benign but not faster than the tuned sequential scan at the GB10-forced C=16 (see section 5); **superseded for paged by the tensor-core M5 path of 0047**. | NEW per-path (`test-backend-ops` 91/91, <=1e-7 NMSE vs CPU ref) |
|
||||
| 0047 | **GDN M5 tensor-core chunked-scan prefill, f32-only re-port, default-ON under paged KV** - the f32/tf32 tensor-core forms of 0031's scan (KK/QK Gram = M2, KS/QS state-boundary 3xtf32 = M3, P*U output = M4, full form-T solve + state-update mma = M5), single build, runtime-selected by `GDN_TC`. Ships **M5 default-on when `LLAMA_KV_PAGED` is set** (`GDN_TC=5` + `GDN_CHUNK_MIN=64`, both env-overridable; OFF/`INT_MAX` when not paged). `GDN_CHUNK_MIN` is the per-call engage threshold and stays > 1 so decode (1 tok/call) keeps the sequential recurrence (at 1 it swallows decode and drops S_TG ~25%); 64 tuned from a {1,32,64,128,256} sweep. The bf16/hybrid dev-tree machinery (STATE_BF16/HYBRID, the dropped 0026 ssm_bf16_tau) and the bf16 CONFIG-C (M8) plus register-resident M6/M7 variants are NOT part of this f32-only series. MoE prefill S_PP +3.5% @npp512 (3x A/B), +17.7% @npp2048; decode S_TG unchanged. | NEW per-path, benign (`test-backend-ops` GATED_DELTA_NET 46/46 default AND force-M5, incl. multi-chunk/tail-chunk/multi-seq; greedy md5 default-on == M5-forced == canonical on the gate prompt: paged-MoE `8cb0ce23`, dense `5951a5b4`; long MoE prompt = one benign greedy flip vs sequential, dense byte-identical) |
|
||||
| 0046 | **GDN prefill geometry gated by scan length** - patch 0022's `(NUM_WARPS=16, COLS_PER_WARP=8)` column-fold of the GDN sequential-recurrence dispatch (`case 128`) is a decode win but was applied UNCONDITIONALLY, so it also hit dense prefill (~-6% vs stock): on a long sequential scan the launch `grid.z` collapses from `S_v/4 = 32` to `S_v/(16*8) = 1` and the SMs starve (profiled: `gated_delta_net` +54% GPU time = the whole dense-prefill regression). Gate the geometry by per-call scan length: long scans (prefill, `n_tokens >= GDN_PREFILL_NTOK`, default 256) take stock's high-grid.z `(4,1)` geometry; short scans (decode) keep the `(16,8)` retune. Recovers dense prefill +7.2% back to stock parity, keeps the decode win. `GDN_PREFILL_NTOK` tunes the crossover; an explicit `GDN_NW`/`GDN_CPW` sweep still overrides (gate yields when either is set), so the one-build %peak A/B harness is unchanged. | yes (patch 0022 proved every `{NW,CPW}` variant byte-identical, so switching geometry by scan length cannot move the md5) |
|
||||
|
||||
### Speculative / MTP investigation (0054, 0055)
|
||||
|
||||
| # | What it does | Bit-exact / effect |
|
||||
|---|---|---|
|
||||
| 0054 | **Disable backend sampling for MTP drafts** - forces server MTP draft generation through the target-side sampler acceptance path instead of letting the draft backend sample independently. This was required for the Phase 14 rollback/prefix safety gate. | yes for canonical non-MTP gates; Phase 14 MTP normalized greedy-prefix gate passed |
|
||||
| 0055 | **Trace speculative batch shapes** - adds default-off `LLAMA_SPEC_SHAPE_TRACE=1` server logs around `server_slot::handle_last_sampled_token()`, reporting normal decode rows and MTP verification `K + 1` rows (`draft`, `outputs`, `spec_i_first`, `spec_i_last`). This is instrumentation only for Phase 18 shape-entropy measurement before any scheduler experiment. | yes (env unset is silent; DGX gates after patch: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`) |
|
||||
| 0056 | **Trace MoE MMQ batch shapes** - adds default-off `LLAMA_MOE_MMQ_SHAPE_TRACE=<n>` logs from the grouped-MMQ host selector, reporting routed assignment count, estimated active experts, density, selected `mmq_x`, `mmq_y`, and stream-k. This is evidence-only instrumentation for sizing structural grouped-MMQ work after Phase 28 rejected launch-bounds/row-tile knobs. | yes (env unset and trace-enabled gates both green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`; trace cap verified with 4 lines) |
|
||||
| 0057 | **Trace MoE MMQ launch shapes** - extends `LLAMA_MOE_MMQ_SHAPE_TRACE=<n>` with bounded `[LLAMA_MOE_MMQ_LAUNCH]` lines from `launch_mul_mat_q`, recording actual `ntiles_dst`, `stream_k_blocks`, tile efficiency, `fixup`, `ntx/nty/ntzw`, and compiled `mmq_x/mmq_y`. This is evidence-only instrumentation to distinguish real stream-k/fixup overhead from small-M kernel-shape cost. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`; Phase 31 n128 trace showed decode and prefill `fixup=0`, `stream_k_blocks == ntiles_dst`) |
|
||||
| 0058 | **Trace MoE small-M MMQ candidates** - adds `LLAMA_MOE_MMQ_SMALL_M_TRACE=<n>` and a host-only classifier for decode-like low-density grouped-MMQ shapes (`ncols_max <= 128`, density `<=4`, `mmq_x_best <=64`). It only counts candidate calls for the next structural tile-policy A/B; no numeric branch is added. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`; Phase 32 n128 trace found 4096 candidates, mostly `mmq_x_best=64/48`) |
|
||||
| 0059 | **Gate MoE small-M MMQ tile policy** - adds default-off `LLAMA_MOE_SMALL_M_TILE=<n>` to cap only classified small-M MoE grouped-MMQ calls. This was used to A/B vLLM-like smaller M blocks without changing default inference. | yes (default-off, tile16, tile8, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`; Phase 33 rejected tile16 and tile8 as slower) |
|
||||
| 0060 | **Trace MoE MMID dispatch routes** - adds default-off `LLAMA_MOE_MMID_ROUTE_TRACE=<n>` around `MUL_MAT_ID` dispatch, classifying each call as `mmvq`, `mmvf`, grouped `mmq`, `mmf`, or host-sync `fallback`. This is evidence-only instrumentation to resolve whether serving hits the per-expert host-sync fallback. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID` `806/806`; Phase 34 n128 trace found `mmq=2776`, `mmvq=1320`, `host_sync=0/4096`) |
|
||||
| 0061 | **Trace regular MUL_MAT dispatch routes** - adds default-off `LLAMA_MUL_MAT_ROUTE_TRACE=<n>` around regular `MUL_MAT`, classifying projection-heavy calls as `vec_f`, `mat_f`, `vec_q`, `mmq`, `batched_cublas`, `op_*`, `fp4_prefill`, or `fwht`. This is evidence-only instrumentation for the `bf16-proj` serving bucket. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT` `1146/1146`, `MUL_MAT_ID` `806/806`; Phase 35 n128 trace found BF16 routes `mat_f=2485`, `op_cublas=1330`) |
|
||||
| 0062 | **Trace cuBLAS subroutes** - adds default-off `LLAMA_CUBLAS_ROUTE_TRACE=<n>` around the generic cuBLAS `MUL_MAT` path, classifying calls as `nvfp4_bf16_tc`, `bf16_tc`, `f16_tc_32f`, `f16_tc_16f`, or `sgemm`. This is evidence-only instrumentation for the Phase 35 `op_cublas` bucket. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT` `1146/1146`, `MUL_MAT_ID` `806/806`; Phase 36 n128 trace found `bf16_tc=5681`, `sgemm=2511`) |
|
||||
| 0063 | **Trace cuBLAS tensor names** - extends `LLAMA_CUBLAS_ROUTE_TRACE=<n>` with `src0`, `src1`, and `dst` names so the `sgemm` bucket can be tied back to graph nodes. | yes (default-off, trace-enabled, and post-serving gates green: MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT` `1146/1146`, `MUL_MAT_ID` `806/806`; Phase 37 n128 trace identified `sgemm` as `ffn_gate_inp* -> ffn_moe_logits/shared_expert_gate`) |
|
||||
|
||||
> **Dropped: patch 0026 (hybrid per-head bf16 SSM state, `ssm_bf16_tau`).** Once
|
||||
> the decode fusions (0028 recurrent-state gather-fusion + 0029 block-table cache)
|
||||
> landed, the bf16-SSM lever bought nothing: a clean re-measurement forcing **all**
|
||||
> gated-DeltaNet heads to bf16 (`tau=100000`) gives **flat** decode (780.6 vs
|
||||
> 780.0 t/s) - the mode engages but adds zero throughput because it is subsumed by
|
||||
> the fusions. It was a precision trade (not bit-exact) plus extra bug surface and
|
||||
> CUDA template-instantiation compile cost with no benefit, so it was removed. See
|
||||
> section 5 ("rejected / flat levers") for the full record.
|
||||
|
||||
---
|
||||
|
||||
## 4. Benchmarks
|
||||
|
||||
Hardware: **GB10 / DGX Spark** (CUDA 13, sm_121). Models: dense
|
||||
**Qwen3.6-27B-NVFP4** and MoE **Qwen3.6-35B-A3B-NVFP4**. Metric: `decode_agg`
|
||||
S_TG (t/s) from `llama-batched-bench`, `-fa on -ngl 99`, `npp 128 / ntg 128`,
|
||||
swept over serving width `npl` in {8, 32, 64, 128}. Plots:
|
||||
[`qwen36_decode_overview.png`](docs/qwen36_decode_overview.png) (both models),
|
||||
[`qwen36_dense_decode_vs_npl.png`](docs/qwen36_dense_decode_vs_npl.png),
|
||||
[`qwen36_moe_decode_vs_npl.png`](docs/qwen36_moe_decode_vs_npl.png); raw data
|
||||
[`final_benchmark.csv`](docs/final_benchmark.csv).
|
||||
|
||||

|
||||
|
||||
> The plot above also shows a third "bf16-tau" llama curve. That was the opt-in
|
||||
> `ssm_bf16_tau` lever (patch 0026), since **dropped** - a clean re-measurement
|
||||
> showed it flat once the decode fusions landed (see section 5). The numbers below
|
||||
> use only **stock** vs **patched** vs **vLLM**.
|
||||
|
||||
> **What was re-measured (2026-06-27).** The two llama columns - **stock** and
|
||||
> **patched** - were re-measured this session on one consistent
|
||||
> `llama-batched-bench` harness. The **vLLM** column is the **prior-session
|
||||
> reference** (kept as-is, *not* re-run this session). Per-run peak
|
||||
> VRAM was *not* re-captured: the GB10's unified Grace-Blackwell LPDDR5x reports
|
||||
> `[N/A]` to `nvidia-smi --query-gpu=memory.used` and the bench does not print it
|
||||
> (the memory-advantage note below is the prior-session finding).
|
||||
|
||||
### (a) + (b) Patched vs stock vs vLLM
|
||||
|
||||
The **stock** column is a separate, unpatched llama.cpp built at this backend's
|
||||
**exact pin (`9d5d882d`)**; the **patched** column is
|
||||
the paged binary, env/flag-toggled (`LLAMA_KV_PAGED=1`, plus
|
||||
`LLAMA_MOE_FORCE_GRAPHS=1` for MoE). Both
|
||||
run on the **same harness**, so "x over stock" is an apples-to-apples measure of
|
||||
the patch series. (Note: the patch series' dominant SSM decode fusions are
|
||||
compiled in, not env-gated - toggling `LLAMA_KV_PAGED` alone on the *patched*
|
||||
binary does **not** reproduce stock; only the separately-built unpatched
|
||||
`9d5d882d` binary does.) The **vLLM** column is a **different harness** (vLLM
|
||||
server + client continuous batching) and a **prior-session reference**, so the
|
||||
cross-engine "% of vLLM" is **indicative, not apples-to-apples**.
|
||||
|
||||
**Dense Qwen3.6-27B-NVFP4** (decode t/s):
|
||||
|
||||
| npl | stock | patched | vLLM (prior) | patched x over stock |
|
||||
|----:|------:|--------:|-------------:|---------------------:|
|
||||
| 8 | 68.3 | 85.3 | 70.4 | 1.25x |
|
||||
| 32 | 119.9 | 211.9 | 211.8 | 1.77x |
|
||||
| 64 | 142.8 | 305.2 | 309.1 | 2.14x |
|
||||
| 128 | 155.1 | 382.1 | 418.8 | 2.46x |
|
||||
|
||||
Dense **patched** is parity-to-ahead of vLLM (121 / 100 / 99 / 91% of vLLM across
|
||||
the widths).
|
||||
|
||||
**MoE Qwen3.6-35B-A3B-NVFP4** (decode t/s):
|
||||
|
||||
| npl | stock | patched | vLLM (prior) | patched x over stock |
|
||||
|----:|------:|--------:|-------------:|---------------------:|
|
||||
| 8 | 186.7 | 230.3 | 256.5 | 1.23x |
|
||||
| 32 | 267.4 | 466.4 | 500.8 | 1.74x |
|
||||
| 64 | 320.5 | 622.4 | 686.1 | 1.94x |
|
||||
| 128 | 347.2 | 784.3 | 882.2 | 2.26x |
|
||||
|
||||
MoE **patched** is 90 / 93 / 91 / 89% of vLLM.
|
||||
|
||||
**Caveat on the vLLM column.** It is a **different harness** and a
|
||||
**prior-session** measurement (not re-run this session), so the cross-engine "% of
|
||||
vLLM" is **indicative, not apples-to-apples**. Memory (prior session): llama uses
|
||||
**1.5-3x lower** memory than vLLM.
|
||||
|
||||
**Takeaway.** Re-measured this session, the patch series gives up to **2.46x
|
||||
(dense) / 2.26x (MoE)** over true-stock `9d5d882d` on the same harness (close to,
|
||||
slightly below, the prior 2.59x / 2.33x - llama was re-measured, vLLM kept).
|
||||
Dense is parity-to-ahead of vLLM; MoE **patched** sits at ~89-93% of the
|
||||
prior-session vLLM. The residual MoE gap is structural (see section 5).
|
||||
|
||||
### (c) Apple Silicon (M4, 16GB Metal) - does the patchset help here?
|
||||
|
||||
Short answer: **no - the wins are CUDA/Blackwell-specific.** Two facts first: the
|
||||
24GB NVFP4 GGUF doesn't fit a 16GB M4 (SSD paging), and on Metal `supports_op`
|
||||
**excludes NVFP4** from `MUL_MAT`/`MUL_MAT_ID`/`GET_ROWS` (FP4 matmuls fall back to
|
||||
CPU - no Apple FP4-MMA). So NVFP4 Qwen3.6 is not a Mac fit; a Metal-native Q4_K is.
|
||||
|
||||
Measured **stock vs patched** (same pin `c299a92c`, both built `-DGGML_METAL=ON`;
|
||||
the 28-patch series **compiles clean on Metal** - the CUDA code is `#if`-guarded),
|
||||
on **Qwen3-8B Q4_K_M** (a dense GQA model that fits 16GB and exercises the *live*
|
||||
Metal features; no Qwen3.6 hybrid GGUF fits 16GB, and the GDN fusions gate off on
|
||||
Metal anyway), `llama-bench` pp512/tg128 t/s:
|
||||
|
||||
| config | pp512 | tg128 |
|
||||
|---|---:|---:|
|
||||
| stock | 226.7 | 20.4 |
|
||||
| patched, paged **off** | 226.7 | 20.3 (= stock) |
|
||||
| patched, paged **on** | 222.6 | 19.8 (~0.97x) |
|
||||
|
||||
Concurrency (`batched-bench`) scales identically to stock (S_TG ~20 -> ~137 at
|
||||
npl32, from llama.cpp's existing batching). **Verdict: neutral-to-slightly-negative
|
||||
on Metal.** Patched-paged-off equals stock; turning paged on is ~0-3% slower
|
||||
decode / ~2-8% slower prefill, because the in-kernel block-table flash-attn read
|
||||
that *recovers* the gather cost is CUDA-only (`fattn-*.cuh`) - on Metal the paged
|
||||
path falls back to a host-side gather, pure overhead over stock's contiguous read.
|
||||
Everything Blackwell-specific (NVFP4, GDN fusions via 0030, occupancy) is inert.
|
||||
So **on Apple Silicon, prefer the stock `llama-cpp` backend.**
|
||||
|
||||
**Vulkan / SYCL** (source analysis): the gated-DeltaNet and SSM_CONV ops DO have
|
||||
upstream kernels on Vulkan and SYCL (as on Metal), so the Qwen3.6 hybrids RUN on
|
||||
all three via the non-fused path. The patchset's fusions are gated off there
|
||||
(0030), so the outcome is the same neutral-to-slightly-negative as Metal - not
|
||||
"won't run". This backend therefore ships **CUDA-only** (where the fusions are
|
||||
live + verified); non-CUDA users should use the stock `llama-cpp` backend. See
|
||||
[`UPSTREAM_LAYER2_SCOPE.md`](docs/UPSTREAM_LAYER2_SCOPE.md) for what native non-CUDA
|
||||
fused kernels would take.
|
||||
|
||||
---
|
||||
|
||||
## 5. Dev notes - what we learned
|
||||
|
||||
**Bit-exact methodology.** Every bit-exact patch is gated two ways: (1) a greedy
|
||||
md5 gate - `llama-completion -m MODEL -ngl 99 -fa on -p "The capital of France
|
||||
is" -n 48 --temp 0 --seed 1 | md5sum`, paged paths prefixed with
|
||||
`LLAMA_KV_PAGED=1` (+ `LLAMA_MOE_FORCE_GRAPHS=1` for paged MoE), on the default
|
||||
chat-template path; and (2) `test-backend-ops` (CUDA0 vs CPU oracle) for every
|
||||
touched op (`SSM_CONV*`, `GATED_DELTA_NET`, `MUL_MAT`, `MUL_MAT_ID`).
|
||||
For DGX work, `paged-inference-gates.sh` runs the canonical MoE/dense transcript
|
||||
md5 checks and selected `test-backend-ops` filters, and refuses to start while
|
||||
docker, `local-ai-worker`, GPU compute processes, or a non-free GPU lock are
|
||||
present.
|
||||
|
||||
For direct `llama-server` MTP serving A/B work, use
|
||||
`paged-mtp-serving-bench.sh`. It runs the same pre/post inference gates, compares
|
||||
baseline vs `--spec-type draft-mtp`, and captures the h2h client summaries plus
|
||||
MTP acceptance lines. Phase 15 rejected current MTP serving on GB10 despite
|
||||
passing safety gates; do not enable it by default.
|
||||
|
||||
**The gate is per-path** (see [`PAGED_BITEXACT_NOTE.md`](docs/PAGED_BITEXACT_NOTE.md)).
|
||||
Dense is bit-exact across paged/non-paged (`5951a5b4`). The **paged MoE** md5
|
||||
(`8cb0ce23`) does **not** byte-match the **non-paged MoE** md5 (`07db32c2`); this
|
||||
is a benign FP-accumulation-order difference of the paged attention reduction,
|
||||
**KL-validated** against the f16 reference: KLD(paged||f16) 0.13600 <=
|
||||
KLD(nonpaged||f16) 0.13660, PPL within +/-0.29, ~zero probability bias - two
|
||||
equivalent FP-reorderings of the same quantized model, not a regression. Future
|
||||
paged-MoE regressions therefore compare to `8cb0ce23`, not `07db32c2`.
|
||||
|
||||
**MoE-parity conclusion** (the residual gap is structural). The two heaviest MoE
|
||||
decode kernels - the GDN-SSM recurrence and the NVFP4-expert GEMM - are llama
|
||||
**wins** after this series (the recurrence runs at 102.6% of vLLM's bandwidth;
|
||||
the GEMM ties vLLM at the LPDDR5x BW floor). The residual gap is **bf16-projection
|
||||
bandwidth + the host scheduling loop**, both at the LPDDR5x floor - not a kernel
|
||||
llama is losing. The MoE GEMM kernel is *not* where the gap lives.
|
||||
|
||||
**Rejected / flat levers** (recorded so they are not re-tried):
|
||||
|
||||
- **Lever 2 - graph/stream coverage: FLAT.** Bit-exact graph coverage was
|
||||
exhausted by 0025; more graph/stream overlap is a no-op or small regression on
|
||||
this model.
|
||||
- **D1 premise "static decode is host-sync-bound on the MoE-routing readback":
|
||||
REFUTED.** The hypothesis was that the dominant decode cost is the device->host
|
||||
readback of MoE routing before launching the per-expert GEMMs (mul_mat_id's
|
||||
per-expert host-loop fallback). Profiling (GB10, q36-35b-a3b-nvfp4, batched-bench
|
||||
npl128) shows the opposite: on NVFP4 the grouped stream-k MMQ id-path is what
|
||||
runs (routing stays device-side), so the host-loop fallback is **never hit** -
|
||||
`cudaStreamSynchronize` count is *identical* with CUDA graphs on vs off (1457
|
||||
either way; only the kernel-launch count changes, ~100k vs ~229k). Steady-decode
|
||||
GPU-busy is **~99%** (1% idle), i.e. static decode is GPU-bound, not idle waiting
|
||||
on a sync. The one actionable residual the profile surfaced - per-step host
|
||||
kernel **re-issue** when the step is not graph-captured - shipped as 0043
|
||||
(default-on full-step decode graph), worth +2.6% (npl128) to +5-13% (npl32). The
|
||||
larger continuous-serving host cost is the graph **rebuild** (0040/0041), and the
|
||||
irreducible floor is the per-step logits-D2H-before-sampling serial point - none
|
||||
of which is the MoE-routing readback.
|
||||
- **Lever 3 - act-quant fusion: FLAT.** The W4A4 act-quant tax is removable only
|
||||
by W4A16 (a precision change, rejected) or a structural kernel rewrite; no
|
||||
further bit-exact lever clears it. 0023 already banks the de-dup.
|
||||
- **Lever 4 - NVFP4 the bf16 GDN/attn projections: REJECTED (KL-gate fail).**
|
||||
Quantizing the projections to NVFP4 costs ~+6% PPL; vLLM deliberately keeps the
|
||||
same bf16 projections. No-ship.
|
||||
- **W4A16-Marlin MoE GEMM: REJECTED.** It would be a precision upgrade nobody
|
||||
needs bought with a ~5% slower kernel; both kernels are already at the BW floor.
|
||||
(The "the win was NVFP4-dense-quant, not the Marlin kernel" dense verdict
|
||||
carries over to MoE.)
|
||||
- **Chunked parallel-scan GDN prefill (patch 0031): the scalar-serial form was
|
||||
FLAT-to-SLOWER at C=16 - the tensor-core M5 form (patch 0047) is the win,
|
||||
now DEFAULT-ON under paged KV.** 0031 implements the upstream "faster pre-fill"
|
||||
TODO - the FLA-style chunked gated-delta-rule (intra-chunk delta rule solved in
|
||||
parallel via the UT-transform + forward substitution, inter-chunk recurrence
|
||||
over n_tokens/C steps), math validated equivalent (numpy f32 NMSE ~1e-13;
|
||||
`test-backend-ops` within the 1e-7 NMSE gate, a NEW per-path result). **But
|
||||
GB10's 99KB dynamic-smem opt-in forces C=16** (the 128x128 f32 state alone is
|
||||
64KB of the all-shared layout); the scalar-serial scan (`GDN_TC=0`) was then
|
||||
pinned to 1 block/SM with serial per-thread dk-reductions and measured **~761
|
||||
t/s chunked vs ~971 t/s sequential (~22% slower)**, grid-starved at low n_seqs.
|
||||
The lesson held: **at this head dim the win needs tensor cores, not just
|
||||
chunking.** Patch 0047 builds those tensor-core forms (KK/QK Gram = M2, KS/QS
|
||||
state-boundary 3xtf32 = M3, P*U output = M4, full form-T solve + state-update
|
||||
mma = M5, all `GDN_TC`-selected in one build) and ships **M5** as the default
|
||||
when `LLAMA_KV_PAGED` is set. It is an f32/tf32-only re-port: the bf16/hybrid
|
||||
dev-tree machinery (from the dropped 0026 ssm_bf16_tau) and the bf16 CONFIG-C
|
||||
(M8) plus register-resident M6/M7 variants are NOT part of this series. M5 is the
|
||||
variant that beats the (already 84.7%-of-peak) sequential scan while staying on
|
||||
the bit-exact gate: MoE prefill S_PP **+3.5% @npp512 (3x interleaved A/B), +17.7%
|
||||
@npp2048**; decode S_TG unchanged (the tuned `GDN_CHUNK_MIN=64` engage threshold
|
||||
is > 1, so the 1-tok decode steps never enter the chunked path - at
|
||||
`GDN_CHUNK_MIN=1` the chunked path swallows decode and collapses S_TG ~25%, the
|
||||
reason the threshold is the lever). Bit-exactness is per-path benign:
|
||||
`test-backend-ops` GATED_DELTA_NET is **94/94** vs CPU with M5 forced (incl.
|
||||
multi-chunk n_tokens up to 256); the greedy md5 default-on == M5-forced ==
|
||||
canonical on the short gate prompt (paged-MoE `8cb0ce23`, dense `5951a5b4`); on
|
||||
a long MoE prompt (where the default fires M5 at >=64 tokens) M5 and the
|
||||
sequential path agree word-for-word until **one** benign greedy token-flip
|
||||
("the User:" vs "the User's Request:"), the dense model not flipping at all -
|
||||
the textbook reduction-order flip greedy amplifies, NMSE-validated. The chunk
|
||||
geometry stays env-selectable (`GDN_TC`/`GDN_CHUNK_C`/`GDN_DV_TILE`) for further
|
||||
tuning; M5 is the shipped default because it wins without losing the canonical gate.
|
||||
- **GDN occupancy retune (patch 0022) was a decode win but an UNCONDITIONAL
|
||||
dense-prefill regression - now gated by scan length (patch 0046).** Patch
|
||||
0022's `(NUM_WARPS=16, COLS_PER_WARP=8)` column-fold of the GDN
|
||||
sequential-recurrence dispatch (`case 128`) raises per-warp memory-level
|
||||
parallelism on the short, wide DECODE scans (small `n_tokens`, large
|
||||
`n_seqs`) - the measured +11.1% dense decode win. Applied unconditionally it
|
||||
also hit the dense PREFILL path, where the scan is long and narrow: the launch
|
||||
`grid.z` collapses from `S_v/4 = 32` to `S_v/(16*8) = 1`, the SMs starve, and
|
||||
profiling attributed the whole ~-6% dense-prefill regression vs stock to
|
||||
`gated_delta_net` (+54% GPU time at the (16,8) geometry). Patch 0046 gates the
|
||||
geometry by per-call scan length: long scans (prefill,
|
||||
`n_tokens >= GDN_PREFILL_NTOK`, default 256) take stock's high-grid.z `(4,1)`
|
||||
geometry; short scans (decode) keep the `(16,8)` retune. That recovers dense
|
||||
prefill +7.2% back to stock parity while keeping the decode win, and it is
|
||||
bit-exact: patch 0022 already proved every selectable `{NUM_WARPS,
|
||||
COLS_PER_WARP}` variant is byte-identical (the sweep cannot change the md5), so
|
||||
switching geometry by scan length cannot move the greedy output. The explicit
|
||||
`GDN_NW`/`GDN_CPW` one-build %peak sweep still overrides (the gate yields when
|
||||
either is set), so the A/B harness is unchanged.
|
||||
|
||||
**Opt-in bf16-SSM fast mode - DROPPED (was patch 0026, `ssm_bf16_tau`).** The
|
||||
design premise - that bf16 KL error concentrates in long-memory heads and can be
|
||||
removed by keeping them f32 - was already shaky: the error scales with the bf16
|
||||
head *count* and saturates (~0.06 MeanKLD / ~91% same-top-p) far below any useful
|
||||
byte saving. The lever was then **removed entirely** once the decode fusions
|
||||
(0028 recurrent-state gather-fusion + 0029 block-table cache) landed: a clean
|
||||
re-measurement that forced **all** gated-DeltaNet heads to bf16 (`tau=100000`,
|
||||
the most aggressive setting) gave **flat** decode throughput - **780.6 vs 780.0
|
||||
t/s**. The mode engages but buys **zero** speed; the earlier "+12%" was subsumed
|
||||
by the fusions. So bf16-tau was a precision trade (not bit-exact) plus extra bug
|
||||
surface and CUDA template-instantiation compile cost with **no** offsetting
|
||||
benefit, and patch 0026 was dropped from the series. Lesson recorded so it is not
|
||||
re-tried: do not reintroduce a per-head SSM-precision lever - the bandwidth it
|
||||
targeted is already recovered by the gather-fusion + block-table cache.
|
||||
|
||||
---
|
||||
|
||||
## 6. Architecture and quant generality
|
||||
|
||||
(From the arch-generality and quant-generality audits.)
|
||||
|
||||
- **15 of 16 optimizations are quant-AGNOSTIC.** Only **0023** (NVFP4
|
||||
activation-quantize de-dup) is NVFP4-specific. The SSM/paged/MMQ optimizations
|
||||
help **any quant** of these models (the GDN recurrence, conv, gather and
|
||||
o_proj-MMQ levers operate on the f32 recurrent state and the routing layout,
|
||||
not on the weight dtype).
|
||||
- **Arch-safe to build everywhere.** NVFP4 use is Blackwell-gated and falls back
|
||||
to dequant on other hardware; the GB10-tuned occupancy params (0022) are
|
||||
perf-only and env-selectable (`GDN_NW` / `GDN_CPW`), so they never change
|
||||
correctness on other GPUs. Patch 0030 makes the fused-op emission CUDA-family +
|
||||
CPU only, so a non-CUDA paged build routes to the safe upstream non-fused path.
|
||||
|
||||
- **What generalizes beyond this backend (upstream candidates).** The *speedups*
|
||||
are CUDA/Blackwell-specific (which is why Metal/Vulkan don't benefit - section
|
||||
4c), but several *findings and ops* are portable and worth upstreaming:
|
||||
- The headline is hardware-independent: on hybrid gated-DeltaNet models, decode
|
||||
is bottlenecked by the recurrent-state **plumbing** (memcpy + gathers, ~67% of
|
||||
the step), not the weight GEMM. The fusions for it (in-place state 0018, gather
|
||||
0019/0028, conv 0021) are bit-exact and already have CPU reference kernels, so
|
||||
they would speed up Qwen3.6 / Qwen3-Next / any hybrid-SSM decode on **every**
|
||||
backend once the ggml ops gain the respective (Metal/Vulkan) kernels - the
|
||||
highest-value upstream contribution.
|
||||
- The o_proj GEMV->MMQ reshape (0020) is a model-graph fix (batch the projection
|
||||
to hit the GEMM path) - arch-agnostic in principle, trivial to upstream.
|
||||
- The paged KV + cross-request prefix sharing + decode-first scheduler align with
|
||||
llama.cpp's own in-progress KV / chunked-prefill work and could inform it.
|
||||
- The per-path bit-exact md5 gate + the weekly upstream-drift canary is a reusable
|
||||
maintenance pattern for any vendored-patch backend.
|
||||
|
||||
---
|
||||
|
||||
## 7. Pin + maintenance policy
|
||||
|
||||
- **Canonical source = the fork branch `mudler/llama.cpp:localai-paged`.** The
|
||||
vendored `patches/paged/*.patch` files are now generated (one `git format-patch`
|
||||
per commit) from that branch, which is the pin commit plus the paged patch
|
||||
commits in order, so there is no more hand-export drift between the dev tree and
|
||||
the shipped series.
|
||||
- **Pinned to llama.cpp `0ed235ea2c17a19fc8238668653946721ed136fd`** (kept == the stock `llama-cpp` pin). The pin
|
||||
is advanced **only** by the manual pin-sync process (this section):
|
||||
rebase the source-only patch series onto the new tip, rebuild on GPU, pass the
|
||||
bit-exact gate on every path (dense + MoE, paged + non-paged) plus
|
||||
`test-backend-ops`, **and confirm the full grpc-server build links on CI**.
|
||||
- **The pin must track the stock pin.** `grpc-server.cpp` is shared with the stock
|
||||
backend and tracks the stock pin, so a paged pin that diverges past an upstream
|
||||
server-API refactor breaks the grpc-server LINK even when the patches are
|
||||
bit-exact. A bump to `c299a92c` (23 commits ahead of stock) was greedy-md5
|
||||
bit-exact but failed to link (undefined `stream_*` server helpers introduced by
|
||||
the refactor), and was reverted to the then-current stock pin. The bit-exact gate alone does not
|
||||
catch this; only the full CI grpc-server build does.
|
||||
- **Decoupled from the nightly auto-bumper.** There is deliberately **no**
|
||||
`bump_deps.yaml` entry for this backend - a naive `LLAMA_VERSION` bump could
|
||||
silently shift the tree out from under the patches.
|
||||
- **Weekly canary.** [`.github/workflows/llama-cpp-paged-canary.yml`](../../../.github/workflows/llama-cpp-paged-canary.yml)
|
||||
(via [`.github/scripts/paged-canary-apply.sh`](../../../.github/scripts/paged-canary-apply.sh))
|
||||
tries the patch series against the latest upstream tip with the build's own
|
||||
strict `git apply`. **Red = upstream drifted past the series -> run a
|
||||
PIN_SYNC** (do not bump the pin blindly), following the policy in this section.
|
||||
|
||||
---
|
||||
|
||||
## 8. Models
|
||||
|
||||
> **Build coverage: CUDA-only.** This backend ships only the CUDA/cublas build
|
||||
> targets (cuda-12, cuda-13, and the nvidia-l4t arm64 cuda-12/cuda-13 Jetson
|
||||
> rows). There are no cpu / vulkan / sycl / hipblas / metal-darwin builds: the
|
||||
> patchset's wins are CUDA/Blackwell-specific (section 4c), so off-CUDA the
|
||||
> backend is neutral-to-negative and non-CUDA users should run the stock
|
||||
> `llama-cpp` backend instead. The `backend/index.yaml` meta-backend resolves
|
||||
> `default`/`nvidia` to a CUDA variant accordingly.
|
||||
|
||||
The benchmarked NVFP4 GGUFs are published and wired into the LocalAI gallery:
|
||||
|
||||
| Gallery entry | Weights (HuggingFace) | Notes |
|
||||
|---|---|---|
|
||||
| `qwen3.6-27b-nvfp4-paged` | [`mudler/Qwen3.6-27B-NVFP4-GGUF`](https://huggingface.co/mudler/Qwen3.6-27B-NVFP4-GGUF) | Dense, native Blackwell NVFP4 (FP4-MMA). |
|
||||
| `qwen3.6-35b-a3b-nvfp4-paged` | [`mudler/Qwen3.6-35B-A3B-NVFP4-GGUF`](https://huggingface.co/mudler/Qwen3.6-35B-A3B-NVFP4-GGUF) | MoE (256 experts, top-8), `file_type MOSTLY_NVFP4`. |
|
||||
|
||||
Both gallery entries set `backend: llama-cpp-localai-paged` and the paged serving config
|
||||
(`paged_kv:true`, `max_batch_tokens`, `kv_unified:false`, `parallel`,
|
||||
`flash_attention:on`, `context_size`). They are bit-exact. The full
|
||||
backend-split + gallery plan is in
|
||||
[`LOCALAI_LLAMACPP_BACKEND_PLAN.md`](docs/LOCALAI_LLAMACPP_BACKEND_PLAN.md).
|
||||
|
||||
---
|
||||
|
||||
## 9. vLLM parity - final state (CLOSED)
|
||||
|
||||
> 2026-07-01 follow-up: the investigation was reopened for MTP safety,
|
||||
> MTP-serving, graph-shape tracing, and a current-stack serving snapshot. Phases
|
||||
> 14-20 are recorded in
|
||||
> [`docs/GB10_PARITY_PHASE0_RESULTS.md`](docs/GB10_PARITY_PHASE0_RESULTS.md) and
|
||||
> [`docs/PARITY_HANDOFF.md`](docs/PARITY_HANDOFF.md). They did not change the
|
||||
> GB10 conclusion: MTP/scheduler shortcuts are rejected, and the latest clean
|
||||
> stack remains below vLLM serving parity.
|
||||
|
||||
The multi-week GB10 (DGX Spark, sm_121) vLLM-parity investigation is **closed**.
|
||||
The standing, never-re-litigate record - full benchmark, every lever and verdict,
|
||||
the structural floors, the parity verdict - is
|
||||
[`docs/VLLM_PARITY_FINAL.md`](docs/VLLM_PARITY_FINAL.md). Summary:
|
||||
|
||||
- **Where we are (GB10, Qwen3.6 NVFP4, vs vLLM 0.23.0).** Decode: dense is
|
||||
**ahead of vLLM at low concurrency (116.7% at N=8)** and both models are
|
||||
bandwidth-floored at **~56-68% of vLLM at high concurrency**. Prefill is
|
||||
**~36% (MoE) / ~43% (dense)** of vLLM. Memory: **1.5-3x lower** than vLLM
|
||||
(NVFP4-resident; vLLM's peak is a fixed ~109-112 GB 0.85-util reservation,
|
||||
paged grows with KV from ~50 GB). Output is bit-exact per-path
|
||||
(`5951a5b4` dense, `8cb0ce23` paged-MoE).
|
||||
- **Why the residual is a hardware ceiling, not missing work.** Decode kernels
|
||||
are already **5.4x more GPU-efficient per token** than vLLM's; the gap is the
|
||||
**LPDDR5x ~273 GB/s** floor. The prefill GEMM is **FP4-MMQ-optimal** (every
|
||||
alternative - 0033 dequant->cuBLAS, 0034 native FP4-MMA, 0035/Marlin W4A16,
|
||||
offline-repack and vLLM-verbatim Marlin - was rejected; bf16 TC peak is ~half
|
||||
FP4 peak, and vLLM itself runs a bf16-Marlin fallback on sm_121). The GDN
|
||||
chunked scan is at the tractable tensor-core win (**M5 tf32**, patch 0047);
|
||||
its residual is the **O(C^2) intra-chunk solve + serial recurrence** (occupancy
|
||||
and dtype proven not the bound: BV -1%, bf16-C64 -18.75%). The serving host
|
||||
loop is **closed** (~0-1% of the wall; padded-decode built + rejected).
|
||||
- **Shipped, bit-exact wins.** FP4-MMQ GEMM, M5 tensor-core GDN prefill (0047),
|
||||
fused residual+RMSNorm (0042), fused GatedRMSNorm+SiLU (0044), GDN-prefill
|
||||
geometry gate (0046), the SSM decode-fusion stack (0018-0022/0028, up to
|
||||
2.46x/2.26x over stock), decode-graph reuse (0040/0043), the memory advantage,
|
||||
and low-N decode lead.
|
||||
- **The path to parity is different hardware.** Datacenter Blackwell (HBM,
|
||||
native tcgen05/CUTLASS FP4) lifts the bandwidth floor and **restores exactly
|
||||
the vLLM advantages that lose on GB10** (FLA blocked-solve GDN, Marlin/CUTLASS
|
||||
grouped FP4, HBM-tuned full-cudagraph decode). Re-run the methodology on new
|
||||
silicon; do not reopen the GB10 levers.
|
||||
|
||||
Latest current-stack MoE serving snapshot (`PTOK=128`, `GEN=64`, current clean
|
||||
DGX mirror `f2521ab12`, artifact
|
||||
`/home/mudler/bench/phase26_audited_snapshot/20260701_053650`). This run
|
||||
includes `hardware.txt` and `gate_summary.tsv`; all pre/post gate rows are
|
||||
`ok`:
|
||||
|
||||
| n | paged decode_agg | vLLM decode_agg | paged/vLLM decode | paged agg | vLLM agg | paged/vLLM agg |
|
||||
|---|------------------|-----------------|-------------------|-----------|----------|----------------|
|
||||
| 8 | 230.8 | 283.2 | 81.5% | 170.6 | 241.6 | 70.6% |
|
||||
| 32 | 420.0 | 609.0 | 69.0% | 254.6 | 466.7 | 54.6% |
|
||||
| 128 | 673.4 | 1025.0 | 65.7% | 324.0 | 656.5 | 49.4% |
|
||||
|
||||
Use `paged-current-serving-snapshot.sh` for future current-stack GB10 serving
|
||||
snapshots. It targets the clean `~/llama-phase6-source` mirror, checks
|
||||
docker/`local-ai-worker`/GPU-idle state, uses the owner-file lock, runs pre/post
|
||||
inference gates, writes `hardware.txt`, emits `gate_summary.tsv`, and emits
|
||||
paged/vLLM ratios.
|
||||
`hardware.txt` records the GPU identity and hardware class so GB10/workstation
|
||||
Blackwell evidence is not confused with a future datacenter-Blackwell rerun.
|
||||
`gate_summary.tsv` records pre/post MoE md5, dense md5, and backend-op checks
|
||||
so an artifact proves inferencing gates without reading full logs.
|
||||
Do not use the stale DGX
|
||||
`~/bench/combined_definitive.sh` without first porting it to the current mirror
|
||||
and lock discipline.
|
||||
|
||||
Phase 28 challenged the remaining low-conflict NVFP4 grouped-MMQ occupancy
|
||||
knobs on the same DGX mirror
|
||||
(`/home/mudler/bench/phase28_mmq_occupancy/20260701_040450`). The only buildable
|
||||
variant, `GGML_CUDA_FP4_MINBLOCKS=2`, was inference-safe before and after
|
||||
serving (MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID 806/806`) but regressed
|
||||
n128 decode serving (`705.1 -> 689.9` decode_agg_tps, `0.9784x`). The row-tile
|
||||
knob `GGML_CUDA_FP4_MMQ_Y=64` failed the NVFP4 writeback compile-time
|
||||
invariant. Do not promote these knobs; grouped-MMQ parity work now requires a
|
||||
structural kernel change, not launch-bounds or row-tile tweaks.
|
||||
|
||||
Phase 29 added the default-off grouped-MMQ shape trace as patch `0056`
|
||||
(`/home/mudler/bench/phase29_mmq_shape_trace/20260701_042428`). The helper was
|
||||
added test-first (`test-cuda-mmq-shape-trace`), compiled under CUDA on DGX, and
|
||||
kept inference stable with the trace disabled and enabled:
|
||||
MoE `8cb0ce23`, dense `5951a5b4`, `MUL_MAT_ID 806/806`. Example trace line:
|
||||
`[LLAMA_MOE_MMQ_SHAPE] type=40 moe=1 ncols_dst=104 nchannels_x=256 ncols_max=13 n_active_est=104 density=1 mmq_x_max=128 mmq_x_lim=64 mmq_x_best=16 mmq_y=128 stream_k=1`.
|
||||
|
||||
Phase 31 extended that trace as patch `0057`
|
||||
(`/home/mudler/bench/phase31_mmq_launch_trace/20260701_064424`) with
|
||||
`[LLAMA_MOE_MMQ_LAUNCH]` lines from `launch_mul_mat_q`. Default-off,
|
||||
trace-enabled, and post-serving gates stayed stable: MoE `8cb0ce23`, dense
|
||||
`5951a5b4`, `MUL_MAT_ID 806/806`. The n128 serving trace showed decode-like
|
||||
`4800/4800` and prefill-like `4920/4920` launch lines with `fixup=0` and
|
||||
`stream_k_blocks == ntiles_dst`, rejecting a no-fixup/no-stream-k shortcut for
|
||||
this workload.
|
||||
|
||||
Phase 32 added the small-M classifier trace as patch `0058`
|
||||
(`/home/mudler/bench/phase32_small_m_classifier/20260701_070127`). Default-off,
|
||||
trace-enabled, and post-serving gates stayed stable: MoE `8cb0ce23`, dense
|
||||
`5951a5b4`, `MUL_MAT_ID 806/806`. The n128 serving trace found 4096 small-M
|
||||
candidate calls: `mmq_x_best=64` 1800, `48` 1096, `40` 360, `32` 360, `16`
|
||||
360, `24` 120. This justifies Phase 33 as a default-off tile-policy A/B
|
||||
(`mmq_x=16`, possibly `8`) rather than a broad kernel rewrite.
|
||||
|
||||
Phase 33 added default-off `LLAMA_MOE_SMALL_M_TILE=<n>` as patch `0059`
|
||||
(`/home/mudler/bench/phase33_small_m_tile_policy/20260701_071136`). The knob is
|
||||
md5/op safe, but both tested values were slower in same-session n128 serving:
|
||||
baseline `672.1` decode_agg_tps, tile16 `640.3` (`0.953x`), tile8 `583.2`
|
||||
(`0.868x`). Do not promote simple smaller `mmq_x` caps for this workload.
|
||||
|
||||
Phase 34 added default-off `LLAMA_MOE_MMID_ROUTE_TRACE=<n>` as patch `0060`
|
||||
(`/home/mudler/bench/phase34_mmid_route_trace/20260701_072737`). Default-off,
|
||||
trace-enabled, and post-serving gates stayed stable: MoE `8cb0ce23`, dense
|
||||
`5951a5b4`, `MUL_MAT_ID 806/806`. Live n128 serving with trace cap 4096 produced
|
||||
`mmq=2776`, `mmvq=1320`, and `host_sync=0/4096`; the top shapes were
|
||||
`mmq ne2=12` (1096), `mmq ne2=18` (480), and `mmvq ne2=8` (360). This refutes
|
||||
host-sync fallback as the current n128 `MUL_MAT_ID` problem; follow-up work should
|
||||
target grouped-MMQ small-M kernel partitioning or another measured bucket.
|
||||
|
||||
Phase 35 added default-off `LLAMA_MUL_MAT_ROUTE_TRACE=<n>` as patch `0061`
|
||||
(`/home/mudler/bench/phase35_mul_mat_route_trace/20260701_074359`). Default-off,
|
||||
trace-enabled, and post-serving gates stayed stable: MoE `8cb0ce23`, dense
|
||||
`5951a5b4`, `MUL_MAT 1146/1146`, `MUL_MAT_ID 806/806`. Live n128 serving with
|
||||
trace cap 8192 produced route counts: `mat_f=2888`, `op_cublas=2292`,
|
||||
`mmq=1328`, `vec_q=1214`, `vec_f=470`. BF16 (`type=30`) dominated the trace
|
||||
with `mat_f=2485` and `op_cublas=1330`; top BF16 shapes were `mat_f ne1=12`
|
||||
(775), `op_cublas ne1=18` (760), and `mat_f ne1=8` (570). Next projection work
|
||||
should trace or optimize the BF16 `op_cublas`/`mat_f` split, not batched cuBLAS.
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user