Files
LocalAI/docs/content/features
mudler-agentandEttore Di Giacinto 38a2aa45fe feat(gallery): add Nimble 9B and CLM decision models, bump vllm.cpp to a19294a9 (#12397)
* feat(gallery): add nimble-9b-vllm-cpp decision model

Add Bespoke Nimble 9B, converted for vllm.cpp and pinned to the weights
commit 52eead25 of mudler/Bespoke-Nimble-9B-vllm-cpp (HEAD only adds the
model card). It is a redistribution of bespokelabs/Bespoke-Nimble-9B
with the LoRA merged into Qwen3.5-9B; config.json names NimbleModel, so
no hf_overrides are needed.

The artifact sits under overrides, where the installer reads it. The
entry sets an 8192-token context, Nimble's own prompt limit, and a KV
pool of 1024 blocks of 32 tokens for 4 sequences (about 1 GiB at 32 KiB
per token for the 8 full-attention layers).

Installed with local-ai models install and served on CPU through the
vllm-cpp backend: the model card's billing request gives billing
(0.986), refund 0.998 and urgency 0.33. Peak resident memory was
18.4 GB, so the description asks for about 20 GB of free RAM.

List the entry in the decisions gallery table. CLM stays out of the
gallery: the pinned engine cannot load the published head layout.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

* feat(gallery): add clm-v0.1-8b-vllm-cpp and bump vllm.cpp for CLM

The CLM checkpoint on mudler/CLM-v0.1-8B-vllm-cpp stores the heads with
the reference's own tensor names (state_head.inp, hidden.N, norms.N,
out). The pinned vllm.cpp 96788348 still expects the old .0/.2/.4/.6
layout and refuses the load with "head.safetensors incomplete for
state_head". vllm.cpp a19294a9 matches the reference layout and adds the
converter that produced the upload, so move the pin there. The ABI stays
at v30.

Add the CLM entry, pinned to the weights commit 0d1903b1 (HEAD only adds
the model card), with a 4096-token context and a KV pool for 4 sequences
(about 2.25 GiB at 144 KiB per token for Qwen3-8B).

Installed with local-ai models install and served on CPU against a
libvllm built at a19294a9: the model card example (john works at google,
entity type) gives person 0.950, the same as the card. Peak resident
memory was 17.9 GB.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-02 09:41:31 +02:00
..