mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-02 19:14:38 -04:00
* chore(vllm-cpp): bump vllm.cpp to 967883486 (ABI v30) Moves the pin from c3bebc357 to 967883486. On top of the Nimble decision adapter and the Qwen3.5 vision-loader fix, this brings Tev1 on /v1/systemone and vllm_decide (opt-in through a "Tev1Model" architecture in config.json), a tokenizer/ subdirectory fallback so the Laya HF snapshot loads as downloaded, a stop-token fix, a logprobs fix under async scheduling and a pinned parakeet.cpp fetch for the diarization build. ABI v30 only adds the diarization and speaker-attributed ASR entry points; no existing struct or signature changed, so the purego mirrors keep their layout and only abiVersion moves to 30. Between 4479dc99f and 967883486 vllm.h changed only in a comment. v30 turns VLLM_CPP_WITH_DIARIZATION on by default. The fetch is pinned now, but ON still downloads parakeet.cpp at configure time and links a second ggml into libvllm for calls this backend never makes, so build with the option off: the symbols stay present as refusing stubs. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-sonnet-5-5 * feat(vllm-cpp): add the hf_overrides engine arg vLLM parity: engine_args.hf_overrides is a JSON object of top-level config.json keys merged over the model directory's own config.json. The main use is opting a published checkpoint into an engine adapter its config does not name, such as {"architectures": ["Tev1Model"]} on the Tev1 snapshots, which declare Qwen3_5ForConditionalGeneration. The C ABI has no override input and the engine reads config.json from the directory it is given, so Load builds a private overlay directory: the merged config.json plus a symlink to every other entry of the model directory, and passes that to the engine. The download is never written. Free, a failed load and the next Load remove the overlay. validModelPath and the DFlash draft resolution still see the real directory. A value that is not a JSON object, a .gguf model or a directory without config.json fails the load instead of being skipped like an unknown engine_args key, because loading the unmodified config would serve a different architecture than the one configured. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-sonnet-5-5 * fix(gallery): nest vllm-cpp artifacts under overrides artifacts: is a model-config key, and the installer reads model-config keys only from overrides:. Five vllm-cpp entries (laya, gliner25-decide, qwen3-vl-4b, cua-s1-forms and gliner2.5) declared it at the entry top level, where it is silently dropped: the install reports success, writes a config whose model is the bare HF repo id and downloads nothing, and vllm-cpp (which does not infer artifacts) then fails the first load with "model path not found". Move each block under overrides:, and add a guard test that refuses a top-level artifacts: key in the index. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-sonnet-5-5 * feat(gallery): add Tev1 4B and 0.8B on vllm-cpp Two decisions entries for Together AI's Tev1 checkpoints, pinned to the current HF revisions. Tev1 is autoregressive: vllm.cpp answers /v1/systemone by scoring the option letters, and the same engine still serves chat completions. The published config.json names Qwen3_5ForConditionalGeneration, so each entry sets hf_overrides: {architectures: [Tev1Model]} to enable the decision route without editing the download. known_usecases is [decisions] only, since a declared decisions list is authoritative for reservation. The descriptions state what was checked: agreement with transformers on CPU over seven questions (4B 7/7, max probability difference 0.0004; 0.8B 6/7 with one near tie), CPU-only for the decision route, and a fine-tune license the model card says is still being finalized, so no license key is set. The Decisions API page lists both entries, drops the note that Tev1 does not serve /v1/systemone and documents the 24-option limit (Ollama allows 26). Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-sonnet-5-5 --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
220 lines
9.7 KiB
Markdown
220 lines
9.7 KiB
Markdown
+++
|
|
disableToc = false
|
|
title = "vllm.cpp backend"
|
|
weight = 16
|
|
url = "/features/vllm-cpp/"
|
|
+++
|
|
|
|
[vllm.cpp](https://github.com/mudler/vllm.cpp) is the LocalAI team's C++20 port of
|
|
vLLM: the same continuous-batching scheduler, paged KV cache and automatic prefix
|
|
caching, with no Python at inference time. LocalAI serves it through the native
|
|
`vllm-cpp` backend, which loads either a HuggingFace safetensors directory or a
|
|
`.gguf` file and applies the chat template, tool-call parsing and reasoning split
|
|
inside the engine.
|
|
|
|
This page covers installing the backend and the models LocalAI ships ready to run
|
|
on it. For the full `engine_args` reference (KV sizing, scheduling policy,
|
|
speculative decoding, LMCache), see the
|
|
[vllm.cpp section of the text generation guide]({{% relref "features/text-generation" %}}#vllmcpp).
|
|
|
|
## Installing
|
|
|
|
```bash
|
|
local-ai backends install vllm-cpp
|
|
```
|
|
|
|
Or install it from the **Backends** page in the web UI. Images are published for
|
|
CPU, CUDA 13, Vulkan, Metal and Jetson L4T.
|
|
|
|
### Which GPUs the CUDA images cover
|
|
|
|
The CUDA images are currently built for **Blackwell-family architectures only**:
|
|
`sm_120a` (RTX 50 series, RTX PRO 6000 Blackwell) and `sm_121a` (GB10 / DGX
|
|
Spark) on x86-64, and `sm_121a` alone on arm64. CUDA 13 is required, so there is
|
|
no CUDA 12 variant.
|
|
|
|
That is narrower than vllm.cpp itself, which builds ten architectures. On an
|
|
Ampere, Ada, Hopper, Jetson Orin or Jetson Thor GPU the CUDA backend installs
|
|
successfully and then fails at the first request with `no kernel image is
|
|
available for execution on the device`. Until the build widens, use the Vulkan
|
|
or CPU image on those cards: `vulkan-vllm-cpp` builds with CUDA off entirely and
|
|
gives them a GPU path, without the NVFP4 and Marlin kernels.
|
|
|
|
## Ready-made models
|
|
|
|
The model gallery carries a curated set of vllm.cpp configurations. Each one
|
|
arrives with the engine settings already applied, so tool calling, the reasoning
|
|
split and speculative decoding work without hand-editing YAML.
|
|
|
|
| Gallery entry | Model | Size | Needs |
|
|
|---|---|---|---|
|
|
| `qwen3.6-27b-nvfp4-vllm-cpp` | Qwen3.6-27B, NVFP4 | 25 GB | Blackwell GPU |
|
|
| `qwen3.6-27b-nvfp4-mtp-vllm-cpp` | the same, with MTP speculative decoding | 25 GB | Blackwell GPU |
|
|
| `qwen3.6-27b-nvfp4-dflash-vllm-cpp` | the same, with DFlash speculative decoding | 28 GB | Blackwell GPU |
|
|
| `qwen3.6-35b-a3b-nvfp4-vllm-cpp` | Qwen3.6-35B-A3B MoE, NVFP4 | 23 GB | Blackwell GPU |
|
|
| `qwen3.6-35b-a3b-nvfp4-mtp-vllm-cpp` | the same, with MTP speculative decoding | 23 GB | Blackwell GPU |
|
|
| `qwen3-coder-30b-a3b-vllm-cpp` | Qwen3-Coder-30B-A3B, bf16 | 57 GB | Blackwell GPU, or CPU |
|
|
| `qwen3-4b-vllm-cpp` | Qwen3-4B, bf16 | 8 GB | CPU, Metal, Vulkan, Blackwell GPU |
|
|
| `qwen3-0.6b-vllm-cpp` | Qwen3-0.6B, bf16 | 1.4 GB | CPU, Metal, Vulkan, Blackwell GPU |
|
|
|
|
```bash
|
|
local-ai models install qwen3-0.6b-vllm-cpp
|
|
```
|
|
|
|
The two small bf16 entries are the ones that run anywhere the backend does,
|
|
including CPU. The NVFP4 entries need a Blackwell-class NVIDIA GPU on two counts:
|
|
NVFP4 has no kernel on older architectures, and the CUDA images are built only
|
|
for Blackwell in any case.
|
|
|
|
Sizing note: every entry sets `num_blocks` to give roughly one to four full
|
|
contexts of KV cache, which is a starting point rather than a tuned value. KV is
|
|
not free; the 4B entry, for instance, spends 144 KiB per token, so its 1024
|
|
blocks are about 4.5 GB on top of the weights. Raise `num_blocks` for more
|
|
concurrency, lower it on a small box.
|
|
|
|
### Why the 27B entries pin a revision
|
|
|
|
The Qwen3.6-27B entries pin their weights to a specific HuggingFace commit rather
|
|
than tracking the repository's default branch. This is deliberate and worth
|
|
understanding before you copy one of these configs.
|
|
|
|
The upstream repository was later re-quantized in place, under the same name, from
|
|
NVFP4 to FP8 W8A8. A config that names the repository without a revision therefore
|
|
resolves to entirely different weights, with different numerics and different
|
|
performance, and nothing about the load reports that anything changed. Pinning is
|
|
what makes the entry reproducible:
|
|
|
|
```yaml
|
|
artifacts:
|
|
- name: model
|
|
target: model
|
|
source:
|
|
type: huggingface
|
|
repo: unsloth/Qwen3.6-27B-NVFP4
|
|
revision: 890bdef7a42feba6d83b6e17a03315c694112f2a
|
|
```
|
|
|
|
The same reasoning applies to any quantized community repository you depend on.
|
|
|
|
### Choosing between the speculative variants
|
|
|
|
Speculative decoding trades memory for decode throughput. All three Qwen3.6-27B
|
|
entries serve the same weights and produce the same quality; they differ only in
|
|
how tokens are proposed.
|
|
|
|
| Entry | Method | Extra weights | Extra memory |
|
|
|---|---|---|---|
|
|
| `qwen3.6-27b-nvfp4-vllm-cpp` | none | none | none |
|
|
| `qwen3.6-27b-nvfp4-mtp-vllm-cpp` | MTP, depth 1 | none, the draft head ships inside the checkpoint | about 3.6 GB |
|
|
| `qwen3.6-27b-nvfp4-dflash-vllm-cpp` | DFlash, 16-token blocks | a separate 3.5 GB drafter | drafter plus draft cache |
|
|
|
|
MTP drafts one token per step from a head that already lives in the target
|
|
checkpoint's own `mtp.*` tensors, so it costs no extra download. DFlash drafts a
|
|
whole 16-token block in one non-autoregressive pass from a separate drafter, which
|
|
is the larger win at the cost of a second checkpoint on disk.
|
|
|
|
Start with the plain entry if you are short on memory, and with the DFlash entry
|
|
if you are not.
|
|
|
|
### Tool calling
|
|
|
|
Every entry above sets `use_tokenizer_template: true` and disables LocalAI's
|
|
Go-side grammar path, so tool calls are detected and parsed by the engine's own
|
|
streaming parsers and arrive as real `tool_calls` on the OpenAI response.
|
|
|
|
The parser is normally auto-detected from the chat template, but one case cannot
|
|
be: Qwen3-Coder's tool dialect is byte-identical on the wire to another family's,
|
|
so template sniffing would pick the wrong parser. The
|
|
`qwen3-coder-30b-a3b-vllm-cpp` entry therefore names it explicitly, and any
|
|
Qwen3-Coder config you write yourself should do the same:
|
|
|
|
```yaml
|
|
engine_args:
|
|
tool_parser: qwen3_coder
|
|
```
|
|
|
|
## Overriding config.json keys (`hf_overrides`)
|
|
|
|
`engine_args.hf_overrides` is a JSON object of top-level `config.json` keys that
|
|
the backend merges over the model directory's own `config.json` before the
|
|
engine loads it, like vLLM's `--hf-overrides`. The main use is to opt a published
|
|
checkpoint into an engine adapter that its config does not name. For example,
|
|
the Tev1 repositories declare `Qwen3_5ForConditionalGeneration`, and vllm.cpp
|
|
serves them as decision models only when the architecture is `Tev1Model`:
|
|
|
|
```yaml
|
|
engine_args:
|
|
hf_overrides:
|
|
architectures: ["Tev1Model"]
|
|
```
|
|
|
|
The downloaded model files do not change. At load the backend creates a private
|
|
temporary directory. It writes the merged `config.json` there and adds a
|
|
symlink for each other entry of the model directory (weights, tokenizer files,
|
|
a `tokenizer/` subdirectory). Then it gives that directory to the engine. When
|
|
the model unloads, the backend removes the directory.
|
|
|
|
Rules:
|
|
|
|
- The merge is top-level only. An override key replaces the whole value of that
|
|
key, including a nested object such as `text_config`.
|
|
- The model must be a directory that contains a `config.json`. A `.gguf` file
|
|
or a directory without `config.json` fails the load.
|
|
- A value that is not a JSON object (an array, a scalar, or JSON that does not
|
|
parse) fails the load. The backend does not ignore it, because loading the
|
|
unchanged config would serve a different architecture than the one you
|
|
configured.
|
|
- An empty object (`{}`) does nothing.
|
|
|
|
## Named entity recognition (GLiNER2.5)
|
|
|
|
The `vllm-cpp` backend serves [GLiNER2.5](https://huggingface.co/fastino/gliner2.5-multi-v1),
|
|
a zero-shot NER and structured-extraction model. Point the backend at the
|
|
safetensors directory and the backend exposes the `TokenClassify` gRPC method,
|
|
which LocalAI maps to its standard NER API surface.
|
|
|
|
Labels are supplied at inference time, not baked into the model config. Set
|
|
them in `engine_args`:
|
|
|
|
```yaml
|
|
engine_args:
|
|
ner_labels: "person,organization,location,date,time,money,quantity"
|
|
ner_threshold: 0.5
|
|
ner_max_width: 12
|
|
```
|
|
|
|
`ner_labels` is a comma-separated list. When omitted, the backend falls back to
|
|
a built-in default set (`person`, `organization`, `location`, `date`, `time`,
|
|
`money`, `quantity`). `ner_threshold` is the sigmoid cutoff (default 0.5);
|
|
`ner_max_width` is the maximum span length in tokens (default 12).
|
|
|
|
The model runs the DeBERTa v2 encoder with disentangled attention on the host
|
|
forward, which is the required contract for pooling models in vllm.cpp. A
|
|
device-resident forward is tracked as a performance optimization, not a
|
|
correctness gap.
|
|
|
|
### Decisions API
|
|
|
|
The `vllm-cpp` backend serves the kev-compatible SystemOne endpoints (the Decisions API): typed
|
|
`choice`, `noul` and `score` questions over a state text, answered by a
|
|
non-generative decision model in one pass. A decision model declares
|
|
`known_usecases: [decisions]`. See [Decisions API]({{% relref "features/decisions" %}})
|
|
for the request shape, the models you can install and the access rules.
|
|
|
|
| Endpoint | Method | Description |
|
|
|---|---|---|
|
|
| `/v1/systemone` | POST | Answer all questions in one pass |
|
|
| `/v1/systemone/permute` | POST | Re-run one choice question under n_perm option orders |
|
|
| `/v1/systemone/separate` | POST | Answer each question in its own pass (N passes) |
|
|
|
|
The GLiNER2.5 zero-shot NER model (`token_classify`) also serves
|
|
`/v1/systemone`, through the NER path, and it is the model to use for
|
|
`/v1/systemone/permute` and `/v1/systemone/separate`, which decision models
|
|
refuse with a `400`. It derives its NER labels from the question definitions, so
|
|
no `ner_labels` configuration is needed.
|
|
|
|
## Beyond text generation
|
|
|
|
The `vllm-cpp` backend also serves MiniMax-H3, which generates video and audio
|
|
jointly. See [Video generation]({{% relref "features/video-generation" %}}#minimax-h3-vllmcpp).
|