The Go bindings mirror vllm.h by hand and refuse a library whose vllm_abi_version differs from what they were written against. Two automated pin bumps (#11174, #11352) moved VLLM_CPP_VERSION onto engines declaring ABI v10 while govllmcpp.go still mirrored v5, so every vllm-cpp image built since then panics at startup on every platform: panic: vllm-cpp: ABI mismatch: library reports v10, backend built against v5 Grow both PODs to the v10 layout: vllm_model_params gains speculative_config, enable_prefix_caching, max_num_batched_tokens, scheduling_policy, kv_transfer_config and enable_jump_forward (88 bytes), vllm_sampling_params gains the v8 logits-processor pair (136 bytes). The offsets in the specs come from offsetof() against the pinned header. All of the new fields are inert when zeroed, so the engine behaves exactly as it did under v5; the backend sets none of them. Nothing cross-checked the two files, which is why a blind pin bump could ship a backend that cannot load. The library build now runs abi-check first: it compares VLLM_ABI_VERSION in the fetched header against abiVersion in govllmcpp.go and fails the build naming both, instead of leaving the mismatch for a user's runtime. Fixes #11379 Assisted-by: Claude:claude-fable-5 golangci-lint Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
4.6 KiB
vllm-cpp backend
LocalAI text-generation backend for vllm.cpp, the LocalAI-team C++20 port of vLLM (paged KV cache, continuous batching, safetensors + GGUF loading, CUDA / CPU / Metal / Vulkan) with no Python at inference time.
The backend dlopens the engine's stable C ABI (libvllm, include/vllm.h,
ABI v10) through purego:
Load->vllm_engine_load: accepts a.gguffile or a HF-style model directory (config.json+ safetensors).context_sizemaps tomax_model_len;options: ["block_size:<n>", "num_blocks:<n>", "max_num_seqs:<n>"]size the KV cache and scheduler admission.Predict->vllm_complete(blocking).PredictStream->vllm_complete_stream; concurrent gRPC requests batch continuously in the engine's shared AsyncLLM scheduler.- Chat / tool calling rides the SAME code path as the llama.cpp autoparser:
with
use_tokenizer_template: truethe backend implementsPredictRich/PredictStreamRichover the ABI v3 chat entry points (vllm_chat/vllm_chat_stream). The ENGINE applies the model's chat template (GGUFtokenizer.chat_templateortokenizer_config.json), decides when a tool call engages (tool_choice: autolowers to a LAZY structural-tag decode constraint;required/named force one), parses tool calls with its streaming Hermes-style parser, and the backend maps eachchat.completion.chunkontoChatDelta/ToolCallDeltaprotos. - Without structured messages the plain path applies:
PredictOptions.Grammar-> the ABI'sstructured_grammar(GBNF) for LocalAI's Go-side grammar-constrained tool calling; JSON-schema / regex / choice constraints are also exposed by the ABI.
The struct mirrors in govllmcpp.go are hand-written against one ABI version,
and the engine refuses to load against any other. Moving VLLM_CPP_VERSION in
the Makefile therefore means updating abiVersion plus the mirrors (and their
offsets in vllmcpp_test.go) in the same change; make abi-check compares the
pinned header against the bindings and the library build runs it first.
Model config example:
name: qwen3-vllm
backend: vllm-cpp
context_size: 8192
parameters:
model: Qwen3-4B # model dir (safetensors) or .gguf file
options:
- max_num_seqs:16
Apple Silicon: the MLX GEMM provider (ON by default, gated to prefill)
BUILD_TYPE=metal builds vllm.cpp's MLX provider for the dense GEMM
(VLLM_CPP_MLX=on, the default here). It is on because upstream now SHAPE-GATES
it to prefill; it was briefly off in this branch's history, and that was correct
at the time for an ungated provider.
The gate matters more than the flag. MLX's steel GEMM wins prefill but loses
decode, because the provider pays an mx::eval synchronisation plus an output
memcpy on every call and decode makes ~112 calls per token. Measured on an
Apple M4, Qwen3-1.7B-bf16 warm at p=512 g=128:
| configuration | prefill TTFT | warm throughput |
|---|---|---|
| MLX gated to prefill (pin >= 89c46aeb) | 524.5 ms | 24.37 tok/s, 97.6% of MLX-LM |
| MLX ungated (older pins) | 537 ms | 12.7 tok/s |
| MLX off | 602 ms | 23.9 tok/s, 95.9% |
Ratios are against an MLX-LM baseline measured INTERLEAVED with ours over four ABBA blocks (its spread 0.34%, ours 0.12%). An earlier revision of this file claimed 99.1%; that used a two-run MLX-LM baseline containing an outlier and overstated us by about 1.5 points.
VLLM_CPP_VERSION and this flag are coupled. Moving the pin back before
89c46aeb while leaving VLLM_CPP_MLX=on would take the middle row — roughly
half throughput. If you roll the pin back, roll the default back with it.
One caveat: MLX's GEMM is not bit-identical to the native kernel, so an MLX build
produces a different greedy sequence than a non-MLX one. That is a property of the
provider, not of the gate, and it predates this packaging. Full disposition in
vllm.cpp docs/BENCHMARKS.md.
Build knobs:
VLLM_CPP_MLX=offbuilds Metal without the provider: ~124 MB smaller, and 96.4% of MLX-LM instead of 99.1%.MLX_VERSIONpins the wheel (default0.29.4). MLX is consumed as the prebuilt pip wheel because building it from source needsxcrun metal, i.e. a full Xcode the macOS runners do not have.
Packaging vendors libmlx.dylib, mlx.metallib and MLX's MIT license into
package/lib/, and rewrites libvllm.dylib's rpath to @loader_path/lib
(re-signing it, since install_name_tool invalidates the signature). The
metallib must stay beside libmlx.dylib: MLX looks for it there.
Testing: make test runs the unit specs; export VLLM_CPP_MODEL=<model> (and
optionally VLLM_CPP_LIBRARY=<libvllm path>) to enable the e2e specs.