This branch opened with VLLM_CPP_MLX=on, justified by an A/B that measured the MLX provider at 1.88x to 2.19x against the native MSL GEMM. That measurement was correct when taken and is now stale: vllm.cpp's own Metal kernels have improved several-fold since, through mma prefill attention, a vectorised decode V accumulation, vectorised attention staging, a fused qk-norm-RoPE preamble and a simdgroup-per-row softmax. The native path MLX was compared against no longer exists. Re-measured on the same Apple M4, in the same binary, with the arms toggled by VT_OP_PROVIDER_DISABLE=mlx, on Qwen3-1.7B-bf16 warm at p=512 g=128: MLX provider ON prefill TTFT 1370 ms warm throughput 11.98 tok/s MLX provider OFF prefill TTFT 1400 ms warm throughput 22.06 tok/s Shipping the previous default would have halved Apple Silicon throughput. MLX's steel GEMM is still about 20% faster than ours in isolation, but the provider pays a per-op mx::eval synchronisation plus an output memcpy, because it cannot write into our buffer. Across prefill's roughly 112 GEMMs that overhead leaves a 2% gain; on decode, where the same synchronisation is paid once per matmul per token, it costs 46%. The option is kept for prefill-dominated workloads, where the margin is small but real. The README section is rewritten rather than patched: it previously presented the stale table as the reason for the default, so leaving it in place would have made the new default look arbitrary. Assisted-by: Claude Code:claude-opus-5 [ClaudeCode] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
vllm-cpp backend
LocalAI text-generation backend for vllm.cpp, the LocalAI-team C++20 port of vLLM (paged KV cache, continuous batching, safetensors + GGUF loading, CUDA / CPU / Metal / Vulkan) with no Python at inference time.
The backend dlopens the engine's stable C ABI (libvllm, include/vllm.h,
ABI v2) through purego:
Load->vllm_engine_load: accepts a.gguffile or a HF-style model directory (config.json+ safetensors).context_sizemaps tomax_model_len;options: ["block_size:<n>", "num_blocks:<n>", "max_num_seqs:<n>"]size the KV cache and scheduler admission.Predict->vllm_complete(blocking).PredictStream->vllm_complete_stream; concurrent gRPC requests batch continuously in the engine's shared AsyncLLM scheduler.- Chat / tool calling rides the SAME code path as the llama.cpp autoparser:
with
use_tokenizer_template: truethe backend implementsPredictRich/PredictStreamRichover the ABI v3 chat entry points (vllm_chat/vllm_chat_stream). The ENGINE applies the model's chat template (GGUFtokenizer.chat_templateortokenizer_config.json), decides when a tool call engages (tool_choice: autolowers to a LAZY structural-tag decode constraint;required/named force one), parses tool calls with its streaming Hermes-style parser, and the backend maps eachchat.completion.chunkontoChatDelta/ToolCallDeltaprotos. - Without structured messages the plain path applies:
PredictOptions.Grammar-> the ABI'sstructured_grammar(GBNF) for LocalAI's Go-side grammar-constrained tool calling; JSON-schema / regex / choice constraints are also exposed by the ABI.
Model config example:
name: qwen3-vllm
backend: vllm-cpp
context_size: 8192
parameters:
model: Qwen3-4B # model dir (safetensors) or .gguf file
options:
- max_num_seqs:16
Apple Silicon: the MLX GEMM provider (OFF by default)
BUILD_TYPE=metal can build vllm.cpp's optional MLX provider for the dense GEMM
(VLLM_CPP_MLX=on). It is OFF by default, because it is currently slower.
This branch originally shipped it ON, on the strength of an A/B that had MLX at 1.88-2.19x against the native MSL GEMM. That measurement was honest when taken and is now stale: vllm.cpp's Metal kernels have since improved several-fold (mma prefill attention, vectorised decode V accumulation, a fused qk-norm-RoPE preamble and more), so the native path no longer resembles the one MLX was compared against.
Re-measured on the same Apple M4, same binary, arms toggled with
VT_OP_PROVIDER_DISABLE=mlx, Qwen3-1.7B-bf16 warm at p=512 g=128:
| prefill TTFT | warm throughput | |
|---|---|---|
| MLX provider ON | 1370 ms | 11.98 tok/s |
| MLX provider OFF | 1400 ms | 22.06 tok/s |
MLX's steel GEMM is still ~20% faster than ours in isolation, but the provider
pays a per-op mx::eval synchronisation plus an output memcpy (it cannot write
into our buffer). On prefill's ~112 GEMMs that overhead leaves +2%; on decode,
where the same sync is paid once per matmul per token, it costs 46%.
Turning it on is therefore only sensible for prefill-dominated workloads, and
even then the margin is small. Full disposition in vllm.cpp docs/BENCHMARKS.md,
"The MLX provider verdict".
Build knobs:
VLLM_CPP_MLX=onbuilds the provider in: ~19 MBlibmlx.dylibplus a ~105 MBmlx.metallib, and currently slower end to end. Off is the default.MLX_VERSIONpins the wheel (default0.29.3). MLX is consumed as the prebuilt pip wheel because building it from source needsxcrun metal, i.e. a full Xcode the macOS runners do not have.
Packaging vendors libmlx.dylib, mlx.metallib and MLX's MIT license into
package/lib/, and rewrites libvllm.dylib's rpath to @loader_path/lib
(re-signing it, since install_name_tool invalidates the signature). The
metallib must stay beside libmlx.dylib: MLX looks for it there.
Testing: make test runs the unit specs; export VLLM_CPP_MODEL=<model> (and
optionally VLLM_CPP_LIBRARY=<libvllm path>) to enable the e2e specs.