mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-30 10:04:32 -04:00
* feat(vllm-cpp): serve MiniMax-H3 video+audio generation
vllm.cpp's C ABI grew a video slice (ABI v12): a second engine handle
loaded from the MiniMax-H3 checkpoint SET, one blocking generate, and a
composed ffmpeg argv the caller execs. This wires that into LocalAI's
existing /video endpoint, so `vllm-cpp` now serves both text and video
and a clip comes back as an MP4 with a real audio track rather than a
silent render.
The video engine is a separate handle rather than a mode of the text
one because H3 is not a model directory: the DiT, the text encoder and
two VAEs are separate artifacts, and vllm.cpp has the two loaders refuse
each other's checkpoints. `Load` takes the video branch when the config
declares any of the video options; `parameters.model` is the DiT and the
rest of the set is named in `options:`.
Three details are worth calling out because getting them wrong is
expensive:
- The partition is DECLARED, not detected. The community quantisations
strip the release metadata and the FL2VA and Ref2VA DiTs are
byte-structurally identical, so the engine refuses to generate until
it is told which it has. Worse, a mismatch does not fail cleanly: a
reference passed to an FL2VA DiT renders for hours and returns a
coloured lattice over the frame. The backend refuses that combination
up front instead.
- ffmpeg comes from the host. libvllm writes frames plus a WAV and
composes the mux argv, then spawns nothing - that process boundary is
upstream's decision. The backend execs it, the same arrangement
vibevoice-cpp uses for transcoding, and ffmpeg also converts a
start_image upload into the binary PPM at the exact output canvas the
engine requires.
- It is slow. Roughly 176 s per denoise step at the default 1344x768
canvas on a 20-SM device, so the 50-step default is a multi-hour job.
Nothing on this path imposes a deadline.
The /video endpoint no longer forces 512x512 when the request omits the
geometry. Every video backend already supplies its own default for a
zero (512x512 for stablediffusion-ggml, 1280x720 for diffusers, 832x480
for longcat-video, 1344x768 for H3), so the hardcoded value only ever
overrode the model's trained canvas with one three of the four were
never trained at.
Moving the engine pin from ABI v10 to v16 also grows the text
vllm_model_params mirror by the v14 device field and the v16 KV-sizing
knobs. LocalAI sets none of them - 0 is the pre-v14 engine byte for byte
- but the struct SIZE is part of the layout contract, so leaving them
out would have vllm_engine_load read past the allocation.
Gallery: `minimax-h3-fl2va-q4` installs the Q4_K_M FL2VA set (~40 GB
across five weight files plus the two VAE configs that carry the latent
statistics).
Assisted-by: Claude:claude-opus-5 golangci-lint yamllint go-vet
* fix(vllm-cpp): unbreak the Darwin build at the new engine pin
src/capi/vllm_c.cpp opens one `extern "C" {` for the whole ABI surface,
so file-local helpers declared inside it inherit C linkage. The video
slice added one that returns std::string, which Apple Clang reports as
-Wreturn-type-c-linkage and vllm.cpp's target-local -Werror turns into a
build failure. GCC and upstream Clang do not diagnose it, so only the
metal-darwin-arm64 job saw it.
Suppress it the same way this Makefile already suppresses Apple Clang's
-Wgnu-folding-constant on the Metal build. The helper is never called
across the boundary so the warning describes no hazard here, but it is a
real upstream wart: the fix belongs in vllm.cpp, hoisting the helper
above the extern "C" block, and this flag should go when a pin carrying
that fix lands.
Assisted-by: Claude:claude-opus-5
* fix(vllm-cpp): patch the engine clone instead of the warning flag
The -Wno-return-type-c-linkage added in the previous commit does nothing.
vllm_cpp_set_warnings adds `-Wall -Wextra -Werror` as PRIVATE target
options, so they land after anything CMAKE_CXX_FLAGS contributes, and
-Wall re-enables the -Wreturn-type group that -Wreturn-type-c-linkage
belongs to. The darwin job failed again on the same line, which is the
evidence: a consumer cannot wave this off from outside the engine.
Position is the only fix, so carry it as a patch against the pinned SHA,
the way longcat-video patches its own upstream. It hoists the helper
above the `extern "C" {` that gives it C linkage; it is file-local and
never called across the boundary, so nothing else moves.
`git apply` is unguarded on purpose: a patch that stops applying must
fail the clone loudly, because the alternative is a pin that silently
ships without a fix it is documented to carry. The patch header names
what retires it - a pin carrying the fix upstream, where it belongs.
Verified by applying the patch with `git apply` to the exact blob at the
pinned SHA and diffing the result against the intended file.
Assisted-by: Claude:claude-opus-5
* chore(vllm-cpp): bump the engine pin to ABI v17 and drop the vendored OrEmpty patch
The OrEmpty linkage fix this backend carried as patches/0001-* landed upstream
(mudler/vllm.cpp#195, 7534da65), so the patch has done its job. It is deleted
rather than left in place: the Makefile applies patches/*.patch unguarded and
documents that "a patch that no longer applies must FAIL the clone", so keeping
it against fixed source would break the build the moment the pin moved. Bumping
the pin and deleting the patch therefore have to be the SAME change.
Pin f921062b -> 776c56f1 (current vllm.cpp main).
That range also carries the engine's ABI v17 (vllm_server_main: the OpenAI server
published on the public surface). registerLib compares the library's
vllm_abi_version against `abiVersion` for EXACT equality, so the constant moves
16 -> 17 in the same commit or every load fails with an ABI mismatch.
The bump is safe for the layout assertions in video_test.go: diffing include/vllm.h
across the two pins shows zero struct-field changes -- v17 adds one function
declaration, the version macro and a doc comment, nothing else -- so every
unsafe.Offsetof in the video params test still holds.
Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
* chore(vllm-cpp): re-pin to pick up the VLLM_CPP_SERVER=OFF link fix
The previous pin carried vllm.cpp's ABI v17 (vllm_server_main) but not the guard
that makes it link when the server is compiled out. This backend builds libvllm
with VLLM_CPP_SERVER off, so the darwin lane failed at the dylib link with
vllm::entrypoints::openai::VllmServerMain undefined.
Fixed upstream in mudler/vllm.cpp#202: the C entry point is now guarded, so the
symbol is still exported (ABI v17 stays resolvable for dlopen) while the
no-server arm reports the missing capability instead of dragging in a translation
unit that was never compiled.
Verified upstream in BOTH arms before re-pinning: SERVER=ON builds and runs, and
SERVER=OFF configures, links, produces libvllm.so, and `nm -D` shows
vllm_server_main exported next to vllm_video_generate and vllm_transcribe.
Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
168 lines
7.8 KiB
Markdown
168 lines
7.8 KiB
Markdown
# vllm-cpp backend
|
|
|
|
LocalAI backend for [vllm.cpp](https://github.com/mudler/vllm.cpp),
|
|
the LocalAI-team C++20 port of vLLM (paged KV cache, continuous batching,
|
|
safetensors + GGUF loading, CUDA / CPU / Metal / Vulkan) with no Python at
|
|
inference time.
|
|
|
|
It serves two things: text generation, and MiniMax-H3 joint video+audio
|
|
generation.
|
|
|
|
The backend dlopens the engine's stable C ABI (`libvllm`, `include/vllm.h`,
|
|
ABI v16) through purego:
|
|
|
|
- `Load` -> `vllm_engine_load`: accepts a `.gguf` file or a HF-style model
|
|
directory (`config.json` + safetensors). `context_size` maps to
|
|
`max_model_len`; `options: ["block_size:<n>", "num_blocks:<n>",
|
|
"max_num_seqs:<n>"]` size the KV cache and scheduler admission.
|
|
- `Predict` -> `vllm_complete` (blocking).
|
|
- `PredictStream` -> `vllm_complete_stream`; concurrent gRPC requests batch
|
|
continuously in the engine's shared AsyncLLM scheduler.
|
|
- Chat / tool calling rides the SAME code path as the llama.cpp autoparser:
|
|
with `use_tokenizer_template: true` the backend implements
|
|
`PredictRich`/`PredictStreamRich` over the ABI v3 chat entry points
|
|
(`vllm_chat` / `vllm_chat_stream`). The ENGINE applies the model's chat
|
|
template (GGUF `tokenizer.chat_template` or `tokenizer_config.json`),
|
|
decides when a tool call engages (`tool_choice: auto` lowers to a LAZY
|
|
structural-tag decode constraint; `required`/named force one), parses tool
|
|
calls with its streaming Hermes-style parser, and the backend maps each
|
|
`chat.completion.chunk` onto `ChatDelta`/`ToolCallDelta` protos.
|
|
- Without structured messages the plain path applies:
|
|
`PredictOptions.Grammar` -> the ABI's `structured_grammar` (GBNF) for
|
|
LocalAI's Go-side grammar-constrained tool calling; JSON-schema / regex /
|
|
choice constraints are also exposed by the ABI.
|
|
|
|
`patches/` carries fixes the pinned engine SHA does not have yet, applied to
|
|
the clone the same way `longcat-video` patches its upstream. `git apply` is
|
|
unguarded on purpose: a patch that stops applying must fail the clone loudly
|
|
rather than leave a pin silently missing a fix it is documented to carry. Each
|
|
patch header says what retires it.
|
|
|
|
The struct mirrors in `govllmcpp.go` are hand-written against one ABI version,
|
|
and the engine refuses to load against any other. Moving `VLLM_CPP_VERSION` in
|
|
the Makefile therefore means updating `abiVersion` plus the mirrors (and their
|
|
offsets in `vllmcpp_test.go`) in the same change; `make abi-check` compares the
|
|
pinned header against the bindings and the library build runs it first.
|
|
|
|
Model config example:
|
|
|
|
```yaml
|
|
name: qwen3-vllm
|
|
backend: vllm-cpp
|
|
context_size: 8192
|
|
parameters:
|
|
model: Qwen3-4B # model dir (safetensors) or .gguf file
|
|
options:
|
|
- max_num_seqs:16
|
|
```
|
|
|
|
## MiniMax-H3 video+audio generation
|
|
|
|
`GenerateVideo` -> `vllm_video_generate` (ABI v12). H3 renders picture and sound
|
|
together, so the output MP4 carries a real AAC track.
|
|
|
|
The video engine is a SECOND handle (`vllm_video_engine`), not a mode of the
|
|
text one, because H3 is a checkpoint SET rather than a model directory: the DiT,
|
|
the text encoder and two VAEs are separate artifacts, and vllm.cpp has the two
|
|
loaders refuse each other's checkpoints. `Load` takes the video branch when the
|
|
model config carries any of the video options below; `parameters.model` is the
|
|
DiT and everything else is named in `options:`.
|
|
|
|
```yaml
|
|
name: minimax-h3-fl2va-q4
|
|
backend: vllm-cpp
|
|
cuda: true
|
|
known_usecases: [video]
|
|
parameters:
|
|
model: minimax-h3/MiniMax-H3-FL2VA-Q4_K_M.gguf
|
|
options:
|
|
- video_encoder:minimax-h3/qwen3vl-32B-MiniMax-H3-Q4_K_M.gguf
|
|
- video_tokenizer:minimax-h3/tokenizer.json
|
|
- video_vae:minimax-h3/video_vae.safetensors
|
|
- video_vae_config:minimax-h3/video_vae_config.json
|
|
- audio_vae:minimax-h3/audio_vae.safetensors
|
|
- audio_vae_config:minimax-h3/audio_vae_config.json
|
|
- video_partition:fl2va
|
|
- video_device:cuda
|
|
- video_dequant_bf16:true
|
|
- video_width:1344
|
|
- video_height:768
|
|
- video_num_frames:124
|
|
```
|
|
|
|
Three things are worth knowing before touching this path.
|
|
|
|
**The partition is declared, not detected, and a mismatch does not fail
|
|
cleanly.** The FL2VA DiT serves `t2va` and `fl2va`; `ref2va` is a different
|
|
checkpoint. The community GGUF/NVFP4 quantisations strip the release metadata
|
|
and the two DiTs are byte-structurally identical, so the engine refuses every
|
|
generate until `video_partition` says which one it has. Handing reference
|
|
conditioning to an FL2VA DiT renders for hours and returns a coloured lattice
|
|
over the frame, so `checkPartitionConditioning` refuses that combination here,
|
|
before the engine is called.
|
|
|
|
**ffmpeg comes from the host.** libvllm writes the frames and the WAV and
|
|
COMPOSES the mux argv, then spawns nothing — that process boundary is upstream's
|
|
decision. `muxVideo` takes the composed argv, substitutes `argv[0]` with the
|
|
resolved binary and execs it; the backend image is `FROM scratch` and carries no
|
|
ffmpeg, the same arrangement `vibevoice-cpp` uses for transcoding. ffmpeg also
|
|
converts a `start_image`/`end_image` upload into the binary PPM at the exact
|
|
output canvas the engine requires, since libvllm vendors neither an image codec
|
|
nor a resampler.
|
|
|
|
**It is slow.** Roughly 176 s per denoise step at 1344x768 on a 20-SM device, so
|
|
the 50-step default is hours. Nothing here imposes a deadline.
|
|
|
|
Geometry mirrors the engine so the two agree: the canvas is truncated onto a
|
|
32-pixel grid, the frame count sits on the 17n+5 grid, and an unspecified canvas
|
|
with a keyframe is derived from that image's aspect on a 768-pixel short edge
|
|
(`MiniMaxH3ResolveShape`, `minimax_h3_planner.cpp`).
|
|
|
|
## Apple Silicon: the MLX GEMM provider (ON by default, gated to prefill)
|
|
|
|
`BUILD_TYPE=metal` builds vllm.cpp's MLX provider for the dense GEMM
|
|
(`VLLM_CPP_MLX=on`, the default here). It is on because upstream now SHAPE-GATES
|
|
it to prefill; it was briefly off in this branch's history, and that was correct
|
|
at the time for an ungated provider.
|
|
|
|
The gate matters more than the flag. MLX's steel GEMM wins prefill but loses
|
|
decode, because the provider pays an `mx::eval` synchronisation plus an output
|
|
memcpy on every call and decode makes ~112 calls *per token*. Measured on an
|
|
Apple M4, Qwen3-1.7B-bf16 warm at p=512 g=128:
|
|
|
|
| configuration | prefill TTFT | warm throughput |
|
|
|---|--:|--:|
|
|
| MLX **gated to prefill** (pin >= 89c46aeb) | **524.5 ms** | **24.37 tok/s, 97.6% of MLX-LM** |
|
|
| MLX ungated (older pins) | 537 ms | 12.7 tok/s |
|
|
| MLX off | 602 ms | 23.9 tok/s, 95.9% |
|
|
|
|
Ratios are against an MLX-LM baseline measured INTERLEAVED with ours over four
|
|
ABBA blocks (its spread 0.34%, ours 0.12%). An earlier revision of this file
|
|
claimed 99.1%; that used a two-run MLX-LM baseline containing an outlier and
|
|
overstated us by about 1.5 points.
|
|
|
|
**`VLLM_CPP_VERSION` and this flag are coupled.** Moving the pin back before
|
|
`89c46aeb` while leaving `VLLM_CPP_MLX=on` would take the middle row — roughly
|
|
half throughput. If you roll the pin back, roll the default back with it.
|
|
|
|
One caveat: MLX's GEMM is not bit-identical to the native kernel, so an MLX build
|
|
produces a different greedy sequence than a non-MLX one. That is a property of the
|
|
provider, not of the gate, and it predates this packaging. Full disposition in
|
|
vllm.cpp `docs/BENCHMARKS.md`.
|
|
|
|
Build knobs:
|
|
|
|
- `VLLM_CPP_MLX=off` builds Metal without the provider: ~124 MB smaller, and
|
|
96.4% of MLX-LM instead of 99.1%.
|
|
- `MLX_VERSION` pins the wheel (default `0.29.4`). MLX is consumed as the
|
|
prebuilt pip wheel because building it from source needs `xcrun metal`, i.e. a
|
|
full Xcode the macOS runners do not have.
|
|
|
|
Packaging vendors `libmlx.dylib`, `mlx.metallib` and MLX's MIT license into
|
|
`package/lib/`, and rewrites `libvllm.dylib`'s rpath to `@loader_path/lib`
|
|
(re-signing it, since `install_name_tool` invalidates the signature). The
|
|
metallib must stay beside `libmlx.dylib`: MLX looks for it there.
|
|
|
|
Testing: `make test` runs the unit specs; export `VLLM_CPP_MODEL=<model>` (and
|
|
optionally `VLLM_CPP_LIBRARY=<libvllm path>`) to enable the e2e specs.
|