mirror of
https://github.com/mudler/LocalAI.git
synced 2026-07-30 09:57:57 -04:00
Bumps VLLM_CPP_VERSION from 9e1c9025 to eec09bed and turns VLLM_CPP_MLX back on. These two must move together, which is why they are one commit. Upstream now shape-gates the MLX provider to prefill: it declines m < 2, which is exactly the decode GEMV. MLX's steel GEMM wins prefill, 524.5 ms of TTFT against 602 for the native path, but loses decode badly because the provider pays an mx::eval synchronisation and an output memcpy on every call while decode makes about 112 calls per token. Ungated it does both; gated it does only the good half. Measured on an Apple M4 with Qwen3-1.7B-bf16 warm at p=512 g=128: MLX gated to prefill (pin >= 89c46aeb) TTFT 524.5 ms 24.40 tok/s, 99.1% of MLX-LM MLX ungated (older pins) TTFT 537 ms 12.7 tok/s MLX off TTFT 602 ms 23.9 tok/s This branch briefly defaulted the provider off, which was the correct call for an ungated provider at the old pin. The gate is what makes on correct again, so the pin and the flag are coupled: rolling VLLM_CPP_VERSION back before 89c46aeb while leaving MLX on would select the middle row and roughly halve throughput. Both the Makefile comment and the README state that dependency explicitly. The bump also brings six Metal kernels landed upstream since the old pin — mma prefill attention, a vectorised decode V accumulation, vectorised attention staging, a fused qk-norm-RoPE preamble, a simdgroup-per-row softmax and a simdgroup-per-head preamble — which take the non-MLX Metal path from 89.4% to 96.4% of MLX-LM on their own. One caveat, recorded in the README: MLX's GEMM is not bit-identical to the native kernel, so an MLX build produces a different greedy sequence than a non-MLX build. That is a property of the provider rather than of the gate and predates this packaging. Assisted-by: Claude Code:claude-opus-5 [ClaudeCode] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>