mirror of
https://github.com/mudler/LocalAI.git
synced 2026-07-30 09:57:57 -04:00
The engine resolves speculative_config.model against a directory containing config.json, or against ~/.cache/huggingface/hub/models--<org>--<repo>/ snapshots/*, and it never downloads. LocalAI keeps models in its own directory, so the repo-id spelling the vLLM docs teach - "z-lab/Qwen3.6-27B-DFlash" - misses the HF cache and dies deep inside the load with "draft checkpoint not found", which reads like a broken checkpoint rather than a model nobody fetched. Resolve it before the load call: the reference as given, then its last path segment under LocalAI's models dir (what LocalAI's own downloader produces), then the whole reference under the models dir. When none resolve, fail there naming both what was asked for and every location tried, so the message says what to do about it. mtp and ngram pass through untouched - neither has a separate draft checkpoint. A speculative_config that does not parse also passes through, because the engine owns config validation and produces the better error. Docs also gain the two limits that were missing and are easy to lose an afternoon to: speculation is Qwen3.5/3.6-only at this engine pin regardless of format, and mtp/dflash need a safetensors target. The latter is a gap in the engine's GGUF loader rather than a property of GGUF - the format carries MTP weights fine, llama.cpp reads them as nextn.* tensors plus a <arch>.nextn_predict_layers key - so the docs say that rather than implying GGUF cannot express it. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude Code:claude-opus-5 [ClaudeCode]