LocalAI

mirror of https://github.com/mudler/LocalAI.git synced 2026-06-11 02:07:27 -04:00

Files

LocalAI [bot] 618e90cd13 feat(gallery): add Gemma 4 QAT family + MTP speculative-decoding pairs (#10215 )

Add the remaining official Google Gemma 4 QAT Q4_0 GGUFs (E2B, E4B,
26B-A4B, 31B) next to the existing 12B entry, each shipping its
multimodal mmproj.

Also add three MTP (Multi-Token Prediction) speculative-decoding bundles
that pair each QAT target with a QAT-matched assistant/drafter head:

  - 12B       <- Janvitos/gemma-4-12B-it-qat-assistant-MTP-Q8_0-GGUF
  - 26B-A4B   <- boxwrench/gemma-4-qat-mtp-assistant-heads
  - 31B       <- boxwrench/gemma-4-qat-mtp-assistant-heads

The assistant heads use the gemma4_assistant architecture and are not
standalone chat models, so each entry bundles the target + draft and
sets draft_model together with the draft-mtp spec options
(spec_type:draft-mtp / spec_n_max:6 / spec_p_min:0.75), matching
MTPSpecOptions() in core/config/mtp.go. QAT-matched heads raise draft
acceptance substantially over generic non-QAT heads.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>

2026-06-08 10:26:42 +02:00

alpaca.yaml

…

arch-function.yaml

…

bge-m3-colbert.yaml

feat(middleware): Model routing, PII filtering, Cloud model proxies (#9802 )

2026-05-25 09:28:27 +02:00

cerbero.yaml

…

chatml-hercules.yaml

…

chatml.yaml

feat(gallery): Speed up load times and clean gallery entries (#9211 )

2026-05-06 14:51:38 +02:00

codellama.yaml

…

command-r.yaml

…

deephermes.yaml

…

deepseek-r1.yaml

feat(gallery): Speed up load times and clean gallery entries (#9211 )

2026-05-06 14:51:38 +02:00

deepseek.yaml

…

dreamshaper.yaml

…

falcon3.yaml

…

flux-ggml.yaml

…

flux.yaml

…

gemma.yaml

feat(gallery): Speed up load times and clean gallery entries (#9211 )

2026-05-06 14:51:38 +02:00

granite3-2.yaml

…

granite4.yaml

feat(gallery): Speed up load times and clean gallery entries (#9211 )

2026-05-06 14:51:38 +02:00

granite.yaml

…

harmony.yaml

feat(gallery): Speed up load times and clean gallery entries (#9211 )

2026-05-06 14:51:38 +02:00

hermes-2-pro-mistral.yaml

…

hermes-vllm.yaml

…

index.yaml

feat(gallery): add Gemma 4 QAT family + MTP speculative-decoding pairs (#10215 )

2026-06-08 10:26:42 +02:00

jamba.yaml

…

kokoros.yaml

feat: Add Kokoros backend (#9212 )

2026-04-08 19:23:16 +02:00

lfm.yaml

feat(realtime): Add Liquid Audio s2s model and assistant mode on talk page (#9801 )

2026-05-13 21:57:27 +02:00

liquid-audio.yaml

feat(realtime): Add Liquid Audio s2s model and assistant mode on talk page (#9801 )

2026-05-13 21:57:27 +02:00

llama3-instruct.yaml

…

llama3.1-instruct-grammar.yaml

…

llama3.1-instruct.yaml

…

llama3.1-reflective.yaml

…

llama3.2-fcall.yaml

…

llama3.2-quantized.yaml

…

llava.yaml

…

ltx-ggml.yaml

feat(stablediffusion-ggml): LTX-2 support + LTX-2.3 GGUF gallery entries (#9980 )

2026-05-25 13:00:28 +02:00

mathstral.yaml

…

mistral-0.3.yaml

…

moondream.yaml

feat(gallery): Speed up load times and clean gallery entries (#9211 )

2026-05-06 14:51:38 +02:00

mudler.yaml

…

nanbeige4.1.yaml

feat(gallery): Speed up load times and clean gallery entries (#9211 )

2026-05-06 14:51:38 +02:00

noromaid.yaml

…

openvino.yaml

feat(gallery): Speed up load times and clean gallery entries (#9211 )

2026-05-06 14:51:38 +02:00

parler-tts.yaml

…

phi-2-chat.yaml

…

phi-2-orange.yaml

…

phi-3-chat.yaml

…

phi-3-vision.yaml

…

phi-4-chat-fcall.yaml

…

phi-4-chat.yaml

…

piper.yaml

…

pocket-tts.yaml

feat(tts): add pocket-tts backend (#8018 )

2026-01-13 23:35:19 +01:00

qwen3-deepresearch.yaml

…

qwen3-openbuddy.yaml

…

qwen3.yaml

fix: tool-call JSON leaks into content with stream+tools on tokenizer-template models (#10052 ) (#10057 )

2026-05-29 10:12:53 +02:00

qwen-fcall.yaml

…

qwen-image.yaml

…

rerankers.yaml

…

rwkv.yaml

…

sd-ggml.yaml

…

sentencetransformers.yaml

…

sglang-gemma-4-e2b-mtp.yaml

feat(sglang): wire engine_args, add cuda13 build, ship MTP gallery demos (#9686 )

2026-05-07 17:27:29 +02:00

sglang-gemma-4-e4b-mtp.yaml

feat(sglang): wire engine_args, add cuda13 build, ship MTP gallery demos (#9686 )

2026-05-07 17:27:29 +02:00

sglang-mimo-7b-mtp.yaml

feat(sglang): wire engine_args, add cuda13 build, ship MTP gallery demos (#9686 )

2026-05-07 17:27:29 +02:00

sglang.yaml

feat(sglang): wire engine_args, add cuda13 build, ship MTP gallery demos (#9686 )

2026-05-07 17:27:29 +02:00

sherpa-onnx-asr.yaml

feat: Add Sherpa ONNX backend for ASR and TTS (#8523 )

2026-04-24 14:40:06 +02:00

sherpa-onnx-tts.yaml

feat: Add Sherpa ONNX backend for ASR and TTS (#8523 )

2026-04-24 14:40:06 +02:00

sherpa-onnx-vad.yaml

feat: Add Sherpa ONNX backend for ASR and TTS (#8523 )

2026-04-24 14:40:06 +02:00

smolvlm.yaml

feat(gallery): Speed up load times and clean gallery entries (#9211 )

2026-05-06 14:51:38 +02:00

stablediffusion3.yaml

…

tuluv2.yaml

…

vibevoice.yaml

…

vicuna-chat.yaml

…

virtual.yaml

…

vllm.yaml

…

wan-ggml.yaml

chore(gallery): fixup wan

2026-04-19 21:31:22 +00:00

whisper-base.yaml

…

wizardlm2.yaml

…

z-image-ggml.yaml

Fix load of z-image-turbo (#9264 )

2026-04-11 08:42:13 +02:00