chore(gallery): add Qwopus Flash V2 variants

Add Q4_K_M and Q8_0 builds with vision and MTP decoding. Pin the
weights and projector to a verified Hugging Face revision.

Assisted-by: Codex:gpt-6
This commit is contained in:
localai-org-maint-bot committed 2026-09-27 12:04:40 +00:00
1 parent c9e822215a
commit 6043e5e0cb
2 files changed
+108

No files matched your search

+14
View File
@@ -48,6 +48,20 @@ To select Q8_0 explicitly, run `local-ai models install mimo-v2.6-distill-qwen-9
The configurations default to 32,768 context tokens and use the model's embedded chat template.
See the [model card](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) for training details.
## Qwopus3.8 Flash V2
Install `qwopus3.8-27b-flash-v2` for the Q4_K_M GGUF build, with Q8_0 available through variant selection:
```bash
local-ai models install qwopus3.8-27b-flash-v2
local-ai models install qwopus3.8-27b-flash-v2 --variant qwopus3.8-27b-flash-v2-q8
```
Both builds use llama.cpp with the embedded chat template, MTP speculative decoding, and the F32 vision projector.
Weights and projector downloads are pinned to a Hugging Face revision and verified with SHA256.
This Apache-2.0 release is a further post-training of Qwopus3.8 Flash for reasoning and agent tasks.
See the [publisher's model card](https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-V2-GGUF) for evaluation details and limitations.
## Hemmingway-1
Install `hemmingway-1` for English text generation with llama.cpp. The gallery groups its Q4_K_M and Q8_0 builds as variants.
+94
View File
@@ -206,6 +206,100 @@
- filename: ds4flash.gguf
uri: https://huggingface.co/unsloth/DeepSeek-V4-Flash-Vision-Exp-GGUF
sha256: 9c46395af7320ec1d68afe81ec7fa1c7060a07117dceabfd977f12a95fa30cdf
- name: "qwopus3.8-27b-flash-v2"
variants:
- model: qwopus3.8-27b-flash-v2-q8
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
urls:
- https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash
- https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-V2-GGUF
description: |
Qwopus3.8-27B-Flash-V2 is a new post-training release for reasoning and agent
workloads. This Q4_K_M GGUF includes the F32 vision projector and uses
llama.cpp's embedded chat template with MTP speculative decoding.
license: "apache-2.0"
tags:
- llm
- gguf
- qwen
- qwen3
- vision
- multimodal
- instruction-tuned
- reasoning
- mtp
icon: https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/CI_Demo/mathv-1327.jpg
overrides:
backend: llama-cpp
function:
automatic_tool_parsing_fallback: true
grammar:
disable: true
known_usecases:
- chat
mmproj: llama-cpp/mmproj/Qwopus3.8-27B-Flash-V2-MTP-Q4_K_M/mmproj-F32.gguf
options:
- use_jinja:true
- spec_type:draft-mtp
- spec_n_max:6
- spec_p_min:0.75
parameters:
model: llama-cpp/models/Qwopus3.8-27B-Flash-V2-MTP-Q4_K_M/Qwopus3.8-27B-Flash-V2-MTP-Q4_K_M.gguf
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/Qwopus3.8-27B-Flash-V2-MTP-Q4_K_M/Qwopus3.8-27B-Flash-V2-MTP-Q4_K_M.gguf
uri: https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-V2-GGUF/resolve/ecb87867b0977dfd1554d2fc54105a802b34345a/Qwopus3.8-27B-Flash-V2-MTP-Q4_K_M.gguf
sha256: 227bedb8ebf4a05e342c99f1f852be19cf0ed394f6cc5901823c07a735ea983e
- filename: llama-cpp/mmproj/Qwopus3.8-27B-Flash-V2-MTP-Q4_K_M/mmproj-F32.gguf
uri: https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-V2-GGUF/resolve/ecb87867b0977dfd1554d2fc54105a802b34345a/mmproj-F32.gguf
sha256: c9d201ea8a2a474ce55cfab6d1e1480d4b2e1574dda976db15aee267072ca4d6
- name: "qwopus3.8-27b-flash-v2-q8"
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
urls:
- https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash
- https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-V2-GGUF
description: |
Qwopus3.8-27B-Flash-V2 is a new post-training release for reasoning and agent
workloads. This Q8_0 GGUF includes the F32 vision projector and uses
llama.cpp's embedded chat template with MTP speculative decoding.
license: "apache-2.0"
tags:
- llm
- gguf
- qwen
- qwen3
- vision
- multimodal
- instruction-tuned
- reasoning
- mtp
icon: https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/CI_Demo/mathv-1327.jpg
overrides:
backend: llama-cpp
function:
automatic_tool_parsing_fallback: true
grammar:
disable: true
known_usecases:
- chat
mmproj: llama-cpp/mmproj/Qwopus3.8-27B-Flash-V2-MTP-Q4_K_M/mmproj-F32.gguf
options:
- use_jinja:true
- spec_type:draft-mtp
- spec_n_max:6
- spec_p_min:0.75
parameters:
model: llama-cpp/models/Qwopus3.8-27B-Flash-V2-MTP-Q8_0/Qwopus3.8-27B-Flash-V2-MTP-Q8_0.gguf
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/Qwopus3.8-27B-Flash-V2-MTP-Q8_0/Qwopus3.8-27B-Flash-V2-MTP-Q8_0.gguf
uri: https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-V2-GGUF/resolve/ecb87867b0977dfd1554d2fc54105a802b34345a/Qwopus3.8-27B-Flash-V2-MTP-Q8_0.gguf
sha256: bc291a2ab2ac209d2cd97f0e0d25bfb98381d4cb4ee4f8baa4cd3c662db95f78
- filename: llama-cpp/mmproj/Qwopus3.8-27B-Flash-V2-MTP-Q4_K_M/mmproj-F32.gguf
uri: https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-V2-GGUF/resolve/ecb87867b0977dfd1554d2fc54105a802b34345a/mmproj-F32.gguf
sha256: c9d201ea8a2a474ce55cfab6d1e1480d4b2e1574dda976db15aee267072ca4d6
- name: "qwopus3.8-27b-flash"
variants:
- model: qwopus3.8-27b-flash-q8