mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-28 09:05:05 -04:00
chore(gallery): tag swift-qwen3.8-27b as mtp and vision, fix license
The entry enables spec_type:draft-mtp, so variant ranking needs the mtp tag. Replace the scraped model-card description, set the Swift Open License v1.0 and link the base model repo. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code]
This commit is contained in:
1 parent
e17039fb30
commit
7340970ae7
1 file changed
+10
-32
+10
-32
@@ -2,44 +2,21 @@
|
||||
- name: "swift-qwen3.8-27b"
|
||||
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
|
||||
urls:
|
||||
- https://huggingface.co/ukisai/Swift-Qwen3.8-27b
|
||||
- https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF
|
||||
description: |
|
||||
Website •
|
||||
Learn more •
|
||||
GGUF •
|
||||
Enterprise licensing
|
||||
|
||||
# Swift-Qwen3.8-27B
|
||||
|
||||
Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B,
|
||||
using **58.3% fewer thinking tokens** while maintaining near-identical performance
|
||||
(**<1% loss**) and as a result getting a **x1.95 speed-up** on several tasks.
|
||||
|
||||
The prompt is a sample from LiveCodeBench v6
|
||||
|
||||
## Training approach
|
||||
|
||||
We built Swift by identifying reasoning-marker tokens that, in our analysis, trigger overthinking in Qwen’s
|
||||
reasoning rollouts. We then fine-tuned Qwen by penalizing usage of those tokens while it reasons.
|
||||
|
||||
Swift produces shorter reasoning traces. In our testing, we also observe fewer overthinking errors.
|
||||
|
||||
For maximum gains, Swift also includes a transfer component derived from
|
||||
BottleCap AI's ThinkingCap-Qwen3.6-27B.
|
||||
|
||||
## Evaluation scope
|
||||
|
||||
> All results below compare the Qwen3.8-27B BF16 base with the same base plus the
|
||||
> Swift adapter.
|
||||
|
||||
## Benchmarks
|
||||
|
||||
...
|
||||
license: "other"
|
||||
Swift-Qwen3.8-27B is UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B.
|
||||
The publisher reports 58.3% fewer thinking tokens with less than 1% quality loss.
|
||||
This Q4_K_M GGUF includes the F16 vision projector and enables MTP speculative decoding.
|
||||
The weights use the Swift Open License v1.0.
|
||||
license: "swift-open-license-1.0"
|
||||
tags:
|
||||
- llm
|
||||
- gguf
|
||||
- reasoning
|
||||
- vision
|
||||
- multimodal
|
||||
- mtp
|
||||
overrides:
|
||||
backend: llama-cpp
|
||||
function:
|
||||
@@ -48,6 +25,7 @@
|
||||
disable: true
|
||||
known_usecases:
|
||||
- chat
|
||||
- vision
|
||||
mmproj: llama-cpp/mmproj/Swift-Qwen3.8-27B-Q4_K_M/mmproj-Swift-Qwen3.8-27B-F16.gguf
|
||||
options:
|
||||
- use_jinja:true
|
||||
|
||||
Reference in new issue
Block a user