diff --git a/gallery/index.yaml b/gallery/index.yaml index 929c337bb..eb9e5e15c 100644 --- a/gallery/index.yaml +++ b/gallery/index.yaml @@ -2,44 +2,21 @@ - name: "swift-qwen3.8-27b" url: "github:mudler/LocalAI/gallery/virtual.yaml@master" urls: + - https://huggingface.co/ukisai/Swift-Qwen3.8-27b - https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF description: | - Website  •  - Learn more  •  - GGUF  •  - Enterprise licensing - - # Swift-Qwen3.8-27B - - Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, - using **58.3% fewer thinking tokens** while maintaining near-identical performance - (**<1% loss**) and as a result getting a **x1.95 speed-up** on several tasks. - - The prompt is a sample from LiveCodeBench v6 - - ## Training approach - - We built Swift by identifying reasoning-marker tokens that, in our analysis, trigger overthinking in Qwen’s - reasoning rollouts. We then fine-tuned Qwen by penalizing usage of those tokens while it reasons. - - Swift produces shorter reasoning traces. In our testing, we also observe fewer overthinking errors. - - For maximum gains, Swift also includes a transfer component derived from - BottleCap AI's ThinkingCap-Qwen3.6-27B. - - ## Evaluation scope - - > All results below compare the Qwen3.8-27B BF16 base with the same base plus the - > Swift adapter. - - ## Benchmarks - - ... - license: "other" + Swift-Qwen3.8-27B is UkisAI's reasoning-efficient fine-tune of Qwen3.8-27B. + The publisher reports 58.3% fewer thinking tokens with less than 1% quality loss. + This Q4_K_M GGUF includes the F16 vision projector and enables MTP speculative decoding. + The weights use the Swift Open License v1.0. + license: "swift-open-license-1.0" tags: - llm - gguf - reasoning + - vision + - multimodal + - mtp overrides: backend: llama-cpp function: @@ -48,6 +25,7 @@ disable: true known_usecases: - chat + - vision mmproj: llama-cpp/mmproj/Swift-Qwen3.8-27B-Q4_K_M/mmproj-Swift-Qwen3.8-27B-F16.gguf options: - use_jinja:true