PTQ1_0 is a Prism-private GGUF type (GGML_TYPE_PTQ1_0 = 143 in the
PrismML llama.cpp fork), so stock llama-cpp cannot load it. Switch to
the bonsai backend like the existing ternary-bonsai-27b entries, and
replace the scraped Qwen3.8 description and icon.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
The entry enables spec_type:draft-mtp, so variant ranking needs the mtp
tag. Replace the scraped model-card description, set the Swift Open
License v1.0 and link the base model repo.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Add IQ4_XS and Q4_K_M GGUF builds with a BF16 vision projector.
Pin verified artifacts and document installation and variant selection.
Assisted-by: Codex:gpt-6
Add Q4_K_M and Q8_0 builds with the F16 vision projector and install docs.
Pin artifact revisions and verify SHA256 against HF LFS metadata and HTTP
headers.
Assisted-by: Codex:gpt-6
Add four text-only llama.cpp builds with pinned download URLs and
verified checksums. Document variant selection and the model license.
Assisted-by: Codex:gpt-6
Add Q4_K_M and Q8_0 builds with the F16 vision projector and pinned
artifact URLs. Document installation and explicit variant selection.
Assisted-by: Codex:GPT-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* fix(gallery): remove invalid Qwen-Image chat entry
The entry sends diffusion weights to llama.cpp as a chat model.
Remove it and document the existing image-generation alternatives.
Assisted-by: Codex:gpt-6
* feat(gallery): add Hemmingway-1 GGUF variants
Add Q4_K_M and Q8_0 builds for llama.cpp with embedded chat templates.
Record the upstream CC BY-NC 4.0 license and installation instructions.
Verify both SHA256 values against Hugging Face LFS metadata and headers.
Assisted-by: Codex:gpt-6
---------
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Add five nemo-speech-cpp gallery entries for the diarization
capability introduced by the NeMo-Speech.cpp bump in #12257:
- nemo-speech-cpp-sortformer-diarization-v2: standalone streaming
Sortformer 4-speaker diarization (nvidia/diar_streaming_sortformer_4spk-v2).
Serves /v1/audio/diarization with known_usecases: [diarization].
- nemo-speech-cpp-nemotron-3.5-asr-streaming: standalone multilingual
streaming ASR (nvidia/nemotron-3.5-asr-streaming-0.6b).
- nemo-speech-cpp-nemotron-3.5-asr-streaming-diarized: Nemotron ASR
with the sortformer attached via the diar_model option, giving
per-word speaker tags on /v1/audio/transcriptions.
- nemo-speech-cpp-parakeet-tdt-0.6b-v3: standalone multilingual ASR,
25 languages (nvidia/parakeet-tdt-0.6b-v3).
- nemo-speech-cpp-parakeet-tdt-0.6b-v3-diarized: Parakeet v3 ASR
with the sortformer attached via the diar_model option, giving
per-word speaker tags on /v1/audio/transcriptions.
No backend code changes: the sortformer to familyDiarization mapping,
the diar_model option, and the MethodDiarize gRPC method already exist.
All five entries verified end-to-end against a running LocalAI instance.
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Add gallery entries for three vllm.cpp-backed models:
- laya-vllm-cpp: multilingual non-autoregressive System 1 decision
model (ModernBERT-large, 421M params). Served via POST /v1/systemone.
Weights from convaiinnovations/laya (Apache-2.0).
- cua-s1-forms-vllm-cpp: one-pass option scorer for GUI form filling.
Served via POST /api/score. Weights from cua-ai/cua-s1-forms (MIT).
- gliner2.5-vllm-cpp: zero-shot named entity recognition and structured
extraction (GLiNER2.5, mDeBERTa-v3-base, 287M params). Served via the
token classification endpoint. Weights from fastino/gliner2.5-multi-v1
(Apache-2.0).
All entries set backend: vllm-cpp and declare CPU+GPU tags.
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* chore(model gallery): 🤖 add new models via gallery agent
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* feat(gallery): add NeoHorse-1-9B GGUF variants
Add the official Q4_K_M, Q5_K_M, and Q8_0 builds with revision-pinned
weights and verified SHA256 values.
Assisted-by: Codex:gpt-6
* chore(model gallery): 🤖 add new models via gallery agent
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* feat(gallery): add Qwen3.8 35B Distill variants
Add Q4_K_M, Q5_K_M, and Q8_0 builds with the vision projector.
Pin publisher revisions and document installation and variant selection.
Assisted-by: Codex:gpt-6
* feat(gallery): add ByteShape Qwen3.8 variants
Offer five ShapeLearn GGUF builds with a vision projector and MTP.
Pin downloads to a verified HF revision and document variant selection.
Assisted-by: Codex:gpt-6
* feat(gallery): add Flash Next GSQ-RCO variants
Offer Q2_0, IQ2_XS, and IQ3_XXS builds with both model shards and the
vision projector. Pin and verify download hashes and document how to
select each variant.
Assisted-by: Codex:GPT-6
* feat(gallery): add Occamy-1.0 GGUF variants
Add the publisher's Q4_K_M and Q8_0 builds with the F16 vision projector.
Link the builds as variants and pin downloads to a verified revision.
Document installation and the source tokenizer's NFC requirement.
Assisted-by: Codex:gpt-6
* fix(gallery): set MiniCPM5 context at the top level
The Q4 and Q8 overrides place context_size inside parameters, where
PredictionOptions ignores it. Move it beside parameters so both
builds use the intended 8,192-token context, matching F16.
Assisted-by: Codex:GPT-6
* feat(gallery): add Hy-MT2 7B GGUF variants
Offer the official Q4_K_M, Q6_K, and Q8_0 builds for translation.
Pin the downloads and document installation and translation prompts.
Assisted-by: Codex:gpt-6
* feat(gallery): add Maple-Preview GGUF variants
Offer four ternary builds through the existing llama.cpp backend.
Use the publisher's CPU settings and embedded chat template.
Pin downloads and verify SHA256 values against two HF metadata sources.
Document installation and explicit variant selection.
Assisted-by: Codex:gpt-6
* feat(gallery): add Qwen3.8 Cyber GGUF variants
Offer IQ4_XS and Q8_0 builds with the matching BF16 vision projector.
Pin download revisions and document automatic and explicit selection.
Assisted-by: Codex:GPT-6
* feat(gallery): add official NeoHorse 4B variants
Offer the official Q5_K_M and BF16 GGUF builds alongside the existing
community quantizations. Pin both downloads and document variant selection.
Assisted-by: Codex:gpt-6
* chore(model gallery): 🤖 add new models via gallery agent
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* chore(model gallery): 🤖 add new models via gallery agent
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* fix(stablediffusion-ggml): bump sd.cpp for Qwen-Image 2.1, fix RPC build
Bump stable-diffusion.cpp to c678dfe70, which adds Qwen-Image 2.1
support (leejet/stable-diffusion.cpp#1994).
The same range pulls a ggml update that adds GGML_OP_SAGE_ATTN but
does not update the GGML_OP_COUNT static_assert in ggml-rpc.h. We build
with SD_RPC=ON and upstream CI does not, so every backend image failed
to compile (see #12170).
Add a sync-rpc-op-count step after checkout. It sets the ggml-rpc.h
assert to the count that ggml.c asserts. The new op is appended before
GGML_OP_COUNT, so existing op ids on the wire do not change, and the
RPC handshake compares only major and minor versions. When upstream
fixes the header, the step does nothing, so future automated bumps
stay green.
Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(gallery): add Qwen-Image 2.1 GGUF for stablediffusion-ggml
Add qwen-image-2.1-q4_k-ggml, with qwen-image-2.1-q8_0-ggml as a
variant, from leejet/Qwen-Image-2.1-GGUF. The config follows the
upstream stable-diffusion.cpp recipe: Qwen3-VL-8B-Instruct text
encoder, the Qwen-Image 2.1 VAE, cfg scale 6, euler sampler.
The bundle also pulls the Qwen3-VL mmproj as llm_vision_path, so that
image editing with reference images works. The text encoder and mmproj
use the same filenames and checksums as the qwen3-vl-8b-instruct
entry, so both entries share the files on disk.
Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Consolidates all 25 pending gallery bot PRs into a single merge to resolve
the conflict cascade — every PR branched from a different point in master
and they all touch gallery/index.yaml, so merging them individually was
blocked by constant conflicts.
Changes:
- gallery/index.yaml: +739 lines (new model entries and fixes)
- docs/content/features/model-gallery.md: +116 lines (new model docs)
- docs/content/features/audio-cpp.md: +12 lines (Sortformer checksum fix)
Entry count: 1597 -> 1890 (293 new entries, no duplicates, YAML validated).
Supersedes: #11986#11992#11994#11996#11999#12002#12017#12019#12021#12025#12027#12029#12032#12036#12037#12038#12041#12042#12043#12047#12050#12064#12065#12066#12118
Assisted-by: MAKI:regolo/glm5.2
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
chore(model gallery): add four Italian community Piper voices
Add the Ugo voice from Einrich99/PiperTTS-UGO-Italian and the Aurora,
Giorgio and Leonardo voices from kirys79/piper_italiano. All four use
the piper backend and are CC BY 4.0.
The kirys79 Giorgio and Leonardo files carry checkpoint names and a
bare .json config. The entries save them as it_IT-<voice>-high.onnx
and .onnx.json, because the piper backend looks for the config at
<model>.onnx.json.
Assisted-by: Claude:claude-opus-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* fix(vulkan): preserve host ICD discovery for packaged backends
Add bundled Mesa manifests through VK_ADD_DRIVER_FILES instead of replacing the system driver list. Merge inherited and model-specific additive paths while preserving explicit operator overrides, with regression coverage.
Assisted-by: Codex:gpt-5 golangci-lint
Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* feat(3d): add Kimodo CPU and Vulkan animation backend
Introduce a distinct animation capability and model-described 3D operations, with a typed /3d/animate API, RPC transport, distributed media staging, permissions, and tracing.
Add a persistent kimodo.cpp adapter, skeleton GLB export, CPU/Vulkan packages, model and backend galleries, importer support, CI builds, and documentation. Adapt Studio inputs to each model and provide real-time skeleton playback, seeking, and history.
Cover backend validation, packaging, API behavior, importer inventories, distributed staging, and Studio workflows. Validate real-model CPU/Vulkan generation and deploy the integration to the local QA instance.
Assisted-by: Codex:gpt-5 golangci-lint
Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Richard Palethorpe <io@richiejp.com>
* feat(kimodocpp): adopt monolithic encoders and resident inference
Update upstream for resident weights, packed execution paths, and cached motion graphs. Default to all 32 text layers while retaining configurable streaming and legacy bundle support.
Use monolithic Q8_0 encoders by default and offer all six published quantizations through the gallery and importer. Refresh pinned hashes, tests, and documentation; remove the obsolete thread patch and ensure cached source checkouts follow the upstream pin.
Validated CPU and Vulkan generation, lower-bit streaming, gallery/importer suites, packaging, lint, and cold/warm Studio generation on localai-dev.
Assisted-by: Codex:gpt-5 golangci-lint
Assisted-by: Codex:gpt-5.6-sol
Signed-off-by: Richard Palethorpe <io@richiejp.com>
---------
Signed-off-by: Richard Palethorpe <io@richiejp.com>
The Ministral 3 14B Reasoning entry inherits a Mistral 0.3 prompt and
JSON parser. Its name-first tool calls can therefore reach clients as
plain text.
Use the embedded template and llama.cpp's native tool parser. Document
migration for installed configurations, which gallery updates do not
rewrite.
Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* [gallery] feat: add Orukeet to the existing NeMo speech backend
Assisted-by: Codex:gpt-6
Signed-off-by: Nathan Roll <nathan@oruk.ai>
* docs(nemo): remove model-specific instructions
Keep the backend guide focused on model families, as requested by mudler.
Remove the gallery limitation instead of restoring an outdated claim.
Assisted-by: Codex:GPT-6
Signed-off-by: Nathan Roll <nathan@oruk.ai>
---------
Signed-off-by: Nathan Roll <nathan@oruk.ai>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The batch of "gallery: apply PR" commits replayed gallery-agent diffs
against a stale base. Each new top-of-file entry overwrote the entry
above it instead of being inserted, which lost seven entries:
- qwen3.8-27b-uncensored-q4/-q8 (#11705, overwritten by #11909)
- qwen3.8-flash-next-uncensored (#11832, overwritten by #11841)
- spark-x2.5-4b-q4/-q6/-q8 (#11923, overwritten by #11926)
- deepseek-v4-flash-vision-exp (#11873): #11927 renamed its name line
to qwopus3.8-27b-flash, which duplicated that entry and failed the
"declares every entry name exactly once" gallery lint on master.
Each restored entry is identical (YAML-equal) to the one in its PR head.
Assisted-by: Claude:claude-opus-5 [Bash]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Let the model-provided tokenizer template format Gemma conversations instead of maintaining a shared inline prompt template.
Assisted-by: Codex:gpt-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>