Files
LocalAI/docs/content/reference/compatibility-table.md
mudler's LocalAI [bot] 2d889e61a6 feat(backend): add magpie-tts-cpp text-to-speech backend (#11115)
* feat(backend): add magpie-tts-cpp text-to-speech backend

Add a Go + purego backend wrapping the magpie-tts.cpp ggml port of NVIDIA's
Magpie TTS Multilingual 357M (encoder + autoregressive decoder over NanoCodec
tokens), producing 22.05 kHz mono audio in 5 baked voices (Aria, Jason, John,
Leo, Sofia; case-insensitive names or indices 0-4) across 9+ languages from a
single self-contained GGUF. Mirrors qwen3-tts-cpp / moss-tts-cpp: dlopen the
static-ggml shared library, bind the flat magpie_tts_capi_* C-API via purego
(no local C shim needed, the upstream .so exports it directly), and serve the
gRPC TTS + TTSStream methods behind base.SingleThread (the C context is not
reentrant across synthesize calls).

The backend CMakeLists translates the Makefile's -DGGML_{CUDA,METAL,VULKAN,HIP}
flags into upstream's MAGPIE_GGML_* toggles (upstream FORCE-overwrites the ggml
cache entries from those), pinned to magpie-tts.cpp v0.1.1
(e3f3dd1ebe22b64e7405f93b519f2d1930712568), which statically links ggml into
libmagpie-tts.so (ldd shows only system libs).

Wires the full registration: backend-matrix.yml (CPU amd64/arm64, CUDA 12/13,
Intel SYCL f16/f32, Vulkan amd64/arm64, ROCm, NVIDIA L4T + L4T CUDA 13, and
Darwin metal), backend/index.yaml metas and image entries, the root Makefile
build targets, the changed-backends backend-filter path mapping, the bump_deps
auto-bump matrix, a test-extra per-backend smoke job, the /backends/known
pref-only importer entry, the backend capabilities map (TTS + TTSStream, no
voice cloning), and the README / compatibility-table docs rows.

Verified locally: unit + e2e Ginkgo suites pass against the real q8_0 GGUF
(22.05 kHz mono WAV, RMS > 0.01), a live gRPC LoadModel + TTS round-trip
returns valid non-silent audio, and the pre-commit gates (make lint,
make test-coverage-check) pass, run manually with LOCALAI_TEST_HTTP_PORT
overriding the locally-occupied 9090.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* gallery: add magpie-tts-cpp model entries (q8_0 + f16)

Add the Magpie TTS Multilingual 357M GGUFs from mudler/magpie-tts.cpp-gguf to
the model gallery: q8_0 (~624 MB, near-lossless, fastest decode, recommended)
with an f16 (~784 MB) variant, both served by the magpie-tts-cpp backend.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* magpie-tts-cpp: bump pin to rewritten upstream v0.1.1 SHA

Upstream history was rewritten to purge accidentally committed build
artifacts; v0.1.1 now resolves to 6f7696cf.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-25 08:38:48 +02:00

14 KiB

+++ disableToc = false title = "Model compatibility table" weight = 24 url = "/model-compatibility/" +++

Besides llama based models, LocalAI is compatible also with other architectures. The table below lists all the backends, compatible models families and the associated repository.

{{% notice note %}}

LocalAI will attempt to automatically load models which are not explicitly configured for a specific backend. You can specify the backend to use by configuring a model with a YAML file. See [the advanced section]({{%relref "advanced" %}}) for more details.

All backends listed here can be installed on demand from the [Backend Gallery]({{%relref "features/backends" %}}). The exact set of acceleration variants published for each backend is defined in backend/index.yaml.

{{% /notice %}}

Text Generation & Language Models

Backend Description Capability Embeddings Streaming Acceleration
llama.cpp LLM inference in C/C++. Supports LLaMA, Mamba, RWKV, Falcon, Starcoder, GPT-2, and many others GPT, Functions yes yes CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
ik_llama.cpp Hard fork of llama.cpp optimized for CPU/hybrid CPU+GPU with IQK quants, custom quant mixes, and MLA for DeepSeek GPT yes yes CPU (AVX2+)
turboquant llama.cpp fork adding the TurboQuant KV-cache quantization scheme GPT yes yes CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Jetson L4T
ds4 DeepSeek V4 Flash single-model inference engine, optimized for Metal and CUDA GPT no yes CPU, CUDA 12/13, Metal, Jetson L4T
vLLM Fast LLM serving with PagedAttention; GPTQ/AWQ/FP8 quantization GPT, Functions, Multimodal no yes CUDA 12/13, ROCm, Intel SYCL, Jetson L4T
vLLM Omni Unified multimodal generation (text, image, video, audio) on top of vLLM Multimodal GPT, Functions no yes CUDA 12/13, ROCm, Jetson L4T
SGLang Fast serving framework for LLMs and vision-language models with speculative decoding GPT, Functions, Multimodal no yes CUDA 12/13, ROCm, Intel SYCL, Jetson L4T
transformers HuggingFace Transformers framework GPT, Embeddings, Multimodal yes yes* CUDA 12/13, ROCm, Intel SYCL, Metal
MLX Apple Silicon LLM inference GPT, Functions no yes CPU, CUDA 12/13, Metal, Jetson L4T
MLX-VLM Vision-Language Models on Apple Silicon Multimodal GPT, Functions no yes CPU, CUDA 12/13, Metal, Jetson L4T
MLX Distributed Distributed LLM inference across multiple Apple Silicon Macs GPT no no CPU, CUDA 12/13, Metal, Jetson L4T
tinygrad Minimalist deep-learning framework with zero runtime dependencies GPT, Embeddings, Multimodal yes yes CPU

Speech-to-Text

Backend Description Acceleration
whisper.cpp OpenAI Whisper in C/C++ CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
faster-whisper Fast Whisper with CTranslate2 CPU, CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T
WhisperX Word-level timestamps and speaker diarization CPU, CUDA 12/13, Metal, Jetson L4T
moonshine Ultra-fast transcription for low-end devices (ONNX) CPU, CUDA 12/13, Metal
parakeet.cpp C++/GGML port of NVIDIA NeMo Parakeet (tdt/ctc/rnnt/hybrid), with cache-aware streaming CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
CrispASR Unified speech engine (whisper.cpp fork) supporting Parakeet, Canary, and many ASR architectures, plus TTS CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
voxtral Voxtral Realtime 4B speech-to-text in pure C CPU, Metal
Qwen3-ASR Qwen3 automatic speech recognition CPU, CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T
NeMo NVIDIA NeMo ASR toolkit CPU, CUDA 12/13, ROCm, Intel SYCL, Metal
sherpa-onnx Sherpa-ONNX ASR (Whisper, Paraformer, SenseVoice) and TTS CPU, CUDA 12, Metal

Text-to-Speech

Backend Description Acceleration
piper Fast neural TTS CPU, Metal
Coqui TTS TTS with 1100+ languages and voice cloning CUDA 12, ROCm, Intel SYCL, Metal
Kokoro Lightweight TTS (82M params) CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T
Kokoros Pure Rust Kokoro TTS via ONNX CPU
Chatterbox Production-grade TTS with emotion control CPU, CUDA 12/13, Metal, Jetson L4T
VibeVoice Real-time TTS with voice cloning CPU, CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T
vibevoice.cpp Native C++/GGML port of VibeVoice for TTS (voice cloning) and long-form ASR with diarization CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
moss-tts.cpp Native C++/GGML port of OpenMOSS MOSS-TTS-Local v1.5: 48 kHz stereo TTS with reference-audio voice cloning CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
magpie-tts.cpp Native C++/GGML port of NVIDIA Magpie TTS Multilingual 357M: 22.05 kHz mono TTS, 5 baked voices, 9+ languages CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
Qwen3-TTS TTS with custom voice, voice design, and voice cloning CPU, CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T
qwentts.cpp Native C++/GGML Qwen3-TTS with streaming, named speakers, and voice design CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
OmniVoice Native C++/GGML TTS with voice cloning, voice design, and streaming CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
fish-speech High-quality TTS with voice cloning CPU, CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T
Pocket TTS Lightweight CPU-efficient TTS with voice cloning CPU, CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T
OuteTTS TTS with custom speaker voices CPU, CUDA 12
faster-qwen3-tts Real-time Qwen3-TTS with CUDA graph capture CPU, CUDA 12/13, Jetson L4T
NeuTTS Air Instant voice cloning, on-device TTS CPU, CUDA 12, ROCm
VoxCPM Expressive end-to-end TTS CPU, CUDA 12/13, ROCm, Intel SYCL, Metal
Kitten TTS Kitten TTS model CPU, Metal
Supertonic Lightning-fast on-device multilingual TTS via ONNX CPU
MLX-Audio Audio models on Apple Silicon CPU, CUDA 12/13, Metal, Jetson L4T
liquid-audio LFM2 end-to-end speech-to-speech, ASR, and TTS CPU, CUDA 12/13, ROCm, Intel SYCL, Jetson L4T

Music & Sound Generation

Backend Description Acceleration
ACE-Step Music generation from text descriptions, lyrics, or audio CPU, CUDA 12/13, ROCm, Intel SYCL, Metal
acestep.cpp ACE-Step 1.5 C++ backend using GGML CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T

Image & Video Generation

Backend Description Acceleration
stable-diffusion.cpp Stable Diffusion, Flux, PhotoMaker, Ideogram in C/C++ CPU, CUDA 12/13, Intel SYCL, Vulkan, Metal, Jetson L4T
diffusers HuggingFace diffusion models (image and video generation) CPU, CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T
vLLM Omni Multimodal generation including text-to-image and text-to-video CUDA 12/13, ROCm, Jetson L4T

Vision, Detection & Recognition

Backend Description Acceleration
RF-DETR Real-time transformer-based object detection (Python) CPU, CUDA 12/13, Intel SYCL, Metal, Jetson L4T
rf-detr.cpp Native RF-DETR object detection and instance segmentation in C/C++ using GGML CPU, CUDA 12/13, Intel SYCL, Vulkan, Jetson L4T
locate-anything.cpp Open-vocabulary object detection and visual grounding (LocateAnything-3B) in C/C++ using GGML CPU, CUDA 12/13, Intel SYCL, Vulkan, Jetson L4T
depth-anything.cpp Depth Anything 3 monocular metric depth + camera pose in C/C++ using GGML CPU, CUDA 12/13, Intel SYCL, Vulkan, Jetson L4T
sam3.cpp Segment Anything (SAM 3/2/EdgeTAM) with text/point/box prompts in C/C++ using GGML CPU, CUDA 12/13, Intel SYCL, Vulkan, Jetson L4T
face-detect.cpp Native face detection, recognition, embedding, demographics and anti-spoofing (SCRFD/ArcFace, YuNet/SFace) in C/C++ using GGML CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
voice-detect.cpp Native speaker (voice) recognition and voice analysis (ECAPA-TDNN, WeSpeaker, ERes2Net, CAM++, wav2vec2) in C/C++ using GGML CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T
insightface Face verification, embedding, and anti-spoofing liveness (ONNX Runtime) CPU, CUDA 12
speaker-recognition Speaker (voice) recognition via SpeechBrain ECAPA-TDNN CPU, CUDA 12, Metal

Audio Processing

Backend Description Acceleration
Silero VAD Voice Activity Detection CPU, Metal
LocalVQE Joint acoustic echo cancellation, noise suppression, and dereverberation in C/C++ using GGML CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Jetson L4T
Opus Audio codec for WebRTC / Realtime API CPU, Metal

Utilities & Other

Backend Description Acceleration
rerankers Document reranking for RAG CUDA 12, ROCm, Intel SYCL, Metal
privacy-filter.cpp Standalone GGML engine for the openai-privacy-filter PII/NER token-classification model family (powers LocalAI's PII redaction tier) CPU, CUDA 13, Vulkan
local-store Local-first vector database for embeddings CPU, Metal
TRL Fine-tuning (SFT, DPO, GRPO, RLOO, KTO, ORPO) CPU, CUDA 12/13
llama.cpp quantization HuggingFace → GGUF model conversion and quantization CPU, Metal

Acceleration Support Summary

GPU Acceleration

  • NVIDIA CUDA: CUDA 12.0, CUDA 13.0 support across most backends
  • AMD ROCm: HIP-based acceleration for AMD GPUs
  • Intel oneAPI: SYCL-based acceleration for Intel GPUs (F16/F32 precision)
  • Vulkan: Cross-platform GPU acceleration
  • Metal: Apple Silicon GPU acceleration (M1/M2/M3+)

Specialized Hardware

  • NVIDIA Jetson (L4T CUDA 12): ARM64 support for embedded AI (AGX Orin, Jetson Nano, Jetson Xavier NX, Jetson AGX Xavier)
  • NVIDIA Jetson (L4T CUDA 13): ARM64 support for embedded AI (DGX Spark)
  • Apple Silicon: Native Metal acceleration for Mac M1/M2/M3+
  • Darwin x86: Intel Mac support

CPU Optimization

  • AVX/AVX2/AVX512: Advanced vector extensions for x86
  • Quantization: 4-bit, 5-bit, 8-bit integer quantization support
  • Mixed Precision: F16/F32 mixed precision support

Note: any backend name listed above can be used in the backend field of the model configuration file (See [the advanced section]({{%relref "advanced" %}})).

  • * Only for CUDA and OpenVINO CPU/XPU acceleration.