Files
LocalAI/backend/go/moss-tts-cpp
LocalAI [bot] 3bb0d1cb49 feat(backend): add moss-tts-cpp text-to-speech backend (#10860)
* feat(backend): add moss-tts-cpp text-to-speech backend

Add a Go + purego backend wrapping the moss-tts.cpp ggml port of the OpenMOSS
MOSS-TTS-Local v1.5 text-to-speech model (GPT-J local transformer decoded through
MOSS-Audio-Tokenizer-v2), producing 48 kHz stereo audio with optional
reference-audio voice cloning. Mirrors the qwen3-tts-cpp backend: dlopen the
static-ggml shared library, bind the moss-tts.cpp C-API via purego, and serve
the gRPC TTS method. A thin C shim holds the pipeline handle and copies engine
PCM into a Go-freeable buffer.

Wires the CI registration: backend-matrix.yml (CPU, CUDA 12/13, Intel SYCL
f16/f32, Vulkan, ROCm, NVIDIA L4T, plus Darwin metal), backend/index.yaml metas
and image entries pointing at mudler/MOSS-TTS-Local-Transformer-v1.5-GGUF, the
root Makefile build targets, and the changed-backends.js path mapping.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: list the moss-tts-cpp backend among the LocalAI-maintained engines

Add moss-tts.cpp to the README "Backends built by us" table, the
Text-to-Speech compatibility table, and the reference-audio voice-cloning
backend list, so the new backend is documented alongside its peers.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(moss-tts-cpp): pin moss-tts.cpp to the squashed single-commit release

moss-tts.cpp history was collapsed to a single commit; repoint MOSSTTS_CPP_VERSION
to ee722b8e9205ee9b1b1c398a4e87e4e393e9be41.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* backend(moss-tts-cpp): add the moss-tts-cpp-development gallery meta

The gallery had the -development image entries but no matching -development
meta anchor (as locate-anything-cpp and depth-anything-cpp have), so the master
build was not installable as a gallery backend. Add moss-tts-cpp-development
mirroring the production meta with the -development capability image names.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-17 09:26:12 +02:00
..

MOSS-TTS C++ backend

This backend runs the OpenMOSS MOSS-TTS-Local (v1.5) GGUF model through moss-tts.cpp, a from-scratch C++/ggml port with no Python at inference time. It generates 48 kHz stereo speech and supports reference-audio voice cloning.

The engine loads three GGUFs: the local transformer (the model), the MOSS-Audio-Tokenizer neural codec, and the text tokenizer. It is loaded via purego (cgo-less dlopen) exactly like qwen3-tts-cpp.

Model configuration

The model path points at the local transformer GGUF. The codec and text tokenizer are auto-discovered as siblings of the model:

  • codec: a *.gguf whose name contains audio + tokenizer (or codec), e.g. moss-audio-tokenizer-v2-f32.gguf
  • tokenizer: the other *.gguf whose name contains tokenizer, e.g. moss-tokenizer-v1_5.gguf
name: moss-tts-cpp
backend: moss-tts-cpp
parameters:
  model: moss-tts-local-v1_5-q8_0.gguf
known_usecases:
  - tts
tts:
  audio_path: voices/default-reference.wav  # optional model-wide clone reference

Override discovery when the filenames are non-standard:

options:
  - "codec:moss-audio-tokenizer-v2-f32.gguf"
  - "tokenizer:moss-tokenizer-v1_5.gguf"
  - "seed:42"

GGUFs for the Local v1.5 model (plus its codec and tokenizer) live at mudler/MOSS-TTS-Local-Transformer-v1.5-GGUF.

Voice cloning

MOSS-TTS-Local has no named speakers; cloning is driven purely by a reference-audio path. Request precedence is: a request voice that ends in a known audio extension (.wav, .flac, .mp3, .ogg, .m4a), then tts.audio_path. The engine decodes the reference itself, so no client-side resampling is required.

API example

curl http://localhost:8080/v1/audio/speech \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "moss-tts-cpp",
    "input": "This request uses a saved reference voice.",
    "voice": "/path/to/reference.wav"
  }' \
  --output speech.wav

Native end-to-end test

The labeled test loads real GGUFs, synthesizes a 48 kHz stereo WAV, and streams audio:

make -C backend/go/moss-tts-cpp moss-tts-cpp

MOSSTTS_MODEL=/path/to/moss-tts-local-v1_5-q8_0.gguf \
MOSSTTS_CODEC=/path/to/moss-audio-tokenizer-v2-f32.gguf \
MOSSTTS_TOKENIZER=/path/to/moss-tokenizer-v1_5.gguf \
MOSSTTS_LIBRARY=backend/go/moss-tts-cpp/libgomosstts-cpp-fallback.so \
  go test ./backend/go/moss-tts-cpp -ginkgo.label-filter=e2e