mirror of
https://github.com/mudler/LocalAI.git
synced 2026-07-30 18:09:05 -04:00
* feat(ui): add voice library workflow Give administrators a production-ready flow to record or upload consented reference audio, manage reusable profiles, inspect API usage, discover compatible models, and hand a saved voice directly to text-to-speech. Assisted-by: Codex:gpt-5 * feat(voice): add managed voice cloning profiles Make reusable reference voices manageable through the admin API instead of requiring model-directory and YAML edits. Discover compatible installed and gallery models from server-side backend capabilities, retain explicit model configuration controls, and stage saved references for supported backends. Expose profile management through REST and MCP, document backend-specific behavior, and cover the workflow from profile creation through real Qwen3-TTS synthesis. Harden the agent-job HTTP test against completion racing cancellation. Assisted-by: Codex:gpt-5 --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
38 lines
1.4 KiB
YAML
38 lines
1.4 KiB
YAML
---
|
|
name: localai
|
|
|
|
config_file: |-
|
|
name: pocket-tts
|
|
backend: pocket-tts
|
|
description: |
|
|
Pocket TTS is a lightweight text-to-speech model designed to run efficiently on CPUs.
|
|
This model supports voice cloning through HuggingFace voice URLs or local audio files.
|
|
|
|
parameters:
|
|
model: ""
|
|
|
|
# TTS configuration
|
|
tts:
|
|
# Advertise saved Voice Library profiles. A per-request saved profile
|
|
# overrides voice/audio_path/default_voice without changing this YAML.
|
|
voice_cloning: true
|
|
# Voice selection - can be:
|
|
# 1. Built-in voice name (e.g., "alba", "marius", "javert", "jean", "fantine", "cosette", "eponine", "azelma")
|
|
# 2. HuggingFace URL (e.g., "hf://kyutai/tts-voices/alba-mackenna/casual.wav")
|
|
# 3. Local file path (relative to model directory or absolute)
|
|
# voice: "azelma"
|
|
# Alternative: use audio_path to specify a voice file directly
|
|
# audio_path: "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
|
|
|
|
known_usecases:
|
|
- tts
|
|
|
|
# Backend-specific options
|
|
# These are passed as "key:value" strings to the backend
|
|
options:
|
|
# Default voice to pre-load (optional)
|
|
# Can be a voice name or HuggingFace URL
|
|
# If set, this voice will be loaded when the model loads for faster first generation
|
|
- "default_voice:azelma"
|
|
# - "default_voice:hf://kyutai/tts-voices/alba-mackenna/casual.wav"
|