* feat(voice): list registered voices and record which encoder made them The voice registry could register, identify and forget but not list, and it did not remember which speaker encoder produced an embedding. Add Metadata.Model and Registry.List, answered from the index the store registry already keeps for Forget. Needed so a backend can be given the registered voices that match its own speaker encoder. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(voice): store the encoder model with a registered voice Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(voice): pick the registered voices that match a speaker model Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(proto): carry known voices and speaker names on diarize and live messages Assisted-by: Claude:claude-haiku-4-5 [Claude Code] * feat(diarization): name speakers from the voice registry When a diarization model has a speaker_model option, the endpoint sends the registered voices made by that encoder to the backend. The backend's name and name_score come back as extra fields next to the normalized SPEAKER_NN speaker, and the speakers summary carries the first name seen for each speaker. RTTM output and results without names are unchanged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(live): pass registered voices to a live session and surface speaker names Live sessions now send the registered voices that match the model's speaker_model to the backend, and each speaker segment carries the name the backend matched. The realtime segment event gains an optional speaker_name field. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): load a speaker model and build per-request voice registries Adds the speaker bindings (ABI v9 and v10, probed separately), the speaker_model, speaker_threshold and speaker_margin options, and a per-request registry builder over the known voices. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): name the speakers in Diarize from the known voices Diarize builds a per-request speaker registry from the known voices when a speaker model is loaded, calls the named C functions, and puts each slot's registered name and score on the segments. The registry is freed on every path. A library without ABI 10 reports Unimplemented instead of dropping the names. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): name speakers in the live scene stream The live scene stream now begins with a known-voice registry when a speaker model is loaded and the live config carries voices, and each closed speaker segment takes its slot's current name from the feed's names map. A segment that closes before its slot is identified has an empty name. The registry is freed after the stream, on every path. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(gallery): speaker naming entries and docs for parakeet-cpp Add three gallery entries that load the WeSpeaker ResNet34 speaker model next to the diarization or realtime scene models, and document speaker names in the voice recognition, diarization, audio to text and realtime pages. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(parakeet-cpp): skip an unusable registered voice instead of failing the request A registered voice with the wrong embedding size, or one the C side refused, failed the whole diarization request, so one legacy voice broke the model for every user. Skip such voices with a warning that does not carry the voice name, and take the plain path when none is left. Also map an exact 0 speaker threshold or margin to a tiny positive value, since the C side reads 0 as "use the default", and fix a stale comment about which contexts Free() walks. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * fix(diarization): warn once per model about voices from another encoder; document the privacy limit The different-encoder warning fired on every request. Log it once per feature and speaker model, then at debug level. Document that the global voice registry lets any caller of a speaker_model model learn matching names, and that skipped wrong-sized voices are logged. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * chore(parakeet-cpp): bump parakeet.cpp to 8c8cec0 (C-API v10) and check speaker naming against the real library The pin moves from 623a968 to 8c8cec0, which brings in everything merged in parakeet.cpp since: the voice identification change (C-API v9, #78) and raw-embedding enroll plus diarize-only speaker naming (C-API v10, #79). New real-library specs (gated on PARAKEET_BACKEND_TEST_SPEAKER_MODEL, _DIAR_MODEL, _WAV and, for the live path, _STREAM_MODEL) name the two speakers of two_speakers.wav from a committed pair of WeSpeaker embeddings, with the voices passed in reversed order. They also check that the float32 threshold reaches C through purego. The shared test loader now registers the v9/v10 and scene symbols as main.go does. The rebase onto origin/master had no conflicts. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
LocalAI Backend Architecture
This directory contains the core backend infrastructure for LocalAI, including the gRPC protocol definition, multi-language Dockerfiles, and language-specific backend implementations.
Overview
LocalAI uses a unified gRPC-based architecture that allows different programming languages to implement AI backends while maintaining consistent interfaces and capabilities. The backend system supports multiple hardware acceleration targets and provides a standardized way to integrate various AI models and frameworks.
Architecture Components
1. Protocol Definition (backend.proto)
The backend.proto file defines the gRPC service interface that all backends must implement. This ensures consistency across different language implementations and provides a contract for communication between LocalAI core and backend services.
Core Services
- Text Generation:
Predict,PredictStreamfor LLM inference - Embeddings:
Embeddingfor text vectorization - Image Generation:
GenerateImagefor stable diffusion and image models - Audio Processing:
AudioTranscription,TTS,SoundGeneration - Video Generation:
GenerateVideofor video synthesis - Object Detection:
Detectfor computer vision tasks - Vector Storage:
StoresSet,StoresGet,StoresFindfor RAG operations - Reranking:
Rerankfor document relevance scoring - Voice Activity Detection:
VADfor audio segmentation
Key Message Types
PredictOptions: Comprehensive configuration for text generationModelOptions: Model loading and configuration parametersResult: Standardized response formatStatusResponse: Backend health and memory usage information
2. Multi-Language Dockerfiles
The backend system provides language-specific Dockerfiles that handle the build environment and dependencies for different programming languages:
Dockerfile.pythonDockerfile.golangDockerfile.llama-cpp
3. Language-Specific Implementations
Python Backends (python/)
- transformers: Hugging Face Transformers framework
- vllm: High-performance LLM inference
- mlx: Apple Silicon optimization
- diffusers: Stable Diffusion models
- longcat-video: CUDA text/image-to-video and speech-driven avatar generation
- Audio: coqui, faster-whisper, funasr, kitten-tts
- Vision: mlx-vlm, rfdetr
- Specialized: rerankers, chatterbox, kokoro
Go Backends (go/)
- whisper: OpenAI Whisper speech recognition in Go with GGML cpp backend (whisper.cpp)
- stablediffusion-ggml: Stable Diffusion in Go with GGML Cpp backend
- piper: Text-to-speech synthesis Golang with C bindings using rhaspy/piper
- local-store: Vector storage backend
- valkey-store: Durable vector storage backend backed by Valkey Search (FT.*)
C++ Backends (cpp/)
- llama-cpp: Llama.cpp integration
- grpc: GRPC utilities and helpers
Hardware Acceleration Support
CUDA (NVIDIA)
- Versions: CUDA 12.x, 13.x
- Features: cuBLAS, cuDNN, TensorRT optimization
- Targets: x86_64, ARM64 (Jetson)
ROCm (AMD)
- Features: HIP, rocBLAS, MIOpen
- Targets: AMD GPUs with ROCm support
Intel
- Features: oneAPI, Intel Extension for PyTorch
- Targets: Intel GPUs, XPUs, CPUs
Vulkan
- Features: Cross-platform GPU acceleration
- Targets: Windows, Linux, Android, macOS
Apple Silicon
- Features: MLX framework, Metal Performance Shaders
- Targets: M1/M2/M3 Macs
Backend Registry (index.yaml)
The index.yaml file serves as a central registry for all available backends, providing:
- Metadata: Name, description, license, icons
- Capabilities: Hardware targets and optimization profiles
- Tags: Categorization for discovery
- URLs: Source code and documentation links
Building Backends
Prerequisites
- Docker with multi-architecture support
- Appropriate hardware drivers (CUDA, ROCm, etc.)
- Build tools (make, cmake, compilers)
Build Commands
Example of build commands with Docker
# Build Python backend
docker build -f backend/Dockerfile.python \
--build-arg BACKEND=transformers \
--build-arg BUILD_TYPE=cublas12 \
--build-arg CUDA_MAJOR_VERSION=12 \
--build-arg CUDA_MINOR_VERSION=0 \
-t localai-backend-transformers .
# Build Go backend
docker build -f backend/Dockerfile.golang \
--build-arg BACKEND=whisper \
--build-arg BUILD_TYPE=cpu \
-t localai-backend-whisper .
# Build C++ backend
docker build -f backend/Dockerfile.llama-cpp \
--build-arg BACKEND=llama-cpp \
--build-arg BUILD_TYPE=cublas12 \
-t localai-backend-llama-cpp .
For ARM64/Mac builds, docker can't be used, and the makefile in the respective backend has to be used.
Build Types
cpu: CPU-only optimizationcublas12,cublas13: CUDA 12.x, 13.x with cuBLAShipblas: ROCm with rocBLASintel: Intel oneAPI optimizationvulkan: Vulkan-based accelerationmetal: Apple Metal optimization
Backend Development
Creating a New Backend
- Choose Language: Select Python, Go, or C++ based on requirements
- Implement Interface: Implement the gRPC service defined in
backend.proto - Add Dependencies: Create appropriate requirements files
- Configure Build: Set up Dockerfile and build scripts
- Register Backend: Add entry to
index.yaml - Test Integration: Verify gRPC communication and functionality
Backend Structure
backend-name/
├── backend.py/go/cpp # Main implementation
├── requirements.txt # Dependencies
├── Dockerfile # Build configuration
├── install.sh # Installation script
├── run.sh # Execution script
├── test.sh # Test script
└── README.md # Backend documentation
Required gRPC Methods
At minimum, backends must implement:
Health()- Service health checkLoadModel()- Model loading and initializationPredict()- Main inference endpointStatus()- Backend status and metrics
Integration with LocalAI Core
Backends communicate with LocalAI core through gRPC:
- Service Discovery: Core discovers available backends
- Model Loading: Core requests model loading via
LoadModel - Inference: Core sends requests via
Predictor specialized endpoints - Streaming: Core handles streaming responses for real-time generation
- Monitoring: Core tracks backend health and performance
Performance Optimization
Memory Management
- Model Caching: Efficient model loading and caching
- Batch Processing: Optimize for multiple concurrent requests
- Memory Pinning: GPU memory optimization for CUDA/ROCm
Hardware Utilization
- Multi-GPU: Support for tensor parallelism
- Mixed Precision: FP16/BF16 for memory efficiency
- Kernel Fusion: Optimized CUDA/ROCm kernels
Troubleshooting
Common Issues
- GRPC Connection: Verify backend service is running and accessible
- Model Loading: Check model paths and dependencies
- Hardware Detection: Ensure appropriate drivers and libraries
- Memory Issues: Monitor GPU memory usage and model sizes
Contributing
When contributing to the backend system:
- Follow Protocol: Implement the exact gRPC interface
- Add Tests: Include comprehensive test coverage
- Document: Provide clear usage examples
- Optimize: Consider performance and resource usage
- Validate: Test across different hardware targets