mirror of
https://github.com/mudler/LocalAI.git
synced 2026-08-01 02:49:51 -04:00
* feat: add Valkey Search vector store backend Add a new built-in Go gRPC store backend 'valkey-store' that implements the four Stores RPCs (Set/Get/Delete/Find) against the Valkey Search module (FT.*) using the pure-Go github.com/valkey-io/valkey-go client. It is selected via the existing per-request 'backend' field on /stores, so there is no proto or HTTP API change, and it mirrors the in-memory local-store while adding persistence across restarts and opt-in HNSW. Each vector is a Valkey HASH keyed by hex(little-endian float32); the index is created lazily on first Set (FLAT+COSINE by default), cosine similarity is derived as 1-distance, and namespaces get a collision-resistant token. Includes unit tests (valkey-go mock) and env-gated integration tests against valkey/valkey-bundle, plus build/matrix/gallery wiring and docs. Assisted-by: Kiro:claude-opus-4.8 golangci-lint Signed-off-by: Daria Korenieva <daric2612@gmail.com> * Address review feedback: recover persisted index dimension, harden Find - Load now recovers the persisted vector DIM from FT.INFO (not just index existence), so a post-restart Set/Find validates against the real DIM instead of silently re-learning a wrong one and dropping mismatched vectors from the index. This also restores Find's dimension check after a restart. - StoresFind treats a dropped/missing index as an empty store (empty result, no error) and clears the stale indexCreated flag, matching local-store's empty-store behaviour. - StoresSet reuses checkDims for its per-key length check so the four RPCs share one dimension-guard implementation. - Add unit tests for FT.INFO dimension recovery, loadIndexState, and the dropped-index Find path. Assisted-by: Kiro:claude-opus-4.8 Signed-off-by: Daria Korenieva <daric2612@gmail.com> * Address review feedback: TLS ServerName/CA, Find nil-check, config fail-fast Addresses external review comments on the valkey-store backend: - StoresFind now rejects a nil/empty query Key before dereferencing it, so a malformed gRPC request can no longer panic the backend. - TLS: derive ServerName (SNI) from the VALKEY_ADDR host so certificate verification works for IP-addressed endpoints, and add VALKEY_TLS_CA_CERT (custom CA bundle) and VALKEY_TLS_SKIP_VERIFY (testing-only) knobs. - Config integer parsing now fails fast on a malformed value (e.g. VALKEY_HNSW_M=1x6) instead of silently defaulting, matching the fail-fast behaviour of the index-algo/distance-metric validation. - Add VALKEY_DB (SELECT n) support for logical-DB isolation. - Cap the human-readable part of a namespace token at 64 chars so a very long model name cannot produce an unbounded key prefix / index name (the appended short hash keeps distinct namespaces collision-free). - Document the KNN-query injection-safety invariant (fields are constants) and why StoresGet uses a single aggregate DoMulti deadline for reads. - Unit tests for the Find nil/empty-key guard, fail-fast HNSW parsing, and VALKEY_DB parsing/validation; docs + .env updated for the new vars. Assisted-by: Kiro:claude-opus-4.8 golangci-lint Signed-off-by: Daria Korenieva <daric2612@gmail.com> * Address review feedback: configure valkey-store via model config richiejp asked that the valkey-store backend take its configuration from a model config rather than process-wide VALKEY_* environment variables, so multiple stores can each have their own Valkey config within one LocalAI process. This removes every env access from the backend and routes config through the model-config seam every other backend uses. - config.go: loadConfig(opts *pb.ModelOptions) now parses the model config `options:` list (key:value strings, split on the first ':') instead of os.Getenv. Option keys mirror the old VALKEY_* names without the prefix (addr, index_algo, distance_metric, ...). Defaults, fail-fast validation and the mandatory client name are unchanged. - store.go: Load threads opts into loadConfig; TLS comments/errors renamed off the VALKEY_* names. - core/backend/stores.go: StoreBackend and NewVectorStore take a *config.ModelConfigLoader, resolve the per-store ModelConfig by store name, and pass its Options (and Backend when unset) to the backend via WithLoadGRPCLoadModelOpts. No config -> default backend + built-in defaults, preserving the zero-config experience. - Endpoints/routes/application: thread the config loader to StoreBackend. - Unit + integration tests: configure via options; the integration test passes addr through the model-config path (VALKEY_ADDR is now only the test harness locating the server). - docs + .env: document the model-config options, drop the env var table. Assisted-by: Kiro:claude-opus-4.8 Signed-off-by: Daria Korenieva <daric2612@gmail.com> * Remove valkey-store informational comment from .env The backend is configured via model config, not env vars — the comment was unnecessary noise in .env. The configuration is already documented in docs/content/features/stores.md. Signed-off-by: Daria Korenieva <daric2612@gmail.com> * feat(valkey-store): gate Load on NamespacePrefix to refuse autoload probing Mirror local-store's pattern: reject model names without store.NamespacePrefix so the model loader's greedy autoload probe cannot bind an arbitrary model name to the vector store backend (the #9287 failure mode). Also adds unit tests for the gate covering: prefixed namespace, prefix alone, unprefixed model name, empty model, and nil opts. Signed-off-by: Daria Korenieva <daric2612@gmail.com> * feat(valkey-store): add username_env/password_env credential indirection Add support for resolving Valkey credentials from environment variables named in the model config, mirroring cloud-proxy's api_key_env pattern. This keeps secrets out of model YAML files and lets distinct store configs each reference their own credentials. Options: username_env / password_env name the env var holding the value. The direct username / password options still work and take precedence when both are set (backward compatible). Includes 5 unit tests and updated stores.md documentation. Signed-off-by: Daria Korenieva <daric2612@gmail.com> * fix: correct rebase artifacts in backend-matrix.yml and Makefile Fix two issues introduced by the conflict-resolution script during the rebase onto master: 1. .github/backend-matrix.yml: valkey-store entries were merged INTO the cloud-proxy entries (duplicate keys in same YAML map items) instead of being separate list items. This broke cloud-proxy Linux builds and the cloud-proxy darwin entry lost its build-type/lang. Fixed by making them standalone entries and restoring cloud-proxy exactly as on master. 2. Makefile: duplicated .NOTPARALLEL and docker-build-backends lines. Collapsed to single lines that are master's current content plus the valkey-store additions. Also adds the three optional pickups from #10801: - /valkey-store in .gitignore (the built binary) - valkey-store row in docs/content/reference/compatibility-table.md - valkey-store line in backend/README.md Signed-off-by: Daria Korenieva <daric2612@gmail.com> --------- Signed-off-by: Daria Korenieva <daric2612@gmail.com> Co-authored-by: Daria Korenieva <daric2612@gmail.com>
213 lines
7.5 KiB
Markdown
213 lines
7.5 KiB
Markdown
# LocalAI Backend Architecture
|
|
|
|
This directory contains the core backend infrastructure for LocalAI, including the gRPC protocol definition, multi-language Dockerfiles, and language-specific backend implementations.
|
|
|
|
## Overview
|
|
|
|
LocalAI uses a unified gRPC-based architecture that allows different programming languages to implement AI backends while maintaining consistent interfaces and capabilities. The backend system supports multiple hardware acceleration targets and provides a standardized way to integrate various AI models and frameworks.
|
|
|
|
## Architecture Components
|
|
|
|
### 1. Protocol Definition (`backend.proto`)
|
|
|
|
The `backend.proto` file defines the gRPC service interface that all backends must implement. This ensures consistency across different language implementations and provides a contract for communication between LocalAI core and backend services.
|
|
|
|
#### Core Services
|
|
|
|
- **Text Generation**: `Predict`, `PredictStream` for LLM inference
|
|
- **Embeddings**: `Embedding` for text vectorization
|
|
- **Image Generation**: `GenerateImage` for stable diffusion and image models
|
|
- **Audio Processing**: `AudioTranscription`, `TTS`, `SoundGeneration`
|
|
- **Video Generation**: `GenerateVideo` for video synthesis
|
|
- **Object Detection**: `Detect` for computer vision tasks
|
|
- **Vector Storage**: `StoresSet`, `StoresGet`, `StoresFind` for RAG operations
|
|
- **Reranking**: `Rerank` for document relevance scoring
|
|
- **Voice Activity Detection**: `VAD` for audio segmentation
|
|
|
|
#### Key Message Types
|
|
|
|
- **`PredictOptions`**: Comprehensive configuration for text generation
|
|
- **`ModelOptions`**: Model loading and configuration parameters
|
|
- **`Result`**: Standardized response format
|
|
- **`StatusResponse`**: Backend health and memory usage information
|
|
|
|
### 2. Multi-Language Dockerfiles
|
|
|
|
The backend system provides language-specific Dockerfiles that handle the build environment and dependencies for different programming languages:
|
|
|
|
- `Dockerfile.python`
|
|
- `Dockerfile.golang`
|
|
- `Dockerfile.llama-cpp`
|
|
|
|
### 3. Language-Specific Implementations
|
|
|
|
#### Python Backends (`python/`)
|
|
- **transformers**: Hugging Face Transformers framework
|
|
- **vllm**: High-performance LLM inference
|
|
- **mlx**: Apple Silicon optimization
|
|
- **diffusers**: Stable Diffusion models
|
|
- **longcat-video**: CUDA text/image-to-video and speech-driven avatar generation
|
|
- **Audio**: coqui, faster-whisper, kitten-tts
|
|
- **Vision**: mlx-vlm, rfdetr
|
|
- **Specialized**: rerankers, chatterbox, kokoro
|
|
|
|
#### Go Backends (`go/`)
|
|
- **whisper**: OpenAI Whisper speech recognition in Go with GGML cpp backend (whisper.cpp)
|
|
- **stablediffusion-ggml**: Stable Diffusion in Go with GGML Cpp backend
|
|
- **piper**: Text-to-speech synthesis Golang with C bindings using rhaspy/piper
|
|
- **local-store**: Vector storage backend
|
|
- **valkey-store**: Durable vector storage backend backed by Valkey Search (FT.*)
|
|
|
|
#### C++ Backends (`cpp/`)
|
|
- **llama-cpp**: Llama.cpp integration
|
|
- **grpc**: GRPC utilities and helpers
|
|
|
|
## Hardware Acceleration Support
|
|
|
|
### CUDA (NVIDIA)
|
|
- **Versions**: CUDA 12.x, 13.x
|
|
- **Features**: cuBLAS, cuDNN, TensorRT optimization
|
|
- **Targets**: x86_64, ARM64 (Jetson)
|
|
|
|
### ROCm (AMD)
|
|
- **Features**: HIP, rocBLAS, MIOpen
|
|
- **Targets**: AMD GPUs with ROCm support
|
|
|
|
### Intel
|
|
- **Features**: oneAPI, Intel Extension for PyTorch
|
|
- **Targets**: Intel GPUs, XPUs, CPUs
|
|
|
|
### Vulkan
|
|
- **Features**: Cross-platform GPU acceleration
|
|
- **Targets**: Windows, Linux, Android, macOS
|
|
|
|
### Apple Silicon
|
|
- **Features**: MLX framework, Metal Performance Shaders
|
|
- **Targets**: M1/M2/M3 Macs
|
|
|
|
## Backend Registry (`index.yaml`)
|
|
|
|
The `index.yaml` file serves as a central registry for all available backends, providing:
|
|
|
|
- **Metadata**: Name, description, license, icons
|
|
- **Capabilities**: Hardware targets and optimization profiles
|
|
- **Tags**: Categorization for discovery
|
|
- **URLs**: Source code and documentation links
|
|
|
|
## Building Backends
|
|
|
|
### Prerequisites
|
|
- Docker with multi-architecture support
|
|
- Appropriate hardware drivers (CUDA, ROCm, etc.)
|
|
- Build tools (make, cmake, compilers)
|
|
|
|
### Build Commands
|
|
|
|
Example of build commands with Docker
|
|
|
|
```bash
|
|
# Build Python backend
|
|
docker build -f backend/Dockerfile.python \
|
|
--build-arg BACKEND=transformers \
|
|
--build-arg BUILD_TYPE=cublas12 \
|
|
--build-arg CUDA_MAJOR_VERSION=12 \
|
|
--build-arg CUDA_MINOR_VERSION=0 \
|
|
-t localai-backend-transformers .
|
|
|
|
# Build Go backend
|
|
docker build -f backend/Dockerfile.golang \
|
|
--build-arg BACKEND=whisper \
|
|
--build-arg BUILD_TYPE=cpu \
|
|
-t localai-backend-whisper .
|
|
|
|
# Build C++ backend
|
|
docker build -f backend/Dockerfile.llama-cpp \
|
|
--build-arg BACKEND=llama-cpp \
|
|
--build-arg BUILD_TYPE=cublas12 \
|
|
-t localai-backend-llama-cpp .
|
|
```
|
|
|
|
For ARM64/Mac builds, docker can't be used, and the makefile in the respective backend has to be used.
|
|
|
|
### Build Types
|
|
|
|
- **`cpu`**: CPU-only optimization
|
|
- **`cublas12`**, **`cublas13`**: CUDA 12.x, 13.x with cuBLAS
|
|
- **`hipblas`**: ROCm with rocBLAS
|
|
- **`intel`**: Intel oneAPI optimization
|
|
- **`vulkan`**: Vulkan-based acceleration
|
|
- **`metal`**: Apple Metal optimization
|
|
|
|
## Backend Development
|
|
|
|
### Creating a New Backend
|
|
|
|
1. **Choose Language**: Select Python, Go, or C++ based on requirements
|
|
2. **Implement Interface**: Implement the gRPC service defined in `backend.proto`
|
|
3. **Add Dependencies**: Create appropriate requirements files
|
|
4. **Configure Build**: Set up Dockerfile and build scripts
|
|
5. **Register Backend**: Add entry to `index.yaml`
|
|
6. **Test Integration**: Verify gRPC communication and functionality
|
|
|
|
### Backend Structure
|
|
|
|
```
|
|
backend-name/
|
|
├── backend.py/go/cpp # Main implementation
|
|
├── requirements.txt # Dependencies
|
|
├── Dockerfile # Build configuration
|
|
├── install.sh # Installation script
|
|
├── run.sh # Execution script
|
|
├── test.sh # Test script
|
|
└── README.md # Backend documentation
|
|
```
|
|
|
|
### Required gRPC Methods
|
|
|
|
At minimum, backends must implement:
|
|
- `Health()` - Service health check
|
|
- `LoadModel()` - Model loading and initialization
|
|
- `Predict()` - Main inference endpoint
|
|
- `Status()` - Backend status and metrics
|
|
|
|
## Integration with LocalAI Core
|
|
|
|
Backends communicate with LocalAI core through gRPC:
|
|
|
|
1. **Service Discovery**: Core discovers available backends
|
|
2. **Model Loading**: Core requests model loading via `LoadModel`
|
|
3. **Inference**: Core sends requests via `Predict` or specialized endpoints
|
|
4. **Streaming**: Core handles streaming responses for real-time generation
|
|
5. **Monitoring**: Core tracks backend health and performance
|
|
|
|
## Performance Optimization
|
|
|
|
### Memory Management
|
|
- **Model Caching**: Efficient model loading and caching
|
|
- **Batch Processing**: Optimize for multiple concurrent requests
|
|
- **Memory Pinning**: GPU memory optimization for CUDA/ROCm
|
|
|
|
### Hardware Utilization
|
|
- **Multi-GPU**: Support for tensor parallelism
|
|
- **Mixed Precision**: FP16/BF16 for memory efficiency
|
|
- **Kernel Fusion**: Optimized CUDA/ROCm kernels
|
|
|
|
## Troubleshooting
|
|
|
|
### Common Issues
|
|
|
|
1. **GRPC Connection**: Verify backend service is running and accessible
|
|
2. **Model Loading**: Check model paths and dependencies
|
|
3. **Hardware Detection**: Ensure appropriate drivers and libraries
|
|
4. **Memory Issues**: Monitor GPU memory usage and model sizes
|
|
|
|
## Contributing
|
|
|
|
When contributing to the backend system:
|
|
|
|
1. **Follow Protocol**: Implement the exact gRPC interface
|
|
2. **Add Tests**: Include comprehensive test coverage
|
|
3. **Document**: Provide clear usage examples
|
|
4. **Optimize**: Consider performance and resource usage
|
|
5. **Validate**: Test across different hardware targets
|