* feat(sycl): make the intel llama.cpp backend self-contained on any host The SYCL backend shipped an incomplete oneAPI runtime AND relied on a host-provided GPU driver, so it only ran inside the build container. On a bare host it died with "libze_loader.so.1 / libdnnl.so.3: cannot open shared object file", and even with the host's Intel driver installed it SIGSEGV'd during SYCL init when the host driver was built against a newer glibc than the backend's bundled loader (rolling-release distros). package_intel_libs now bundles the complete, coherent oneAPI runtime (the missing MKL ILP64 / sycl_blas / tbb_thread + oneDNN + the dlopen'd UR adapters, plus a sweep of the backend binaries' own direct deps) and the Intel GPU userspace driver (libze_intel_gpu + libigdrcl + IGC + gmm) with its OpenCL ICD manifest, mirroring how package_vulkan_libs bundles Mesa. run.sh points the Level Zero and OpenCL loaders at the bundled driver, and install-base-deps.sh installs it in the SYCL build image. Bundling the driver is safe across kernels because it talks to the host i915/xe via the stable DRM UAPI (unlike NVIDIA's kernel-locked userspace). Validated on Arch (glibc 2.43, i915): the backend loads and runs on an Iris Xe with no host Intel packages installed. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me> * fix(sycl): install a driver that exists, and let the user choose their own The driver install added earlier in this branch asked apt for intel-level-zero-gpu, which is not a package in Ubuntu 24.04. apt fails outright on an unknown name, so neither driver was installed, nothing was there to copy, and the images carried no driver at all. It now comes from Intel's own repository, which has 25.18 for this Ubuntu release, against 23.43 from late 2023 in the Ubuntu archive. The archive driver does not know any card released since, so a machine with a recent Intel GPU would end up carrying a driver that cannot drive it. Anything that goes wrong during that install fails the build on purpose: an unreachable repository is a passing problem that a retry fixes, while quietly carrying a different driver, or none, is a difference nobody would notice until a user reports an idle GPU. run.sh used to overwrite whatever driver the user had chosen. Level Zero uses only the driver it is given, so on a machine with a card too new for the carried driver, the GPU would go unused with no way back. Both that setting and the OpenCL one are now left alone when already set, and the docs say how to point a backend at the machine's own driver. The OpenCL setting also used to be applied whenever the backend held a driver list, even when the driver it named had not been copied, which leaves OpenCL with nothing instead of falling back to the machine's own driver. It now requires the copied driver to be present, and the packaging leaves out the list entry of any driver it did not copy. The oneAPI images list a processor-only OpenCL library, which was being carried with nothing behind it. Two more corrections in the packaging. The scan for libraries a program is linked against only looked at files named llama-cpp-*, so turboquant and bonsai, which are also built for Intel GPUs, were left with the incomplete set of libraries this branch set out to fix; it now looks at every program in the directory. And a build that should carry a driver but ends up without one now says so, which is what a stale prebuilt base image looks like: such a backend still runs on a machine that has its own driver, so nothing fails and the only other symptom is a user reporting an idle GPU. Backends now also ask the driver to report how much graphics memory is free, without which llama.cpp reads zero on an integrated GPU, since such a chip shares the system memory instead of having its own. turboquant and bonsai get the same run.sh handling as llama.cpp. The driver is only carried by the builds that start through run.sh, because run.sh is what points Level Zero and OpenCL at it. The Python backends for Intel GPUs start differently and would never load it, so they keep using the machine's own driver rather than carrying several hundred megabytes they cannot use. Checked in a container on Ubuntu 24.04: the install brings driver 25.18 with the files where the packaging expects them, an unreachable repository fails the build, and the copied set resolves on its own once the machine's Intel packages are moved away. Assisted-by: Claude:claude-opus-5 Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me> * fix(ci): rebuild every Linux backend when the GPU packaging script changes scripts/build/package-gpu-libs.sh decides which GPU libraries end up inside an image. The filter that builds the backend matrix listed it as an input of the Python images only, so changing it rebuilt no Go and no C++ backend, even though those run it from their own package.sh. A packaging fix aimed at the Intel llama.cpp backend could merge and reach no image, which is the same failure this rule was written to prevent. Assisted-by: Claude:claude-opus-5 Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me> * fix(sycl): carry only the driver Level Zero uses, not the OpenCL one llama.cpp reaches an Intel GPU through Level Zero, which hands the driver programs that are already compiled and so needs only the back end of the graphics compiler. The OpenCL driver can be handed source code instead, so it needs the compiler's front end as well, and that arrives with its own copy of clang. Carrying it cost about 139 MB in every backend built for Intel GPUs, and took the carried set from 123 MB to 261 MB. Nothing here takes that path. No LocalAI code selects an OpenCL device, each backend image holds one backend, and the documentation never described OpenCL as a way to run models: the only mentions are a stale clblas row in the BUILD_TYPE table, for a llama.cpp backend that no longer exists and that no build matrix entry uses, and the sycl-ls troubleshooting hint. Before this branch the packaging carried the OpenCL loader and adapter but no driver, so the path could not work in a released image either. There is nobody to keep working. The driver list that OpenCL reads is no longer carried, and run.sh no longer sets OCL_ICD_VENDORS, so OpenCL inside a container keeps using whatever the image provides rather than being pointed at a directory with no driver in it. Checked in a container against the real 25.18 driver: the carried set is 123 MB with nothing unresolved, and Level Zero still reports the GPU with the machine's own Intel packages moved out of the way. Neither the Level Zero driver nor the compiler back end names the front end or clang among the libraries it opens by name, so the leaner set is complete for this path. Assisted-by: Claude:claude-opus-5 Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me> --------- Signed-off-by: Dimitris Karakasilis <dimitris@karakasilis.me> Co-authored-by: localai-org-maint-bot <bot-opensource@localaisrl.com>
LocalAI Backend Architecture
This directory contains the core backend infrastructure for LocalAI, including the gRPC protocol definition, multi-language Dockerfiles, and language-specific backend implementations.
Overview
LocalAI uses a unified gRPC-based architecture that allows different programming languages to implement AI backends while maintaining consistent interfaces and capabilities. The backend system supports multiple hardware acceleration targets and provides a standardized way to integrate various AI models and frameworks.
Architecture Components
1. Protocol Definition (backend.proto)
The backend.proto file defines the gRPC service interface that all backends must implement. This ensures consistency across different language implementations and provides a contract for communication between LocalAI core and backend services.
Core Services
- Text Generation:
Predict,PredictStreamfor LLM inference - Embeddings:
Embeddingfor text vectorization - Image Generation:
GenerateImagefor stable diffusion and image models - Audio Processing:
AudioTranscription,TTS,SoundGeneration - Video Generation:
GenerateVideofor video synthesis - Object Detection:
Detectfor computer vision tasks - Vector Storage:
StoresSet,StoresGet,StoresFindfor RAG operations - Reranking:
Rerankfor document relevance scoring - Voice Activity Detection:
VADfor audio segmentation
Key Message Types
PredictOptions: Comprehensive configuration for text generationModelOptions: Model loading and configuration parametersResult: Standardized response formatStatusResponse: Backend health and memory usage information
2. Multi-Language Dockerfiles
The backend system provides language-specific Dockerfiles that handle the build environment and dependencies for different programming languages:
Dockerfile.pythonDockerfile.golangDockerfile.llama-cpp
3. Language-Specific Implementations
Python Backends (python/)
- transformers: Hugging Face Transformers framework
- vllm: High-performance LLM inference
- mlx: Apple Silicon optimization
- diffusers: Stable Diffusion models
- longcat-video: CUDA text/image-to-video and speech-driven avatar generation
- Audio: coqui, faster-whisper, kitten-tts
- Vision: mlx-vlm, rfdetr
- Specialized: rerankers, chatterbox, kokoro
Go Backends (go/)
- whisper: OpenAI Whisper speech recognition in Go with GGML cpp backend (whisper.cpp)
- stablediffusion-ggml: Stable Diffusion in Go with GGML Cpp backend
- piper: Text-to-speech synthesis Golang with C bindings using rhaspy/piper
- local-store: Vector storage backend
- valkey-store: Durable vector storage backend backed by Valkey Search (FT.*)
C++ Backends (cpp/)
- llama-cpp: Llama.cpp integration
- grpc: GRPC utilities and helpers
Hardware Acceleration Support
CUDA (NVIDIA)
- Versions: CUDA 12.x, 13.x
- Features: cuBLAS, cuDNN, TensorRT optimization
- Targets: x86_64, ARM64 (Jetson)
ROCm (AMD)
- Features: HIP, rocBLAS, MIOpen
- Targets: AMD GPUs with ROCm support
Intel
- Features: oneAPI, Intel Extension for PyTorch
- Targets: Intel GPUs, XPUs, CPUs
Vulkan
- Features: Cross-platform GPU acceleration
- Targets: Windows, Linux, Android, macOS
Apple Silicon
- Features: MLX framework, Metal Performance Shaders
- Targets: M1/M2/M3 Macs
Backend Registry (index.yaml)
The index.yaml file serves as a central registry for all available backends, providing:
- Metadata: Name, description, license, icons
- Capabilities: Hardware targets and optimization profiles
- Tags: Categorization for discovery
- URLs: Source code and documentation links
Building Backends
Prerequisites
- Docker with multi-architecture support
- Appropriate hardware drivers (CUDA, ROCm, etc.)
- Build tools (make, cmake, compilers)
Build Commands
Example of build commands with Docker
# Build Python backend
docker build -f backend/Dockerfile.python \
--build-arg BACKEND=transformers \
--build-arg BUILD_TYPE=cublas12 \
--build-arg CUDA_MAJOR_VERSION=12 \
--build-arg CUDA_MINOR_VERSION=0 \
-t localai-backend-transformers .
# Build Go backend
docker build -f backend/Dockerfile.golang \
--build-arg BACKEND=whisper \
--build-arg BUILD_TYPE=cpu \
-t localai-backend-whisper .
# Build C++ backend
docker build -f backend/Dockerfile.llama-cpp \
--build-arg BACKEND=llama-cpp \
--build-arg BUILD_TYPE=cublas12 \
-t localai-backend-llama-cpp .
For ARM64/Mac builds, docker can't be used, and the makefile in the respective backend has to be used.
Build Types
cpu: CPU-only optimizationcublas12,cublas13: CUDA 12.x, 13.x with cuBLAShipblas: ROCm with rocBLASintel: Intel oneAPI optimizationvulkan: Vulkan-based accelerationmetal: Apple Metal optimization
Backend Development
Creating a New Backend
- Choose Language: Select Python, Go, or C++ based on requirements
- Implement Interface: Implement the gRPC service defined in
backend.proto - Add Dependencies: Create appropriate requirements files
- Configure Build: Set up Dockerfile and build scripts
- Register Backend: Add entry to
index.yaml - Test Integration: Verify gRPC communication and functionality
Backend Structure
backend-name/
├── backend.py/go/cpp # Main implementation
├── requirements.txt # Dependencies
├── Dockerfile # Build configuration
├── install.sh # Installation script
├── run.sh # Execution script
├── test.sh # Test script
└── README.md # Backend documentation
Required gRPC Methods
At minimum, backends must implement:
Health()- Service health checkLoadModel()- Model loading and initializationPredict()- Main inference endpointStatus()- Backend status and metrics
Integration with LocalAI Core
Backends communicate with LocalAI core through gRPC:
- Service Discovery: Core discovers available backends
- Model Loading: Core requests model loading via
LoadModel - Inference: Core sends requests via
Predictor specialized endpoints - Streaming: Core handles streaming responses for real-time generation
- Monitoring: Core tracks backend health and performance
Performance Optimization
Memory Management
- Model Caching: Efficient model loading and caching
- Batch Processing: Optimize for multiple concurrent requests
- Memory Pinning: GPU memory optimization for CUDA/ROCm
Hardware Utilization
- Multi-GPU: Support for tensor parallelism
- Mixed Precision: FP16/BF16 for memory efficiency
- Kernel Fusion: Optimized CUDA/ROCm kernels
Troubleshooting
Common Issues
- GRPC Connection: Verify backend service is running and accessible
- Model Loading: Check model paths and dependencies
- Hardware Detection: Ensure appropriate drivers and libraries
- Memory Issues: Monitor GPU memory usage and model sizes
Contributing
When contributing to the backend system:
- Follow Protocol: Implement the exact gRPC interface
- Add Tests: Include comprehensive test coverage
- Document: Provide clear usage examples
- Optimize: Consider performance and resource usage
- Validate: Test across different hardware targets