LocalAI/docs/content/features/backends.md at 13f59f082299a3248e79b87f18c7eb2c4bc809fc

mirror of https://github.com/mudler/LocalAI.git synced 2026-06-18 21:58:58 -04:00

Files

LocalAI [bot] 13f59f0822 docs: document the privacy-filter.cpp backend (#10386 )

docs: document the privacy-filter.cpp backend in README and compatibility table

The privacy-filter.cpp backend (#10360) was registered in backend/index.yaml
and referenced from the PII feature docs, but was missing from the backend
catalog surfaces. Add it to the README "Backends built by us" table, the
compatibility table (Utilities & Other, CPU/CUDA 13/Vulkan), and the backend
type list in the backends feature doc.

Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>

2026-06-18 15:07:01 +02:00

5.7 KiB

Raw Blame History

title, description, weight, url

title	description	weight	url
Backends	Learn how to use, manage, and develop backends in LocalAI	4	/backends/

LocalAI supports a variety of backends that can be used to run different types of AI models. There are core Backends which are included, and there are containerized applications that provide the runtime environment for specific model types, such as LLMs, diffusion models, or text-to-speech models.

Available Backends

LocalAI ships 60+ backends covering text generation, speech-to-text, text-to-speech, music and sound generation, image and video generation, vision and object detection, audio processing, reranking, fine-tuning, and more. Each one is published as an on-demand OCI image with the appropriate acceleration variants (CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T).

For the complete list of backends, the model families they support, and their acceleration targets, see the [Backend & Model Compatibility Table]({{%relref "reference/compatibility-table" %}}). The authoritative source is backend/index.yaml, and the same catalog is browsable in the web UI under the Backends section.

Managing Backends in the UI

The LocalAI web interface provides an intuitive way to manage your backends:

Navigate to the "Backends" section in the navigation menu
Browse available backends from configured galleries
Use the search bar to find specific backends by name, description, or type
Filter backends by type using the quick filter buttons (LLM, Diffusion, TTS, Whisper)
Install or delete backends with a single click
Monitor installation progress in real-time

Each backend card displays:

Backend name and description
Type of models it supports
Installation status
Action buttons (Install/Delete)
Additional information via the info button

Backend Galleries

Backend galleries are repositories that contain backend definitions. They work similarly to model galleries but are specifically for backends.

Adding a Backend Gallery

You can add backend galleries by specifying the Environment Variable LOCALAI_BACKEND_GALLERIES:

export LOCALAI_BACKEND_GALLERIES='[{"name":"my-gallery","url":"https://raw.githubusercontent.com/username/repo/main/backends"}]'

The URL needs to point to a valid yaml file, for example:

- name: "test-backend"
  uri: "quay.io/image/tests:localai-backend-test"
  alias: "foo-backend"

Where URI is the path to an OCI container image.

Backend Gallery Structure

A backend gallery is a collection of YAML files, each defining a backend. Here's an example structure:

name: "llm-backend"
description: "A backend for running LLM models"
uri: "quay.io/username/llm-backend:latest"
alias: "llm"
tags:
  - "llm"
  - "text-generation"

Pre-installing Backends

You can pre-install backends when starting LocalAI using the LOCALAI_EXTERNAL_BACKENDS environment variable:

export LOCALAI_EXTERNAL_BACKENDS="llm-backend,diffusion-backend"
local-ai run

Creating a Backend

To create a new backend, you need to:

Create a container image that implements the LocalAI backend interface
Define a backend YAML file
Publish your backend to a container registry

Backend Container Requirements

Your backend container should:

Implement the LocalAI backend interface (gRPC or HTTP)
Handle model loading and inference
Support the required model types
Include necessary dependencies
Have a top level run.sh file that will be used to run the backend
Pushed to a registry so can be used in a gallery

Getting started

For getting started, see the available backends in LocalAI here: https://github.com/mudler/LocalAI/tree/master/backend .

For Python based backends there is a template that can be used as starting point: https://github.com/mudler/LocalAI/tree/master/backend/python/common/template .
For Golang based backends, you can see the piper backend as an example: https://github.com/mudler/LocalAI/tree/master/backend/go/piper
For C++ based backends, you can see the llama-cpp backend as an example: https://github.com/mudler/LocalAI/tree/master/backend/cpp/llama-cpp

Publishing Your Backend

Build your container image:

docker build -t quay.io/username/my-backend:latest .

Push to a container registry:

docker push quay.io/username/my-backend:latest

Add your backend to a gallery:
- Create a YAML entry in your gallery repository
- Include the backend definition
- Make the gallery accessible via HTTP/HTTPS

Backend Types

LocalAI supports various types of backends:

LLM Backends: For running language models (e.g., llama.cpp, vLLM, SGLang, transformers, MLX)
Speech-to-Text Backends: For transcription (e.g., whisper.cpp, parakeet.cpp, faster-whisper, NeMo)
Text-to-Speech Backends: For speech synthesis (e.g., piper, Kokoro, VibeVoice, Qwen3-TTS)
Sound Generation Backends: For music and audio generation (e.g., ACE-Step)
Image & Video Generation Backends: For diffusion models (e.g., stable-diffusion.cpp, diffusers)
Vision & Detection Backends: For object detection, segmentation, depth, and face/voice recognition (e.g., rf-detr.cpp, locate-anything.cpp, sam3.cpp, insightface)
Audio Processing Backends: For voice activity detection and audio enhancement (e.g., Silero VAD, LocalVQE)
Utility Backends: For reranking, PII/NER token classification, fine-tuning, quantization, and vector storage (e.g., rerankers, privacy-filter.cpp, TRL, local-store)

See the [Backend & Model Compatibility Table]({{%relref "reference/compatibility-table" %}}) for the full catalog.

5.7 KiB Raw Blame History