mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-29 09:35:02 -04:00
chore: merge master into distributed transport PR
Keep the version-reporting import from master and omit the unused sanitize import after the transport changes. Assisted-by: Codex:GPT-6
This commit is contained in:
commit
4151ceda85
51 files changed
+1785
-72
No files matched your search
@@ -74,6 +74,8 @@ When using `--models-config-file`, you can define multiple models as a list:
|
||||
backend: llama-cpp
|
||||
```
|
||||
|
||||
LocalAI changes only config files that are inside the models directory. If the file from `--models-config-file` is outside the models directory, you cannot view, edit, pin, enable or disable its models from the web UI or the model admin API. Edit the file directly, then restart LocalAI.
|
||||
|
||||
## Core Configuration Fields
|
||||
|
||||
### Basic Model Settings
|
||||
|
||||
@@ -243,7 +243,7 @@ The devices in the following list have been tested with `hipblas` images.
|
||||
|
||||
1. Check your GPU LLVM target is compatible with the version of ROCm. This can be found in the [LLVM Docs](https://llvm.org/docs/AMDGPUUsage.html).
|
||||
2. Check which ROCm version is compatible with your LLVM target and your chosen OS (pay special attention to supported kernel versions). See the [ROCm compatibility matrix](https://rocm.docs.amd.com/en/latest/compatibility/compatibility-matrix.html).
|
||||
3. Install your chosen version of the `dkms` and `rocm` (it is recommended that the native package manager be used for this process for any OS as version changes are executed more easily via this method if updates are required). Take care to restart after installing `amdgpu-dkms` and before installing `rocm`, for details regarding this see the [ROCm installation documentation](https://rocm.docs.amd.com/projects/install-on-linux/en/latest/how-to/native-install/index.html).
|
||||
3. Install your chosen version of the `dkms` and `rocm` (it is recommended that the native package manager be used for this process for any OS as version changes are executed more easily via this method if updates are required). Take care to restart after installing `amdgpu-dkms` and before installing `rocm`, for details regarding this see the [ROCm installation documentation](https://rocm.docs.amd.com/projects/install-on-linux/en/latest/install/install-methods/package-manager-index.html).
|
||||
4. Deploy. Yes it's that easy.
|
||||
|
||||
#### Setup Example (Docker/containerd)
|
||||
|
||||
@@ -740,6 +740,19 @@ Set `LOCALAI_DISTRIBUTED_SHARED_MODELS=true` (or `--distributed-shared-models`)
|
||||
|
||||
This flag is a contract you assert: all nodes must mount identical paths. Leave it off (the default) when workers have independent models directories - the frontend stages files to them over HTTP (or S3) as described above.
|
||||
|
||||
### Which files are staged
|
||||
|
||||
The frontend stages the files that the model config names (`parameters.model`, `mmproj`, draft model, LoRA adapters and similar fields). It also stages every other file that the model declares:
|
||||
|
||||
- The `files:` of the gallery entry or `/import-model` import that installed the model. LocalAI records these in `._gallery_<name>.yaml` next to the model config.
|
||||
- The `download_files:` of the model config.
|
||||
|
||||
A backend can read files that the config does not name. For example, llama.cpp opens all shards of a split GGUF (`<name>-00002-of-00004.gguf` and the rest) from the directory of the first shard. The worker cannot see the frontend's models directory, so it gets only the files that the frontend stages.
|
||||
|
||||
If you write a model config by hand and the model has files like these, list them under `download_files:`. If you do not, the worker gets only the first shard and the load fails with `failed to load GGUF split`.
|
||||
|
||||
The file sizes used for the load deadline and for the disk headroom check include all of these files.
|
||||
|
||||
### Model artifact staging
|
||||
|
||||
For managed Hugging Face artifacts, the controller resolves the repository and
|
||||
|
||||
@@ -98,7 +98,7 @@ LLAMACPP_GRPC_SERVERS="address1:port,address2:port" local-ai run
|
||||
```
|
||||
The workload on the LocalAI server will then be distributed across the specified nodes.
|
||||
|
||||
Alternatively, you can build the RPC workers/server following the llama.cpp [README](https://github.com/ggerganov/llama.cpp/blob/master/examples/rpc/README.md), which is compatible with LocalAI.
|
||||
Alternatively, you can build the RPC workers/server following the llama.cpp [README](https://github.com/ggml-org/llama.cpp/blob/master/tools/rpc/README.md), which is compatible with LocalAI.
|
||||
|
||||
## Manual example (worker)
|
||||
|
||||
|
||||
@@ -373,7 +373,7 @@ curl $LOCALAI/models/apply -H "Content-Type: application/json" -d '{
|
||||
where:
|
||||
- `localai` is the repository. It is optional and can be omitted. If the repository is omitted LocalAI will search the model by name in all the repositories. In the case the same model name is present in both galleries the first match wins.
|
||||
- `bert-embeddings` is the model name in the gallery
|
||||
(read its [config here](https://github.com/mudler/LocalAI/tree/master/gallery/blob/main/bert-embeddings.yaml)).
|
||||
(read its [config here](https://github.com/mudler/LocalAI/blob/master/gallery/index.yaml)).
|
||||
|
||||
### Model variants
|
||||
|
||||
|
||||
@@ -587,7 +587,7 @@ The `llama.cpp` backend supports additional configuration options that can be sp
|
||||
|--------|------|-------------|---------|
|
||||
| `use_jinja` or `jinja` | boolean | Enable Jinja2 template processing for chat templates. When enabled, the backend uses Jinja2-based chat templates from the model for formatting messages. | `use_jinja:true` |
|
||||
| `context_shift` | boolean | Enable context shifting, which allows the model to dynamically adjust context window usage. | `context_shift:true` |
|
||||
| `cache_ram` | integer | Size budget in MiB for the **server-side prompt cache** (a host-RAM store of idle slot KV states that's reloaded on a prompt-prefix hit, see [upstream PR #16391](https://github.com/ggml-org/llama.cpp/pull/16391)). Default: `-1` (no limit). `0` disables the prompt cache entirely. Together with `kv_unified` and `cache_idle_slots` this is what makes a repeated system prompt skip prefill on subsequent calls. | `cache_ram:4096` |
|
||||
| `cache_ram` | integer | Size budget in MiB for the **server-side prompt cache** (a host-RAM store of idle slot KV states that's reloaded on a prompt-prefix hit, see [upstream PR #16391](https://github.com/ggml-org/llama.cpp/pull/16391)). Default: `8192` MiB (llama.cpp default). `-1` removes the limit. `0` disables the prompt cache entirely. Together with `kv_unified` and `cache_idle_slots` this is what makes a repeated system prompt skip prefill on subsequent calls. | `cache_ram:4096` |
|
||||
| `parallel` or `n_parallel` | integer | Enable parallel request processing. When set to a value greater than 1, enables continuous batching for handling multiple requests concurrently. | `parallel:4` |
|
||||
| `grpc_servers` or `rpc_servers` | string | Comma-separated list of gRPC server addresses for distributed inference. Allows distributing workload across multiple llama.cpp workers. | `grpc_servers:localhost:50051,localhost:50052` |
|
||||
| `fit_params` or `fit` | boolean | Enable auto-adjustment of model/context parameters to fit available device memory. Default: `true`. | `fit_params:true` |
|
||||
@@ -643,7 +643,7 @@ Agents, coding assistants, and Anthropic/OpenAI-compatible CLIs typically resend
|
||||
|
||||
| Setting | Default | Role |
|
||||
|---|---|---|
|
||||
| `cache_ram:N` | `-1` (no limit) | Allocates the host-side prompt cache. `0` disables it. |
|
||||
| `cache_ram:N` | `8192` (llama.cpp default) | Allocates the host-side prompt cache. `0` disables it. |
|
||||
| `kv_unified:true` | `true` | Single unified KV buffer (**prerequisite** for idle-slot saving). |
|
||||
| `cache_idle_slots:true` | `true` | Persists the idle slot's KV into the prompt cache on task switch. |
|
||||
|
||||
@@ -658,6 +658,8 @@ options:
|
||||
|
||||
Set `cache_ram:0` to opt out of the prompt cache entirely (saves host RAM at the cost of re-prefilling repeated prompts).
|
||||
|
||||
`cache_ram:-1` removes the limit. With idle-slot saving on, every distinct prompt then leaves its slot state in host RAM, so a workload with many different prompts (classification, ingestion) grows the backend by roughly the KV size of each prompt until the host runs out of memory.
|
||||
|
||||
#### Reference
|
||||
|
||||
- [llama](https://github.com/ggerganov/llama.cpp)
|
||||
@@ -988,6 +990,41 @@ options:
|
||||
The full list of registered parsers lives in `sglang.srt.function_call`
|
||||
and `sglang.srt.parser.reasoning_parser`.
|
||||
|
||||
#### Reasoning defaults and token budgets
|
||||
|
||||
Set SGLang reasoning options in the model's `options:` list:
|
||||
|
||||
```yaml
|
||||
options:
|
||||
- reasoning_parser:qwen3
|
||||
- thinking_budget:512
|
||||
- reasoning_default:on
|
||||
engine_args:
|
||||
enable_strict_thinking: true
|
||||
```
|
||||
|
||||
`thinking_budget` sets a positive integer token budget for reasoning on each request.
|
||||
Invalid, zero, and negative values produce a warning and leave the budget unset.
|
||||
SGLang requires `engine_args.enable_strict_thinking: true` to enforce the budget.
|
||||
LocalAI warns if you configure a budget without that engine option.
|
||||
Keep the budget well below the `max_tokens` of your requests: if `max_tokens` is reached first,
|
||||
the budget never triggers and the whole reply can be spent on reasoning, leaving the answer empty.
|
||||
|
||||
`reasoning_default:on` or `reasoning_default:off` sets the default for LocalAI's tokenizer chat template.
|
||||
Request metadata `enable_thinking` set to `"true"` or `"false"` overrides this default.
|
||||
An explicit prompt bypasses tokenizer template rendering.
|
||||
When no default or request override is set, the template keeps its own behavior.
|
||||
|
||||
LocalAI signals required reasoning when the rendered prompt ends with the configured parser's opening reasoning token.
|
||||
An explicit output grammar disables this detection.
|
||||
Configure a reasoning parser that matches your model.
|
||||
|
||||
The backend reads these options when it loads the model.
|
||||
`POST /models/reload` rereads model configuration files but does not update options in an already loaded backend.
|
||||
Restarting only the backend does not reread configuration files.
|
||||
Restart LocalAI after changing these options to reload both the configuration and the backend.
|
||||
|
||||
|
||||
### vllm.cpp
|
||||
|
||||
[vllm.cpp](https://github.com/mudler/vllm.cpp) is the LocalAI team's C++ port of
|
||||
|
||||
@@ -108,6 +108,8 @@ docker run -ti --name local-ai -p 8080:8080 --runtime nvidia --gpus all localai/
|
||||
|
||||
## Using Compose
|
||||
|
||||
The repository's `docker-compose.yaml` installs `phi-2-chat` from the model gallery by default. Change its `command` list to select a different gallery model.
|
||||
|
||||
For a more manageable setup, especially with persistent volumes, use Docker Compose or Podman Compose:
|
||||
|
||||
### Using CDI (Container Device Interface) - Recommended for NVIDIA Container Toolkit 1.14+
|
||||
|
||||
@@ -25,17 +25,17 @@ Here's an example to initiate the **phi-2** model:
|
||||
docker run -p 8080:8080 localai/localai:{{< version >}} https://gist.githubusercontent.com/mudler/ad601a0488b497b69ec549150d9edd18/raw/a8a8869ef1bb7e3830bf5c0bae29a0cce991ff8d/phi-2.yaml
|
||||
```
|
||||
|
||||
You can also check all the embedded models configurations [here](https://github.com/mudler/LocalAI/tree/master/embedded/models).
|
||||
You can also check all the embedded models configurations [here](https://github.com/mudler/LocalAI/tree/master/gallery).
|
||||
|
||||
{{% notice tip %}}
|
||||
The model configurations used in the quickstart are accessible here: [https://github.com/mudler/LocalAI/tree/master/embedded/models](https://github.com/mudler/LocalAI/tree/master/embedded/models). Contributions are welcome; please feel free to submit a Pull Request.
|
||||
The model configurations used in the quickstart are accessible here: [https://github.com/mudler/LocalAI/tree/master/gallery](https://github.com/mudler/LocalAI/tree/master/gallery). Contributions are welcome; please feel free to submit a Pull Request.
|
||||
|
||||
The `phi-2` model configuration from the quickstart is expanded from [https://github.com/mudler/LocalAI/blob/master/examples/configurations/phi-2.yaml](https://github.com/mudler/LocalAI/blob/master/examples/configurations/phi-2.yaml).
|
||||
The `phi-2` model configuration from the quickstart is expanded from [https://github.com/mudler/LocalAI-examples/blob/main/configurations/phi-2.yaml](https://github.com/mudler/LocalAI-examples/blob/main/configurations/phi-2.yaml).
|
||||
{{% /notice %}}
|
||||
|
||||
## Example: Customizing the Prompt Template
|
||||
|
||||
To modify the prompt template, create a Github gist or a Pastebin file, and copy the content from [https://github.com/mudler/LocalAI/blob/master/examples/configurations/phi-2.yaml](https://github.com/mudler/LocalAI/blob/master/examples/configurations/phi-2.yaml). Alter the fields as needed:
|
||||
To modify the prompt template, create a Github gist or a Pastebin file, and copy the content from [https://github.com/mudler/LocalAI-examples/blob/main/configurations/phi-2.yaml](https://github.com/mudler/LocalAI-examples/blob/main/configurations/phi-2.yaml). Alter the fields as needed:
|
||||
|
||||
```yaml
|
||||
name: phi-2
|
||||
|
||||
@@ -98,7 +98,7 @@ availability may lag upstream releases.
|
||||
- [AnythingLLM](https://github.com/Mintplex-Labs/anything-llm)
|
||||
- [Logseq GPT3 OpenAI plugin](https://github.com/briansunter/logseq-plugin-gpt3-openai)
|
||||
- [CodeGPT (JetBrains)](https://plugins.jetbrains.com/plugin/21056-codegpt) - Custom OpenAI-compatible endpoints
|
||||
- [Wave Terminal](https://docs.waveterm.dev/features/supportedLLMs/localai) - Native LocalAI support
|
||||
- [Wave Terminal](https://docs.waveterm.dev/ai-presets) - Native LocalAI support
|
||||
- [Obsidian BMO Chatbot](https://github.com/longy2k/obsidian-bmo-chatbot)
|
||||
- [spark](https://github.com/cedriking/spark)
|
||||
- [openops (Mattermost)](https://github.com/mattermost/openops)
|
||||
|
||||
@@ -45,7 +45,7 @@ All backends listed here can be installed on demand from the [Backend Gallery]({
|
||||
| [moonshine](https://github.com/moonshine-ai/moonshine) | Ultra-fast transcription for low-end devices (ONNX) | CPU, CUDA 12/13, Metal |
|
||||
| [parakeet.cpp](https://github.com/mudler/parakeet.cpp) | C++/GGML port of NVIDIA NeMo Parakeet (tdt/ctc/rnnt/hybrid), with cache-aware streaming | CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T |
|
||||
| [CrispASR](https://github.com/CrispStrobe/CrispASR) | Unified speech engine (whisper.cpp fork) supporting Parakeet, Canary, and many ASR architectures, plus TTS | CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T |
|
||||
| [voxtral](https://github.com/mudler/voxtral.c) | Voxtral Realtime 4B speech-to-text in pure C | CPU, Metal |
|
||||
| [voxtral](https://github.com/antirez/voxtral.c) | Voxtral Realtime 4B speech-to-text in pure C | CPU, Metal |
|
||||
| [Qwen3-ASR](https://github.com/QwenLM/Qwen3-ASR) | Qwen3 automatic speech recognition | CPU, CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T |
|
||||
| [NeMo](https://github.com/NVIDIA/NeMo) | NVIDIA NeMo ASR toolkit | CPU, CUDA 12/13, ROCm, Intel SYCL, Metal |
|
||||
| [sherpa-onnx](https://k2-fsa.github.io/sherpa/onnx/) | Sherpa-ONNX ASR (Whisper, Paraformer, SenseVoice) and TTS | CPU, CUDA 12, Metal |
|
||||
@@ -70,10 +70,10 @@ All backends listed here can be installed on demand from the [Backend Gallery]({
|
||||
| [OmniVoice](https://github.com/ServeurpersoCom/omnivoice.cpp) | Native C++/GGML TTS with voice cloning, voice design, and streaming | CPU, CUDA 12/13, ROCm, Intel SYCL, Vulkan, Metal, Jetson L4T |
|
||||
| [fish-speech](https://github.com/fishaudio/fish-speech) | High-quality TTS with voice cloning | CPU, CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T |
|
||||
| [Pocket TTS](https://github.com/kyutai-labs/pocket-tts) | Lightweight CPU-efficient TTS with voice cloning | CPU, CUDA 12/13, ROCm, Intel SYCL, Metal, Jetson L4T |
|
||||
| [OuteTTS](https://github.com/OuteAI/outetts) | TTS with custom speaker voices | CPU, CUDA 12 |
|
||||
| [OuteTTS](https://github.com/edwko/OuteTTS) | TTS with custom speaker voices | CPU, CUDA 12 |
|
||||
| [faster-qwen3-tts](https://github.com/andimarafioti/faster-qwen3-tts) | Real-time Qwen3-TTS with CUDA graph capture | CPU, CUDA 12/13, Jetson L4T |
|
||||
| [NeuTTS Air](https://github.com/neuphonic/neutts-air) | Instant voice cloning, on-device TTS | CPU, CUDA 12, ROCm |
|
||||
| [VoxCPM](https://github.com/ModelBest/VoxCPM) | Expressive end-to-end TTS | CPU, CUDA 12/13, ROCm, Intel SYCL, Metal |
|
||||
| [VoxCPM](https://github.com/OpenBMB/VoxCPM) | Expressive end-to-end TTS | CPU, CUDA 12/13, ROCm, Intel SYCL, Metal |
|
||||
| [Kitten TTS](https://github.com/KittenML/KittenTTS) | Kitten TTS model | CPU, Metal |
|
||||
| [Supertonic](https://github.com/supertone-inc/supertonic) | Lightning-fast on-device multilingual TTS via ONNX | CPU |
|
||||
| [MLX-Audio](https://github.com/Blaizzy/mlx-audio) | Audio models on Apple Silicon | CPU, CUDA 12/13, Metal, Jetson L4T |
|
||||
|
||||
Reference in new issue
Block a user