mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-29 01:25:03 -04:00
fix(llama-cpp): keep llama.cpp's default cache_ram instead of no limit (#12297)
grpc-server.cpp forced params.cache_ram_mib = -1 (no limit) since #7009. Since v4.3 kv_unified and cache_idle_slots are on by default, so every distinct prompt now leaves its slot KV state in the host-side prompt cache, and without a limit the backend grows until the host runs out of memory. Measured on gfx1151 (Strix Halo, 128 GB), llama-cpp backend, one request at a time, 100 distinct prompts of ~2000 characters plus a fixed system prompt, max_tokens 200: model cache_ram RSS loaded -> after 100 gemma-4-26B-A4B (q8_0 KV) -1 (default) 1.4 GB -> 25.3 GB Qwen3.6-35B-A3B (q8_0 KV) -1 (default) 1.1 GB -> 19.5 GB gemma-4-26B-A4B -1, same prompt 100x 1.4 GB -> 1.6 GB gemma-4-26B-A4B 4096 1.4 GB -> 5.4 GB (flat from request 20 on, same latency) Qwen3.6-35B-A3B 4096 1.1 GB -> 5.1 GB (flat) The memory is not released when idle. In production a document classification pass pushed the daily chat model to 34 GB RSS overnight. Drop the override so llama.cpp's own default (8192 MiB) applies; the cache_ram option still accepts -1 for users who want no limit. Update both docs tables (the option reference and the prompt-cache table) and note what -1 does. Assisted-by: Claude:claude-opus-5-5 Assisted-by: Codex:GPT-6 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
This commit is contained in:
1 parent
50c284fdcc
commit
bc1d9924de
2 files changed
+9
-5
No files matched your search
@@ -539,8 +539,10 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
|
||||
|
||||
// Initialize ctx_shift to false by default (can be overridden by options)
|
||||
params.ctx_shift = false;
|
||||
// Initialize cache_ram_mib to -1 by default (no limit, can be overridden by options)
|
||||
params.cache_ram_mib = -1;
|
||||
// cache_ram_mib keeps llama.cpp's own default (8192 MiB) unless overridden by
|
||||
// options. It used to be forced to -1 (no limit): since kv_unified and
|
||||
// cache_idle_slots are on by default, every distinct prompt then leaves its
|
||||
// slot state in host RAM and the backend grows without bound.
|
||||
// Initialize n_parallel to 1 by default (can be overridden by options)
|
||||
params.n_parallel = 1;
|
||||
// Initialize grpc_servers to empty (can be overridden by options)
|
||||
@@ -656,7 +658,7 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
|
||||
try {
|
||||
params.cache_ram_mib = std::stoi(optval_str);
|
||||
} catch (const std::exception& e) {
|
||||
// If conversion fails, keep default value (-1)
|
||||
// If conversion fails, keep the default value
|
||||
}
|
||||
}
|
||||
} else if (!strcmp(optname, "parallel") || !strcmp(optname, "n_parallel")) {
|
||||
|
||||
@@ -586,7 +586,7 @@ The `llama.cpp` backend supports additional configuration options that can be sp
|
||||
|--------|------|-------------|---------|
|
||||
| `use_jinja` or `jinja` | boolean | Enable Jinja2 template processing for chat templates. When enabled, the backend uses Jinja2-based chat templates from the model for formatting messages. | `use_jinja:true` |
|
||||
| `context_shift` | boolean | Enable context shifting, which allows the model to dynamically adjust context window usage. | `context_shift:true` |
|
||||
| `cache_ram` | integer | Size budget in MiB for the **server-side prompt cache** (a host-RAM store of idle slot KV states that's reloaded on a prompt-prefix hit, see [upstream PR #16391](https://github.com/ggml-org/llama.cpp/pull/16391)). Default: `-1` (no limit). `0` disables the prompt cache entirely. Together with `kv_unified` and `cache_idle_slots` this is what makes a repeated system prompt skip prefill on subsequent calls. | `cache_ram:4096` |
|
||||
| `cache_ram` | integer | Size budget in MiB for the **server-side prompt cache** (a host-RAM store of idle slot KV states that's reloaded on a prompt-prefix hit, see [upstream PR #16391](https://github.com/ggml-org/llama.cpp/pull/16391)). Default: `8192` MiB (llama.cpp default). `-1` removes the limit. `0` disables the prompt cache entirely. Together with `kv_unified` and `cache_idle_slots` this is what makes a repeated system prompt skip prefill on subsequent calls. | `cache_ram:4096` |
|
||||
| `parallel` or `n_parallel` | integer | Enable parallel request processing. When set to a value greater than 1, enables continuous batching for handling multiple requests concurrently. | `parallel:4` |
|
||||
| `grpc_servers` or `rpc_servers` | string | Comma-separated list of gRPC server addresses for distributed inference. Allows distributing workload across multiple llama.cpp workers. | `grpc_servers:localhost:50051,localhost:50052` |
|
||||
| `fit_params` or `fit` | boolean | Enable auto-adjustment of model/context parameters to fit available device memory. Default: `true`. | `fit_params:true` |
|
||||
@@ -642,7 +642,7 @@ Agents, coding assistants, and Anthropic/OpenAI-compatible CLIs typically resend
|
||||
|
||||
| Setting | Default | Role |
|
||||
|---|---|---|
|
||||
| `cache_ram:N` | `-1` (no limit) | Allocates the host-side prompt cache. `0` disables it. |
|
||||
| `cache_ram:N` | `8192` (llama.cpp default) | Allocates the host-side prompt cache. `0` disables it. |
|
||||
| `kv_unified:true` | `true` | Single unified KV buffer (**prerequisite** for idle-slot saving). |
|
||||
| `cache_idle_slots:true` | `true` | Persists the idle slot's KV into the prompt cache on task switch. |
|
||||
|
||||
@@ -657,6 +657,8 @@ options:
|
||||
|
||||
Set `cache_ram:0` to opt out of the prompt cache entirely (saves host RAM at the cost of re-prefilling repeated prompts).
|
||||
|
||||
`cache_ram:-1` removes the limit. With idle-slot saving on, every distinct prompt then leaves its slot state in host RAM, so a workload with many different prompts (classification, ingestion) grows the backend by roughly the KV size of each prompt until the host runs out of memory.
|
||||
|
||||
#### Reference
|
||||
|
||||
- [llama](https://github.com/ggerganov/llama.cpp)
|
||||
|
||||
Reference in new issue
Block a user