mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-29 09:35:02 -04:00
fix(llama-cpp): keep llama.cpp's default cache_ram instead of no limit (#12297)
grpc-server.cpp forced params.cache_ram_mib = -1 (no limit) since #7009. Since v4.3 kv_unified and cache_idle_slots are on by default, so every distinct prompt now leaves its slot KV state in the host-side prompt cache, and without a limit the backend grows until the host runs out of memory. Measured on gfx1151 (Strix Halo, 128 GB), llama-cpp backend, one request at a time, 100 distinct prompts of ~2000 characters plus a fixed system prompt, max_tokens 200: model cache_ram RSS loaded -> after 100 gemma-4-26B-A4B (q8_0 KV) -1 (default) 1.4 GB -> 25.3 GB Qwen3.6-35B-A3B (q8_0 KV) -1 (default) 1.1 GB -> 19.5 GB gemma-4-26B-A4B -1, same prompt 100x 1.4 GB -> 1.6 GB gemma-4-26B-A4B 4096 1.4 GB -> 5.4 GB (flat from request 20 on, same latency) Qwen3.6-35B-A3B 4096 1.1 GB -> 5.1 GB (flat) The memory is not released when idle. In production a document classification pass pushed the daily chat model to 34 GB RSS overnight. Drop the override so llama.cpp's own default (8192 MiB) applies; the cache_ram option still accepts -1 for users who want no limit. Update both docs tables (the option reference and the prompt-cache table) and note what -1 does. Assisted-by: Claude:claude-opus-5-5 Assisted-by: Codex:GPT-6 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
This commit is contained in:
1 parent
50c284fdcc
commit
bc1d9924de
2 files changed
+9
-5
No files matched your search
@@ -539,8 +539,10 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
|
||||
|
||||
// Initialize ctx_shift to false by default (can be overridden by options)
|
||||
params.ctx_shift = false;
|
||||
// Initialize cache_ram_mib to -1 by default (no limit, can be overridden by options)
|
||||
params.cache_ram_mib = -1;
|
||||
// cache_ram_mib keeps llama.cpp's own default (8192 MiB) unless overridden by
|
||||
// options. It used to be forced to -1 (no limit): since kv_unified and
|
||||
// cache_idle_slots are on by default, every distinct prompt then leaves its
|
||||
// slot state in host RAM and the backend grows without bound.
|
||||
// Initialize n_parallel to 1 by default (can be overridden by options)
|
||||
params.n_parallel = 1;
|
||||
// Initialize grpc_servers to empty (can be overridden by options)
|
||||
@@ -656,7 +658,7 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
|
||||
try {
|
||||
params.cache_ram_mib = std::stoi(optval_str);
|
||||
} catch (const std::exception& e) {
|
||||
// If conversion fails, keep default value (-1)
|
||||
// If conversion fails, keep the default value
|
||||
}
|
||||
}
|
||||
} else if (!strcmp(optname, "parallel") || !strcmp(optname, "n_parallel")) {
|
||||
|
||||
Reference in new issue
Block a user