fix(llama-cpp): keep llama.cpp's default cache_ram instead of no limit (#12297)

grpc-server.cpp forced params.cache_ram_mib = -1 (no limit) since #7009.
Since v4.3 kv_unified and cache_idle_slots are on by default, so every
distinct prompt now leaves its slot KV state in the host-side prompt
cache, and without a limit the backend grows until the host runs out of
memory.

Measured on gfx1151 (Strix Halo, 128 GB), llama-cpp backend, one request
at a time, 100 distinct prompts of ~2000 characters plus a fixed system
prompt, max_tokens 200:

  model                        cache_ram     RSS loaded -> after 100
  gemma-4-26B-A4B (q8_0 KV)    -1 (default)  1.4 GB -> 25.3 GB
  Qwen3.6-35B-A3B (q8_0 KV)    -1 (default)  1.1 GB -> 19.5 GB
  gemma-4-26B-A4B              -1, same prompt 100x  1.4 GB -> 1.6 GB
  gemma-4-26B-A4B              4096          1.4 GB -> 5.4 GB (flat from
                                             request 20 on, same latency)
  Qwen3.6-35B-A3B              4096          1.1 GB -> 5.1 GB (flat)

The memory is not released when idle. In production a document
classification pass pushed the daily chat model to 34 GB RSS overnight.

Drop the override so llama.cpp's own default (8192 MiB) applies; the
cache_ram option still accepts -1 for users who want no limit. Update
both docs tables (the option reference and the prompt-cache table) and
note what -1 does.


Assisted-by: Claude:claude-opus-5-5
Assisted-by: Codex:GPT-6

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
This commit is contained in:
Stefan Walcz authored and GitHub committed 2026-09-28 11:37:24 +02:00
1 parent 50c284fdcc
commit bc1d9924de
2 files changed
+9 -5

No files matched your search

+5 -3
View File
@@ -539,8 +539,10 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
// Initialize ctx_shift to false by default (can be overridden by options)
params.ctx_shift = false;
// Initialize cache_ram_mib to -1 by default (no limit, can be overridden by options)
params.cache_ram_mib = -1;
// cache_ram_mib keeps llama.cpp's own default (8192 MiB) unless overridden by
// options. It used to be forced to -1 (no limit): since kv_unified and
// cache_idle_slots are on by default, every distinct prompt then leaves its
// slot state in host RAM and the backend grows without bound.
// Initialize n_parallel to 1 by default (can be overridden by options)
params.n_parallel = 1;
// Initialize grpc_servers to empty (can be overridden by options)
@@ -656,7 +658,7 @@ static void params_parse(server_context& /*ctx_server*/, const backend::ModelOpt
try {
params.cache_ram_mib = std::stoi(optval_str);
} catch (const std::exception& e) {
// If conversion fails, keep default value (-1)
// If conversion fails, keep the default value
}
}
} else if (!strcmp(optname, "parallel") || !strcmp(optname, "n_parallel")) {