mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-29 17:44:30 -04:00
grpc-server.cpp forced params.cache_ram_mib = -1 (no limit) since #7009. Since v4.3 kv_unified and cache_idle_slots are on by default, so every distinct prompt now leaves its slot KV state in the host-side prompt cache, and without a limit the backend grows until the host runs out of memory. Measured on gfx1151 (Strix Halo, 128 GB), llama-cpp backend, one request at a time, 100 distinct prompts of ~2000 characters plus a fixed system prompt, max_tokens 200: model cache_ram RSS loaded -> after 100 gemma-4-26B-A4B (q8_0 KV) -1 (default) 1.4 GB -> 25.3 GB Qwen3.6-35B-A3B (q8_0 KV) -1 (default) 1.1 GB -> 19.5 GB gemma-4-26B-A4B -1, same prompt 100x 1.4 GB -> 1.6 GB gemma-4-26B-A4B 4096 1.4 GB -> 5.4 GB (flat from request 20 on, same latency) Qwen3.6-35B-A3B 4096 1.1 GB -> 5.1 GB (flat) The memory is not released when idle. In production a document classification pass pushed the daily chat model to 34 GB RSS overnight. Drop the override so llama.cpp's own default (8192 MiB) applies; the cache_ram option still accepts -1 for users who want no limit. Update both docs tables (the option reference and the prompt-cache table) and note what -1 does. Assisted-by: Claude:claude-opus-5-5 Assisted-by: Codex:GPT-6 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>