LocalAI/backend/cpp/llama-cpp/patches/paged at da67fd87e2f4c2aa0c8cc686b1ddba510f8f2911 - LocalAI - Gitea: Git with a cup of tea

mirror/LocalAI

mirror of https://github.com/mudler/LocalAI.git synced 2026-06-25 09:09:07 -04:00

Files

History

Ettore Di Giacinto da67fd87e2 docs(paged): A.2 CUDA-graph decode lever measurement and gap diagnosis

Phase 1 measures the CUDA-graph lever on the paged decode (q36-27b-nvfp4
dense, GB10 sm_121, fusion off). The 4-cell decode_agg {stock,paged} x
{graphs on,off} is flat within ~1%: the graphs-on win is +0.13% at npl128
and +1.1% at npl32 (both within run noise). The default paged decode is not
eager: it captures and replays graphs with a 256-token reset cadence
identical to stock non-paged (block-table ne0 = GGML_PAD(n_gather,256) only
steps at 256-token boundaries); only the gather fallback grows n_gather every
step and runs pure eager. 'graphs reused=0' was a uid fast-path false negative
(llama rebuilds the cgraph each step, so the reuse log never fires while the
graph still replays via the instance path).

nsys (reliable eager trace, plus the captured trace re-run with
--cuda-graph-trace=node to defeat nsys omitting graph-internal kernels, an
artifact that otherwise reads 0.3% busy) shows the steady decode is 99.4-99.5%
GPU-busy. Idle is ~0.6% of the step: 0.37% within-step launch gaps (the only
thing graphs remove, cut to 0.11% when captured) plus a 0.24% between-step
host gap (~2ms per step). Throughput is identical on/off.

Verdict: CUDA-graphing the paged decode is not a throughput lever; the decode
is GPU-compute-bound and the 2.6x gap to vLLM (148 vs 391) is in the per-step
GPU kernel work (FP4 GEMM + attention at batch 128), not launch overhead or
the host loop.

Assisted-by: Claude:opus-4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

2026-06-24 21:26:16 +00:00

..

0001-vendor-paged-kv-manager.patch

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

0002-paged-kv-block-placement-env-LLAMA_KV_PAGED.patch

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

0003-gather-read-plan.md

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

0003-paged-gather-read-env-LLAMA_KV_PAGED.patch

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

0004-paged-on-demand-block-allocation-env-LLAMA_KV_PAGED.patch

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

0006-paged-cross-request-prefix-caching-env-LLAMA_KV_PAGED.patch

feat(llama-cpp/paged): cross-request prefix caching patch 0006

2026-06-22 10:14:27 +00:00

0007-paged-engine-prefix-recompute-skip-env-LLAMA_KV_PAGED.patch

feat(llama-cpp/paged): engine-level prefix recompute-skip (patch 0007)

2026-06-22 10:47:10 +00:00

0008-paged-server-cross-request-prefix-share-env-LLAMA_KV_PAGED.patch

feat(paged): wire cross-request prefix share into llama-server (patch 0008)

2026-06-22 15:03:16 +00:00

0009-paged-in-kernel-decode-read-env-LLAMA_KV_PAGED-patch.patch

paged: in-kernel decode read patch 0009 (kill the gather regression)

2026-06-22 18:04:09 +00:00

0010-paged-tile-in-kernel-read-and-dispatch-guard-env-LLAMA_KV_PAGED.patch

feat(paged): tile in-kernel decode read + dispatch guard (patch 0010)

2026-06-22 20:37:12 +00:00

0011-paged-decode-route-GQA-grouped-tile-kernel-by-defaul.patch

feat(paged): route GQA-grouped tile kernel by default for paged decode (patch 0011)

2026-06-22 22:38:28 +00:00

0012-paged-mask-pad-invariant-assert.patch

feat(paged): assert mask-pad invariant for the paged tile route (patch 0012)

2026-06-23 09:13:08 +00:00

0013-paged-decoupled-prefill-token-budget.patch

feat(paged): add patch 0013 decoupled per-step prefill-token budget

2026-06-23 09:55:32 +00:00

0014-paged-expert-aware-moe-token-tile-cap.patch

feat(paged): mirror patch 0014 - expert-aware MoE token-tile cap

2026-06-23 13:49:15 +00:00

0015-paged-expert-density-aware-moe-token-tile-auto-select.patch

feat(paged): mirror MoE token-tile density-aware auto-select (patch 0015)

2026-06-23 19:04:55 +00:00

0016-paged-dynamic-prefill-budget-continuous-batch.patch

feat(llama-cpp/paged): dynamic decode-first prefill budget (patch 0016, continuous-batch P1)

2026-06-24 07:48:20 +00:00

0017-fp4-gemm-decode-tile-tune.patch

docs(paged): mirror FP4 decode-GEMM track-B P0 gate + P1 kill-gate results (patch 0017)

2026-06-24 17:58:00 +00:00

A2_CUDAGRAPH_DECODE.md

docs(paged): A.2 CUDA-graph decode lever measurement and gap diagnosis

2026-06-24 21:26:16 +00:00

ADDITIVE_DESIGN.md

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

CONTINUOUS_BATCH_SCHEDULER_SCOPE.md

docs(paged): adversarial review of the continuous-batch scheduler scope

2026-06-23 22:48:31 +00:00

DECODE_GAP_STUDY.md

docs(paged): decode-step gap study vs vLLM on GB10

2026-06-22 15:44:24 +00:00

FP4_GEMM_SCOPE_B.md

docs(paged): adversarial review of track-B FP4-GEMM parity go/no-go

2026-06-24 14:31:35 +00:00

GDN_DECODE_VERIFY.md

docs(paged): verify llama.cpp GDN decode is O(1)-in-context, not a 2.4x lever

2026-06-24 11:21:44 +00:00

MOE_DENSITY_AUTO_TILE.md

feat(paged): mirror MoE token-tile density-aware auto-select (patch 0015)

2026-06-23 19:04:55 +00:00

MOE_GROUPED_GEMM_SCOPE.md

docs(paged): scope durable grouped FP4-MMA MoE GEMM port for GB10

2026-06-23 13:17:03 +00:00

MOE_TOKEN_TILE_CAP.md

feat(paged): mirror patch 0014 - expert-aware MoE token-tile cap

2026-06-23 13:49:15 +00:00

P1_DYNAMIC_BUDGET_RESULTS.md

docs(paged): staggered-arrival evaluation of patch 0016 dynamic budget

2026-06-24 10:56:13 +00:00

PAGED_BENCH.md

docs(llama-cpp/paged): GPU 0007 re-run + shared-prefix benchmark results

2026-06-22 12:59:09 +00:00

PAGED_GPU_VERIFY.md

docs(paged): record GPU correctness + CUDA backend-build verification

2026-06-22 11:50:01 +00:00

PAGED_VLLM_APPLES.md

docs(paged): apples-to-apples paged llama.cpp vs vLLM (batched+NVFP4+prefix cache)

2026-06-22 14:16:52 +00:00

PAGED_VLLM_COMPARE.md

docs(paged): stock GPU batch-shape determinism + vLLM shared-prefix comparison

2026-06-22 13:48:01 +00:00

QWEN36_NVFP4_BENCH.md

docs(paged): fair re-run verdict - synthesize NVFP4 llama vs vLLM scorecard

2026-06-23 21:39:22 +00:00

SERVER_SWEEP.md

docs(paged): GB10 head-to-head server sweep (llama-server vs vLLM)

2026-06-23 12:22:15 +00:00

THROUGHPUT_B_P1_RESULTS.md

docs(paged): mirror FP4 decode-GEMM track-B P0 gate + P1 kill-gate results (patch 0017)

2026-06-24 17:58:00 +00:00

VLLM_DECODE_GROUNDING.md

docs(paged): ground vLLM 0.23.0 eager-decode architecture vs llama.cpp

2026-06-24 07:44:07 +00:00