LocalAI/backend/cpp/llama-cpp/patches/paged at 8925c009b75ee7f37914810cff438948a402e7e4 - LocalAI - Gitea: Git with a cup of tea

mirror/LocalAI

mirror of https://github.com/mudler/LocalAI.git synced 2026-06-24 00:28:55 -04:00

Files

History

Ettore Di Giacinto 8925c009b7 docs(paged): scope durable grouped FP4-MMA MoE GEMM port for GB10

Build-ready plan (not implemented) for matching/beating vLLM MoE
grouped-GEMM efficiency on GB10 sm_121 for Qwen3-30B-A3B mxfp4.

Honest reframe: the grouped GEMM the mission scoped to build already
exists upstream and runs on GB10 for mxfp4 - should_use_mmq() routes
MUL_MAT_ID to the grouped mmq path, which already contains both vLLM
building blocks (mm_ids_helper moe_align/scatter + a persistent stream-k
FP4-MMA grouped GEMM). The npl128 cliff was a since-fixed regression, not
a batched-bench artifact; re-measured decode is monotonic 85->1771 t/s.

The one structural gap is M-tile sizing: ggml maximizes mmq_x over the
aggregate token count while vLLM uses a small per-expert BLOCK_SIZE_M, so
each tiny per-expert M-tile is 3-6% filled at decode density. Scope is a
surgical two-step delta (expert-aware mmq_x selection; block-padded
moe_align), the parity gate (test_mul_mat_id bit-exact + ragged small-M),
and a phased plan gated behind the GB10 W4A16 occupancy wall.

Assisted-by: Claude:opus-4.8 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

2026-06-23 13:17:03 +00:00

..

0001-vendor-paged-kv-manager.patch

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

0002-paged-kv-block-placement-env-LLAMA_KV_PAGED.patch

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

0003-gather-read-plan.md

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

0003-paged-gather-read-env-LLAMA_KV_PAGED.patch

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

0004-paged-on-demand-block-allocation-env-LLAMA_KV_PAGED.patch

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

0006-paged-cross-request-prefix-caching-env-LLAMA_KV_PAGED.patch

feat(llama-cpp/paged): cross-request prefix caching patch 0006

2026-06-22 10:14:27 +00:00

0007-paged-engine-prefix-recompute-skip-env-LLAMA_KV_PAGED.patch

feat(llama-cpp/paged): engine-level prefix recompute-skip (patch 0007)

2026-06-22 10:47:10 +00:00

0008-paged-server-cross-request-prefix-share-env-LLAMA_KV_PAGED.patch

feat(paged): wire cross-request prefix share into llama-server (patch 0008)

2026-06-22 15:03:16 +00:00

0009-paged-in-kernel-decode-read-env-LLAMA_KV_PAGED-patch.patch

paged: in-kernel decode read patch 0009 (kill the gather regression)

2026-06-22 18:04:09 +00:00

0010-paged-tile-in-kernel-read-and-dispatch-guard-env-LLAMA_KV_PAGED.patch

feat(paged): tile in-kernel decode read + dispatch guard (patch 0010)

2026-06-22 20:37:12 +00:00

0011-paged-decode-route-GQA-grouped-tile-kernel-by-defaul.patch

feat(paged): route GQA-grouped tile kernel by default for paged decode (patch 0011)

2026-06-22 22:38:28 +00:00

0012-paged-mask-pad-invariant-assert.patch

feat(paged): assert mask-pad invariant for the paged tile route (patch 0012)

2026-06-23 09:13:08 +00:00

0013-paged-decoupled-prefill-token-budget.patch

feat(paged): add patch 0013 decoupled per-step prefill-token budget

2026-06-23 09:55:32 +00:00

ADDITIVE_DESIGN.md

build(llama-cpp): isolate paged patches in patches/paged/ behind LLAMA_PAGED flag (default on)

2026-06-22 09:22:36 +00:00

DECODE_GAP_STUDY.md

docs(paged): decode-step gap study vs vLLM on GB10

2026-06-22 15:44:24 +00:00

MOE_GROUPED_GEMM_SCOPE.md

docs(paged): scope durable grouped FP4-MMA MoE GEMM port for GB10

2026-06-23 13:17:03 +00:00

PAGED_BENCH.md

docs(llama-cpp/paged): GPU 0007 re-run + shared-prefix benchmark results

2026-06-22 12:59:09 +00:00

PAGED_GPU_VERIFY.md

docs(paged): record GPU correctness + CUDA backend-build verification

2026-06-22 11:50:01 +00:00

PAGED_VLLM_APPLES.md

docs(paged): apples-to-apples paged llama.cpp vs vLLM (batched+NVFP4+prefix cache)

2026-06-22 14:16:52 +00:00

PAGED_VLLM_COMPARE.md

docs(paged): stock GPU batch-shape determinism + vLLM shared-prefix comparison

2026-06-22 13:48:01 +00:00

SERVER_SWEEP.md

docs(paged): GB10 head-to-head server sweep (llama-server vs vLLM)

2026-06-23 12:22:15 +00:00