From 975ff5ca42b5e2fd4d0966ea4470a6c3559f4439 Mon Sep 17 00:00:00 2001 From: Ettore Di Giacinto Date: Wed, 30 Sep 2026 22:28:07 +0000 Subject: [PATCH] feat(gallery): add kev-0.8b-vllm-cpp decision model Add the kev 0.8B decision model, converted for vllm.cpp and pinned to revision c17e7366 of mudler/kev-0.8b-vllm-cpp. It is a redistribution of jaredpalmer/kev-0.8b with the LoRA merged and the PointerHead stored as head.safetensors, so only the vllm-cpp backend can load it. The artifact sits under overrides, where the installer reads it. The entry sets a 2048-token context and an explicit KV pool: with the default 4096-token context the CPU KV pool holds only 4064 tokens and the load fails. List the entry in the decisions gallery table and drop kev from the list of decision models that are not gallery entries yet. Signed-off-by: Ettore Di Giacinto Assisted-by: Claude Code:claude-sonnet-5-5 --- docs/content/features/decisions.md | 6 ++-- gallery/index.yaml | 54 ++++++++++++++++++++++++++++++ 2 files changed, 58 insertions(+), 2 deletions(-) diff --git a/docs/content/features/decisions.md b/docs/content/features/decisions.md index 2893d47dc..2feb30c44 100644 --- a/docs/content/features/decisions.md +++ b/docs/content/features/decisions.md @@ -107,10 +107,12 @@ Install one from the gallery and filter on the `decisions` tag: | `gliner25-decide-vllm-cpp` | GLiNER2.5-Decide | DeBERTa-v3-large with a classification head, about 2 GB | | `tev1-4b-vllm-cpp` | Tev1 4B | Autoregressive Qwen3.5-4B fine-tune that answers with an option letter, about 9.3 GB | | `tev1-0.8b-vllm-cpp` | Tev1 0.8B | Autoregressive Qwen3.5-0.8B fine-tune that answers with an option letter, about 1.8 GB | +| `kev-0.8b-vllm-cpp` | kev 0.8B | Qwen3.5-0.8B-Base with a merged LoRA and a PointerHead readout, converted for vllm.cpp only, about 1.53 GB | The engine, [vllm.cpp]({{% relref "features/vllm-cpp" %}}), also supports the -kev, CLM and xor decision models. Those checkpoints need a conversion step, so -they are not gallery entries yet. +CLM and xor decision models. Those checkpoints need a conversion step, so +they are not gallery entries yet. The kev entry installs a checkpoint that was +already converted with the vllm.cpp `convert-kev.py` script. Tev1 is an autoregressive decision model. The engine answers each question by scoring the option letters, so its `confidence` is the entropy measure Ollama diff --git a/gallery/index.yaml b/gallery/index.yaml index 104585e01..eb37a7d12 100644 --- a/gallery/index.yaml +++ b/gallery/index.yaml @@ -63821,6 +63821,60 @@ type: huggingface repo: togethercomputer/Tev1-0.8B-experimental revision: 6bb2dff14b38fea90ddb14d870166ccaf77374e9 +- name: kev-0.8b-vllm-cpp + url: github:mudler/LocalAI/gallery/virtual.yaml@master + urls: + - https://huggingface.co/mudler/kev-0.8b-vllm-cpp + - https://huggingface.co/jaredpalmer/kev-0.8b + - https://github.com/mudler/vllm.cpp + description: | + kev is a System 1 decision model by Jared Palmer. It answers typed choice, + noul and score questions about a text state with one scoring pass per + question. It does not generate text. The model is a frozen + Qwen3.5-0.8B-Base backbone, a rank-16 LoRA adapter and a PointerHead + readout. + + This entry installs a converted redistribution of jaredpalmer/kev-0.8b: + the LoRA is merged into the BF16 backbone, the head is stored as + head.safetensors, and config.json names the KevModel architecture. The + checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers, + vLLM and llama.cpp cannot load it. + + In LocalAI, serve it via POST /v1/systemone. The vllm.cpp project records + PointerHead golden-vector tests (25 cases) and a 5-case end-to-end + comparison against the kev reference server as equal. The upload itself + was smoke-tested with one request on CPU; there is no accuracy benchmark + and no GPU run. The entry sets a 2048-token context and an explicit KV + pool, because the default 4096-token context does not fit the default + CPU KV pool and the load fails. BF16 weights, about 1.53 GB, pinned to a + revision. + license: apache-2.0 + tags: + - decisions + - systemone + - vllm-cpp + - cpu + - gpu + size: 1.53GB + last_checked: "2026-09-30" + overrides: + backend: vllm-cpp + known_usecases: + - decisions + context_size: 2048 + engine_args: + block_size: 32 + num_blocks: 256 + max_num_seqs: 4 + parameters: + model: mudler/kev-0.8b-vllm-cpp + artifacts: + - name: model + target: model + source: + type: huggingface + repo: mudler/kev-0.8b-vllm-cpp + revision: c17e73666ded1e9d284470eae7e0de9a27294e77 - name: qwen3-vl-4b-vllm-cpp url: github:mudler/LocalAI/gallery/virtual.yaml@master urls: