Merge pull request #12391 from mudler/feat/gallery-kev-decisions

feat(gallery): add kev-0.8b on vllm-cpp as a decisions model
This commit is contained in:
Ettore Di Giacinto committed 2026-09-30 23:04:53 +00:00
commit 62e0827cba
2 files changed
+58 -2

No files matched your search

+4 -2
View File
@@ -107,10 +107,12 @@ Install one from the gallery and filter on the `decisions` tag:
| `gliner25-decide-vllm-cpp` | GLiNER2.5-Decide | DeBERTa-v3-large with a classification head, about 2 GB |
| `tev1-4b-vllm-cpp` | Tev1 4B | Autoregressive Qwen3.5-4B fine-tune that answers with an option letter, about 9.3 GB |
| `tev1-0.8b-vllm-cpp` | Tev1 0.8B | Autoregressive Qwen3.5-0.8B fine-tune that answers with an option letter, about 1.8 GB |
| `kev-0.8b-vllm-cpp` | kev 0.8B | Qwen3.5-0.8B-Base with a merged LoRA and a PointerHead readout, converted for vllm.cpp only, about 1.53 GB |
The engine, [vllm.cpp]({{% relref "features/vllm-cpp" %}}), also supports the
kev, CLM and xor decision models. Those checkpoints need a conversion step, so
they are not gallery entries yet.
CLM and xor decision models. Those checkpoints need a conversion step, so
they are not gallery entries yet. The kev entry installs a checkpoint that was
already converted with the vllm.cpp `convert-kev.py` script.
Tev1 is an autoregressive decision model. The engine answers each question by
scoring the option letters, so its `confidence` is the entropy measure Ollama
+54
View File
@@ -63821,6 +63821,60 @@
type: huggingface
repo: togethercomputer/Tev1-0.8B-experimental
revision: 6bb2dff14b38fea90ddb14d870166ccaf77374e9
- name: kev-0.8b-vllm-cpp
url: github:mudler/LocalAI/gallery/virtual.yaml@master
urls:
- https://huggingface.co/mudler/kev-0.8b-vllm-cpp
- https://huggingface.co/jaredpalmer/kev-0.8b
- https://github.com/mudler/vllm.cpp
description: |
kev is a System 1 decision model by Jared Palmer. It answers typed choice,
noul and score questions about a text state with one scoring pass per
question. It does not generate text. The model is a frozen
Qwen3.5-0.8B-Base backbone, a rank-16 LoRA adapter and a PointerHead
readout.
This entry installs a converted redistribution of jaredpalmer/kev-0.8b:
the LoRA is merged into the BF16 backbone, the head is stored as
head.safetensors, and config.json names the KevModel architecture. The
checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers,
vLLM and llama.cpp cannot load it.
In LocalAI, serve it via POST /v1/systemone. The vllm.cpp project records
PointerHead golden-vector tests (25 cases) and a 5-case end-to-end
comparison against the kev reference server as equal. The upload itself
was smoke-tested with one request on CPU; there is no accuracy benchmark
and no GPU run. The entry sets a 2048-token context and an explicit KV
pool, because the default 4096-token context does not fit the default
CPU KV pool and the load fails. BF16 weights, about 1.53 GB, pinned to a
revision.
license: apache-2.0
tags:
- decisions
- systemone
- vllm-cpp
- cpu
- gpu
size: 1.53GB
last_checked: "2026-09-30"
overrides:
backend: vllm-cpp
known_usecases:
- decisions
context_size: 2048
engine_args:
block_size: 32
num_blocks: 256
max_num_seqs: 4
parameters:
model: mudler/kev-0.8b-vllm-cpp
artifacts:
- name: model
target: model
source:
type: huggingface
repo: mudler/kev-0.8b-vllm-cpp
revision: c17e73666ded1e9d284470eae7e0de9a27294e77
- name: qwen3-vl-4b-vllm-cpp
url: github:mudler/LocalAI/gallery/virtual.yaml@master
urls: