mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-02 19:14:38 -04:00
Merge pull request #12391 from mudler/feat/gallery-kev-decisions
feat(gallery): add kev-0.8b on vllm-cpp as a decisions model
This commit is contained in:
commit
62e0827cba
2 files changed
+58
-2
No files matched your search
@@ -107,10 +107,12 @@ Install one from the gallery and filter on the `decisions` tag:
|
||||
| `gliner25-decide-vllm-cpp` | GLiNER2.5-Decide | DeBERTa-v3-large with a classification head, about 2 GB |
|
||||
| `tev1-4b-vllm-cpp` | Tev1 4B | Autoregressive Qwen3.5-4B fine-tune that answers with an option letter, about 9.3 GB |
|
||||
| `tev1-0.8b-vllm-cpp` | Tev1 0.8B | Autoregressive Qwen3.5-0.8B fine-tune that answers with an option letter, about 1.8 GB |
|
||||
| `kev-0.8b-vllm-cpp` | kev 0.8B | Qwen3.5-0.8B-Base with a merged LoRA and a PointerHead readout, converted for vllm.cpp only, about 1.53 GB |
|
||||
|
||||
The engine, [vllm.cpp]({{% relref "features/vllm-cpp" %}}), also supports the
|
||||
kev, CLM and xor decision models. Those checkpoints need a conversion step, so
|
||||
they are not gallery entries yet.
|
||||
CLM and xor decision models. Those checkpoints need a conversion step, so
|
||||
they are not gallery entries yet. The kev entry installs a checkpoint that was
|
||||
already converted with the vllm.cpp `convert-kev.py` script.
|
||||
|
||||
Tev1 is an autoregressive decision model. The engine answers each question by
|
||||
scoring the option letters, so its `confidence` is the entropy measure Ollama
|
||||
|
||||
@@ -63821,6 +63821,60 @@
|
||||
type: huggingface
|
||||
repo: togethercomputer/Tev1-0.8B-experimental
|
||||
revision: 6bb2dff14b38fea90ddb14d870166ccaf77374e9
|
||||
- name: kev-0.8b-vllm-cpp
|
||||
url: github:mudler/LocalAI/gallery/virtual.yaml@master
|
||||
urls:
|
||||
- https://huggingface.co/mudler/kev-0.8b-vllm-cpp
|
||||
- https://huggingface.co/jaredpalmer/kev-0.8b
|
||||
- https://github.com/mudler/vllm.cpp
|
||||
description: |
|
||||
kev is a System 1 decision model by Jared Palmer. It answers typed choice,
|
||||
noul and score questions about a text state with one scoring pass per
|
||||
question. It does not generate text. The model is a frozen
|
||||
Qwen3.5-0.8B-Base backbone, a rank-16 LoRA adapter and a PointerHead
|
||||
readout.
|
||||
|
||||
This entry installs a converted redistribution of jaredpalmer/kev-0.8b:
|
||||
the LoRA is merged into the BF16 backbone, the head is stored as
|
||||
head.safetensors, and config.json names the KevModel architecture. The
|
||||
checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers,
|
||||
vLLM and llama.cpp cannot load it.
|
||||
|
||||
In LocalAI, serve it via POST /v1/systemone. The vllm.cpp project records
|
||||
PointerHead golden-vector tests (25 cases) and a 5-case end-to-end
|
||||
comparison against the kev reference server as equal. The upload itself
|
||||
was smoke-tested with one request on CPU; there is no accuracy benchmark
|
||||
and no GPU run. The entry sets a 2048-token context and an explicit KV
|
||||
pool, because the default 4096-token context does not fit the default
|
||||
CPU KV pool and the load fails. BF16 weights, about 1.53 GB, pinned to a
|
||||
revision.
|
||||
license: apache-2.0
|
||||
tags:
|
||||
- decisions
|
||||
- systemone
|
||||
- vllm-cpp
|
||||
- cpu
|
||||
- gpu
|
||||
size: 1.53GB
|
||||
last_checked: "2026-09-30"
|
||||
overrides:
|
||||
backend: vllm-cpp
|
||||
known_usecases:
|
||||
- decisions
|
||||
context_size: 2048
|
||||
engine_args:
|
||||
block_size: 32
|
||||
num_blocks: 256
|
||||
max_num_seqs: 4
|
||||
parameters:
|
||||
model: mudler/kev-0.8b-vllm-cpp
|
||||
artifacts:
|
||||
- name: model
|
||||
target: model
|
||||
source:
|
||||
type: huggingface
|
||||
repo: mudler/kev-0.8b-vllm-cpp
|
||||
revision: c17e73666ded1e9d284470eae7e0de9a27294e77
|
||||
- name: qwen3-vl-4b-vllm-cpp
|
||||
url: github:mudler/LocalAI/gallery/virtual.yaml@master
|
||||
urls:
|
||||
|
||||
Reference in new issue
Block a user