Add the kev 0.8B decision model, converted for vllm.cpp and pinned to
revision c17e7366 of mudler/kev-0.8b-vllm-cpp. It is a redistribution of
jaredpalmer/kev-0.8b with the LoRA merged and the PointerHead stored as
head.safetensors, so only the vllm-cpp backend can load it.
The artifact sits under overrides, where the installer reads it. The
entry sets a 2048-token context and an explicit KV pool: with the
default 4096-token context the CPU KV pool holds only 4064 tokens and the
load fails.
List the entry in the decisions gallery table and drop kev from the
list of decision models that are not gallery entries yet.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
* chore(vllm-cpp): bump vllm.cpp to 967883486 (ABI v30)
Moves the pin from c3bebc357 to 967883486. On top of the Nimble decision
adapter and the Qwen3.5 vision-loader fix, this brings Tev1 on
/v1/systemone and vllm_decide (opt-in through a "Tev1Model" architecture
in config.json), a tokenizer/ subdirectory fallback so the Laya HF
snapshot loads as downloaded, a stop-token fix, a logprobs fix under async
scheduling and a pinned parakeet.cpp fetch for the diarization build.
ABI v30 only adds the diarization and speaker-attributed ASR entry
points; no existing struct or signature changed, so the purego mirrors
keep their layout and only abiVersion moves to 30. Between 4479dc99f and
967883486 vllm.h changed only in a comment.
v30 turns VLLM_CPP_WITH_DIARIZATION on by default. The fetch is pinned
now, but ON still downloads parakeet.cpp at configure time and links a
second ggml into libvllm for calls this backend never makes, so build
with the option off: the symbols stay present as refusing stubs.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
* feat(vllm-cpp): add the hf_overrides engine arg
vLLM parity: engine_args.hf_overrides is a JSON object of top-level
config.json keys merged over the model directory's own config.json. The
main use is opting a published checkpoint into an engine adapter its
config does not name, such as {"architectures": ["Tev1Model"]} on the
Tev1 snapshots, which declare Qwen3_5ForConditionalGeneration.
The C ABI has no override input and the engine reads config.json from
the directory it is given, so Load builds a private overlay directory:
the merged config.json plus a symlink to every other entry of the model
directory, and passes that to the engine. The download is never written.
Free, a failed load and the next Load remove the overlay.
validModelPath and the DFlash draft resolution still see the real
directory.
A value that is not a JSON object, a .gguf model or a directory without
config.json fails the load instead of being skipped like an unknown
engine_args key, because loading the unmodified config would serve a
different architecture than the one configured.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
* fix(gallery): nest vllm-cpp artifacts under overrides
artifacts: is a model-config key, and the installer reads model-config
keys only from overrides:. Five vllm-cpp entries (laya, gliner25-decide,
qwen3-vl-4b, cua-s1-forms and gliner2.5) declared it at the entry top
level, where it is silently dropped: the install reports success, writes
a config whose model is the bare HF repo id and downloads nothing, and
vllm-cpp (which does not infer artifacts) then fails the first load with
"model path not found".
Move each block under overrides:, and add a guard test that refuses a
top-level artifacts: key in the index.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
* feat(gallery): add Tev1 4B and 0.8B on vllm-cpp
Two decisions entries for Together AI's Tev1 checkpoints, pinned to the
current HF revisions. Tev1 is autoregressive: vllm.cpp answers
/v1/systemone by scoring the option letters, and the same engine still
serves chat completions. The published config.json names
Qwen3_5ForConditionalGeneration, so each entry sets
hf_overrides: {architectures: [Tev1Model]} to enable the decision route
without editing the download. known_usecases is [decisions] only, since
a declared decisions list is authoritative for reservation.
The descriptions state what was checked: agreement with transformers on
CPU over seven questions (4B 7/7, max probability difference 0.0004;
0.8B 6/7 with one near tie), CPU-only for the decision route, and a
fine-tune license the model card says is still being finalized, so no
license key is set.
The Decisions API page lists both entries, drops the note that Tev1
does not serve /v1/systemone and documents the 24-option limit (Ollama
allows 26).
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
All three routes now validate the request before it reaches a model: body
size (413 over 64 KiB), state, question count, blank ids, option and level
counts, and noul criteria keys. Forwarded decision requests skipped this
before, so a malformed question surfaced as a backend error.
The docs claimed the wire shape matches Ollama's. Field names and question
types do; confidence, error shape, keep_alive and state rendering differ, and
the docs now say so.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The usecase describes what a model can do, and the category is the Decisions
API. SystemOne stays as the wire contract: the /v1/systemone routes, the
Score RPC question_type and the swagger tag are unchanged. The usecase,
flag, auth feature, UI label, gallery tags and docs page are now decisions.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>