fix(systemone): route NER models to the NER path and refuse decision models on permute and separate

vllm_decide refuses NER architectures and the NER entry point refuses decision
architectures, so each model kind 500ed on half of the routes. A token_classify
model now goes to the NER path on /v1/systemone, and /permute and /separate
return 400 for decision models. Docs and instructions state which kind serves
which route.

Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
Ettore Di Giacinto committed 2026-09-30 11:55:26 +00:00
1 parent c7f278dd0d
commit b3d65fd538
5 files changed
+116 -10

No files matched your search

+14 -5
View File
@@ -21,6 +21,13 @@ project and match the `/v1/systemone` endpoint that Ollama added in 0.35.
| `/v1/systemone/permute` | POST | Re-run one choice question under `n_perm` option orders |
| `/v1/systemone/separate` | POST | Answer each question in its own pass |
Which route a model can serve depends on its kind:
| Model kind | `/v1/systemone` | `/permute` and `/separate` |
|---|---|---|
| Decision model (`systemone`), such as Laya or GLiNER2.5-Decide | Yes | No, returns `400` |
| Zero-shot NER model (`token_classify`), such as GLiNER2.5 | Yes, through the NER path | Yes |
## Question types
| Type | Answer | Fields in the answer |
@@ -58,8 +65,8 @@ curl http://localhost:8080/v1/systemone -H "Content-Type: application/json" -d '
}'
```
Every answer carries a `confidence` value, and the response reports token usage
and `latency_ms`.
Answers from a decision model carry a `confidence` value, and the response
reports token usage and `latency_ms`. The NER path does not report token usage.
## Choosing a model
@@ -78,9 +85,11 @@ parameters:
`systemone` is never guessed, and a model that declares it is not listed as a
chat, completion or embeddings model. A model that declares usecases without
`systemone` or `token_classify` gets a `400` from these endpoints that names the
missing usecase. A config that declares no usecases at all keeps working, so
setups that predate the flag are not broken. Models that declare `token_classify`
are served by the zero-shot NER path.
missing usecase. A model that declares `token_classify` and not `systemone` is
served by the zero-shot NER path. A vllm-cpp config that declares no usecases is
treated as a decision model, so setups that predate the flag keep working, but a
config that declares only `chat` (as an older `laya` gallery entry did) now gets
the `400` and needs `known_usecases: [systemone]`.
Install one from the gallery and filter on the `systemone` tag:
+5 -3
View File
@@ -174,9 +174,11 @@ for the request shape, the models you can install and the access rules.
| `/v1/systemone/permute` | POST | Re-run one choice question under n_perm option orders |
| `/v1/systemone/separate` | POST | Answer each question in its own pass (N passes) |
The GLiNER2.5 zero-shot NER model (`token_classify`) also serves these
endpoints. It derives its NER labels from the question definitions, so no
`ner_labels` configuration is needed.
The GLiNER2.5 zero-shot NER model (`token_classify`) also serves
`/v1/systemone`, through the NER path, and it is the model to use for
`/v1/systemone/permute` and `/v1/systemone/separate`, which decision models
refuse with a `400`. It derives its NER labels from the question definitions, so
no `ner_labels` configuration is needed.
## Beyond text generation