mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-02 19:14:38 -04:00
docs: document the systemone usecase and decisions API
Assisted-by: Claude Code:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
1 parent
c2af8d56ea
commit
c7f278dd0d
3 files changed
+117
-11
No files matched your search
@@ -1066,7 +1066,9 @@ known_usecases:
|
||||
- embeddings
|
||||
```
|
||||
|
||||
Available flags: `chat`, `completion`, `edit`, `embeddings`, `rerank`, `image`, `transcript`, `tts`, `sound_generation`, `tokenize`, `vad`, `video`, `detection`, `llm` (combination of CHAT, COMPLETION, EDIT).
|
||||
Available flags: `chat`, `completion`, `edit`, `embeddings`, `rerank`, `image`, `transcript`, `tts`, `sound_generation`, `tokenize`, `vad`, `video`, `detection`, `score`, `token_classify`, `systemone`, `llm` (combination of CHAT, COMPLETION, EDIT).
|
||||
|
||||
`systemone` marks a model as a decision model for the [SystemOne API]({{% relref "features/systemone" %}}) (`POST /v1/systemone`). It is never guessed, and a model that declares it is not listed as a chat, completion or embeddings model.
|
||||
|
||||
`token_classify` marks a model as a token-classification (NER) provider for the PII filter (e.g. an `openai-privacy-filter` GGUF). Declare it explicitly together with `embeddings: true` (the classifier loads via TOKEN_CLS pooling). It runs on the dedicated `privacy-filter` backend (`backend/cpp/privacy-filter`), a standalone GGML engine for the `openai-privacy-filter` family - separate from `llama-cpp`, which no longer carries the token-classification path.
|
||||
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
+++
|
||||
disableToc = false
|
||||
title = "SystemOne decisions"
|
||||
weight = 66
|
||||
url = "/features/systemone/"
|
||||
+++
|
||||
|
||||
SystemOne is an API for fast, typed decisions. You send a piece of text (the
|
||||
*state*) and a set of named questions. A decision model answers each question
|
||||
with a value and a confidence, in one pass. The model does not generate text, so
|
||||
there is nothing to parse and no free-form output to validate.
|
||||
|
||||
The request and response shapes follow the [kev](https://github.com/jaredpalmer/kev)
|
||||
project and match the `/v1/systemone` endpoint that Ollama added in 0.35.
|
||||
|
||||
## Endpoints
|
||||
|
||||
| Endpoint | Method | Description |
|
||||
|---|---|---|
|
||||
| `/v1/systemone` | POST | Answer all questions in one pass |
|
||||
| `/v1/systemone/permute` | POST | Re-run one choice question under `n_perm` option orders |
|
||||
| `/v1/systemone/separate` | POST | Answer each question in its own pass |
|
||||
|
||||
## Question types
|
||||
|
||||
| Type | Answer | Fields in the answer |
|
||||
|---|---|---|
|
||||
| `choice` | One option out of a named set | `choice`, `probabilities`, `confidence` |
|
||||
| `noul` | Yes, no or unknown for a statement | `noul` (0 to 1), `entities` |
|
||||
| `score` | One level on a scale | `score`, `legend`, `probabilities`, `confidence` |
|
||||
|
||||
## Example
|
||||
|
||||
```bash
|
||||
curl http://localhost:8080/v1/systemone -H "Content-Type: application/json" -d '{
|
||||
"model": "laya-vllm-cpp",
|
||||
"state": "My order arrived broken and I want my money back. This is the second time.",
|
||||
"questions": {
|
||||
"team": {
|
||||
"type": "choice",
|
||||
"instructions": "Which team should handle this ticket?",
|
||||
"criteria": {
|
||||
"billing": "Payments, invoices and refunds",
|
||||
"shipping": "Delivery and damaged goods",
|
||||
"product": "Questions about how the product works"
|
||||
}
|
||||
},
|
||||
"refund_requested": {
|
||||
"type": "noul",
|
||||
"instructions": "The customer explicitly asks for a refund"
|
||||
},
|
||||
"urgency": {
|
||||
"type": "score",
|
||||
"instructions": "How urgent is this ticket?",
|
||||
"criteria": ["not urgent", "somewhat urgent", "urgent", "critical"]
|
||||
}
|
||||
}
|
||||
}'
|
||||
```
|
||||
|
||||
Every answer carries a `confidence` value, and the response reports token usage
|
||||
and `latency_ms`.
|
||||
|
||||
## Choosing a model
|
||||
|
||||
A model can serve SystemOne only if it is a decision model. Declare the usecase
|
||||
in the model config:
|
||||
|
||||
```yaml
|
||||
name: laya
|
||||
backend: vllm-cpp
|
||||
known_usecases:
|
||||
- systemone
|
||||
parameters:
|
||||
model: convaiinnovations/laya
|
||||
```
|
||||
|
||||
`systemone` is never guessed, and a model that declares it is not listed as a
|
||||
chat, completion or embeddings model. A model that declares usecases without
|
||||
`systemone` or `token_classify` gets a `400` from these endpoints that names the
|
||||
missing usecase. A config that declares no usecases at all keeps working, so
|
||||
setups that predate the flag are not broken. Models that declare `token_classify`
|
||||
are served by the zero-shot NER path.
|
||||
|
||||
Install one from the gallery and filter on the `systemone` tag:
|
||||
|
||||
| Gallery entry | Model | Notes |
|
||||
|---|---|---|
|
||||
| `laya-vllm-cpp` | Laya | ModernBERT-large, non-autoregressive, about 800 MB |
|
||||
| `gliner25-decide-vllm-cpp` | GLiNER2.5-Decide | DeBERTa-v3-large with a classification head, about 2 GB |
|
||||
|
||||
The engine, [vllm.cpp]({{% relref "features/vllm-cpp" %}}), also supports the
|
||||
kev, CLM and xor decision models. Those checkpoints need a conversion step, so
|
||||
they are not gallery entries yet.
|
||||
|
||||
Tev1 is an autoregressive decision model. It answers through chat completions
|
||||
and does not serve `/v1/systemone` yet.
|
||||
|
||||
## Access control
|
||||
|
||||
When authentication is on, the three routes need the `systemone` feature. It is
|
||||
on by default for every user, like the other API features, and an administrator
|
||||
can turn it off per user.
|
||||
@@ -160,22 +160,23 @@ forward, which is the required contract for pooling models in vllm.cpp. A
|
||||
device-resident forward is tracked as a performance optimization, not a
|
||||
correctness gap.
|
||||
|
||||
### SystemOne structured-extraction API
|
||||
### SystemOne decision API
|
||||
|
||||
The `vllm-cpp` backend also exposes kev-compatible SystemOne endpoints that
|
||||
turn zero-shot NER into structured question answering. These mirror the API
|
||||
from the [kev](https://github.com/jaredpalmer/kev) project:
|
||||
The `vllm-cpp` backend serves the kev-compatible SystemOne endpoints: typed
|
||||
`choice`, `noul` and `score` questions over a state text, answered by a
|
||||
non-generative decision model in one pass. A decision model declares
|
||||
`known_usecases: [systemone]`. See [SystemOne decisions]({{% relref "features/systemone" %}})
|
||||
for the request shape, the models you can install and the access rules.
|
||||
|
||||
| Endpoint | Method | Description |
|
||||
|---|---|---|
|
||||
| `/v1/systemone` | POST | Answer all questions in one NER pass |
|
||||
| `/v1/systemone` | POST | Answer all questions in one pass |
|
||||
| `/v1/systemone/permute` | POST | Re-run one choice question under n_perm option orders |
|
||||
| `/v1/systemone/separate` | POST | Answer each question in its own NER pass (N passes) |
|
||||
| `/v1/systemone/separate` | POST | Answer each question in its own pass (N passes) |
|
||||
|
||||
Each question has a `type` of `noul` (binary entity presence), `choice` (pick
|
||||
one option), or `score` (pick one level). The `model` field in the request body
|
||||
selects the NER model. Labels are derived from the question definition, so no
|
||||
`ner_labels` configuration is needed for these endpoints.
|
||||
The GLiNER2.5 zero-shot NER model (`token_classify`) also serves these
|
||||
endpoints. It derives its NER labels from the question definitions, so no
|
||||
`ner_labels` configuration is needed.
|
||||
|
||||
## Beyond text generation
|
||||
|
||||
|
||||
Reference in new issue
Block a user