diff --git a/docs/content/advanced/model-configuration.md b/docs/content/advanced/model-configuration.md index 5cfa74ccd..10e71ef96 100644 --- a/docs/content/advanced/model-configuration.md +++ b/docs/content/advanced/model-configuration.md @@ -1066,7 +1066,9 @@ known_usecases: - embeddings ``` -Available flags: `chat`, `completion`, `edit`, `embeddings`, `rerank`, `image`, `transcript`, `tts`, `sound_generation`, `tokenize`, `vad`, `video`, `detection`, `llm` (combination of CHAT, COMPLETION, EDIT). +Available flags: `chat`, `completion`, `edit`, `embeddings`, `rerank`, `image`, `transcript`, `tts`, `sound_generation`, `tokenize`, `vad`, `video`, `detection`, `score`, `token_classify`, `systemone`, `llm` (combination of CHAT, COMPLETION, EDIT). + +`systemone` marks a model as a decision model for the [SystemOne API]({{% relref "features/systemone" %}}) (`POST /v1/systemone`). It is never guessed, and a model that declares it is not listed as a chat, completion or embeddings model. `token_classify` marks a model as a token-classification (NER) provider for the PII filter (e.g. an `openai-privacy-filter` GGUF). Declare it explicitly together with `embeddings: true` (the classifier loads via TOKEN_CLS pooling). It runs on the dedicated `privacy-filter` backend (`backend/cpp/privacy-filter`), a standalone GGML engine for the `openai-privacy-filter` family - separate from `llama-cpp`, which no longer carries the token-classification path. diff --git a/docs/content/features/systemone.md b/docs/content/features/systemone.md new file mode 100644 index 000000000..d58a3d436 --- /dev/null +++ b/docs/content/features/systemone.md @@ -0,0 +1,103 @@ ++++ +disableToc = false +title = "SystemOne decisions" +weight = 66 +url = "/features/systemone/" ++++ + +SystemOne is an API for fast, typed decisions. You send a piece of text (the +*state*) and a set of named questions. A decision model answers each question +with a value and a confidence, in one pass. The model does not generate text, so +there is nothing to parse and no free-form output to validate. + +The request and response shapes follow the [kev](https://github.com/jaredpalmer/kev) +project and match the `/v1/systemone` endpoint that Ollama added in 0.35. + +## Endpoints + +| Endpoint | Method | Description | +|---|---|---| +| `/v1/systemone` | POST | Answer all questions in one pass | +| `/v1/systemone/permute` | POST | Re-run one choice question under `n_perm` option orders | +| `/v1/systemone/separate` | POST | Answer each question in its own pass | + +## Question types + +| Type | Answer | Fields in the answer | +|---|---|---| +| `choice` | One option out of a named set | `choice`, `probabilities`, `confidence` | +| `noul` | Yes, no or unknown for a statement | `noul` (0 to 1), `entities` | +| `score` | One level on a scale | `score`, `legend`, `probabilities`, `confidence` | + +## Example + +```bash +curl http://localhost:8080/v1/systemone -H "Content-Type: application/json" -d '{ + "model": "laya-vllm-cpp", + "state": "My order arrived broken and I want my money back. This is the second time.", + "questions": { + "team": { + "type": "choice", + "instructions": "Which team should handle this ticket?", + "criteria": { + "billing": "Payments, invoices and refunds", + "shipping": "Delivery and damaged goods", + "product": "Questions about how the product works" + } + }, + "refund_requested": { + "type": "noul", + "instructions": "The customer explicitly asks for a refund" + }, + "urgency": { + "type": "score", + "instructions": "How urgent is this ticket?", + "criteria": ["not urgent", "somewhat urgent", "urgent", "critical"] + } + } +}' +``` + +Every answer carries a `confidence` value, and the response reports token usage +and `latency_ms`. + +## Choosing a model + +A model can serve SystemOne only if it is a decision model. Declare the usecase +in the model config: + +```yaml +name: laya +backend: vllm-cpp +known_usecases: + - systemone +parameters: + model: convaiinnovations/laya +``` + +`systemone` is never guessed, and a model that declares it is not listed as a +chat, completion or embeddings model. A model that declares usecases without +`systemone` or `token_classify` gets a `400` from these endpoints that names the +missing usecase. A config that declares no usecases at all keeps working, so +setups that predate the flag are not broken. Models that declare `token_classify` +are served by the zero-shot NER path. + +Install one from the gallery and filter on the `systemone` tag: + +| Gallery entry | Model | Notes | +|---|---|---| +| `laya-vllm-cpp` | Laya | ModernBERT-large, non-autoregressive, about 800 MB | +| `gliner25-decide-vllm-cpp` | GLiNER2.5-Decide | DeBERTa-v3-large with a classification head, about 2 GB | + +The engine, [vllm.cpp]({{% relref "features/vllm-cpp" %}}), also supports the +kev, CLM and xor decision models. Those checkpoints need a conversion step, so +they are not gallery entries yet. + +Tev1 is an autoregressive decision model. It answers through chat completions +and does not serve `/v1/systemone` yet. + +## Access control + +When authentication is on, the three routes need the `systemone` feature. It is +on by default for every user, like the other API features, and an administrator +can turn it off per user. diff --git a/docs/content/features/vllm-cpp.md b/docs/content/features/vllm-cpp.md index 74e3c0d38..6dc67c0b9 100644 --- a/docs/content/features/vllm-cpp.md +++ b/docs/content/features/vllm-cpp.md @@ -160,22 +160,23 @@ forward, which is the required contract for pooling models in vllm.cpp. A device-resident forward is tracked as a performance optimization, not a correctness gap. -### SystemOne structured-extraction API +### SystemOne decision API -The `vllm-cpp` backend also exposes kev-compatible SystemOne endpoints that -turn zero-shot NER into structured question answering. These mirror the API -from the [kev](https://github.com/jaredpalmer/kev) project: +The `vllm-cpp` backend serves the kev-compatible SystemOne endpoints: typed +`choice`, `noul` and `score` questions over a state text, answered by a +non-generative decision model in one pass. A decision model declares +`known_usecases: [systemone]`. See [SystemOne decisions]({{% relref "features/systemone" %}}) +for the request shape, the models you can install and the access rules. | Endpoint | Method | Description | |---|---|---| -| `/v1/systemone` | POST | Answer all questions in one NER pass | +| `/v1/systemone` | POST | Answer all questions in one pass | | `/v1/systemone/permute` | POST | Re-run one choice question under n_perm option orders | -| `/v1/systemone/separate` | POST | Answer each question in its own NER pass (N passes) | +| `/v1/systemone/separate` | POST | Answer each question in its own pass (N passes) | -Each question has a `type` of `noul` (binary entity presence), `choice` (pick -one option), or `score` (pick one level). The `model` field in the request body -selects the NER model. Labels are derived from the question definition, so no -`ner_labels` configuration is needed for these endpoints. +The GLiNER2.5 zero-shot NER model (`token_classify`) also serves these +endpoints. It derives its NER labels from the question definitions, so no +`ner_labels` configuration is needed. ## Beyond text generation