fix(models): fallback to application config default context size in /v1/models/capabilities (#12202) (#12216)

* fix(models): fallback to application config default context size (#12202)

Honor appConfig.ContextSize in /v1/models/capabilities when model context_size is unset.

* docs(models): explain context size fallback

Describe the application default used by capability discovery and
preserve the distinction between total context and per-request limits.

Assisted-by: Codex:GPT-6

* fix(models): apply the default context size only when context_size is unset

The request path applies the application default context size only
when a model leaves context_size unset. An explicit 0 or -1 falls
through to the backend fallback. The capabilities endpoint now does
the same, so it reports the value the backend uses.

Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
authored and GitHub committed 2026-09-28 00:52:56 +02:00
1 parent eec532704c
commit f82efdb43b
3 files changed
+38

No files matched your search

+5
View File
@@ -141,6 +141,11 @@ curl http://localhost:8080/api/instructions/config-management?format=json
An additive, LocalAI-specific superset of `/v1/models`. It returns the same set of models but enriches each entry with the **capabilities** the model supports and the **input/output modalities** it accepts and produces. Use it to decide, before sending a request, whether a given model can take an image, audio, or video attachment directly - or whether the input needs converting/transcribing first.
The reported `context_size` uses a positive model-level value first.
If the model does not set `context_size`, it uses **Settings → Performance → Default Context Size** when positive.
Otherwise, it uses the backend fallback of 4096 tokens.
For llama.cpp with separate KV caches, the reported value accounts for the number of parallel slots.
Because it is purely additive, clients that only understand `/v1/models` keep working unchanged; they simply never call this route.
```bash