Explain decision-model uses and the current gallery choices.
Link the setup guide and include a bounded API example with access
requirements and model-confidence caveats.
Assisted-by: OpenAI:undisclosed [shell]
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* feat(gallery): add nimble-9b-vllm-cpp decision model
Add Bespoke Nimble 9B, converted for vllm.cpp and pinned to the weights
commit 52eead25 of mudler/Bespoke-Nimble-9B-vllm-cpp (HEAD only adds the
model card). It is a redistribution of bespokelabs/Bespoke-Nimble-9B
with the LoRA merged into Qwen3.5-9B; config.json names NimbleModel, so
no hf_overrides are needed.
The artifact sits under overrides, where the installer reads it. The
entry sets an 8192-token context, Nimble's own prompt limit, and a KV
pool of 1024 blocks of 32 tokens for 4 sequences (about 1 GiB at 32 KiB
per token for the 8 full-attention layers).
Installed with local-ai models install and served on CPU through the
vllm-cpp backend: the model card's billing request gives billing
(0.986), refund 0.998 and urgency 0.33. Peak resident memory was
18.4 GB, so the description asks for about 20 GB of free RAM.
List the entry in the decisions gallery table. CLM stays out of the
gallery: the pinned engine cannot load the published head layout.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
* feat(gallery): add clm-v0.1-8b-vllm-cpp and bump vllm.cpp for CLM
The CLM checkpoint on mudler/CLM-v0.1-8B-vllm-cpp stores the heads with
the reference's own tensor names (state_head.inp, hidden.N, norms.N,
out). The pinned vllm.cpp 96788348 still expects the old .0/.2/.4/.6
layout and refuses the load with "head.safetensors incomplete for
state_head". vllm.cpp a19294a9 matches the reference layout and adds the
converter that produced the upload, so move the pin there. The ABI stays
at v30.
Add the CLM entry, pinned to the weights commit 0d1903b1 (HEAD only adds
the model card), with a 4096-token context and a KV pool for 4 sequences
(about 2.25 GiB at 144 KiB per token for Qwen3-8B).
Installed with local-ai models install and served on CPU against a
libvllm built at a19294a9: the model card example (john works at google,
entity type) gives person 0.950, the same as the card. Peak resident
memory was 17.9 GB.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
When the first result of a streamed request is an error (for example a
prompt that exceeds the context), PredictStream wrote the error message
as a Reply and only then returned the error status. LocalAI treated that
Reply as the first token: it sent the assistant role chunk and the error
text as `content` on an HTTP 200 stream. Because a chunk had already been
written, the pre-stream HTTP error path from #12204 never triggered, so
streaming clients still got a 200 with the error as model output, while
the same request without streaming correctly returns a 400.
Return the error only as the gRPC status. The e2e backend suite gets a
`context_overflow` capability (enabled for llama-cpp) that streams an
over-long prompt and asserts an error status with no content.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
When an Anthropic model declines to answer, the Messages API returns
HTTP 200 with `stop_reason: "refusal"` and empty `content`. The
translate mode mapped that to a normal reply with no content, so the
OpenAI-compatible response looked like a successful completion
(`finish_reason: "stop"`, empty message). Routers, agents and UIs could
not tell "the model declined" from "the model had nothing to say", and
no fallback was triggered.
Return an explicit error for `stop_reason: "refusal"` in both the
non-streaming path and the streaming path (`message_delta`). Regular
replies, including empty `end_turn` replies, are unchanged.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
AudioTranscription caught every exception and returned an empty
TranscriptResult. A failed diarization step therefore discarded a
transcript that was already finished: with an HF token that has not
accepted the terms of the gated pyannote pipeline, the download fails
with 403 and every transcription came back as an empty text with
HTTP 200.
Diarization now degrades: if it fails, the transcript is returned
without speaker labels and the reason is logged. Any other failure
aborts the call with INTERNAL instead of pretending success with an
empty text.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
The environment fallback was applied whenever n_parallel was still 1
after option parsing. An explicit `parallel: 1` in the model YAML is
indistinguishable from the default that way, so it was replaced by
LLAMACPP_PARALLEL. The docs say options in the YAML take precedence
over environment variables; a single model could not be forced to one
slot while the global variable was set.
Track whether the options set the slot count and resolve it in a small
helper (parallel_params.h): option first, then LLAMACPP_PARALLEL, then
1. The helper gets a standalone unit test picked up by
`make test-backend-cpp`.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
* feat(schema): validate portable speaker profiles
Add the versioned profile schema for explicit speaker enrollment.
Validate compatibility against separately supplied loaded-encoder metadata.
Reject unusable speakers, invalid vectors, and inconsistent clean spans.
This slice does not change HTTP routes, backend integration, or the UI.
Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet): export profiles with transcripts
Export opt-in speaker profiles and trusted encoder metadata.
Replay registrations by ID so duplicate display names keep independent
vectors.
Use one profile-capable diarization for slots, names, and clean spans.
Assign timestamped ASR words to those slots without a second diarization.
Preserve legacy opt-out and no-ASR behavior, and propagate failures.
Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(audio): enroll portable speaker profiles
Gate profile exports with voice-recognition permission and validate
registration against metadata from the loaded encoder. Preserve audio
enrollment and independent registrations with duplicate display names.
Exclude diarization and registration exchanges before API trace capture
so persisted traces cannot retain profile vectors or JSON audio.
Defer candidate dimensions to trusted loaded metadata. Sort candidates
by registration ID so incompatible profiles cannot suppress legacy voices
through registry iteration order. Keep portable identity checks closed
when trusted metadata is unavailable.
Test persisted traces, explicit slot zero, and selection through offline
and live transport. Document privacy and the ephemeral registry lifecycle.
Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(ui): remember speakers from diarization
Add a Studio page for diarization and opt-in speaker profiles. Preview
clean intervals from the original recording before explicit registration.
Join profiles by raw speaker labels, preserve duplicate names, and relabel
turns only after a successful save. Discard stale results when the model
or recording changes. Share registration metadata with voice management
without storing vectors or recordings from this flow.
Document permissions and the global, ephemeral registry. Cover enrollment,
permissions, previews, and asynchronous races with mocked Playwright tests.
Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs: clarify HTTP speaker enrollment support
Replace the stale enrollment limitation with the current HTTP workflow.
Distinguish native transport from explicit registration and link its docs.
Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(parakeet): pin merged speaker profile support
Use the merged commit from mudler/parakeet.cpp#80.
Its tree matches the previously accepted native pin.
Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs: add diarization enrollment setup example
Connect the existing gallery modes to the speaker enrollment workflow.
Show installation, private profile export, explicit raw-slot registration,
and later recognition without another export.
Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs(blog): explain diarization speaker profiles
Put the diarization walkthrough on the LocalAI website in the feature PR.
Cover the three gallery modes, explicit enrollment, and privacy limits.
Link setup instructions and keep availability conditional on feature support.
Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs(blog): focus diarization on everyday use
Explain what users can do with recordings before the setup steps.
Replace the technical walkthrough with a short Studio guide and link
readers to the existing reference for model names and developer use.
Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs(blog): lead with speaker capabilities
Present speaker recognition through everyday uses and a short UI flow.
Keep technical reference details in the existing documentation.
Assisted-by: OpenAI:unknown
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(diarization): satisfy Go lint checks
Avoid copying protobuf message state when extending backend status, check the multipart reader close result, and document the focused testing.T lint exemptions.
Assisted-by: nib:gpt-5.6-sol
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
FunASR left Transformers and Hugging Face Hub unconstrained, which let uv backtrack to tokenizers 0.10.3 without Python 3.12 wheels. Keep Transformers on the supported 4.x range, including the Intel upgrade profile.\n\nAssisted-by: nib:gpt-5.6-sol\nSigned-off-by: Ettore Di Giacinto <mudler@localai.io>
Remove watchdog tracking when backends stop, crash, or fail to start.
Validate eviction addresses under the model lifecycle lock so delayed
shutdowns cannot stop a replacement backend. Preserve replacement size
estimates when removing an old address.
Add lifecycle regression tests and document shutdown behavior.
Fixes#12331
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
* ⬆️ Update ggml-org/llama.cpp
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(llama-cpp): migrate score batches to the new API
The pinned engine removes common_batch_add and the raw batch view.
Use common_batch entries and llama_process for score suffix decoding.
Read shared-prefix scores from the current common_batch view.
Validation: reproduce both compiler errors on the original patch.
The patched server context and complete grpc-server translation unit
pass g++ -std=c++17 -fsyntax-only with generated protobuf headers.
Assisted-by: Codex:gpt-6
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The chat bot examples moved to mudler/LocalAI-examples; the old paths in
LocalAGI and LocalAI no longer exist.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Pratik Gandhi <travpreneur@gmail.com>
* feat(voice): list registered voices and record which encoder made them
The voice registry could register, identify and forget but not list, and
it did not remember which speaker encoder produced an embedding. Add
Metadata.Model and Registry.List, answered from the index the store
registry already keeps for Forget. Needed so a backend can be given the
registered voices that match its own speaker encoder.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* feat(voice): store the encoder model with a registered voice
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* feat(voice): pick the registered voices that match a speaker model
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* feat(proto): carry known voices and speaker names on diarize and live messages
Assisted-by: Claude:claude-haiku-4-5 [Claude Code]
* feat(diarization): name speakers from the voice registry
When a diarization model has a speaker_model option, the endpoint sends
the registered voices made by that encoder to the backend. The backend's
name and name_score come back as extra fields next to the normalized
SPEAKER_NN speaker, and the speakers summary carries the first name seen
for each speaker. RTTM output and results without names are unchanged.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* feat(live): pass registered voices to a live session and surface speaker names
Live sessions now send the registered voices that match the model's
speaker_model to the backend, and each speaker segment carries the name
the backend matched. The realtime segment event gains an optional
speaker_name field.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* feat(parakeet-cpp): load a speaker model and build per-request voice registries
Adds the speaker bindings (ABI v9 and v10, probed separately), the
speaker_model, speaker_threshold and speaker_margin options, and a
per-request registry builder over the known voices.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* feat(parakeet-cpp): name the speakers in Diarize from the known voices
Diarize builds a per-request speaker registry from the known voices when a
speaker model is loaded, calls the named C functions, and puts each slot's
registered name and score on the segments. The registry is freed on every
path. A library without ABI 10 reports Unimplemented instead of dropping
the names.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* feat(parakeet-cpp): name speakers in the live scene stream
The live scene stream now begins with a known-voice registry when a
speaker model is loaded and the live config carries voices, and each
closed speaker segment takes its slot's current name from the feed's
names map. A segment that closes before its slot is identified has an
empty name. The registry is freed after the stream, on every path.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* feat(gallery): speaker naming entries and docs for parakeet-cpp
Add three gallery entries that load the WeSpeaker ResNet34 speaker model
next to the diarization or realtime scene models, and document speaker
names in the voice recognition, diarization, audio to text and realtime
pages.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* fix(parakeet-cpp): skip an unusable registered voice instead of failing the request
A registered voice with the wrong embedding size, or one the C side
refused, failed the whole diarization request, so one legacy voice broke
the model for every user. Skip such voices with a warning that does not
carry the voice name, and take the plain path when none is left.
Also map an exact 0 speaker threshold or margin to a tiny positive value,
since the C side reads 0 as "use the default", and fix a stale comment
about which contexts Free() walks.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* fix(diarization): warn once per model about voices from another encoder; document the privacy limit
The different-encoder warning fired on every request. Log it once per
feature and speaker model, then at debug level. Document that the global
voice registry lets any caller of a speaker_model model learn matching
names, and that skipped wrong-sized voices are logged.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
* chore(parakeet-cpp): bump parakeet.cpp to 8c8cec0 (C-API v10) and check speaker naming against the real library
The pin moves from 623a968 to 8c8cec0, which brings in everything merged
in parakeet.cpp since: the voice identification change (C-API v9, #78) and
raw-embedding enroll plus diarize-only speaker naming (C-API v10, #79).
New real-library specs (gated on PARAKEET_BACKEND_TEST_SPEAKER_MODEL,
_DIAR_MODEL, _WAV and, for the live path, _STREAM_MODEL) name the two
speakers of two_speakers.wav from a committed pair of WeSpeaker embeddings,
with the voices passed in reversed order. They also check that the float32
threshold reaches C through purego. The shared test loader now registers
the v9/v10 and scene symbols as main.go does.
The rebase onto origin/master had no conflicts.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Add pinned Q4_K_M and Q6_K builds for text chat with llama.cpp.
Document installation and explicit variant selection.
Assisted-by: Codex:gpt-6
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
Add the kev 0.8B decision model, converted for vllm.cpp and pinned to
revision c17e7366 of mudler/kev-0.8b-vllm-cpp. It is a redistribution of
jaredpalmer/kev-0.8b with the LoRA merged and the PointerHead stored as
head.safetensors, so only the vllm-cpp backend can load it.
The artifact sits under overrides, where the installer reads it. The
entry sets a 2048-token context and an explicit KV pool: with the
default 4096-token context the CPU KV pool holds only 4064 tokens and the
load fails.
List the entry in the decisions gallery table and drop kev from the
list of decision models that are not gallery entries yet.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
G304 flags reading a path built from a variable. The directory is the model
directory from the operator's own model config, not a request input, so it is
annotated the way the other backends do it.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(vllm-cpp): bump vllm.cpp to 967883486 (ABI v30)
Moves the pin from c3bebc357 to 967883486. On top of the Nimble decision
adapter and the Qwen3.5 vision-loader fix, this brings Tev1 on
/v1/systemone and vllm_decide (opt-in through a "Tev1Model" architecture
in config.json), a tokenizer/ subdirectory fallback so the Laya HF
snapshot loads as downloaded, a stop-token fix, a logprobs fix under async
scheduling and a pinned parakeet.cpp fetch for the diarization build.
ABI v30 only adds the diarization and speaker-attributed ASR entry
points; no existing struct or signature changed, so the purego mirrors
keep their layout and only abiVersion moves to 30. Between 4479dc99f and
967883486 vllm.h changed only in a comment.
v30 turns VLLM_CPP_WITH_DIARIZATION on by default. The fetch is pinned
now, but ON still downloads parakeet.cpp at configure time and links a
second ggml into libvllm for calls this backend never makes, so build
with the option off: the symbols stay present as refusing stubs.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
* feat(vllm-cpp): add the hf_overrides engine arg
vLLM parity: engine_args.hf_overrides is a JSON object of top-level
config.json keys merged over the model directory's own config.json. The
main use is opting a published checkpoint into an engine adapter its
config does not name, such as {"architectures": ["Tev1Model"]} on the
Tev1 snapshots, which declare Qwen3_5ForConditionalGeneration.
The C ABI has no override input and the engine reads config.json from
the directory it is given, so Load builds a private overlay directory:
the merged config.json plus a symlink to every other entry of the model
directory, and passes that to the engine. The download is never written.
Free, a failed load and the next Load remove the overlay.
validModelPath and the DFlash draft resolution still see the real
directory.
A value that is not a JSON object, a .gguf model or a directory without
config.json fails the load instead of being skipped like an unknown
engine_args key, because loading the unmodified config would serve a
different architecture than the one configured.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
* fix(gallery): nest vllm-cpp artifacts under overrides
artifacts: is a model-config key, and the installer reads model-config
keys only from overrides:. Five vllm-cpp entries (laya, gliner25-decide,
qwen3-vl-4b, cua-s1-forms and gliner2.5) declared it at the entry top
level, where it is silently dropped: the install reports success, writes
a config whose model is the bare HF repo id and downloads nothing, and
vllm-cpp (which does not infer artifacts) then fails the first load with
"model path not found".
Move each block under overrides:, and add a guard test that refuses a
top-level artifacts: key in the index.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
* feat(gallery): add Tev1 4B and 0.8B on vllm-cpp
Two decisions entries for Together AI's Tev1 checkpoints, pinned to the
current HF revisions. Tev1 is autoregressive: vllm.cpp answers
/v1/systemone by scoring the option letters, and the same engine still
serves chat completions. The published config.json names
Qwen3_5ForConditionalGeneration, so each entry sets
hf_overrides: {architectures: [Tev1Model]} to enable the decision route
without editing the download. known_usecases is [decisions] only, since
a declared decisions list is authoritative for reservation.
The descriptions state what was checked: agreement with transformers on
CPU over seven questions (4B 7/7, max probability difference 0.0004;
0.8B 6/7 with one near tie), CPU-only for the decision route, and a
fine-tune license the model card says is still being finalized, so no
license key is set.
The Decisions API page lists both entries, drops the note that Tev1
does not serve /v1/systemone and documents the 24-option limit (Ollama
allows 26).
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-sonnet-5-5
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Adds a Ginkgo contract suite for the standalone (LocalAGI-backed) agent service in core/services/agentpool, with a fake OpenAI-compatible LLM and a harness that boots a real AgentPoolService. It pins CRUD, per-user isolation, export and import, pause and resume, the chat and SSE event contract, status and observables, persistence and the raw-key agent lookup. No production code changes.
Specs that pin a known gap are named known gap or known defect, so later phases flip them on purpose.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
All three routes now validate the request before it reaches a model: body
size (413 over 64 KiB), state, question count, blank ids, option and level
counts, and noul criteria keys. Forwarded decision requests skipped this
before, so a malformed question surfaced as a backend error.
The docs claimed the wire shape matches Ollama's. Field names and question
types do; confidence, error shape, keep_alive and state rendering differ, and
the docs now say so.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The usecase describes what a model can do, and the category is the Decisions
API. SystemOne stays as the wire contract: the /v1/systemone routes, the
Score RPC question_type and the swagger tag are unchanged. The usecase,
flag, auth feature, UI label, gallery tags and docs page are now decisions.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The React UI read every non-media attachment with file.text(). For a PDF
that decodes the binary bytes as UTF-8, so the model received raw
"%PDF ... stream ... endobj" noise instead of the document. The legacy
Alpine UI ran pdf.js; that step was not ported when the React UI
replaced it, but both file pickers still advertise .pdf.
Add a shared readAttachmentText helper that routes PDFs through
pdfjs-dist and reads other files as before. pdf.js and its worker load
on first use, so the main bundle does not grow. A PDF that cannot be
parsed or has no text layer (scanned, encrypted, damaged) is rejected
with a toast instead of being attached as an empty or garbage file.
Cover the chat and home paths with Playwright specs that build a real
PDF in the test.
Assisted-by: Claude Code:claude-sonnet-5-5 [playwright] [eslint]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
nib v0.12.0 called xlog.SetLogger with a *slog.Logger, which does not
build against the xlog v0.0.6 LocalAI uses. v0.12.1 fixes that
(mudler/nib#137).
nib's ApprovalMode is now a named string type, so the chat tests
compare against nibtypes.ApprovalAuto and ApprovalPrompt instead of
untyped strings, and run.go sets the exported constant.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
vllm_decide refuses NER architectures and the NER entry point refuses decision
architectures, so each model kind 500ed on half of the routes. A token_classify
model now goes to the NER path on /v1/systemone, and /permute and /separate
return 400 for decision models. Docs and instructions state which kind serves
which route.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
laya declares the systemone usecase instead of chat. The gated Qwen3.6 27B
NVFP4 entries gain vision; the 35B-A3B entries gain it as experimental
because image input is not token-gated. Adds GLiNER2.5-Decide and
Qwen3-VL-4B. A guard test keeps capability tags and known_usecases in
agreement for every vllm-cpp entry.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Adds a default-on systemone route feature for the three /v1/systemone
routes and an /api/instructions area for them. No MCP tool is added: the
endpoints run inference and are not admin install/edit actions, and the
route-map test still passes.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
A chat-only model now gets a 400 naming known_usecases: [systemone] instead
of a backend error. Configs declaring no usecases and token_classify models
stay allowed so existing laya and GLiNER setups keep working.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Explicit-only, reserving usecase like score and token_classify: a declared
list is authoritative and the heuristic never guesses it. vllm-cpp now
lists systemone and vision as possible usecases.
Assisted-by: Claude Code:claude-sonnet-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>