* fix(schema): preserve SystemOne image inputs
Assisted-by: OpenAI
* test(schema): follow Ginkgo conventions for decision inputs
Assisted-by: OpenAI
* feat(llama-cpp): dispatch native decisions through Score
Upgrade the stock dependency and reconcile Score/TTS patches. Reuse native decision parsing, tasks, formatting and response-reader cleanup; preserve ordinary scoring admission and guard older dependencies.
Assisted-by: OpenAI
* refactor(systemone): share request and model validation
Assisted-by: OpenAI:gpt-5
* fix(systemone): preserve HTTP wire-byte validation limit
Keep structural validation separate from the serialized internal request bound so HTML escaping cannot reject valid HTTP payloads.
Assisted-by: OpenAI:gpt-5
* feat(systemone): bound images and account native decisions
Preserve public wire limits independently from router serialization. Reject unsupported NER images, map native request/capability errors, and stamp explicit usage once. Advertise decisions for stock llama-cpp.
Assisted-by: OpenAI
* fix(systemone): record usage on registered native route
Exercise real registration and billing with a mock native backend. Reject empty native responses, malformed image URLs, trailing JSON, and wire overflow including whitespace.
Assisted-by: OpenAI
* feat(router): add lazy native decision transport
Bind named models through internal ModelSystemOne calls with shared validation and bounded abandoned operations. Remove request and echoed-error contents from decision traces.
Assisted-by: OpenAI:gpt-5
* feat(router): classify overlapping policies with native decisions
Ask independent noul questions, validate probabilities and preserve first-superset routing. Wire the central factory with config-sensitive invalidation and cancellation-safe resolution. Document native framing and bounded operation limits.
Assisted-by: OpenAI:gpt-5
* feat(gallery): add pinned Julia-1 native decision model
Add a separate text-only llama-cpp Q8 entry with pinned Apache-2.0 source provenance and checksum. Installed using the gallery installer and exercised choice, score and noul on CPU.
Assisted-by: OpenAI
* test(router): verify native decisions through central factory
Add an opt-in real-model Ginkgo integration covering the native Go loader and C++ transport, token usage, independent overlapping labels, and candidate selection. Document owned-server execution and the intentionally non-quality threshold.
Assisted-by: Codex:gpt-5
* fix(llama-cpp): align upstream pin and preserve decision signatures
Advance to bed0a856 without losing the automated upstream bump. Detect full-request fill_task support at compile time and forward every question for Nimble framing while retaining the earlier native signature. Preserve reconciled SCORE/TTS patches; add standalone compatibility coverage.
Assisted-by: Codex:gpt-5
* feat(gallery): add native decision family defaults
Pin Laya, Kev-4B, lev, OpenJev and Nimble artifacts. Verify Laya/Kev/lev gallery installs and CPU contracts on both native pins; clearly mark OpenJev/Nimble runtime validation pending and their noncommercial licenses.
Assisted-by: OpenAI
* docs(decisions): clarify integrated Nimble prerequisite
Record the exact combined backend pin while retaining pending OpenJev and Nimble installation/runtime validation status.
Assisted-by: Codex:gpt-5
* fix(gallery): indent native decision model sequences
Match repository yamllint indentation for Laya, Kev, lev and OpenJev list fields. Parsed gallery data is unchanged; reproduce CI gallery lint failure before the whitespace-only fix and pass the same command afterward.
Assisted-by: Codex:gpt-5
* docs(decisions): record OpenJev and Nimble CPU validation
Record gallery installation, checksum/metadata verification and multiquestion native smoke results on bed0a856. Retain noncommercial and text-only limitations without accuracy or deterministic-output claims.
Assisted-by: OpenAI
* fix(ui): expose native Decisions router classifiers
Select classifier models using metadata-driven capability routing, retain tuned thresholds, and validate native decision selections before saving. Cover both native backends and create/save/reopen in the real React editor.
Assisted-by: Codex:gpt-5
* fix(router): exclude aliases from native decision discovery
Check the originally named config before advertising native Decisions eligibility. Retain target capability inheritance for ordinary generation aliases. Exercise the actual capabilities endpoint with native models on both backends, aliases, and disabled models.
Assisted-by: Codex:gpt-5
* feat(systemone): share bounded multimodal input validation
Preserve text wire limits while admitting bounded PNG/JPEG decision input. Share collection and header validation across internal and public callers and keep the native runner response budget independent.
Assisted-by: OpenAI:API-assistant
* fix(systemone): bound admission lifetimes and validate complete images
Retain shared admission leases through actual work completion, including abandoned internal operations. Decode bounded image pixels, cap public native responses before usage stamping, and preserve oversized malformed text status precedence.
Assisted-by: OpenAI:API-assistant
* fix(router): classify images before media fetching
Preserve ordered structured probes for native decisions. Defer OpenAI
media preparation until routing selects the served model, so rejected
decision URLs cannot trigger downloads before shared validation.
Guard direct image collection with context-aware shared admission.
Keep text classifiers and embedding caches from discarding image input.
Retain fail-closed classifier configuration and runtime fallback policy.
Add middleware, typed-content, admission, cancellation and cache tests.
Assisted-by: OpenAI:API-assistant
* fix(router): bound extraction before serialization
Check probe budgets before copying text or marshaling message state.
Count JSON escaping so oversized internal inputs fail before allocation.
Preserve typed Anthropic blocks through selected-model conversion and
fallback. Keep retry coverage in Ginkgo without global test registration.
Assisted-by: OpenAI
* fix(router): bound supported probe serialization
Arbitrary structs can bypass the probe budget through pointer marshalers,
string tags, and promoted fields. Accept concrete chat schema types and
plain JSON values instead of emulating arbitrary struct serialization.
Budget escaped direct prompts before marshaling so raw length cannot hide
serialized expansion. Preserve runtime fallback and reject oversized
input before invoking the decision runner.
Add Ginkgo allocation, boundary, and marshaler invocation regressions.
Six-package tests, three-package race tests, and full-T2 delta lint pass.
Assisted-by: OpenAI:GPT-5 golangci-lint
* feat(decisions): enable bounded OpenJev images
Validate native decision images before permissive media parsing and pixel
allocation. Require both decision image support and a vision projector;
missing or audio-only projectors cannot silently become text decisions.
Pin the OpenJev Q8 projector and document its license and disk footprint.
Add native safety tests, canonical limit parity, gallery and load-option
checks, and a reproducible CPU direct-RPC contrasting-image smoke.
Assisted-by: OpenAI:GPT-5
* fix(decisions): reject incomplete image streams
stb accepts corrupt PNG Adler checksums and truncated JPEG scans.
Use bounded zlib validation and strict libjpeg decoding before parsing.
Keep dimension and aggregate pixel checks ahead of decoder allocations.
Wire decoder dependencies into native builds and runtime packaging.
Add regressions for appended EOI and embedded marker bypasses.
Assisted-by: OpenAI:GPT-5
* fix(ci): gate native decision image validation
Run the decoder security tests outside the stdlib-only native suite.
Fetch vendor headers at the backend pin and provision decoder dependencies.
Gate Go limit parity and production CMake wiring without model downloads.
Assisted-by: OpenAI:GPT-5
* test(decisions): cover multimodal public API paths
Exercise shared image contracts through the registered HTTP routes and
external mock backend. Add opt-in cached gallery installation and real
OpenJev image decisions through SystemOne and both routing APIs.
Assisted-by: Codex:gpt-5
* test(decisions): assert isolation and cache bypass
Observe external RPC calls and compare complete classifier history.
Winner-only and cache-miss checks could hide dropped history or cache use.
Give real inference its own application and model directory so shared
backend mappings and loaded processes cannot affect mixed suite order.
Assisted-by: OpenAI:ChatGPT
* test(decisions): isolate fixture globals
Disable optional global services in the isolated HTTP fixture and register
cleanup before setup assertions. Verify meter provider identity survives
fixture creation and destruction.
Snapshot observed usage before assertions so failures cannot retain the
mutex. Require a successful usage stamp before checking error responses.
Assisted-by: Codex:gpt-5 golangci-lint
* fix(application): honor optional telemetry controls
Skip failover gauge registration when metrics are disabled. Register
against the application meter rather than looking up the global provider.
Allow embedders to retain the bounded routing log without billing stats.
Keep the existing default when stats are disabled. The isolated HTTP
fixture uses this option without losing its native router assertions.
Assisted-by: Codex:gpt-5 golangci-lint
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
When the first result of a streamed request is an error (for example a
prompt that exceeds the context), PredictStream wrote the error message
as a Reply and only then returned the error status. LocalAI treated that
Reply as the first token: it sent the assistant role chunk and the error
text as `content` on an HTTP 200 stream. Because a chunk had already been
written, the pre-stream HTTP error path from #12204 never triggered, so
streaming clients still got a 200 with the error as model output, while
the same request without streaming correctly returns a 400.
Return the error only as the gRPC status. The e2e backend suite gets a
`context_overflow` capability (enabled for llama-cpp) that streams an
over-long prompt and asserts an error status with no content.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
The environment fallback was applied whenever n_parallel was still 1
after option parsing. An explicit `parallel: 1` in the model YAML is
indistinguishable from the default that way, so it was replaced by
LLAMACPP_PARALLEL. The docs say options in the YAML take precedence
over environment variables; a single model could not be forced to one
slot while the global variable was set.
Track whether the options set the slot count and resolve it in a small
helper (parallel_params.h): option first, then LLAMACPP_PARALLEL, then
1. The helper gets a standalone unit test picked up by
`make test-backend-cpp`.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
* ⬆️ Update ggml-org/llama.cpp
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(llama-cpp): migrate score batches to the new API
The pinned engine removes common_batch_add and the raw batch view.
Use common_batch entries and llama_process for score suffix decoding.
Read shared-prefix scores from the current common_batch view.
Validation: reproduce both compiler errors on the original patch.
The patched server context and complete grpc-server translation unit
pass g++ -std=c++17 -fsyntax-only with generated protobuf headers.
Assisted-by: Codex:gpt-6
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
grpc-server.cpp forced params.cache_ram_mib = -1 (no limit) since #7009.
Since v4.3 kv_unified and cache_idle_slots are on by default, so every
distinct prompt now leaves its slot KV state in the host-side prompt
cache, and without a limit the backend grows until the host runs out of
memory.
Measured on gfx1151 (Strix Halo, 128 GB), llama-cpp backend, one request
at a time, 100 distinct prompts of ~2000 characters plus a fixed system
prompt, max_tokens 200:
model cache_ram RSS loaded -> after 100
gemma-4-26B-A4B (q8_0 KV) -1 (default) 1.4 GB -> 25.3 GB
Qwen3.6-35B-A3B (q8_0 KV) -1 (default) 1.1 GB -> 19.5 GB
gemma-4-26B-A4B -1, same prompt 100x 1.4 GB -> 1.6 GB
gemma-4-26B-A4B 4096 1.4 GB -> 5.4 GB (flat from
request 20 on, same latency)
Qwen3.6-35B-A3B 4096 1.1 GB -> 5.1 GB (flat)
The memory is not released when idle. In production a document
classification pass pushed the daily chat model to 34 GB RSS overnight.
Drop the override so llama.cpp's own default (8192 MiB) applies; the
cache_ram option still accepts -1 for users who want no limit. Update
both docs tables (the option reference and the prompt-cache table) and
note what -1 does.
Assisted-by: Claude:claude-opus-5-5
Assisted-by: Codex:GPT-6
Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
* ⬆️ Update TheTom/llama-cpp-turboquant
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(turboquant): drop the upstreamed D512 patch
Upstream c0e227c guards D512 declarations, dispatch, and instances with
GGML_USE_HIP. This prevents the CUDA shared-memory overflow that our
patch addressed. The old patch now rejects the guarded source.
Remove the obsolete patch for the pinned a3d5603d revision. The remaining
patch series applies successfully, and the build-target test passes.
Assisted-by: Codex:gpt-6
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
The previous patch only removed DECL_FATTN_VEC_CASE_D512 for turbo2_0
and turbo3_0 V cache types. turbo4_0 also overflows shared memory
(0x10100 bytes > 0xc000 max), causing ptxas errors on CUDA 12/13.
Additionally, the previous patch was incomplete: it only removed the
template instantiations but not the dispatch calls in fattn.cu or the
extern declarations in fattn-vec.cuh. This caused linker errors
(undefined reference to ggml_cuda_flash_attn_ext_vec_case_d512).
This patch removes all three layers for all turbo V types:
- Template instance .cu files (DECL_FATTN_VEC_CASE_D512)
- Dispatch calls in fattn.cu (FATTN_VEC_CASE_D512)
- Extern declarations in fattn-vec.cuh (extern DECL_FATTN_VEC_CASE_D512)
Signed-off-by: mudler <mudler@localai.io>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
* ⬆️ Update TheTom/llama-cpp-turboquant
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(turboquant): patch D512 flash-attn shared memory overflow
turboquant 4deec55 added DECL_FATTN_VEC_CASE_D512 for TURBO2_0 and
TURBO3_0 V cache types. The D=512 kernel template with these types
allocates 65 KB of shared memory, exceeding the 48 KB GPU limit:
ptxas error: Entry function uses too much shared data
(0x10100 bytes, 0xc000 max)
Carry the fix as a patch under backend/cpp/turboquant/patches/ until
TheTom/llama-cpp-turboquant#386 is merged upstream.
TURBO4_0 (4-bit) does not overflow and is left unchanged.
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Models whose options carry no explicit backend: open their session on the
CPU backend even in accelerator images. The gallery entries carry
backend:best since #11892; this covers hand-written model configurations
the same way, per deployment: the environment variable supplies the
fallback, an explicit backend: option always wins (merged beside the
existing threads and maingpu fallbacks), and validation reuses the
option parser.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>