Commit Graph
2181 Commits
Author SHA1 Message Date
Ettore Di Giacinto 53ce024efc fix(localai-proxy): clear the gosec findings
Code scanning flagged seven issues in the new backend:

- G115 text.go: tool-call indexes and tokenize lengths come from the
  upstream server as int and were cast straight to int32. Add clampInt32 so
  an absurd upstream value saturates instead of wrapping.
- G115 live.go: the int16 -> uint16 cast in PCM16 encoding is a deliberate
  two's-complement reinterpretation of an already clamped sample; mark it
  with #nosec and say so.
- G304 proxy.go, media.go, client.go: api_key_file comes from the model
  config, and the media input and output paths are files core staged or
  chose for the call. None are caller-supplied. Clean the paths and add
  #nosec with that reason, as core/gallery and the sound classification
  endpoint already do.
- G306 media.go: write generated media 0o600. Core runs as the same user
  and serves the file itself.

gosec reports 0 issues for backend/go/localai-proxy and
core/services/failover. The G104 once reported for failover/prober.go is
no longer present.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto fe8519d7b5 fix(localai-proxy): treat upstream_url as the server root, ignoring /v1
The failover prober cuts upstream_url at /v1, but Load kept the path, so
an upstream_url ending in /v1 probed healthy while every request went to
/v1/v1/... and got a 404, which never trips the target. Cut the path at
/v1 in Load too, with a warning.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto e66d87b887 fix(localai-proxy): bound live transcription setup and track in-flight turns precisely
A caller that gives up before the ready ack now ends the call with Canceled
and closes the upstream socket, and setup has a 3 minute default bound when
request_timeout_seconds is unset, so a hung upstream cannot hold the call
or block failover. A closing session now waits only for a turn the upstream
committed (or is still speaking), not for turns it discarded. Audio held
before the ready ack is capped at 5 s, and NaN samples become silence.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto c2375a7db8 feat(localai-proxy): bridge live transcription to the upstream realtime API
AudioTranscriptionLive opens <upstream>/v1/realtime?model=<realtime_pipeline>
as a transcription session with server VAD, forwards PCM as base64 PCM16
appends, and maps transcription deltas and completions to Delta/Eou. Closing
the send side waits briefly for an in-flight utterance, then sends the final
text. An upstream error, failed transcription or disconnect ends the stream
with Unavailable so failover reopens on the next target.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 37023939ad fix(localai-proxy): wrap bare base64 image inputs, block unreachable Depth exports, and validate download URLs
Detect/Depth/FaceVerify/FaceAnalyze forwarded the bare base64 core hands the
backend, but the upstream's own REST handlers only accept a URL or a
data:...;base64, string, so every real call 400'd. Wrap the payload as a
data URI (sniffing its MIME type) before sending it.

Depth requests for exports/dst now return Unimplemented: those files are
written to the upstream's own local disk and are unreachable from here, so
failover should move to a local target instead.

Generation replies that hand back a URL are now re-fetched by path only,
checked against the upstream's known generated-content prefixes, instead of
stripping the configured base as a literal string prefix — the old approach
broke (or silently trusted an arbitrary host) the moment the upstream
advertised a different base via LOCALAI_BASE_URL, a reverse proxy, or
X-Forwarded-Host.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto be9c039e65 feat(localai-proxy): serve image, video, 3D and vision APIs remotely
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 87fd7da531 feat(localai-proxy): serve speech, transcription and audio APIs remotely
The proxy now forwards TTS, streaming TTS, sound generation,
transcription (plain and streaming), diarization, VAD, sound
detection and audio transforms to the upstream LocalAI.

Streaming TTS passes the upstream WAV bytes through unchanged. A
streaming transcription that stops before its final frame, or sends
an error frame, fails with Unavailable instead of ending as a short
success. Transcription always sends diarize, because the upstream
treats a missing field as true. Audio transforms also download the
separation stems the upstream names and write them beside Dst.
Sound generation from a source clip returns Unimplemented, because
the REST endpoint has no field for the clip.

The multipart helper now takes repeated fields and several files.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 882f51fd62 fix(failover): trip a rate-limited or exhausted target instead of skipping it
Treating ResourceExhausted as a capability gap skipped the target
without counting a failure, so a target that stays rate limited or out
of memory kept its traffic. It is now an ordinary retryable failure:
the request moves to the next target and the exhausted one trips.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 6df1767133 fix(localai-proxy): keep rerank, stream errors and chat intact through the proxy
Rerank no longer sends top_n 0, which the upstream rejects. A
mid-stream upstream error frame now fails the call instead of ending
it as a short success. Temperature 0 is forwarded. An upstream 429
becomes ResourceExhausted, which failover skips like Unimplemented.

A localai-proxy config sends its own name upstream when upstream_model
is unset, and a chat proxy defaults to the tokenizer template so chat
reaches the upstream as messages.

Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
Ettore Di Giacinto 6d5c9600d7 feat(localai-proxy): add a backend that serves text APIs from a remote LocalAI
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-27 07:42:21 +00:00
localai-org-maint-botandmudler 9fa672faee chore: ⬆️ Update CrispStrobe/CrispASR to 6b78932d09765406ba0e0154d95bc6289246ceee (#12273)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:56 +02:00
localai-org-maint-botandmudler d270c2823c chore: ⬆️ Update PrismML-Eng/llama.cpp to adfffbe41b2cabcd51fff326ab045662265062bb (#12271)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:43 +02:00
localai-org-maint-botandmudler 8f29d5d271 chore: ⬆️ Update ikawrakow/ik_llama.cpp to 1aaf7105be6e55a97fa4a9fd6f5bd362b08436dc (#12270)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:32 +02:00
localai-org-maint-botandmudler ded329854c chore: ⬆️ Update 0xShug0/audio.cpp to e79205f3e0083d04e812e1a4a376f71be97e9a22 (#12269)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-26 09:01:15 +02:00
mudler-agentandEttore Di Giacinto 42c58a5838 feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29 (#12247)
* feat(vllm-cpp): unify decision pipeline through Score RPC with vllm_decide ABI v29

Replace the model-specific SystemOne gRPC approach with a generic Score
RPC extension. The pre-existing Score RPC (previously unused by any
backend) now carries question_type and response_json fields:

- question_type="systemone" routes kev/laya decision-pipeline requests
  through the unified vllm_decide C ABI (v29), returning the full
  response JSON in response_json.
- question_type empty routes cua-s1-forms candidate scoring through the
  same vllm_decide ABI, returning CandidateScore probabilities.

The vllm-cpp backend's Score() method calls vllm_decide and dispatches
by architecture internally. The /v1/systemone HTTP endpoint checks
whether the model's backend supports Score; if so, it forwards the raw
request JSON and returns the backend response as-is. Other backends
fall through to the existing NER-based path.

This mirrors the vllm.cpp C ABI refactor (PR #3301) that replaced
vllm_systemone + vllm_score with a single vllm_decide function. The
purego bindings bump abiVersion from 27 to 29 and resolve vllm_decide
and vllm_decide_free symbols.

Also fixes validModelPath to accept cua-s1-forms.json and
rl_agent_config.json alongside config.json, matching the engine's
model_loader.cpp config-filename ordering.

AI-Assisted: true
Assisted-by: Maki:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore: ⬆️ Update mudler/vllm.cpp to e28ec46c6 (fix macOS -Werror build)

Bumps vllm.cpp to e28ec46c6 which fixes a -Wnull-conversion error in
qwen3_5.cpp:12483 that broke the macOS Metal CI build under -Werror.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-25 23:27:27 +02:00
localai-org-maint-botandmudler c511b6dadf chore: ⬆️ Update ggml-org/llama.cpp to 84e76d8a23162eca70490da131945ebec1f09bf4 (#12258)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 14:59:38 +02:00
localai-org-maint-botandmudler 1768dac662 chore: ⬆️ Update PrismML-Eng/llama.cpp to 842b1880415d6f508f03b789e5ce70194def7bfd (#12250)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:29:04 +02:00
localai-org-maint-botandmudler 4668e9d141 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 8ab42195a05a9d48a3942b17568c1f3a876e133a (#12249)
⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:08:07 +02:00
localai-org-maint-botandmudler c678654d3e chore: ⬆️ Update 0xShug0/audio.cpp to 857de2366ed74bdb2c37f85259089e3a0a6b8cb0 (#12248)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:07:58 +02:00
localai-org-maint-botandmudler 1b961c0aca chore: ⬆️ Update ikawrakow/ik_llama.cpp to 20f7a72edd7049fe5a87eef2b5e9a50ae109ca4b (#12251)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:07:50 +02:00
localai-org-maint-botandmudler b010fd9b46 chore: ⬆️ Update ggml-org/whisper.cpp to d09f61a708f3487afa956ff578e60eae5e7a233c (#12253)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 09:04:37 +02:00
localai-org-maint-botandmudler b600f1b34d chore: ⬆️ Update leejet/stable-diffusion.cpp to b167b942f77ecb17e7f78e163a8c32ff7ac95c10 (#12255)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:57:57 +02:00
localai-org-maint-botandmudler a51bce57d6 chore: ⬆️ Update NVIDIA/NeMo-Speech.cpp to 97a15afa5caa9bce5baaa86c1184103877af4101 (#12257)
⬆️ Update NVIDIA/NeMo-Speech.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:18:18 +02:00
localai-org-maint-botandmudler 690a95afea chore: ⬆️ Update CrispStrobe/CrispASR to acc08e3bd3e5c17a3852115f3efa0e1ab30bc47a (#12259)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-25 08:18:04 +02:00
localai-org-maint-botandmudler 3a63699f5c chore: ⬆️ Update vllm-project/vllm cu130 wheel to 0.30.0 (#12214)
⬆️ Update vllm-project/vllm cu130 wheel

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 21:20:07 +02:00
localai-org-maint-botandmudler c6a574ef1b chore: ⬆️ Update vllm-metal (darwin) to v0.30.0 (#12225)
⬆️ Update vllm-project/vllm-metal (darwin)

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 21:19:30 +02:00
localai-org-maint-botandmudler 11b3b184ea chore: ⬆️ Update PrismML-Eng/llama.cpp to 0324c66521960d67aa7da8687fb1453a79a6565c (#12226)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 21:19:15 +02:00
localai-org-maint-botandmudler e143f14551 chore: ⬆️ Update ggml-org/whisper.cpp to a664346ea5c6dddff3e61a2b7b32dd4514613f50 (#12227)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 21:19:02 +02:00
localai-org-maint-botandmudler fdd19c7f76 chore: ⬆️ Update ggml-org/llama.cpp to d2e54583c7452353eb35d40431281f6ee984332f (#12228)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 21:18:47 +02:00
localai-org-maint-botandmudler df45e6cab8 chore: ⬆️ Update 0xShug0/audio.cpp to 9bdd1d908bbd128e9eb405f5a8e38d0defb84c72 (#12224)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 17:05:19 +02:00
localai-org-maint-botandmudler 8588020d77 chore: ⬆️ Update mudler/vllm.cpp to b24f8094cba9b4f02df71bcff8d41ddc7e88b4ef (#12229)
⬆️ Update mudler/vllm.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 09:32:52 +02:00
localai-org-maint-botandmudler 694e1ea3cc chore: ⬆️ Update CrispStrobe/CrispASR to 97a35a6e519fda1835f3c8353516384aa8710b8c (#12230)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 09:32:33 +02:00
localai-org-maint-botandmudler 91b462db24 chore: ⬆️ Update ikawrakow/ik_llama.cpp to f3d6e6e3020ddfebad60113845bf521620766da5 (#12233)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-24 09:05:08 +02:00
mudler-agentandEttore Di Giacinto 81eaca8768 fix(turboquant): extend D512 flash-attn patch to all turbo V types (#12234)
The previous patch only removed DECL_FATTN_VEC_CASE_D512 for turbo2_0
and turbo3_0 V cache types. turbo4_0 also overflows shared memory
(0x10100 bytes > 0xc000 max), causing ptxas errors on CUDA 12/13.

Additionally, the previous patch was incomplete: it only removed the
template instantiations but not the dispatch calls in fattn.cu or the
extern declarations in fattn-vec.cuh. This caused linker errors
(undefined reference to ggml_cuda_flash_attn_ext_vec_case_d512).

This patch removes all three layers for all turbo V types:
- Template instance .cu files (DECL_FATTN_VEC_CASE_D512)
- Dispatch calls in fattn.cu (FATTN_VEC_CASE_D512)
- Extern declarations in fattn-vec.cuh (extern DECL_FATTN_VEC_CASE_D512)

Signed-off-by: mudler <mudler@localai.io>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-24 08:01:28 +02:00
2ccda5ba92 chore: ⬆️ Update TheTom/llama-cpp-turboquant to 4deec5587b2963af00bdf80884f3337e02eb7d64 (#12154)
* ⬆️ Update TheTom/llama-cpp-turboquant

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(turboquant): patch D512 flash-attn shared memory overflow

turboquant 4deec55 added DECL_FATTN_VEC_CASE_D512 for TURBO2_0 and
TURBO3_0 V cache types. The D=512 kernel template with these types
allocates 65 KB of shared memory, exceeding the 48 KB GPU limit:

  ptxas error: Entry function uses too much shared data
  (0x10100 bytes, 0xc000 max)

Carry the fix as a patch under backend/cpp/turboquant/patches/ until
TheTom/llama-cpp-turboquant#386 is merged upstream.

TURBO4_0 (4-bit) does not overflow and is left unchanged.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-23 21:12:38 +00:00
9ad18c8c67 chore: ⬆️ Update ServeurpersoCom/omnivoice.cpp to 3ac485d0688fc684f5dcf2c95283b745220f9dcc (#12191)
* ⬆️ Update ServeurpersoCom/omnivoice.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>

* fix(omnivoice-cpp): workaround GGML_SOURCE_DIR for nested builds

omnivoice.cpp 3ac485d changed from a relative add_subdirectory(ggml)
to CMAKE_SOURCE_DIR-based path resolution. CMAKE_SOURCE_DIR points to
the top-level project, not the current subdirectory, so when
omnivoice is consumed via add_subdirectory() the build fails:

  add_subdirectory given source ".../omnivoice-cpp/ggml"
  which is not an existing directory.

Set GGML_SOURCE_DIR to the correct path before add_subdirectory so
the upstream code picks it up from the cache. This is a workaround
until ServeurpersoCom/omnivoice.cpp#20 is merged upstream.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-23 22:13:58 +02:00
localai-org-maint-botandEttore Di Giacinto d71a22f1ae chore: ⬆️ Update mudler/vllm.cpp to d4738d241271b4d10134a6499f97337f20fcf8ce (#12175)
vllm.cpp d4738d2 bumped VLLM_ABI_VERSION from 26 to 27. The
abi-check target caught the mismatch: the Go struct mirrors in
govllmcpp.go still declared v26.

Update abiVersion, the header comment, and the test expectation
to v27. The struct layout did not change between the two versions,
so no offset adjustments are needed.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-23 21:21:53 +02:00
21c5495a99 feat(vllm-cpp): add GLiNER2.5 NER via TokenClassify (#12140)
* feat(vllm-cpp): add GLiNER2.5 NER via TokenClassify

Wire the vllm-cpp backend to the C ABI NER surface (vllm_gliner_ner,
ABI v27) so LocalAI can serve zero-shot named entity recognition through
the existing TokenClassify gRPC method.

backend.go: TokenClassify method on *VllmCpp calls vllm_gliner_ner with
the text and labels, copies the C-owned entity array into protobuf
TokenClassifyEntity messages, and frees the result.

govllmcpp.go: cNerEntity and cNerResult Go POD mirrors matching the C
structs; vllmGlinerNer and vllmNerResultFree purego bindings; abiVersion
bumped 26 -> 27.

options.go: ner_labels, ner_threshold, ner_max_width parsed from
engine_args.

pkg/grpc: ClassifyModel interface and TokenClassify server handler
(follows the Embedding locking pattern).

core/config: vllm-cpp backend declares MethodTokenClassify and
UsecaseTokenClassify.

docs/content/features/vllm-cpp.md: NER section documenting the
engine_args keys and the host-forward contract.

Assisted-by: MAKI:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): correct NER pointer lint directive

Use the govet directive for the C-owned NER array, matching the other
purego pointer conversions. The array remains valid until its deferred
free; the misspelled directive caused CI to flag this conversion.

Assisted-by: Codex:gpt-6 golangci-lint
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(vllm-cpp): add kev-compatible SystemOne API endpoints

Add POST /v1/systemone, /v1/systemone/permute, and
/v1/systemone/separate to LocalAI, mirroring the kev project's
structured-extraction API. Each endpoint runs zero-shot NER over the
rendered state text and builds kev-compatible answers for three question
types: noul (binary entity presence), choice (pick one option), and
score (pick one level).

The TokenClassifyRequest proto gains a `repeated string labels` field so
each question can supply its own labels at inference time, and
TokenClassifier gains TokenClassifyWithLabels for per-call label
selection. The vllm-cpp backend uses request labels when non-empty,
falling back to configured ner_labels then the built-in defaults.

Helpers (renderState, softmax, choiceConfidence, scoreConfidence, r2)
are ported from kev/api.py and mirrored in vllm.cpp's api_server.cpp so
both servers produce the same answer shape.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): suppress gosec G404 on seeded permutation RNG

The SystemOne permute endpoint uses math/rand with a caller-supplied
seed for reproducible option permutations, matching kev's random.seed.
gosec flags this as G404 (weak RNG). Add #nosec with a comment naming
the intent: this is reproducibility, not cryptography.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(vllm-cpp): bump vllm.cpp pin to GLiNER2.5 merge commit

Advance VLLM_CPP_VERSION from f3cd97e to 5058268d, the commit that
landed GLiNER2.5 zero-shot NER support (PR #3224) in vllm.cpp. This
brings the DeBERTa v2 encoder, GLiNER2 boundary head, C ABI NER
functions, and server endpoints into the LocalAI vllm-cpp backend.
The ABI version (27) and Go struct mirrors already match.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(vllm-cpp): use instruction text as NER label in SystemOne handler

The SystemOne handler was passing question IDs as NER labels for noul
questions and bare key names for choice questions, so the model never
matched any entities. Port the label mapping from vllm.cpp's
ParseSystemOneBody:

- noul: use the rendered instructions field (with instr alias) as the
  NER label, not the question ID
- choice: use optionText(name, desc) — "name: description" or "name"
  when the description is null/empty — not the bare key
- score: already correct (rendered criteria text)
- permute: shuffle indices and build parallel key/label arrays so the
  NER call uses the optionText labels while the response is keyed by
  the original option names

Also add the instructions field to the SystemOneQuestion schema struct
(accepted alongside the instr backward-compat alias).

Verified end-to-end against the real GLiNER2.5 model: noul questions
now find "Apple Inc. is" (organization, 0.999) and "Tim Cook is"
(person, 0.852) where they previously returned zero entities.

Following-Agents-Protocol: true
AI-Assisted: true
Assisted-by: AGENT:regolo/glm5.2 [maki]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-09-23 12:31:28 +02:00
X c5216b36af fix(vllm-omni): remove invalid syntax in test.py (#12135)
Signed-off-by: hendrixx-cnc <tjhendrx@icloud.com>
2026-09-23 12:24:21 +02:00
pos-ei-don 0b09cff673 fix(sglang): support msgspec-based ServerArgs (sglang >= 0.5.20) (#12155)
sglang 0.5.20 moved its config tier from dataclasses to msgspec.Struct
(sgl-project/sglang#38753). _apply_engine_args validates engine_args keys
via dataclasses.fields(ServerArgs), which raises TypeError there. That
call runs on every LoadModel, so no model loads at all on the sglang
backend once sglang >= 0.5.20 is installed, and the error surfaces as a
generic "Unexpected <class 'TypeError'>" that does not name the cause.

Introspect both shapes: msgspec structs carry their field names in
__struct_fields__, so key validation and the close-match suggestion keep
working, and older dataclass-based sglang stays supported.

Adds a test that pins the msgspec path with a stand-in, so it is covered
regardless of which sglang version is installed.

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
2026-09-23 12:23:35 +02:00
Plamen K. Kosseff e7306a087a feat(audio-cpp): AUDIOCPP_DEFAULT_BACKEND fallback for models without a backend option (#12133)
Models whose options carry no explicit backend: open their session on the
CPU backend even in accelerator images. The gallery entries carry
backend:best since #11892; this covers hand-written model configurations
the same way, per deployment: the environment variable supplies the
fallback, an explicit backend: option always wins (merged beside the
existing threads and maingpu fallbacks), and validation reuses the
option parser.

Assisted-by: Claude:claude-fable-5

Signed-off-by: Plamen K. Kosseff <p.kosseff@gmail.com>
2026-09-23 12:16:31 +02:00
localai-org-maint-botandmudler 2e279920b0 chore: ⬆️ Update CrispStrobe/CrispASR to 18d74132d22fa7c967181720310d4ba1df9b7bf0 (#12213)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-23 11:04:36 +02:00
localai-org-maint-botandmudler c93bd4da85 chore: ⬆️ Update leejet/stable-diffusion.cpp to c92d73c408515c94beef32161bb5960764fde7a0 (#12212)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-23 11:04:23 +02:00
localai-org-maint-botandmudler 01c60b22e6 chore: ⬆️ Update 0xShug0/audio.cpp to 1ee4ce8275997a7dcf0e2a5dc3410e509b898d6d (#12211)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-23 11:04:08 +02:00
localai-org-maint-botandmudler d8bc8445c9 chore: ⬆️ Update ggml-org/whisper.cpp to a44e07845931421bb6f3447ce0010ed9dc76a118 (#12209)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-23 09:31:33 +02:00
localai-org-maint-botandmudler 90e6cae60e chore: ⬆️ Update PrismML-Eng/llama.cpp to bdc23b56b4458b9f1655aec5287f3ab56ee8daaa (#12207)
⬆️ Update PrismML-Eng/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-23 09:28:12 +02:00
localai-org-maint-botandmudler 82f7b25766 chore: ⬆️ Update ikawrakow/ik_llama.cpp to c5b5773bed338c5f3b985d277764a4d780b83d42 (#12210)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-23 09:26:14 +02:00
localai-org-maint-botandmudler 74b5cf1311 chore: ⬆️ Update ggml-org/llama.cpp to 709fe755dfa810d77e2ac386292b29648b536864 (#12208)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-23 09:26:01 +02:00
localai-org-maint-botandmudler f197c35a0f chore: ⬆️ Update ggml-org/whisper.cpp to 307869af285d7f6f689ba100b3515e2d1b3feb05 (#12200)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-22 12:57:12 +02:00
localai-org-maint-botandmudler 2d55aafb86 chore: ⬆️ Update ggml-org/llama.cpp to 58367713a6935c0810103378144008df32e3d5db (#12197)
⬆️ Update ggml-org/llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-09-22 12:57:00 +02:00