mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-29 09:35:02 -04:00
b449ad3828b0949ecbcc00db75626ff89a1002c0
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
61f4f67b75 |
sglang backend: pass through thinking_budget + require_reasoning (#12193)
* sglang backend: pass through thinking_budget + require_reasoning sglang's raw Engine.async_generate() API (which this backend calls directly, bypassing sglang's own OpenAI server) supports a precise, tokenizer-derived reasoning-length budget via sampling_params["custom_params"]["thinking_budget"] plus require_reasoning=True, gated behind --enable-strict-thinking. Neither was reachable through LocalAI: this backend built sampling_params only from a fixed field mapping (temperature, top_p, ...) with no custom_params key, and never passed require_reasoning to async_generate at all. - LoadModel now reads a model-level "thinking_budget" option (same mechanism as the existing tool_parser/reasoning_parser options), and _build_sampling_params adds it as custom_params.thinking_budget on every request when configured. - _new_reasoning_parser already derives, from the rendered prompt, whether the model's chat template pre-opened a reasoning block (Qwen3-style templates append <think> to the prompt instead of letting the model emit it) -- the same signal sglang's own OpenAI server computes from per-template config to decide require_reasoning. This backend has no template manager, so it now returns that signal too and _predict forwards it to async_generate(require_reasoning=...). Verified against production (NVFP4, sm_121, Qwen3.6-35B-A3B) via a raw Engine.async_generate() call bypassing this backend: 301 reasoning tokens against a 300-token budget, clean completion, ~27s. Not yet verified through this backend's own gRPC path end-to-end (no local CUDA/sglang environment available here) -- existing + new unit tests in test.py cover the pure-Python merge/passthrough logic only. Scope note: require_reasoning is derived only from the existing prompt-suffix heuristic, not sglang's full per-template _get_reasoning_from_request decision tree (minimax-m3/hunyuan special cases etc.) -- this backend has no template manager to evaluate that tree against, and the prompt-suffix check is the one heuristic already validated in this file (test_reasoning_parser_forced_when_template_prefills_think_tag). Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> * sglang backend: honour a model-level reasoning_default A model YAML can already carry "parameters: reasoning_effort:", but that value only reaches this backend when a *caller* sets it per request (the Go side turns it into Metadata["enable_thinking"]). As a model-level default it is silently dropped: a config reading "reasoning_effort: none" still produces full reasoning on every request, so the config says one thing and the model does another. That gap is expensive in practice. On a self-hosted Qwen3.6-35B-A3B the reasoning phase consumed the entire max_tokens budget before any content was produced - 90% of code completions came back empty at max_tokens=768, and the server log filled with "backend produced only reasoning, retrying". The config looked like reasoning was off the whole time. This adds "reasoning_default:off" (or ":on") on the same model-level options: mechanism as thinking_budget. A per-request value always wins; the default only fills in when the request is silent. Measured on the stack above (sglang 0.5.20, NVFP4, GB10/sm_121) after applying it: default (nothing set) -> 0 chars reasoning, 27 tokens "reasoning_effort": "none" -> 0 chars reasoning, 27 tokens metadata enable_thinking=true -> capped at the 512-token thinking_budget, 541 tokens total, finish_reason stop Tests: three cases added to backend/python/sglang/test.py covering the default, per-request override in both directions, and the unconfigured case (which must leave the template untouched). Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> * sglang backend: validate thinking_budget instead of crashing LoadModel Addresses the review on this PR: - `int(thinking_budget)` raised on values like "5000.0" or "abc" and took LoadModel down. The option is now parsed by _parse_thinking_budget(): integral numbers in any spelling are accepted, anything else is ignored with a warning on stderr. - Zero and negative budgets are ignored with a warning instead of being passed to sglang, where they have no defined meaning. Turning reasoning off is what reasoning_default:off is for. - A load-time warning when thinking_budget is set but enable_strict_thinking is not in engine_args, since sglang then ignores the budget silently. - Tests for integral spellings, unset, zero, negative, non-integer and the strict-thinking warning. Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> * docs(sglang): explain reasoning options Document the reasoning budget, strict-thinking requirement, and precedence of request metadata over the model-level default. Also note that the budget has to stay well below max_tokens (otherwise it never triggers and the reply can end up empty), and that POST /models/reload or a backend-only restart does not pick up changed options; LocalAI itself has to be restarted. Assisted-by: Codex:GPT-6 Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> * docs(sglang): clarify configuration reloads Distinguish rereading model configuration from updating a running backend. Keep the full LocalAI restart recommendation for changed reasoning options. Assisted-by: Codex:GPT-6 * sglang backend: only pass require_reasoning when sglang supports it Engine.async_generate() gained the require_reasoning keyword in sglang 0.5.13 and takes no **kwargs. The CPU profile builds v0.5.11 from source and the other profiles only set a >=0.5.11 floor, so passing the keyword unconditionally made every request fail with TypeError. Detect support once at import time, as the file already does for sampling_seed. enable_strict_thinking first appears in sglang 0.5.12; fix the comment. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code] --------- Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <localai-org-maint-bot@users.noreply.github.com> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
0b09cff673 |
fix(sglang): support msgspec-based ServerArgs (sglang >= 0.5.20) (#12155)
sglang 0.5.20 moved its config tier from dataclasses to msgspec.Struct (sgl-project/sglang#38753). _apply_engine_args validates engine_args keys via dataclasses.fields(ServerArgs), which raises TypeError there. That call runs on every LoadModel, so no model loads at all on the sglang backend once sglang >= 0.5.20 is installed, and the error surfaces as a generic "Unexpected <class 'TypeError'>" that does not name the cause. Introspect both shapes: msgspec structs carry their field names in __struct_fields__, so key validation and the close-match suggestion keep working, and older dataclass-based sglang stays supported. Adds a test that pins the msgspec path with a stand-in, so it is covered regardless of which sglang version is installed. Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> |
||
|
|
4894056380 |
fix(sglang): force reasoning when the template prefills the think tag
Qwen3-style chat templates append the opening <think> tag to the *prompt*
when thinking is enabled. The model therefore never generates it and emits
only the reasoning text plus the closing </think>.
sglang's ReasoningParser keys off the opening tag:
in_reasoning = self._in_reasoning or self.think_start_token in text
if not in_reasoning:
return StreamingParseResult(normal_text=text)
so with such a template the entire completion — reasoning and answer, the
raw </think> in between — is returned as content and reasoning_content
stays empty, no matter how reasoning_parser is configured.
sglang's own OpenAI server handles this via
force_reasoning = (self.template_manager.force_reasoning
or self._get_reasoning_from_request(request))
This backend has no template manager, so derive the same signal from the
rendered prompt: if it ends with the detector's think_start_token, the tag
was prefilled and the parser is constructed with force_reasoning=True.
Structured decoding is the exception, and it matters: a grammar applies
from the first token, so the model cannot emit the closing tag even though
the template opened the block. The whole completion is schema output and
belongs in content — forcing there files it as reasoning and returns an
empty answer. Measured against a JSON-schema code audit: 10107 characters
of "reasoning", zero content. sglang's own server keeps the two apart for
the same reason; its grammar backend owns the reasoning prefix when a
reasoning parser is configured.
force_reasoning is only passed when it is meant to be True, so detector
defaults (DeepSeek-R1 already defaults to True) are untouched, and a
prompt without a prefilled tag behaves exactly as before — which matters,
because forcing unconditionally makes an answer generated with thinking
off disappear into reasoning_content.
The construction is factored into _new_reasoning_parser() so the streaming
and non-streaming paths, which previously built the parser separately,
cannot drift apart.
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
|
||
|
|
9319450aa6 |
fix(python-backends): re-attach media markers under use_tokenizer_template (#11621)
With `template.use_tokenizer_template: true` the sglang and vllm backends
render the prompt themselves via `tokenizer.apply_chat_template()`, and they
hand it plain string content. A chat template only emits the model's own media
tokens when the content is a list of parts, so the rendered prompt carries no
`<|vision_start|><|image_pad|><|vision_end|>`. The pixels do reach the engine
(`image_data` / `multi_modal_data`), but both engines locate them by scanning
the prompt for that token, so they are discarded silently: HTTP 200, no
warning, and the model answers as if no image had been attached.
Add `attach_media_parts()` to the shared `python_utils` helper and call it in
both backends: the last user turn is rebuilt as
`[{"type": "image"} * n, {"type": "video"} * n, {"type": "text", ...}]` before
templating, which makes the template emit the placeholders. The pixels keep
travelling out of band exactly as before.
Text-only requests are untouched - with no media the helper returns None and
the original string-content path runs unchanged. If a template cannot iterate
content parts (a text-only model), the parts render is caught and the request
falls back to the previous string-content prompt instead of failing.
Signed-off-by: Tai An <antai12232931@outlook.com>
|
||
|
|
84db1e6430 |
fix(backends): preserve an explicit seed of 0 in sglang and vllm
#11772 exempted Temperature from the zero-filter in both backend adapters, because proto3 has no field presence and an explicit 0 is indistinguishable from "unset". Seed has exactly the same property and is still filtered: if proto_field != "Temperature" and value in (None, 0, 0.0, [], False, ""): continue A caller pinning `"seed": 0` for a reproducible run therefore gets a random seed instead, with no error and no log line — the one case where the failure is invisible precisely because the request looked deliberate. Both adapters now share a named tuple of fields whose zero is meaningful, so the next one is added in one place rather than as a second special case. Deliberately left filtered: top_k, top_p, min_p and the penalties. Their zero is not a value a caller means — sglang disables top_k with -1, not 0, so forwarding 0 there would turn a default into an invalid argument. Verified on the sglang backend (Qwen3.5-MoE, arm64): with the temperature fix alone, two identical requests at temperature 0 are byte-identical, but pinning seed 0 has no effect until this change. Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> |
||
|
|
62f1c0ca7f |
feat(gallery): add PhoneLLM variants (#11772)
Add complete vLLM and SGLang entries with their exact tool parsers. Preserve an explicit zero temperature in both backend adapters. Assisted-by: Codex:gpt-5 Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com> |
||
|
|
a760a7ab4b |
fix(backends): honor enable_thinking=false in sglang and vllm (#11715)
Those backends only forwarded the flag when it was "true", so "false" never reached apply_chat_template and Qwen3 kept thinking on. Signed-off-by: lei_lei <96427312+leilei3167@users.noreply.github.com> |
||
|
|
e62221b020 |
fix(sglang): implement Status RPC to unblock backend-monitor polling (#10867)
The sglang Python backend inherits the default Status RPC from
backend_pb2_grpc.BackendServicer, which raises NotImplementedError.
LocalAI's backend-monitor polls /backend.Backend/Status periodically on
every registered backend; when the call fails, /backend/monitor returns
HTTP 500 and downstream inference requests to the sglang backend are
blocked even though the model is loaded and answering directly via the
gRPC endpoint.
Add a minimal Status shim that mirrors the existing Health method and
returns StatusResponse{state=READY} unconditionally. This unblocks the
monitor path; a state-aware follow-up (UNINITIALIZED during load, BUSY
under active inference) is left for a subsequent change.
Reproduced on DGX Spark (GB10, arm64-l4t-cuda-13 image) with the sglang
v0.5.15 backend and Qwen3-Coder-Next-NVFP4-GB10; verified locally that
patching the shim in place immediately restores /backend/monitor and
inference across the sglang slot.
|
||
|
|
bcc41219f7 |
feat: materialize Hugging Face model artifacts (#10825)
* feat(config): add model artifact source contract Assisted-by: Codex:GPT-5 [Codex] * feat(downloader): add authenticated raw-byte progress Assisted-by: Codex:GPT-5 [Codex] * feat(huggingface): resolve immutable snapshot manifests Assisted-by: Codex:GPT-5 [Codex] * feat(models): add artifact storage primitives Assisted-by: Codex:GPT-5 [Codex] * feat(models): materialize pinned Hugging Face snapshots Assisted-by: Codex:GPT-5 [Codex] * feat(models): bind managed snapshots at runtime Assisted-by: Codex:GPT-5 [Codex] * feat(gallery): materialize model artifacts during install Assisted-by: Codex:GPT-5 [Codex] * feat(gallery): declare managed Hugging Face artifacts Assisted-by: Codex:GPT-5 [Codex] * feat(models): preload managed model artifacts Assisted-by: Codex:GPT-5 [Codex] * fix(gallery): retain shared artifact caches on delete Assisted-by: Codex:GPT-5 [Codex] * feat(models): report artifact acquisition progress Assisted-by: Codex:GPT-5 [Codex] * refactor(backends): load managed models from ModelFile Assisted-by: Codex:GPT-5 [Codex] * refactor(backends): load staged speech model snapshots Assisted-by: Codex:GPT-5 [Codex] * refactor(backends): use staged snapshots in engine backends Assisted-by: Codex:GPT-5 [Codex] * test(distributed): cover staged artifact snapshots Assisted-by: Codex:GPT-5 [Codex] * docs: explain managed model artifacts Assisted-by: Codex:GPT-5 [Codex] * docs: add product design context Assisted-by: Codex:GPT-5 [Codex] * feat(ui): show model artifact download progress Assisted-by: Codex:GPT-5 [Codex] * Eagerly materialize Hugging Face artifacts Materialize HF-backed model references as managed GGUF artifacts during load, with lazy download retained only as fallback. Assisted-by: Codex:GPT-5 [shell] * Refactor HF downloads through a shared executor Assisted-by: Codex:GPT-5 [shell] * drop Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> |
||
|
|
6ab29ec8b9 |
fix(sglang): parse tool_call function arguments before applying the chat template (#10558)
OpenAI wire format carries `function.arguments` as a JSON-encoded string, but chat templates (e.g. Qwen3-Coder) iterate over it as a mapping. The vllm backend already parses arguments before applying the chat template (PR #10256); this mirrors that fix in the sglang backend. Without this fix the second turn of any tool-using session (assistant returns tool_calls, user posts `role:"tool"` result, model is invoked with arguments still as a string) crashes inside transformers' Jinja chat-template rendering with: TypeError: Can only get item pairs from a mapping. File ".../transformers/utils/chat_template_utils.py", in render_jinja_template File ".../jinja2/filters.py", in do_items raise TypeError("Can only get item pairs from a mapping.") Reproduced on `lmsysorg/sglang:v0.5.14` via LocalAI v4.5.4 with `saricles/Qwen3-Coder-Next-NVFP4-GB10` (W4A4 NVFP4 / compressed-tensors) on NVIDIA DGX Spark (GB10, sm_121). After the patch, a tool-call roundtrip (assistant tool_calls -> tool result -> assistant final answer) returns http=200 with the expected follow-up content; no behaviour change on requests that don't carry tool_calls. Signed-off-by: Poseidon <philipp.wacker@ibf-solutions.com> Co-authored-by: Poseidon <philipp.wacker@ibf-solutions.com> |
||
|
|
c894d9c826 |
feat(sglang): wire engine_args, add cuda13 build, ship MTP gallery demos (#9686)
Bring the sglang Python backend up to feature parity with vllm by adding
the same engine_args:-map plumbing the vLLM backend already has. Any
ServerArgs field (~380 in sglang 0.5.11) becomes settable from a model
YAML, including the speculative-decoding flags needed for Multi-Token
Prediction. Validation matches the vllm backend's: keys are checked
against dataclasses.fields(ServerArgs), unknown keys raise ValueError
with a difflib close-match suggestion at LoadModel time, and the typed
ModelOptions fields keep their existing meaning with engine_args
overriding them.
Backend code:
* backend/python/sglang/backend.py: add _apply_engine_args, import
dataclasses/difflib/ServerArgs, call from LoadModel; rename Seed ->
sampling_seed (sglang 0.5.11 renamed the SamplingParams field).
* backend/python/sglang/test.py + test.sh + Makefile: six unit tests
exercising the helper directly (no engine load required).
Build / CI / backend gallery (cuda13 + l4t13 paths are now first-class):
* backend/python/sglang/install.sh: add --prerelease=allow because
sglang 0.5.11 hard-pins flash-attn-4 which only ships beta wheels;
add --index-strategy=unsafe-best-match for cublas12 so the cu128
torch index wins over default-PyPI's cu130; new pyproject.toml-driven
l4t13 install path so [tool.uv.sources] can pin torch/torchvision/
torchaudio/sglang to the jetson-ai-lab index without forcing every
transitive PyPI dep through the L4T mirror's flaky proxy (mirrors the
equivalent fix in backend/python/vllm/install.sh).
* backend/python/sglang/pyproject.toml (new): L4T project spec with
explicit-source jetson-ai-lab index. Replaces requirements-l4t13.txt
for the l4t13 BUILD_PROFILE; other profiles still go through the
requirements-*.txt pipeline via libbackend.sh's installRequirements.
* backend/python/sglang/requirements-l4t13.txt: removed; superseded
by pyproject.toml.
* backend/python/sglang/requirements-cublas{12,13}{,-after}.txt: pin
sglang>=0.5.11 (Gemma 4 floor); add cu130 torch index for cublas13
(new files) and cu128 torch index for cublas12 (default PyPI now
ships cu130 torch wheels by default and breaks cu12 hosts).
* backend/index.yaml: add cuda13-sglang and cuda13-sglang-development
capability mappings + image entries pointing at
quay.io/.../-gpu-nvidia-cuda-13-sglang.
* .github/workflows/backend.yml: new cublas13 sglang matrix entry,
mirroring vllm's cuda13 build.
Model gallery + docs:
* gallery/sglang.yaml: base sglang config template, mirrors vllm.yaml.
* gallery/sglang-gemma-4-{e2b,e4b}-mtp.yaml: Gemma 4 MTP demos
transcribed verbatim from the SGLang Gemma 4 cookbook MTP commands.
* gallery/sglang-mimo-7b-mtp.yaml: MiMo-7B-RL with built-in MTP heads
+ online fp8 weight quantization, verified end-to-end on a 16 GB
RTX 5070 Ti at ~88 tok/s. Uses mem_fraction_static: 0.7 because the
MTP draft worker's vocab embedding is loaded unquantised and OOMs
the static reservation at sglang's 0.85 default.
* gallery/index.yaml: three new entries (gemma-4-e2b-it:sglang-mtp,
gemma-4-e4b-it:sglang-mtp, mimo-7b-mtp:sglang).
* docs/content/features/text-generation.md: new SGLang section with
setup, engine_args reference, MTP demos, version requirements.
* .agents/sglang-backend.md (new): agent one-pager covering the flat
ServerArgs structure, the typed-vs-engine_args precedence, the
speculative-decoding cheatsheet, and the mem_fraction_static gotcha
documented above.
* AGENTS.md: index entry for the new agent doc.
Known limitation: the two Gemma 4 MTP gallery entries ship a recipe
that doesn't yet run on stock libraries. The drafter checkpoints
(google/gemma-4-{E2B,E4B}-it-assistant) declare
model_type: gemma4_assistant / Gemma4AssistantForCausalLM, which
neither transformers (<=5.6.0, including the SGLang cookbook's pinned
commit 91b1ab1f... and main HEAD) nor sglang's own model registry
(<=0.5.11) registers as of 2026-05-06. They will start working when
HF or sglang upstream registers the architecture -- no LocalAI
changes needed. The MiMo MTP demo and the non-MTP Gemma 4 paths work
today on this build (verified on RTX 5070 Ti, 16 GB).
Assisted-by: Claude:claude-opus-4-7 [Read] [Edit] [Bash] [WebFetch] [WebSearch]
Signed-off-by: Richard Palethorpe <io@richiejp.com>
|
||
|
|
b4e30692a2 |
feat(backends): add sglang (#9359)
* feat(backends): add sglang Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(sglang): force AVX-512 CXXFLAGS and disable CI e2e job sgl-kernel's shm.cpp uses __m512 AVX-512 intrinsics unconditionally; -march=native fails on CI runners without AVX-512 in /proc/cpuinfo. Force -march=sapphirerapids so the build always succeeds, matching sglang upstream's docker/xeon.Dockerfile recipe. The resulting binary still requires an AVX-512 capable CPU at runtime, so disable tests-sglang-grpc in test-extra.yml for the same reason tests-vllm-grpc is disabled. Local runs with make test-extra-backend-sglang still work on hosts with the right SIMD baseline. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(sglang): patch CMakeLists.txt instead of CXXFLAGS for AVX-512 CXXFLAGS with -march=sapphirerapids was being overridden by add_compile_options(-march=native) in sglang's CPU CMakeLists.txt, since CMake appends those flags after CXXFLAGS. Sed-patch the CMakeLists.txt directly after cloning to replace -march=native. --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> |