mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-29 17:44:30 -04:00
* sglang backend: pass through thinking_budget + require_reasoning sglang's raw Engine.async_generate() API (which this backend calls directly, bypassing sglang's own OpenAI server) supports a precise, tokenizer-derived reasoning-length budget via sampling_params["custom_params"]["thinking_budget"] plus require_reasoning=True, gated behind --enable-strict-thinking. Neither was reachable through LocalAI: this backend built sampling_params only from a fixed field mapping (temperature, top_p, ...) with no custom_params key, and never passed require_reasoning to async_generate at all. - LoadModel now reads a model-level "thinking_budget" option (same mechanism as the existing tool_parser/reasoning_parser options), and _build_sampling_params adds it as custom_params.thinking_budget on every request when configured. - _new_reasoning_parser already derives, from the rendered prompt, whether the model's chat template pre-opened a reasoning block (Qwen3-style templates append <think> to the prompt instead of letting the model emit it) -- the same signal sglang's own OpenAI server computes from per-template config to decide require_reasoning. This backend has no template manager, so it now returns that signal too and _predict forwards it to async_generate(require_reasoning=...). Verified against production (NVFP4, sm_121, Qwen3.6-35B-A3B) via a raw Engine.async_generate() call bypassing this backend: 301 reasoning tokens against a 300-token budget, clean completion, ~27s. Not yet verified through this backend's own gRPC path end-to-end (no local CUDA/sglang environment available here) -- existing + new unit tests in test.py cover the pure-Python merge/passthrough logic only. Scope note: require_reasoning is derived only from the existing prompt-suffix heuristic, not sglang's full per-template _get_reasoning_from_request decision tree (minimax-m3/hunyuan special cases etc.) -- this backend has no template manager to evaluate that tree against, and the prompt-suffix check is the one heuristic already validated in this file (test_reasoning_parser_forced_when_template_prefills_think_tag). Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> * sglang backend: honour a model-level reasoning_default A model YAML can already carry "parameters: reasoning_effort:", but that value only reaches this backend when a *caller* sets it per request (the Go side turns it into Metadata["enable_thinking"]). As a model-level default it is silently dropped: a config reading "reasoning_effort: none" still produces full reasoning on every request, so the config says one thing and the model does another. That gap is expensive in practice. On a self-hosted Qwen3.6-35B-A3B the reasoning phase consumed the entire max_tokens budget before any content was produced - 90% of code completions came back empty at max_tokens=768, and the server log filled with "backend produced only reasoning, retrying". The config looked like reasoning was off the whole time. This adds "reasoning_default:off" (or ":on") on the same model-level options: mechanism as thinking_budget. A per-request value always wins; the default only fills in when the request is silent. Measured on the stack above (sglang 0.5.20, NVFP4, GB10/sm_121) after applying it: default (nothing set) -> 0 chars reasoning, 27 tokens "reasoning_effort": "none" -> 0 chars reasoning, 27 tokens metadata enable_thinking=true -> capped at the 512-token thinking_budget, 541 tokens total, finish_reason stop Tests: three cases added to backend/python/sglang/test.py covering the default, per-request override in both directions, and the unconfigured case (which must leave the template untouched). Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> * sglang backend: validate thinking_budget instead of crashing LoadModel Addresses the review on this PR: - `int(thinking_budget)` raised on values like "5000.0" or "abc" and took LoadModel down. The option is now parsed by _parse_thinking_budget(): integral numbers in any spelling are accepted, anything else is ignored with a warning on stderr. - Zero and negative budgets are ignored with a warning instead of being passed to sglang, where they have no defined meaning. Turning reasoning off is what reasoning_default:off is for. - A load-time warning when thinking_budget is set but enable_strict_thinking is not in engine_args, since sglang then ignores the budget silently. - Tests for integral spellings, unset, zero, negative, non-integer and the strict-thinking warning. Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> * docs(sglang): explain reasoning options Document the reasoning budget, strict-thinking requirement, and precedence of request metadata over the model-level default. Also note that the budget has to stay well below max_tokens (otherwise it never triggers and the reply can end up empty), and that POST /models/reload or a backend-only restart does not pick up changed options; LocalAI itself has to be restarted. Assisted-by: Codex:GPT-6 Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> * docs(sglang): clarify configuration reloads Distinguish rereading model configuration from updating a running backend. Keep the full LocalAI restart recommendation for changed reasoning options. Assisted-by: Codex:GPT-6 * sglang backend: only pass require_reasoning when sglang supports it Engine.async_generate() gained the require_reasoning keyword in sglang 0.5.13 and takes no **kwargs. The CPU profile builds v0.5.11 from source and the other profiles only set a >=0.5.11 floor, so passing the keyword unconditionally made every request fail with TypeError. Detect support once at import time, as the file already does for sampling_seed. enable_strict_thinking first appears in sglang 0.5.12; fix the comment. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code] --------- Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com> Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <localai-org-maint-bot@users.noreply.github.com> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
370 lines
16 KiB
Python
370 lines
16 KiB
Python
"""Unit tests for the sglang backend.
|
|
|
|
Helper-level tests run without launching the gRPC server or loading model
|
|
weights — they only exercise the pure-Python helpers on
|
|
``BackendServicer``. They do still require ``sglang`` to be importable
|
|
because ``_apply_engine_args`` validates keys against ``ServerArgs``
|
|
(a dataclass up to sglang 0.5.19, a ``msgspec.Struct`` from 0.5.20 on).
|
|
"""
|
|
import unittest
|
|
|
|
|
|
|
|
def _request(metadata=None):
|
|
"""Minimal stand-in for a PredictOptions request in reasoning tests."""
|
|
from types import SimpleNamespace
|
|
|
|
return SimpleNamespace(Metadata=metadata or {})
|
|
|
|
class TestSglangHelpers(unittest.TestCase):
|
|
"""Tests for the pure helpers on BackendServicer (no gRPC, no engine)."""
|
|
|
|
def _servicer(self):
|
|
import sys
|
|
import os
|
|
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
|
from backend import BackendServicer # noqa: E402
|
|
return BackendServicer()
|
|
|
|
def test_parse_options(self):
|
|
servicer = self._servicer()
|
|
opts = servicer._parse_options([
|
|
"tool_parser:hermes",
|
|
"reasoning_parser:deepseek_r1",
|
|
"invalid_no_colon",
|
|
"key_with_colons:a:b:c",
|
|
])
|
|
self.assertEqual(opts["tool_parser"], "hermes")
|
|
self.assertEqual(opts["reasoning_parser"], "deepseek_r1")
|
|
self.assertEqual(opts["key_with_colons"], "a:b:c")
|
|
self.assertNotIn("invalid_no_colon", opts)
|
|
|
|
def test_apply_engine_args_known_keys(self):
|
|
"""User-supplied JSON merges into the kwargs dict; pre-set typed
|
|
fields stay put when not overridden."""
|
|
import json as _json
|
|
servicer = self._servicer()
|
|
base = {
|
|
"model_path": "facebook/opt-125m",
|
|
"mem_fraction_static": 0.7,
|
|
}
|
|
extras = _json.dumps({
|
|
"trust_remote_code": True,
|
|
"speculative_algorithm": "EAGLE",
|
|
"speculative_num_steps": 1,
|
|
})
|
|
out = servicer._apply_engine_args(base, extras)
|
|
self.assertIs(out, base) # in-place merge — same dict back
|
|
self.assertTrue(out["trust_remote_code"])
|
|
self.assertEqual(out["speculative_algorithm"], "EAGLE")
|
|
self.assertEqual(out["speculative_num_steps"], 1)
|
|
self.assertEqual(out["model_path"], "facebook/opt-125m")
|
|
self.assertEqual(out["mem_fraction_static"], 0.7)
|
|
|
|
def test_apply_engine_args_engine_args_overrides_typed_fields(self):
|
|
"""engine_args wins over previously-set typed kwargs (vLLM precedence)."""
|
|
import json as _json
|
|
servicer = self._servicer()
|
|
base = {"model_path": "facebook/opt-125m", "mem_fraction_static": 0.7}
|
|
out = servicer._apply_engine_args(
|
|
base, _json.dumps({"mem_fraction_static": 0.5}),
|
|
)
|
|
self.assertEqual(out["mem_fraction_static"], 0.5)
|
|
|
|
def test_apply_engine_args_unknown_key_raises(self):
|
|
"""Typo'd key raises ValueError with a close-match suggestion."""
|
|
import json as _json
|
|
servicer = self._servicer()
|
|
base = {"model_path": "facebook/opt-125m"}
|
|
with self.assertRaises(ValueError) as ctx:
|
|
servicer._apply_engine_args(
|
|
base, _json.dumps({"trust_remotecode": True}),
|
|
)
|
|
msg = str(ctx.exception)
|
|
self.assertIn("trust_remotecode", msg)
|
|
self.assertIn("trust_remote_code", msg)
|
|
|
|
def test_apply_engine_args_msgspec_serverargs(self):
|
|
"""sglang >= 0.5.20 exposes ServerArgs as a msgspec.Struct instead of a
|
|
dataclass; the field names then live in ``__struct_fields__``.
|
|
|
|
Pinned with a stand-in so the msgspec path is covered no matter which
|
|
sglang version happens to be installed.
|
|
"""
|
|
import json as _json
|
|
servicer = self._servicer()
|
|
import backend as backend_mod
|
|
|
|
class _StructLikeServerArgs:
|
|
__struct_fields__ = ("model_path", "mem_fraction_static",
|
|
"trust_remote_code")
|
|
|
|
original = backend_mod.ServerArgs
|
|
backend_mod.ServerArgs = _StructLikeServerArgs
|
|
try:
|
|
out = servicer._apply_engine_args(
|
|
{}, _json.dumps({"trust_remote_code": True}),
|
|
)
|
|
self.assertTrue(out["trust_remote_code"])
|
|
with self.assertRaises(ValueError) as ctx:
|
|
servicer._apply_engine_args(
|
|
{}, _json.dumps({"mem_fraction_statik": 0.7}),
|
|
)
|
|
msg = str(ctx.exception)
|
|
self.assertIn("mem_fraction_statik", msg)
|
|
self.assertIn("mem_fraction_static", msg)
|
|
finally:
|
|
backend_mod.ServerArgs = original
|
|
|
|
def test_apply_engine_args_empty_passthrough(self):
|
|
"""Empty / None engine_args returns the kwargs dict untouched."""
|
|
servicer = self._servicer()
|
|
base = {"model_path": "facebook/opt-125m"}
|
|
self.assertIs(servicer._apply_engine_args(base, ""), base)
|
|
self.assertIs(servicer._apply_engine_args(base, None), base)
|
|
|
|
def test_apply_engine_args_invalid_json_raises(self):
|
|
servicer = self._servicer()
|
|
with self.assertRaises(ValueError) as ctx:
|
|
servicer._apply_engine_args({}, "not-json")
|
|
self.assertIn("not valid JSON", str(ctx.exception))
|
|
|
|
def test_apply_engine_args_non_object_raises(self):
|
|
servicer = self._servicer()
|
|
with self.assertRaises(ValueError) as ctx:
|
|
servicer._apply_engine_args({}, "[1,2,3]")
|
|
self.assertIn("must be a JSON object", str(ctx.exception))
|
|
|
|
def test_build_prompt_forwards_enable_thinking(self):
|
|
from types import SimpleNamespace
|
|
|
|
class Tok:
|
|
def __init__(self):
|
|
self.kwargs = None
|
|
|
|
def apply_chat_template(self, messages, **kwargs):
|
|
self.kwargs = kwargs
|
|
return "PROMPT"
|
|
|
|
def kwargs_for(metadata):
|
|
servicer = self._servicer()
|
|
tok = Tok()
|
|
servicer.tokenizer = tok
|
|
msg = SimpleNamespace(
|
|
role="user", content="hi", name="",
|
|
tool_call_id="", reasoning_content="", tool_calls="",
|
|
)
|
|
req = SimpleNamespace(
|
|
Prompt="", UseTokenizerTemplate=True,
|
|
Messages=[msg], Tools="", Metadata=metadata,
|
|
)
|
|
self.assertEqual(servicer._build_prompt(req), "PROMPT")
|
|
return tok.kwargs
|
|
|
|
self.assertIs(kwargs_for({"enable_thinking": "true"})["enable_thinking"], True)
|
|
# "false" used to be dropped, so Qwen3 kept thinking on
|
|
self.assertIs(kwargs_for({"enable_thinking": "false"})["enable_thinking"], False)
|
|
self.assertNotIn("enable_thinking", kwargs_for({}))
|
|
self.assertIs(kwargs_for({"enable_thinking": "FALSE"})["enable_thinking"], False)
|
|
|
|
def test_reasoning_parser_forced_when_template_prefills_think_tag(self):
|
|
"""Qwen3's template puts ``<think>`` in the prompt, so the completion
|
|
never contains it. Without force_reasoning the detector treats the whole
|
|
completion as normal text and reasoning_content stays empty."""
|
|
servicer = self._servicer()
|
|
servicer.reasoning_parser_name = "qwen3"
|
|
|
|
# What the model actually emits when the prompt ends in "<think>".
|
|
completion = "adding two and two</think>4"
|
|
|
|
forced, require_reasoning = servicer._new_reasoning_parser(
|
|
False, prompt="user: hi\n<think>\n"
|
|
)
|
|
self.assertTrue(require_reasoning)
|
|
reasoning, content = forced.parse_non_stream(completion)
|
|
self.assertEqual(reasoning, "adding two and two")
|
|
self.assertEqual(content, "4")
|
|
|
|
# No prefilled tag in the prompt: detector default, unchanged behaviour.
|
|
unforced, require_reasoning = servicer._new_reasoning_parser(False, prompt="user: hi\n")
|
|
self.assertFalse(require_reasoning)
|
|
reasoning, content = unforced.parse_non_stream(completion)
|
|
self.assertFalse(reasoning)
|
|
self.assertEqual(content, completion)
|
|
|
|
def test_reasoning_parser_not_forced_when_thinking_is_off(self):
|
|
"""Thinking off means no ``<think>`` in the prompt either, so the answer
|
|
must not be swallowed into reasoning_content."""
|
|
servicer = self._servicer()
|
|
servicer.reasoning_parser_name = "qwen3"
|
|
|
|
parser, require_reasoning = servicer._new_reasoning_parser(False, prompt="user: primes?\n")
|
|
self.assertFalse(require_reasoning)
|
|
reasoning, content = parser.parse_non_stream("2,3,5,7,11")
|
|
self.assertFalse(reasoning)
|
|
self.assertEqual(content, "2,3,5,7,11")
|
|
|
|
def test_grammar_constrained_output_is_not_forced_into_reasoning(self):
|
|
"""Structured decoding applies from the first token, so the model cannot
|
|
emit the closing tag even though the template opened the block. The whole
|
|
completion is schema output and must stay in content."""
|
|
servicer = self._servicer()
|
|
servicer.reasoning_parser_name = "qwen3"
|
|
|
|
schema_out = '{"findings": [{"line": 42, "issue": "off-by-one"}]}'
|
|
parser, require_reasoning = servicer._new_reasoning_parser(
|
|
False, prompt="audit this\n<think>\n", grammar_constrained=True,
|
|
)
|
|
self.assertFalse(require_reasoning)
|
|
reasoning, content = parser.parse_non_stream(schema_out)
|
|
self.assertFalse(reasoning)
|
|
self.assertEqual(content, schema_out)
|
|
|
|
def test_reasoning_parser_absent_without_configured_parser(self):
|
|
servicer = self._servicer()
|
|
servicer.reasoning_parser_name = None
|
|
parser, require_reasoning = servicer._new_reasoning_parser(False, prompt="<think>")
|
|
self.assertIsNone(parser)
|
|
self.assertFalse(require_reasoning)
|
|
|
|
def test_reasoning_default_off_applies_when_request_is_silent(self):
|
|
"""A model configured with reasoning_default:off must render with
|
|
thinking disabled even when the request carries no enable_thinking -
|
|
that is the whole point: `parameters: reasoning_effort:` never
|
|
reaches this backend, so without this the config lies about the
|
|
default."""
|
|
servicer = self._servicer()
|
|
servicer.reasoning_default = "off"
|
|
self.assertIs(servicer._thinking_default(_request(metadata={})), False)
|
|
|
|
def test_request_metadata_overrides_reasoning_default(self):
|
|
"""A per-request value always wins over the model-level default -
|
|
in both directions."""
|
|
servicer = self._servicer()
|
|
servicer.reasoning_default = "off"
|
|
self.assertIs(
|
|
servicer._thinking_default(_request(metadata={"enable_thinking": "true"})),
|
|
True,
|
|
)
|
|
servicer.reasoning_default = "on"
|
|
self.assertIs(
|
|
servicer._thinking_default(_request(metadata={"enable_thinking": "false"})),
|
|
False,
|
|
)
|
|
|
|
def test_no_reasoning_default_leaves_template_untouched(self):
|
|
"""Unconfigured must stay unconfigured: returning None means the
|
|
backend adds no enable_thinking kwarg at all, so the template keeps
|
|
whatever default it ships with."""
|
|
servicer = self._servicer()
|
|
self.assertIsNone(servicer._thinking_default(_request(metadata={})))
|
|
|
|
def test_thinking_budget_added_to_sampling_params_as_custom_params(self):
|
|
"""The model-level thinking_budget option (set from LoadModel's
|
|
Options, mirroring tool_parser/reasoning_parser) must ride along as
|
|
sampling_params['custom_params']['thinking_budget'] on every
|
|
request — that's the only field sglang's --enable-strict-thinking
|
|
grammar backend reads to bound the reasoning length."""
|
|
from types import SimpleNamespace
|
|
|
|
servicer = self._servicer()
|
|
servicer.thinking_budget = 512
|
|
request = SimpleNamespace(
|
|
Temperature=0.7, N=0, PresencePenalty=0, FrequencyPenalty=0,
|
|
RepetitionPenalty=0, TopP=0, TopK=0, MinP=0, Seed=0,
|
|
StopPrompts=[], StopTokenIds=[], IgnoreEOS=False, Tokens=0,
|
|
MinTokens=0, SkipSpecialTokens=False, Grammar="",
|
|
)
|
|
params = servicer._build_sampling_params(request)
|
|
self.assertEqual(params["custom_params"], {"thinking_budget": 512})
|
|
|
|
def test_no_thinking_budget_means_no_custom_params_key(self):
|
|
"""Unconfigured is unconfigured: no thinking_budget option must not
|
|
add an empty/None custom_params that could clobber a sglang-side
|
|
--preferred-sampling-params default (see sglang#40634)."""
|
|
from types import SimpleNamespace
|
|
|
|
servicer = self._servicer()
|
|
request = SimpleNamespace(
|
|
Temperature=0.7, N=0, PresencePenalty=0, FrequencyPenalty=0,
|
|
RepetitionPenalty=0, TopP=0, TopK=0, MinP=0, Seed=0,
|
|
StopPrompts=[], StopTokenIds=[], IgnoreEOS=False, Tokens=0,
|
|
MinTokens=0, SkipSpecialTokens=False, Grammar="",
|
|
)
|
|
params = servicer._build_sampling_params(request)
|
|
self.assertNotIn("custom_params", params)
|
|
|
|
def test_thinking_budget_accepts_integral_spellings(self):
|
|
"""YAML options arrive as strings; "512" and "512.0" both mean 512."""
|
|
servicer = self._servicer()
|
|
self.assertEqual(servicer._parse_thinking_budget("512"), 512)
|
|
self.assertEqual(servicer._parse_thinking_budget("512.0"), 512)
|
|
self.assertEqual(servicer._parse_thinking_budget(" 64 "), 64)
|
|
self.assertEqual(servicer._parse_thinking_budget(256), 256)
|
|
|
|
def test_thinking_budget_unset_is_none(self):
|
|
servicer = self._servicer()
|
|
self.assertIsNone(servicer._parse_thinking_budget(None))
|
|
self.assertIsNone(servicer._parse_thinking_budget(""))
|
|
|
|
def test_thinking_budget_zero_and_negative_are_ignored(self):
|
|
"""No defined meaning in sglang -- ignored, not passed through."""
|
|
servicer = self._servicer()
|
|
self.assertIsNone(servicer._parse_thinking_budget("0"))
|
|
self.assertIsNone(servicer._parse_thinking_budget("-100"))
|
|
|
|
def test_thinking_budget_non_integer_does_not_raise(self):
|
|
"""A bad value must not crash LoadModel for the whole model."""
|
|
servicer = self._servicer()
|
|
self.assertIsNone(servicer._parse_thinking_budget("12.5"))
|
|
self.assertIsNone(servicer._parse_thinking_budget("lots"))
|
|
|
|
def test_warns_when_budget_set_without_strict_thinking(self):
|
|
servicer = self._servicer()
|
|
self.assertIn(
|
|
"enable_strict_thinking",
|
|
servicer._strict_thinking_warning(512, {"model_path": "x"}),
|
|
)
|
|
self.assertIsNone(servicer._strict_thinking_warning(512, {"enable_strict_thinking": True}))
|
|
self.assertIsNone(servicer._strict_thinking_warning(None, {}))
|
|
|
|
def test_explicit_zero_temperature_and_seed_are_preserved(self):
|
|
"""Temperature=0 is greedy decoding and 0 is a valid seed — neither is
|
|
an unset value. A dropped seed turns a reproducible request random."""
|
|
from types import SimpleNamespace
|
|
|
|
servicer = self._servicer()
|
|
import sys as _sys
|
|
_SEED_KEY_FOR_TEST = _sys.modules["backend"]._SEED_KEY
|
|
request = SimpleNamespace(
|
|
Temperature=0,
|
|
N=0,
|
|
PresencePenalty=0,
|
|
FrequencyPenalty=0,
|
|
RepetitionPenalty=0,
|
|
TopP=0,
|
|
TopK=0,
|
|
MinP=0,
|
|
Seed=0,
|
|
StopPrompts=[],
|
|
StopTokenIds=[],
|
|
IgnoreEOS=False,
|
|
Tokens=0,
|
|
MinTokens=0,
|
|
SkipSpecialTokens=False,
|
|
Grammar="",
|
|
)
|
|
|
|
params = servicer._build_sampling_params(request)
|
|
self.assertEqual(params["temperature"], 0)
|
|
self.assertEqual(params[_SEED_KEY_FOR_TEST], 0)
|
|
# Other protobuf-default scalar fields must remain filtered. top_k=0 in
|
|
# particular is not a value sglang accepts (-1 disables it), so it must
|
|
# keep falling through to the engine default.
|
|
self.assertNotIn("top_p", params)
|
|
self.assertNotIn("top_k", params)
|
|
|
|
|
|
if __name__ == "__main__":
|
|
unittest.main()
|