Files
LocalAI/backend/python/sglang/test.py
T
61f4f67b75 sglang backend: pass through thinking_budget + require_reasoning (#12193)
* sglang backend: pass through thinking_budget + require_reasoning

sglang's raw Engine.async_generate() API (which this backend calls
directly, bypassing sglang's own OpenAI server) supports a precise,
tokenizer-derived reasoning-length budget via
sampling_params["custom_params"]["thinking_budget"] plus
require_reasoning=True, gated behind --enable-strict-thinking. Neither
was reachable through LocalAI: this backend built sampling_params only
from a fixed field mapping (temperature, top_p, ...) with no custom_params
key, and never passed require_reasoning to async_generate at all.

- LoadModel now reads a model-level "thinking_budget" option (same
  mechanism as the existing tool_parser/reasoning_parser options), and
  _build_sampling_params adds it as custom_params.thinking_budget on
  every request when configured.
- _new_reasoning_parser already derives, from the rendered prompt, whether
  the model's chat template pre-opened a reasoning block (Qwen3-style
  templates append <think> to the prompt instead of letting the model
  emit it) -- the same signal sglang's own OpenAI server computes from
  per-template config to decide require_reasoning. This backend has no
  template manager, so it now returns that signal too and _predict
  forwards it to async_generate(require_reasoning=...).

Verified against production (NVFP4, sm_121, Qwen3.6-35B-A3B) via a raw
Engine.async_generate() call bypassing this backend: 301 reasoning
tokens against a 300-token budget, clean completion, ~27s. Not yet
verified through this backend's own gRPC path end-to-end (no local
CUDA/sglang environment available here) -- existing + new unit tests in
test.py cover the pure-Python merge/passthrough logic only.

Scope note: require_reasoning is derived only from the existing
prompt-suffix heuristic, not sglang's full per-template
_get_reasoning_from_request decision tree (minimax-m3/hunyuan special
cases etc.) -- this backend has no template manager to evaluate that
tree against, and the prompt-suffix check is the one heuristic already
validated in this file (test_reasoning_parser_forced_when_template_prefills_think_tag).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: honour a model-level reasoning_default

A model YAML can already carry "parameters: reasoning_effort:", but that
value only reaches this backend when a *caller* sets it per request (the Go
side turns it into Metadata["enable_thinking"]). As a model-level default it
is silently dropped: a config reading "reasoning_effort: none" still produces
full reasoning on every request, so the config says one thing and the model
does another.

That gap is expensive in practice. On a self-hosted Qwen3.6-35B-A3B the
reasoning phase consumed the entire max_tokens budget before any content was
produced - 90% of code completions came back empty at max_tokens=768, and the
server log filled with "backend produced only reasoning, retrying". The
config looked like reasoning was off the whole time.

This adds "reasoning_default:off" (or ":on") on the same model-level
options: mechanism as thinking_budget. A per-request value always wins; the
default only fills in when the request is silent.

Measured on the stack above (sglang 0.5.20, NVFP4, GB10/sm_121) after
applying it:
  default (nothing set)          -> 0 chars reasoning, 27 tokens
  "reasoning_effort": "none"     -> 0 chars reasoning, 27 tokens
  metadata enable_thinking=true  -> capped at the 512-token thinking_budget,
                                    541 tokens total, finish_reason stop

Tests: three cases added to backend/python/sglang/test.py covering the
default, per-request override in both directions, and the unconfigured case
(which must leave the template untouched).

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* sglang backend: validate thinking_budget instead of crashing LoadModel

Addresses the review on this PR:

- `int(thinking_budget)` raised on values like "5000.0" or "abc" and took
  LoadModel down. The option is now parsed by _parse_thinking_budget():
  integral numbers in any spelling are accepted, anything else is ignored
  with a warning on stderr.
- Zero and negative budgets are ignored with a warning instead of being
  passed to sglang, where they have no defined meaning. Turning reasoning
  off is what reasoning_default:off is for.
- A load-time warning when thinking_budget is set but enable_strict_thinking
  is not in engine_args, since sglang then ignores the budget silently.
- Tests for integral spellings, unset, zero, negative, non-integer and the
  strict-thinking warning.

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): explain reasoning options

Document the reasoning budget, strict-thinking requirement, and
precedence of request metadata over the model-level default.

Also note that the budget has to stay well below max_tokens (otherwise
it never triggers and the reply can end up empty), and that
POST /models/reload or a backend-only restart does not pick up changed
options; LocalAI itself has to be restarted.

Assisted-by: Codex:GPT-6
Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>

* docs(sglang): clarify configuration reloads

Distinguish rereading model configuration from updating a running backend. Keep the full LocalAI restart recommendation for changed reasoning options.

Assisted-by: Codex:GPT-6

* sglang backend: only pass require_reasoning when sglang supports it

Engine.async_generate() gained the require_reasoning keyword in sglang
0.5.13 and takes no **kwargs. The CPU profile builds v0.5.11 from source
and the other profiles only set a >=0.5.11 floor, so passing the keyword
unconditionally made every request fail with TypeError. Detect support
once at import time, as the file already does for sampling_seed.

enable_strict_thinking first appears in sglang 0.5.12; fix the comment.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5-5 [Claude Code]

---------

Signed-off-by: pos-ei-don <1822533+pos-ei-don@users.noreply.github.com>
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <localai-org-maint-bot@users.noreply.github.com>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-28 04:49:37 +02:00

370 lines
16 KiB
Python

"""Unit tests for the sglang backend.
Helper-level tests run without launching the gRPC server or loading model
weights — they only exercise the pure-Python helpers on
``BackendServicer``. They do still require ``sglang`` to be importable
because ``_apply_engine_args`` validates keys against ``ServerArgs``
(a dataclass up to sglang 0.5.19, a ``msgspec.Struct`` from 0.5.20 on).
"""
import unittest
def _request(metadata=None):
"""Minimal stand-in for a PredictOptions request in reasoning tests."""
from types import SimpleNamespace
return SimpleNamespace(Metadata=metadata or {})
class TestSglangHelpers(unittest.TestCase):
"""Tests for the pure helpers on BackendServicer (no gRPC, no engine)."""
def _servicer(self):
import sys
import os
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from backend import BackendServicer # noqa: E402
return BackendServicer()
def test_parse_options(self):
servicer = self._servicer()
opts = servicer._parse_options([
"tool_parser:hermes",
"reasoning_parser:deepseek_r1",
"invalid_no_colon",
"key_with_colons:a:b:c",
])
self.assertEqual(opts["tool_parser"], "hermes")
self.assertEqual(opts["reasoning_parser"], "deepseek_r1")
self.assertEqual(opts["key_with_colons"], "a:b:c")
self.assertNotIn("invalid_no_colon", opts)
def test_apply_engine_args_known_keys(self):
"""User-supplied JSON merges into the kwargs dict; pre-set typed
fields stay put when not overridden."""
import json as _json
servicer = self._servicer()
base = {
"model_path": "facebook/opt-125m",
"mem_fraction_static": 0.7,
}
extras = _json.dumps({
"trust_remote_code": True,
"speculative_algorithm": "EAGLE",
"speculative_num_steps": 1,
})
out = servicer._apply_engine_args(base, extras)
self.assertIs(out, base) # in-place merge — same dict back
self.assertTrue(out["trust_remote_code"])
self.assertEqual(out["speculative_algorithm"], "EAGLE")
self.assertEqual(out["speculative_num_steps"], 1)
self.assertEqual(out["model_path"], "facebook/opt-125m")
self.assertEqual(out["mem_fraction_static"], 0.7)
def test_apply_engine_args_engine_args_overrides_typed_fields(self):
"""engine_args wins over previously-set typed kwargs (vLLM precedence)."""
import json as _json
servicer = self._servicer()
base = {"model_path": "facebook/opt-125m", "mem_fraction_static": 0.7}
out = servicer._apply_engine_args(
base, _json.dumps({"mem_fraction_static": 0.5}),
)
self.assertEqual(out["mem_fraction_static"], 0.5)
def test_apply_engine_args_unknown_key_raises(self):
"""Typo'd key raises ValueError with a close-match suggestion."""
import json as _json
servicer = self._servicer()
base = {"model_path": "facebook/opt-125m"}
with self.assertRaises(ValueError) as ctx:
servicer._apply_engine_args(
base, _json.dumps({"trust_remotecode": True}),
)
msg = str(ctx.exception)
self.assertIn("trust_remotecode", msg)
self.assertIn("trust_remote_code", msg)
def test_apply_engine_args_msgspec_serverargs(self):
"""sglang >= 0.5.20 exposes ServerArgs as a msgspec.Struct instead of a
dataclass; the field names then live in ``__struct_fields__``.
Pinned with a stand-in so the msgspec path is covered no matter which
sglang version happens to be installed.
"""
import json as _json
servicer = self._servicer()
import backend as backend_mod
class _StructLikeServerArgs:
__struct_fields__ = ("model_path", "mem_fraction_static",
"trust_remote_code")
original = backend_mod.ServerArgs
backend_mod.ServerArgs = _StructLikeServerArgs
try:
out = servicer._apply_engine_args(
{}, _json.dumps({"trust_remote_code": True}),
)
self.assertTrue(out["trust_remote_code"])
with self.assertRaises(ValueError) as ctx:
servicer._apply_engine_args(
{}, _json.dumps({"mem_fraction_statik": 0.7}),
)
msg = str(ctx.exception)
self.assertIn("mem_fraction_statik", msg)
self.assertIn("mem_fraction_static", msg)
finally:
backend_mod.ServerArgs = original
def test_apply_engine_args_empty_passthrough(self):
"""Empty / None engine_args returns the kwargs dict untouched."""
servicer = self._servicer()
base = {"model_path": "facebook/opt-125m"}
self.assertIs(servicer._apply_engine_args(base, ""), base)
self.assertIs(servicer._apply_engine_args(base, None), base)
def test_apply_engine_args_invalid_json_raises(self):
servicer = self._servicer()
with self.assertRaises(ValueError) as ctx:
servicer._apply_engine_args({}, "not-json")
self.assertIn("not valid JSON", str(ctx.exception))
def test_apply_engine_args_non_object_raises(self):
servicer = self._servicer()
with self.assertRaises(ValueError) as ctx:
servicer._apply_engine_args({}, "[1,2,3]")
self.assertIn("must be a JSON object", str(ctx.exception))
def test_build_prompt_forwards_enable_thinking(self):
from types import SimpleNamespace
class Tok:
def __init__(self):
self.kwargs = None
def apply_chat_template(self, messages, **kwargs):
self.kwargs = kwargs
return "PROMPT"
def kwargs_for(metadata):
servicer = self._servicer()
tok = Tok()
servicer.tokenizer = tok
msg = SimpleNamespace(
role="user", content="hi", name="",
tool_call_id="", reasoning_content="", tool_calls="",
)
req = SimpleNamespace(
Prompt="", UseTokenizerTemplate=True,
Messages=[msg], Tools="", Metadata=metadata,
)
self.assertEqual(servicer._build_prompt(req), "PROMPT")
return tok.kwargs
self.assertIs(kwargs_for({"enable_thinking": "true"})["enable_thinking"], True)
# "false" used to be dropped, so Qwen3 kept thinking on
self.assertIs(kwargs_for({"enable_thinking": "false"})["enable_thinking"], False)
self.assertNotIn("enable_thinking", kwargs_for({}))
self.assertIs(kwargs_for({"enable_thinking": "FALSE"})["enable_thinking"], False)
def test_reasoning_parser_forced_when_template_prefills_think_tag(self):
"""Qwen3's template puts ``<think>`` in the prompt, so the completion
never contains it. Without force_reasoning the detector treats the whole
completion as normal text and reasoning_content stays empty."""
servicer = self._servicer()
servicer.reasoning_parser_name = "qwen3"
# What the model actually emits when the prompt ends in "<think>".
completion = "adding two and two</think>4"
forced, require_reasoning = servicer._new_reasoning_parser(
False, prompt="user: hi\n<think>\n"
)
self.assertTrue(require_reasoning)
reasoning, content = forced.parse_non_stream(completion)
self.assertEqual(reasoning, "adding two and two")
self.assertEqual(content, "4")
# No prefilled tag in the prompt: detector default, unchanged behaviour.
unforced, require_reasoning = servicer._new_reasoning_parser(False, prompt="user: hi\n")
self.assertFalse(require_reasoning)
reasoning, content = unforced.parse_non_stream(completion)
self.assertFalse(reasoning)
self.assertEqual(content, completion)
def test_reasoning_parser_not_forced_when_thinking_is_off(self):
"""Thinking off means no ``<think>`` in the prompt either, so the answer
must not be swallowed into reasoning_content."""
servicer = self._servicer()
servicer.reasoning_parser_name = "qwen3"
parser, require_reasoning = servicer._new_reasoning_parser(False, prompt="user: primes?\n")
self.assertFalse(require_reasoning)
reasoning, content = parser.parse_non_stream("2,3,5,7,11")
self.assertFalse(reasoning)
self.assertEqual(content, "2,3,5,7,11")
def test_grammar_constrained_output_is_not_forced_into_reasoning(self):
"""Structured decoding applies from the first token, so the model cannot
emit the closing tag even though the template opened the block. The whole
completion is schema output and must stay in content."""
servicer = self._servicer()
servicer.reasoning_parser_name = "qwen3"
schema_out = '{"findings": [{"line": 42, "issue": "off-by-one"}]}'
parser, require_reasoning = servicer._new_reasoning_parser(
False, prompt="audit this\n<think>\n", grammar_constrained=True,
)
self.assertFalse(require_reasoning)
reasoning, content = parser.parse_non_stream(schema_out)
self.assertFalse(reasoning)
self.assertEqual(content, schema_out)
def test_reasoning_parser_absent_without_configured_parser(self):
servicer = self._servicer()
servicer.reasoning_parser_name = None
parser, require_reasoning = servicer._new_reasoning_parser(False, prompt="<think>")
self.assertIsNone(parser)
self.assertFalse(require_reasoning)
def test_reasoning_default_off_applies_when_request_is_silent(self):
"""A model configured with reasoning_default:off must render with
thinking disabled even when the request carries no enable_thinking -
that is the whole point: `parameters: reasoning_effort:` never
reaches this backend, so without this the config lies about the
default."""
servicer = self._servicer()
servicer.reasoning_default = "off"
self.assertIs(servicer._thinking_default(_request(metadata={})), False)
def test_request_metadata_overrides_reasoning_default(self):
"""A per-request value always wins over the model-level default -
in both directions."""
servicer = self._servicer()
servicer.reasoning_default = "off"
self.assertIs(
servicer._thinking_default(_request(metadata={"enable_thinking": "true"})),
True,
)
servicer.reasoning_default = "on"
self.assertIs(
servicer._thinking_default(_request(metadata={"enable_thinking": "false"})),
False,
)
def test_no_reasoning_default_leaves_template_untouched(self):
"""Unconfigured must stay unconfigured: returning None means the
backend adds no enable_thinking kwarg at all, so the template keeps
whatever default it ships with."""
servicer = self._servicer()
self.assertIsNone(servicer._thinking_default(_request(metadata={})))
def test_thinking_budget_added_to_sampling_params_as_custom_params(self):
"""The model-level thinking_budget option (set from LoadModel's
Options, mirroring tool_parser/reasoning_parser) must ride along as
sampling_params['custom_params']['thinking_budget'] on every
request — that's the only field sglang's --enable-strict-thinking
grammar backend reads to bound the reasoning length."""
from types import SimpleNamespace
servicer = self._servicer()
servicer.thinking_budget = 512
request = SimpleNamespace(
Temperature=0.7, N=0, PresencePenalty=0, FrequencyPenalty=0,
RepetitionPenalty=0, TopP=0, TopK=0, MinP=0, Seed=0,
StopPrompts=[], StopTokenIds=[], IgnoreEOS=False, Tokens=0,
MinTokens=0, SkipSpecialTokens=False, Grammar="",
)
params = servicer._build_sampling_params(request)
self.assertEqual(params["custom_params"], {"thinking_budget": 512})
def test_no_thinking_budget_means_no_custom_params_key(self):
"""Unconfigured is unconfigured: no thinking_budget option must not
add an empty/None custom_params that could clobber a sglang-side
--preferred-sampling-params default (see sglang#40634)."""
from types import SimpleNamespace
servicer = self._servicer()
request = SimpleNamespace(
Temperature=0.7, N=0, PresencePenalty=0, FrequencyPenalty=0,
RepetitionPenalty=0, TopP=0, TopK=0, MinP=0, Seed=0,
StopPrompts=[], StopTokenIds=[], IgnoreEOS=False, Tokens=0,
MinTokens=0, SkipSpecialTokens=False, Grammar="",
)
params = servicer._build_sampling_params(request)
self.assertNotIn("custom_params", params)
def test_thinking_budget_accepts_integral_spellings(self):
"""YAML options arrive as strings; "512" and "512.0" both mean 512."""
servicer = self._servicer()
self.assertEqual(servicer._parse_thinking_budget("512"), 512)
self.assertEqual(servicer._parse_thinking_budget("512.0"), 512)
self.assertEqual(servicer._parse_thinking_budget(" 64 "), 64)
self.assertEqual(servicer._parse_thinking_budget(256), 256)
def test_thinking_budget_unset_is_none(self):
servicer = self._servicer()
self.assertIsNone(servicer._parse_thinking_budget(None))
self.assertIsNone(servicer._parse_thinking_budget(""))
def test_thinking_budget_zero_and_negative_are_ignored(self):
"""No defined meaning in sglang -- ignored, not passed through."""
servicer = self._servicer()
self.assertIsNone(servicer._parse_thinking_budget("0"))
self.assertIsNone(servicer._parse_thinking_budget("-100"))
def test_thinking_budget_non_integer_does_not_raise(self):
"""A bad value must not crash LoadModel for the whole model."""
servicer = self._servicer()
self.assertIsNone(servicer._parse_thinking_budget("12.5"))
self.assertIsNone(servicer._parse_thinking_budget("lots"))
def test_warns_when_budget_set_without_strict_thinking(self):
servicer = self._servicer()
self.assertIn(
"enable_strict_thinking",
servicer._strict_thinking_warning(512, {"model_path": "x"}),
)
self.assertIsNone(servicer._strict_thinking_warning(512, {"enable_strict_thinking": True}))
self.assertIsNone(servicer._strict_thinking_warning(None, {}))
def test_explicit_zero_temperature_and_seed_are_preserved(self):
"""Temperature=0 is greedy decoding and 0 is a valid seed — neither is
an unset value. A dropped seed turns a reproducible request random."""
from types import SimpleNamespace
servicer = self._servicer()
import sys as _sys
_SEED_KEY_FOR_TEST = _sys.modules["backend"]._SEED_KEY
request = SimpleNamespace(
Temperature=0,
N=0,
PresencePenalty=0,
FrequencyPenalty=0,
RepetitionPenalty=0,
TopP=0,
TopK=0,
MinP=0,
Seed=0,
StopPrompts=[],
StopTokenIds=[],
IgnoreEOS=False,
Tokens=0,
MinTokens=0,
SkipSpecialTokens=False,
Grammar="",
)
params = servicer._build_sampling_params(request)
self.assertEqual(params["temperature"], 0)
self.assertEqual(params[_SEED_KEY_FOR_TEST], 0)
# Other protobuf-default scalar fields must remain filtered. top_k=0 in
# particular is not a value sglang accepts (-1 disables it), so it must
# keep falling through to the engine default.
self.assertNotIn("top_p", params)
self.assertNotIn("top_k", params)
if __name__ == "__main__":
unittest.main()