mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-05 04:24:39 -04:00
When the first result of a streamed request is an error (for example a prompt that exceeds the context), PredictStream wrote the error message as a Reply and only then returned the error status. LocalAI treated that Reply as the first token: it sent the assistant role chunk and the error text as `content` on an HTTP 200 stream. Because a chunk had already been written, the pre-stream HTTP error path from #12204 never triggered, so streaming clients still got a 200 with the error as model output, while the same request without streaming correctly returns a 400. Return the error only as the gRPC status. The e2e backend suite gets a `context_overflow` capability (enabled for llama-cpp) that streams an over-long prompt and asserts an error status with no content. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>