fix(llama-cpp): do not stream the error text as content on pre-stream failures (#12425)

When the first result of a streamed request is an error (for example a
prompt that exceeds the context), PredictStream wrote the error message
as a Reply and only then returned the error status. LocalAI treated that
Reply as the first token: it sent the assistant role chunk and the error
text as `content` on an HTTP 200 stream. Because a chunk had already been
written, the pre-stream HTTP error path from #12204 never triggered, so
streaming clients still got a 200 with the error as model output, while
the same request without streaming correctly returns a 400.

Return the error only as the gRPC status. The e2e backend suite gets a
`context_overflow` capability (enabled for llama-cpp) that streams an
over-long prompt and asserts an error status with no content.

Assisted-by: Claude:claude-opus-5-5

Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
This commit is contained in:
Stefan Walcz authored and GitHub committed 2026-10-02 09:23:57 +02:00
1 parent 1056c62f4c
commit c3bea567fe
4 files changed
+44 -4

No files matched your search

+3 -3
View File
@@ -2176,10 +2176,10 @@ public:
// connection is closed
return grpc::Status(grpc::StatusCode::CANCELLED, "Request cancelled by client");
} else if (first_result->is_error()) {
// Return the error only as the status. Writing it as a Reply first
// made it the first content chunk: LocalAI streamed the error text
// as assistant output on an HTTP 200 instead of failing the request.
json error_json = first_result->to_json();
backend::Reply reply;
reply.set_message(error_json.value("message", ""));
writer->Write(reply);
return grpc::Status(grpc::StatusCode::INTERNAL, error_json.value("message", "Error occurred"));
}