mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-05 04:24:39 -04:00
fix(llama-cpp): do not stream the error text as content on pre-stream failures (#12425)
When the first result of a streamed request is an error (for example a prompt that exceeds the context), PredictStream wrote the error message as a Reply and only then returned the error status. LocalAI treated that Reply as the first token: it sent the assistant role chunk and the error text as `content` on an HTTP 200 stream. Because a chunk had already been written, the pre-stream HTTP error path from #12204 never triggered, so streaming clients still got a 200 with the error as model output, while the same request without streaming correctly returns a 400. Return the error only as the gRPC status. The e2e backend suite gets a `context_overflow` capability (enabled for llama-cpp) that streams an over-long prompt and asserts an error status with no content. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Stefan Walcz <stefan.walcz@walcz.de>
This commit is contained in:
1 parent
1056c62f4c
commit
c3bea567fe
4 files changed
+44
-4
No files matched your search
@@ -2176,10 +2176,10 @@ public:
|
||||
// connection is closed
|
||||
return grpc::Status(grpc::StatusCode::CANCELLED, "Request cancelled by client");
|
||||
} else if (first_result->is_error()) {
|
||||
// Return the error only as the status. Writing it as a Reply first
|
||||
// made it the first content chunk: LocalAI streamed the error text
|
||||
// as assistant output on an HTTP 200 instead of failing the request.
|
||||
json error_json = first_result->to_json();
|
||||
backend::Reply reply;
|
||||
reply.set_message(error_json.value("message", ""));
|
||||
writer->Write(reply);
|
||||
return grpc::Status(grpc::StatusCode::INTERNAL, error_json.value("message", "Error occurred"));
|
||||
}
|
||||
|
||||
|
||||
Reference in new issue
Block a user