mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-29 17:44:30 -04:00
The proxy sent every chat and completion request with context.Background, and the gRPC server gave rich backends no context at all. When a client disconnected, or failover gave up on the target, the upstream kept generating to the end, which costs tokens on a paid or shared upstream. A silent upstream held the backend forever. Add the optional AIModelRichContext interface. The gRPC server prefers it and passes the call's context, like the Score and Rerank extensions. The proxy implements it, so the upstream request ends with the gRPC call. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code]