ResourceExhausted is retried and trips the target, which is right for a
rate limit or an out-of-memory backend. A payload over the gRPC message
cap is ResourceExhausted too, but every target rejects it the same way,
so it tripped the whole chain. Classify it as a request error.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The non-OpenAI endpoints (depth, detection, face_*, voice_*, images,
video, 3d) map a backend's gRPC Unimplemented to an echo 501 without the
gRPC status. The retry loop did not see a capability gap, and IsRetryable
is false for 501, so the client got 501 and the next target was never
tried. Treat a returned or written 501 as a capability gap.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Treating ResourceExhausted as a capability gap skipped the target
without counting a failure, so a target that stays rate limited or out
of memory kept its traffic. It is now an ordinary retryable failure:
the request moves to the next target and the exhausted one trips.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Rerank no longer sends top_n 0, which the upstream rejects. A
mid-stream upstream error frame now fails the call instead of ending
it as a short success. Temperature 0 is forwarded. An upstream 429
becomes ResourceExhausted, which failover skips like Unimplemented.
A localai-proxy config sends its own name upstream when upstream_model
is unset, and a chat proxy defaults to the tokenizer template so chat
reaches the upstream as messages.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>