mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-28 17:15:02 -04:00
fix: return correct HTTP status codes for saturation and no-nodes-available (#12113)
* feat: return 429 when backends are saturated When backends are at capacity (per-model max_concurrent or the process-wide --max-concurrent-backend-requests ceiling), the response was 503. The OpenAI SDK, litellm, and most agent harnesses key on 429 for rate-limit backoff and treat 503 as a hard error. Both saturation paths now return 429 with the existing Retry-After header and type: "rate_limit_error" in the JSON body. The per-model admission middleware keeps admission_rejected as the code field so existing alerts that match on it still fire. Non-saturation 503s are unchanged: model cold-loading (with progress body), model-load failure cooldown, PII detector fail-closed, and classifier unavailable. These mean "not ready" rather than "busy". Assisted-by: AGENT:regolo/glm5.2 [TOOL] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix: return 503 when scheduler has no available nodes When the scheduler cannot find any healthy node to serve a model — all nodes are full and eviction cannot free a slot, or a node_selector excludes every candidate — the error fell through to 500. A 500 tells clients something is broken when the condition is transient and retryable. The router now wraps these errors with a new ErrNoAvailableNodes sentinel. The HTTP error handler maps it to 503 via applyNoAvailableNodes, following the same pattern as applyBackendAdmission (429). Unrelated scheduler errors (DB timeouts, registry lookups) still return 500. Three return sites are wrapped: - resolveSelectorCandidates: selector matches zero healthy nodes - scheduleNewModel eviction-busy: all models have in-flight requests - scheduleNewModel eviction-failed: eviction itself errored The existing scheduleAndLoad wrapper ("no available nodes: %w") preserves the sentinel through the chain via errors.Is, as does ModelRouterAdapter. Assisted-by: AGENT:regolo/glm5.2 [TOOL] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(http): use Ginkgo for admission tests Replace forbidden testing.T calls with Ginkgo and Gomega so the lint check accepts the admission handler tests. Assisted-by: Codex:GPT-6 forbidigo * fix(middleware): show the recorded status for admission rejections The admission audit row now records 429, but the Middleware page still printed a hard-coded 503. Read the status from the event, and update the two package comments that still said 503. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-5-5 [Claude Code] --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
This commit is contained in:
13 files changed
+144
-19
No files matched your search
@@ -88,7 +88,9 @@ The `/v1/responses` endpoint returns errors with this structure:
|
||||
| 404 | Not Found | Model or resource does not exist |
|
||||
| 409 | Conflict | Resource already exists (e.g., duplicate token) |
|
||||
| 422 | Unprocessable Entity | Validation failed (e.g., invalid parameter range) |
|
||||
| 429 | Too Many Requests | All backends are saturated (per-model `max_concurrent` or process-wide `--max-concurrent-backend-requests` ceiling reached). Includes a `Retry-After` header and `type: "rate_limit_error"` so OpenAI-compatible clients and harnesses back off automatically |
|
||||
| 500 | Internal Server Error | Backend inference failure, unexpected server errors |
|
||||
| 503 | Service Unavailable | No healthy node available to serve the model (cluster is full, eviction cannot free a slot, or a `node_selector` excludes all candidates). Also used during model-load cooldown and while a model is still cold-loading. Retryable |
|
||||
|
||||
## Global Error Handling
|
||||
|
||||
|
||||
@@ -95,7 +95,7 @@ For more information on VRAM management, see [VRAM and Memory Management]({{%rel
|
||||
| Parameter | Default | Description | Environment Variable |
|
||||
|-----------|---------|-------------|----------------------|
|
||||
| `--address` | `:8080` | Bind address for the API server | `$LOCALAI_ADDRESS`, `$ADDRESS` |
|
||||
| `--max-concurrent-backend-requests` | `1024` | Process-wide ceiling for concurrent backend inference operations. Excess inference receives HTTP 503 with `Retry-After`; UI and administrative endpoints remain available | `$LOCALAI_MAX_CONCURRENT_BACKEND_REQUESTS`, `$MAX_CONCURRENT_BACKEND_REQUESTS` |
|
||||
| `--max-concurrent-backend-requests` | `1024` | Process-wide ceiling for concurrent backend inference operations. Excess inference receives HTTP 429 with `Retry-After`; UI and administrative endpoints remain available | `$LOCALAI_MAX_CONCURRENT_BACKEND_REQUESTS`, `$MAX_CONCURRENT_BACKEND_REQUESTS` |
|
||||
| `--cors` | `false` | Enable CORS (Cross-Origin Resource Sharing) | `$LOCALAI_CORS`, `$CORS` |
|
||||
| `--cors-allow-origins` | | Comma-separated list of allowed CORS origins | `$LOCALAI_CORS_ALLOW_ORIGINS`, `$CORS_ALLOW_ORIGINS` |
|
||||
| `--disable-csrf` | `false` | Disable CSRF middleware (enabled by default) | `$LOCALAI_DISABLE_CSRF` |
|
||||
|
||||
@@ -21,8 +21,9 @@ The left column is the literal string as it appears in the LocalAI server log (o
|
||||
| `grpc service not ready` | The backend process was spawned but its gRPC server did not become healthy in time (slow start, crash on startup, or the process died while loading). When a local backend has already exited, the error includes its exit code and last stderr line. | Use the included stderr diagnostic when present; otherwise check the log lines just above. A crash here often means out of memory, a missing shared library, or an incompatible CPU (see `SIGILL`). Increase available RAM/VRAM or pick a smaller quantization. |
|
||||
| `failed to load model: ...` | Returned by the load endpoints and several feature paths (voice, realtime, audio transform) when the model config could not be resolved or the backend load failed. | Confirm the model name exists (`local-ai models list`) and its YAML is valid. The trailing text carries the specific reason. |
|
||||
| HTTP `503` with a `Retry-After` header, after a load failed | Model-load failure cooldown. After a model fails to load, LocalAI refuses new load attempts for that model for a short window so a client that keeps polling a broken model does not respawn a crashing backend on every request. The window starts at `--model-load-failure-cooldown` (default `10s`) and doubles per consecutive failure up to 5m; it resets on the first success. | Fix the underlying load failure (see the rows above), then wait out the `Retry-After` seconds before retrying, or restart LocalAI to clear the cooldown. Set `--model-load-failure-cooldown 0` (or `LOCALAI_MODEL_LOAD_FAILURE_COOLDOWN=0`) to disable the cooldown entirely. See {{% relref "reference/cli-reference" %}}. |
|
||||
| HTTP `503` with a `Retry-After` header, under load | Per-model concurrency limit reached. When a model config sets a `MaxConcurrent` limit, extra requests are rejected with `503` and a `Retry-After` (whole seconds, floor 1) instead of queueing. | Retry after the advised delay, raise the model's concurrency limit, or run more replicas. |
|
||||
| HTTP `503` when backend inference is saturated | The process-wide `--max-concurrent-backend-requests` backend-execution ceiling is full. This protects inference and in-flight backend-trace memory without blocking UI or administrative endpoints. | Retry after the advised delay, reduce inference concurrency, raise the limit if the host has capacity, or add replicas. |
|
||||
| HTTP `503` with `no available nodes` or `no healthy nodes match selector` | The scheduler could not find any healthy node to serve the model. All nodes are full and eviction cannot free a slot, or a `node_selector` in the model's scheduling config excludes every candidate. | Retry after a node becomes available or an in-flight request completes and frees a slot. In a cluster, add nodes or replicas. If a selector is set, confirm at least one healthy node matches it. |
|
||||
| HTTP `429` with a `Retry-After` header, under load | Per-model concurrency limit reached. When a model config sets a `MaxConcurrent` limit, extra requests are rejected with `429` and a `Retry-After` (whole seconds, floor 1) instead of queueing. | Retry after the advised delay, raise the model's concurrency limit, or run more replicas. |
|
||||
| HTTP `429` when backend inference is saturated | The process-wide `--max-concurrent-backend-requests` backend-execution ceiling is full. This protects inference and in-flight backend-trace memory without blocking UI or administrative endpoints. | Retry after the advised delay, reduce inference concurrency, raise the limit if the host has capacity, or add replicas. |
|
||||
| `invalid pitch` (with CUDA) | The prompt exceeded the model's context size. | Reduce the prompt length, or raise the model's context size (`context_size:` in the model YAML). |
|
||||
| `SIGILL` (illegal instruction) on startup | The prebuilt backend binary uses CPU instructions your CPU does not have (for example AVX512, AVX2, F16C, FMA). | Rebuild the backend for your CPU. In a container, set `REBUILD=true` and disable the unsupported instructions, for example `CMAKE_ARGS="-DGGML_F16C=OFF -DGGML_AVX512=OFF -DGGML_AVX2=OFF -DGGML_FMA=OFF" make build`. |
|
||||
| CUDA / VRAM out of memory (backend log shows `out of memory`, `CUDA error: out of memory`, or the process is killed loading) | The model plus its KV cache does not fit in GPU memory. | Use a smaller quantization, reduce `context_size:`, offload fewer layers to the GPU (lower `gpu_layers:`), or free VRAM held by other processes. On multi-GPU hosts, confirm the model is not trying to load entirely onto one device. |
|
||||
|
||||
Reference in new issue
Block a user