mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-29 17:44:30 -04:00
docs: document localai-proxy and distributed failover limits
Add the localai-proxy known limits (no grammar or media forwarding, TTS streams that end cleanly after an upstream failure, the /v1 path in upstream_url), state that the Unimplemented skip covers the APIs that answer HTTP 501, and describe a NATS-partitioned leader and pin re-sync in distributed mode. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
1 parent
4a7c96f533
commit
1b6b4b806a
2 files changed
+21
-1
No files matched your search
@@ -288,6 +288,14 @@ share one failover state:
|
||||
the next target when a target fails during the request.
|
||||
- If a frontend cannot start the shared state, it logs an error and manages
|
||||
failover alone, as a single LocalAI instance does.
|
||||
- Leadership follows the database lock, not NATS. A leader that loses its NATS
|
||||
connection but keeps its database connection stays the leader. Until NATS
|
||||
recovers, the other frontends keep the last decisions they received from it,
|
||||
and its new decisions do not reach them.
|
||||
- Each frontend reads the pins from the database again every 30 seconds and
|
||||
after a NATS reconnect, and applies any change within 10 seconds. A frontend
|
||||
that missed a pin or an unpin catches up in this way. If a pin cannot be written to the database, LocalAI returns an
|
||||
error and restores the previous pin.
|
||||
|
||||
## Limits
|
||||
|
||||
|
||||
@@ -313,7 +313,9 @@ capability gap, not a broken target: audio encoding and decoding,
|
||||
audio-to-audio streams, token classification (PII NER), model metadata,
|
||||
fine-tuning, quantization and model export fall in this bucket. A failover
|
||||
chain skips a target that returns `Unimplemented` and tries the next target,
|
||||
but does not mark the target down.
|
||||
but does not mark the target down. This applies to every API, also to the APIs
|
||||
that report `Unimplemented` to the client as HTTP `501` (images, video, 3D,
|
||||
detection, depth, face and voice).
|
||||
|
||||
Errors from the upstream: a 5xx response (other than 501) or a connection
|
||||
failure becomes `Unavailable`, and a failover chain marks the target down. A
|
||||
@@ -336,6 +338,16 @@ Known limits:
|
||||
transcriptions through the proxy never set it. Live transcription through
|
||||
`realtime_pipeline` sets `eou` at the end of each utterance.
|
||||
- Sound generation from a source audio file is not supported.
|
||||
- Chat and completions do not forward grammars, so JSON mode and other
|
||||
grammar-constrained output are not enforced by the upstream. Images, audio
|
||||
and video attached to messages are not forwarded either. The backend logs a
|
||||
warning for each request that loses one of these fields.
|
||||
- Streamed TTS cannot detect an upstream synthesis failure that ends the
|
||||
stream cleanly. The client receives the audio produced so far as a complete
|
||||
response, and a failover chain does not retry it. A stream that is cut off
|
||||
is reported as an error.
|
||||
- `upstream_url` is the root of the upstream server. If it has a `/v1` path,
|
||||
the backend removes `/v1` and everything after it, and logs a warning.
|
||||
|
||||
## Limitations
|
||||
|
||||
|
||||
Reference in new issue
Block a user