Add the localai-proxy known limits (no grammar or media forwarding,
TTS streams that end cleanly after an upstream failure, the /v1 path in
upstream_url), state that the Unimplemented skip covers the APIs that
answer HTTP 501, and describe a NATS-partitioned leader and pin
re-sync in distributed mode.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Add the distributed-aware state contributor rule: any feature that
keeps runtime state must choose shared (syncstate), single-runner
(advisorylock), stateless, or documented per-instance behaviour, so it
behaves correctly across multiple frontends instead of diverging
silently. Also sweeps the failover/localai-proxy docs for gaps found
along the way: the UI (chain editor field, health strip, overview
page, chain badge), the 429->ResourceExhausted trip and 501->skip
mappings, and a spec correction for the live-transcription bridge's
actual close behavior.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The e2e suite now registers the localai-proxy binary and points proxy
models back at the test server itself, so a request leaves LocalAI
through the backend, returns over REST and is answered by a mock model.
Chat, embeddings, TTS and transcription through the proxy return the
upstream model's answer; a chain whose proxy target's upstream model
fails to load serves from the local target; and a realtime pipeline
whose LLM stage is a chain on a remote target completes a turn, then
switches to the local target with a localai.model.failover trip event
when a gate in front of the upstream starts answering 503.
The docs describe the localai-proxy backend next to cloud-proxy and add
a per-stage remote LocalAI example to the failover page.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
The Anthropic translate provider builds the upstream request from scratch and
never emitted cache_control, so prompt caching was impossible for OpenAI-format
clients routed through cloud-proxy — even though the entire system prompt + tools
prefix is re-sent on every agentic turn.
Add an opt-in cache_prompt flag (ProxyOptions.cache_prompt; model YAML
proxy.cache_prompt: true). On a translate+anthropic model, buildAnthropicRequest
injects cache_control:{type:ephemeral} on the stable prefix — the system block,
the last tool, and the last message block (at most 3 of Anthropic's 4 allowed
breakpoints). Anthropic then serves the repeated prefix at the cache-read rate
(0.1x input) on subsequent calls, cutting cost on multi-turn/agentic workloads.
No effect in passthrough mode, for non-Anthropic providers, or when unset.
System is widened to any so it can carry the block form required to attach
cache_control, while still marshalling as a bare string when caching is off.
Adds a unit test asserting exactly three breakpoints when on and none when off,
and documents the option in docs/content/operations/cloud-proxy.md.
Assisted-by: Claude:opus-4.8
Signed-off-by: stefanwalcz <stefan.walcz@walcz.de>