Add the localai-proxy known limits (no grammar or media forwarding, TTS streams that end cleanly after an upstream failure, the /v1 path in upstream_url), state that the Unimplemented skip covers the APIs that answer HTTP 501, and describe a NATS-partitioned leader and pin re-sync in distributed mode. Assisted-by: Claude:claude-opus-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
15 KiB
+++ title = "Cloud passthrough proxy" weight = 28 toc = true url = "/features/cloud-proxy/" description = "Forward requests to OpenAI, Anthropic, or any compatible provider" tags = ["Proxy", "Cloud", "Routing", "Advanced"] categories = ["Features"] +++
LocalAI can forward chat-completion and Anthropic Messages requests to an
external provider instead of running them through the local gRPC backend
pipeline. Configure a model with backend: cloud-proxy and a proxy.upstream_url,
and LocalAI bypasses templating, MCP injection, and the local model loader
entirely - the upstream sees the body the client sent (with only the top-level
model field optionally rewritten).
The streaming PII filter still runs over the upstream's SSE stream, so cloud egress remains subject to the same redaction rules a local model would apply.
When to use this
- Mix local and cloud models in the same LocalAI instance - clients hit one endpoint, LocalAI dispatches per model.
- Apply LocalAI's auth, usage tracking, and PII redaction to cloud traffic before the body leaves the network.
- Use the intelligent router to send small or simple prompts to a local model and complex ones to Claude or GPT-4o.
To fall back to a local model when the upstream is down, list the proxy model in a [failover chain]({{%relref "features/model-failover" %}}).
How it works
- Request hits LocalAI on
/v1/chat/completions(OpenAI-shaped) or/v1/messages(Anthropic-shaped). - The standard auth and routing middleware runs.
- Per-model PII redaction runs request-side as it would for any model.
- The handler detects the
cloud-proxybackend in passthrough mode and loads the cloud-proxy gRPC backend, which owns the outbound HTTP. - The backend POSTs the body to
proxy.upstream_urlwith provider-aware authentication, then streams the SSE response back to core. - The streaming PII filter rewrites per-token text in flight; the upstream's event names and metadata pass through unchanged.
Passthrough mode is wire-format-faithful - it does not translate request shapes between providers. A client posting an OpenAI-shaped body to an Anthropic upstream will get a confused upstream. Use the matching wire format, or switch to translate mode (below).
Configuration
The cloud-proxy backend has one knob - the provider it should authenticate against - and two modes:
proxy.mode |
What it does | When to use |
|---|---|---|
passthrough (default) |
Forwards the request body verbatim to upstream_url. Client must speak the upstream's wire format. |
Same wire format on both ends. |
translate |
Backend converts internal proto to the upstream's wire format. Client can speak OpenAI-shaped requests to an Anthropic upstream, etc. | Cross-format adaptation. |
proxy.provider selects the auth scheme and (in translate mode) the wire
format. Supported values: openai, anthropic.
API keys are loaded from either an environment variable (api_key_env) or a
file (api_key_file). The key never appears in the config file or the admin
UI; pick whichever fits your secret-management setup.
OpenAI passthrough
name: gpt-4o-proxy
backend: cloud-proxy
# When set, replaces the client's "model" field before forwarding.
# Useful when the LocalAI alias differs from the upstream's canonical name.
proxy:
mode: passthrough
provider: openai
upstream_url: https://api.openai.com/v1/chat/completions
api_key_env: OPENAI_API_KEY
upstream_model: gpt-4o
request_timeout_seconds: 120
# PII filtering defaults to ON for cloud-proxy backends. Override by setting
# pii.enabled: false explicitly. Per-pattern action overrides go in
# pii.patterns; see the Middleware admin page or the Middleware feature doc.
pii:
enabled: true
Then start LocalAI with the API key in the environment:
export OPENAI_API_KEY=sk-...
local-ai run
Clients hit http://localhost:8080/v1/chat/completions with "model": "gpt-4o-proxy"
and the request lands on OpenAI's API.
Anthropic passthrough
name: claude-sonnet-proxy
backend: cloud-proxy
proxy:
mode: passthrough
provider: anthropic
upstream_url: https://api.anthropic.com/v1/messages
api_key_env: ANTHROPIC_API_KEY
upstream_model: claude-3-5-sonnet-20241022
request_timeout_seconds: 300
pii:
enabled: true
# Block - not just mask - leaked credentials before they reach the upstream.
patterns:
- id: api_key_prefix
action: block
Anthropic clients hit http://localhost:8080/v1/messages with
"model": "claude-sonnet-proxy".
Other OpenAI-compatible providers
Most third-party providers (Together, Groq, DeepInfra, OpenRouter, …) speak
the OpenAI chat-completions wire format. Use provider: openai with the
provider's URL and API key:
name: llama-3-70b-via-together
backend: cloud-proxy
proxy:
mode: passthrough
provider: openai
upstream_url: https://api.together.xyz/v1/chat/completions
api_key_env: TOGETHER_API_KEY
upstream_model: meta-llama/Llama-3-70b-chat-hf
Translate mode
In translate mode the cloud-proxy backend converts LocalAI's internal proto to the provider's wire format. This lets a client speak one shape (e.g. OpenAI Chat Completions) against an upstream that expects another (e.g. Anthropic Messages).
name: claude-via-openai-clients
backend: cloud-proxy
proxy:
mode: translate
provider: anthropic
upstream_url: https://api.anthropic.com/v1/messages
api_key_env: ANTHROPIC_API_KEY
upstream_model: claude-3-5-sonnet-20241022
Translate mode currently routes only pure-text completions - tool calls,
image blocks, and per-request usage tokens are dropped through the
internal Predict() signature. Use passthrough mode when your clients need
the upstream's full feature set.
Anthropic prompt caching
proxy.cache_prompt: true makes the translator add Anthropic
prompt-cache
breakpoints (cache_control: {type: ephemeral}) to the stable prefix of every
request: the system block, the last tool, and the final message block (at most
three of Anthropic's four allowed breakpoints). Anthropic then serves that
repeated prefix at the cache-read rate (~0.1x input) on subsequent calls, which
sharply cuts cost on agentic or multi-turn workloads that re-send a large,
unchanging system-plus-tools prefix each turn.
The flag only applies with mode: translate and provider: anthropic; it has
no effect in passthrough mode, for other providers, or when unset (the system
field is then still emitted as a bare string).
name: claude-cached
backend: cloud-proxy
proxy:
mode: translate
provider: anthropic
upstream_url: https://api.anthropic.com/v1/messages
api_key_env: ANTHROPIC_API_KEY
upstream_model: claude-3-5-sonnet-20241022
cache_prompt: true
Loading secrets from a file
api_key_file is an alternative to api_key_env when your secret manager
mounts keys as files (e.g. Kubernetes secrets, Docker secrets, Vault Agent):
proxy:
api_key_file: /run/secrets/openai_api_key
The file is read at backend load time and trimmed of surrounding whitespace.
api_key_env and api_key_file are mutually exclusive.
Combining with the intelligent router
A router model can spread traffic across local and cloud candidates. The score classifier reads the policy descriptions and routes per request:
name: smart-router
router:
classifier: score
classifier_model: arch-router-1.5b
fallback: qwen-3-7b-local
activation_threshold: 0.40
policies:
- label: casual
description: small talk, greetings, short answers
- label: code
description: writing or debugging code in any programming language
- label: heavy-reasoning
description: long-form analysis, complex math, multi-step reasoning
candidates:
- model: qwen-3-7b-local
labels: [casual]
- model: gpt-4o-proxy
labels: [casual, code]
- model: claude-sonnet-proxy
labels: [casual, code, heavy-reasoning]
The router rewrites input.Model to the chosen candidate; per-model PII,
ACLs, and the cloud-proxy fork all run against the resolved target.
See [Middleware: PII filtering and intelligent routing]({{< relref "middleware.md" >}}) for the full router and PII-filter reference.
Proxying to another LocalAI (localai-proxy)
cloud-proxy forwards chat and Messages requests only. To serve a model from
another LocalAI instance for every API it has, use backend: localai-proxy.
The backend receives the request from the local pipeline like any other
backend and sends it to the REST API of the upstream LocalAI. Because it is a
normal backend, a localai-proxy model can be a stage of a realtime pipeline
or a target of a [failover chain]({{% relref "features/model-failover" %}}).
name: remote-llm
backend: localai-proxy
known_usecases: [chat]
proxy:
# Base URL of the upstream LocalAI. Do not add /v1 or an endpoint path:
# the backend adds the path for each API.
upstream_url: https://argus.lan:8080
# The model name on the upstream. When empty, the name of this config.
upstream_model: gemma-3-12b
# Optional. The upstream API key, from an environment variable
# (or api_key_file). Sent as "Authorization: Bearer <key>".
api_key_env: ARGUS_API_KEY
# Optional. Time limit for each non-streaming request. Streams have no limit.
request_timeout_seconds: 120
A model that does live transcription in a realtime pipeline also names a
realtime pipeline on the upstream. The backend opens a transcription session
on the upstream /v1/realtime endpoint with that pipeline:
name: remote-stt
backend: localai-proxy
known_usecases: [transcript]
options:
- realtime_pipeline:asr-pipeline
proxy:
upstream_url: https://argus.lan:8080
upstream_model: parakeet
Set known_usecases on every localai-proxy model. Failover uses it to match
targets, and LocalAI cannot guess the usecases of a remote model. For a chat
model, known_usecases: [chat] has one more effect: LocalAI sends the chat
messages to the upstream /v1/chat/completions endpoint, and the upstream
applies its own chat template, tool parsing and reasoning parsing. Without
chat, or when the config has its own templates, LocalAI renders the prompt
locally and sends it to /v1/completions. proxy.mode and proxy.provider
have no effect on this backend.
Supported APIs:
- Text: chat and completions (also streamed), embeddings, rerank, tokenize, detokenize, score.
- Audio: TTS (also streamed), sound generation, transcription (also streamed),
live transcription (with
realtime_pipeline), diarization, VAD, sound classification, audio transformations. - Image, video and 3D: image generation, upscaling, video generation, 3D generation and animation. The backend downloads the files that the upstream generates.
- Vision: object detection, depth, face verification and analysis, voice verification, analysis and embeddings.
- Stores: set, get, delete, find.
Methods that have no REST API on the upstream return the gRPC error
Unimplemented ("localai-proxy: has no upstream counterpart"). The
upstream returning 501 Not Implemented maps to the same code. Both mean a
capability gap, not a broken target: audio encoding and decoding,
audio-to-audio streams, token classification (PII NER), model metadata,
fine-tuning, quantization and model export fall in this bucket. A failover
chain skips a target that returns Unimplemented and tries the next target,
but does not mark the target down. This applies to every API, also to the APIs
that report Unimplemented to the client as HTTP 501 (images, video, 3D,
detection, depth, face and voice).
Errors from the upstream: a 5xx response (other than 501) or a connection
failure becomes Unavailable, and a failover chain marks the target down. A
4xx response becomes InvalidArgument, and LocalAI returns it to the client
without a retry or a trip — except 429 Too Many Requests, which becomes
ResourceExhausted: the request itself is fine, the upstream is just out of
capacity, so a failover chain retries it on the next target and trips the
rate-limited one, moving traffic off it until it recovers.
Known limits:
- Voice-profile paths pass through unresolved. When LocalAI resolves a TTS voice to a local file (for example a voice clone reference), the backend sends that path to the upstream, where it does not exist. Use voices that the upstream knows by name.
- Depth exports are not supported. The upstream writes them to its own disk,
so a depth request with exports or a destination file returns
Unimplemented. Depth maps and points without exports work. - The REST transcription API has no end-of-utterance (
eou) flag, so transcriptions through the proxy never set it. Live transcription throughrealtime_pipelinesetseouat the end of each utterance. - Sound generation from a source audio file is not supported.
- Chat and completions do not forward grammars, so JSON mode and other grammar-constrained output are not enforced by the upstream. Images, audio and video attached to messages are not forwarded either. The backend logs a warning for each request that loses one of these fields.
- Streamed TTS cannot detect an upstream synthesis failure that ends the stream cleanly. The client receives the audio produced so far as a complete response, and a failover chain does not retry it. A stream that is cut off is reported as an error.
upstream_urlis the root of the upstream server. If it has a/v1path, the backend removes/v1and everything after it, and logs a warning.
Limitations
- Passthrough does no wire-shape translation. Use
mode: translate(with the constraints documented above) or send requests that match the upstream's format. - No output-side PII for non-streaming responses. Streaming responses are filtered in flight; buffered responses pass through verbatim. Request-side PII covers both.
- No retry or backoff. Transient upstream failures bubble up to the client
as
502 Bad Gateway. - No request shape validation. If the upstream rejects the body, its error envelope is forwarded to the client unchanged.
Operational notes
- Cloud-proxy backends load like any other gRPC backend - they consume one process per loaded model and appear in the backend management view, but they hold no GPU memory.
- Usage stats and the trace log capture cloud-proxy requests like any other
request. Token counts come from the upstream's
usagefield when present. - Set
request_timeout_secondsdefensively - a hung upstream otherwise ties up an HTTP handler until the client disconnects.
