From 4a5c5e51b75478c656de93974eadc88ffd804daf Mon Sep 17 00:00:00 2001 From: Ettore Di Giacinto Date: Wed, 5 Aug 2026 07:33:08 +0000 Subject: [PATCH] website: say plainly that engines are swappable behind the same API The runtime section described the small core and on-demand backends but never stated the simple fact readers look for: one model can run on llama.cpp while the next loads on vLLM, SGLang or MLX, behind the same endpoint, and switching is one line in the model's config. Signed-off-by: Ettore Di Giacinto Assisted-by: Claude Code:claude-fable-5 --- website/layouts/index.html | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/website/layouts/index.html b/website/layouts/index.html index 9cec35949..0d5c01609 100644 --- a/website/layouts/index.html +++ b/website/layouts/index.html @@ -39,7 +39,8 @@

The runtime

Everything else plugs into LocalAI.

One binary with an OpenAI-compatible API in front of it. Point an existing client at it and the calls keep working, except now the model is on your machine. It also speaks the Anthropic, Ollama and ElevenLabs APIs, so most tools need a URL change and nothing else.

-

Underneath, a small core pulls each engine in as a separate backend, only when a model asks for it. That is why one install covers this much ground without becoming a 9 GB download.

+

The engine behind that API is swappable. One model can run on llama.cpp while the next loads on vLLM, SGLang or MLX, and the client never notices: same endpoint, same request, different engine underneath. Switching is one line in the model's config.

+

A small core pulls each engine in as a separate backend, only when a model asks for it. That is why one install covers this much ground without becoming a 9 GB download.

OpenAI APIAnthropic APIOllama APIElevenLabs APIRealtime over WebRTC