mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-21 21:54:52 -04:00
Node placement and replica rules could only name a model, so an operator who pinned "llama3" to the GPU tier had to rewrite the rule whenever a different model took over that job. An alias already gives a stable name for whichever model serves it, and a rule on that name makes it a deployment slot: repoint the alias and the placement follows. A rule keeps the name the operator chose. Reads resolve that name through the config loader to the model the rule governs, so the reconciler counts, schedules and trims replicas of the target, and the router finds an alias-keyed rule from the target it is already routing. An alias that resolves to nothing governs nothing loadable, so the reconciler skips it and the write paths refuse it. A replica is shared by every name that resolves to it, so only one rule can decide where it runs. The REST and MCP write paths reject a rule whose target another rule already governs. A pair that arrives some other way, such as a seed file or an alias repointed onto a model that already has a rule, resolves in favour of the rule named after the model itself and then the oldest, and the rest are listed as shadowed. The eviction guard is the exception: it matches rules to replicas in raw SQL inside a locking transaction and cannot resolve an alias. It reads a stored target that the reconciler refreshes each tick, and falls back to the rule's own name when that target is empty. Assisted-by: Claude:claude-opus-5 golangci-lint eslint Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
99 lines
3.9 KiB
Markdown
99 lines
3.9 KiB
Markdown
|
|
+++
|
|
disableToc = false
|
|
title = "Model Aliases"
|
|
weight = 14
|
|
url = "/features/model-aliases/"
|
|
+++
|
|
|
|
A **model alias** is a model name that redirects all traffic to another
|
|
configured model. Declare `gpt-4` as an alias of `my-llama-3` and every client
|
|
calling `gpt-4` is served by `my-llama-3` with no client reconfiguration: the
|
|
clients keep their existing model name while you control what answers them on
|
|
the server side.
|
|
|
|
## Declaring an alias
|
|
|
|
Create a minimal config file in your models directory:
|
|
|
|
```yaml
|
|
name: gpt-4
|
|
alias: my-llama-3
|
|
```
|
|
|
|
That is the whole config: a `name` (the alias clients call) and an `alias` key
|
|
(the target that actually serves the request).
|
|
|
|
## Rules and behavior
|
|
|
|
- The target (`my-llama-3`) must be an existing, non-alias, enabled model. You
|
|
cannot point an alias at a missing model, a disabled model, or another alias
|
|
(no chains).
|
|
- Aliases are 1:1. One alias maps to exactly one target.
|
|
- The target can be swapped live by editing the config file, calling the API,
|
|
using the UI, or asking the assistant. No restart is required.
|
|
- Both `gpt-4` and `my-llama-3` appear in `GET /v1/models`.
|
|
- Responses echo the requested alias: a call to `gpt-4` returns `gpt-4` in the
|
|
response `model` field, not the target name.
|
|
- Usage accounting records both sides: requested `gpt-4`, served `my-llama-3`.
|
|
- Aliases work for every modality (chat, embeddings, audio, images, and so on).
|
|
|
|
## Managing aliases
|
|
|
|
You can create, swap, and remove aliases from any of the management surfaces.
|
|
|
|
### Web UI
|
|
|
|
Open **Add Model** and pick the **Alias / Routing** template, then set a name
|
|
and a target. To re-point an existing alias, edit it and change the target.
|
|
|
|
### REST API
|
|
|
|
- Create: `POST /models/import`
|
|
- Swap the target: `PATCH /api/models/config-json/:name`
|
|
- List all aliases: `GET /api/aliases`
|
|
- Delete: `POST /models/delete/:name`
|
|
|
|
### Assistant and MCP
|
|
|
|
The LocalAI Assistant (and the MCP server) expose the same operations as tools:
|
|
`set_alias`, `list_aliases`, and `delete_model`.
|
|
|
|
{{% notice note %}}
|
|
**You cannot turn an existing real model into an alias.** If you run `set_alias`
|
|
(or `PATCH /api/models/config-json/:name`) against a name that is already a real,
|
|
non-alias model, the request is **rejected**. An alias is a pure redirect, so it
|
|
must not carry a `backend` or `parameters.model`; a real model does, and merging
|
|
an `alias` onto it produces an invalid config that validation refuses with
|
|
`alias config ... must not set backend or parameters.model`. This is intentional:
|
|
it stops a stray `set_alias` call from clobbering a model that is serving.
|
|
|
|
To add an alias, point a **new** name at the target instead of reusing an
|
|
existing model's name. Re-pointing an **existing alias** at a different target
|
|
is fully supported and is the live-swap path: the alias config has no backend of
|
|
its own, so swapping its target stays a valid pure redirect.
|
|
{{% /notice %}}
|
|
|
|
## Aliases as deployment slots (distributed mode)
|
|
|
|
In [distributed mode]({{%relref "features/distributed-mode" %}}) an alias can
|
|
carry a scheduling rule. `POST /api/nodes/scheduling` accepts an alias for
|
|
`model_name`, and the rule then governs whatever model the alias points at:
|
|
|
|
```bash
|
|
curl -X POST http://frontend:8080/api/nodes/scheduling \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"model_name": "production", "node_selector": {"tier": "gpu"}, "min_replicas": 2}'
|
|
```
|
|
|
|
Re-point `production` and the placement policy follows it, so the alias behaves
|
|
as a stable slot whose contents you can swap. Because a replica is shared by
|
|
every name that resolves to it, only one rule may govern a given model at a
|
|
time. See [Scheduling a model alias]({{%relref "features/distributed-mode" %}}#scheduling-a-model-alias).
|
|
|
|
## Limits
|
|
|
|
Aliases are a static 1:1 redirect. For classifier-based or load-balanced
|
|
selection across several downstream models, use the intelligent router in the
|
|
[Middleware]({{%relref "operations/middleware" %}}) feature instead.
|