Commit Graph
4 Commits
Author SHA1 Message Date
mudler-agentandEttore Di Giacinto 895d50385f fix(distributed): converge model configs across frontends (#12558)
* fix(galleryop): announce model changes before the preload

After a gallery install or delete, the replica that ran it replaced its
config loader, then preloaded every installed model, and only then
published the models invalidation. The preload does remote lookups and
checksums for each model, so on a large models directory peers learned
about the change minutes after the originator listed it. When the
preload failed or the operation was cancelled, the event was never sent.

Publish the invalidation, and apply the delete lifecycle, as soon as
the loader holds the new set. The preload still runs afterwards with
its own error handling. Its failure is reported on the operation, but
it no longer rolls back a deletion that peers have already applied.

Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(distributed): resync model configs from the models directory

Frontends refresh their model configs only when a models invalidation
arrives on NATS. NATS keeps no history, so a frontend that is
disconnected when the message is published never applies the change.
It keeps serving the old config, for example an alias that points at
the previous model, until some later change happens to touch it.

Each frontend now reruns the peer reconcile against the shared models
directory after every NATS reconnect, and every
--model-config-resync-interval (default 30s) when a config file
changed. The pass names no model, so only models whose file changed
get a revision transition, and an unchanged directory costs one read
of the config files.

The reconcile replaced the whole loader with a parse of the models
directory, which dropped models loaded with --config-file and
published a deletion revision for them. Configs defined outside the
directory are now kept, both there and after a gallery install.

Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(nodes): stop stale frontends from retargeting alias rules

A scheduling rule keyed by an alias derives its target from the alias
mapping of the frontend that reads it. Every frontend keeps its own
copy of the model configs, so after an alias is repointed a frontend
that has not reloaded it still resolves the old target. Two frontends
then rewrote the rule's stored target_model against each other on
alternate reconciler ticks, and the outdated one scaled up the model
the alias used to point at.

The registry already records the accepted config revision of each
model. A frontend now derives a rule's target from its own alias
mapping only when its config revision for the rule's name matches
that record. Otherwise it keeps the stored target_model: it neither
writes the column nor reconciles replicas of the old target. The check
reads the database only for a rule whose stored and derived targets
differ. With no accepted revision on record, the old behaviour stays.

Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-08 09:56:12 +02:00
Ettore Di Giacinto 1dc3aeef87 fix(distributed): resolve config revisions through one entry point
A model's revision is published by administration and checked against on
every inference request. Those were computed by separate code: the
request path resolves through the loader, while each publisher hashed
whatever ModelConfig it happened to hold. By then SetDefaults had folded
in the GGUF guess and app-level options, so the published value was one
no request would ever carry and the model became unroutable until the
row was deleted by hand.

Fixing the publishers one at a time did not hold. Three rounds each
found another: the startup resync, then a saved edit and a toggle, then
a rename and the peer-change path.

ModelConfigLoader.RevisionFor is now the only way to obtain a revision,
and the raw hash is unexported, so a caller outside this package cannot
hash a config it holds. A publisher and a request agree by construction
rather than by two implementations happening to match.

The request path no longer falls back to hashing its merged config
either: an unstamped config is routed without a revision rather than
with a wrong one.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [golangci-lint]
2026-08-24 19:11:16 +00:00
mudler's LocalAI [bot]andEttore Di Giacinto 82c191afad fix(distributed): keep model replicas config-consistent (#11664)
* docs: design configurable copy buffering

Document the context-aware copy buffer option and its validation plan.

Assisted-by: Codex:gpt-5

* docs: design durable distributed staging operations

Assisted-by: Codex:gpt-5

* docs: design distributed model config revisions

Assisted-by: Codex:GPT-5 [apply_patch] [exec_command]

* feat(config): add stable model revisions

Hash typed model configuration and effective protobuf options deterministically for distributed revision comparisons.

Assisted-by: Codex:GPT-5 [apply_patch] [exec_command]

* feat(worker): acknowledge exact model stops

Assisted-by: Codex:GPT-5 [apply_patch] [exec_command]

* feat(nodes): track model config revisions

Assisted-by: Codex:GPT-5 [apply_patch]

* fix(distributed): retry quarantined model cleanup

Stop quarantined replicas by exact process identity, retain failed cleanup as durable capped retries, and compare-and-delete only the claimed registry row. Process one sufficiently leased row at a time so multiple frontends cannot duplicate slow cleanup work.

Assisted-by: Codex:gpt-5

* fix(distributed): bind loads to config revisions

Assisted-by: Codex: GPT-5 [OpenAI Codex]

* fix(modeladmin): apply config revisions consistently

Route model edits, patches, state changes, deletion, and peer refreshes through the same revision lifecycle. Quarantine stale replicas before exact cleanup and report durable pending cleanup without failing successful config writes.

Assisted-by: Codex: GPT-5 [OpenAI Codex]

* feat(distributed): expose model config revision state

Document replica revision observability and durable cleanup behavior. Keep pending cleanup explicit in model mutation responses and verify endpoint contracts expose revision state without serialized load options.

Assisted-by: Codex:GPT-5 [OpenAI Codex]

* test(distributed): cover model revision convergence

Exercise cross-frontend quarantine, stale replay rejection, exact cleanup retry, worker re-registration, and current-generation replica convergence against the distributed PostgreSQL harness.

Assisted-by: Codex:gpt-5

* fix(distributed): pass config revision CI checks

Keep configured gallery sources out of authoritative runtime snapshots only after validating their real schema, and harden rollback snapshots against symlink races and non-regular files.

Assisted-by: Codex: GPT-5 [OpenAI Codex]

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-22 22:44:03 +02:00
LocalAI [bot]andEttore Di Giacinto 64150ca7ab fix(distributed): broadcast admin model-config changes across replicas (#10540)
In distributed mode the admin model endpoints (/models/edit, /models/import,
/models/toggle-state and the PATCH config-json endpoint) wrote the YAML to the
shared models dir but reloaded only the local replica's in-memory
ModelConfigLoader. With multiple frontend replicas behind one service, a save
landed on whichever replica handled the request; peers kept serving their stale
in-memory view, so a load-balanced request was a coin-flip between old and new
config (a created alias visible on one replica and missing on the other, an
edited alias target diverging, etc.).

The NATS cache-invalidation channel (SubjectCacheInvalidateModels +
OnModelsChanged) already existed for the gallery install/delete path; these
admin endpoints simply never published on it. Wire them up via a new
GalleryService.BroadcastModelsChanged helper (no-op in standalone mode).

Also fix delete propagation: LoadModelConfigsFromPath is additive and never
drops an entry whose file is gone, so the subscriber hook (which only reloaded
from disk) could not propagate a removal. ApplyRemoteChange now honors the
event op - pruning the element on "delete" and reloading otherwise - and shuts
down any running instance of the affected model so the new config takes effect.
This closes the same latent gap on the gallery delete path.


Assisted-by: Claude:claude-opus-4-8 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-06-27 01:36:57 +02:00