mirror of
https://github.com/mudler/LocalAI.git
synced 2026-07-30 09:57:57 -04:00
feat(distributed): cache staged-artifact hashes and publish load lifecycle Every load request re-hashed every staged artifact on the controller (probeExisting and the upload path both re-read the full file), which for a large multi-file model on NAS-backed storage is minutes of pure re-reading per request even when nothing changed - observed as ~9 minutes of "Upload skipped (file already exists with matching hash)" before every avatar generation. Cache the local hash in the same .sha256 sidecar the worker-side transfer server already maintains, invalidated whenever the sidecar is older than the file. The whole staging+loading phase was also invisible: the NodeModel row was only written after LoadModel succeeded, so /api/nodes and the UI showed nothing while a cold load spent 10+ minutes staging - indistinguishable from nothing happening. Publish the lifecycle instead: "staging" as soon as the node is chosen, "loading" when the checkpoint load starts, and the existing "loaded" on success, with the row removed on any failure so a dead load does not leave a phantom replica. The early row also reserves the replica slot against concurrent schedulers. The nodes view already renders non-loaded states on model chips; style "staging" like "loading". Audited every state-filtered registry/router query: eviction, routing, reconciler and idle-model queries all filter state='loaded' explicitly, so the new transitional rows are visible to observability surfaces but inert to scheduling decisions (except slot occupancy, intentionally). Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>