mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-10 07:47:29 -04:00
The worker answered model.unload by calling Free on the first backend process in its map, whatever model the request named. Every idle scale-down or LRU eviction of one model could therefore empty another model's backend on the same node. LocalAI still counted that model as loaded, so its next request failed. A parakeet diarization model then returned 501 "speaker profiles require a loaded speaker encoder" until someone reloaded it by hand. Resolve the target from the model name (all replicas), prefer an address when the request carries one, and free nothing for an unknown model. Also let parakeet-cpp Diarize check the diarization model before the speaker-profile capability. A backend with nothing loaded now answers FailedPrecondition, which LocalAI treats as a stale replica and reloads, instead of a final Unimplemented. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>