mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-13 06:45:26 -04:00
* fix(advisorylock): set statement_timeout alongside lock_timeout
WithLockCtx already overrides a deployment-wide lock_timeout on its
dedicated connection so a blocking pg_advisory_lock() waits its turn
instead of failing with 55P03. statement_timeout aborts that exact same
statement independently, with SQLSTATE 57014, and was not overridden.
Production roles commonly carry statement_timeout=60s. Any guarded
section longer than that (a cold model load stages for tens of minutes)
therefore killed every concurrent waiter:
advisorylock: acquiring lock 9003261067483446873: ERROR: canceling
statement due to statement timeout (SQLSTATE 57014)
Derive it from the same context budget as lock_timeout, with a matching
RESET so the pooled connection is returned clean.
Assisted-by: Claude Opus 5 [claude-code]
* feat(distributed): add ModelLoadJob, the durable cold-load record
A cold load in distributed mode is a long-running background job, but it
was modelled as a synchronous side effect of an inference request: the
whole of it (backend install, multi-GB staging, checkpoint load) ran
inside the per-model advisory lock. Loading a 35.7 GB GGUF held that lock
for ~20 minutes, so every concurrent request for the same model blocked
on pg_advisory_lock and died at the role's 60s statement_timeout.
Introduce the row that lets the lock shrink to a decision. Exactly one
ModelLoadJob may be active per tracking key; that uniqueness — not the
lifetime of a lock — is what de-duplicates concurrent loaders across
replicas. ClaimLoadJob does its read-then-write under the advisory lock
and nothing else: no network, file or gRPC I/O inside the guarded
section, so a claim costs milliseconds no matter how long the resulting
load takes.
LastProgress is a heartbeat rather than a byte counter. A checkpoint load
legitimately moves zero bytes for many minutes, so a reaper keyed on byte
movement would reclaim a healthy job mid-load; byte progress stays the
concern of load_deadline.go. A job whose heartbeat stops for longer than
the orphan window is reclaimable, so a replica killed mid-load cannot
wedge a model permanently.
Failed jobs keep their row for a short grace so an immediately-following
request reports the real cause instead of silently starting a fresh load
of a model that just failed.
No caller yet — the router moves onto this in the next commit.
Assisted-by: Claude Opus 5 [claude-code]
* refactor(distributed): run cold loads as jobs, outside the advisory lock
Route wrapped the entire cold load — node selection, backend install,
multi-GB staging and the remote LoadModel — in the per-model advisory
lock. The lock's job is to de-duplicate concurrent loaders, a decision
that takes milliseconds; holding it for the tens of minutes the resulting
work takes is what turned a dedup mechanism into a cluster-wide outage
for that model.
Split it into a claim and a run. The claim is the only thing left inside
the lock. The run is a background job owned by the claiming replica and
bounded by the same progress-extended deadline as before; every other
request for that model — local or on another replica — attaches as a
waiter and is served the moment the model is ready, with no duplicate
load and no lock contention.
Waiters share one broadcast rather than an ordered queue: they all want
the identical outcome, so ordering them would add fairness machinery that
changes no result. The local channel wakes same-replica waiters instantly
and a 2s DB poll is the authority, because a waiter on another replica
has no channel to close. On wake a waiter re-runs the warm path rather
than trusting the signal — the model may have been evicted in between.
A waiter whose client disconnects returns immediately and the job keeps
running; it belongs to the job record, not to the request. A failure is
recorded on the row so every waiter reports the real cause, and the row
survives briefly so the next request does not read "no job" as "not
loading" and start a duplicate load of a model that just failed.
The runner heartbeats the row on a fixed interval whether or not bytes
are moving, which is what keeps a legitimately silent checkpoint load
from being reclaimed as an orphan. Phase (installing/staging/loading) and
placement ride to the heartbeat on the context, the same seam
load_deadline.go already uses, so single-host paths are untouched.
Non-distributed mode (no DB) keeps the inline load exactly as it was.
Assisted-by: Claude Opus 5 [claude-code]
* feat(distributed): bound the wait for a loading model and answer with progress
A request whose model is cold-loading now attaches to the running job and
is served the moment the model is ready. That wait has to be bounded: a
held HTTP request cannot survive real infrastructure, and an ingress or LB
idle timeout kills a twenty-minute request regardless of what LocalAI
does.
New LOCALAI_MODEL_LOAD_WAIT (default 60s) bounds the CALLER, never the
load — the job keeps running either way. On expiry the request gets 503
with Retry-After and a structured body naming the model, the node, the
phase, byte progress and an ETA. The `error` envelope keeps OpenAI
clients working; `loading` is additive so they ignore it.
The ETA comes from the job's own observed rate and is omitted rather than
guessed until enough bytes have moved for that rate to mean anything: a
confidently wrong ETA on a twenty-minute wait is worse than none.
Retry-After is that ETA when known, clamped to [5s, 300s], and the wait
budget otherwise.
LOCALAI_MODEL_LOAD_WAIT=0 waits unbounded, for deployments with no proxy
in front. Zero in the config struct still means "unset, use the default",
so the CLI records the operator's zero as ModelLoadWaitUnbounded rather
than losing the distinction.
The distributed branch of ModelLoader.loadModel wrapped the router's
error with %s, which flattened it to a string. Use %w: the typed error is
what the HTTP layer keys the 503 off.
Assisted-by: Claude Opus 5 [claude-code]
* feat(api): add GET /api/models/{id}/load-status
A client that receives 503 while a model stages onto a worker needs
somewhere to poll. This returns the same `loading` object the 503 carries
— phase, node, byte progress and ETA — or 404 when no load is running.
Read-only and observability-shaped, so it is deliberately neither
admin-gated nor feature-gated: it explains a 503 the caller just
received, and hiding that behind a per-modality feature would make the
explanation for a failed image request depend on chat permissions. It
also gets no MCP tool, since there is nothing here an admin would manage
conversationally.
Registered on the surfaces from .agents/api-endpoints-and-auth.md: the
swagger block (existing `models` tag, so /api/instructions needs no new
area), the endpoint discovery maps in RegisterLocalAIRoutes, regenerated
swagger, and the distributed-mode docs page. No FLAG_* usecase is
involved, so capabilities.js is unchanged.
Assisted-by: Claude Opus 5 [claude-code]
* feat(ui): show cold-load progress in Chat and retry when the model is ready
A chat request for a model that is still staging onto a worker now gets a
503 carrying live progress instead of an error. Render it: the composer
shows the phase (installing / staging / loading), the node, the percent
and the ETA, then polls load-status and re-sends the request the moment
the model is ready.
Reuses the staging progress idiom the page already had rather than
inventing a second one — the two sources are folded into one
loadProgress, with the load job winning because it is authoritative
across frontend replicas and knows the phase, where the staging operation
only knows about a byte transfer this replica happens to be performing.
Waiting is bounded (three send attempts, ~30 min of polling each), so a
load that never finishes still surfaces as an error rather than as a
spinner nobody questions. An aborted generation stops the polling too.
Assisted-by: Claude Opus 5 [claude-code]
* fix(distributed): check warm-path cleanup errors
The router moved legacy cleanup calls onto newly linted lines. Report
cleanup failures while preserving the fallback to a cold load.
Assisted-by: Codex:gpt-5 [golangci-lint]
---------
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
693 lines
24 KiB
Go
693 lines
24 KiB
Go
package model
|
|
|
|
import (
|
|
"context"
|
|
"errors"
|
|
"fmt"
|
|
"maps"
|
|
"os"
|
|
"path/filepath"
|
|
"strings"
|
|
"sync"
|
|
"sync/atomic"
|
|
"time"
|
|
|
|
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
|
|
"github.com/mudler/LocalAI/pkg/system"
|
|
"github.com/mudler/LocalAI/pkg/utils"
|
|
|
|
"github.com/mudler/xlog"
|
|
)
|
|
|
|
// new idea: what if we declare a struct of these here, and use a loop to check?
|
|
|
|
// TODO: Split ModelLoader and TemplateLoader? Just to keep things more organized. Left together to share a mutex until I look into that. Would split if we separate directories for .bin/.yaml and .tmpl
|
|
// ModelUnloadHook is called when a model is about to be unloaded.
|
|
// The model name is passed as the argument.
|
|
type ModelUnloadHook func(modelName string)
|
|
|
|
// RemoteModelUnloader handles unloading models from remote backend nodes.
|
|
// In distributed mode, this is implemented by the SmartRouter.
|
|
// When ShutdownModel is called for a model with no local process,
|
|
// RemoteModelUnloader.UnloadRemoteModel is called to tell the remote node to free it.
|
|
type RemoteModelUnloader interface {
|
|
UnloadRemoteModel(modelName string) error
|
|
}
|
|
|
|
// RemoteModelContextUnloader is the optional cancellation/force extension used
|
|
// by distributed workers. Keeping it separate preserves compatibility with
|
|
// existing RemoteModelUnloader implementations.
|
|
type RemoteModelContextUnloader interface {
|
|
UnloadRemoteModelContext(ctx context.Context, modelName string, force bool) error
|
|
}
|
|
|
|
// RemoteModelPresenceChecker reports whether any node in the cluster currently
|
|
// holds the model.
|
|
//
|
|
// Unloading is deliberately idempotent: UnloadRemoteModel succeeds when no node
|
|
// has the model, because cleanup paths (model deletion, config edits, watchdog
|
|
// eviction) legitimately run against an already-unloaded model and must not
|
|
// fail. That makes the unload result unable to distinguish "stopped it" from
|
|
// "there was nothing to stop", so the one caller that needs the distinction —
|
|
// ShutdownModel, which must report 404 rather than a misleading success for a
|
|
// model loaded nowhere — asks separately, before unloading.
|
|
type RemoteModelPresenceChecker interface {
|
|
HasRemoteModel(ctx context.Context, modelName string) (bool, error)
|
|
}
|
|
|
|
// ModelRouter is a callback that routes model loading to a remote node
|
|
// instead of starting a local process. When set on the ModelLoader,
|
|
// grpcModel() will delegate to this function before attempting local loading.
|
|
type ModelRouter func(ctx context.Context, backend, modelID, modelName, modelFile string,
|
|
opts *pb.ModelOptions, parallel bool) (*Model, error)
|
|
|
|
// BackendLoadEvent describes one actual backend load attempt: a backend
|
|
// process spawn (or remote-address attach) followed by its LoadModel RPC.
|
|
// Cache hits and loads coalesced onto another goroutine's in-flight attempt
|
|
// never produce an event, so observers see real loads only. Distributed-mode
|
|
// routing is excluded too: there grpcModel runs per inference request and the
|
|
// worker node owns the actual load.
|
|
type BackendLoadEvent struct {
|
|
ModelID string
|
|
ModelName string
|
|
// Backend is the alias-resolved backend string (e.g. "parakeet-cpp").
|
|
Backend string
|
|
// BackendURI is the resolved runtime serving the load: the installed
|
|
// backend's launcher path (which names the variant directory) or a
|
|
// remote gRPC address. This is what identifies WHICH build served the
|
|
// model — a stale installed backend is invisible in the model config
|
|
// but obvious here.
|
|
BackendURI string
|
|
Duration time.Duration
|
|
Err error
|
|
}
|
|
|
|
type ModelLoader struct {
|
|
ModelPath string
|
|
mu sync.Mutex
|
|
store ModelStore
|
|
loading map[string]chan struct{} // tracks models currently being loaded
|
|
operations *modelOperationLocks
|
|
wd *WatchDog
|
|
externalBackends map[string]string
|
|
lruEvictionMaxRetries int // Maximum number of retries when waiting for busy models
|
|
lruEvictionRetryInterval time.Duration // Interval between retries when waiting for busy models
|
|
onUnloadHooks []ModelUnloadHook
|
|
loadObserver func(BackendLoadEvent)
|
|
remoteUnloader RemoteModelUnloader
|
|
modelRouter ModelRouter // distributed mode: route to remote node
|
|
backendLogs *BackendLogStore
|
|
backendLoggingEnabled atomic.Bool
|
|
// stoppingProcs marks backend processes that LocalAI is stopping on
|
|
// purpose (model unload / graceful shutdown), keyed by the
|
|
// *process.Process pointer. The exit-watcher goroutine in startProcess
|
|
// consults it to decide whether an exit is an expected stop or a crash —
|
|
// the exit code can't, since a child killed by our own SIGTERM/SIGKILL
|
|
// reports -1, indistinguishable from a signal-induced crash.
|
|
stoppingProcs sync.Map
|
|
// loadFailures records, per modelID, the cooldown window applied after a
|
|
// failed load so that a client repeatedly polling a broken model does not
|
|
// spawn (and leak) a fresh backend process on every request. Guarded by mu.
|
|
loadFailures map[string]*loadFailureState
|
|
loadFailureBaseCooldown time.Duration // first cooldown after a failure
|
|
loadFailureMaxCooldown time.Duration // cap for the exponential backoff
|
|
}
|
|
|
|
// loadFailureState tracks consecutive load failures for a single modelID and
|
|
// the instant at which its cooldown window expires.
|
|
type loadFailureState struct {
|
|
consecutive int
|
|
cooldownUntil time.Time
|
|
}
|
|
|
|
// ModelLoadCooldownError is returned when a model load is skipped because a
|
|
// recent attempt failed and the per-model cooldown window has not yet elapsed.
|
|
// The HTTP layer maps it to 503 with a Retry-After header so a polling client
|
|
// backs off instead of triggering a fresh backend start on every request.
|
|
type ModelLoadCooldownError struct {
|
|
ModelID string
|
|
RetryAfter time.Duration
|
|
}
|
|
|
|
func (e *ModelLoadCooldownError) Error() string {
|
|
return fmt.Sprintf("model %q load is in cooldown after a recent failure; retry after %s",
|
|
e.ModelID, e.RetryAfter.Round(time.Second))
|
|
}
|
|
|
|
// NewModelLoader creates a new ModelLoader instance.
|
|
// LRU eviction is now managed through the WatchDog component.
|
|
func NewModelLoader(system *system.SystemState) *ModelLoader {
|
|
nml := &ModelLoader{
|
|
ModelPath: system.Model.ModelsPath,
|
|
store: NewInMemoryModelStore(),
|
|
loading: make(map[string]chan struct{}),
|
|
operations: newModelOperationLocks(),
|
|
externalBackends: make(map[string]string),
|
|
lruEvictionMaxRetries: 30, // Default: 30 retries
|
|
lruEvictionRetryInterval: 1 * time.Second, // Default: 1 second
|
|
backendLogs: NewBackendLogStore(1000),
|
|
loadFailures: make(map[string]*loadFailureState),
|
|
loadFailureBaseCooldown: 10 * time.Second, // Default: 10s after the first failure
|
|
loadFailureMaxCooldown: 5 * time.Minute, // Default cap for the backoff
|
|
}
|
|
|
|
return nml
|
|
}
|
|
|
|
// SetLoadFailureCooldown configures the per-model load-failure backoff. base is
|
|
// the cooldown applied after the first failure; it doubles on each consecutive
|
|
// failure up to max. base is authoritative (base <= 0 disables the cooldown
|
|
// entirely); max only overrides the current cap when > 0.
|
|
func (ml *ModelLoader) SetLoadFailureCooldown(base, max time.Duration) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
if base < 0 {
|
|
base = 0
|
|
}
|
|
ml.loadFailureBaseCooldown = base
|
|
if max > 0 {
|
|
ml.loadFailureMaxCooldown = max
|
|
}
|
|
if ml.loadFailureMaxCooldown < ml.loadFailureBaseCooldown {
|
|
ml.loadFailureMaxCooldown = ml.loadFailureBaseCooldown
|
|
}
|
|
}
|
|
|
|
// cooldownRemaining returns how long the modelID's load cooldown still has to
|
|
// run, or 0 if there is none. Callers must hold ml.mu.
|
|
func (ml *ModelLoader) cooldownRemaining(modelID string) time.Duration {
|
|
st := ml.loadFailures[modelID]
|
|
if st == nil {
|
|
return 0
|
|
}
|
|
if remaining := time.Until(st.cooldownUntil); remaining > 0 {
|
|
return remaining
|
|
}
|
|
return 0
|
|
}
|
|
|
|
// recordLoadFailure grows the modelID's consecutive-failure count and arms the
|
|
// next cooldown window using exponential backoff capped at loadFailureMaxCooldown.
|
|
func (ml *ModelLoader) recordLoadFailure(modelID string) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
if ml.loadFailureBaseCooldown <= 0 {
|
|
return // cooldown disabled
|
|
}
|
|
st := ml.loadFailures[modelID]
|
|
if st == nil {
|
|
st = &loadFailureState{}
|
|
ml.loadFailures[modelID] = st
|
|
}
|
|
st.consecutive++
|
|
// base * 2^(consecutive-1), clamped. Cap the shift to avoid overflowing
|
|
// the Duration; anything past the cap collapses to loadFailureMaxCooldown.
|
|
shift := st.consecutive - 1
|
|
if shift > 20 {
|
|
shift = 20
|
|
}
|
|
backoff := ml.loadFailureBaseCooldown * (1 << shift)
|
|
if backoff <= 0 || backoff > ml.loadFailureMaxCooldown {
|
|
backoff = ml.loadFailureMaxCooldown
|
|
}
|
|
st.cooldownUntil = time.Now().Add(backoff)
|
|
}
|
|
|
|
// clearLoadFailure resets the modelID's failure state after a successful load.
|
|
// Callers must hold ml.mu.
|
|
func (ml *ModelLoader) clearLoadFailure(modelID string) {
|
|
delete(ml.loadFailures, modelID)
|
|
}
|
|
|
|
// GetLoadingCount returns the number of models currently being loaded
|
|
func (ml *ModelLoader) GetLoadingCount() int {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
return len(ml.loading)
|
|
}
|
|
|
|
// OnModelUnload registers a hook that is called when a model is unloaded.
|
|
func (ml *ModelLoader) OnModelUnload(hook ModelUnloadHook) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
ml.onUnloadHooks = append(ml.onUnloadHooks, hook)
|
|
}
|
|
|
|
func (ml *ModelLoader) SetWatchDog(wd *WatchDog) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
ml.wd = wd
|
|
}
|
|
|
|
// SetLoadObserver registers a callback fired after every actual backend load
|
|
// attempt, successful or not. See BackendLoadEvent for what counts as one.
|
|
func (ml *ModelLoader) SetLoadObserver(obs func(BackendLoadEvent)) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
ml.loadObserver = obs
|
|
}
|
|
|
|
func (ml *ModelLoader) notifyLoadObserver(ev BackendLoadEvent) {
|
|
ml.mu.Lock()
|
|
obs := ml.loadObserver
|
|
ml.mu.Unlock()
|
|
if obs != nil {
|
|
obs(ev)
|
|
}
|
|
}
|
|
|
|
// SetRemoteUnloader sets the handler for unloading models on remote nodes.
|
|
// In distributed mode, this should be set to the SmartRouter adapter.
|
|
func (ml *ModelLoader) SetRemoteUnloader(u RemoteModelUnloader) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
ml.remoteUnloader = u
|
|
}
|
|
|
|
// SetModelRouter sets the distributed model router callback.
|
|
// When set, grpcModel() will delegate to this function before attempting local loading.
|
|
func (ml *ModelLoader) SetModelRouter(r ModelRouter) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
ml.modelRouter = r
|
|
}
|
|
|
|
// SetModelStore replaces the default in-memory model store.
|
|
// In distributed mode this is called with a DistributedModelStore.
|
|
func (ml *ModelLoader) SetModelStore(s ModelStore) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
ml.store = s
|
|
}
|
|
|
|
func (ml *ModelLoader) GetWatchDog() *WatchDog {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
return ml.wd
|
|
}
|
|
|
|
func (ml *ModelLoader) BackendLogs() *BackendLogStore {
|
|
return ml.backendLogs
|
|
}
|
|
|
|
func (ml *ModelLoader) SetBackendLoggingEnabled(enabled bool) {
|
|
ml.backendLoggingEnabled.Store(enabled)
|
|
}
|
|
|
|
func (ml *ModelLoader) BackendLoggingEnabled() bool {
|
|
return ml.backendLoggingEnabled.Load()
|
|
}
|
|
|
|
// SetLRUEvictionRetrySettings updates the LRU eviction retry settings
|
|
func (ml *ModelLoader) SetLRUEvictionRetrySettings(maxRetries int, retryInterval time.Duration) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
ml.lruEvictionMaxRetries = maxRetries
|
|
ml.lruEvictionRetryInterval = retryInterval
|
|
}
|
|
|
|
func (ml *ModelLoader) ExistsInModelPath(s string) bool {
|
|
return utils.ExistsInPath(ml.ModelPath, s)
|
|
}
|
|
|
|
func (ml *ModelLoader) SetExternalBackend(name, uri string) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
ml.externalBackends[name] = uri
|
|
}
|
|
|
|
func (ml *ModelLoader) DeleteExternalBackend(name string) {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
delete(ml.externalBackends, name)
|
|
}
|
|
|
|
func (ml *ModelLoader) GetExternalBackend(name string) string {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
return ml.externalBackends[name]
|
|
}
|
|
|
|
func (ml *ModelLoader) GetAllExternalBackends(o *Options) map[string]string {
|
|
ml.mu.Lock()
|
|
defer ml.mu.Unlock()
|
|
backends := make(map[string]string)
|
|
maps.Copy(backends, ml.externalBackends)
|
|
if o != nil {
|
|
maps.Copy(backends, o.externalBackends)
|
|
}
|
|
return backends
|
|
}
|
|
|
|
var knownFilesToSkip []string = []string{
|
|
"MODEL_CARD",
|
|
"README",
|
|
"README.md",
|
|
}
|
|
|
|
var knownModelsNameSuffixToSkip []string = []string{
|
|
".tmpl",
|
|
".keep",
|
|
".yaml",
|
|
".yml",
|
|
".json",
|
|
".txt",
|
|
".pt",
|
|
".onnx",
|
|
".md",
|
|
".ds_store",
|
|
".",
|
|
".safetensors",
|
|
".bin",
|
|
".gguf",
|
|
".ggml",
|
|
".ckpt",
|
|
".zip",
|
|
".tag",
|
|
".bak",
|
|
".partial",
|
|
".tar.gz",
|
|
}
|
|
|
|
func (ml *ModelLoader) ListFilesInModelPath() ([]string, error) {
|
|
files, err := os.ReadDir(ml.ModelPath)
|
|
if err != nil {
|
|
return []string{}, err
|
|
}
|
|
|
|
models := []string{}
|
|
FILE:
|
|
for _, file := range files {
|
|
|
|
for _, skip := range knownFilesToSkip {
|
|
if strings.EqualFold(file.Name(), skip) {
|
|
continue FILE
|
|
}
|
|
}
|
|
|
|
// Skip templates, YAML, .keep, .json, .DS_Store, and other non-model files.
|
|
// Use case-insensitive matching so e.g. CACHEDIR.TAG is caught by ".tag".
|
|
lowerName := strings.ToLower(file.Name())
|
|
for _, skip := range knownModelsNameSuffixToSkip {
|
|
if strings.HasSuffix(lowerName, skip) {
|
|
continue FILE
|
|
}
|
|
}
|
|
// Skip backup files created by LocalAI or huggingface_hub (e.g. model.yaml.bak-pre-gpumem072).
|
|
if strings.Contains(lowerName, ".bak") {
|
|
continue FILE
|
|
}
|
|
|
|
// Skip directories
|
|
if file.IsDir() {
|
|
continue
|
|
}
|
|
|
|
models = append(models, file.Name())
|
|
}
|
|
|
|
return models, nil
|
|
}
|
|
|
|
func (ml *ModelLoader) ListLoadedModels() []*Model {
|
|
ml.mu.Lock()
|
|
store := ml.store
|
|
ml.mu.Unlock()
|
|
|
|
models := []*Model{}
|
|
store.Range(func(_ string, m *Model) bool {
|
|
models = append(models, m)
|
|
return true
|
|
})
|
|
|
|
return models
|
|
}
|
|
|
|
func (ml *ModelLoader) LoadModel(modelID, modelName string, loader func(string, string, string) (*Model, error)) (*Model, error) {
|
|
return ml.LoadModelWithFile(modelID, modelName, modelName, loader)
|
|
}
|
|
|
|
func (ml *ModelLoader) LoadModelWithFile(modelID, modelName, modelFileName string, loader func(string, string, string) (*Model, error)) (*Model, error) {
|
|
if modelFileName == "" {
|
|
modelFileName = modelName
|
|
}
|
|
release := ml.operations.acquire(modelID, false)
|
|
defer release()
|
|
return ml.loadModel(modelID, modelName, modelFileName, loader, true)
|
|
}
|
|
|
|
// loadModel is the implementation behind LoadModelWithFile. checkCooldown gates fresh,
|
|
// independent load triggers behind the per-model failure cooldown; it is set to
|
|
// false for the coalesced retry of an in-flight burst (a follower whose leader
|
|
// just failed), which is not a new trigger and should still get its one retry.
|
|
func (ml *ModelLoader) loadModel(modelID, modelName, modelFileName string, loader func(string, string, string) (*Model, error), checkCooldown bool) (*Model, error) {
|
|
ml.mu.Lock()
|
|
distributed := ml.modelRouter != nil
|
|
ml.mu.Unlock()
|
|
|
|
if distributed {
|
|
// Distributed mode: SmartRouter must run per inference request so
|
|
// PickBestReplica (core/services/nodes/replicapicker.go) picks the
|
|
// least-loaded replica each time. The cached *Model returned from a
|
|
// previous call holds a client wrapper bound to one (nodeID,
|
|
// replicaIndex), so reusing it pins every subsequent request to the
|
|
// node that won the very first pick — defeating per-replica load
|
|
// balancing. Bypass the cache and the loading-coalesce map; the
|
|
// router does its own coalescing for first-time loads (advisory DB
|
|
// lock + singleflight on backend.install RPC), so concurrent first
|
|
// requests still produce a single worker-side install.
|
|
//
|
|
// TODO(distributed-cache): if profiling shows the per-request
|
|
// FindAndLockNodeWithModel SELECT FOR UPDATE becomes a hot path
|
|
// under burst load, replace this branch with a per-modelID cache
|
|
// that holds a *list* of replicas (refreshed every ~5s in
|
|
// background) and picks per call via PickBestReplica against
|
|
// locally-tracked in-flight counters. Same policy, no DB round-trip
|
|
// per inference. Trade-off: cross-frontend in-flight visibility
|
|
// becomes eventually consistent, acceptable for 1-3 frontend
|
|
// deployments.
|
|
modelFile := filepath.Join(ml.ModelPath, modelFileName)
|
|
model, err := loader(modelID, modelName, modelFile)
|
|
if err != nil {
|
|
// %w, not %s: the router reports a still-loading model as a typed
|
|
// error that the HTTP layer turns into 503 plus live progress, and
|
|
// that only survives an unbroken chain.
|
|
return nil, fmt.Errorf("failed to route model with internal loader: %w", err)
|
|
}
|
|
if model == nil {
|
|
return nil, fmt.Errorf("loader didn't return a model")
|
|
}
|
|
// Record the latest mapping so DistributedModelStore.Range, shutdown,
|
|
// and listing endpoints see a representative entry. The DB is the
|
|
// source of truth for cluster-wide state; the local store is just a
|
|
// stub for in-process callers.
|
|
ml.mu.Lock()
|
|
ml.store.Set(modelID, model)
|
|
ml.mu.Unlock()
|
|
return model, nil
|
|
}
|
|
|
|
// Check if we already have a loaded model
|
|
if model := ml.checkIsLoaded(modelID); model != nil {
|
|
return model, nil
|
|
}
|
|
|
|
ml.mu.Lock()
|
|
// A concurrent leader can publish its model between the health check above
|
|
// and this lock acquisition. It is fresh by definition, so avoid starting a
|
|
// duplicate load without performing another network probe under ml.mu.
|
|
if model, ok := ml.store.Get(modelID); ok {
|
|
ml.mu.Unlock()
|
|
return model, nil
|
|
}
|
|
|
|
// If a recent load attempt for this model failed, short-circuit fresh load
|
|
// triggers until the cooldown elapses. This stops a client that keeps
|
|
// polling a broken model from spawning (and leaking) a new backend process
|
|
// on every request. The coalesced follower-retry below passes
|
|
// checkCooldown=false so an in-flight burst still gets its one retry.
|
|
if checkCooldown {
|
|
if retryAfter := ml.cooldownRemaining(modelID); retryAfter > 0 {
|
|
ml.mu.Unlock()
|
|
xlog.Debug("Model load in cooldown after a recent failure", "modelID", modelID, "retryAfter", retryAfter)
|
|
return nil, &ModelLoadCooldownError{ModelID: modelID, RetryAfter: retryAfter}
|
|
}
|
|
}
|
|
|
|
// Check if another goroutine is already loading this model
|
|
if loadingChan, isLoading := ml.loading[modelID]; isLoading {
|
|
ml.mu.Unlock()
|
|
// Wait for the other goroutine to finish loading
|
|
xlog.Debug("Waiting for model to be loaded by another request", "modelID", modelID)
|
|
<-loadingChan
|
|
// Now check if the model is loaded
|
|
model := ml.checkIsLoaded(modelID)
|
|
if model != nil {
|
|
return model, nil
|
|
}
|
|
// If still not loaded, the other goroutine failed. Retry once as part of
|
|
// this burst, bypassing the cooldown gate (we are not a new trigger).
|
|
return ml.loadModel(modelID, modelName, modelFileName, loader, false)
|
|
}
|
|
|
|
// Mark this model as loading (create a channel that will be closed when done)
|
|
loadingChan := make(chan struct{})
|
|
ml.loading[modelID] = loadingChan
|
|
ml.mu.Unlock()
|
|
|
|
// Ensure we clean up the loading state when done
|
|
defer func() {
|
|
ml.mu.Lock()
|
|
delete(ml.loading, modelID)
|
|
close(loadingChan)
|
|
ml.mu.Unlock()
|
|
}()
|
|
|
|
// Load the model (this can take a long time, no lock held)
|
|
modelFile := filepath.Join(ml.ModelPath, modelFileName)
|
|
xlog.Debug("Loading model in memory from file", "file", modelFile)
|
|
|
|
model, err := loader(modelID, modelName, modelFile)
|
|
if err != nil {
|
|
ml.recordLoadFailure(modelID)
|
|
return nil, fmt.Errorf("failed to load model with internal loader: %s", err)
|
|
}
|
|
|
|
if model == nil {
|
|
ml.recordLoadFailure(modelID)
|
|
return nil, fmt.Errorf("loader didn't return a model")
|
|
}
|
|
|
|
// Add to models map and reset any prior failure cooldown for this model.
|
|
ml.mu.Lock()
|
|
ml.clearLoadFailure(modelID)
|
|
ml.store.Set(modelID, model)
|
|
ml.mu.Unlock()
|
|
|
|
return model, nil
|
|
}
|
|
|
|
func (ml *ModelLoader) ShutdownModel(modelName string) error {
|
|
ctx, cancel := context.WithTimeout(context.Background(), gracefulShutdownTimeout)
|
|
defer cancel()
|
|
err := ml.shutdownModel(ctx, modelName, false)
|
|
if errors.Is(err, ErrModelBusy) && forceBackendShutdown {
|
|
xlog.Warn("Model remained busy until the graceful shutdown deadline; forcing shutdown", "model", modelName)
|
|
return ml.ShutdownModelForce(modelName)
|
|
}
|
|
return err
|
|
}
|
|
|
|
// ShutdownModelForce stops a backend without waiting for an in-flight gRPC
|
|
// call to finish first. It is used by the watchdog's busy-killer, which only
|
|
// fires once a backend has been stuck on a call past the busy timeout — the
|
|
// graceful ShutdownModel intentionally waits for active calls until its
|
|
// bounded deadline. See deleteProcess.
|
|
func (ml *ModelLoader) ShutdownModelForce(modelName string) error {
|
|
ctx, cancel := context.WithTimeout(context.Background(), forcedShutdownTimeout)
|
|
defer cancel()
|
|
return ml.shutdownModel(ctx, modelName, true)
|
|
}
|
|
|
|
// ShutdownModelContext is the cancellation-aware lifecycle primitive used by
|
|
// both graceful and forced shutdown wrappers.
|
|
func (ml *ModelLoader) ShutdownModelContext(ctx context.Context, modelName string, force bool) error {
|
|
return ml.shutdownModel(ctx, modelName, force)
|
|
}
|
|
|
|
func (ml *ModelLoader) shutdownModel(ctx context.Context, modelName string, force bool) error {
|
|
release, err := ml.operations.acquireContext(ctx, modelName, true)
|
|
if err != nil {
|
|
return fmt.Errorf("waiting to shut down model %q: %w", modelName, err)
|
|
}
|
|
defer release()
|
|
return ml.deleteProcess(ctx, modelName, force)
|
|
}
|
|
|
|
// isResident reports whether the model store already holds an entry for
|
|
// modelID. Unlike checkIsLoaded it never probes the backend and never evicts,
|
|
// so it is safe to call on a per-request hot path: it answers "have we been
|
|
// here before?", which is exactly what the load logging needs to distinguish a
|
|
// cold load from routine traffic against a resident model.
|
|
func (ml *ModelLoader) isResident(modelID string) bool {
|
|
ml.mu.Lock()
|
|
store := ml.store
|
|
ml.mu.Unlock()
|
|
|
|
_, ok := store.Get(modelID)
|
|
return ok
|
|
}
|
|
|
|
func (ml *ModelLoader) CheckIsLoaded(s string) *Model {
|
|
release := ml.operations.acquire(s, false)
|
|
defer release()
|
|
return ml.checkIsLoaded(s)
|
|
}
|
|
|
|
func (ml *ModelLoader) checkIsLoaded(s string) *Model {
|
|
ml.mu.Lock()
|
|
store := ml.store
|
|
wd := ml.wd
|
|
ml.mu.Unlock()
|
|
|
|
m, ok := store.Get(s)
|
|
if !ok {
|
|
return nil
|
|
}
|
|
|
|
// Health checks are per model. Serializing them here prevents a cold health
|
|
// cache from generating a burst of identical probes without blocking any
|
|
// other model's lifecycle.
|
|
m.healthMu.Lock()
|
|
defer m.healthMu.Unlock()
|
|
|
|
xlog.Debug("Model already loaded in memory", "model", s)
|
|
|
|
// Skip the gRPC health check if the model was recently verified.
|
|
// This avoids serializing concurrent requests behind ml.mu while each
|
|
// one does a network round-trip (especially costly in distributed mode).
|
|
if m.IsRecentlyHealthy() {
|
|
xlog.Debug("Model health check cached, skipping gRPC probe", "model", s)
|
|
return m
|
|
}
|
|
|
|
client := m.GRPC(false, wd)
|
|
|
|
xlog.Debug("Checking model availability", "model", s)
|
|
cTimeout, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
|
|
defer cancel()
|
|
|
|
alive, err := client.HealthCheck(cTimeout)
|
|
if !alive {
|
|
xlog.Warn("GRPC Model not responding", "error", err)
|
|
xlog.Warn("Deleting the process in order to recreate it")
|
|
process := m.Process()
|
|
if process == nil {
|
|
// Remote/distributed model — no local process to check.
|
|
// Only evict on definitive connection errors (node is down).
|
|
// Timeouts may mean the node is busy, so keep the model cached.
|
|
if isConnectionError(err) {
|
|
xlog.Warn("Remote model unreachable (connection error), removing from cache", "model", s, "error", err)
|
|
if delErr := ml.deleteProcess(cTimeout, s, false); delErr != nil {
|
|
xlog.Error("error cleaning up remote model", "error", delErr, "model", s)
|
|
}
|
|
return nil
|
|
}
|
|
xlog.Warn("Remote model health check failed (possible timeout), keeping cached", "model", s, "error", err)
|
|
return m
|
|
}
|
|
if !process.IsAlive() {
|
|
xlog.Debug("GRPC Process is not responding", "model", s)
|
|
// stop and delete the process, this forces to re-load the model and re-create again the service
|
|
err := ml.deleteProcess(cTimeout, s, true)
|
|
if err != nil {
|
|
xlog.Error("error stopping process", "error", err, "process", s)
|
|
}
|
|
return nil
|
|
}
|
|
}
|
|
|
|
m.MarkHealthy()
|
|
return m
|
|
}
|