Compare commits

...

2 Commits

Author SHA1 Message Date
Ettore Di Giacinto
1df5a3ef7f fix(model-artifacts): keep the companion option present on a remote load, even unresolved
Follow-up to the companion-persistence fix. On a live distributed cluster the
managed base_model companion option still failed to reach the remote worker
(nvidia-thor) even with the companion resolved-and-persisted in the config:
the backend logged "Downloading required files for meituan-longcat/LongCat-Video"
and failed "base_model must point to a LongCat-Video checkpoint". Symlinking the
companion into place did not help (the option was simply absent from the worker's
LoadModel), while an explicit absolute base_model in options: worked as a control.

Trace of where the remote ModelOptions is built and whether the companion is
present there:

- The *pb.ModelOptions the worker's LoadModel consumes is built on the
  CONTROLLER by grpcModelOpts (core/backend/options.go) -> withCompanionArtifactOptions,
  set as gRPCOptions, and sent by direct gRPC via FileStagingClient.LoadModel. It
  is NOT rebuilt on the worker. The reconciler's replica scale-up instead replays
  a Postgres-stored proto blob, which already carries whatever grpcModelOpts
  produced.
- withCompanionArtifactOptions is the ONLY builder of ModelOptions.Options in the
  tree, and it emits base_model iff the config's companion artifact has
  Resolved != nil. Staging preserves the option and derives the worker ModelPath
  as the nested per-model staged root, so a resolved companion resolves under it
  without a download (verified end to end; the path-nesting angle is a red
  herring here).

So the option is absent only when the config the loader is serving from carries
the companion WITHOUT a resolved snapshot (its resolved state not reaching the
serving config, e.g. a peer-replica reload from local disk or a config loaded
before resolution). In that state the old code emitted NOTHING for the companion,
and longcat-video fell back to its OWN hardcoded default (BASE_MODEL_ID), which
is exactly the observed download-and-fail.

Fix: an unresolved-but-declared companion no longer vanishes. It now falls back
to its DECLARED source repository id, so the backend fetches the artifact the
config actually asked for instead of a hardcoded default; the resolved snapshot
path (the staged, no-download fast path) is still preferred whenever the
companion is resolved, so the single-node and healthy distributed paths are
unchanged. The fallback logs a warning naming the artifact and repo, and the
router now logs the exact option strings crossing to the worker at debug, so a
recurrence is diagnosable in one load instead of by inference.

Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-23 16:34:56 +00:00
Ettore Di Giacinto
bb2ed02cca fix(model-artifacts): persist companion artifacts, not just the primary
A managed model can declare companion artifacts (LongCat-Video-Avatar-1.5
pulls its tokenizer, text encoder and VAE from the separate LongCat-Video
base repo via a target: companion artifact). preloadOne resolves the whole
set in memory, but the binding written back to disk carried only the
primary: persistArtifactBinding marshalled []Spec{result.Spec} and replaced
the entire artifacts: list with it, silently dropping every companion.

In a single process the loss is invisible because the in-memory config keeps
the companion. It bites on the next controller restart: the config reloads
from the mangled file with the primary alone, so withCompanionArtifactOptions
finds no resolved companion and synthesizes no base_model option. The remote
longcat-video backend then never receives base_model, falls back to
BASE_MODEL_ID and downloads the repo itself ("Downloading required files for
meituan-longcat/LongCat-Video"), failing the load with "base_model must point
to a LongCat-Video checkpoint".

This is why an explicit base_model:<path> added to the config options works
where the managed companion does not: an explicit option lives in options:,
which is never rewritten, while the managed companion lives in artifacts:,
which the binding overwrote.

Persist the full resolved set (primary + every companion), and widen
bindingNeedsPersistence to compare the whole artifact list so a companion
resolving for the first time still triggers a write. The single-node path is
unaffected: there the in-memory config already carried the companion, and the
staging/ModelPath resolution for a remote worker (nested per-model staged
root, #10949) is unchanged and already correct once the option is generated.

Assisted-by: Claude Code:claude-opus-4-8[1m] [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-07-23 14:18:25 +00:00
7 changed files with 182 additions and 25 deletions

View File

@@ -113,11 +113,34 @@ var _ = Describe("companion artifact backend options", func() {
Expect(opts.Options).To(Equal([]string{"attention_backend:sdpa"}))
})
It("skips a companion that has not been resolved yet", func() {
It("names the source repository when the companion is not resolved yet", func() {
// A companion that reaches load time WITHOUT a resolved snapshot must not
// vanish silently: emitting no option lets the backend fall back to its own
// hardcoded default, which is how a distributed longcat-video worker ended
// up trying to load the wrong base model and failing "base_model must point
// to a LongCat-Video checkpoint". Naming the DECLARED repository instead
// points the backend at the artifact the config actually asked for. The
// snapshot path (the staged, no-download fast path) is still preferred
// whenever the companion IS resolved.
cfg := configWithCompanion()
cfg.Artifacts[1].Resolved = nil
opts := grpcModelOpts(cfg, "/models")
_, found := optionValue(opts.Options, "base_model")
Expect(found).To(BeFalse())
value, found := optionValue(opts.Options, "base_model")
Expect(found).To(BeTrue())
Expect(value).To(Equal("meituan-longcat/LongCat-Video"))
// The fallback is a repo reference, never a models-relative snapshot path.
Expect(value).ToNot(ContainSubstring(".artifacts"))
})
It("prefers the resolved snapshot path over the source repository", func() {
opts := grpcModelOpts(configWithCompanion(), "/models")
value, found := optionValue(opts.Options, "base_model")
Expect(found).To(BeTrue())
expected, err := modelartifacts.RelativeSnapshotPath(companionKey)
Expect(err).NotTo(HaveOccurred())
Expect(value).To(Equal(expected))
// The resolved fast path must never degrade to a bare repo id.
Expect(value).ToNot(Equal("meituan-longcat/LongCat-Video"))
})
})

View File

@@ -294,6 +294,13 @@ func EffectiveBatchSize(c config.ModelConfig) int {
//
// An option the author set explicitly always wins: pinning a companion to a
// local checkout has to beat the managed snapshot.
//
// A companion that is declared but NOT resolved falls back to its source
// repository id rather than being dropped: a dropped companion is invisible to
// the backend, which then loads its own hardcoded default and fails far away
// from the cause. The repo-id fallback trades the staging fast path (the weights
// are fetched on the worker) for correctness, and logs a warning so the missing
// controller-side resolution is diagnosable.
func withCompanionArtifactOptions(options []string, artifacts []modelartifacts.Spec) []string {
configured := make(map[string]struct{}, len(options))
for _, option := range options {
@@ -306,19 +313,46 @@ func withCompanionArtifactOptions(options []string, artifacts []modelartifacts.S
// reallocate away from) the config's own slice.
combined := slices.Clone(options)
for _, artifact := range artifacts {
if artifact.Target != modelartifacts.TargetCompanion || artifact.Resolved == nil {
if artifact.Target != modelartifacts.TargetCompanion {
continue
}
if _, exists := configured[artifact.Name]; exists {
xlog.Debug("keeping the configured companion option over the managed snapshot", "artifact", artifact.Name)
continue
}
snapshot, err := modelartifacts.RelativeSnapshotPath(artifact.Resolved.CacheKey)
if err != nil {
xlog.Warn("skipping companion artifact with an unusable cache key", "artifact", artifact.Name, "error", err)
// Preferred fast path: a resolved companion is surfaced as its staged,
// models-relative snapshot directory. Staging materializes exactly this
// path on a remote worker and the backend resolves it under its own
// ModelPath, so the weights are never fetched again at load time.
if artifact.Resolved != nil {
if snapshot, err := modelartifacts.RelativeSnapshotPath(artifact.Resolved.CacheKey); err == nil {
xlog.Debug("surfacing resolved companion snapshot to the backend", "artifact", artifact.Name, "path", snapshot)
combined = append(combined, artifact.Name+":"+snapshot)
continue
} else {
xlog.Warn("companion artifact has an unusable cache key; falling back to its source repository", "artifact", artifact.Name, "error", err)
}
}
// Fallback: the companion reached load time without a resolved snapshot
// (its resolved state never made it into the config the loader is serving
// from, e.g. after a controller restart or a peer-replica config reload).
// Emitting nothing here is what makes the failure so hard to see: the
// backend then falls back to its OWN hardcoded default companion, which on
// a distributed longcat-video worker meant fetching the wrong base model
// and failing "base_model must point to a LongCat-Video checkpoint". Name
// the DECLARED repository instead, so the backend at least fetches the
// artifact the config actually asked for. It is a warn because it means the
// no-download fast path was lost: the controller-side materialization or
// persistence for this companion needs investigating.
if repo := strings.TrimSpace(artifact.Source.Repo); repo != "" {
xlog.Warn("companion artifact is not resolved on the controller; the backend will fetch it by repository id (no staging fast path)",
"artifact", artifact.Name, "repo", repo)
combined = append(combined, artifact.Name+":"+repo)
continue
}
combined = append(combined, artifact.Name+":"+snapshot)
xlog.Warn("companion artifact is neither resolved nor has a source repository; the backend will get no option for it", "artifact", artifact.Name)
}
return combined
}

View File

@@ -10,7 +10,15 @@ import (
"github.com/mudler/LocalAI/pkg/modelartifacts"
)
func persistArtifactBinding(fileName, modelName string, result modelartifacts.Result) error {
// persistArtifactBinding writes the resolved artifact set back into a model's
// config document. It replaces the whole `artifacts:` list, so the caller must
// pass EVERY artifact the model declares — the primary and all companions — not
// just the one that triggered the write. Persisting only the primary silently
// dropped companions from disk, and on the next controller restart the reloaded
// config had no companion at all: withCompanionArtifactOptions then synthesized
// no companion option and a remote backend fell back to fetching the companion
// repo itself, failing the load (the distributed longcat-video base_model bug).
func persistArtifactBinding(fileName, modelName string, artifacts []modelartifacts.Spec) error {
data, err := os.ReadFile(fileName)
if err != nil {
return err
@@ -24,7 +32,7 @@ func persistArtifactBinding(fileName, modelName string, result modelartifacts.Re
return err
}
artifactValue := &yaml.Node{}
encoded, err := yaml.Marshal([]modelartifacts.Spec{result.Spec})
encoded, err := yaml.Marshal(artifacts)
if err != nil {
return err
}

View File

@@ -25,19 +25,16 @@ var _ = Describe("artifact binding persistence", func() {
sibling_only: true
parameters: {model: sibling.gguf}
`), 0644)).To(Succeed())
result := modelartifacts.Result{
RelativePath: ".artifacts/huggingface/0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef/snapshot",
Spec: modelartifacts.Spec{
Name: "model", Target: "model",
Source: modelartifacts.Source{Type: "huggingface", Repo: "owner/repo", Revision: "main"},
Resolved: &modelartifacts.Resolved{
Endpoint: "https://huggingface.co",
Revision: "0123456789abcdef0123456789abcdef01234567",
CacheKey: "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef",
},
primary := modelartifacts.Spec{
Name: "model", Target: "model",
Source: modelartifacts.Source{Type: "huggingface", Repo: "owner/repo", Revision: "main"},
Resolved: &modelartifacts.Resolved{
Endpoint: "https://huggingface.co",
Revision: "0123456789abcdef0123456789abcdef01234567",
CacheKey: "0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef",
},
}
Expect(persistArtifactBinding(fileName, "managed", result)).To(Succeed())
Expect(persistArtifactBinding(fileName, "managed", []modelartifacts.Spec{primary})).To(Succeed())
updated, err := os.ReadFile(fileName)
Expect(err).NotTo(HaveOccurred())
Expect(string(updated)).To(ContainSubstring("name: sibling"))
@@ -47,4 +44,48 @@ var _ = Describe("artifact binding persistence", func() {
Expect(string(updated)).To(ContainSubstring("cache_key: 0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef"))
Expect(string(updated)).To(ContainSubstring("revision: 0123456789abcdef0123456789abcdef01234567"))
})
It("writes back every artifact it is given, primary and companion", func() {
// The binding replaces the whole artifacts list, so a companion is only
// retained if it is passed in. Dropping it here is what lost companions
// on a controller restart (the distributed longcat-video base_model bug).
fileName := filepath.Join(GinkgoT().TempDir(), "models.yaml")
Expect(os.WriteFile(fileName, []byte(`
- name: avatar
backend: longcat-video
artifacts:
- name: model
target: model
source: {type: huggingface, repo: owner/avatar}
- name: base_model
target: companion
source: {type: huggingface, repo: owner/base}
parameters: {model: owner/avatar}
`), 0644)).To(Succeed())
primaryKey := "1111111111111111111111111111111111111111111111111111111111111111"
companionKey := "2222222222222222222222222222222222222222222222222222222222222222"
resolved := func(repo, key string) modelartifacts.Spec {
return modelartifacts.Spec{
Name: "x", Target: "companion",
Source: modelartifacts.Source{Type: "huggingface", Repo: repo, Revision: "main"},
Resolved: &modelartifacts.Resolved{
Endpoint: "https://huggingface.co",
Revision: "0123456789abcdef0123456789abcdef01234567",
CacheKey: key,
},
}
}
primary := resolved("owner/avatar", primaryKey)
primary.Name, primary.Target = "model", "model"
companion := resolved("owner/base", companionKey)
companion.Name = "base_model"
Expect(persistArtifactBinding(fileName, "avatar", []modelartifacts.Spec{primary, companion})).To(Succeed())
updated, err := os.ReadFile(fileName)
Expect(err).NotTo(HaveOccurred())
Expect(string(updated)).To(ContainSubstring("name: base_model"))
Expect(string(updated)).To(ContainSubstring(primaryKey))
Expect(string(updated)).To(ContainSubstring(companionKey))
})
})

View File

@@ -435,12 +435,15 @@ func (bcl *ModelConfigLoader) PreloadWithContext(ctx context.Context, modelPath
bcl.Unlock()
continue
}
if artifactResult != nil && bindingNeedsPersistence(current, *artifactResult) && current.modelConfigFile != "" {
// Persist the WHOLE resolved artifact set (primary + every companion),
// not just the primary result: writing back only the primary dropped
// companions from disk and lost them on the next restart.
if artifactResult != nil && bindingNeedsPersistence(current, updated.Artifacts) && current.modelConfigFile != "" {
modelartifacts.ReportProgress(ctx, modelartifacts.ProgressEvent{
Phase: modelartifacts.PhasePersisting,
Artifact: artifactResult.Spec.Name,
})
if err := persistArtifactBinding(current.modelConfigFile, current.Name, *artifactResult); err != nil {
if err := persistArtifactBinding(current.modelConfigFile, current.Name, updated.Artifacts); err != nil {
bcl.Unlock()
return err
}
@@ -549,8 +552,14 @@ func (bcl *ModelConfigLoader) preloadOne(
return updated, artifactResult, nil
}
func bindingNeedsPersistence(current ModelConfig, result modelartifacts.Result) bool {
return len(current.Artifacts) == 0 || !reflect.DeepEqual(current.Artifacts[0], result.Spec)
// bindingNeedsPersistence reports whether the freshly resolved artifact set
// differs from what is currently on the config, and so has to be written back.
// It compares the WHOLE set, not just the primary: a companion that resolved
// for the first time (or changed) must trigger a write even when the primary is
// unchanged, or its resolved state would never reach disk and would be lost on
// the next restart.
func bindingNeedsPersistence(current ModelConfig, resolved []modelartifacts.Spec) bool {
return !reflect.DeepEqual(current.Artifacts, resolved)
}
func (bcl *ModelConfigLoader) displayPreloadedModel(config ModelConfig) {

View File

@@ -111,6 +111,39 @@ parameters: {model: meituan-longcat/LongCat-Video-Avatar-1.5}
Expect(loaded.ModelFileName()).To(ContainSubstring(loaded.Artifacts[0].Resolved.CacheKey))
})
It("keeps every resolved artifact in the persisted file across a reload", func() {
// Regression for the distributed longcat-video companion loss: a
// controller resolves the primary and companion in memory, but if the
// binding it writes back to disk carries only the primary, the companion
// is gone the moment the process restarts and reloads the file. With no
// companion in the config, withCompanionArtifactOptions synthesizes no
// base_model option, so the remote backend falls back to downloading the
// base repo itself and fails ("base_model must point to a LongCat-Video
// checkpoint"). The persisted document, reloaded fresh, must still name
// the companion.
modelsPath := GinkgoT().TempDir()
configPath := filepath.Join(modelsPath, "avatar.yaml")
Expect(os.WriteFile(configPath, []byte(companionConfig), 0644)).To(Succeed())
fake := &companionMaterializer{}
loader := NewModelConfigLoader(modelsPath, WithArtifactMaterializer(fake))
Expect(loader.LoadModelConfigsFromPath(modelsPath)).To(Succeed())
Expect(loader.PreloadWithContext(context.Background(), modelsPath)).To(Succeed())
// A fresh loader models the restart: it only ever sees what was written
// back to disk, never the in-memory state the first loader held.
reloaded := NewModelConfigLoader(modelsPath, WithArtifactMaterializer(&companionMaterializer{}))
Expect(reloaded.LoadModelConfigsFromPath(modelsPath)).To(Succeed())
persisted, found := reloaded.GetModelConfig("avatar")
Expect(found).To(BeTrue())
Expect(persisted.Artifacts).To(HaveLen(2))
Expect(persisted.Artifacts[1].Name).To(Equal("base_model"))
Expect(persisted.Artifacts[1].Target).To(Equal(modelartifacts.TargetCompanion))
Expect(persisted.Artifacts[1].Resolved).ToNot(BeNil())
Expect(persisted.Artifacts[1].Resolved.CacheKey).ToNot(BeEmpty())
})
It("fails the load when an explicitly declared companion cannot be acquired", func() {
// Explicit artifacts are all-or-nothing: a config that names a companion
// is asserting the backend needs it, so silently loading without it

View File

@@ -342,6 +342,15 @@ func (r *SmartRouter) scheduleAndLoad(ctx context.Context, backendType, tracking
xlog.Info("Loading model on remote node", "node", node.Name, "model", modelName, "addr", backendAddr,
"payloadBytes", payloadBytes, "loadBudget", loadTimeout)
// The exact option strings that cross to the worker. A managed companion
// (e.g. longcat-video's base_model) rides here as a key:value option, and
// its absence is otherwise invisible until the backend fails far away
// having fetched the wrong weights. Logged at debug so a load that
// "downloaded the base model" can be traced to whether the option was
// present in what the worker actually received.
xlog.Debug("Remote LoadModel options", "node", node.Name, "model", modelName,
"options", loadOpts.Options, "modelPath", loadOpts.ModelPath, "modelFile", loadOpts.ModelFile)
// The cold-load hold above this call extends on STAGING progress, and
// the remote LoadModel reports none — so once the last byte lands the
// hold expires a stall window later and would cancel a load that is