mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-05 12:34:43 -04:00
* feat(messaging): add shared subject rules Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(messaging): cover BroadcastRoots, ControlRoots and SubjectRoot Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(messaging): add Broadcaster and enforce subject rules in every carrier Broadcaster is the fan-out half of MessagingClient. The NATS client and the in-memory FakeBus now refuse a subject outside the served roots and any wildcard other than a whole single token, and FakeBus shares MatchSubject instead of its own copy. FakeBus Unsubscribe now removes its own subscription instead of the first one with the same subject. A shared conformance suite in messagingtest runs against both carriers. The distributed e2e specs that used invented test.* subjects, and the one that subscribed with a > filter, now use subjects from subjects.go. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: depend on Broadcaster where only publish and subscribe are used Narrowed to messaging.Broadcaster: nodes/staging_progress.go, nodes/install_progress_publisher.go, galleryop/operation.go, galleryop/service.go, agentpool/user_services.go, agentpool/agent_jobs.go, openresponses/store.go, openresponses/sync.go, syncstate/syncstate.go, finetune/service.go, quantization/service.go and failover/distsync/distsync.go. SubscribeJSON now takes a Broadcaster because it only calls Subscribe, which lets the narrowed consumers use it. Stayed wide: worker/supervisor.go, because its client field also serves the SubscribeReply handlers in worker/lifecycle.go. The request/reply, queue and wiring files (nodes/unloader.go, nodes/file_stager_s3.go, jobs/dispatcher.go, agents/dispatcher.go, agents/events.go, worker/file_staging.go, cli/agent_worker.go, http/app.go) are unchanged by design. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): name the no-route condition and confine the carrier error Consumers matched nats.ErrNoResponders, which names an absence, to demote a node. They now match ErrNoRoute, the control path maps the carrier's failure onto it, and timeouts and worker refusals are pinned as not being no-route. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs(nodes): state which FileStager implementations return ErrNoRoute Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): build backend clients through one node-aware seam Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: describe the distributed transport seams Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: correct comments that overclaim after the seams refactor Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(agent-worker): refuse an unserved LOCALAI_AGENT_SUBJECT at startup The messaging client now refuses a subject whose root no carrier serves. An agent worker started with a custom LOCALAI_AGENT_SUBJECT such as tenant-a.agent.execute used to start and then wait on a subject the frontend never publishes to. After the subject rules landed it exited at subscribe time with an error that did not name the setting. Behaviour change: the worker now checks LOCALAI_AGENT_SUBJECT before it registers or connects, and exits with an error that names the variable and says to use a served subject under the agent root, for example agent.execute. The served roots are not widened: a custom root was never delivered by the frontend, and a wider set would reopen the drift the subject rules exist to close. The flag help and the agent worker docs state the constraint. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(nodes): pin the reactions to ErrNoRoute Three callers react to ErrNoRoute and had no spec: the reconciler's upgrade drain falls back to the legacy forced install, the reconciler marks the node unhealthy when a pending op has no route, and the backend-op fan-out marks the node unhealthy. Each spec drives the real caller with a scripted no-responders reply and reads the result from the registry or the recorded requests. A fourth spec pins the other side: a pending op that times out leaves the node healthy and only counts the attempt, so mapping timeouts onto ErrNoRoute would fail here. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(messaging): pin client subject checks and fail the carrier suite in CI Add specs that call Publish, Request, Subscribe, QueueSubscribe, SubscribeReply and QueueSubscribeReply on a client with no connection. Each call must return ErrUnservedSubject for bogus.thing and ErrUnsupportedWildcard for jobs.>. This proves that the subject check runs before the connection is used, and needs no server. The NATS conformance suite is the only check that runs the subject rules against a real carrier. Before this change it skipped without output when Docker was missing. Now it fails when CI is set, so a Linux runner without Docker cannot hide it. It still skips on local runs and on macOS CI, which has no Docker. Add SubjectNodeBackendInstallProgress to the list of constructors that must build served subjects, and ask contributors to extend the list. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: state what ErrNoRoute may change, and group the distributed guides The seams note said MarkUnhealthy was the only state change allowed on ErrNoRoute. A pending backend op still records the failed attempt, counts toward the reconciler's retry limit and is dead-lettered after the maximum attempts. The note now says that MarkUnhealthy is the only change to the node's own state, and that the per-op accounting is not a verdict about the node. The note also documents that the NATS conformance run fails under CI when Docker is missing. The distributed-seams row moves next to the distributed-state row in the topics table. The liveness ping spec header now says no route is a reason to skip the worker, not proof that the worker is gone. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): give the backend client factory the node id Mechanical: the method gains a nodeID parameter and the eight test fakes are updated. No behaviour change. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): drop the optional node-aware factory The node id is now in the main method, so the optional interface and its helper had no behaviour of their own. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): dial backend probes through the client factory Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(nodes): dial workers' file servers through a per-node dialer Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(http): proxy backend logs through the per-node worker dialer The admin backend-logs proxy (list, lines and the WebSocket stream) now reaches a worker through the same per-node dialer as the HTTP file stager, so every frontend-to-worker dial goes through one seam. The shared direct dialer keeps alive for 15s where the proxy used 30s. Harmless for requests bounded at 15s. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(http): keep the backend-logs proxy independent of the admin connection The proxy request had no context before the dialer change and is bounded only by its 15s timeout. Keep it that way. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: move the worker control payloads to workerctl Mechanical move of the request and reply structs, the install progress event and the file payloads out of messaging. The verbs no longer belong to one carrier. No alias is left behind. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(worker): serve the lifecycle verbs through a controlServer The worker registers one handler per verb and a NATS server maps each verb to its subject. Registration errors now name the verb. node.stop is served with SubscribeReply, which is identical on the wire because the handler never replies. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(worker): report install progress through the control sink Install and upgrade now emit download progress through the sink the control server hands them. The debounce and the terminal flush stay in the handler path, built over that sink by the new nodes.NewDebouncedInstallProgressSink, which replaces NewDebouncedInstallProgressPublisher. The subject and payload on the wire are unchanged. The supervisor no longer holds the bus, and installFn and upgradeFn let specs drive both verbs without a gallery. The malformed-request log lines are restored for install, upgrade, backend.delete, model.unload, model.stop and model.delete, with the reply bytes unchanged. The signal adapter is renamed noReply, which also lets worker.go import os/signal without an alias again. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(worker): serve the file-staging verbs through a controlServer An empty list-dir answer is now {} rather than {"files":null}, because the typed reply omits an empty Files slice. The frontend decodes both to a nil slice in nodes/file_stager_s3.go. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(messaging): add WorkQueue and the NATS producer This is the producer side of the competing-consumer seam. The work kinds map one to one to today's subjects and queue groups: task to jobs.new and mcp-ci to jobs.mcp-ci.new (both in group workers), agent-run to agent.execute (group agent-workers). Enqueue publishes the payload as Publish does today, with one JSON marshal. FakeBus now records queue groups and keeps reply handlers so later specs can pin and drive them. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(messaging): add the NATS WorkConsumer An in-flight limit of one runs the handler inline on the delivery goroutine, as the MCP CI consumer does today. Any other limit spawns per delivery, as the agent consumer does. Queue groups are unchanged. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: publish queued work through WorkQueue The job dispatcher, the agent pool and the agent scheduler enqueue through messaging.WorkQueue; the NATS implementation publishes to the same subjects as before. DistributedServices builds the queue next to the NATS client and hands it to the dispatcher and the agent pool, whose distributed mode switch now reads a non-nil WorkQueue. The unused AgentPoolService.SetNATSClient is removed. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: consume queued work through WorkConsumer The agent dispatcher and the MCP CI consumer register through messaging.WorkConsumer. The NATS implementation keeps the inline one-at-a-time model for MCP CI and the per-delivery model for agent runs. handleMCPCIJob reports on the events publisher the carrier hands it instead of a captured client. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: delete the consumers nothing in production reached jobs.new has a producer and no production consumer, and the agent dispatcher's Dispatch was only called from tests. Publishing jobs.new is unchanged. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(mcp): send MCP requests to agent workers through AgentControl Timeouts still honour only the deadline, not cancellation, exactly as today. The NATS no-responders error maps to ErrNoRoute and a timeout does not. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(agent-worker): serve MCP requests and backend.stop through agentRPCServer The agent worker's MCP tool and discovery reply subscriptions and its backend stop listener move behind an unexported agentRPCServer interface, served on NATS by nodes.NATSAgentRPCServer. The handlers become typed mcp.ToolHandler and mcp.DiscoveryHandler values that answer every failure with a reply carrying Error. Queue group (agent-workers), inline execution on the delivery goroutine, the background handler context, the unmarshal error reply texts and the reply-less backend stop subscription are unchanged. The backend stop handler takes the decoded backend name, so it can still close that backend's MCP sessions. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor(messaging): remove helpers that only tests used BroadcastRoots, ControlRoots and SubjectRoot had no production caller. The roots spec now asserts every served root through ValidateSubject instead. MatchSubject moves back into the test support package, the only place that used it, with its table. NATSAgentRPCServer drops the subscription list it stored and never read, and NewNATSAgentRPCServer gets a doc comment. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test(mcp): round trip the agent RPC server over a real NATS server One spec sends a tool request and a discovery request through NATSAgentControl to NATSAgentRPCServer and checks that the handlers see the decoded requests and the replies come back. It also puts an undecodable body on the tool subject and checks the server answers with an unmarshal error instead of leaving the requester to time out. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: describe the distributed transport seams The developer note now lists the final seams: fan-out, queues, both halves of the control verbs and of agent RPC, and the dial. It records the open items a second carrier has to handle. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * test: pin the in-flight limit each queue consumer asks for The work queue specs pin what Consume does for a given limit, but nothing pinned which limit each production consumer passes. Changing the agent worker's MCP CI limit from 1 to 0 would have let MCP CI jobs run concurrently on each worker with every test green. Move the MCP CI Consume call into startMCPCIConsumer with the same wiring and pin that it asks for (WorkMCPCI, 1). Pin that NATSDispatcher.Start asks for (WorkAgentRun, maxConcurrent) for several limits. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * refactor: remove helpers the branch left without a caller SubjectJobCancelWildcard lost its last subscriber when the frontend stopped listening on jobs.*.cancel; the NATS permissions and conformance suite spell the subject out, so nothing reads the constant. decodeBackendStopRequest returned a stopAll flag that production dropped and only a test read. decodeBackendStop is now the single decoder with the same semantics: an empty body is stop-all, an empty Backend is stop-all, malformed JSON is an error. stopBackends still derives stop-all from Backend, so no reply changes. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(messaging): keep an explicitly empty agent queue a plain subscription Before the work queue seam the agent worker passed LOCALAI_AGENT_QUEUE straight to QueueSubscribe, so an explicitly empty value made a plain subscription and every agent worker ran every agent run. WithAgentRunRoute replaced an empty queue with agent-workers, which silently changed that. Keep the queue as given once the option is applied. An empty subject still falls back to agent.execute, since it never had a meaning of its own. The flag default stays agent-workers, so only an explicitly empty value reaches this. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: correct comments and record the PR B notes Fix the recordingFactory comment (it also records the parallel flag), document that a negative maxInFlight is unbounded and that Unsubscribe from a handler deadlocks, and say a permanently undecodable payload returns nil. Record controlHandler's undecodable return as a kept exception, and add the second carrier notes to the developer note: the reconciler has no ClientFactory option, the logs proxy honours HTTP_PROXY, verbs one carrier serves need an opt-out, terminal replies come from the result event, and agent runs publish through the NATS-bound EventBridge, which is not an additive change. Assisted-by: Claude:claude-sonnet-5-5 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
842 lines
27 KiB
Go
842 lines
27 KiB
Go
package quantization
|
|
|
|
import (
|
|
"cmp"
|
|
"context"
|
|
"encoding/json"
|
|
"fmt"
|
|
"os"
|
|
"path/filepath"
|
|
"regexp"
|
|
"slices"
|
|
"strings"
|
|
"sync"
|
|
"time"
|
|
|
|
"github.com/google/uuid"
|
|
"github.com/mudler/LocalAI/core/config"
|
|
"github.com/mudler/LocalAI/core/gallery/importers"
|
|
"github.com/mudler/LocalAI/core/schema"
|
|
"github.com/mudler/LocalAI/core/services/distributed"
|
|
"github.com/mudler/LocalAI/core/services/messaging"
|
|
"github.com/mudler/LocalAI/core/services/syncstate"
|
|
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
|
|
"github.com/mudler/LocalAI/pkg/model"
|
|
"github.com/mudler/LocalAI/pkg/utils"
|
|
"github.com/mudler/xlog"
|
|
"gopkg.in/yaml.v3"
|
|
)
|
|
|
|
// QuantizationService manages quantization jobs and their lifecycle.
|
|
type QuantizationService struct {
|
|
appConfig *config.ApplicationConfig
|
|
modelLoader *model.ModelLoader
|
|
configLoader *config.ModelConfigLoader
|
|
|
|
// mu serializes the read-modify-write of job values. The SyncedMap guards its
|
|
// own map structure, but a job is a pointer mutated in place (e.g. the import
|
|
// goroutine), so the service still needs a lock to keep those field updates and
|
|
// the subsequent Set atomic with respect to readers.
|
|
mu sync.Mutex
|
|
|
|
// jobs is the cross-replica job store: an in-memory map kept consistent across
|
|
// replicas via NATS, optionally read-through to PostgreSQL in distributed mode.
|
|
jobs *syncstate.SyncedMap[string, *schema.QuantizationJob]
|
|
|
|
// progressMu guards progressSubs.
|
|
//
|
|
// A backend's per-job progress stream has a single destructive consumer: the
|
|
// backend pops each update off one queue and hands it to whoever is reading.
|
|
// So the service opens that stream exactly once per job — in watchProgress,
|
|
// started by StartJob — and fans the updates out in-process to the SSE clients
|
|
// registered here. Opening a second stream per client would make the two
|
|
// readers race for the same updates.
|
|
progressMu sync.Mutex
|
|
progressSubs map[string][]chan *schema.QuantizationProgressEvent
|
|
}
|
|
|
|
// progressSubBuffer is the per-subscriber event buffer. It absorbs a client that
|
|
// is briefly slow; a client that falls further behind drops events rather than
|
|
// stalling the single reader of the backend stream.
|
|
const progressSubBuffer = 64
|
|
|
|
// isTerminalStatus reports whether a job status is final, i.e. no further
|
|
// progress update will follow.
|
|
func isTerminalStatus(status string) bool {
|
|
return status == "stopped" || status == "completed" || status == "failed"
|
|
}
|
|
|
|
// NewQuantizationService creates a new QuantizationService. In distributed mode
|
|
// pass the shared NATS client and PostgreSQL store so jobs stay consistent across
|
|
// replicas; pass nil for both in standalone mode, where the disk Loader hydrates
|
|
// the map and there is nothing to broadcast.
|
|
func NewQuantizationService(
|
|
appConfig *config.ApplicationConfig,
|
|
modelLoader *model.ModelLoader,
|
|
configLoader *config.ModelConfigLoader,
|
|
nats messaging.Broadcaster,
|
|
store *distributed.QuantStore,
|
|
) *QuantizationService {
|
|
s := &QuantizationService{
|
|
appConfig: appConfig,
|
|
modelLoader: modelLoader,
|
|
configLoader: configLoader,
|
|
progressSubs: make(map[string][]chan *schema.QuantizationProgressEvent),
|
|
}
|
|
|
|
// Only attach a Store interface when a concrete store exists, otherwise the
|
|
// SyncedMap would see a non-nil interface wrapping a nil pointer and try to
|
|
// hydrate/write through a nil DB.
|
|
var syncStore syncstate.Store[string, *schema.QuantizationJob]
|
|
if store != nil {
|
|
syncStore = &quantStoreAdapter{store: store}
|
|
}
|
|
|
|
s.jobs = syncstate.New(syncstate.Config[string, *schema.QuantizationJob]{
|
|
Name: "quant.jobs",
|
|
Key: func(j *schema.QuantizationJob) string { return j.ID },
|
|
Nats: nats,
|
|
Store: syncStore,
|
|
Loader: s.loadJobsFromDisk, // ignored when Store is set (distributed mode)
|
|
})
|
|
|
|
// Hydrate + subscribe. A hydrate failure must not take the server down: log and
|
|
// continue degraded (standalone), mirroring the FineTune/OpCache wiring.
|
|
if err := s.jobs.Start(appConfig.Context); err != nil {
|
|
xlog.Warn("Quantization SyncedMap start failed; running degraded", "error", err)
|
|
}
|
|
return s
|
|
}
|
|
|
|
// Close releases the SyncedMap subscription and background workers.
|
|
func (s *QuantizationService) Close() error {
|
|
return s.jobs.Close()
|
|
}
|
|
|
|
// quantizationBaseDir returns the base directory for quantization job data.
|
|
func (s *QuantizationService) quantizationBaseDir() string {
|
|
return filepath.Join(s.appConfig.DataPath, "quantization")
|
|
}
|
|
|
|
// jobDir returns the directory for a specific job.
|
|
func (s *QuantizationService) jobDir(jobID string) string {
|
|
return filepath.Join(s.quantizationBaseDir(), jobID)
|
|
}
|
|
|
|
// saveJobState persists a job's state to disk as state.json.
|
|
func (s *QuantizationService) saveJobState(job *schema.QuantizationJob) {
|
|
dir := s.jobDir(job.ID)
|
|
if err := os.MkdirAll(dir, 0750); err != nil {
|
|
xlog.Error("Failed to create quantization job directory", "job_id", job.ID, "error", err)
|
|
return
|
|
}
|
|
|
|
data, err := json.MarshalIndent(job, "", " ")
|
|
if err != nil {
|
|
xlog.Error("Failed to marshal quantization job state", "job_id", job.ID, "error", err)
|
|
return
|
|
}
|
|
|
|
statePath := filepath.Join(dir, "state.json")
|
|
if err := os.WriteFile(statePath, data, 0640); err != nil {
|
|
xlog.Error("Failed to write quantization job state", "job_id", job.ID, "error", err)
|
|
}
|
|
}
|
|
|
|
// loadJobsFromDisk scans the quantization directory for persisted jobs and
|
|
// returns them. It is the SyncedMap Loader used in standalone mode (no DB); the
|
|
// returned slice hydrates the map on Start.
|
|
func (s *QuantizationService) loadJobsFromDisk(_ context.Context) ([]*schema.QuantizationJob, error) {
|
|
baseDir := s.quantizationBaseDir()
|
|
entries, err := os.ReadDir(baseDir)
|
|
if err != nil {
|
|
// Directory doesn't exist yet — that's fine, start empty.
|
|
return nil, nil
|
|
}
|
|
|
|
var jobs []*schema.QuantizationJob
|
|
for _, entry := range entries {
|
|
if !entry.IsDir() {
|
|
continue
|
|
}
|
|
statePath := filepath.Join(baseDir, entry.Name(), "state.json")
|
|
data, err := os.ReadFile(statePath)
|
|
if err != nil {
|
|
continue
|
|
}
|
|
|
|
var job schema.QuantizationJob
|
|
if err := json.Unmarshal(data, &job); err != nil {
|
|
xlog.Warn("Failed to parse quantization job state", "path", statePath, "error", err)
|
|
continue
|
|
}
|
|
|
|
// Jobs that were running when we shut down are now stale
|
|
if job.Status == "queued" || job.Status == "downloading" || job.Status == "converting" || job.Status == "quantizing" {
|
|
job.Status = "stopped"
|
|
job.Message = "Server restarted while job was running"
|
|
}
|
|
|
|
// Imports that were in progress are now stale
|
|
if job.ImportStatus == "importing" {
|
|
job.ImportStatus = "failed"
|
|
job.ImportMessage = "Server restarted while import was running"
|
|
}
|
|
|
|
jobs = append(jobs, &job)
|
|
}
|
|
|
|
if len(jobs) > 0 {
|
|
xlog.Info("Loaded persisted quantization jobs", "count", len(jobs))
|
|
}
|
|
return jobs, nil
|
|
}
|
|
|
|
// StartJob starts a new quantization job.
|
|
func (s *QuantizationService) StartJob(ctx context.Context, userID string, req schema.QuantizationJobRequest) (*schema.QuantizationJobResponse, error) {
|
|
s.mu.Lock()
|
|
defer s.mu.Unlock()
|
|
|
|
jobID := uuid.New().String()
|
|
|
|
backendName := req.Backend
|
|
if backendName == "" {
|
|
backendName = "llama-cpp-quantization"
|
|
}
|
|
|
|
quantType := req.QuantizationType
|
|
if quantType == "" {
|
|
quantType = "q4_k_m"
|
|
}
|
|
|
|
// Always use DataPath for output — not user-configurable
|
|
outputDir := filepath.Join(s.quantizationBaseDir(), jobID)
|
|
|
|
// Build gRPC request
|
|
grpcReq := &pb.QuantizationRequest{
|
|
Model: req.Model,
|
|
QuantizationType: quantType,
|
|
OutputDir: outputDir,
|
|
JobId: jobID,
|
|
ExtraOptions: req.ExtraOptions,
|
|
}
|
|
|
|
// Load the quantization backend (per-job model ID so multiple jobs can run concurrently)
|
|
modelID := backendName + "-quantize-" + jobID
|
|
backendModel, err := s.modelLoader.Load(
|
|
model.WithBackendString(backendName),
|
|
model.WithModel(backendName),
|
|
model.WithModelID(modelID),
|
|
)
|
|
if err != nil {
|
|
return nil, fmt.Errorf("failed to load backend %s: %w", backendName, err)
|
|
}
|
|
|
|
// Start quantization via gRPC
|
|
result, err := backendModel.StartQuantization(ctx, grpcReq)
|
|
if err != nil {
|
|
return nil, fmt.Errorf("failed to start quantization: %w", err)
|
|
}
|
|
if !result.Success {
|
|
return nil, fmt.Errorf("quantization failed to start: %s", result.Message)
|
|
}
|
|
|
|
// Track the job
|
|
job := &schema.QuantizationJob{
|
|
ID: jobID,
|
|
UserID: userID,
|
|
Model: req.Model,
|
|
Backend: backendName,
|
|
ModelID: modelID,
|
|
QuantizationType: quantType,
|
|
Status: "queued",
|
|
OutputDir: outputDir,
|
|
ExtraOptions: req.ExtraOptions,
|
|
CreatedAt: time.Now().UTC().Format(time.RFC3339),
|
|
Config: &req,
|
|
}
|
|
// Set write-through persists to PostgreSQL (distributed) and broadcasts to
|
|
// peer replicas; the disk state.json is written separately for restart
|
|
// recovery / standalone hydrate.
|
|
if err := s.jobs.Set(ctx, job); err != nil {
|
|
return nil, fmt.Errorf("failed to persist job: %w", err)
|
|
}
|
|
s.saveJobState(job)
|
|
|
|
// Consume the backend's progress stream for the lifetime of the job, not for
|
|
// the lifetime of a client's SSE connection: a job that runs with nobody
|
|
// attached must still reach "completed" in the store and in state.json. The
|
|
// request ctx is done as soon as this HTTP handler returns, so the watcher
|
|
// rides the application context instead.
|
|
go s.watchProgress(s.appConfig.Context, jobID, backendName, modelID)
|
|
|
|
return &schema.QuantizationJobResponse{
|
|
ID: jobID,
|
|
Status: "queued",
|
|
Message: result.Message,
|
|
}, nil
|
|
}
|
|
|
|
// GetJob returns a quantization job by ID.
|
|
func (s *QuantizationService) GetJob(userID, jobID string) (*schema.QuantizationJob, error) {
|
|
s.mu.Lock()
|
|
defer s.mu.Unlock()
|
|
|
|
job, ok := s.jobs.Get(jobID)
|
|
if !ok {
|
|
return nil, fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
if userID != "" && job.UserID != userID {
|
|
return nil, fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
return job, nil
|
|
}
|
|
|
|
// ListJobs returns all jobs for a user, sorted by creation time (newest first).
|
|
func (s *QuantizationService) ListJobs(userID string) []*schema.QuantizationJob {
|
|
s.mu.Lock()
|
|
defer s.mu.Unlock()
|
|
|
|
var result []*schema.QuantizationJob
|
|
for _, job := range s.jobs.List() {
|
|
if userID == "" || job.UserID == userID {
|
|
result = append(result, job)
|
|
}
|
|
}
|
|
|
|
slices.SortFunc(result, func(a, b *schema.QuantizationJob) int {
|
|
return cmp.Compare(b.CreatedAt, a.CreatedAt)
|
|
})
|
|
|
|
return result
|
|
}
|
|
|
|
// StopJob stops a running quantization job.
|
|
func (s *QuantizationService) StopJob(ctx context.Context, userID, jobID string) error {
|
|
s.mu.Lock()
|
|
job, ok := s.jobs.Get(jobID)
|
|
if !ok {
|
|
s.mu.Unlock()
|
|
return fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
if userID != "" && job.UserID != userID {
|
|
s.mu.Unlock()
|
|
return fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
s.mu.Unlock()
|
|
|
|
// Kill the backend process directly
|
|
stopModelID := job.ModelID
|
|
if stopModelID == "" {
|
|
stopModelID = job.Backend + "-quantize"
|
|
}
|
|
s.modelLoader.ShutdownModel(stopModelID)
|
|
|
|
s.mu.Lock()
|
|
job.Status = "stopped"
|
|
job.Message = "Quantization stopped by user"
|
|
if err := s.jobs.Set(ctx, job); err != nil {
|
|
xlog.Warn("Failed to persist stopped job", "job_id", jobID, "error", err)
|
|
}
|
|
s.saveJobState(job)
|
|
s.mu.Unlock()
|
|
|
|
// Release clients attached to the progress stream: the backend process is gone,
|
|
// so the watcher will not see a terminal update to forward.
|
|
s.publishProgress(jobID, &schema.QuantizationProgressEvent{
|
|
JobID: jobID,
|
|
Status: "stopped",
|
|
Message: "Quantization stopped by user",
|
|
})
|
|
|
|
return nil
|
|
}
|
|
|
|
// DeleteJob removes a quantization job and its associated data from disk.
|
|
func (s *QuantizationService) DeleteJob(userID, jobID string) error {
|
|
s.mu.Lock()
|
|
job, ok := s.jobs.Get(jobID)
|
|
if !ok {
|
|
s.mu.Unlock()
|
|
return fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
if userID != "" && job.UserID != userID {
|
|
s.mu.Unlock()
|
|
return fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
|
|
// Reject deletion of actively running jobs
|
|
activeStatuses := map[string]bool{
|
|
"queued": true, "downloading": true, "converting": true, "quantizing": true,
|
|
}
|
|
if activeStatuses[job.Status] {
|
|
s.mu.Unlock()
|
|
return fmt.Errorf("cannot delete job %s: currently %s (stop it first)", jobID, job.Status)
|
|
}
|
|
if job.ImportStatus == "importing" {
|
|
s.mu.Unlock()
|
|
return fmt.Errorf("cannot delete job %s: import in progress", jobID)
|
|
}
|
|
|
|
importModelName := job.ImportModelName
|
|
// Delete write-through removes the DB row (distributed) and broadcasts the
|
|
// removal to peer replicas. DeleteJob has no ctx, so use Background.
|
|
if err := s.jobs.Delete(context.Background(), jobID); err != nil {
|
|
xlog.Warn("Failed to delete job from store", "job_id", jobID, "error", err)
|
|
}
|
|
s.mu.Unlock()
|
|
|
|
// Remove job directory (state.json, output files)
|
|
jobDir := s.jobDir(jobID)
|
|
if err := os.RemoveAll(jobDir); err != nil {
|
|
xlog.Warn("Failed to remove quantization job directory", "job_id", jobID, "path", jobDir, "error", err)
|
|
}
|
|
|
|
// If an imported model exists, clean it up too
|
|
if importModelName != "" {
|
|
modelsPath := s.appConfig.SystemState.Model.ModelsPath
|
|
modelDir := filepath.Join(modelsPath, importModelName)
|
|
configPath := filepath.Join(modelsPath, importModelName+".yaml")
|
|
|
|
if err := os.RemoveAll(modelDir); err != nil {
|
|
xlog.Warn("Failed to remove imported model directory", "path", modelDir, "error", err)
|
|
}
|
|
if err := os.Remove(configPath); err != nil && !os.IsNotExist(err) {
|
|
xlog.Warn("Failed to remove imported model config", "path", configPath, "error", err)
|
|
}
|
|
|
|
// Reload model configs
|
|
if err := s.configLoader.LoadModelConfigsFromPath(modelsPath, s.appConfig.ToConfigLoaderOptions()...); err != nil {
|
|
xlog.Warn("Failed to reload configs after delete", "error", err)
|
|
}
|
|
}
|
|
|
|
xlog.Info("Deleted quantization job", "job_id", jobID)
|
|
return nil
|
|
}
|
|
|
|
// watchProgress is the single reader of a job's backend progress stream. It
|
|
// records every transition on the job — in the cross-replica store and in
|
|
// state.json — and republishes it to the clients attached via StreamProgress.
|
|
//
|
|
// Recording here rather than in StreamProgress is the point: the backend hands
|
|
// each update to one consumer, so while StreamProgress was that consumer a job's
|
|
// state only advanced while somebody was watching it.
|
|
func (s *QuantizationService) watchProgress(ctx context.Context, jobID, backendName, modelID string) {
|
|
backendModel, err := s.modelLoader.Load(
|
|
model.WithBackendString(backendName),
|
|
model.WithModel(backendName),
|
|
model.WithModelID(modelID),
|
|
)
|
|
if err != nil {
|
|
xlog.Warn("Failed to load backend for quantization progress", "job_id", jobID, "error", err)
|
|
return
|
|
}
|
|
|
|
err = backendModel.QuantizationProgress(ctx, &pb.QuantizationProgressRequest{
|
|
JobId: jobID,
|
|
}, func(update *pb.QuantizationProgressUpdate) {
|
|
s.publishProgress(jobID, s.applyProgressUpdate(ctx, jobID, update))
|
|
})
|
|
if err != nil {
|
|
xlog.Warn("Quantization progress stream ended with an error", "job_id", jobID, "error", err)
|
|
}
|
|
|
|
// On shutdown leave the job alone: loadJobsFromDisk already reports jobs that
|
|
// were running at exit as stopped.
|
|
if ctx.Err() != nil {
|
|
return
|
|
}
|
|
|
|
// A stream that ends without a terminal update means the backend is gone and
|
|
// nothing further will arrive. Record that instead of leaving the job in a
|
|
// running state forever — which is the failure this watcher exists to prevent —
|
|
// and release any client still waiting on a terminal event.
|
|
s.mu.Lock()
|
|
j, ok := s.jobs.Get(jobID)
|
|
stale := ok && !isTerminalStatus(j.Status)
|
|
if stale {
|
|
j.Status = "failed"
|
|
if j.Message == "" {
|
|
j.Message = "Backend progress stream ended before the job reported a result"
|
|
}
|
|
if err := s.jobs.Set(ctx, j); err != nil {
|
|
xlog.Warn("Failed to persist orphaned job state", "job_id", jobID, "error", err)
|
|
}
|
|
s.saveJobState(j)
|
|
}
|
|
s.mu.Unlock()
|
|
|
|
if stale {
|
|
s.publishProgress(jobID, &schema.QuantizationProgressEvent{
|
|
JobID: jobID,
|
|
Status: "failed",
|
|
Message: "Backend progress stream ended before the job reported a result",
|
|
})
|
|
}
|
|
}
|
|
|
|
// applyProgressUpdate records a backend progress update on the job and returns
|
|
// the event to hand to subscribers.
|
|
func (s *QuantizationService) applyProgressUpdate(ctx context.Context, jobID string, update *pb.QuantizationProgressUpdate) *schema.QuantizationProgressEvent {
|
|
s.mu.Lock()
|
|
if j, ok := s.jobs.Get(jobID); ok {
|
|
// Don't let progress updates overwrite terminal states
|
|
if !isTerminalStatus(j.Status) {
|
|
j.Status = update.Status
|
|
}
|
|
if update.Message != "" {
|
|
j.Message = update.Message
|
|
}
|
|
if update.OutputFile != "" {
|
|
j.OutputFile = update.OutputFile
|
|
}
|
|
if err := s.jobs.Set(ctx, j); err != nil {
|
|
xlog.Warn("Failed to persist progress update", "job_id", jobID, "error", err)
|
|
}
|
|
s.saveJobState(j)
|
|
}
|
|
s.mu.Unlock()
|
|
|
|
// Convert extra metrics
|
|
extraMetrics := make(map[string]float32, len(update.ExtraMetrics))
|
|
for k, v := range update.ExtraMetrics {
|
|
extraMetrics[k] = v
|
|
}
|
|
|
|
return &schema.QuantizationProgressEvent{
|
|
JobID: update.JobId,
|
|
ProgressPercent: update.ProgressPercent,
|
|
Status: update.Status,
|
|
Message: update.Message,
|
|
OutputFile: update.OutputFile,
|
|
ExtraMetrics: extraMetrics,
|
|
}
|
|
}
|
|
|
|
// subscribeProgress registers a channel to receive a job's progress events.
|
|
func (s *QuantizationService) subscribeProgress(jobID string) chan *schema.QuantizationProgressEvent {
|
|
ch := make(chan *schema.QuantizationProgressEvent, progressSubBuffer)
|
|
s.progressMu.Lock()
|
|
s.progressSubs[jobID] = append(s.progressSubs[jobID], ch)
|
|
s.progressMu.Unlock()
|
|
return ch
|
|
}
|
|
|
|
// unsubscribeProgress removes a channel registered by subscribeProgress. The
|
|
// channel is never closed, so a publish racing with an unsubscribe cannot send
|
|
// on a closed channel.
|
|
func (s *QuantizationService) unsubscribeProgress(jobID string, ch chan *schema.QuantizationProgressEvent) {
|
|
s.progressMu.Lock()
|
|
defer s.progressMu.Unlock()
|
|
|
|
subs := s.progressSubs[jobID]
|
|
for i, c := range subs {
|
|
if c == ch {
|
|
s.progressSubs[jobID] = append(subs[:i], subs[i+1:]...)
|
|
break
|
|
}
|
|
}
|
|
if len(s.progressSubs[jobID]) == 0 {
|
|
delete(s.progressSubs, jobID)
|
|
}
|
|
}
|
|
|
|
// publishProgress fans an event out to a job's subscribers.
|
|
func (s *QuantizationService) publishProgress(jobID string, event *schema.QuantizationProgressEvent) {
|
|
s.progressMu.Lock()
|
|
subs := append([]chan *schema.QuantizationProgressEvent(nil), s.progressSubs[jobID]...)
|
|
s.progressMu.Unlock()
|
|
|
|
for _, ch := range subs {
|
|
select {
|
|
case ch <- event:
|
|
default:
|
|
// A subscriber that cannot keep up must not stall the reader that is
|
|
// recording job state for everyone else.
|
|
xlog.Warn("Dropping quantization progress event for a slow subscriber", "job_id", jobID)
|
|
}
|
|
}
|
|
}
|
|
|
|
// StreamProgress calls the callback for each progress event of a job until it
|
|
// reaches a terminal status or ctx is done. It is a pure reader: the job's own
|
|
// watcher owns the backend stream and the state transitions.
|
|
func (s *QuantizationService) StreamProgress(ctx context.Context, userID, jobID string, callback func(event *schema.QuantizationProgressEvent)) error {
|
|
s.mu.Lock()
|
|
job, ok := s.jobs.Get(jobID)
|
|
if !ok {
|
|
s.mu.Unlock()
|
|
return fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
if userID != "" && job.UserID != userID {
|
|
s.mu.Unlock()
|
|
return fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
s.mu.Unlock()
|
|
|
|
ch := s.subscribeProgress(jobID)
|
|
defer s.unsubscribeProgress(jobID, ch)
|
|
|
|
// Re-read the job after subscribing: it may have finished between the lookup
|
|
// above and the subscription, and no further event would ever arrive. Jobs
|
|
// restored from disk after a restart are terminal too, and have no watcher.
|
|
s.mu.Lock()
|
|
current, ok := s.jobs.Get(jobID)
|
|
terminal := ok && isTerminalStatus(current.Status)
|
|
var final *schema.QuantizationProgressEvent
|
|
if terminal {
|
|
final = &schema.QuantizationProgressEvent{
|
|
JobID: current.ID,
|
|
Status: current.Status,
|
|
Message: current.Message,
|
|
OutputFile: current.OutputFile,
|
|
}
|
|
}
|
|
s.mu.Unlock()
|
|
if terminal {
|
|
callback(final)
|
|
return nil
|
|
}
|
|
|
|
for {
|
|
select {
|
|
case <-ctx.Done():
|
|
return ctx.Err()
|
|
case event := <-ch:
|
|
callback(event)
|
|
if isTerminalStatus(event.Status) {
|
|
return nil
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
// sanitizeQuantModelName replaces non-alphanumeric characters with hyphens and lowercases.
|
|
func sanitizeQuantModelName(s string) string {
|
|
re := regexp.MustCompile(`[^a-zA-Z0-9\-]`)
|
|
s = re.ReplaceAllString(s, "-")
|
|
s = regexp.MustCompile(`-+`).ReplaceAllString(s, "-")
|
|
s = strings.Trim(s, "-")
|
|
return strings.ToLower(s)
|
|
}
|
|
|
|
// inferenceBackendFor returns the backend that can load what a quantization
|
|
// backend produced.
|
|
//
|
|
// The gallery publishes a quantizer as a release channel of the engine that
|
|
// runs its output: "llama-cpp-quantization" is llama.cpp's quantizer, and the
|
|
// GGUF it writes is served by "llama-cpp". The suffix is a channel marker and
|
|
// carries no engine information, so stripping it yields the backend to pin in
|
|
// the imported model's config. Names that carry no channel suffix (a backend
|
|
// that both quantizes and serves, such as "rocmfp4") are already the engine
|
|
// name and pass through unchanged, as do pinned hardware variants
|
|
// ("rocm-rocmfp4"), which are valid values for a config's `backend:`.
|
|
func inferenceBackendFor(quantBackend string) string {
|
|
return strings.TrimSuffix(config.NormalizeBackendName(quantBackend), "-quantization")
|
|
}
|
|
|
|
// ImportModel imports a quantized model into LocalAI asynchronously.
|
|
func (s *QuantizationService) ImportModel(ctx context.Context, userID, jobID string, req schema.QuantizationImportRequest) (string, error) {
|
|
s.mu.Lock()
|
|
job, ok := s.jobs.Get(jobID)
|
|
if !ok {
|
|
s.mu.Unlock()
|
|
return "", fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
if userID != "" && job.UserID != userID {
|
|
s.mu.Unlock()
|
|
return "", fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
if job.Status != "completed" {
|
|
s.mu.Unlock()
|
|
return "", fmt.Errorf("job %s is not completed (status: %s)", jobID, job.Status)
|
|
}
|
|
if job.ImportStatus == "importing" {
|
|
s.mu.Unlock()
|
|
return "", fmt.Errorf("import already in progress for job %s", jobID)
|
|
}
|
|
if job.OutputFile == "" {
|
|
s.mu.Unlock()
|
|
return "", fmt.Errorf("no output file for job %s", jobID)
|
|
}
|
|
s.mu.Unlock()
|
|
|
|
// Compute model name
|
|
modelName := req.Name
|
|
if modelName == "" {
|
|
base := sanitizeQuantModelName(job.Model)
|
|
if base == "" {
|
|
base = "model"
|
|
}
|
|
shortID := jobID
|
|
if len(shortID) > 8 {
|
|
shortID = shortID[:8]
|
|
}
|
|
modelName = base + "-" + job.QuantizationType + "-" + shortID
|
|
}
|
|
|
|
// Compute output path in models directory
|
|
modelsPath := s.appConfig.SystemState.Model.ModelsPath
|
|
outputPath := filepath.Join(modelsPath, modelName)
|
|
|
|
// Check for name collision
|
|
configPath := filepath.Join(modelsPath, modelName+".yaml")
|
|
if err := utils.VerifyPath(modelName+".yaml", modelsPath); err != nil {
|
|
return "", fmt.Errorf("invalid model name: %w", err)
|
|
}
|
|
if _, err := os.Stat(configPath); err == nil {
|
|
return "", fmt.Errorf("model %q already exists, choose a different name", modelName)
|
|
}
|
|
|
|
// Create output directory
|
|
if err := os.MkdirAll(outputPath, 0750); err != nil {
|
|
return "", fmt.Errorf("failed to create output directory: %w", err)
|
|
}
|
|
|
|
// Set import status
|
|
s.mu.Lock()
|
|
job.ImportStatus = "importing"
|
|
job.ImportMessage = ""
|
|
job.ImportModelName = ""
|
|
if err := s.jobs.Set(ctx, job); err != nil {
|
|
xlog.Warn("Failed to persist import start", "job_id", jobID, "error", err)
|
|
}
|
|
s.saveJobState(job)
|
|
s.mu.Unlock()
|
|
|
|
// Launch the import in a background goroutine
|
|
go func() {
|
|
s.setImportMessage(job, "Copying quantized model...")
|
|
|
|
// Copy the output file to the models directory
|
|
srcFile := job.OutputFile
|
|
dstFile := filepath.Join(outputPath, filepath.Base(srcFile))
|
|
|
|
srcData, err := os.ReadFile(srcFile)
|
|
if err != nil {
|
|
s.setImportFailed(job, fmt.Sprintf("failed to read output file: %v", err))
|
|
return
|
|
}
|
|
if err := os.WriteFile(dstFile, srcData, 0644); err != nil {
|
|
s.setImportFailed(job, fmt.Sprintf("failed to write model file: %v", err))
|
|
return
|
|
}
|
|
|
|
s.setImportMessage(job, "Generating model configuration...")
|
|
|
|
// Auto-import: detect format and generate config
|
|
cfg, err := importers.ImportLocalPath(outputPath, modelName)
|
|
if err != nil {
|
|
s.setImportFailed(job, fmt.Sprintf("model copied to %s but config generation failed: %v", outputPath, err))
|
|
return
|
|
}
|
|
|
|
cfg.Name = modelName
|
|
|
|
// The importer detects the file format and defaults to llama-cpp for any
|
|
// GGUF. That is wrong for a model this service just quantized with a
|
|
// backend stock llama.cpp cannot read: the job knows which backend
|
|
// produced the file, so pin that one instead of the detected default.
|
|
if backend := inferenceBackendFor(job.Backend); backend != "" {
|
|
cfg.Backend = backend
|
|
}
|
|
if job.QuantizationType != "" {
|
|
cfg.Description = "Quantized model (" + job.QuantizationType + ", GGUF)"
|
|
}
|
|
|
|
// Write YAML config
|
|
yamlData, err := yaml.Marshal(cfg)
|
|
if err != nil {
|
|
s.setImportFailed(job, fmt.Sprintf("failed to marshal config: %v", err))
|
|
return
|
|
}
|
|
if err := os.WriteFile(configPath, yamlData, 0644); err != nil {
|
|
s.setImportFailed(job, fmt.Sprintf("failed to write config file: %v", err))
|
|
return
|
|
}
|
|
|
|
s.setImportMessage(job, "Registering model with LocalAI...")
|
|
|
|
// Reload configs so the model is immediately available
|
|
if err := s.configLoader.LoadModelConfigsFromPath(modelsPath, s.appConfig.ToConfigLoaderOptions()...); err != nil {
|
|
xlog.Warn("Failed to reload configs after import", "error", err)
|
|
}
|
|
if err := s.configLoader.Preload(modelsPath); err != nil {
|
|
xlog.Warn("Failed to preload after import", "error", err)
|
|
}
|
|
|
|
xlog.Info("Quantized model imported and registered", "job_id", jobID, "model_name", modelName)
|
|
|
|
// Runs after the HTTP request returns, so use Background rather than the
|
|
// (now likely cancelled) request ctx for the write-through.
|
|
s.mu.Lock()
|
|
job.ImportStatus = "completed"
|
|
job.ImportModelName = modelName
|
|
job.ImportMessage = ""
|
|
if err := s.jobs.Set(context.Background(), job); err != nil {
|
|
xlog.Warn("Failed to persist import completion", "job_id", jobID, "error", err)
|
|
}
|
|
s.saveJobState(job)
|
|
s.mu.Unlock()
|
|
}()
|
|
|
|
return modelName, nil
|
|
}
|
|
|
|
// setImportMessage updates the import message and persists the job state. Called
|
|
// from the background import goroutine, so it uses Background for write-through.
|
|
func (s *QuantizationService) setImportMessage(job *schema.QuantizationJob, msg string) {
|
|
s.mu.Lock()
|
|
job.ImportMessage = msg
|
|
if err := s.jobs.Set(context.Background(), job); err != nil {
|
|
xlog.Warn("Failed to persist import message", "job_id", job.ID, "error", err)
|
|
}
|
|
s.saveJobState(job)
|
|
s.mu.Unlock()
|
|
}
|
|
|
|
// setImportFailed sets the import status to failed with a message.
|
|
func (s *QuantizationService) setImportFailed(job *schema.QuantizationJob, message string) {
|
|
xlog.Error("Quantization import failed", "job_id", job.ID, "error", message)
|
|
s.mu.Lock()
|
|
job.ImportStatus = "failed"
|
|
job.ImportMessage = message
|
|
if err := s.jobs.Set(context.Background(), job); err != nil {
|
|
xlog.Warn("Failed to persist import failure", "job_id", job.ID, "error", err)
|
|
}
|
|
s.saveJobState(job)
|
|
s.mu.Unlock()
|
|
}
|
|
|
|
// GetOutputPath returns the path to the quantized model file and a download name.
|
|
func (s *QuantizationService) GetOutputPath(userID, jobID string) (string, string, error) {
|
|
s.mu.Lock()
|
|
job, ok := s.jobs.Get(jobID)
|
|
if !ok {
|
|
s.mu.Unlock()
|
|
return "", "", fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
if userID != "" && job.UserID != userID {
|
|
s.mu.Unlock()
|
|
return "", "", fmt.Errorf("job not found: %s", jobID)
|
|
}
|
|
if job.Status != "completed" {
|
|
s.mu.Unlock()
|
|
return "", "", fmt.Errorf("job not completed (status: %s)", job.Status)
|
|
}
|
|
if job.OutputFile == "" {
|
|
s.mu.Unlock()
|
|
return "", "", fmt.Errorf("no output file for job %s", jobID)
|
|
}
|
|
outputFile := job.OutputFile
|
|
s.mu.Unlock()
|
|
|
|
if _, err := os.Stat(outputFile); os.IsNotExist(err) {
|
|
return "", "", fmt.Errorf("output file not found: %s", outputFile)
|
|
}
|
|
|
|
downloadName := filepath.Base(outputFile)
|
|
return outputFile, downloadName, nil
|
|
}
|