mirror of
https://github.com/mudler/LocalAI.git
synced 2026-08-01 11:00:24 -04:00
* feat(backend): add depth-anything (Depth Anything 3) C++/ggml backend + gallery Mirrors the locate-anything-cpp backend to register a new depth-anything backend that wraps the Depth Anything 3 ggml port (depth-anything.cpp) via purego (cgo-less, no Python at inference). - backend/go/depth-anything-cpp/: gRPC backend (Load + Predict + GenerateImage), purego binding to the da_capi_* C ABI, CMake/Makefile/run/package/test scripts building depth-anything.cpp's DA_SHARED static .so per CPU variant. - backend/index.yaml: depth-anything backend meta + all hardware-variant capability entries (cpu/cuda12/cuda13/intel-sycl-f32+f16/vulkan/nvidia-l4t). - gallery/index.yaml: 8 Depth Anything 3 GGUF models (base q4_k/q8_0/f16/f32, small, large, giant, mono-large). - .github/backend-matrix.yml: one build entry per hardware variant. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(depth): typed Depth RPC + REST endpoint exposing full DA3 data Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(depth): pin depth-anything.cpp to e0b6814 (ABI 3 dense C-API) The Depth RPC handler calls da_capi_depth_dense / da_capi_points (C-API ABI 3); pin the native build to the commit that exports them. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(depth): pin depth-anything.cpp to v0.1.0 release (b515c31) Repoint the native version from the now-orphaned e0b6814 to the b515c31 release commit, kept alive by the upstream v0.1.0 tag. C-API is unchanged (da_capi_abi_version == 3). Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(depth): wire depth-anything-cpp into build, CI bump, and importer The backend dir, gallery index, and CI build-matrix were present but the backend was never wired into the integration points that adding-backends.md requires: - root Makefile: add to .NOTPARALLEL, the test-extra chain, a BACKEND_* definition, the docker-build target eval, and docker-build-backends (mirrors parakeet-cpp; the backend's own Makefile already documented that its `test` target is driven by test-extra). - bump_deps.yaml: register the DEPTHANYTHING_VERSION pin so the daily auto-bump bot tracks mudler/depth-anything.cpp master (it cannot see an unregistered Makefile pin). - import form: add a preference-only KnownBackend entry so depth-anything is selectable at /import-model (mirrors sam3-cpp; no reliable GGUF auto-detect signal, so pref-only per the doc's default). changed-backends.js needs no entry: the generic golang suffix branch already resolves backend/go/depth-anything-cpp/. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(depth): auto-detect importer for depth-anything GGUFs Replace the preference-only entry with a real auto-detect importer (mirrors parakeet-cpp / locate-anything): - DepthAnythingImporter matches a .gguf whose name carries a depth-anything token (depth-anything-<size>-<quant>.gguf), so /import-model recognises mudler/depth-anything.cpp-gguf repos and direct GGUF URLs without an explicit backend preference. preferences.backend= "depth-anything" still forces it. - Registered before LlamaCPPImporter so its GGUF bundles aren't claimed by the generic .gguf importer; the narrow name match means it cannot claim arbitrary llama GGUFs or the upstream safetensors PyTorch repos. - Multi-quant repos pick the smallest quant by default (q4_k -> ... -> f32, depth stays >0.998 corr even at q4_k); quantizations preference overrides. - Drops the now-redundant knownPrefOnlyBackends entry (importer-backed backends are not listed there, matching parakeet-cpp). - Table-driven Ginkgo test covers detection, negative cases (llama GGUF, upstream safetensors), default/override/fallback quant pick, and direct URL import. 10/10 specs pass. Assisted-by: Claude:claude-opus-4-8 Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(depth): check conn.Close error in grpc Depth client (errcheck) The new Depth() client method used a bare `defer conn.Close()`. golangci-lint runs with new-from-merge-base, so although the 39 sibling methods use the same bare form (grandfathered), the newly added line trips errcheck. Drop the result explicitly to satisfy the linter. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-4-8 * fix(depth): bump depth-anything.cpp to v0.1.1 (embeddable CMake) v0.1.0 (b515c31) used ${CMAKE_SOURCE_DIR} for its include dirs, which points at the parent project when built via add_subdirectory() as this backend does, so the container build failed with missing stb_image.h / da_gguf_keys.h. v0.1.1 (2d42897) switches to project-relative paths. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-4-8 * fix(depth): resolve gosec findings in the backend wrapper The code-scanning gate flagged three new failure-level alerts in godepthanythingcpp.go (gosec runs with -no-fail; GitHub gates on new alerts): - G301: export dirs were created with 0o755. Tighten to 0o750 (no world access needed for backend-written export output). - G304: writeDepthPNG creates req.GetDst(). That path is chosen by the LocalAI core as the intended output destination (same pattern every image backend uses), not attacker input, so annotate with #nosec G304 and document why. The remaining G103 "audit unsafe" notes on the unsafe.Slice C-buffer copies are warning-level (the same purego interop whisper/parakeet use) and do not gate the check, per the supertonic exclusion precedent in secscan.yaml. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-4-8 * fix(depth): bump depth-anything.cpp to v0.1.2 (CUDA cross-build arch) v0.1.1 forced CMAKE_CUDA_ARCHITECTURES=native, which breaks the GPU-less l4t/cublas CI builds (nvcc "Unsupported gpu architecture 'compute_'" on CMake 3.22). v0.1.2 (442eea4) drops the override and lets ggml pick its default cross-build arch list. Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Assisted-by: Claude:claude-opus-4-8 --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
245 lines
10 KiB
Go
245 lines
10 KiB
Go
package nodes
|
|
|
|
import (
|
|
"context"
|
|
"sync"
|
|
"time"
|
|
|
|
"github.com/mudler/LocalAI/pkg/grpc"
|
|
"github.com/mudler/LocalAI/pkg/grpc/grpcerrors"
|
|
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
|
|
"github.com/mudler/xlog"
|
|
ggrpc "google.golang.org/grpc"
|
|
)
|
|
|
|
// InFlightTrackingClient wraps a grpc.Backend and tracks active inference requests
|
|
// in the NodeRegistry. This allows the router's eviction logic to know which models
|
|
// are actively serving and should not be unloaded.
|
|
//
|
|
// Per-replica: a single tracker instance is bound to (nodeID, modelName, replicaIndex).
|
|
// The router constructs one tracker per Route() result, so each in-flight tick lands
|
|
// on the correct row even when multiple replicas of the same model live on the same node.
|
|
type InFlightTrackingClient struct {
|
|
grpc.Backend // embed for passthrough of untracked methods
|
|
registry InFlightTracker
|
|
nodeID string
|
|
modelName string
|
|
replicaIndex int
|
|
|
|
firstOnce sync.Once // guards onFirstComplete
|
|
onFirstComplete func() // called once after the first tracked inference call completes
|
|
}
|
|
|
|
// NewInFlightTrackingClient wraps a gRPC backend client with in-flight tracking.
|
|
func NewInFlightTrackingClient(inner grpc.Backend, registry InFlightTracker, nodeID, modelName string, replicaIndex int) *InFlightTrackingClient {
|
|
return &InFlightTrackingClient{
|
|
Backend: inner,
|
|
registry: registry,
|
|
nodeID: nodeID,
|
|
modelName: modelName,
|
|
replicaIndex: replicaIndex,
|
|
}
|
|
}
|
|
|
|
// OnFirstComplete registers a callback that fires once after the first tracked
|
|
// inference call completes. This is used to release the initial in-flight
|
|
// reservation (set during model load) after the triggering request finishes,
|
|
// so that in-flight returns to 0 when the model is idle.
|
|
func (c *InFlightTrackingClient) OnFirstComplete(fn func()) {
|
|
c.onFirstComplete = fn
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) track(ctx context.Context) func() {
|
|
if err := c.registry.IncrementInFlight(ctx, c.nodeID, c.modelName, c.replicaIndex); err != nil {
|
|
xlog.Warn("Failed to increment in-flight counter", "node", c.nodeID, "model", c.modelName, "replica", c.replicaIndex, "error", err)
|
|
return func() {}
|
|
}
|
|
return func() {
|
|
decCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
|
|
defer cancel()
|
|
c.registry.DecrementInFlight(decCtx, c.nodeID, c.modelName, c.replicaIndex)
|
|
// Release the initial reservation after the first inference call completes
|
|
if c.onFirstComplete != nil {
|
|
c.firstOnce.Do(c.onFirstComplete)
|
|
}
|
|
}
|
|
}
|
|
|
|
// reconcile self-heals stale routing: when a backend reports that the model is
|
|
// no longer loaded (the process survived but the model was evicted, while the
|
|
// registry still lists it as loaded), it drops the replica row so the next
|
|
// request triggers a fresh load instead of routing back here. Without this the
|
|
// model stays unreachable until the controller restarts. The original error is
|
|
// returned unchanged.
|
|
func (c *InFlightTrackingClient) reconcile(err error) error {
|
|
if !grpcerrors.IsModelNotLoaded(err) {
|
|
return err
|
|
}
|
|
rmCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
|
|
defer cancel()
|
|
if rmErr := c.registry.RemoveNodeModel(rmCtx, c.nodeID, c.modelName, c.replicaIndex); rmErr != nil {
|
|
xlog.Warn("Failed to drop stale replica after model-not-loaded",
|
|
"node", c.nodeID, "model", c.modelName, "replica", c.replicaIndex, "error", rmErr)
|
|
} else {
|
|
xlog.Warn("Backend reports model not loaded; dropped stale replica so the next request reloads",
|
|
"node", c.nodeID, "model", c.modelName, "replica", c.replicaIndex)
|
|
}
|
|
return err
|
|
}
|
|
|
|
// --- Tracked inference methods ---
|
|
|
|
func (c *InFlightTrackingClient) Predict(ctx context.Context, in *pb.PredictOptions, opts ...ggrpc.CallOption) (*pb.Reply, error) {
|
|
defer c.track(ctx)()
|
|
reply, err := c.Backend.Predict(ctx, in, opts...)
|
|
return reply, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) PredictStream(ctx context.Context, in *pb.PredictOptions, f func(reply *pb.Reply), opts ...ggrpc.CallOption) error {
|
|
defer c.track(ctx)()
|
|
return c.reconcile(c.Backend.PredictStream(ctx, in, f, opts...))
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) Embeddings(ctx context.Context, in *pb.PredictOptions, opts ...ggrpc.CallOption) (*pb.EmbeddingResult, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.Embeddings(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) GenerateImage(ctx context.Context, in *pb.GenerateImageRequest, opts ...ggrpc.CallOption) (*pb.Result, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.GenerateImage(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) GenerateVideo(ctx context.Context, in *pb.GenerateVideoRequest, opts ...ggrpc.CallOption) (*pb.Result, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.GenerateVideo(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) TTS(ctx context.Context, in *pb.TTSRequest, opts ...ggrpc.CallOption) (*pb.Result, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.TTS(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) TTSStream(ctx context.Context, in *pb.TTSRequest, f func(reply *pb.Reply), opts ...ggrpc.CallOption) error {
|
|
defer c.track(ctx)()
|
|
return c.reconcile(c.Backend.TTSStream(ctx, in, f, opts...))
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) SoundGeneration(ctx context.Context, in *pb.SoundGenerationRequest, opts ...ggrpc.CallOption) (*pb.Result, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.SoundGeneration(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) AudioTranscription(ctx context.Context, in *pb.TranscriptRequest, opts ...ggrpc.CallOption) (*pb.TranscriptResult, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.AudioTranscription(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) AudioTranscriptionStream(ctx context.Context, in *pb.TranscriptRequest, f func(chunk *pb.TranscriptStreamResponse), opts ...ggrpc.CallOption) error {
|
|
defer c.track(ctx)()
|
|
return c.reconcile(c.Backend.AudioTranscriptionStream(ctx, in, f, opts...))
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) Detect(ctx context.Context, in *pb.DetectOptions, opts ...ggrpc.CallOption) (*pb.DetectResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.Detect(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) Depth(ctx context.Context, in *pb.DepthRequest, opts ...ggrpc.CallOption) (*pb.DepthResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.Depth(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) Rerank(ctx context.Context, in *pb.RerankRequest, opts ...ggrpc.CallOption) (*pb.RerankResult, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.Rerank(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) VAD(ctx context.Context, in *pb.VADRequest, opts ...ggrpc.CallOption) (*pb.VADResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.VAD(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) Diarize(ctx context.Context, in *pb.DiarizeRequest, opts ...ggrpc.CallOption) (*pb.DiarizeResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.Diarize(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) FaceVerify(ctx context.Context, in *pb.FaceVerifyRequest, opts ...ggrpc.CallOption) (*pb.FaceVerifyResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.FaceVerify(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) FaceAnalyze(ctx context.Context, in *pb.FaceAnalyzeRequest, opts ...ggrpc.CallOption) (*pb.FaceAnalyzeResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.FaceAnalyze(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) VoiceVerify(ctx context.Context, in *pb.VoiceVerifyRequest, opts ...ggrpc.CallOption) (*pb.VoiceVerifyResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.VoiceVerify(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) VoiceAnalyze(ctx context.Context, in *pb.VoiceAnalyzeRequest, opts ...ggrpc.CallOption) (*pb.VoiceAnalyzeResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.VoiceAnalyze(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) VoiceEmbed(ctx context.Context, in *pb.VoiceEmbedRequest, opts ...ggrpc.CallOption) (*pb.VoiceEmbedResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.VoiceEmbed(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) TokenClassify(ctx context.Context, in *pb.TokenClassifyRequest, opts ...ggrpc.CallOption) (*pb.TokenClassifyResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.TokenClassify(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) Score(ctx context.Context, in *pb.ScoreRequest, opts ...ggrpc.CallOption) (*pb.ScoreResponse, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.Score(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) AudioEncode(ctx context.Context, in *pb.AudioEncodeRequest, opts ...ggrpc.CallOption) (*pb.AudioEncodeResult, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.AudioEncode(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) AudioDecode(ctx context.Context, in *pb.AudioDecodeRequest, opts ...ggrpc.CallOption) (*pb.AudioDecodeResult, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.AudioDecode(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
func (c *InFlightTrackingClient) AudioTransform(ctx context.Context, in *pb.AudioTransformRequest, opts ...ggrpc.CallOption) (*pb.AudioTransformResult, error) {
|
|
defer c.track(ctx)()
|
|
res, err := c.Backend.AudioTransform(ctx, in, opts...)
|
|
return res, c.reconcile(err)
|
|
}
|
|
|
|
// AudioTransformStream, AudioToAudioStream and Forward are deliberately left as
|
|
// embedded passthrough: they return a stream client and the inference spans the
|
|
// stream's lifetime, not the constructor call. Wrapping the constructor with
|
|
// track() would increment and immediately decrement (and fire onFirstComplete)
|
|
// before any audio flows. Tracking those correctly needs the done() func tied to
|
|
// stream close, which the current Backend interface doesn't surface here.
|