mirror of
https://github.com/mudler/LocalAI.git
synced 2026-07-30 18:09:05 -04:00
backend(audio-cpp): correct the status, lifetime and state contracts of LoadedModel
An environment fault during session creation was reported as UNIMPLEMENTED. A missing libggml-cpu-*.so surfaced to the client as 'family silero_vad advertises vad/offline but refused to create the session: Failed to initialize CPU backend', which tells LocalAI the model cannot do this and must never be retried, and sends an operator hunting a capability bug instead of a packaging one. A throw from create_task_session is now a plain runtime_error, so it maps to INTERNAL. Only a null return, where the family genuinely declined, stays a CapabilityError. The model.'s task: option was parsed and then dropped: it lived in a local that died at the end of LoadModel and had no route to RequestShape::pinned_task. LoadedModel now keeps it and exposes pinned_task(). The global model becomes a shared_ptr reached through snapshot(). An audio RPC runs for seconds and cannot hold g_model_mu for its duration, so under a unique_ptr a Free arriving mid-request would destroy the model underneath it. Handlers now take a counted reference and whichever finishes last does the teardown, outside the lock. session_for documents the streaming state contract rather than resetting the session itself. Resetting on a cache hit was tried first and is not possible: silero_vad throws 'session prepare() must be called before Silero VAD reset()', so it would turn an ordinary second fetch into a hard error. start_stream's base implementation is already a reset, so a caller that runs prepare then start_stream per stream gets a clean session; a probe against the bundled silero_vad confirms an identical replay when it does and a carried-over stream when it does not. Also: an unknown backend: name is rejected before the model loads rather than after; MainGPU is parsed instead of passed through std::atoi, which turned 'gpu1' into device 0 silently; and device carries a device_set flag, because 0 is both the default and a real device index, so MainGPU was overriding an explicit device:0 that the neighbouring threads: handling promises will win. Assisted-by: Claude:claude-opus-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
committed by
localai-org-maint-bot
parent
b4153fc6eb
commit
e2cad2bf3d
@@ -5,6 +5,12 @@
|
||||
namespace audiocpp_backend {
|
||||
namespace {
|
||||
|
||||
// PRECONDITION: `suffix` must already be lowercase. Both sides are folded, so
|
||||
// this reads as symmetric, but only `value` can carry case in practice and a
|
||||
// caller passing ".GGUF" would still work today for that reason alone. Do not
|
||||
// rely on it: the fold on the suffix side is the only thing standing between
|
||||
// this and a helper that answers false for every input, and it is not covered
|
||||
// by any test, because with a lowercase suffix no input can distinguish it.
|
||||
bool ends_with_ci(const std::string &value, const std::string &suffix) {
|
||||
if (value.size() <= suffix.size()) {
|
||||
return false; // a bare ".gguf" is an extension, not a model file
|
||||
|
||||
@@ -27,6 +27,7 @@
|
||||
#include <cstdlib>
|
||||
#include <exception>
|
||||
#include <iostream>
|
||||
#include <limits>
|
||||
#include <memory>
|
||||
#include <mutex>
|
||||
#include <string>
|
||||
@@ -52,8 +53,47 @@ std::atomic<bool> g_shutdown_requested{false};
|
||||
// the pointer itself, not the model: swapping or dropping it races with the
|
||||
// handlers that read it, whereas the model's own concurrency is the inference
|
||||
// lane's job.
|
||||
//
|
||||
// shared_ptr, not unique_ptr. An audio RPC runs for seconds and cannot hold
|
||||
// g_model_mu for its duration, so it has to work from a reference taken under
|
||||
// the lock and used after releasing it. Under a unique_ptr, a Free or a reload
|
||||
// arriving mid-request destroys the model out from under that reference. Every
|
||||
// handler instead takes a counted reference through snapshot(), so Free drops
|
||||
// the global's reference and whichever request finishes last destroys the model.
|
||||
std::mutex g_model_mu;
|
||||
std::unique_ptr<audiocpp_backend::LoadedModel> g_model;
|
||||
std::shared_ptr<audiocpp_backend::LoadedModel> g_model;
|
||||
|
||||
// The only correct way for a handler to reach the model. Returns by value, so
|
||||
// the caller owns a reference for as long as its local lives, and null when no
|
||||
// model is loaded. Never return LoadedModel& from here: that is the shape that
|
||||
// reintroduces the use-after-free.
|
||||
std::shared_ptr<audiocpp_backend::LoadedModel> snapshot() {
|
||||
std::lock_guard<std::mutex> lock(g_model_mu);
|
||||
return g_model;
|
||||
}
|
||||
|
||||
// Parses ModelOptions.MainGPU into a device index.
|
||||
//
|
||||
// Not std::atoi: it returns 0 for anything unparseable, so "gpu1" or a device
|
||||
// UUID would silently become device 0 and the model would load on the wrong
|
||||
// device with no diagnostic anywhere. A refusal the operator can read beats a
|
||||
// wrong answer they cannot see.
|
||||
int parse_device_index(const std::string &value) {
|
||||
size_t consumed = 0;
|
||||
long parsed = 0;
|
||||
try {
|
||||
parsed = std::stol(value, &consumed);
|
||||
} catch (const std::exception &) {
|
||||
consumed = 0;
|
||||
}
|
||||
if (consumed != value.size() || parsed < 0 ||
|
||||
parsed > std::numeric_limits<int>::max()) {
|
||||
throw audiocpp_backend::ConfigError(
|
||||
"audio-cpp: main_gpu must be a non-negative device index, got '" +
|
||||
value + "'");
|
||||
}
|
||||
return static_cast<int>(parsed);
|
||||
}
|
||||
|
||||
// Maps a thrown exception onto the gRPC status the client should see.
|
||||
GStatus to_status(const std::exception &err) {
|
||||
@@ -82,9 +122,8 @@ public:
|
||||
|
||||
GStatus Status(ServerContext *, const backend::HealthMessage *,
|
||||
backend::StatusResponse *response) override {
|
||||
std::lock_guard<std::mutex> lock(g_model_mu);
|
||||
response->set_state(g_model ? backend::StatusResponse::READY
|
||||
: backend::StatusResponse::UNINITIALIZED);
|
||||
response->set_state(snapshot() ? backend::StatusResponse::READY
|
||||
: backend::StatusResponse::UNINITIALIZED);
|
||||
return GStatus::OK;
|
||||
}
|
||||
|
||||
@@ -102,15 +141,18 @@ public:
|
||||
if (parsed.options.threads == 0 && request->threads() > 0) {
|
||||
parsed.options.threads = static_cast<int>(request->threads());
|
||||
}
|
||||
// MainGPU carries the device index for GPU backends.
|
||||
if (parsed.options.device == 0 && !request->maingpu().empty()) {
|
||||
parsed.options.device = std::atoi(request->maingpu().c_str());
|
||||
// MainGPU carries the device index for GPU backends, and is the
|
||||
// fallback: an explicit device: option wins, matching threads above.
|
||||
// device_set rather than `device == 0`, because 0 is a real device
|
||||
// index and the value alone cannot say whether anyone chose it.
|
||||
if (!parsed.options.device_set && !request->maingpu().empty()) {
|
||||
parsed.options.device = parse_device_index(request->maingpu());
|
||||
}
|
||||
|
||||
const std::string path = audiocpp_backend::resolve_model_path(
|
||||
request->modelpath(), request->modelfile(), request->model());
|
||||
|
||||
auto loaded = std::make_unique<audiocpp_backend::LoadedModel>(
|
||||
auto loaded = std::make_shared<audiocpp_backend::LoadedModel>(
|
||||
path, parsed.options);
|
||||
|
||||
// Read everything the reply and the log line need while this scope
|
||||
@@ -120,10 +162,15 @@ public:
|
||||
const std::string variant = loaded->variant();
|
||||
const std::string capabilities =
|
||||
audiocpp_backend::describe_capabilities(loaded->capabilities());
|
||||
std::shared_ptr<audiocpp_backend::LoadedModel> replaced;
|
||||
{
|
||||
std::lock_guard<std::mutex> lock(g_model_mu);
|
||||
replaced = std::move(g_model);
|
||||
g_model = std::move(loaded);
|
||||
}
|
||||
// The previous model, if any, is dropped outside the lock, for the
|
||||
// same reason Free does it: teardown must not hold up Status.
|
||||
replaced.reset();
|
||||
|
||||
// Logged once at load time so an operator can see what the model
|
||||
// can do without making a request. ModelMetadata is deliberately
|
||||
@@ -147,9 +194,19 @@ public:
|
||||
|
||||
GStatus Free(ServerContext *, const backend::HealthMessage *,
|
||||
backend::Result *result) override {
|
||||
std::lock_guard<std::mutex> lock(g_model_mu);
|
||||
// Destroying LoadedModel releases its sessions first, then the model.
|
||||
g_model.reset();
|
||||
// Drops the global's reference. Any handler still running holds its own
|
||||
// from snapshot(), so the model is destroyed by whichever of them
|
||||
// finishes last rather than underneath one of them. Destroying
|
||||
// LoadedModel releases its sessions first, then the model.
|
||||
std::shared_ptr<audiocpp_backend::LoadedModel> released;
|
||||
{
|
||||
std::lock_guard<std::mutex> lock(g_model_mu);
|
||||
released = std::move(g_model);
|
||||
}
|
||||
// Released outside the lock: if this is the last reference, the model
|
||||
// and its sessions are torn down here, and that must not block Status
|
||||
// or a fresh LoadModel behind g_model_mu.
|
||||
released.reset();
|
||||
result->set_success(true);
|
||||
return GStatus::OK;
|
||||
}
|
||||
|
||||
@@ -286,6 +286,22 @@ LoadedModel::LoadedModel(const std::string &resolved_path,
|
||||
throw ConfigError("audio-cpp: unknown audio.cpp family '" + family + "'");
|
||||
}
|
||||
|
||||
// Session options are built BEFORE the load, because parse_backend_type
|
||||
// rejects an unknown backend name. Validating after the load would make
|
||||
// `backend:cudaa` cost a full model load, on a fault a string comparison
|
||||
// could have caught.
|
||||
session_options_.backend.type = parse_backend_type(options.backend);
|
||||
session_options_.backend.device = options.device;
|
||||
if (options.threads > 0) {
|
||||
session_options_.backend.threads = options.threads;
|
||||
}
|
||||
for (const auto &entry : options.session_options) {
|
||||
session_options_.options[entry.first] = entry.second;
|
||||
}
|
||||
|
||||
pinned_task_ = options.task;
|
||||
wait_budget_ceiling_ms_ = options.busy_timeout_ms;
|
||||
|
||||
engine::runtime::ModelLoadRequest request;
|
||||
request.model_path = std::filesystem::path(resolved_path);
|
||||
request.family_hint = family;
|
||||
@@ -315,17 +331,6 @@ LoadedModel::LoadedModel(const std::string &resolved_path,
|
||||
languages_ = engine_caps.languages;
|
||||
supports_timestamps_ = engine_caps.supports_timestamps;
|
||||
capabilities_ = to_capabilities(family, engine_caps);
|
||||
|
||||
session_options_.backend.type = parse_backend_type(options.backend);
|
||||
session_options_.backend.device = options.device;
|
||||
if (options.threads > 0) {
|
||||
session_options_.backend.threads = options.threads;
|
||||
}
|
||||
for (const auto &entry : options.session_options) {
|
||||
session_options_.options[entry.first] = entry.second;
|
||||
}
|
||||
|
||||
wait_budget_ceiling_ms_ = options.busy_timeout_ms;
|
||||
}
|
||||
|
||||
LoadedModel::Session LoadedModel::session_for(Rpc rpc, const RequestShape &shape) {
|
||||
@@ -336,7 +341,8 @@ LoadedModel::Session LoadedModel::session_for(Rpc rpc, const RequestShape &shape
|
||||
|
||||
const SessionKey key{static_cast<int>(route.task), static_cast<int>(route.mode)};
|
||||
auto found = sessions_.find(key);
|
||||
if (found == sessions_.end()) {
|
||||
const bool cache_hit = found != sessions_.end();
|
||||
if (!cache_hit) {
|
||||
engine::runtime::TaskSpec spec;
|
||||
spec.task = to_engine_task(route.task);
|
||||
spec.mode = to_engine_mode(route.mode);
|
||||
@@ -344,13 +350,23 @@ LoadedModel::Session LoadedModel::session_for(Rpc rpc, const RequestShape &shape
|
||||
try {
|
||||
created = model_->create_task_session(spec, session_options_);
|
||||
} catch (const std::exception &err) {
|
||||
throw CapabilityError(
|
||||
// NOT a CapabilityError. The family said it supports this pair, and
|
||||
// a throw from here is overwhelmingly an environment fault: a ggml
|
||||
// backend .so that package.sh did not ship, an out of memory, a CUDA
|
||||
// device that is not there. UNIMPLEMENTED would tell LocalAI and
|
||||
// every client "this model cannot do this, never retry", and send an
|
||||
// operator hunting a capability bug instead of a packaging one. A
|
||||
// plain runtime_error maps to INTERNAL, which is what a fixable
|
||||
// deployment fault should look like.
|
||||
throw std::runtime_error(
|
||||
std::string("audio-cpp: family '") + capabilities_.family +
|
||||
"' advertises " + task_name(route.task) + "/" +
|
||||
mode_name(route.mode) + " but refused to create the session: " +
|
||||
err.what());
|
||||
}
|
||||
if (created == nullptr) {
|
||||
// A null return with no throw is the family declining, which is a
|
||||
// genuine capability answer and stays UNIMPLEMENTED.
|
||||
throw CapabilityError(std::string("audio-cpp: family '") +
|
||||
capabilities_.family +
|
||||
"' returned no session for " +
|
||||
@@ -373,6 +389,18 @@ LoadedModel::Session LoadedModel::session_for(Rpc rpc, const RequestShape &shape
|
||||
"' advertises " + task_name(route.task) +
|
||||
"/streaming but its session is not streaming");
|
||||
}
|
||||
// Deliberately NOT reset here, though a cached streaming session does
|
||||
// carry state across chunks. reset() is not callable at this point:
|
||||
// silero_vad's implementation throws "session prepare() must be called
|
||||
// before Silero VAD reset()", so resetting on a cache hit would turn an
|
||||
// ordinary second fetch into a hard error, which is worse than the leak
|
||||
// it would prevent.
|
||||
//
|
||||
// The state is instead cleared by the sequence every streaming caller
|
||||
// owes anyway. IStreamingVoiceTaskSession::start_stream's base
|
||||
// implementation IS a call to reset(), so a caller that runs
|
||||
// prepare(...) then start_stream(...) at the top of each stream gets a
|
||||
// clean session for free. See the STATE CONTRACT in loaded_model.h.
|
||||
} else {
|
||||
session.offline =
|
||||
dynamic_cast<engine::runtime::IOfflineVoiceTaskSession *>(raw);
|
||||
|
||||
@@ -81,13 +81,41 @@ public:
|
||||
return session_options_;
|
||||
}
|
||||
|
||||
// The model's `task:` option, empty when unset. Every handler must copy it
|
||||
// into RequestShape::pinned_task before calling session_for: routing is
|
||||
// otherwise derived from the RPC alone, and this is the option's only route
|
||||
// from the load to the request that honours it.
|
||||
const std::string &pinned_task() const noexcept { return pinned_task_; }
|
||||
|
||||
// Routes the RPC and returns the cached session, creating it on first use.
|
||||
// Throws CapabilityError when this family cannot serve the RPC.
|
||||
// Throws CapabilityError when this family cannot serve the RPC, and a plain
|
||||
// runtime_error when it can but the session could not be built, which is an
|
||||
// environment fault rather than a capability answer.
|
||||
//
|
||||
// Call it only while holding this model's lane. The session cache is not
|
||||
// itself synchronised, and it does not need to be: the lane admits one
|
||||
// caller at a time, which is the same constraint the sessions themselves
|
||||
// impose.
|
||||
//
|
||||
// STATE CONTRACT, and it is the CALLER'S to honour. Sessions are cached per
|
||||
// (task, mode), so a streaming session is normally the same warm object the
|
||||
// previous stream used, carrying that stream's state. session_for hands it
|
||||
// back as it is.
|
||||
//
|
||||
// Every streaming caller must therefore begin a stream with
|
||||
//
|
||||
// session.streaming->prepare(build_preparation_request(...));
|
||||
// session.streaming->start_stream(request);
|
||||
//
|
||||
// in that order. start_stream's base implementation is a call to reset(),
|
||||
// which is what clears the previous stream, and reset() is only legal after
|
||||
// prepare(): silero_vad throws "session prepare() must be called before
|
||||
// Silero VAD reset()" otherwise. That ordering constraint is also why
|
||||
// session_for cannot do this for you. Skipping it does not raise an error,
|
||||
// it silently continues the previous stream.
|
||||
//
|
||||
// Offline sessions need no such care: their interface has no reset and
|
||||
// run() takes a whole request.
|
||||
Session session_for(Rpc rpc, const RequestShape &shape);
|
||||
|
||||
// Takes the inference lane, or throws LaneUnavailable. Serializes runs
|
||||
@@ -139,6 +167,7 @@ private:
|
||||
std::string variant_;
|
||||
std::string description_;
|
||||
std::vector<std::string> languages_;
|
||||
std::string pinned_task_;
|
||||
bool supports_timestamps_ = false;
|
||||
int wait_budget_ceiling_ms_ = 0;
|
||||
};
|
||||
|
||||
@@ -116,6 +116,7 @@ ParsedOptions parse_model_options(const std::vector<std::string> &entries) {
|
||||
"integer, got '" + value + "'";
|
||||
return parsed;
|
||||
}
|
||||
parsed.options.device_set = true;
|
||||
} else if (key == "threads") {
|
||||
if (!parse_non_negative_int(value, parsed.options.threads)) {
|
||||
parsed.error = "audio-cpp: option 'threads' needs a non-negative "
|
||||
|
||||
@@ -20,6 +20,11 @@ struct ModelOptions {
|
||||
// ggml backend: cpu, cuda, vulkan, metal, best.
|
||||
std::string backend = "cpu";
|
||||
int device = 0;
|
||||
// True once a `device:` entry has been seen. 0 is both the default and a
|
||||
// legitimate device index, so the value alone cannot tell an explicit
|
||||
// `device:0` from an unset option, and a caller merging in its own fallback
|
||||
// would silently override the explicit choice.
|
||||
bool device_set = false;
|
||||
// 0 means "let the runtime decide".
|
||||
int threads = 0;
|
||||
std::string model_spec_override;
|
||||
|
||||
Reference in New Issue
Block a user