mirror of
https://github.com/mudler/LocalAI.git
synced 2026-09-30 01:54:31 -04:00
* fix(config): do not read a TTS speaker-encoder mmproj as vision support Qwen3-TTS on llama-cpp ships an mmproj holding the speaker encoder and code predictor. VisionSupported() treated any non-empty MMProj as proof of image input, so every such model would be advertised as vision-capable. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(llama-cpp): add TTS request option parsing helper Validates text and speaker reference presence and strictly parses the top_k / top_p per-request params, in a header with no llama.cpp or gRPC dependencies so the standalone C++ unit test gate picks it up. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(llama-cpp): range-check the TTS top_k and top_p request params Format validation alone let NaN, infinity and out-of-range values through. The consumer copies both values into the audio generation input unconditionally and only guards its separate sampler assignment with "> 0", a test NaN also fails, so a NaN reached llama.cpp with the guard never firing. top_k must now be >= 0 and top_p must fall within 0.0 to 1.0 inclusive, with the bound written as a negated in-range test so NaN is rejected rather than silently accepted. Also cover the two checks the suite could not previously kill: the whole-string check in the float parser and the int32 range check. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(llama-cpp): bump pin to f9e832c10 and carry the TTS server task Picks up ggml-org/llama.cpp#26254 (Qwen3-TTS via mtmd) and #26536 (the short-input audio chunk fix). Adds 0002-add-server-task-type-tts.patch, the server-side half of the still-draft #26603, so TTS runs through the slot scheduler instead of racing it. Remove that patch when #26603 merges. The patch is rebased on top of the score patch: its tokenize-switch hunk collided with the SERVER_TASK_TYPE_SCORE case, and its lone SRV_WRN call passes no variadic argument, which the macro cannot expand. The score patch itself needed no refresh. Also fixes fallout from the bump in grpc-server.cpp: upstream dropped the per-slot n_ctx argument from server_schema::eval_llama_cmpl_schema. Only the schema branch loses it, since forks predating the server-schema split still expect the old argument list. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(llama-cpp): implement the TTS and TTSStream RPCs Both were declared in backend.proto but unimplemented. They now submit a SERVER_TASK_TYPE_TTS task and drain the response reader, the same shape PredictStream uses. The streaming path emits a leading sample_rate message and then raw PCM, because ModelTTSStream builds the WAV header itself; the non-streaming path emits a complete WAV to the requested dst. The streamed samples are converted from the pipeline's float32 to signed 16-bit first. MTMD_HELPER_GEN_AUDIO_OUTTYPE_PCM hands back floats, while the header ModelTTSStream writes announces 16-bit samples, so shipping the floats verbatim would decode as noise. prepare.sh and CMakeLists.txt now stage tts_request_options.h alongside the other grpc-server helpers, and register its standalone test with ctest the way passthrough_options_test is registered. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(llama-cpp): mask non-codec tokens for Qwen3-TTS generation The Qwen3-TTS gen-audio pipeline maps a sampled backbone token to a codebook row with an unchecked subtraction, in mtmd-helper-gen.cpp: inp.code0 = sampled - codec_0; For ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF the vocab is 155008 tokens, <|codec_0|> is 151936 and the codec codes end at 153983. The model's own tokenizer.ggml.suppress_tokens holds 1023 ids covering 153984..155007, every special above the codec range except <|codec_eos_token|> (154086) which stays reachable as the stop token. Nothing masks the text range 0..151935, so the backbone can sample a text token at any step, the subtraction goes negative, and ggml_compute_forward_get_rows aborts the whole backend process on GGML_ASSERT(i01 >= 0 && i01 < ne01). Complete the mask upstream started: bias every token below <|codec_0|> to -INFINITY for TTS tasks so only codec codes and the codec EOS remain reachable. The biases are appended to task.params.sampling.logit_bias, which common_sampler_init already merges with the model's suppress tokens into one llama_sampler_init_logit_bias, so no sampler is added to the chain. Measured cost is 0.082 ms per sampled token and 1.16 MB, set against a forward pass in the multi-millisecond range. It lands in launch_slot_with_task rather than in a route handler so that llama.cpp's own POST /tts and LocalAI's TTS/TTSStream RPCs are both covered, and <|codec_0|> is resolved from the vocab rather than hardcoded so a model without it is left alone. This is reproducible with upstream's own llama-tts and no LocalAI code loaded, aborting at frame 55 on Q4_K_M and frame 71 on Q8_0, so it is neither a quantization artifact nor an artifact of the gRPC adapter. Two further defects in the same draft pipeline still prevent end-to-end audio; they are independent of this one and are recorded in the task report for an upstream bug report. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * chore(llama-cpp): bump pin to 9de0fcf2b and drop the TTS codec mask Upstream fixed the Qwen3-TTS abort in ggml-org/llama.cpp c8e03ce81 ("mtmd/ggml: add ggml_build_forward_order", #26649), landed one hour after the previous pin. ggml_build_forward_expand marks a tensor and all its ancestors for compute, so using it as a pure ordering hint defeated ggml_build_forward_select and made GEN_WAV calls execute the GEN_CODE branch against a stale inp_code0, hitting the get_rows bound assert in ggml_compute_forward_get_rows. That single defect accounts for every abort seen on this model, so 0003-mask-non-codec-tokens-for-tts.patch is removed rather than rebased. The mask changed the observed behavior, but it was perturbing a graph ordering bug rather than fixing a sampling one: at the new pin the whole path works without it. Keeping it would have meant carrying a 152k-entry logit bias, and rebasing it on every pin bump, for no benefit. Verified at 9de0fcf2b with only 0001 and 0002 applied, which both apply clean with no fuzz and needed no rebase: non-streaming HTTP 200, 410924 bytes, 8.56 s RIFF (little-endian) data, WAVE audio, Microsoft PCM, 16 bit, mono 24000 Hz streaming HTTP 200, 560684 bytes, 11.68 s, exactly one RIFF at byte 0, same format, which also exercises the float32-to-s16 conversion at runtime for the first time Pristine unpatched llama-tts at the same pin now also completes, 130 frames to a valid WAV, where it aborted at frame 55 before. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(llama-cpp): clear the TTS slot sequence between requests Only the first TTS request in a backend process succeeded. Every later one failed instantly, in about 0.13 s, with "TTS prompt processing failed" from step_prompt, regardless of streaming or non-streaming and regardless of the text. With LOCALAI_SINGLE_ACTIVE_BACKEND=true the process is kept alive between requests, so a deployment would have served exactly one utterance per backend start. The cause is missing KV hygiene, not anything in the gRPC adapter. TTS slots never enter the shared batch: pre_decode() returns early for them and process_tts_slots() drives them instead, so they skip the prompt-cache bookkeeping that clears a slot's sequence between requests. Nothing in the gen-audio path makes up for it: mtmd_helper_gen_audio_reset only clears host-side buffers, and the pipeline always decodes from position 0 into the sequence identified by slot.id. So the second task on a slot writes positions 0..N over the first task's tokens and llama_decode fails. Fix is one call to slot.prompt_clear(), the same helper the normal path uses, in the SERVER_TASK_TYPE_TTS branch of launch_slot_with_task before set_input. It goes into 0002 rather than a new patch file because it is a defect in the code that patch introduces, and the header now records it as ours so we know whether it still needs carrying if #26603 merges without it. Verified in one backend process, different text on every request: three consecutive non-streaming requests, three consecutive streaming requests, and an interleaved non-streaming, streaming, non-streaming, streaming run. All ten returned HTTP 200 with RIFF ... WAVE audio, Microsoft PCM, 16 bit, mono 24000 Hz, the streamed ones carrying exactly one RIFF header at byte 0, and every output measured as real speech rather than silence or a truncated fragment. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(llama-cpp): expose max_frames for TTS requests The Qwen3-TTS backbone does not always emit <|codec_eos_token|>, and when it does not, generation runs to upstream's 512-frame n_predict default. At the model's 12.5 Hz frame rate that is 40.96 s of audio, which a short input can trigger: one request in this session produced 40.96 s for a ten-word sentence. prepareTTSTask hardcoded n_predict to -1, so callers had no way to bound it. Add a max_frames key alongside top_k and top_p, parsed with the same strict whole-string parsing so a typo is an error rather than a silently truncated value, and rejected with a field-naming message when negative. 0 keeps the existing sentinel convention and means unset, so a request that omits it behaves exactly as before. Named max_frames rather than n_predict because frames are what the parameter means at a TTS endpoint: one frame is 0.08 s of audio. The 512-frame default is deliberately unchanged. Lowering it would truncate legitimately long inputs, which is a worse failure than an occasionally overlong one. Verified end to end on one text of thirty words: max_frames=25 HTTP 200, 96044 bytes, 2.00 s, exactly 25 frames max_frames=50 HTTP 200, 192044 bytes, 4.00 s, exactly 50 frames no max_frames HTTP 200, 572204 bytes, 11.92 s, stopped at its own codec EOS after 149 frames, unchanged behavior max_frames=-1 InvalidArgument "max_frames must be >= 0, got \"-1\"" max_frames=many InvalidArgument "max_frames must be an integer, got \"many\"" Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(llama-cpp): send the TTS sample rate up front, and tidy three review items Four items from the Task 4 review. Streaming first-byte latency. TTSStream sent the sample-rate reply only once the first audio result arrived, and a chunk needs a whole 72-frame window, roughly 5.8 s of audio and far longer in wall time on CPU. The Go side blocks on that reply before it can emit the WAV header, so a streaming client sat at zero bytes for the whole stretch. The rate is a property of the loaded model and is available synchronously from mtmd_gen_audio_get_info, so it now goes out immediately after post_task and the rate_sent bookkeeping is gone. Measured on a warm model, first byte drops from 30.48 s to 0.014 s, and the output is still a valid WAV with exactly one RIFF header at byte 0. Unchecked close. The non-streaming path ignored ofstream::close(), so a failure that only surfaces on flush was reported as success while leaving a truncated file at dst. It now returns INTERNAL like the other write failures. Wrong comment on set_lang. gen_audio::inp::get() already maps a stored blank to nullptr, so our guard is behavior-preserving, not behavior-fixing. The comment claimed otherwise; the code was right. Repetition penalty. penalty_last_n = -1 is inert at this pin, because llama_sampler_init_penalties clamps it with std::max(penalty_last_n, 0) and then builds a disabled sampler, so the 1.05 penalty never applies. Upstream's README attributes looping to a missing repeat_penalty, so it was worth testing as a root-cause fix for the model running to the frame cap. Dropping the line lets the sampling default of 64 apply, which was confirmed in the sampler chain trace as penalty_last_n = 64 with repeat_penalty = 1.050. Over 15 uncapped short requests each way it did not help: 0 of 15 ran to the cap with the penalty inert, 1 of 15 with it active. Both lines are therefore kept for parity with upstream's draft, and a comment now records that the pair is inert and why, so the next reader does not believe a penalty is applied. max_frames remains the way to bound output. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * build(llama-cpp): let unpatched forks opt out of the TTS task turboquant and bonsai copy grpc-server.cpp into llama.cpp forks that do not carry our patches. disable-tts-task.sh injects the same kind of preprocessor switch disable-score-task.sh already uses, so those builds answer UNIMPLEMENTED rather than failing to compile. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(config): keep a TTS speaker-encoder projector out of vision detection Task 1 exempted a declared-TTS model's mmproj from VisionSupported, but the first real gallery entry with an mmproj still came back vision-capable through two paths the earlier fix did not close. GuessUsecases has no FLAG_VISION branch, so it falls through to true for any chat-ish model. That is not just a wrong answer at the call site: syncKnownUsecasesFromString rewrites KnownUsecaseStrings from HasUsecases, and the loader calls it more than once per config file, so the guessed FLAG_VISION is written out and parsed back into KnownUsecases as if the operator had declared it. Give GuessUsecases a FLAG_VISION branch that defers to the same explicit signals VisionSupported uses. Second, llama.cpp builds an mtmd context for the speaker-encoder projector and reports its media marker on the first chat probe, which resurrected vision after the model had been used once. Apply the same declared-TTS exemption to MediaMarker that the mmproj check already had. Verified against the qwen3-tts-llamacpp-q4 gallery entry: no vision capability and no image input modality, before load, after a TTS request, and after a chat probe. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * feat(gallery): add Qwen3-TTS entries for the llama-cpp backend Two entries over upstream's own GGUF conversion, Q8_0 and Q4_K_M, each pairing a backbone with the Q8_0 projector. Named to sit alongside the existing qwen3-tts-cpp entries rather than replace them. Also tags the llama-cpp backend text-to-speech / TTS so the backend browser surfaces the capability. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * docs: cover Qwen3-TTS on the llama-cpp backend Adds the gallery variants, the two-file mmproj configuration, the required voice reference, and the language and sampling knobs. Also corrects the streaming-support list, which named only voxcpm. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(config): register llama-cpp as a TTS and voice-cloning backend The branch taught the llama-cpp backend to serve Qwen3-TTS and shipped two gallery entries for it, but never told the capability table. llama-cpp still declared only the text RPCs and usecases, so: - VoiceCloningForModel returned nil at the capability check, before it ever reached the model's own tts.voice_cloning override, and /tts answered 400 "selected model does not support reference-audio voice cloning" for any localai://voice-profiles/... voice. No model YAML could opt back in. - GET /api/backends/usecases did not list tts for llama-cpp, so the gallery greyed out the TTS filter for the entries this branch adds. - The React TTS page saw voice_cloning: null and kept both models out of the Voice Library. Add the TTS RPCs and usecase, and the reference-audio contract. The contract needs narrowing, because the per-backend switch in VoiceCloningForModel ends in a permissive default: an unnarrowed entry would have advertised reference-audio cloning on every GGUF chat model in the gallery. Narrow on the declared TTS usecase rather than the model name. The TTS checkpoints are the only llama-cpp models carrying known_usecases: [tts]; name matching would have to guess at third-party repacks, and "base", the substring the neighbouring Qwen and vLLM cases key on, is a routine word in text-model names. The check reads the declared bit directly instead of going through HasUsecases, which falls through to GuessUsecases and would hand the decision to a heuristic that never had a llama.cpp TTS model in mind. DefaultUsecases stays [chat]: a bare GGUF served by llama.cpp is a chat model, and both the gallery filter and the importer read that field. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(gallery): declare what nemotron-3-nano-omni actually accepts The entry is backend: vllm-omni with known_usecases: [chat, completion], no mmproj and no media marker, so it used to report vision only through the blanket GuessUsecases fallthrough that the vision branch in this branch removed. Nemotron 3 Nano Omni is a multimodal understanding model: image, video and audio in, text out. Declaring that is what the sibling vllm-omni-qwen3-omni-30b already does. known_usecases gains vision only. FLAG_VIDEO is video GENERATION, an output modality, and this model generates none; video and audio input belong in known_input_modalities, which is where AudioInputSupported and VideoInputSupported read them from. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(importers): import a Qwen3-TTS GGUF repo as TTS, not chat The llama-cpp importer hardcodes known_usecases: [chat] and assigns any mmproj-matching file as a vision projector, so ggml-org/Qwen3-TTS-12Hz-1.7B- Base-GGUF imported as a chat model with vision. Both fields were wrong, and the model was unreachable from /tts and from the Voice Library. Filenames cannot fix this. A Qwen3-TTS repo has the exact shape of a vision repo, one backbone GGUF plus one mmproj-*.gguf, so the projector's own header is the only honest signal: mtmd writes clip.has_gen_audio_encoder for the projectors it can drive as a speech pipeline and refuses to build one without it. Probe the selected mmproj for that flag, reusing the range-fetch the MTP detection already does, and declare tts when it is set. The mmproj assignment then stops reading as vision on its own, since a declared-TTS model already exempts its projector from vision detection. The probe is best-effort like the MTP one: a network blip leaves the chat default in place rather than failing the import. Verified against the real artifacts on disk: the Qwen3-TTS projector reports gen-audio, its backbone does not. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> * fix(llama-cpp): stop non-TTS models crashing on the new pin Two regressions, both hit every ordinary llama-cpp model and neither was caught locally because every test on this branch loaded a TTS model. The first is a null dereference. server_slot::tts_ctx::reset() called mtmd_helper_gen_audio_reset() unconditionally, but the gen-audio pipeline is only allocated for models carrying a gen-audio mmproj, and upstream's implementation reads ctx->pipeline before null-checking anything. Since server_slot::reset() runs during slot initialization for every model, any non-TTS model segfaulted the backend the moment it loaded. Guard the call on the is_supported() predicate already defined beside it, and keep the plain field resets unconditional. The second is unrelated to TTS and came in with the pin bump. PredictOptions.Penalty is a bare proto float, so a caller that names no repetition penalty sends 0 rather than omitting the field. Since 9de0fcf2b, common_sampler_init() rejects a non-positive penalty_repeat outright because it would divide logits by zero, turning every such request into "Failed to initialize samplers". Treat 0 as unset and leave llama.cpp's own neutral default in place. Verified with the same suite CI runs, which is what caught both: tests/e2e-backends passes 6 of 6 including the load and predict specs that were red. Qwen3-TTS still synthesises on both paths, 24 kHz mono 16-bit WAV with exactly one RIFF header on the streamed output. Assisted-by: Claude:claude-fable-5 [Claude Code] Signed-off-by: Ettore Di Giacinto <mudler@localai.io> --------- Signed-off-by: Ettore Di Giacinto <mudler@localai.io> Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
1115 lines
48 KiB
Go
1115 lines
48 KiB
Go
package config
|
|
|
|
import (
|
|
"slices"
|
|
"strings"
|
|
|
|
"github.com/mudler/LocalAI/pkg/model"
|
|
)
|
|
|
|
// Usecase name constants — the canonical string values used in gallery entries,
|
|
// model configs (known_usecases), and UsecaseInfoMap keys.
|
|
const (
|
|
UsecaseChat = "chat"
|
|
UsecaseCompletion = "completion"
|
|
UsecaseEdit = "edit"
|
|
UsecaseVision = "vision"
|
|
UsecaseEmbeddings = "embeddings"
|
|
UsecaseTokenize = "tokenize"
|
|
UsecaseImage = "image"
|
|
UsecaseVideo = "video"
|
|
Usecase3D = "3d"
|
|
UsecaseTranscript = "transcript"
|
|
UsecaseTTS = "tts"
|
|
UsecaseSoundGeneration = "sound_generation"
|
|
UsecaseRerank = "rerank"
|
|
UsecaseDetection = "detection"
|
|
UsecaseDepth = "depth"
|
|
UsecaseVAD = "vad"
|
|
UsecaseAudioTransform = "audio_transform"
|
|
UsecaseDiarization = "diarization"
|
|
UsecaseSoundClassification = "sound_classification"
|
|
UsecaseRealtimeAudio = "realtime_audio"
|
|
UsecaseFaceRecognition = "face_recognition"
|
|
UsecaseSpeakerRecognition = "speaker_recognition"
|
|
UsecaseTokenClassify = "token_classify"
|
|
UsecaseScore = "score"
|
|
)
|
|
|
|
// GRPCMethod identifies a Backend service RPC from backend.proto.
|
|
type GRPCMethod string
|
|
|
|
const (
|
|
MethodPredict GRPCMethod = "Predict"
|
|
MethodPredictStream GRPCMethod = "PredictStream"
|
|
MethodEmbedding GRPCMethod = "Embedding"
|
|
MethodGenerateImage GRPCMethod = "GenerateImage"
|
|
MethodUpscaleImage GRPCMethod = "UpscaleImage"
|
|
MethodGenerateVideo GRPCMethod = "GenerateVideo"
|
|
MethodGenerate3D GRPCMethod = "Generate3D"
|
|
MethodAudioTranscription GRPCMethod = "AudioTranscription"
|
|
MethodTTS GRPCMethod = "TTS"
|
|
MethodTTSStream GRPCMethod = "TTSStream"
|
|
MethodSoundGeneration GRPCMethod = "SoundGeneration"
|
|
MethodTokenizeString GRPCMethod = "TokenizeString"
|
|
MethodDetect GRPCMethod = "Detect"
|
|
MethodDepth GRPCMethod = "Depth"
|
|
MethodRerank GRPCMethod = "Rerank"
|
|
MethodVAD GRPCMethod = "VAD"
|
|
MethodAudioTransform GRPCMethod = "AudioTransform"
|
|
MethodDiarize GRPCMethod = "Diarize"
|
|
MethodSoundDetection GRPCMethod = "SoundDetection"
|
|
MethodAudioToAudioStream GRPCMethod = "AudioToAudioStream"
|
|
MethodFaceVerify GRPCMethod = "FaceVerify"
|
|
MethodFaceAnalyze GRPCMethod = "FaceAnalyze"
|
|
MethodVoiceVerify GRPCMethod = "VoiceVerify"
|
|
MethodVoiceEmbed GRPCMethod = "VoiceEmbed"
|
|
MethodVoiceAnalyze GRPCMethod = "VoiceAnalyze"
|
|
MethodTokenClassify GRPCMethod = "TokenClassify"
|
|
MethodScore GRPCMethod = "Score"
|
|
)
|
|
|
|
// UsecaseInfo describes a single known_usecase value and how it maps
|
|
// to the gRPC backend API.
|
|
type UsecaseInfo struct {
|
|
// Flag is the ModelConfigUsecase bitmask value.
|
|
Flag ModelConfigUsecase
|
|
// GRPCMethod is the primary Backend service RPC this usecase maps to.
|
|
GRPCMethod GRPCMethod
|
|
// IsModifier is true when this usecase doesn't map to its own gRPC RPC
|
|
// but modifies how another RPC behaves (e.g., vision uses Predict with images).
|
|
IsModifier bool
|
|
// DependsOn names the usecase(s) this modifier requires (e.g., "chat").
|
|
DependsOn string
|
|
// Description is a human/LLM-readable explanation of what this usecase means.
|
|
Description string
|
|
}
|
|
|
|
// UsecaseInfoMap maps each known_usecase string to its gRPC and semantic info.
|
|
var UsecaseInfoMap = map[string]UsecaseInfo{
|
|
UsecaseChat: {
|
|
Flag: FLAG_CHAT,
|
|
GRPCMethod: MethodPredict,
|
|
Description: "Conversational/instruction-following via the Predict RPC with chat templates.",
|
|
},
|
|
UsecaseCompletion: {
|
|
Flag: FLAG_COMPLETION,
|
|
GRPCMethod: MethodPredict,
|
|
Description: "Text completion via the Predict RPC with a completion template.",
|
|
},
|
|
UsecaseEdit: {
|
|
Flag: FLAG_EDIT,
|
|
GRPCMethod: MethodPredict,
|
|
Description: "Text editing via the Predict RPC with an edit template.",
|
|
},
|
|
UsecaseVision: {
|
|
Flag: FLAG_VISION,
|
|
GRPCMethod: MethodPredict,
|
|
IsModifier: true,
|
|
DependsOn: UsecaseChat,
|
|
Description: "The model accepts images alongside text in the Predict RPC. For llama-cpp this requires an mmproj file.",
|
|
},
|
|
UsecaseEmbeddings: {
|
|
Flag: FLAG_EMBEDDINGS,
|
|
GRPCMethod: MethodEmbedding,
|
|
Description: "Vector embedding generation via the Embedding RPC.",
|
|
},
|
|
UsecaseTokenize: {
|
|
Flag: FLAG_TOKENIZE,
|
|
GRPCMethod: MethodTokenizeString,
|
|
Description: "Tokenization via the TokenizeString RPC without running inference.",
|
|
},
|
|
UsecaseImage: {
|
|
Flag: FLAG_IMAGE,
|
|
GRPCMethod: MethodGenerateImage,
|
|
Description: "Image generation via the GenerateImage RPC (Stable Diffusion, Flux, etc.).",
|
|
},
|
|
UsecaseVideo: {
|
|
Flag: FLAG_VIDEO,
|
|
GRPCMethod: MethodGenerateVideo,
|
|
Description: "Video generation via the GenerateVideo RPC, with optional image or audio conditioning when supported by the backend.",
|
|
},
|
|
Usecase3D: {
|
|
Flag: FLAG_3D,
|
|
GRPCMethod: MethodGenerate3D,
|
|
Description: "Image-conditioned 3D asset generation via the Generate3D RPC — a binary glTF (GLB) mesh with optional PBR material (TRELLIS.2).",
|
|
},
|
|
UsecaseTranscript: {
|
|
Flag: FLAG_TRANSCRIPT,
|
|
GRPCMethod: MethodAudioTranscription,
|
|
Description: "Speech-to-text via the AudioTranscription RPC.",
|
|
},
|
|
UsecaseTTS: {
|
|
Flag: FLAG_TTS,
|
|
GRPCMethod: MethodTTS,
|
|
Description: "Text-to-speech via the TTS RPC.",
|
|
},
|
|
UsecaseSoundGeneration: {
|
|
Flag: FLAG_SOUND_GENERATION,
|
|
GRPCMethod: MethodSoundGeneration,
|
|
Description: "Music/sound generation via the SoundGeneration RPC (not speech).",
|
|
},
|
|
UsecaseRerank: {
|
|
Flag: FLAG_RERANK,
|
|
GRPCMethod: MethodRerank,
|
|
Description: "Document reranking via the Rerank RPC.",
|
|
},
|
|
UsecaseDetection: {
|
|
Flag: FLAG_DETECTION,
|
|
GRPCMethod: MethodDetect,
|
|
Description: "Object detection via the Detect RPC with bounding boxes.",
|
|
},
|
|
UsecaseDepth: {
|
|
Flag: FLAG_DEPTH,
|
|
GRPCMethod: MethodDepth,
|
|
Description: "Per-pixel metric depth, camera pose and 3D point cloud via the Depth RPC (Depth Anything 3).",
|
|
},
|
|
UsecaseVAD: {
|
|
Flag: FLAG_VAD,
|
|
GRPCMethod: MethodVAD,
|
|
Description: "Voice activity detection via the VAD RPC.",
|
|
},
|
|
UsecaseAudioTransform: {
|
|
Flag: FLAG_AUDIO_TRANSFORM,
|
|
GRPCMethod: MethodAudioTransform,
|
|
Description: "Audio-in / audio-out transformations (echo cancellation, noise suppression, dereverberation, voice conversion) via the AudioTransform RPC.",
|
|
},
|
|
UsecaseDiarization: {
|
|
Flag: FLAG_DIARIZATION,
|
|
GRPCMethod: MethodDiarize,
|
|
Description: "Speaker diarization (who-spoke-when, per-speaker segments) via the Diarize RPC.",
|
|
},
|
|
UsecaseSoundClassification: {
|
|
Flag: FLAG_SOUND_CLASSIFICATION,
|
|
GRPCMethod: MethodSoundDetection,
|
|
Description: "Sound-event classification / audio tagging (scored AudioSet labels like baby cry, glass breaking, alarms) via the SoundDetection RPC.",
|
|
},
|
|
UsecaseRealtimeAudio: {
|
|
Flag: FLAG_REALTIME_AUDIO,
|
|
GRPCMethod: MethodAudioToAudioStream,
|
|
Description: "Self-contained any-to-any audio model for the Realtime API — accepts microphone audio and emits speech + transcript (+ optional function calls) from a single backend via the AudioToAudioStream RPC.",
|
|
},
|
|
UsecaseFaceRecognition: {
|
|
Flag: FLAG_FACE_RECOGNITION,
|
|
GRPCMethod: MethodFaceVerify,
|
|
Description: "Face recognition — verify identity, analyze attributes (age/gender/emotion) via FaceVerify and FaceAnalyze RPCs.",
|
|
},
|
|
UsecaseSpeakerRecognition: {
|
|
Flag: FLAG_SPEAKER_RECOGNITION,
|
|
GRPCMethod: MethodVoiceVerify,
|
|
Description: "Speaker recognition — verify identity, embed and analyze voice via VoiceVerify, VoiceEmbed and VoiceAnalyze RPCs.",
|
|
},
|
|
UsecaseTokenClassify: {
|
|
Flag: FLAG_TOKEN_CLASSIFY,
|
|
GRPCMethod: MethodTokenClassify,
|
|
Description: "Per-token classification (NER) via the TokenClassify RPC — the PII detector tier. Declared explicitly via known_usecases; never auto-guessed, since the token-classification head is not useful as general generation or embeddings.",
|
|
},
|
|
UsecaseScore: {
|
|
Flag: FLAG_SCORE,
|
|
GRPCMethod: MethodScore,
|
|
Description: "Joint log-probability scoring of candidate continuations via the Score RPC. Declared explicitly via known_usecases and usable alongside generation usecases.",
|
|
},
|
|
}
|
|
|
|
// BackendCapability describes which gRPC methods and usecases a backend supports.
|
|
// Derived from reviewing actual implementations in backend/go/ and backend/python/.
|
|
type BackendCapability struct {
|
|
// GRPCMethods lists the Backend service RPCs this backend implements.
|
|
GRPCMethods []GRPCMethod
|
|
// PossibleUsecases lists all usecase strings this backend can support.
|
|
PossibleUsecases []string
|
|
// DefaultUsecases lists the conservative safe defaults.
|
|
DefaultUsecases []string
|
|
// AcceptsImages indicates multimodal image input in Predict.
|
|
AcceptsImages bool
|
|
// AcceptsVideos indicates multimodal video input in Predict.
|
|
AcceptsVideos bool
|
|
// AcceptsAudios indicates multimodal audio input in Predict.
|
|
AcceptsAudios bool
|
|
// AudioTransformInputMono16k declares that this backend's AudioTransform
|
|
// input must be folded to 16 kHz mono 16-bit WAV before it is handed over.
|
|
//
|
|
// Opt-IN, and the default of false means "hand the backend the upload as
|
|
// it is". The /audio/transform endpoint used to fold EVERY upload to
|
|
// 16 kHz mono, which is what LocalVQE wants for acoustic echo cancellation
|
|
// and what no source separation model can survive: htdemucs and
|
|
// mel_band_roformer refuse any rate but their checkpoint's own (44.1 kHz
|
|
// for every published one) and work in stereo, so every separation request
|
|
// made through the HTTP API failed with an INTERNAL raised inside the
|
|
// engine, while a direct gRPC call worked. Declaring the need rather than
|
|
// defaulting to it means a backend that wants the fold says so and a
|
|
// backend that does not needs no entry here at all.
|
|
AudioTransformInputMono16k bool
|
|
// VoiceCloning describes the backend's per-request reference-audio
|
|
// contract. Model variants that share a backend may narrow this further;
|
|
// use VoiceCloningForModel for UI/API decisions.
|
|
VoiceCloning *VoiceCloningCapability
|
|
// Description is a human-readable summary of the backend.
|
|
Description string
|
|
}
|
|
|
|
// VoiceCloningCapability is the model-facing contract for reusable reference
|
|
// voices. The first release intentionally accepts only browser-normalizable
|
|
// PCM WAV so every advertised backend sees the same input shape.
|
|
type VoiceCloningCapability struct {
|
|
ReferenceTranscriptRequired bool `json:"reference_transcript_required"`
|
|
AcceptedAudioFormats []string `json:"accepted_audio_formats"`
|
|
}
|
|
|
|
func referenceVoiceCloning() *VoiceCloningCapability {
|
|
return &VoiceCloningCapability{
|
|
ReferenceTranscriptRequired: true,
|
|
AcceptedAudioFormats: []string{"audio/wav"},
|
|
}
|
|
}
|
|
|
|
// BackendCapabilities maps each backend name (as used in model configs and gallery
|
|
// entries) to its verified capabilities. This is the single source of truth for
|
|
// what each backend supports.
|
|
//
|
|
// Backend names use hyphens (e.g., "llama-cpp") matching the gallery convention.
|
|
// Use NormalizeBackendName() for names with dots (e.g., "llama.cpp").
|
|
var BackendCapabilities = map[string]BackendCapability{
|
|
// --- LLM / text generation backends ---
|
|
// llama.cpp also serves Qwen3-TTS, so TTS is in the union below. It is NOT
|
|
// in DefaultUsecases: a bare GGUF served by llama.cpp is a chat model, and
|
|
// the TTS models declare known_usecases: [tts]. VoiceCloning is likewise
|
|
// narrowed per model in VoiceCloningForModel, since the vast majority of
|
|
// llama-cpp models in the gallery are text LLMs that clone nothing.
|
|
"llama-cpp": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodPredictStream, MethodEmbedding, MethodTokenizeString, MethodScore, MethodTTS, MethodTTSStream},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseEdit, UsecaseEmbeddings, UsecaseTokenize, UsecaseVision, UsecaseScore, UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseChat},
|
|
AcceptsImages: true, // requires mmproj
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "llama.cpp GGUF models: LLM inference with optional vision via mmproj, and Qwen3-TTS speech with reference-audio cloning",
|
|
},
|
|
// privacy-filter is the standalone GGML engine (backend/cpp/privacy-filter,
|
|
// wrapping privacy-filter.cpp) for the openai-privacy-filter PII/NER token
|
|
// classifier — the dedicated TokenClassify path that replaces the
|
|
// patched-llama.cpp route. Never auto-guessed; declared explicitly via
|
|
// known_usecases: [token_classify].
|
|
"privacy-filter": {
|
|
GRPCMethods: []GRPCMethod{MethodTokenClassify},
|
|
PossibleUsecases: []string{UsecaseTokenClassify},
|
|
DefaultUsecases: []string{UsecaseTokenClassify},
|
|
Description: "privacy-filter.cpp — standalone GGML backend for openai-privacy-filter PII/NER token classification",
|
|
},
|
|
"vllm": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodPredictStream, MethodEmbedding},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseEmbeddings, UsecaseVision},
|
|
DefaultUsecases: []string{UsecaseChat},
|
|
AcceptsImages: true,
|
|
AcceptsVideos: true,
|
|
Description: "vLLM engine — high-throughput LLM serving with optional multimodal",
|
|
},
|
|
"sglang": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodPredictStream, MethodTokenizeString},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseTokenize, UsecaseVision},
|
|
DefaultUsecases: []string{UsecaseChat},
|
|
AcceptsImages: true,
|
|
Description: "SGLang — fast LLM inference with structured generation and optional vision",
|
|
},
|
|
// vllm-cpp serves two mutually exclusive engine handles from one backend:
|
|
// a text engine, and MiniMax-H3's video+audio engine when the model config
|
|
// declares the H3 checkpoint set. Both usecases are possible, and chat is
|
|
// the default because a config that says nothing is a text model.
|
|
//
|
|
// AcceptsImages is the fl2va keyframe (start_image/end_image), the same
|
|
// reason longcat-video declares it; the text path takes no image input.
|
|
"vllm-cpp": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodPredictStream, MethodGenerateVideo},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseVideo},
|
|
DefaultUsecases: []string{UsecaseChat},
|
|
AcceptsImages: true,
|
|
Description: "vllm.cpp — the LocalAI team's C++20 port of vLLM; text generation plus MiniMax-H3 video+audio generation",
|
|
},
|
|
"vllm-omni": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodPredictStream, MethodGenerateImage, MethodGenerateVideo, MethodTTS},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseImage, UsecaseVideo, UsecaseTTS, UsecaseVision},
|
|
DefaultUsecases: []string{UsecaseChat},
|
|
AcceptsImages: true,
|
|
AcceptsVideos: true,
|
|
AcceptsAudios: true,
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "vLLM omni-modal — supports text, image, video generation and TTS",
|
|
},
|
|
"transformers": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodPredictStream, MethodEmbedding, MethodTTS, MethodSoundGeneration},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseEmbeddings, UsecaseTTS, UsecaseSoundGeneration},
|
|
DefaultUsecases: []string{UsecaseChat},
|
|
Description: "HuggingFace transformers — general-purpose Python inference",
|
|
},
|
|
"mlx": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodPredictStream, MethodEmbedding},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseEmbeddings},
|
|
DefaultUsecases: []string{UsecaseChat},
|
|
Description: "Apple MLX framework — optimized for Apple Silicon",
|
|
},
|
|
"mlx-distributed": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodPredictStream, MethodEmbedding},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseEmbeddings},
|
|
DefaultUsecases: []string{UsecaseChat},
|
|
Description: "MLX distributed inference across multiple Apple Silicon devices",
|
|
},
|
|
"mlx-vlm": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodPredictStream, MethodEmbedding},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseEmbeddings, UsecaseVision},
|
|
DefaultUsecases: []string{UsecaseChat, UsecaseVision},
|
|
AcceptsImages: true,
|
|
AcceptsAudios: true,
|
|
Description: "MLX vision-language models with multimodal input",
|
|
},
|
|
"mlx-audio": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodTTS},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseChat},
|
|
Description: "MLX audio models — text generation and TTS",
|
|
},
|
|
|
|
// --- Image/video generation backends ---
|
|
"diffusers": {
|
|
GRPCMethods: []GRPCMethod{MethodGenerateImage, MethodUpscaleImage, MethodGenerateVideo},
|
|
PossibleUsecases: []string{UsecaseImage, UsecaseVideo},
|
|
DefaultUsecases: []string{UsecaseImage},
|
|
Description: "HuggingFace diffusers — Stable Diffusion, Flux, video generation",
|
|
},
|
|
"longcat-video": {
|
|
GRPCMethods: []GRPCMethod{MethodGenerateVideo},
|
|
PossibleUsecases: []string{UsecaseVideo},
|
|
DefaultUsecases: []string{UsecaseVideo},
|
|
AcceptsImages: true,
|
|
AcceptsAudios: true,
|
|
Description: "LongCat-Video — text, image, and audio-conditioned avatar video generation on NVIDIA CUDA",
|
|
},
|
|
"stablediffusion": {
|
|
GRPCMethods: []GRPCMethod{MethodGenerateImage},
|
|
PossibleUsecases: []string{UsecaseImage},
|
|
DefaultUsecases: []string{UsecaseImage},
|
|
Description: "Stable Diffusion native backend",
|
|
},
|
|
"stablediffusion-ggml": {
|
|
GRPCMethods: []GRPCMethod{MethodGenerateImage},
|
|
PossibleUsecases: []string{UsecaseImage},
|
|
DefaultUsecases: []string{UsecaseImage},
|
|
Description: "Stable Diffusion via GGML quantized models",
|
|
},
|
|
|
|
// --- 3D generation backends ---
|
|
"trellis2cpp": {
|
|
GRPCMethods: []GRPCMethod{MethodGenerate3D},
|
|
PossibleUsecases: []string{Usecase3D},
|
|
DefaultUsecases: []string{Usecase3D},
|
|
Description: "trellis2.cpp — C++/GGML port of Microsoft TRELLIS.2: single-image to textured 3D mesh (GLB)",
|
|
},
|
|
|
|
// --- Speech-to-text backends ---
|
|
"whisper": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription, MethodVAD},
|
|
PossibleUsecases: []string{UsecaseTranscript, UsecaseVAD},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
Description: "OpenAI Whisper — speech recognition and voice activity detection",
|
|
},
|
|
"faster-whisper": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription},
|
|
PossibleUsecases: []string{UsecaseTranscript},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
Description: "CTranslate2-accelerated Whisper for faster transcription",
|
|
},
|
|
"whisperx": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription},
|
|
PossibleUsecases: []string{UsecaseTranscript},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
Description: "WhisperX — Whisper with word-level timestamps and speaker diarization",
|
|
},
|
|
"moonshine": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription},
|
|
PossibleUsecases: []string{UsecaseTranscript},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
Description: "Moonshine speech recognition",
|
|
},
|
|
"nemo": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription},
|
|
PossibleUsecases: []string{UsecaseTranscript},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
Description: "NVIDIA NeMo speech recognition",
|
|
},
|
|
"parakeet-cpp": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription},
|
|
PossibleUsecases: []string{UsecaseTranscript},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
Description: "NVIDIA NeMo Parakeet ASR (parakeet.cpp)",
|
|
},
|
|
// nemo-speech-cpp is one gRPC server in front of four NeMo-Speech.cpp model
|
|
// families, picked at load time from the GGUF general.architecture key, so
|
|
// PossibleUsecases is their UNION and no single model serves all of it: an
|
|
// asr model transcribes (and diarizes, when a Sortformer model is attached
|
|
// through options), a sortformer model only diarizes, a magpietts model only
|
|
// synthesizes, and a Riva-Translate model only answers Predict.
|
|
//
|
|
// UsecaseChat sits alongside UsecaseCompletion for the translation family
|
|
// because Predict and PredictStream are exactly the RPCs /v1/chat/completions
|
|
// drives, and chat is what a translation model is useful through: each turn
|
|
// goes in as the prompt and comes back translated. The flag is not a gate on
|
|
// any endpoint (a request naming the model explicitly is served either way);
|
|
// what it buys is being eligible as the default chat model when a request
|
|
// names none (core/http/routes/openai.go) and appearing in the React UI's
|
|
// chat model picker (CAP_CHAT in react-ui/src/utils/capabilities.js).
|
|
//
|
|
// Leaving it out is not neutral: chat is a gallery filter key and completion
|
|
// is not (usecaseFilters in core/http/routes/ui_api.go), so
|
|
// GET /api/backends/usecases would grey the Chat filter out and hide a
|
|
// Riva-Translate gallery entry from the one filter that fits it.
|
|
//
|
|
// DefaultUsecases is transcript alone because that is the only family whose
|
|
// weights a bare `backend: nemo-speech-cpp` config is likely to name; a model
|
|
// of any other family should pin its own known_usecases.
|
|
//
|
|
// No VoiceCloning key: MagpieTTS synthesizes from baked speaker ids, not from
|
|
// a reference clip, so advertising cloning would accept a `voice:
|
|
// "profile:<id>"` request the backend cannot serve.
|
|
"nemo-speech-cpp": {
|
|
GRPCMethods: []GRPCMethod{
|
|
MethodAudioTranscription, MethodDiarize,
|
|
MethodTTS, MethodTTSStream,
|
|
MethodPredict, MethodPredictStream,
|
|
},
|
|
PossibleUsecases: []string{
|
|
UsecaseTranscript, UsecaseDiarization, UsecaseTTS,
|
|
UsecaseCompletion, UsecaseChat,
|
|
},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
Description: "NVIDIA NeMo-Speech.cpp: one server for Nemotron ASR (offline, streaming and live), Sortformer diarization, MagpieTTS synthesis and Riva-Translate translation; the model's GGUF architecture decides which",
|
|
},
|
|
"qwen-asr": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription},
|
|
PossibleUsecases: []string{UsecaseTranscript},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
Description: "Qwen automatic speech recognition",
|
|
},
|
|
"voxtral": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription},
|
|
PossibleUsecases: []string{UsecaseTranscript},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
Description: "Voxtral speech recognition",
|
|
},
|
|
"vibevoice": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription, MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTranscript, UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTranscript, UsecaseTTS},
|
|
Description: "VibeVoice — bidirectional speech (transcription and synthesis)",
|
|
},
|
|
"vibevoice-cpp": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription, MethodTTS, MethodTTSStream},
|
|
PossibleUsecases: []string{UsecaseTranscript, UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTranscript, UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "VibeVoice C++ — bidirectional speech, C++ backend with streaming TTS",
|
|
},
|
|
"sherpa-onnx": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription, MethodTTS, MethodTTSStream, MethodVAD},
|
|
PossibleUsecases: []string{UsecaseTranscript, UsecaseTTS, UsecaseVAD},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
Description: "Sherpa-ONNX — multi-model speech toolkit (ASR, TTS, VAD)",
|
|
},
|
|
// audio-cpp is one gRPC server in front of ~30 audio.cpp model families, so
|
|
// PossibleUsecases is their UNION and no single model serves all of it:
|
|
// which RPCs a given model answers is decided by the family baked into its
|
|
// GGUF, and every audio-cpp gallery entry pins its own known_usecases.
|
|
//
|
|
// VoiceCloning is not decoration here. VoiceCloningForModel returns nil as
|
|
// soon as the backend has no capability entry, BEFORE it consults the
|
|
// model's own tts.voice_cloning override, so without this key a
|
|
// `voice: "profile:<id>"` request is refused with 400 for every audio-cpp
|
|
// model and no model YAML can rescue it, on a backend that ships
|
|
// audio-cpp-chatterbox, whose family advertises cloning and not plain TTS,
|
|
// so a reference clip is the only way to use it at all.
|
|
//
|
|
// Deliberately NOT AudioTransformInputMono16k: the families this backend
|
|
// reaches through AudioTransform are separation and conversion (htdemucs,
|
|
// mel_band_roformer, seed_vc), which refuse any rate but their
|
|
// checkpoint's own and work from the stereo image.
|
|
"audio-cpp": {
|
|
GRPCMethods: []GRPCMethod{
|
|
MethodTTS, MethodTTSStream, MethodAudioTranscription,
|
|
MethodVAD, MethodDiarize, MethodSoundGeneration, MethodAudioTransform,
|
|
},
|
|
PossibleUsecases: []string{
|
|
UsecaseTTS, UsecaseTranscript, UsecaseVAD, UsecaseDiarization,
|
|
UsecaseSoundGeneration, UsecaseAudioTransform,
|
|
},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "audio.cpp native engine: one server for TTS, voice cloning, ASR, forced alignment, VAD, diarization, source separation and music generation; the model's family decides which",
|
|
},
|
|
|
|
// --- TTS backends ---
|
|
"piper": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
Description: "Piper — fast neural TTS optimized for Raspberry Pi",
|
|
},
|
|
"kokoro": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
Description: "Kokoro TTS",
|
|
},
|
|
"coqui": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "Coqui TTS — multi-speaker neural synthesis",
|
|
},
|
|
"kitten-tts": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
Description: "Kitten TTS",
|
|
},
|
|
"outetts": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
Description: "OuteTTS",
|
|
},
|
|
"pocket-tts": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "Pocket TTS — lightweight text-to-speech",
|
|
},
|
|
"qwen-tts": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "Qwen TTS",
|
|
},
|
|
"qwen3-tts-cpp": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS, MethodTTSStream},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "Qwen3 TTS C++ - text-to-speech with streaming, named speakers, voice design and cloning (qwentts.cpp / GGML)",
|
|
},
|
|
"magpie-tts-cpp": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS, MethodTTSStream},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
Description: "Magpie TTS C++ - NVIDIA Magpie TTS Multilingual 357M with 5 baked voices and 9+ languages (magpie-tts.cpp / GGML)",
|
|
},
|
|
"faster-qwen3-tts": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "Faster Qwen3 TTS — accelerated Qwen TTS",
|
|
},
|
|
"fish-speech": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "Fish Speech TTS",
|
|
},
|
|
"neutts": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "NeuTTS — neural text-to-speech",
|
|
},
|
|
"chatterbox": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "Chatterbox TTS",
|
|
},
|
|
"voxcpm": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS, MethodTTSStream},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "VoxCPM TTS with streaming support",
|
|
},
|
|
"omnivoice-cpp": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS, MethodTTSStream},
|
|
PossibleUsecases: []string{UsecaseTTS},
|
|
DefaultUsecases: []string{UsecaseTTS},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "OmniVoice C++ — multilingual TTS with streaming voice cloning and voice design",
|
|
},
|
|
"crispasr": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTranscription, MethodTTS, MethodTTSStream, MethodVAD},
|
|
PossibleUsecases: []string{UsecaseTranscript, UsecaseTTS, UsecaseVAD},
|
|
DefaultUsecases: []string{UsecaseTranscript},
|
|
VoiceCloning: referenceVoiceCloning(),
|
|
Description: "CrispASR GGUF runtime — speech recognition, VAD, and model-dependent TTS",
|
|
},
|
|
|
|
// --- Sound generation backends ---
|
|
"ace-step": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS, MethodSoundGeneration},
|
|
PossibleUsecases: []string{UsecaseTTS, UsecaseSoundGeneration},
|
|
DefaultUsecases: []string{UsecaseSoundGeneration},
|
|
Description: "ACE-Step — music and sound generation",
|
|
},
|
|
"acestep-cpp": {
|
|
GRPCMethods: []GRPCMethod{MethodSoundGeneration},
|
|
PossibleUsecases: []string{UsecaseSoundGeneration},
|
|
DefaultUsecases: []string{UsecaseSoundGeneration},
|
|
Description: "ACE-Step C++ — native sound generation",
|
|
},
|
|
"transformers-musicgen": {
|
|
GRPCMethods: []GRPCMethod{MethodTTS, MethodSoundGeneration},
|
|
PossibleUsecases: []string{UsecaseTTS, UsecaseSoundGeneration},
|
|
DefaultUsecases: []string{UsecaseSoundGeneration},
|
|
Description: "Meta MusicGen via transformers — music generation from text",
|
|
},
|
|
|
|
// --- Any-to-any audio backends ---
|
|
"liquid-audio": {
|
|
GRPCMethods: []GRPCMethod{MethodPredict, MethodPredictStream, MethodAudioTranscription, MethodTTS, MethodAudioToAudioStream, MethodVAD},
|
|
PossibleUsecases: []string{UsecaseChat, UsecaseCompletion, UsecaseTranscript, UsecaseTTS, UsecaseRealtimeAudio, UsecaseVAD},
|
|
DefaultUsecases: []string{UsecaseRealtimeAudio, UsecaseChat, UsecaseTranscript, UsecaseTTS, UsecaseVAD},
|
|
AcceptsAudios: true,
|
|
Description: "LFM2 / LFM2.5-Audio — self-contained any-to-any audio model for the Realtime API; also exposes chat, transcription, TTS and a stub energy-based VAD endpoint",
|
|
},
|
|
|
|
// --- Audio transform backends ---
|
|
"localvqe": {
|
|
GRPCMethods: []GRPCMethod{MethodAudioTransform},
|
|
PossibleUsecases: []string{UsecaseAudioTransform},
|
|
DefaultUsecases: []string{UsecaseAudioTransform},
|
|
// The model is trained on 16 kHz mono speech and its AEC needs the
|
|
// input and the loopback reference in the same shape, so the endpoint
|
|
// keeps folding uploads for this backend.
|
|
AudioTransformInputMono16k: true,
|
|
Description: "LocalVQE — joint AEC, noise suppression, and dereverberation for 16 kHz mono speech",
|
|
},
|
|
|
|
// --- Utility backends ---
|
|
"rerankers": {
|
|
GRPCMethods: []GRPCMethod{MethodRerank},
|
|
PossibleUsecases: []string{UsecaseRerank},
|
|
DefaultUsecases: []string{UsecaseRerank},
|
|
Description: "Cross-encoder reranking models",
|
|
},
|
|
"rfdetr": {
|
|
GRPCMethods: []GRPCMethod{MethodDetect},
|
|
PossibleUsecases: []string{UsecaseDetection},
|
|
DefaultUsecases: []string{UsecaseDetection},
|
|
Description: "RF-DETR object detection",
|
|
},
|
|
"rfdetr-cpp": {
|
|
GRPCMethods: []GRPCMethod{MethodDetect},
|
|
PossibleUsecases: []string{UsecaseDetection},
|
|
DefaultUsecases: []string{UsecaseDetection},
|
|
Description: "RF-DETR C++ object detection",
|
|
},
|
|
"depth-anything": {
|
|
GRPCMethods: []GRPCMethod{MethodDepth, MethodPredict, MethodGenerateImage},
|
|
PossibleUsecases: []string{UsecaseDepth},
|
|
DefaultUsecases: []string{UsecaseDepth},
|
|
AcceptsImages: true,
|
|
Description: "Depth Anything 3 C++ — per-pixel metric depth, camera pose and 3D point cloud",
|
|
},
|
|
|
|
// --- Face and speaker recognition backends ---
|
|
"insightface": {
|
|
GRPCMethods: []GRPCMethod{MethodEmbedding, MethodDetect, MethodFaceVerify, MethodFaceAnalyze},
|
|
PossibleUsecases: []string{UsecaseEmbeddings, UsecaseDetection, UsecaseFaceRecognition},
|
|
DefaultUsecases: []string{UsecaseFaceRecognition},
|
|
AcceptsImages: true,
|
|
Description: "InsightFace — face detection, embedding, verification and attribute analysis",
|
|
},
|
|
"speaker-recognition": {
|
|
GRPCMethods: []GRPCMethod{MethodVoiceVerify, MethodVoiceEmbed, MethodVoiceAnalyze},
|
|
PossibleUsecases: []string{UsecaseSpeakerRecognition},
|
|
DefaultUsecases: []string{UsecaseSpeakerRecognition},
|
|
Description: "Speaker recognition — voice identity verification and analysis",
|
|
},
|
|
"voice-detect": {
|
|
GRPCMethods: []GRPCMethod{MethodVoiceVerify, MethodVoiceEmbed, MethodVoiceAnalyze},
|
|
PossibleUsecases: []string{UsecaseSpeakerRecognition},
|
|
DefaultUsecases: []string{UsecaseSpeakerRecognition},
|
|
Description: "voice-detect.cpp: C++/ggml speaker embedding, verification and voice analysis (age/gender/emotion)",
|
|
},
|
|
"face-detect": {
|
|
GRPCMethods: []GRPCMethod{MethodEmbedding, MethodDetect, MethodFaceVerify, MethodFaceAnalyze},
|
|
PossibleUsecases: []string{UsecaseEmbeddings, UsecaseDetection, UsecaseFaceRecognition},
|
|
DefaultUsecases: []string{UsecaseFaceRecognition},
|
|
AcceptsImages: true,
|
|
Description: "face-detect.cpp: C++/ggml face detection, embedding, verification and attribute analysis",
|
|
},
|
|
"silero-vad": {
|
|
GRPCMethods: []GRPCMethod{MethodVAD},
|
|
PossibleUsecases: []string{UsecaseVAD},
|
|
DefaultUsecases: []string{UsecaseVAD},
|
|
Description: "Silero VAD — voice activity detection",
|
|
},
|
|
}
|
|
|
|
// NormalizeBackendName converts backend names to the canonical hyphenated form
|
|
// used in gallery entries (e.g., "llama.cpp" → "llama-cpp").
|
|
func NormalizeBackendName(backend string) string {
|
|
return strings.ReplaceAll(backend, ".", "-")
|
|
}
|
|
|
|
// galleryChannelSuffixes are the release-channel suffixes appended to a backend
|
|
// name in the gallery ("llama-cpp" vs "llama-cpp-development" vs
|
|
// "llama-cpp-quantization"). They carry no engine information, so they are
|
|
// stripped before any family or capability lookup falls back.
|
|
var galleryChannelSuffixes = []string{"-development", "-quantization"}
|
|
|
|
// galleryHardwarePrefixes are the acceleration prefixes the gallery prepends
|
|
// when it publishes one concrete backend image per hardware capability behind a
|
|
// meta name: "cpu-localvqe", "vulkan-localvqe", "cuda12-audio-cpp",
|
|
// "metal-darwin-arm64-llama-cpp". They carry no engine information either, and
|
|
// an operator may pin any of them in a model config's `backend:`.
|
|
//
|
|
// This is the exhaustive set present in backend/index.yaml, longest first so
|
|
// "cuda13-nvidia-l4t-arm64-" is tried before "cuda13-" and
|
|
// "intel-sycl-f16-" before "intel-". Stripping is a FALLBACK only (see
|
|
// GetBackendCapability), so a backend whose real name happened to start with
|
|
// one of these would still be found by its exact name first.
|
|
var galleryHardwarePrefixes = []string{
|
|
"cuda13-nvidia-l4t-arm64-",
|
|
"metal-darwin-arm64-",
|
|
"nvidia-l4t-arm64-",
|
|
"intel-sycl-f16-",
|
|
"intel-sycl-f32-",
|
|
"nvidia-l4t-",
|
|
"vulkan-",
|
|
"cuda12-",
|
|
"cuda13-",
|
|
"metal-",
|
|
"intel-",
|
|
"rocm-",
|
|
"cpu-",
|
|
}
|
|
|
|
// stripBackendVariant reduces a concrete gallery backend name to the meta name
|
|
// its capabilities are registered under: "vulkan-localvqe-development" becomes
|
|
// "localvqe". Returns the input unchanged when nothing matches.
|
|
//
|
|
// Both halves are needed. A pinned variant carries a hardware prefix, a release
|
|
// channel carries a suffix, and backend/index.yaml ships names with both.
|
|
func stripBackendVariant(name string) string {
|
|
for _, suffix := range galleryChannelSuffixes {
|
|
if strings.HasSuffix(name, suffix) {
|
|
name = strings.TrimSuffix(name, suffix)
|
|
break
|
|
}
|
|
}
|
|
for _, prefix := range galleryHardwarePrefixes {
|
|
if strings.HasPrefix(name, prefix) {
|
|
return strings.TrimPrefix(name, prefix)
|
|
}
|
|
}
|
|
return name
|
|
}
|
|
|
|
// IsLlamaCppBackend reports whether a backend name refers to a build of the
|
|
// llama.cpp gRPC server. The gallery ships one concrete backend per hardware
|
|
// capability ("vulkan-llama-cpp", "cuda12-llama-cpp", "metal-llama-cpp", ...)
|
|
// behind the "llama-cpp" meta name, and an operator may pin any of them in a
|
|
// model config. They all run the same server, so anything gated on "is this
|
|
// llama.cpp" must accept the whole family: an exact match against "llama-cpp"
|
|
// silently skips every pinned variant (see #10945, where skipping the media
|
|
// marker probe broke all vision requests).
|
|
//
|
|
// The empty name matches too: it is the GGUF auto-detect path, which resolves
|
|
// to llama.cpp.
|
|
//
|
|
// ik-llama.cpp is deliberately excluded. It is a separate engine with its own
|
|
// gRPC server that happens to share the "-llama-cpp" suffix.
|
|
func IsLlamaCppBackend(backend string) bool {
|
|
name := NormalizeBackendName(backend)
|
|
if name == "" {
|
|
return true
|
|
}
|
|
for _, suffix := range galleryChannelSuffixes {
|
|
name = strings.TrimSuffix(name, suffix)
|
|
}
|
|
if strings.HasSuffix(name, "ik-llama-cpp") {
|
|
return false
|
|
}
|
|
return name == "llama-cpp" || strings.HasSuffix(name, "-llama-cpp")
|
|
}
|
|
|
|
// nonLlamaSamplerBackends lists backends whose native sampler defaults differ
|
|
// from llama.cpp's, so LocalAI must NOT inject llama.cpp's top_k=40 default for
|
|
// them (issue #6632). mlx_lm's intended default is top_k=0 (disabled) and mlx
|
|
// does not remap 0->40, so shipping 40 silently changes sampling for clients
|
|
// that omit top_k. Leaving TopK nil lets the wire value default to 0.
|
|
//
|
|
// This is intentionally a small allow-list of KNOWN non-llama backends: empty
|
|
// and unknown backends fall through to the llama.cpp default to preserve the
|
|
// GGUF auto-detect path's behavior.
|
|
var nonLlamaSamplerBackends = map[string]struct{}{
|
|
"mlx": {},
|
|
"mlx-vlm": {},
|
|
"mlx-distributed": {},
|
|
}
|
|
|
|
// UsesLlamaSamplerDefaults reports whether a backend should receive llama.cpp's
|
|
// sampler defaults (e.g. top_k=40). Empty/unknown backends return true so the
|
|
// GGUF auto-detect path (which resolves to llama.cpp) keeps today's behavior;
|
|
// only the known non-llama backends in nonLlamaSamplerBackends return false.
|
|
func UsesLlamaSamplerDefaults(backend string) bool {
|
|
if backend == "" {
|
|
return true
|
|
}
|
|
_, isNonLlama := nonLlamaSamplerBackends[NormalizeBackendName(backend)]
|
|
return !isNonLlama
|
|
}
|
|
|
|
// UsesLlamaCppServingOptions reports whether a backend understands llama.cpp's
|
|
// serving-tuning model options - the free-form option strings cache_reuse /
|
|
// n_cache_reuse (cross-request KV-prefix reuse) and parallel / n_parallel
|
|
// (concurrent slots). These are llama.cpp server flags; LocalAI injects them as
|
|
// defaults, but a backend that strictly validates its options (e.g.
|
|
// longcat-video) rejects an unknown one with "unknown model option(s)" at
|
|
// LoadModel. Only the llama.cpp backend - and the empty/auto-detect case, which
|
|
// resolves to llama.cpp from a GGUF file, mirroring how llamaCppDefaults is
|
|
// registered - should receive them.
|
|
//
|
|
// This is an allow-list on purpose (unlike UsesLlamaSamplerDefaults's
|
|
// deny-list): these options are meaningful to no other backend, so a new
|
|
// backend defaults to NOT getting them rather than breaking the same way.
|
|
func UsesLlamaCppServingOptions(backend string) bool {
|
|
switch NormalizeBackendName(backend) {
|
|
case "", "llama-cpp":
|
|
return true
|
|
}
|
|
return false
|
|
}
|
|
|
|
// GetBackendCapability returns the capability info for a backend, or nil if unknown.
|
|
// Handles backend name normalization.
|
|
//
|
|
// A PINNED GALLERY VARIANT RESOLVES TO ITS META NAME. The gallery ships one
|
|
// image per hardware capability ("cpu-localvqe", "vulkan-localvqe",
|
|
// "metal-localvqe") and an operator may put any of them in a model's
|
|
// `backend:`. They are the same engine, so an exact-match-only lookup silently
|
|
// downgraded every pinned model to "unknown backend": vulkan-localvqe lost the
|
|
// 16 kHz mono fold that /audio/transform used to apply unconditionally and
|
|
// started failing inside LocalVQE, and a pinned audio-cpp variant would lose
|
|
// its voice-cloning contract the same way. Same class of bug as #10945, same
|
|
// answer as IsLlamaCppBackend.
|
|
//
|
|
// Exact match FIRST, so a backend genuinely registered under a variant-looking
|
|
// name keeps its own entry and stripping can never shadow it.
|
|
func GetBackendCapability(backend string) *BackendCapability {
|
|
capability, _ := resolveBackendCapability(backend)
|
|
return capability
|
|
}
|
|
|
|
// resolveBackendCapability is GetBackendCapability plus the key the entry was
|
|
// found under. Callers that then branch on backend identity MUST use that key,
|
|
// not the name they passed in. VoiceCloningForModel is the reason this exists:
|
|
// its per-backend switch encodes which model variants of a backend can clone,
|
|
// and keying it on the caller's spelling meant "cuda12-vibevoice-cpp" resolved
|
|
// the capability by stripping but missed the "vibevoice-cpp" case, falling
|
|
// through to the permissive default and advertising cloning for the 0.5B model
|
|
// that cannot do it.
|
|
func resolveBackendCapability(backend string) (*BackendCapability, string) {
|
|
name := NormalizeBackendName(backend)
|
|
if cap, ok := BackendCapabilities[name]; ok {
|
|
return &cap, name
|
|
}
|
|
if base := stripBackendVariant(name); base != name {
|
|
if cap, ok := BackendCapabilities[base]; ok {
|
|
return &cap, base
|
|
}
|
|
}
|
|
return nil, name
|
|
}
|
|
|
|
// AudioTransformRequiresMono16kInput reports whether /audio/transform must fold
|
|
// uploads to 16 kHz mono before handing them to this backend.
|
|
//
|
|
// False for an unknown backend, which is the safe answer: an unregistered
|
|
// backend gets its upload unchanged, so a model that needs the file intact
|
|
// (source separation, voice conversion at 44.1 kHz) works without an entry
|
|
// here, and one that needs the fold cannot get it by accident.
|
|
//
|
|
// Pinned gallery variants are covered: GetBackendCapability strips the hardware
|
|
// prefix and the release-channel suffix, so cpu-localvqe, vulkan-localvqe and
|
|
// metal-localvqe all fold exactly as "localvqe" does. They must, because the
|
|
// usecase gate does not stand in for this one: naming a model explicitly makes
|
|
// BuildFilteredFirstAvailableDefaultModel return before it filters.
|
|
func AudioTransformRequiresMono16kInput(backend string) bool {
|
|
capability := GetBackendCapability(backend)
|
|
return capability != nil && capability.AudioTransformInputMono16k
|
|
}
|
|
|
|
// llmAutoLoadUsecases are the usecases that mark a backend able to serve a
|
|
// text/LLM GGUF model. A GGUF model that declares no explicit backend must only
|
|
// be auto-tried against backends carrying one of these usecases - never against
|
|
// audio/codec/image backends (e.g. opus) that happen to be installed alongside
|
|
// it (see issue #9287).
|
|
var llmAutoLoadUsecases = []string{
|
|
UsecaseChat,
|
|
UsecaseCompletion,
|
|
UsecaseEdit,
|
|
UsecaseEmbeddings,
|
|
}
|
|
|
|
// isLLMCapableForAutoLoad reports whether the named backend is known to serve
|
|
// text/LLM models, for pkg/model's GGUF backend auto-detection (#9287). Backends
|
|
// absent from the capability table are treated as not LLM-capable.
|
|
func isLLMCapableForAutoLoad(name string) bool {
|
|
capability := GetBackendCapability(name)
|
|
if capability == nil {
|
|
return false
|
|
}
|
|
for _, u := range capability.PossibleUsecases {
|
|
if slices.Contains(llmAutoLoadUsecases, u) {
|
|
return true
|
|
}
|
|
}
|
|
return false
|
|
}
|
|
|
|
func init() {
|
|
// Wire the LLM-capability filter into pkg/model's GGUF backend
|
|
// auto-detection. pkg/model is a lower-level package and must not import
|
|
// core/config (that would form a core/config -> pkg/model -> core/config
|
|
// import cycle), so core/config registers the predicate here instead (#9287).
|
|
model.RegisterLLMCapableBackendFunc(isLLMCapableForAutoLoad)
|
|
}
|
|
|
|
// VoiceCloningForModel returns the reference-audio contract only when the
|
|
// installed model variant can honor it. Several backends serve both Base
|
|
// (voice cloning) and CustomVoice/VoiceDesign models, so backend name alone is
|
|
// deliberately insufficient. Operators with custom filenames can opt in or
|
|
// out explicitly with tts.voice_cloning; the model option spelling remains a
|
|
// compatibility fallback for configurations created before the typed field.
|
|
func VoiceCloningForModel(cfg *ModelConfig) *VoiceCloningCapability {
|
|
if cfg == nil {
|
|
return nil
|
|
}
|
|
capability, backend := resolveBackendCapability(cfg.Backend)
|
|
if capability == nil || capability.VoiceCloning == nil {
|
|
return nil
|
|
}
|
|
if cfg.VoiceCloning != nil {
|
|
if !*cfg.VoiceCloning {
|
|
return nil
|
|
}
|
|
return cloneVoiceCloningCapability(capability.VoiceCloning)
|
|
}
|
|
|
|
if enabled, explicit := voiceCloningOverride(cfg.Options); explicit {
|
|
if !enabled {
|
|
return nil
|
|
}
|
|
return cloneVoiceCloningCapability(capability.VoiceCloning)
|
|
}
|
|
|
|
identity := strings.ToLower(strings.Join([]string{cfg.Name, cfg.Model, strings.Join(cfg.Options, " ")}, " "))
|
|
supported := false
|
|
switch backend {
|
|
case "qwen3-tts-cpp", "qwen-tts", "vllm-omni":
|
|
supported = strings.Contains(identity, "base") || strings.Contains(identity, "voiceclone") || strings.Contains(identity, "voice_clone")
|
|
case "vibevoice-cpp":
|
|
// Realtime 0.5B consumes a precomputed .gguf voice prompt; the 1.5B
|
|
// path consumes raw WAV references per request.
|
|
supported = strings.Contains(identity, "1.5b")
|
|
case "coqui":
|
|
supported = strings.Contains(identity, "xtts") || strings.Contains(identity, "your_tts")
|
|
case "crispasr":
|
|
supported = strings.Contains(identity, "f5-tts") || strings.Contains(identity, "f5_tts")
|
|
case "llama-cpp":
|
|
// llama.cpp is overwhelmingly a text-LLM backend that happens to also
|
|
// serve Qwen3-TTS, so the permissive default below would advertise
|
|
// reference-audio cloning on every GGUF chat model in the gallery.
|
|
// Narrow on the declared usecase rather than the model name: the TTS
|
|
// checkpoints are the only llama-cpp models that carry
|
|
// known_usecases: [tts], name matching would have to guess at
|
|
// third-party GGUF repacks, and "base" (the substring the Qwen and
|
|
// vLLM cases key on) is a routine word in text-model names.
|
|
//
|
|
// Deliberately reads the declared bit instead of HasUsecases, which
|
|
// falls through to GuessUsecases and would hand the decision to a
|
|
// heuristic that never had a llama.cpp TTS model in mind.
|
|
supported = cfg.KnownUsecases != nil && (*cfg.KnownUsecases&FLAG_TTS) == FLAG_TTS
|
|
default:
|
|
supported = true
|
|
}
|
|
if !supported {
|
|
return nil
|
|
}
|
|
return cloneVoiceCloningCapability(capability.VoiceCloning)
|
|
}
|
|
|
|
func voiceCloningOverride(options []string) (enabled, explicit bool) {
|
|
for _, option := range options {
|
|
parts := strings.FieldsFunc(option, func(r rune) bool { return r == ':' || r == '=' })
|
|
if len(parts) != 2 || !strings.EqualFold(strings.TrimSpace(parts[0]), "voice_cloning") {
|
|
continue
|
|
}
|
|
switch strings.ToLower(strings.TrimSpace(parts[1])) {
|
|
case "true", "1", "yes", "on":
|
|
return true, true
|
|
case "false", "0", "no", "off":
|
|
return false, true
|
|
}
|
|
}
|
|
return false, false
|
|
}
|
|
|
|
func cloneVoiceCloningCapability(capability *VoiceCloningCapability) *VoiceCloningCapability {
|
|
if capability == nil {
|
|
return nil
|
|
}
|
|
clone := *capability
|
|
clone.AcceptedAudioFormats = slices.Clone(capability.AcceptedAudioFormats)
|
|
return &clone
|
|
}
|
|
|
|
// PossibleUsecasesForBackend returns all usecases a backend can support.
|
|
// Returns nil if the backend is unknown.
|
|
func PossibleUsecasesForBackend(backend string) []string {
|
|
if cap := GetBackendCapability(backend); cap != nil {
|
|
return cap.PossibleUsecases
|
|
}
|
|
return nil
|
|
}
|
|
|
|
// DefaultUsecasesForBackendCap returns the conservative default usecases.
|
|
// Returns nil if the backend is unknown.
|
|
func DefaultUsecasesForBackendCap(backend string) []string {
|
|
if cap := GetBackendCapability(backend); cap != nil {
|
|
return cap.DefaultUsecases
|
|
}
|
|
return nil
|
|
}
|
|
|
|
// IsValidUsecaseForBackend checks whether a usecase is in a backend's possible set.
|
|
// Returns true for unknown backends (permissive fallback).
|
|
func IsValidUsecaseForBackend(backend, usecase string) bool {
|
|
cap := GetBackendCapability(backend)
|
|
if cap == nil {
|
|
return true // unknown backend — don't restrict
|
|
}
|
|
return slices.Contains(cap.PossibleUsecases, usecase)
|
|
}
|
|
|
|
// AllBackendNames returns a sorted list of all known backend names.
|
|
func AllBackendNames() []string {
|
|
names := make([]string, 0, len(BackendCapabilities))
|
|
for name := range BackendCapabilities {
|
|
names = append(names, name)
|
|
}
|
|
slices.Sort(names)
|
|
return names
|
|
}
|