mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-03 03:24:34 -04:00
* feat(parakeet-cpp): load diarization and CED models and companions
Repin PARAKEET_VERSION to parakeet.cpp PR #75's head, which adds
parakeet_capi_model_kind (ABI v8). Bind the new diarization, sound
event and combined scene stream C symbols through the same
purego.Dlsym probe pattern already used for the batched JSON entry
point, so the backend still loads against an older libparakeet.so.
Load now classifies the loaded GGUF by role (ASR, diarization or
sound) via parakeet_capi_model_kind and can load up to two companion
models from Options[] (asr_model:, diarization_model:, sound_model:,
paths resolved against opts.ModelPath), verifying each companion's
kind and freeing every context opened so far on any failure. Free
releases the primary and every companion. AudioTranscription now
names the loaded role when it is not ASR instead of a generic model
not loaded error. The dynamic batcher starts only when an ASR context
ends up loaded, primary or companion.
This is groundwork only: the Diarize and SoundDetection RPCs and the
live scene stream that actually use these new roles land in later
commits.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): reset role fields on a failed companion load
loadRoles' freeLoaded only released the C contexts it had opened; it
left ctxPtr/diarCtx/tagCtx and companions pointing at those now-freed
contexts, so a later Free() on the same instance would double-free.
Zero all four alongside the CppFree calls.
Also route AudioTranscriptionStream and AudioTranscriptionLive through
notASRError when ctxPtr is unset but a diarization or sound model is
loaded, matching AudioTranscription: both used to return the generic
model-not-loaded error instead of naming the loaded role.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): add speaker diarization
Implement the Diarize RPC for the parakeet-cpp Go backend, wired to
Nemotron-3-Diarization through libparakeet.so's diarization C-API.
Plain diarization uses parakeet_capi_diarize_pcm; when include_text is
set and an ASR companion is loaded, parakeet_capi_transcribe_and_
diarize_json fills each segment's text instead. Speaker labels are the
decimal index, or "unknown" for -1 (no diarized speaker overlaps).
min_duration_off merges same-speaker segments across a short gap
before min_duration_on drops the segments still too short, then ids
are renumbered. num_speakers/min_speakers/max_speakers/clustering_
threshold have no Sortformer equivalent and are logged at debug
instead of rejected.
Verified against the real Nemotron-3-Diarization + parakeet-tdt_ctc-
110m checkpoints on the two_speakers.wav fixture: correct A-B-A-B
speaker segmentation and matching speaker-attributed transcripts.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): add sound event detection
Wire the SoundDetection RPC to the CED tagger context (p.tagCtx)
loaded by Task 1's role classification. It runs the whole clip
through a one-shot parakeet_capi_sound_stream_* session (window
10s, hop 10s, top_k set to the tagger's class count so every
drained window carries a full score list), averages each class's
score across the drained windows, sorts descending, then applies
the request's threshold and top_k (0 keeps every class).
No tagCtx returns FailedPrecondition; a libparakeet.so missing the
sound_stream symbols returns Unimplemented. Every C call runs under
engineMu, and the stream is always freed, even when a feed or drain
call fails partway through.
Verified against a real ced-tiny-q8_0.gguf on the rooster.wav demo
clip: "Chicken, rooster" tops the list at score 0.91.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): cancel sound detection mid-feed, shrink the lock
SoundDetection now checks ctx before each 10 s feed slice (mirroring
driver.go's feedSlices) and returns Canceled if the caller gave up,
so a long clip can be interrupted instead of feeding to completion
regardless. The stream is still freed on every path, cancellation
included.
Also narrow engineMu to the C calls: the drained JSON document is
now decoded after the lock is released, splitting soundStreamScores
into a locked soundStreamDrain (opts, begin, feed, drain, free) and
an unlocked json.Unmarshal.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): stream speaker and sound events during live transcription
Add two additive proto fields, LiveSpeakerSegment and LiveSoundEvent,
repeated on TranscriptLiveResponse. When a diarization or sound
companion model is loaded, AudioTranscriptionLive now runs a no-ASR
scene stream (parakeet_capi_scene_stream_begin) beside the ASR
streaming session, feeding it the same PCM slices and forwarding any
closed speaker or sound events alongside the matching ASR delta, or
on their own when a slice has no ASR output.
The scene stream is freed and reopened on a mid-stream Config reset,
flushed with is_last before the closing FinalResult, and degrades
gracefully (a warning, not an error) when begin or a later feed call
fails, so live transcription keeps working ASR-only. Existing live
behavior is unchanged when no companion is configured, and no scene
C call is made in that case.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): keep scene events off the ASR critical path in live
Emit each slice's ASR result right after the ASR feed, before the
scene feed for that slice runs, so a companion diarization/sound
model never adds scene compute latency in front of the delta or
<EOU> that drives realtime turn detection. Closed speakers/sounds go
out afterward as their own response, so a slice with both now
produces two responses, ASR first. The live feed log line now
reports ASR and scene wall time separately.
Re-check the diarization/sound contexts a scene stream was begun
with against the live contexts before every feed, under the same
lock: Free() can race between an ASR feed and the matching scene
feed and free the model the stream borrows. A mismatch now returns
without touching the C side. Freeing the stream itself stays
unconditional; the scene stream's destructor only releases its own
buffers and never touches the borrowed contexts.
Also recover a panicking stub inside the live test goroutine instead
of crashing the test binary, and reset the live decode-lag tracker on
a mid-stream config reset, matching what its own comment already
promised.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(realtime): surface live speaker and sound events
Carry the backend's closed speaker segments and sound events
(TranscriptLiveResponse fields 7/8) through LiveTranscriptionEvent
as LiveSpeakerSegment/LiveSoundEvent (nanoseconds mapped to
seconds), and forward them from the semantic_vad live path.
Each speaker segment emits
conversation.item.input_audio_transcription.segment with speaker,
start, end and empty text under the turn's item id. Each sound
event emits conversation.item.sound_detection with one tag
(label, score = peak, index) and the event's new optional
start/end seconds fields, omitted when unset so the existing
unary/windowed sound-detection path is unaffected.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(realtime): keep start/end on a zero-second transcription segment
ConversationItemInputAudioTranscriptionSegmentEvent.Start/End used
omitempty, so a speaker segment starting at 0.0s dropped its
"start" key. Nothing emitted this event before the live scene-event
path, so drop omitempty: the segment always carries real times.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(gallery): add parakeet-cpp diarization, CED and realtime scene models
Add gallery entries for the new parakeet-cpp capabilities: standalone
Nemotron-3-Diarization, the same paired with the Parakeet TDT+CTC
110M ASR model for speaker-attributed text, CED-Tiny and CED-Base
sound classifiers, and a realtime scene bundle combining the
streaming EOU ASR model with diarization and sound companions.
SHA256 taken from the Hub API; licenses from each model card
(openmdw-1.1 for Nemotron-3-Diarization, apache-2.0 for CED,
cc-by-4.0 for the Parakeet ASR models).
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs: document parakeet-cpp diarization, sound detection and live scene events
Cover the new parakeet-cpp capabilities across the feature pages:
Nemotron-3-Diarization as a diarization backend (with and without
speaker text, the ignored speaker-count hints, the Sortformer
voice-like-sound quirk), CED as a sound classification backend, the
asr_model/diarization_model/sound_model/diarization_latency companion
options, and the realtime live speaker/sound events (event shapes,
the speech-turn-only limitation, and using this or
pipeline.sound_detection but not both).
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(gallery): correct the realtime-scene license and wording nits
parakeet-cpp-realtime-scene mistakenly copied cc-by-4.0 from the
existing realtime_eou_120m-v1 entry; the model card lists the NVIDIA
open model license instead. Switch to the gallery's usual spelling
for that license and keep the diarization/CED licenses called out in
the description.
Also: audio-diarization.md now says getting per-segment text needs
both an asr_model companion and include_text=true on the request, and
audio-to-text.md's option table reads "Use on" (a pairing the loader
does not enforce) instead of "Allowed on".
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): reject a companion role that duplicates the primary's
loadRoles let a companion option (asr_model:/diarization_model:/
sound_model:) assign into a role field the primary already occupied,
for example asr_model: on an already-ASR primary. The companion's
context silently overwrote ctxPtr/diarCtx/tagCtx, and Free() only
walks those three fields, so the original primary context was never
freed again.
Reject a companion whose role the primary already holds before its
GGUF is even loaded, freeing everything loadRoles opened so far, the
same way a wrong-kind companion is already rejected.
Also warn, rather than silently fall through, when
parakeet_capi_model_kind reports PARAKEET_MODEL_KIND_NONE for a
successfully loaded primary; the primary is still treated as ASR,
matching today's behavior.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): cap live scene sound score retention
sceneBegin started the live diarization/sound companion stream with
the C API's default sound options, whose top_k keeps 5 scores per
window forever until drained. The live scene path never drains sound
scores (only the offline SoundDetection RPC does, with its own fresh
stream), so this window queue on the C side grew for the whole
session's lifetime.
Set opts.Sound.TopK = 0 before starting the scene stream: this
disables score retention while leaving sound event detection (onset/
offset), which the live path actually consumes, unaffected.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): merge diarization segments per speaker, harden Diarize
mergeCloseSegments only compared neighbors in the single start-sorted
segment list, so two same-speaker segments never merged once another
speaker's turn fell between them (A, B, A): the short B segment broke
the adjacency the merge relied on. Group segments by speaker first,
merge within each speaker's own start-ordered run, then re-sort the
result by start so interleaved speakers come back out in timeline
order.
Also harden Diarize's entry points the same way streamFeedDoc/
sceneFeed already are: diarizeCall re-checks p.diarCtx (and, on the
include_text path, p.ctxPtr) under engineMu right before the C call,
so a Free() racing between Diarize's own checks and the lock can no
longer reach the C side with a freed context. When the include_text
call returns NULL, last_error is now read from both contexts and
whichever came back non-empty is reported, since either side of the
pairing can be the one that failed. A WAV decode failure is reported
as InvalidArgument instead of an unwrapped/untyped error.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): harden SoundDetection's engine checks
soundStreamDrain ran every C call under engineMu but never re-checked
p.tagCtx there, so a Free() racing between SoundDetection's own
tagCtx==0 check and this lock could still reach the C side with a
freed context. Re-check p.tagCtx under the lock and return
ModelNotLoaded when it was cleared, mirroring diarizeCall's own
re-check. A WAV decode failure is now reported as InvalidArgument
instead of an unwrapped/untyped error.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* test(parakeet-cpp): cover a mid-session scene feed failure
feedSlicesScene already degrades gracefully when a scene feed call
fails mid-session: it frees the broken stream and carries the ASR-only
session forward. Add a spec covering that path end to end: the scene
stream is freed exactly once, later audio slices still produce ASR
responses, and no speaker/sound events appear before or after the
failure.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* docs: fix the parakeet-cpp companion role table and realtime scene docs
audio-to-text.md's companion option table read "Use on" with a note
that the loader did not enforce the pairing; it now rejects a
companion whose role duplicates the primary's, so restore the
"Allowed on" wording and describe the real enforcement.
openai-realtime.md's live speaker/sound section claimed a mid-stream
session.update resets the companion stream and that it flushes on
session close; neither happens, since the realtime core opens one
live stream (and so one scene stream) per speech turn and closes it
at that turn's commit, with no mid-stream Config in between. Document
that lifecycle instead, state precisely that start/end are seconds
from the start of the turn's own audio, and note that the diarization
model starts a fresh session every turn, so a speaker index is only
meaningful within one turn. The example sound tag ("Rooster", index
17) did not match any real CED label; index 17 in ced-tiny-q8_0.gguf
is "Baby laughter". Replaced with "Chicken, rooster" at its real
index, 99.
Assisted-by: Claude:claude-sonnet-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(parakeet-cpp): use CED's real index for Chicken, rooster
The scene feed comment and the live test's canned document gave
"Chicken, rooster" index 365. In CED's AudioSet label list it is 99,
which is also what the realtime docs show.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(realtime): call the test event accessor
The scene-event tests range over a method instead of its returned slice.
Call the synchronized accessor so the OpenAI test package compiles.
Assisted-by: Codex:gpt-6
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(parakeet-cpp): pin parakeet.cpp master with sound events
mudler/parakeet.cpp#75 (sound events, scene stream, model kinds) and
#74 (the missing <algorithm> include that broke the image builds) are
on master now. Pin 6dea76a instead of the #75 PR head, and update the
header comment the bump bot reads.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* chore(parakeet-cpp): pin parakeet.cpp with ced.cpp on main
parakeet.cpp #76 moved its ced.cpp submodule from the head of
localai-org/ced.cpp#3 (a branch-only commit) to ced.cpp main, where
#3 landed with an identical tree. Pin 623a968 so the image builds no
longer depend on that branch.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(transcription): carry speaker labels on words and streamed segments
A diarizing backend could label transcript segments, but two paths
dropped the label: TranscriptWord had no speaker field, so live
transcription words and word-level timestamps could not carry one, and
the stream=true transcript.text.done event left the speaker out of
its segments.
TranscriptWord gains an optional speaker (proto field 4, additive).
It flows through the live event and result mapping, the JSON word
output of the endpoint and the CLI, and transcript.text.done now
includes a segment's speaker when there is one. Empty labels are
omitted, so responses without diarization are unchanged.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
(cherry picked from commit 2f0049f979)
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(importers): detect the parakeet.cpp diarization GGUF
The Nemotron-3-Diarization GGUFs are published in
mudler/parakeet-cpp-gguf as nemotron-3-diarization-<quant>.gguf. The
parakeet-cpp importer did not recognise that name, so a direct
`local-ai models import` of the file fell through to another importer.
A direct URL to the file now imports with the diarization usecase. A
repo import still picks ASR weights when the repo also ships the
diarization model, and falls back to the diarization weights only
when there are no others.
Ported from #12323.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(config): advertise diarization and sound detection for parakeet-cpp
The capability table listed parakeet-cpp as transcription only, though
the backend now answers Diarize (Nemotron-3-Diarization) and
SoundDetection (CED) depending on the model kind it loads.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(parakeet-cpp): label transcript segments with the diarization companion
A diarization_model companion only fed live speaker events and
Diarize; /v1/audio/transcriptions ignored it.
With the companion attached and diarize=true (the OpenAI endpoint's
default), unary transcription now labels each segment with its
speaker and splits segments at speaker turns; with word timestamps
each word carries its speaker. The stream=true final result labels
each utterance with the speaker who said most of it. Both use the
checkpoint's own diarization over the whole clip, as NeMo's diarize()
does. Words take the speaker whose segments overlap them most, or the
nearest segment within 0.5 s, the same rule as parakeet.cpp's
speaker-attributed ASR.
Docs: the diarization_model row and a paragraph on transcript
speakers; Nemotron-3-Diarization handles up to 8 speakers.
Ported from #12323.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(realtime): speaker segments from committed-turn transcription
Speaker events reached a realtime session only from the live
semantic_vad path, which needs a cache-aware streaming transcription
model. Committed-turn transcription (server_vad, or any offline
model) always asked the backend for diarize=false and dropped the
segments' speakers.
pipeline.diarization (off by default) asks the transcription model for
speaker labels on each committed turn and emits every labelled segment
as a conversation.item.input_audio_transcription.segment event, with
its text, before the turn's completed event. It is opt-in because some
backends fail a diarization request they cannot serve.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(gallery): add parakeet-cpp-realtime-scene-tdt
parakeet-cpp-realtime-scene pairs the streaming EOU model with the
diarization and CED companions; its speaker and sound events need a
cache-aware streaming model. This entry does the same with Parakeet
TDT 0.6B v3 (multilingual, offline) for realtime under server_vad:
set it as both transcription and sound_detection and turn on
pipeline.diarization, and each committed turn gets speaker segments
and sound tags from one parakeet-cpp backend.
Files and sha256 match the Hub and are shared with the existing TDT v3,
diarization and CED-Tiny entries. A real-model spec checks the
combination on a clip with two speakers and a rooster: A-B-A-B speaker
turns, and "Chicken, rooster" among the sound tags. The test loader
now binds the sound entry points like main.go.
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* feat(gallery): add CED-Base variants of the parakeet-cpp scene models
parakeet-cpp-realtime-scene and parakeet-cpp-realtime-scene-tdt ship
with CED-Tiny. The -base variants use CED-Base (86M), which tags sounds
more confidently (on the rooster clip "Crowing" 0.65 against 0.49 for
Tiny).
Measured on CPU over a 37 s clip: the live diarization + sound stream
runs at 0.125 of real time with CED-Base against 0.103 with CED-Tiny,
because diarization dominates; sound detection per committed turn costs
0.031 against 0.005. The realtime docs list both and note that any CED
size works as sound_model.
Files and sha256 match the Hub and are shared with the existing
parakeet-cpp-ced-base entry. The TDT variant passes the real-model
scene spec with CED-Base (A-B-A-B speakers, "Chicken, rooster" found).
Assisted-by: Claude:claude-opus-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
* fix(config): register pipeline.diarization in the config metadata
TestAllFieldsHaveRegistryEntries fails on the branch because the new
pipeline.diarization field has no registry entry. Add one so the model
editor shows it as a toggle next to the sound detection options.
Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
---------
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
1528 lines
57 KiB
Protocol Buffer
1528 lines
57 KiB
Protocol Buffer
syntax = "proto3";
|
|
|
|
option go_package = "github.com/go-skynet/LocalAI/pkg/grpc/proto";
|
|
option java_multiple_files = true;
|
|
option java_package = "io.skynet.localai.backend";
|
|
option java_outer_classname = "LocalAIBackend";
|
|
|
|
package backend;
|
|
|
|
service Backend {
|
|
rpc Health(HealthMessage) returns (Reply) {}
|
|
rpc Free(HealthMessage) returns (Result) {}
|
|
rpc Predict(PredictOptions) returns (Reply) {}
|
|
rpc LoadModel(ModelOptions) returns (Result) {}
|
|
rpc PredictStream(PredictOptions) returns (stream Reply) {}
|
|
rpc Embedding(PredictOptions) returns (EmbeddingResult) {}
|
|
rpc GenerateImage(GenerateImageRequest) returns (Result) {}
|
|
rpc UpscaleImage(UpscaleImageRequest) returns (Result) {}
|
|
rpc GenerateVideo(GenerateVideoRequest) returns (Result) {}
|
|
rpc Generate3D(Generate3DRequest) returns (Result) {}
|
|
rpc Animate3D(Animate3DRequest) returns (Result) {}
|
|
rpc AudioTranscription(TranscriptRequest) returns (TranscriptResult) {}
|
|
rpc AudioTranscriptionStream(TranscriptRequest) returns (stream TranscriptStreamResponse) {}
|
|
// AudioTranscriptionLive is the bidirectional live-microphone ASR RPC. The
|
|
// first message MUST carry a Config; subsequent messages carry Audio frames
|
|
// (mono float PCM at config.sample_rate, 16 kHz default). After a
|
|
// successful open the backend replies with a single ready ack
|
|
// (TranscriptLiveResponse{ready:true}); backends or models without
|
|
// cache-aware streaming support return UNIMPLEMENTED instead. Newly
|
|
// finalized text streams back as deltas; eou=true marks the model's
|
|
// end-of-utterance token. One stream spans many utterances (the decoder
|
|
// resets itself after each EOU). Closing the send side finalizes: the
|
|
// backend flushes the decoder tail and emits a terminal message carrying
|
|
// final_result. A second Config mid-stream resets the decode session.
|
|
rpc AudioTranscriptionLive(stream TranscriptLiveRequest) returns (stream TranscriptLiveResponse) {}
|
|
rpc TTS(TTSRequest) returns (Result) {}
|
|
rpc TTSStream(TTSRequest) returns (stream Reply) {}
|
|
rpc SoundGeneration(SoundGenerationRequest) returns (Result) {}
|
|
rpc TokenizeString(PredictOptions) returns (TokenizationResponse) {}
|
|
rpc Detokenize(DetokenizeRequest) returns (DetokenizeResponse) {}
|
|
rpc Status(HealthMessage) returns (StatusResponse) {}
|
|
rpc Detect(DetectOptions) returns (DetectResponse) {}
|
|
// SoundDetection runs an audio-tagging / sound-event-classification model
|
|
// (e.g. CED over the AudioSet ontology) on a clip and returns scored labels.
|
|
rpc SoundDetection(SoundDetectionRequest) returns (SoundDetectionResponse) {}
|
|
rpc Depth(DepthRequest) returns (DepthResponse) {}
|
|
rpc FaceVerify(FaceVerifyRequest) returns (FaceVerifyResponse) {}
|
|
rpc FaceAnalyze(FaceAnalyzeRequest) returns (FaceAnalyzeResponse) {}
|
|
rpc VoiceVerify(VoiceVerifyRequest) returns (VoiceVerifyResponse) {}
|
|
rpc VoiceAnalyze(VoiceAnalyzeRequest) returns (VoiceAnalyzeResponse) {}
|
|
rpc VoiceEmbed(VoiceEmbedRequest) returns (VoiceEmbedResponse) {}
|
|
|
|
rpc StoresSet(StoresSetOptions) returns (Result) {}
|
|
rpc StoresDelete(StoresDeleteOptions) returns (Result) {}
|
|
rpc StoresGet(StoresGetOptions) returns (StoresGetResult) {}
|
|
rpc StoresFind(StoresFindOptions) returns (StoresFindResult) {}
|
|
|
|
rpc Rerank(RerankRequest) returns (RerankResult) {}
|
|
|
|
// TokenClassify runs a token-classification (NER) model on the
|
|
// supplied text and returns each detected entity span. Used by the
|
|
// PII redactor's optional NER tier — the regex tier still handles
|
|
// formatted hits cheaply, while this catches names, locations, and
|
|
// other unformatted PII that regex misses.
|
|
rpc TokenClassify(TokenClassifyRequest) returns (TokenClassifyResponse) {}
|
|
|
|
// Score evaluates the model's joint log-probability of each
|
|
// supplied candidate continuation given a shared prompt. The
|
|
// prompt's KV cache is computed once and reused across candidates.
|
|
// Used for routing-policy multi-label classification, reranking,
|
|
// calibrated confidence, and reward-model scoring — any task where
|
|
// the consumer wants the model's confidence in a pre-specified
|
|
// continuation rather than a generated one.
|
|
rpc Score(ScoreRequest) returns (ScoreResponse) {}
|
|
|
|
rpc GetMetrics(MetricsRequest) returns (MetricsResponse);
|
|
|
|
rpc VAD(VADRequest) returns (VADResponse) {}
|
|
|
|
rpc Diarize(DiarizeRequest) returns (DiarizeResponse) {}
|
|
|
|
rpc AudioEncode(AudioEncodeRequest) returns (AudioEncodeResult) {}
|
|
rpc AudioDecode(AudioDecodeRequest) returns (AudioDecodeResult) {}
|
|
|
|
rpc AudioTransform(AudioTransformRequest) returns (AudioTransformResult) {}
|
|
rpc AudioTransformStream(stream AudioTransformFrameRequest) returns (stream AudioTransformFrameResponse) {}
|
|
// AudioToAudioStream is the bidirectional any-to-any S2S RPC. Backends
|
|
// that load a speech-to-speech model consume input audio frames and emit
|
|
// interleaved audio + transcript + tool-call deltas as typed events.
|
|
// Backends without S2S support return UNIMPLEMENTED.
|
|
rpc AudioToAudioStream(stream AudioToAudioRequest) returns (stream AudioToAudioResponse) {}
|
|
|
|
rpc ModelMetadata(ModelOptions) returns (ModelMetadataResponse) {}
|
|
|
|
// Fine-tuning RPCs
|
|
rpc StartFineTune(FineTuneRequest) returns (FineTuneJobResult) {}
|
|
rpc FineTuneProgress(FineTuneProgressRequest) returns (stream FineTuneProgressUpdate) {}
|
|
rpc StopFineTune(FineTuneStopRequest) returns (Result) {}
|
|
rpc ListCheckpoints(ListCheckpointsRequest) returns (ListCheckpointsResponse) {}
|
|
rpc ExportModel(ExportModelRequest) returns (Result) {}
|
|
|
|
// Quantization RPCs
|
|
rpc StartQuantization(QuantizationRequest) returns (QuantizationJobResult) {}
|
|
rpc QuantizationProgress(QuantizationProgressRequest) returns (stream QuantizationProgressUpdate) {}
|
|
rpc StopQuantization(QuantizationStopRequest) returns (Result) {}
|
|
|
|
// Forward proxies a raw HTTP request to an upstream provider. The
|
|
// cloud-proxy backend implements this for passthrough-mode model
|
|
// configs: the client wire format is preserved end-to-end (no
|
|
// translation through internal proto), which means new provider
|
|
// fields work the day they ship. Translation-mode proxies use the
|
|
// standard Predict/PredictStream RPCs instead. Backends that don't
|
|
// support this return UNIMPLEMENTED.
|
|
//
|
|
// The request is bidirectionally streamed so large bodies can flow
|
|
// without buffering. In practice the first ForwardRequest carries
|
|
// path, method, headers, and the initial body chunk; subsequent
|
|
// messages append body chunks. The first ForwardReply carries the
|
|
// upstream status and response headers; subsequent messages stream
|
|
// body chunks (SSE frames or chunked transfer). Cancellation of the
|
|
// gRPC context closes the upstream connection.
|
|
rpc Forward(stream ForwardRequest) returns (stream ForwardReply) {}
|
|
|
|
}
|
|
|
|
// Define the empty request
|
|
message MetricsRequest {}
|
|
|
|
message MetricsResponse {
|
|
int32 slot_id = 1;
|
|
string prompt_json_for_slot = 2; // Stores the prompt as a JSON string.
|
|
float tokens_per_second = 3;
|
|
int32 tokens_generated = 4;
|
|
int32 prompt_tokens_processed = 5;
|
|
}
|
|
|
|
// TokenClassifyRequest carries the text to classify plus an optional
|
|
// score threshold. The transformers backend interprets threshold as
|
|
// the minimum confidence to include in the response; 0 = include all.
|
|
message TokenClassifyRequest {
|
|
string text = 1;
|
|
float threshold = 2;
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 3;
|
|
// Labels overrides the backend's configured entity labels for this
|
|
// request. Empty means "use the model's configured labels" (the
|
|
// default for PII detection, where labels are fixed at load time).
|
|
// Non-empty enables zero-shot per-request label selection (kev /
|
|
// SystemOne: each question type supplies its own labels).
|
|
repeated string labels = 4;
|
|
}
|
|
|
|
// TokenClassifyEntity is one detected entity span. Byte offsets are
|
|
// into the original UTF-8 text — start..end is a half-open range that
|
|
// addresses the substring corresponding to entity_group.
|
|
//
|
|
// entity_group follows HuggingFace's aggregated-tag convention (e.g.
|
|
// "PER", "LOC", "ORG", or a PII-specific label like "EMAIL" /
|
|
// "SSN" depending on the model). The redactor's per-pattern action
|
|
// map keys off this string.
|
|
message TokenClassifyEntity {
|
|
string entity_group = 1;
|
|
int32 start = 2;
|
|
int32 end = 3;
|
|
float score = 4;
|
|
string text = 5;
|
|
}
|
|
|
|
message TokenClassifyResponse {
|
|
repeated TokenClassifyEntity entities = 1;
|
|
}
|
|
|
|
// ScoreRequest carries one shared prompt and one or more continuations
|
|
// to score against it. The backend tokenises the prompt once and reuses
|
|
// the resulting KV cache across all candidates in this request.
|
|
message ScoreRequest {
|
|
string prompt = 1;
|
|
repeated string candidates = 2;
|
|
// Return per-token logprobs for each candidate when true. Default
|
|
// false to keep the wire response small; the joint log_prob field
|
|
// covers the common ranking case.
|
|
bool include_token_logprobs = 3;
|
|
// When true, the response also populates length_normalized_log_prob
|
|
// (joint log-prob divided by candidate token count). Useful when
|
|
// candidates differ in length and the consumer wants a per-token
|
|
// measure comparable across them (PMI-style scoring).
|
|
bool length_normalize = 4;
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 5;
|
|
// Byte length of the prompt prefix that stays identical across
|
|
// repeated scoring calls (e.g. a classifier's option-list system
|
|
// prompt — everything before the per-turn probe text). Backends that
|
|
// snapshot state (hybrid/recurrent models cannot rewind otherwise)
|
|
// use it to place a reuse point exactly at the boundary, so the next
|
|
// call re-processes only the tokens after it. 0 means unknown.
|
|
int32 stable_prefix_len = 6;
|
|
// question_type signals a decision-pipeline request (kev/laya)
|
|
// rather than plain candidate scoring. When set to "systemone", the
|
|
// backend treats prompt as the raw /v1/systemone request JSON and
|
|
// returns the full response in ScoreResponse.response_json. Empty
|
|
// means plain candidate scoring.
|
|
string question_type = 7;
|
|
}
|
|
|
|
// CandidateScore is one row in the ScoreResponse, matching by index
|
|
// the candidate in ScoreRequest.candidates.
|
|
message CandidateScore {
|
|
// Sum of log P(token_i | prompt, candidate_token_<i) across the
|
|
// candidate's tokens. The primary ranking signal.
|
|
double log_prob = 1;
|
|
// log_prob / num_tokens — populated when length_normalize=true on
|
|
// the request.
|
|
double length_normalized_log_prob = 2;
|
|
// Per-token detail — populated when include_token_logprobs=true.
|
|
repeated TokenLogProb tokens = 3;
|
|
// Number of tokens the backend tokenised this candidate into, after
|
|
// any backend-specific normalisation (e.g. leading-space handling).
|
|
int32 num_tokens = 4;
|
|
}
|
|
|
|
message TokenLogProb {
|
|
string token = 1;
|
|
double log_prob = 2;
|
|
}
|
|
|
|
message ScoreResponse {
|
|
repeated CandidateScore candidates = 1;
|
|
// response_json carries the full decision-pipeline JSON response when
|
|
// question_type is set on the request. Empty for plain scoring.
|
|
string response_json = 2;
|
|
}
|
|
|
|
message RerankRequest {
|
|
string query = 1;
|
|
repeated string documents = 2;
|
|
int32 top_n = 3;
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 4;
|
|
}
|
|
|
|
message RerankResult {
|
|
Usage usage = 1;
|
|
repeated DocumentResult results = 2;
|
|
}
|
|
|
|
message Usage {
|
|
int32 total_tokens = 1;
|
|
int32 prompt_tokens = 2;
|
|
}
|
|
|
|
message DocumentResult {
|
|
int32 index = 1;
|
|
string text = 2;
|
|
float relevance_score = 3;
|
|
}
|
|
|
|
message StoresKey {
|
|
repeated float Floats = 1;
|
|
}
|
|
|
|
message StoresValue {
|
|
bytes Bytes = 1;
|
|
}
|
|
|
|
message StoresSetOptions {
|
|
repeated StoresKey Keys = 1;
|
|
repeated StoresValue Values = 2;
|
|
}
|
|
|
|
message StoresDeleteOptions {
|
|
repeated StoresKey Keys = 1;
|
|
}
|
|
|
|
message StoresGetOptions {
|
|
repeated StoresKey Keys = 1;
|
|
}
|
|
|
|
message StoresGetResult {
|
|
repeated StoresKey Keys = 1;
|
|
repeated StoresValue Values = 2;
|
|
}
|
|
|
|
message StoresFindOptions {
|
|
StoresKey Key = 1;
|
|
int32 TopK = 2;
|
|
}
|
|
|
|
message StoresFindResult {
|
|
repeated StoresKey Keys = 1;
|
|
repeated StoresValue Values = 2;
|
|
repeated float Similarities = 3;
|
|
}
|
|
|
|
message HealthMessage {}
|
|
|
|
// The request message containing the user's name.
|
|
message PredictOptions {
|
|
string Prompt = 1;
|
|
int32 Seed = 2;
|
|
int32 Threads = 3;
|
|
int32 Tokens = 4;
|
|
int32 TopK = 5;
|
|
int32 Repeat = 6;
|
|
int32 Batch = 7;
|
|
int32 NKeep = 8;
|
|
float Temperature = 9;
|
|
float Penalty = 10;
|
|
bool F16KV = 11;
|
|
bool DebugMode = 12;
|
|
repeated string StopPrompts = 13;
|
|
bool IgnoreEOS = 14;
|
|
float TailFreeSamplingZ = 15;
|
|
float TypicalP = 16;
|
|
float FrequencyPenalty = 17;
|
|
float PresencePenalty = 18;
|
|
int32 Mirostat = 19;
|
|
float MirostatETA = 20;
|
|
float MirostatTAU = 21;
|
|
bool PenalizeNL = 22;
|
|
string LogitBias = 23;
|
|
bool MLock = 25;
|
|
bool MMap = 26;
|
|
bool PromptCacheAll = 27;
|
|
bool PromptCacheRO = 28;
|
|
string Grammar = 29;
|
|
string MainGPU = 30;
|
|
string TensorSplit = 31;
|
|
float TopP = 32;
|
|
string PromptCachePath = 33;
|
|
bool Debug = 34;
|
|
repeated int32 EmbeddingTokens = 35;
|
|
string Embeddings = 36;
|
|
float RopeFreqBase = 37;
|
|
float RopeFreqScale = 38;
|
|
float NegativePromptScale = 39;
|
|
string NegativePrompt = 40;
|
|
int32 NDraft = 41;
|
|
repeated string Images = 42;
|
|
bool UseTokenizerTemplate = 43;
|
|
repeated Message Messages = 44;
|
|
repeated string Videos = 45;
|
|
repeated string Audios = 46;
|
|
string CorrelationId = 47;
|
|
string Tools = 48; // JSON array of available tools/functions for tool calling
|
|
string ToolChoice = 49; // JSON string or object specifying tool choice behavior
|
|
int32 Logprobs = 50; // Number of top logprobs to return (maps to OpenAI logprobs parameter)
|
|
int32 TopLogprobs = 51; // Number of top logprobs to return per token (maps to OpenAI top_logprobs parameter)
|
|
map<string, string> Metadata = 52; // Generic per-request metadata (e.g., enable_thinking)
|
|
float MinP = 53; // Minimum probability sampling threshold (0.0 = disabled)
|
|
|
|
// ModelIdentity names the model this request is for, so a backend can reject
|
|
// a request that reached it by mistake instead of answering from whatever
|
|
// model it happens to hold. In distributed mode a worker can recycle a
|
|
// stopped backend's gRPC port for a different model's backend, and a
|
|
// liveness-only health probe cannot tell that apart from a valid cached
|
|
// route (#10952).
|
|
//
|
|
// The value is the controller's ModelConfig.Model, the SAME expression that
|
|
// produces ModelOptions.Model at LoadModel time, so the two are equal by
|
|
// construction rather than by convention.
|
|
//
|
|
// Empty means "no identity supplied": backends MUST skip the check. That
|
|
// keeps an old controller talking to a new backend working, and covers
|
|
// callers that legitimately synthesize a PredictOptions internally.
|
|
//
|
|
// Do NOT reuse TTSRequest.model or SoundGenerationRequest.model for this
|
|
// purpose. FileStagingClient already rewrites those to worker-local absolute
|
|
// paths (core/services/nodes/file_staging_client.go), so in distributed mode
|
|
// they already differ from the load-time value and comparing them would
|
|
// reject valid requests. Extending identity to those RPCs needs a separate
|
|
// field carrying the untranslated value - which is exactly what
|
|
// TTSRequest.ModelIdentity and SoundGenerationRequest.ModelIdentity are.
|
|
//
|
|
// Every other request message that reaches a backend through the distributed
|
|
// router now carries the same ModelIdentity field, populated from the same
|
|
// ModelConfig.Model. FileStagingClient rewrites Src/Dst/Voice/Model/
|
|
// StartImage/EndImage/Audio and never ModelIdentity, so what the backend
|
|
// compares is always what the controller sent.
|
|
string ModelIdentity = 54;
|
|
|
|
// 24 was never assigned; reserve it so it is not silently reused.
|
|
reserved 24;
|
|
}
|
|
|
|
// ToolCallDelta represents an incremental tool call update from the C++ parser.
|
|
// Used for both streaming (partial diffs) and non-streaming (final tool calls).
|
|
message ToolCallDelta {
|
|
int32 index = 1; // tool call index (0-based)
|
|
string id = 2; // tool call ID (e.g., "call_abc123")
|
|
string name = 3; // function name (set on first appearance)
|
|
string arguments = 4; // arguments chunk (incremental in streaming, full in non-streaming)
|
|
}
|
|
|
|
// ChatDelta represents incremental content/reasoning/tool_call updates parsed by the C++ backend.
|
|
message ChatDelta {
|
|
string content = 1; // content text delta
|
|
string reasoning_content = 2; // reasoning/thinking text delta
|
|
repeated ToolCallDelta tool_calls = 3; // tool call deltas
|
|
}
|
|
|
|
// The response message containing the result
|
|
message Reply {
|
|
bytes message = 1;
|
|
int32 tokens = 2;
|
|
int32 prompt_tokens = 3;
|
|
double timing_prompt_processing = 4;
|
|
double timing_token_generation = 5;
|
|
bytes audio = 6;
|
|
bytes logprobs = 7; // JSON-encoded logprobs data matching OpenAI format
|
|
repeated ChatDelta chat_deltas = 8; // Parsed chat deltas from C++ autoparser (streaming + non-streaming)
|
|
}
|
|
|
|
message GrammarTrigger {
|
|
string word = 1;
|
|
}
|
|
|
|
message ModelOptions {
|
|
string Model = 1;
|
|
int32 ContextSize = 2;
|
|
int32 Seed = 3;
|
|
int32 NBatch = 4;
|
|
bool F16Memory = 5;
|
|
bool MLock = 6;
|
|
bool MMap = 7;
|
|
bool VocabOnly = 8;
|
|
bool LowVRAM = 9;
|
|
bool Embeddings = 10;
|
|
bool NUMA = 11;
|
|
int32 NGPULayers = 12;
|
|
string MainGPU = 13;
|
|
string TensorSplit = 14;
|
|
int32 Threads = 15;
|
|
float RopeFreqBase = 17;
|
|
float RopeFreqScale = 18;
|
|
float RMSNormEps = 19;
|
|
int32 NGQA = 20;
|
|
string ModelFile = 21;
|
|
|
|
|
|
|
|
// Diffusers
|
|
string PipelineType = 26;
|
|
string SchedulerType = 27;
|
|
bool CUDA = 28;
|
|
float CFGScale = 29;
|
|
bool IMG2IMG = 30;
|
|
string CLIPModel = 31;
|
|
string CLIPSubfolder = 32;
|
|
int32 CLIPSkip = 33;
|
|
string ControlNet = 48;
|
|
|
|
string Tokenizer = 34;
|
|
|
|
// LLM (llama.cpp)
|
|
string LoraBase = 35;
|
|
string LoraAdapter = 36;
|
|
float LoraScale = 42;
|
|
|
|
bool NoMulMatQ = 37;
|
|
string DraftModel = 39;
|
|
|
|
string AudioPath = 38;
|
|
|
|
// vllm
|
|
string Quantization = 40;
|
|
float GPUMemoryUtilization = 50;
|
|
bool TrustRemoteCode = 51;
|
|
bool EnforceEager = 52;
|
|
int32 SwapSpace = 53;
|
|
int32 MaxModelLen = 54;
|
|
int32 TensorParallelSize = 55;
|
|
string LoadFormat = 58;
|
|
bool DisableLogStatus = 66;
|
|
string DType = 67;
|
|
int32 LimitImagePerPrompt = 68;
|
|
int32 LimitVideoPerPrompt = 69;
|
|
int32 LimitAudioPerPrompt = 70;
|
|
|
|
string MMProj = 41;
|
|
|
|
string RopeScaling = 43;
|
|
float YarnExtFactor = 44;
|
|
float YarnAttnFactor = 45;
|
|
float YarnBetaFast = 46;
|
|
float YarnBetaSlow = 47;
|
|
|
|
string Type = 49;
|
|
|
|
string FlashAttention = 56;
|
|
bool NoKVOffload = 57;
|
|
|
|
string ModelPath = 59;
|
|
|
|
repeated string LoraAdapters = 60;
|
|
repeated float LoraScales = 61;
|
|
|
|
repeated string Options = 62;
|
|
|
|
string CacheTypeKey = 63;
|
|
string CacheTypeValue = 64;
|
|
|
|
repeated GrammarTrigger GrammarTriggers = 65;
|
|
|
|
bool Reranking = 71;
|
|
|
|
repeated string Overrides = 72;
|
|
|
|
// EngineArgs carries a JSON-encoded map of backend-native engine arguments
|
|
// applied verbatim to the backend's engine constructor (e.g. vLLM AsyncEngineArgs).
|
|
// Unknown keys produce an error at LoadModel time.
|
|
string EngineArgs = 73;
|
|
string OriginalConfigFile = 77;
|
|
|
|
// EnvVars carries environment variables to be passed to the backend process.
|
|
map<string, string> EnvVars = 76;
|
|
|
|
// Proxy carries the cloud-proxy backend's per-model configuration.
|
|
// Empty for non-proxy backends.
|
|
ProxyOptions Proxy = 74;
|
|
|
|
// EnableScore reserves backend resources for the Score RPC. It is derived
|
|
// from the model's explicit `known_usecases: [score]` declaration so models
|
|
// that never score retain their ordinary serving footprint.
|
|
bool EnableScore = 75;
|
|
}
|
|
|
|
// ProxyOptions configures the cloud-proxy backend. UpstreamURL and
|
|
// Mode are always meaningful; Provider only matters in translate mode.
|
|
// The two api_key_* fields are mutually exclusive and resolved by the
|
|
// backend at LoadModel — core forwards the references rather than the
|
|
// plaintext key.
|
|
message ProxyOptions {
|
|
string upstream_url = 1;
|
|
string mode = 2;
|
|
string provider = 3;
|
|
string api_key_env = 4;
|
|
string api_key_file = 5;
|
|
string upstream_model = 6;
|
|
int32 request_timeout_seconds = 7;
|
|
// cache_prompt enables automatic Anthropic prompt-cache breakpoints
|
|
// (cache_control: ephemeral) on the stable prefix — system, tools, and
|
|
// the last message block — when translating to the Anthropic provider.
|
|
// Cuts input cost on repeated/agentic calls (cache read = 0.1x). Only
|
|
// meaningful for mode=translate + provider=anthropic; ignored otherwise.
|
|
bool cache_prompt = 8;
|
|
}
|
|
|
|
message Result {
|
|
string message = 1;
|
|
bool success = 2;
|
|
// JSON object containing arbitrary backend response metadata. Application
|
|
// conventions (such as metadata.usage) are independent of this schema.
|
|
bytes metadata = 4;
|
|
}
|
|
|
|
// EmbeddingLayout describes whether embeddings contains one final vector or
|
|
// a matrix of per-token vectors. Go-side pooling must never infer this from
|
|
// tokens/dim alone: a one-token raw matrix and a final vector have the same
|
|
// shape.
|
|
enum EmbeddingLayout {
|
|
EMBEDDING_LAYOUT_UNSPECIFIED = 0;
|
|
EMBEDDING_LAYOUT_FINAL = 1;
|
|
EMBEDDING_LAYOUT_PER_TOKEN = 2;
|
|
}
|
|
|
|
message EmbeddingResult {
|
|
repeated float embeddings = 1;
|
|
// Shape of the payload above: dim is the embedding width, tokens is the
|
|
// number of vectors packed into `embeddings` (1 when the backend pooled
|
|
// server-side, N with pooling:none; total across prompts if a request
|
|
// carried several). tokens=0/dim=0 means the backend predates shape
|
|
// reporting. prompt_tokens is the number of prompt tokens evaluated, for
|
|
// usage accounting.
|
|
int32 tokens = 2;
|
|
int32 dim = 3;
|
|
int32 prompt_tokens = 4;
|
|
EmbeddingLayout layout = 5;
|
|
}
|
|
|
|
message TranscriptRequest {
|
|
string dst = 2;
|
|
string language = 3;
|
|
uint32 threads = 4;
|
|
bool translate = 5;
|
|
bool diarize = 6;
|
|
string prompt = 7;
|
|
float temperature = 8;
|
|
repeated string timestamp_granularities = 9;
|
|
bool stream = 10;
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 11;
|
|
}
|
|
|
|
message TranscriptResult {
|
|
repeated TranscriptSegment segments = 1;
|
|
string text = 2;
|
|
string language = 3;
|
|
float duration = 4;
|
|
// True when the decode ended on the model's end-of-utterance special token
|
|
// (<EOU>/<EOB>, emitted by cache-aware streaming models such as
|
|
// parakeet_realtime_eou_120m-v1). The marker itself is stripped from text.
|
|
bool eou = 5;
|
|
}
|
|
|
|
message TranscriptStreamResponse {
|
|
string delta = 1;
|
|
TranscriptResult final_result = 2;
|
|
}
|
|
|
|
// === AudioTranscriptionLive messages =====================================
|
|
|
|
message TranscriptLiveRequest {
|
|
oneof payload {
|
|
TranscriptLiveConfig config = 1;
|
|
TranscriptLiveAudio audio = 2;
|
|
}
|
|
}
|
|
|
|
message TranscriptLiveConfig {
|
|
string language = 1; // "" => model default
|
|
int32 sample_rate = 2; // 0 => 16000; backends may reject others
|
|
map<string, string> params = 3; // backend-specific tuning
|
|
}
|
|
|
|
message TranscriptLiveAudio {
|
|
repeated float pcm = 1; // mono PCM in [-1,1] at config.sample_rate
|
|
}
|
|
|
|
message TranscriptLiveResponse {
|
|
bool ready = 1; // open ack: sent once, before any delta
|
|
string delta = 2; // newly-finalized text since previous response
|
|
bool eou = 3; // <EOU> fired during this feed (the user yielded the turn)
|
|
repeated TranscriptWord words = 4; // words finalized by this feed (stream-relative ns)
|
|
TranscriptResult final_result = 5; // terminal message only, after the send side closes
|
|
bool eob = 6; // <EOB> fired: a backchannel ("uh-huh") ended — NOT a turn boundary
|
|
repeated LiveSpeakerSegment speakers = 7; // closed speaker segments from a companion diarization/scene stream
|
|
repeated LiveSoundEvent sounds = 8; // closed sound events from a companion sound/scene stream
|
|
}
|
|
|
|
message LiveSpeakerSegment {
|
|
string speaker = 1; // decimal speaker index
|
|
int64 start = 2; // stream-relative nanoseconds
|
|
int64 end = 3;
|
|
}
|
|
|
|
message LiveSoundEvent {
|
|
string label = 1;
|
|
int32 index = 2;
|
|
float peak = 3;
|
|
int64 start = 4; // stream-relative nanoseconds
|
|
int64 end = 5;
|
|
}
|
|
|
|
message TranscriptWord {
|
|
int64 start = 1;
|
|
int64 end = 2;
|
|
string text = 3;
|
|
string speaker = 4; // backend speaker label when diarizing; empty otherwise
|
|
}
|
|
|
|
message TranscriptSegment {
|
|
int32 id = 1;
|
|
int64 start = 2;
|
|
int64 end = 3;
|
|
string text = 4;
|
|
repeated int32 tokens = 5;
|
|
string speaker = 6;
|
|
repeated TranscriptWord words = 7;
|
|
}
|
|
|
|
message GenerateImageRequest {
|
|
int32 height = 1;
|
|
int32 width = 2;
|
|
int32 step = 4;
|
|
int32 seed = 5;
|
|
string positive_prompt = 6;
|
|
string negative_prompt = 7;
|
|
string dst = 8;
|
|
string src = 9;
|
|
|
|
// Diffusers
|
|
string EnableParameters = 10;
|
|
int32 CLIPSkip = 11;
|
|
|
|
// Reference images for models that support them (e.g., Flux Kontext)
|
|
repeated string ref_images = 12;
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 13;
|
|
}
|
|
|
|
message UpscaleImageRequest {
|
|
string src = 1; // input image path
|
|
string dst = 2; // output image path
|
|
int32 scale = 3; // upscale factor (e.g. 2 or 4)
|
|
}
|
|
|
|
message GenerateVideoRequest {
|
|
string prompt = 1;
|
|
string negative_prompt = 2; // Negative prompt for video generation
|
|
string start_image = 3; // Path or base64 encoded image for the start frame
|
|
string end_image = 4; // Path or base64 encoded image for the end frame
|
|
int32 width = 5;
|
|
int32 height = 6;
|
|
int32 num_frames = 7; // Number of frames to generate
|
|
int32 fps = 8; // Frames per second
|
|
int32 seed = 9;
|
|
float cfg_scale = 10; // Classifier-free guidance scale
|
|
int32 step = 11; // Number of inference steps
|
|
string dst = 12; // Output path for the generated video
|
|
string audio = 13; // Path to staged audio for audio-conditioned video
|
|
// Backend-specific per-request generation parameters. Values are strings
|
|
// and are validated/coerced by the selected backend.
|
|
map<string, string> params = 14;
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 15;
|
|
}
|
|
|
|
// A named conditioning input. Media references are staged local paths by the
|
|
// time they reach the backend; text is passed verbatim.
|
|
message AnimationInput {
|
|
string type = 1; // text, image, video, or mesh
|
|
string data = 2;
|
|
}
|
|
|
|
message Animate3DRequest {
|
|
map<string, AnimationInput> inputs = 1;
|
|
string dst = 2;
|
|
map<string, string> params = 3;
|
|
string model_identity = 4;
|
|
}
|
|
|
|
message Generate3DRequest {
|
|
string src = 1; // Path to the staged conditioning image (3D generation is image-conditioned)
|
|
string dst = 2; // Output path for the generated binary glTF (.glb) asset
|
|
int32 seed = 3; // <=0 lets the backend pick a random seed
|
|
int32 step = 4; // Flow sampling steps; <=0 uses the backend default
|
|
float cfg_scale = 5; // Classifier-free guidance scale; <=0 uses the backend default
|
|
int32 texture_steps = 6; // Texture flow sampling steps; <=0 uses the backend default
|
|
string quality = 7; // Mesh pipeline: ""|"auto"|"coarse"|"512"|"1024"
|
|
string background = 8; // Conditioning-image background handling: ""|"auto"|"keep"|"black"|"white"
|
|
// Backend-specific per-request generation parameters. Values are strings
|
|
// and are validated/coerced by the selected backend.
|
|
map<string, string> params = 9;
|
|
}
|
|
|
|
message TTSRequest {
|
|
string text = 1;
|
|
string model = 2;
|
|
string dst = 3;
|
|
string voice = 4;
|
|
optional string language = 5;
|
|
// instructions is a free-form, per-request style/voice description (maps to
|
|
// the OpenAI `instructions` field). Backends that support expressive synthesis
|
|
// (e.g. Qwen3-TTS CustomVoice/VoiceDesign) prefer this over the static YAML
|
|
// option when set; backends that don't simply ignore it.
|
|
optional string instructions = 6;
|
|
// params carries optional, backend-specific per-request generation parameters
|
|
// (e.g. Chatterbox exaggeration/cfg_weight/temperature). Values are strings and
|
|
// coerced by the backend; unset leaves the backend's configured defaults.
|
|
map<string, string> params = 7;
|
|
// ModelIdentity is a SEPARATE field from `model` above and carries the
|
|
// UNTRANSLATED controller-side ModelConfig.Model, so a backend can reject a
|
|
// request that reached it through a stale distributed route (#10952).
|
|
//
|
|
// `model` cannot be reused for this: FileStagingClient.TTS/.TTSStream and the
|
|
// SoundGeneration path rewrite it into a worker-local absolute path
|
|
// (core/services/nodes/file_staging_client.go), while the load-time value is
|
|
// untranslated. In distributed mode - exactly the configuration this guards -
|
|
// the two already differ, so comparing them would reject valid requests.
|
|
//
|
|
// Empty means "no identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 8;
|
|
}
|
|
|
|
message VADRequest {
|
|
repeated float audio = 1;
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 2;
|
|
}
|
|
|
|
message VADSegment {
|
|
float start = 1;
|
|
float end = 2;
|
|
}
|
|
|
|
message VADResponse {
|
|
repeated VADSegment segments = 1;
|
|
}
|
|
|
|
// --- Speaker diarization messages ---
|
|
//
|
|
// Pure speaker diarization: "who spoke when". Returns time-stamped segments
|
|
// labelled with cluster IDs (the same string for the same speaker across
|
|
// segments). Some backends (e.g. vibevoice.cpp) produce diarization as a
|
|
// by-product of ASR and may also fill in `text` per segment; backends with a
|
|
// dedicated diarization pipeline (e.g. sherpa-onnx pyannote) leave `text`
|
|
// empty and emit only the segmentation.
|
|
|
|
message DiarizeRequest {
|
|
string dst = 1; // path to audio file (HTTP layer materialises uploads to a temp file)
|
|
uint32 threads = 2;
|
|
string language = 3; // optional; only meaningful for transcription-bundling backends
|
|
int32 num_speakers = 4; // exact speaker count if known (>0 forces); 0 = auto
|
|
int32 min_speakers = 5; // hint when auto-detecting; 0 = unset
|
|
int32 max_speakers = 6; // hint when auto-detecting; 0 = unset
|
|
float clustering_threshold = 7; // distance threshold when num_speakers unknown; 0 = backend default
|
|
float min_duration_on = 8; // discard segments shorter than this (seconds); 0 = backend default
|
|
float min_duration_off = 9; // merge gaps shorter than this (seconds); 0 = backend default
|
|
bool include_text = 10; // when the backend can emit per-segment transcript for free, ask it to populate `text`
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 11;
|
|
}
|
|
|
|
message DiarizeSegment {
|
|
int32 id = 1;
|
|
float start = 2; // seconds
|
|
float end = 3; // seconds
|
|
string speaker = 4; // backend-emitted speaker label (e.g. "0", "SPEAKER_00")
|
|
string text = 5; // optional per-segment transcript (empty unless include_text and supported)
|
|
}
|
|
|
|
message DiarizeResponse {
|
|
repeated DiarizeSegment segments = 1;
|
|
int32 num_speakers = 2; // count of distinct speaker labels in `segments`
|
|
float duration = 3; // total audio duration in seconds (0 if unknown)
|
|
string language = 4; // optional, when the backend bundles transcription
|
|
}
|
|
|
|
message SoundGenerationRequest {
|
|
string text = 1;
|
|
string model = 2;
|
|
string dst = 3;
|
|
optional float duration = 4;
|
|
optional float temperature = 5;
|
|
optional bool sample = 6;
|
|
optional string src = 7;
|
|
optional int32 src_divisor = 8;
|
|
optional bool think = 9;
|
|
optional string caption = 10;
|
|
optional string lyrics = 11;
|
|
optional int32 bpm = 12;
|
|
optional string keyscale = 13;
|
|
optional string language = 14;
|
|
optional string timesignature = 15;
|
|
optional bool instrumental = 17;
|
|
// ModelIdentity is a SEPARATE field from `model` above and carries the
|
|
// UNTRANSLATED controller-side ModelConfig.Model, so a backend can reject a
|
|
// request that reached it through a stale distributed route (#10952).
|
|
//
|
|
// `model` cannot be reused for this: FileStagingClient.TTS/.TTSStream and the
|
|
// SoundGeneration path rewrite it into a worker-local absolute path
|
|
// (core/services/nodes/file_staging_client.go), while the load-time value is
|
|
// untranslated. In distributed mode - exactly the configuration this guards -
|
|
// the two already differ, so comparing them would reject valid requests.
|
|
//
|
|
// Empty means "no identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 18;
|
|
}
|
|
|
|
message TokenizationResponse {
|
|
int32 length = 1;
|
|
repeated int32 tokens = 2;
|
|
}
|
|
|
|
message DetokenizeRequest {
|
|
repeated int32 tokens = 1;
|
|
}
|
|
|
|
message DetokenizeResponse {
|
|
string content = 1;
|
|
}
|
|
|
|
message MemoryUsageData {
|
|
uint64 total = 1;
|
|
map<string, uint64> breakdown = 2;
|
|
}
|
|
|
|
message StatusResponse {
|
|
enum State {
|
|
UNINITIALIZED = 0;
|
|
BUSY = 1;
|
|
READY = 2;
|
|
ERROR = -1;
|
|
}
|
|
State state = 1;
|
|
MemoryUsageData memory = 2;
|
|
}
|
|
|
|
message Message {
|
|
string role = 1;
|
|
string content = 2;
|
|
// Optional fields for OpenAI-compatible message format
|
|
string name = 3; // Tool name (for tool messages)
|
|
string tool_call_id = 4; // Tool call ID (for tool messages)
|
|
string reasoning_content = 5; // Reasoning content (for thinking models)
|
|
string tool_calls = 6; // Tool calls as JSON string (for assistant messages with tool calls)
|
|
}
|
|
|
|
message DetectOptions {
|
|
string src = 1;
|
|
string prompt = 2; // Text prompt (for SAM 3 PCS mode)
|
|
repeated float points = 3; // Point coordinates as [x1, y1, label1, x2, y2, label2, ...] (label: 1=pos, 0=neg)
|
|
repeated float boxes = 4; // Box coordinates as [x1, y1, x2, y2, ...]
|
|
float threshold = 5; // Detection confidence threshold
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 6;
|
|
}
|
|
|
|
message Detection {
|
|
float x = 1;
|
|
float y = 2;
|
|
float width = 3;
|
|
float height = 4;
|
|
float confidence = 5;
|
|
string class_name = 6;
|
|
bytes mask = 7; // PNG-encoded binary segmentation mask
|
|
}
|
|
|
|
message DetectResponse {
|
|
repeated Detection Detections = 1;
|
|
}
|
|
|
|
// --- Sound-event classification / audio tagging messages (CED) ---
|
|
|
|
message SoundDetectionRequest {
|
|
string src = 1; // audio file path (LocalAI writes the upload to disk)
|
|
int32 top_k = 2; // number of top tags to return (0 = all classes)
|
|
float threshold = 3; // optional: drop tags scoring below this
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 4;
|
|
}
|
|
|
|
message SoundClass {
|
|
string label = 1; // AudioSet class name, e.g. "Baby cry, infant cry"
|
|
float score = 2; // per-class probability (multi-label, independent)
|
|
int32 index = 3; // class index in the model ontology
|
|
}
|
|
|
|
message SoundDetectionResponse {
|
|
repeated SoundClass detections = 1; // score-descending
|
|
}
|
|
|
|
// --- Depth estimation messages (Depth Anything 3) ---
|
|
|
|
message DepthRequest {
|
|
string src = 1; // input image (filesystem path or base64-encoded payload)
|
|
string dst = 2; // optional output directory for exports (glb/colmap)
|
|
bool include_depth = 3; // return the per-pixel metric depth map
|
|
bool include_confidence = 4; // return the per-pixel confidence map (DualDPT)
|
|
bool include_pose = 5; // return camera extrinsics/intrinsics (DualDPT)
|
|
bool include_sky = 6; // return the per-pixel sky map (mono models)
|
|
bool include_points = 7; // back-project to a 3D point cloud (DualDPT)
|
|
float points_conf_thresh = 8; // keep points with confidence >= this threshold
|
|
repeated string exports = 9; // requested exports: "glb", "colmap"
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 10;
|
|
}
|
|
|
|
message DepthResponse {
|
|
int32 width = 1; // processed depth-map width
|
|
int32 height = 2; // processed depth-map height
|
|
repeated float depth = 3; // width*height row-major metric depth
|
|
repeated float confidence = 4; // width*height row-major confidence (DualDPT)
|
|
repeated float sky = 5; // width*height row-major sky map (mono)
|
|
repeated float extrinsics = 6; // 12 floats, 3x4 row-major (world-to-camera)
|
|
repeated float intrinsics = 7; // 9 floats, 3x3 row-major
|
|
int32 num_points = 8; // number of 3D points
|
|
repeated float points = 9; // num_points*3 xyz, world space
|
|
bytes point_colors = 10; // num_points*3 uint8 rgb
|
|
repeated string export_paths = 11; // paths written for the requested exports
|
|
bool is_metric = 12; // depth is in metric units
|
|
}
|
|
|
|
// --- Face recognition messages ---
|
|
|
|
message FacialArea {
|
|
float x = 1;
|
|
float y = 2;
|
|
float w = 3;
|
|
float h = 4;
|
|
}
|
|
|
|
message FaceVerifyRequest {
|
|
string img1 = 1; // base64-encoded image
|
|
string img2 = 2; // base64-encoded image
|
|
float threshold = 3; // cosine-distance threshold; 0 = use backend default
|
|
bool anti_spoofing = 4; // run MiniFASNet liveness on each image; failed liveness forces verified=false
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 5;
|
|
}
|
|
|
|
message FaceVerifyResponse {
|
|
bool verified = 1;
|
|
float distance = 2; // 1 - cosine_similarity
|
|
float threshold = 3;
|
|
float confidence = 4; // 0-100
|
|
string model = 5; // e.g. "buffalo_l"
|
|
FacialArea img1_area = 6;
|
|
FacialArea img2_area = 7;
|
|
float processing_time_ms = 8;
|
|
bool img1_is_real = 9; // anti-spoofing result when enabled
|
|
float img1_antispoof_score = 10;
|
|
bool img2_is_real = 11;
|
|
float img2_antispoof_score = 12;
|
|
}
|
|
|
|
message FaceAnalyzeRequest {
|
|
string img = 1; // base64-encoded image
|
|
repeated string actions = 2; // subset of ["age","gender","emotion","race"]; empty = all-supported
|
|
bool anti_spoofing = 3;
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 4;
|
|
}
|
|
|
|
message FaceAnalysis {
|
|
FacialArea region = 1;
|
|
float face_confidence = 2;
|
|
float age = 3;
|
|
string dominant_gender = 4; // "Man" | "Woman"
|
|
map<string, float> gender = 5;
|
|
string dominant_emotion = 6; // reserved; empty in MVP
|
|
map<string, float> emotion = 7;
|
|
string dominant_race = 8; // not populated
|
|
map<string, float> race = 9;
|
|
bool is_real = 10; // anti-spoofing result when enabled
|
|
float antispoof_score = 11;
|
|
}
|
|
|
|
message FaceAnalyzeResponse {
|
|
repeated FaceAnalysis faces = 1;
|
|
}
|
|
|
|
// --- Voice (speaker) recognition messages ---
|
|
//
|
|
// Analogous to the Face* messages above, but for speaker biometrics.
|
|
// Audio fields accept a filesystem path (same convention as
|
|
// TranscriptRequest.dst). The HTTP layer materialises base64 / URL /
|
|
// data-URI inputs to a temp file before calling the gRPC backend.
|
|
|
|
message VoiceVerifyRequest {
|
|
string audio1 = 1; // path to first audio clip
|
|
string audio2 = 2; // path to second audio clip
|
|
float threshold = 3; // cosine-distance threshold; 0 = use backend default
|
|
bool anti_spoofing = 4; // reserved for future AASIST bolt-on
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 5;
|
|
}
|
|
|
|
message VoiceVerifyResponse {
|
|
bool verified = 1;
|
|
float distance = 2; // 1 - cosine_similarity
|
|
float threshold = 3;
|
|
float confidence = 4; // 0-100
|
|
string model = 5; // e.g. "speechbrain/spkrec-ecapa-voxceleb"
|
|
float processing_time_ms = 6;
|
|
}
|
|
|
|
message VoiceAnalyzeRequest {
|
|
string audio = 1; // path to audio clip
|
|
repeated string actions = 2; // subset of ["age","gender","emotion"]; empty = all-supported
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 3;
|
|
}
|
|
|
|
message VoiceAnalysis {
|
|
float start = 1; // segment start time in seconds (0 if single-utterance)
|
|
float end = 2; // segment end time in seconds
|
|
float age = 3;
|
|
string dominant_gender = 4;
|
|
map<string, float> gender = 5;
|
|
string dominant_emotion = 6;
|
|
map<string, float> emotion = 7;
|
|
}
|
|
|
|
message VoiceAnalyzeResponse {
|
|
repeated VoiceAnalysis segments = 1;
|
|
}
|
|
|
|
message VoiceEmbedRequest {
|
|
string audio = 1; // path to audio clip
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 2;
|
|
}
|
|
|
|
message VoiceEmbedResponse {
|
|
repeated float embedding = 1;
|
|
string model = 2;
|
|
}
|
|
|
|
message ToolFormatMarkers {
|
|
string format_type = 1; // "json_native", "tag_with_json", "tag_with_tagged"
|
|
|
|
// Tool section markers
|
|
string section_start = 2; // e.g., "<tool_call>", "[TOOL_CALLS]"
|
|
string section_end = 3; // e.g., "</tool_call>"
|
|
string per_call_start = 4; // e.g., "<|tool_call_begin|>"
|
|
string per_call_end = 5; // e.g., "<|tool_call_end|>"
|
|
|
|
// Function name markers (TAG_WITH_JSON / TAG_WITH_TAGGED)
|
|
string func_name_prefix = 6; // e.g., "<function="
|
|
string func_name_suffix = 7; // e.g., ">"
|
|
string func_close = 8; // e.g., "</function>"
|
|
|
|
// Argument markers (TAG_WITH_TAGGED)
|
|
string arg_name_prefix = 9; // e.g., "<param="
|
|
string arg_name_suffix = 10; // e.g., ">"
|
|
string arg_value_prefix = 11;
|
|
string arg_value_suffix = 12; // e.g., "</param>"
|
|
string arg_separator = 13; // e.g., "\n"
|
|
|
|
// JSON format fields (JSON_NATIVE)
|
|
string name_field = 14; // e.g., "name"
|
|
string args_field = 15; // e.g., "arguments"
|
|
string id_field = 16; // e.g., "id"
|
|
bool fun_name_is_key = 17;
|
|
bool tools_array_wrapped = 18;
|
|
reserved 19;
|
|
|
|
// Reasoning markers
|
|
string reasoning_start = 20; // e.g., "<think>"
|
|
string reasoning_end = 21; // e.g., "</think>"
|
|
|
|
// Content markers
|
|
string content_start = 22;
|
|
string content_end = 23;
|
|
|
|
// Args wrapper markers
|
|
string args_start = 24; // e.g., "<args>"
|
|
string args_end = 25; // e.g., "</args>"
|
|
|
|
// JSON parameter ordering
|
|
string function_field = 26; // e.g., "function" (wrapper key in JSON)
|
|
repeated string parameter_order = 27;
|
|
|
|
// Generated ID field (alternative field name for generated IDs)
|
|
string gen_id_field = 28; // e.g., "call_id"
|
|
|
|
// Call ID markers (position and delimiters for tool call IDs)
|
|
string call_id_position = 29; // "none", "pre_func_name", "between_func_and_args", "post_args"
|
|
string call_id_prefix = 30; // e.g., "[CALL_ID]"
|
|
string call_id_suffix = 31; // e.g., ""
|
|
}
|
|
|
|
message AudioEncodeRequest {
|
|
bytes pcm_data = 1;
|
|
int32 sample_rate = 2;
|
|
int32 channels = 3;
|
|
map<string, string> options = 4;
|
|
}
|
|
|
|
message AudioEncodeResult {
|
|
repeated bytes frames = 1;
|
|
int32 sample_rate = 2;
|
|
int32 samples_per_frame = 3;
|
|
}
|
|
|
|
message AudioDecodeRequest {
|
|
repeated bytes frames = 1;
|
|
map<string, string> options = 2;
|
|
}
|
|
|
|
message AudioDecodeResult {
|
|
bytes pcm_data = 1;
|
|
int32 sample_rate = 2;
|
|
int32 samples_per_frame = 3;
|
|
}
|
|
|
|
// Generic audio transform: an audio-in, audio-out operation, optionally
|
|
// conditioned on a second reference signal. Concrete transforms include
|
|
// AEC + noise suppression + dereverberation (LocalVQE), voice conversion
|
|
// (reference = target speaker), pitch shifting, etc.
|
|
message AudioTransformRequest {
|
|
string audio_path = 1; // required, primary input file path
|
|
string reference_path = 2; // optional auxiliary; empty => zero-fill
|
|
string dst = 3; // required, output file path
|
|
map<string, string> params = 4; // backend-specific tuning
|
|
// ModelIdentity names the model this request is for; see
|
|
// PredictOptions.ModelIdentity for the full rationale. Empty means "no
|
|
// identity supplied" and backends MUST skip the check.
|
|
string ModelIdentity = 5;
|
|
}
|
|
|
|
// One named output of a transform that produces several from a single run.
|
|
// Source separation is the case that needs it: htdemucs yields drums, bass,
|
|
// other and vocals from one pass over the input.
|
|
message AudioTransformStem {
|
|
string name = 1; // the model's own stem id, e.g. "vocals"
|
|
string dst = 2; // path of the file written for that stem
|
|
}
|
|
|
|
message AudioTransformResult {
|
|
string dst = 1;
|
|
int32 sample_rate = 2;
|
|
int32 samples = 3;
|
|
bool reference_provided = 4;
|
|
// Every named output the run produced, in the model's own order, including
|
|
// the one copied into dst. Empty for a transform with a single output.
|
|
//
|
|
// It exists because dst carries one file while separation produces several,
|
|
// and running the model once per stem would cost four full separations of
|
|
// the same audio. The backend runs once, writes each stem beside dst, and
|
|
// names them here; without this field the other stems are on disk but no
|
|
// caller can find them, which is the same as not having produced them.
|
|
repeated AudioTransformStem stems = 5;
|
|
}
|
|
|
|
// Bidirectional streaming audio transform. The first message MUST carry a
|
|
// Config; subsequent messages carry Frames. A second Config mid-stream
|
|
// resets streaming state before the next frame.
|
|
message AudioTransformFrameRequest {
|
|
oneof payload {
|
|
AudioTransformStreamConfig config = 1;
|
|
AudioTransformFrame frame = 2;
|
|
}
|
|
}
|
|
|
|
message AudioTransformStreamConfig {
|
|
enum SampleFormat {
|
|
F32_LE = 0;
|
|
S16_LE = 1;
|
|
}
|
|
SampleFormat sample_format = 1;
|
|
int32 sample_rate = 2; // 0 => backend default
|
|
int32 frame_samples = 3; // 0 => backend default
|
|
map<string, string> params = 4;
|
|
bool reset = 5; // reset streaming state before next frame
|
|
}
|
|
|
|
message AudioTransformFrame {
|
|
bytes audio_pcm = 1; // frame_samples samples in stream's format
|
|
bytes reference_pcm = 2; // empty => zero-fill (silent reference)
|
|
}
|
|
|
|
message AudioTransformFrameResponse {
|
|
bytes pcm = 1;
|
|
int64 frame_index = 2;
|
|
}
|
|
|
|
// === AudioToAudioStream messages =========================================
|
|
//
|
|
// Bidirectional stream between the LocalAI core and an any-to-any audio
|
|
// model. The client opens the stream with a Config payload, then alternates
|
|
// Frame (input audio) and Control (turn boundaries, function-call results,
|
|
// session updates) payloads. The server streams back typed events: audio
|
|
// frames carry PCM in `pcm`; transcript / tool-call deltas carry JSON in
|
|
// `meta`; the stream ends with a `response.done` (success) or `error` event.
|
|
|
|
message AudioToAudioRequest {
|
|
oneof payload {
|
|
AudioToAudioConfig config = 1;
|
|
AudioToAudioFrame frame = 2;
|
|
AudioToAudioControl control = 3;
|
|
}
|
|
}
|
|
|
|
message AudioToAudioConfig {
|
|
// PCM format for client→server audio. 0 => backend default
|
|
// (16 kHz for the LFM2-Audio Conformer encoder).
|
|
int32 input_sample_rate = 1;
|
|
// Preferred server→client audio rate. 0 => backend default
|
|
// (24 kHz for the LFM2-Audio vocoder).
|
|
int32 output_sample_rate = 2;
|
|
// Optional system prompt override. Empty => backend chooses based on
|
|
// mode (e.g. "Respond with interleaved text and audio.").
|
|
string system_prompt = 3;
|
|
// Optional baked-voice id. Models that only ship a fixed set of
|
|
// voices (e.g. LFM2-Audio: us_male/us_female/uk_male/uk_female) match
|
|
// this against their voice table; an empty string keeps the default.
|
|
string voice = 4;
|
|
// JSON-encoded array of tool definitions in OpenAI Chat Completions
|
|
// format. Empty => no tools.
|
|
string tools = 5;
|
|
// Free-form sampling / decoding parameters (temperature, top_k,
|
|
// max_new_tokens, audio_top_k, etc).
|
|
map<string, string> params = 6;
|
|
// True => reset any session-scoped state before processing further
|
|
// frames on this stream. The first Config implicitly resets.
|
|
bool reset = 7;
|
|
}
|
|
|
|
message AudioToAudioFrame {
|
|
// Raw PCM s16le mono at config.input_sample_rate. Empty pcm + end_of_input
|
|
// is a valid "user finished speaking" marker without trailing audio.
|
|
bytes pcm = 1;
|
|
// Marks the last frame of a user turn. The backend may begin emitting
|
|
// a response immediately after seeing this.
|
|
bool end_of_input = 2;
|
|
}
|
|
|
|
message AudioToAudioControl {
|
|
// Free-form control event names. Initial set:
|
|
// "input_audio_buffer.commit" — user finished speaking
|
|
// "response.cancel" — abort in-flight generation
|
|
// "conversation.item.create" — inject a non-audio item (e.g.
|
|
// function_call_output as JSON in
|
|
// `payload`)
|
|
// "session.update" — re-configure mid-stream
|
|
string event = 1;
|
|
// Event-specific JSON payload.
|
|
bytes payload = 2;
|
|
}
|
|
|
|
message AudioToAudioResponse {
|
|
// Event identifies what this frame carries. Mirrors the OpenAI Realtime
|
|
// API server-event names where applicable. Initial set:
|
|
// "response.audio.delta"
|
|
// "response.audio_transcript.delta"
|
|
// "response.function_call_arguments.delta"
|
|
// "response.function_call_arguments.done"
|
|
// "response.done"
|
|
// "error"
|
|
string event = 1;
|
|
// Populated when event = response.audio.delta.
|
|
bytes pcm = 2;
|
|
// Populated alongside pcm to identify its rate. 0 => same as the
|
|
// session's negotiated output_sample_rate.
|
|
int32 sample_rate = 3;
|
|
// JSON payload for non-PCM events (transcript chunk, tool args, error
|
|
// body).
|
|
bytes meta = 4;
|
|
// Monotonic per-stream counter, useful for client reordering and
|
|
// debugging.
|
|
int64 sequence = 5;
|
|
}
|
|
|
|
message ModelMetadataResponse {
|
|
bool supports_thinking = 1;
|
|
string rendered_template = 2; // The rendered chat template with enable_thinking=true (empty if not applicable)
|
|
ToolFormatMarkers tool_format = 3; // Auto-detected tool format markers from differential template analysis
|
|
string media_marker = 4; // Marker the backend expects in the prompt for each multimodal input (images/audio/video). Empty when the backend does not use a marker.
|
|
}
|
|
|
|
// Fine-tuning messages
|
|
|
|
message FineTuneRequest {
|
|
// Model identification
|
|
string model = 1; // HF model name or local path
|
|
string training_type = 2; // "lora", "loha", "lokr", "full" — what parameters to train
|
|
string training_method = 3; // "sft", "dpo", "grpo", "rloo", "reward", "kto", "orpo", "network_training"
|
|
|
|
// Adapter config (universal across LoRA/LoHa/LoKr for LLM + diffusion)
|
|
int32 adapter_rank = 10; // LoRA rank (r), default 16
|
|
int32 adapter_alpha = 11; // scaling factor, default 16
|
|
float adapter_dropout = 12; // default 0.0
|
|
repeated string target_modules = 13; // layer names to adapt
|
|
|
|
// Universal training hyperparameters
|
|
float learning_rate = 20; // default 2e-4
|
|
int32 num_epochs = 21; // default 3
|
|
int32 batch_size = 22; // default 2
|
|
int32 gradient_accumulation_steps = 23; // default 4
|
|
int32 warmup_steps = 24; // default 5
|
|
int32 max_steps = 25; // 0 = use epochs
|
|
int32 save_steps = 26; // 0 = only save final
|
|
float weight_decay = 27; // default 0.01
|
|
bool gradient_checkpointing = 28;
|
|
string optimizer = 29; // adamw_8bit, adamw, sgd, adafactor, prodigy
|
|
int32 seed = 30; // default 3407
|
|
string mixed_precision = 31; // fp16, bf16, fp8, no
|
|
|
|
// Dataset
|
|
string dataset_source = 40; // HF dataset ID, local file/dir path
|
|
string dataset_split = 41; // train, test, etc.
|
|
|
|
// Output
|
|
string output_dir = 50;
|
|
string job_id = 51; // client-assigned or auto-generated
|
|
|
|
// Resume training from a checkpoint
|
|
string resume_from_checkpoint = 55; // path to checkpoint dir to resume from
|
|
|
|
// Backend-specific AND method-specific extensibility
|
|
map<string, string> extra_options = 60;
|
|
}
|
|
|
|
message FineTuneJobResult {
|
|
string job_id = 1;
|
|
bool success = 2;
|
|
string message = 3;
|
|
}
|
|
|
|
message FineTuneProgressRequest {
|
|
string job_id = 1;
|
|
}
|
|
|
|
message FineTuneProgressUpdate {
|
|
string job_id = 1;
|
|
int32 current_step = 2;
|
|
int32 total_steps = 3;
|
|
float current_epoch = 4;
|
|
float total_epochs = 5;
|
|
float loss = 6;
|
|
float learning_rate = 7;
|
|
float grad_norm = 8;
|
|
float eval_loss = 9;
|
|
float eta_seconds = 10;
|
|
float progress_percent = 11;
|
|
string status = 12; // queued, caching, loading_model, loading_dataset, training, saving, completed, failed, stopped
|
|
string message = 13;
|
|
string checkpoint_path = 14; // set when a checkpoint is saved
|
|
string sample_path = 15; // set when a sample is generated (video/image backends)
|
|
map<string, float> extra_metrics = 16; // method-specific metrics
|
|
}
|
|
|
|
message FineTuneStopRequest {
|
|
string job_id = 1;
|
|
bool save_checkpoint = 2;
|
|
}
|
|
|
|
message ListCheckpointsRequest {
|
|
string output_dir = 1;
|
|
}
|
|
|
|
message ListCheckpointsResponse {
|
|
repeated CheckpointInfo checkpoints = 1;
|
|
}
|
|
|
|
message CheckpointInfo {
|
|
string path = 1;
|
|
int32 step = 2;
|
|
float epoch = 3;
|
|
float loss = 4;
|
|
string created_at = 5;
|
|
}
|
|
|
|
message ExportModelRequest {
|
|
string checkpoint_path = 1;
|
|
string output_path = 2;
|
|
string export_format = 3; // lora, loha, lokr, merged_16bit, merged_4bit, gguf, diffusers
|
|
string quantization_method = 4; // for GGUF: q4_k_m, q5_k_m, q8_0, f16, etc.
|
|
string model = 5; // base model name (for merge operations)
|
|
map<string, string> extra_options = 6;
|
|
}
|
|
|
|
// Quantization messages
|
|
|
|
message QuantizationRequest {
|
|
string model = 1; // HF model name or local path
|
|
string quantization_type = 2; // q4_k_m, q5_k_m, q8_0, f16, etc.
|
|
string output_dir = 3; // where to write output files
|
|
string job_id = 4; // client-assigned job ID
|
|
map<string, string> extra_options = 5; // hf_token, custom flags, etc.
|
|
}
|
|
|
|
message QuantizationJobResult {
|
|
string job_id = 1;
|
|
bool success = 2;
|
|
string message = 3;
|
|
}
|
|
|
|
message QuantizationProgressRequest {
|
|
string job_id = 1;
|
|
}
|
|
|
|
message QuantizationProgressUpdate {
|
|
string job_id = 1;
|
|
float progress_percent = 2;
|
|
string status = 3; // queued, downloading, converting, quantizing, completed, failed, stopped
|
|
string message = 4;
|
|
string output_file = 5; // set when completed — path to the output GGUF file
|
|
map<string, float> extra_metrics = 6; // e.g. file_size_mb, compression_ratio
|
|
}
|
|
|
|
message QuantizationStopRequest {
|
|
string job_id = 1;
|
|
}
|
|
|
|
// ForwardHeader is one HTTP header on the request or response. Headers
|
|
// like Authorization are typically injected by the backend (from the
|
|
// resolved API key) rather than passed through from the client.
|
|
message ForwardHeader {
|
|
string name = 1;
|
|
string value = 2;
|
|
}
|
|
|
|
// ForwardRequest is a streamed HTTP request to the upstream. First
|
|
// message carries path/method/headers; subsequent messages carry
|
|
// body_chunk only. All fields except body_chunk are honoured on the
|
|
// first message and ignored thereafter.
|
|
message ForwardRequest {
|
|
string path = 1; // e.g. "/v1/chat/completions" — appended to the model's upstream_url
|
|
string method = 2; // usually "POST"
|
|
repeated ForwardHeader headers = 3;
|
|
bytes body_chunk = 4;
|
|
}
|
|
|
|
// ForwardReply is a streamed HTTP response from the upstream. First
|
|
// message carries status/headers; subsequent messages carry body_chunk
|
|
// only. SSE responses arrive as a sequence of body_chunk frames; the
|
|
// caller is responsible for any parsing.
|
|
message ForwardReply {
|
|
int32 status = 1;
|
|
repeated ForwardHeader headers = 2;
|
|
bytes body_chunk = 3;
|
|
}
|