backend(audio-cpp): carry transform text through params, add the s2s entry

AudioTransform is audio-in / audio-out and its proto message has no text
field, but not every task it routes to is audio-only. vevo2's speech-to-speech
route is a text and prosody route: session.cpp:897 fills refs.target_text from
request.text_input and nowhere else, and the run refuses without one with
"Vevo2 text/prosody route requires text_input or target_text". The params map
is the only channel this RPC has that reaches the engine, so the text travels
through it and apply_transform_text_input unpacks it after the params have
been copied into task.options.

Before this, s2s was not awkward to reach through /audio/transform, it was
unreachable, and it was the last audio.cpp task kind with a real model and no
way to get to it.

target_text is canonical and text is its alias, the order vevo2's own option
table declares them in, so a request setting both gets the canonical one
rather than whichever the map happened to store first. An empty value falls
through to the next candidate instead of ending the search. language rides
along only when a text was found: on its own it conditions nothing, and
manufacturing a text_input for it would route a plain separation request
carrying a language hint through the text path. The keys are left in
task.options rather than erased, because vevo2's loader advertises target_text
as a request option and a family reading it there keeps working.

Nine tests, all confirmed failing on behaviour against a stub that returned
false before the implementation was written. Verified end to end afterwards:
vevo2-q8_0 with task:s2s and params[text] returns a 5.12 s 24 kHz output whose
closed-loop citrinet transcription is exact, and htdemucs separation with no
text param still returns its four stems, with and without params[stem].

audio-cpp-vevo2-speech-to-speech ships that route. Every audio.cpp task kind
with a loadable family now has a gallery entry; spk remains the only gap and
has no family upstream at all.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude:claude-opus-5 [Claude Code]
This commit is contained in:
Ettore Di Giacinto
2026-07-27 14:22:21 +00:00
committed by localai-org-maint-bot
parent 445a4434b0
commit ed14796a15
5 changed files with 229 additions and 0 deletions

View File

@@ -247,4 +247,34 @@ build_sound_generation_request(const backend::SoundGenerationRequest &request,
return task;
}
bool apply_transform_text_input(engine::runtime::TaskRequest &task) {
// Canonical first, alias second, and an empty value falls through to the
// next candidate rather than ending the search: a caller who sent
// target_text="" and text="the real one" meant the second one.
static const char *const kTextKeys[] = {"target_text", "text"};
std::string text;
for (const char *key : kTextKeys) {
const auto found = task.options.find(key);
if (found != task.options.end() && !found->second.empty()) {
text = found->second;
break;
}
}
if (text.empty()) {
return false;
}
engine::runtime::Transcript transcript;
transcript.text = std::move(text);
// Inside the has-text branch on purpose. See the header: a language on its
// own conditions nothing and must not manufacture a text_input.
const auto language = task.options.find("language");
if (language != task.options.end()) {
transcript.language = language->second;
}
task.text_input = std::move(transcript);
return true;
}
} // namespace audiocpp_backend

View File

@@ -74,4 +74,35 @@ engine::runtime::TaskRequest
build_sound_generation_request(const backend::SoundGenerationRequest &request,
std::optional<engine::runtime::AudioBuffer> source_audio);
// Lifts a text-conditioned transform route's text out of the request params
// into TaskRequest.text_input, and reports whether it set one.
//
// WHY THIS EXISTS. AudioTransform is an audio-in / audio-out RPC and its proto
// message has no text field, but not every task it routes to is audio-only.
// vevo2's speech-to-speech and prosody routes read their text from
// request.text_input (src/models/vevo2/session.cpp fills refs.target_text from
// exactly there and nowhere else) and refuse the run without one: "Vevo2
// text/prosody route requires text_input or target_text". The params map is the
// only channel AudioTransform has that reaches the engine, so the text travels
// through it and is unpacked here. Without this, s2s is not merely awkward to
// reach through this RPC, it is unreachable.
//
// CALL IT AFTER the params have been copied into task.options, and note that it
// does NOT erase the keys it reads. vevo2's loader advertises "target_text" in
// its own documented request-option table, so a family that looks there keeps
// finding it; the copy in text_input is what the session actually reads today.
//
// "target_text" is canonical and "text" is its alias, the same order vevo2's
// option table declares them in. A request setting both gets target_text, so
// the canonical spelling wins rather than whichever the map happened to store
// first. An empty value is not a text: it means the caller sent the key with
// nothing in it, and a family asked to vocalise "" should say so itself rather
// than be handed an empty Transcript that looks deliberate.
//
// "language" rides along when a text was found, and only then. On its own it
// conditions nothing, and setting text_input for it alone would turn a plain
// separation request that happened to carry a language hint into a text-routed
// one.
bool apply_transform_text_input(engine::runtime::TaskRequest &task);
} // namespace audiocpp_backend

View File

@@ -444,6 +444,105 @@ static void test_sound_generation_full() {
"sound: src passes through at its own rate and channel count");
}
// ---------------------------------------------------------------------------
// apply_transform_text_input
//
// AudioTransform has no text field on the wire, so a text-conditioned route
// (vevo2's speech-to-speech) can only be reached if the text travels as a
// param and is unpacked into text_input. Every assertion below pins a spelling
// or a precedence that a family actually depends on, not a shape that merely
// looks tidy.
static void test_transform_text_absent() {
engine::runtime::TaskRequest task;
task.options["stem"] = "vocals";
check(!apply_transform_text_input(task),
"transform text: reports false when no text key is present");
check(!task.text_input.has_value(),
"transform text: a request with no text keeps text_input unset");
}
static void test_transform_text_canonical_key() {
engine::runtime::TaskRequest task;
task.options["target_text"] = "sing this line";
check(apply_transform_text_input(task), "transform text: target_text reports true");
check(task.text_input.has_value() && task.text_input->text == "sing this line",
"transform text: target_text becomes text_input.text");
check(has_key(task.options, "target_text"),
"transform text: target_text survives in options for families that read it there");
}
static void test_transform_text_alias_key() {
engine::runtime::TaskRequest task;
task.options["text"] = "say this instead";
check(apply_transform_text_input(task), "transform text: text alias reports true");
check(task.text_input.has_value() && task.text_input->text == "say this instead",
"transform text: the text alias becomes text_input.text");
}
static void test_transform_text_canonical_wins() {
engine::runtime::TaskRequest task;
task.options["target_text"] = "canonical";
task.options["text"] = "alias";
check(apply_transform_text_input(task), "transform text: both keys reports true");
check(task.text_input.has_value() && task.text_input->text == "canonical",
"transform text: target_text wins over text, not whichever hashed first");
}
static void test_transform_text_empty_is_not_a_text() {
engine::runtime::TaskRequest task;
task.options["target_text"] = "";
check(!apply_transform_text_input(task),
"transform text: an empty target_text reports false");
check(!task.text_input.has_value(),
"transform text: an empty target_text leaves text_input unset");
}
static void test_transform_text_empty_canonical_falls_through_to_alias() {
engine::runtime::TaskRequest task;
task.options["target_text"] = "";
task.options["text"] = "the real one";
check(apply_transform_text_input(task),
"transform text: an empty canonical key does not mask a usable alias");
check(task.text_input.has_value() && task.text_input->text == "the real one",
"transform text: the alias is used when the canonical key is empty");
}
static void test_transform_text_language_rides_along() {
engine::runtime::TaskRequest task;
task.options["target_text"] = "vocalise me";
task.options["language"] = "ja";
check(apply_transform_text_input(task), "transform text: text plus language reports true");
check(task.text_input.has_value() && task.text_input->language == "ja",
"transform text: language lands on the Transcript alongside the text");
check(has_key(task.options, "language"),
"transform text: language survives in options too");
}
static void test_transform_language_alone_is_not_a_text() {
engine::runtime::TaskRequest task;
task.options["language"] = "ja";
check(!apply_transform_text_input(task),
"transform text: a language with no text reports false");
check(!task.text_input.has_value(),
"transform text: a language alone must not route a separation request through text");
}
static void test_transform_text_preserves_other_inputs() {
engine::runtime::TaskRequest task;
engine::runtime::AudioBuffer audio;
audio.sample_rate = 44100;
audio.channels = 2;
audio.samples = {0.1f, 0.2f, 0.3f, 0.4f};
task.audio_input = audio;
task.options["target_text"] = "keep the audio";
check(apply_transform_text_input(task), "transform text: with audio present reports true");
check(task.audio_input.has_value() && task.audio_input->samples.size() == 4 &&
task.audio_input->sample_rate == 44100,
"transform text: the source audio is untouched");
}
int main() {
test_voice_is_reference_file();
test_tts_shape();
@@ -457,6 +556,15 @@ int main() {
test_tts_language_and_params();
test_sound_generation_minimal();
test_sound_generation_full();
test_transform_text_absent();
test_transform_text_canonical_key();
test_transform_text_alias_key();
test_transform_text_canonical_wins();
test_transform_text_empty_is_not_a_text();
test_transform_text_empty_canonical_falls_through_to_alias();
test_transform_text_language_rides_along();
test_transform_language_alone_is_not_a_text();
test_transform_text_preserves_other_inputs();
if (failures != 0) {
fprintf(stderr, "%d check(s) failed\n", failures);

View File

@@ -1001,6 +1001,14 @@ public:
task.options[param.first] = param.second;
}
// AFTER the loop, so it reads exactly what the caller sent. This is
// what makes a text-conditioned route reachable through an RPC whose
// message has no text field; see apply_transform_text_input's header
// for why vevo2's speech-to-speech route is unreachable without it.
// A request carrying no text key is untouched, so separation and
// voice conversion pay nothing for this.
audiocpp_backend::apply_transform_text_input(task);
// Refused from the ROUTE, before the file reads and before the run.
// Only source separation produces named stems, and the route says
// whether this is separation without running anything: the identical

View File

@@ -47169,3 +47169,55 @@
- filename: audio-cpp/seed-vc-mlx-q8_0.gguf
sha256: f3e2469eb29762af96f9663ce62eab6e8c40f3798c6dc0835e7613f22b4b75f7
uri: huggingface://audio-cpp/audio.cpp-gguf/SeedVC-MLX-GGUF/seed-vc-mlx-q8_0.gguf
- name: audio-cpp-vevo2-speech-to-speech
url: github:mudler/LocalAI/gallery/virtual.yaml@master
urls:
- https://huggingface.co/audio-cpp/audio.cpp-gguf
- https://huggingface.co/RMSnow/Vevo2
- https://github.com/0xShug0/audio.cpp
description: |
Vevo2 (audio.cpp, Q8_0) - speech to speech, served by the audio-cpp backend
through /audio/transform. Post the source speech as `audio`, the target
speaker's clip as `reference`, and the words being spoken as
`params[text]`.
The text is not optional and not a hint. Vevo2's speech-to-speech route is
a text and prosody route: it reads the line it is resynthesising and
refuses the run without one. /audio/transform has no text field on the
wire, so the params map carries it and the backend lifts it into the
engine's text input; `params[target_text]` is the same thing under the name
Vevo2's own option table uses. Add `params[language]` when the line is not
English.
This entry pins task:s2s. Vevo2 also advertises plain voice conversion,
which needs no text and which auto-routing picks first, so without the pin
the text would be carried and then ignored. Drop the option to get voice
conversion from the same weights.
Q8_0 because it is the only build that fits alongside the rest of the
gallery; upstream records the original and 16-bit builds as the cleaner
ones, so expect drift on some routes.
license: other
tags:
- audio-cpp
- vevo2
- speech-to-speech
- voice-conversion
- audio-transform
- gguf
- ggml
- quantized
last_checked: "2026-07-27"
overrides:
backend: audio-cpp
known_usecases:
- audio_transform
name: audio-cpp-vevo2-speech-to-speech
options:
- task:s2s
parameters:
model: audio-cpp/vevo2-q8_0.gguf
files:
- filename: audio-cpp/vevo2-q8_0.gguf
sha256: f80a70facaaecfcf1aa417ef16ef091318c5e789b24c16a741233c64a72fcee8
uri: huggingface://audio-cpp/audio.cpp-gguf/Vevo2-GGUF/vevo2-q8_0.gguf