feat(parakeet-cpp): VAD (Silero and the Moondream head), vad_model option and gallery entries (#12463)

* feat(parakeet-cpp): implement the VAD call and add the vad_model option

The backend now serves the VAD gRPC call (POST /vad and /v1/vad) with the
standalone VAD of libparakeet. It accepts a Silero VAD GGUF as the model
file, or an ASR model with a VAD head (Moondream Ultra and Redux). The
request audio is float32 PCM at 16 kHz; the response lists the speech
segments in seconds, like the silero-vad backend. A model with neither
fails the request with the library message.

The vad_threshold, vad_min_pause, vad_min_speech, vad_speech_pad and
vad_max_segment options tune the segmenter. Unset values keep the
defaults of the detector in use, and a bad value fails the load.

The vad_model option names a Silero GGUF, resolved against the models
directory like the other companion files. It lets any ASR model cut long
audio at pauses through parakeet_capi_transcribe_path_json_vad_with, and
it implies vad. vad:true alone still uses the model's own head.

The new symbols are probed like the existing optional ones. A library
without them still loads; the feature that needs one fails with a clear
message only when it is used.

Assisted-by: Claude:claude-sonnet-5-5 [go test]

* feat(gallery): add parakeet-cpp VAD entries and a v3 plus Silero example

Add VAD-only entries for the parakeet-cpp backend: the VAD heads of
Moondream Redux (packed, CPU) and Ultra (Q8_0), which share their files
with the existing ASR entries, and Silero VAD v6.2.3 as a GGUF (MIT,
Silero Team). The parakeet-cpp-vad entry installs Silero; it has no variants,
because variant ranking prefers the larger build that fits and these are
different detectors.

Add parakeet-cpp-tdt-0.6b-v3-silero-vad, a v3 entry that sets vad_model
so long audio is cut at pauses by Silero.

The Silero GGUF entries point at the intended Hugging Face URL of the
file; the existing silero-vad entries are unchanged. A test checks the
usecases, the shared files and the default entry and the vad_model reference.

Assisted-by: Claude:claude-sonnet-5-5 [go test]

* docs: describe parakeet-cpp VAD and the vad_model option

Document the VAD endpoint on the parakeet-cpp backend (Silero GGUF and
the VAD heads of Moondream Ultra and Redux), the vad_* tuning options,
and the vad_model option that lets an ASR model without a VAD head cut
long audio with Silero.

Assisted-by: Claude:claude-sonnet-5-5

* chore(parakeet-cpp): bump parakeet.cpp to 6165e3d

Pin the release that adds the standalone VAD (Ultra/Redux head and
Silero) and the C API calls the backend now uses.

Assisted-by: Claude:claude-sonnet-5-5

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
mudler-agentandEttore Di Giacinto authored and GitHub committed 2026-10-04 01:02:09 +02:00
1 parent 00faf4ec26
commit 6c718fdda7
10 files changed
+807 -12

No files matched your search

+27
View File
@@ -270,6 +270,33 @@ options:
`vad:true` applies to offline transcription only and bypasses dynamic batching, because the batched entry point has no VAD variant. Streaming is not affected. A model without a VAD head fails each request with `model has no VAD head`, and a `libparakeet.so` that is too old to export the VAD entry point fails the load. Remove the option for models that have no VAD head.
### Cutting long audio with Silero (`vad_model`)
A model without a VAD head, such as `parakeet-cpp-tdt-0.6b-v3` or a Nemotron model, can cut long audio with [Silero VAD](https://github.com/snakers4/silero-vad) instead. Name a Silero GGUF in the `vad_model` option. The path is resolved against the models directory, like the other companion files. `vad_model` implies `vad`:
```yaml
name: parakeet-v3-silero
backend: parakeet-cpp
parameters:
model: parakeet-cpp/tdt-0.6b-v3-f16.gguf
options:
- vad_model:parakeet-cpp/silero-vad-f16.gguf # Silero GGUF that cuts long audio at pauses
- vad_min_pause:0.3 # optional, seconds
```
The gallery entry `parakeet-cpp-tdt-0.6b-v3-silero-vad` installs both files with this configuration. Audio of 30 seconds or less is transcribed whole and the VAD does not run. `vad:true` alone keeps meaning "use the model's own head". With `vad_model` set, the Silero model is used even if the ASR model has a head.
The segmenter options below apply to both `vad:true` and `vad_model`. Each is optional; an unset value keeps the default of the detector in use, and a bad value fails the load:
| Option | Unit | Meaning |
|---|---|---|
| `vad_threshold` | 0 to 1 | A frame is speech when its probability is at least this |
| `vad_min_pause` | seconds | A silence this long separates two pieces |
| `vad_min_speech` | seconds | Shorter speech runs are dropped |
| `vad_max_segment` | seconds | Cap on the length of a piece (default 30) |
`vad_speech_pad` (seconds) pads each region and only affects the [VAD endpoint]({{%relref "features/voice-activity-detection" %}}). `vad_model` needs a `libparakeet.so` that exports `parakeet_capi_transcribe_path_json_vad_with`; an older library fails the load with a message that names it.
## See also
- [Audio Transform]({{< relref "audio-transform.md" >}}) - clean up the audio (echo cancellation, noise suppression, dereverberation) before passing it to a transcription model.
@@ -117,6 +117,38 @@ Malformed values, negative durations, and NaN thresholds are ignored; the defaul
Reload the model (or restart LocalAI) after changing these options.
## parakeet-cpp backend
The `parakeet-cpp` backend serves the same endpoint. It runs one of two detectors:
- **Silero VAD** from a GGUF file (gallery entry `parakeet-cpp-silero-vad-f16`, 1.3 MB). One probability per 32 ms.
- **The VAD head** of a Moondream Ultra or Redux model (gallery entries `parakeet-cpp-vad-moondream-ultra-q8_0` and `parakeet-cpp-vad-moondream-redux-packed`). One probability per 80 ms. The packed Redux file runs on CPU only.
The entry `parakeet-cpp-vad` installs Silero. The detectors differ and are not variants of one model, so install the entry of the VAD head by name if you want it. The request is the same as above: `audio` is 16 kHz mono float32 PCM, and the response lists `segments` with `start` and `end` in seconds. An ASR model that has no VAD head fails the request with `model has no VAD head`.
```yaml
name: parakeet-vad
backend: parakeet-cpp
known_usecases:
- vad
parameters:
model: parakeet-cpp/silero-vad-f16.gguf
options:
- vad_threshold:0.5
- vad_min_pause:0.1
```
All options are optional. An unset value keeps the default of the detector in use (the library defaults differ between Silero and the head):
| Option | Unit | Silero default | Head default | Description |
|--------|------|---------------:|-------------:|-------------|
| `vad_threshold` | 0 to 1 | `0.5` | `0.5` | Speech probability threshold |
| `vad_min_pause` | seconds | `0.1` | `0.2` | A silence this long separates two segments; shorter gaps merge |
| `vad_min_speech` | seconds | `0.25` | `0.1` | Shorter speech runs are dropped |
| `vad_speech_pad` | seconds | `0.03` | `0` | Padding added around each segment |
Option names differ from the Silero backend above (`min_silence_duration_ms` and `speech_pad_ms` are in milliseconds there). The same options tune transcription with `vad:true` or `vad_model`; see [audio to text]({{%relref "features/audio-to-text" %}}). Requests on one loaded model run one at a time.
## Detection Parameters
The Silero VAD backend uses the following internal defaults (overridable via `options` above):