* feat(parakeet-cpp): load bundle GGUF files and use their components by role A bundle GGUF holds several models (ASR, VAD, diarization, sound events, speaker encoder) in one file, each with its own licence. Detect a bundle at load through parakeet_capi_bundle_components_json and open components with parakeet_capi_load_component. The three symbols are probed together, so an older libparakeet.so still loads plain files as before. The only ASR component is the primary model; bundle_asr:<name> picks one when there are several. A Silero VAD component of the primary bundle is loaded without an option and serves /v1/vad and vad:true. The diar, ced and voice components load on request: diar_component, sound_component and speaker_component, or a companion option (diarization_model, sound_model, speaker_model, vad_model) that names a bundle, even the model file itself. vad_component picks a VAD component and implies vad:true. A role the bundle cannot fill fails the load with the component list, and a diarization or sound request on a model without that role names the bundle components. Every existing option and single-file model behaves as before. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * chore(parakeet-cpp): bump parakeet.cpp to 781a973 Brings in the bundle GGUF format and its C-API (parakeet_capi_load_component, parakeet_capi_bundle_components_json, parakeet_capi_load_error). Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] * feat(parakeet-cpp): gallery entries for the bundle GGUF files, docs Add parakeet-cpp-bundle-small (338 MB: Parakeet TDT+CTC 110M, Nemotron-3- Diarization, CED-Small, WeSpeaker ResNet34-LM, Silero VAD), -standard (1.1 GB, Parakeet TDT 0.6B v3 instead of the 110M model) and -moondream-redux (215 MB: packed Redux and Silero VAD, CPU only). One install serves transcription, VAD, diarization, sound events and speaker naming through the component options. The existing single-purpose entries stay. A bundle has no single licence, so the entries use license: other and state the licence and credit of each component in the description, with the upstream inconsistency of the CED licence. The docs get a section on bundles in audio-to-text with the entries, the roles, the options and the licence notice, and pointers from the VAD, diarization and sound classification pages. A gallery test checks the file names, checksums, usecases and options. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code] --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
6.8 KiB
+++ disableToc = false title = "Voice Activity Detection (VAD)" weight = 35 url = "/features/voice-activity-detection/" +++
Voice Activity Detection (VAD) identifies segments of speech in audio data. LocalAI provides a /v1/vad endpoint powered by the Silero VAD backend.
The [audio.cpp backend]({{%relref "features/audio-cpp" %}}) also serves this endpoint, and ships the silero_vad and marblenet_vad assets inside its own package, so VAD works there with nothing to download (model: bundled:silero_vad plus the family:silero_vad option).
API
- Method:
POST - Endpoints:
/v1/vad,/vad
Request
The request body is JSON with the following fields:
| Parameter | Type | Required | Description |
|---|---|---|---|
model |
string |
Yes | Model name (e.g. silero-vad) |
audio |
float32[] |
Yes | Array of audio samples (16kHz PCM float) |
Response
Returns a JSON object with detected speech segments:
| Field | Type | Description |
|---|---|---|
segments |
array |
List of detected speech segments |
segments[].start |
float |
Start time in seconds |
segments[].end |
float |
End time in seconds |
Usage
Example request
The /v1/vad endpoint expects the audio field to be an array of raw
16kHz mono PCM samples as float32 values, so the request body is usually
built from a real audio file rather than typed by hand.
First convert any audio file to 16kHz mono with ffmpeg:
ffmpeg -i input.mp3 -ar 16000 -ac 1 -f wav speech.wav
Then load the samples and POST them (this snippet needs
pip install soundfile numpy requests):
import soundfile as sf
import numpy as np
import requests
audio, sample_rate = sf.read("speech.wav")
if audio.ndim > 1:
audio = audio.mean(axis=1) # downmix to mono
samples = audio.astype(np.float32).tolist()
response = requests.post(
"http://localhost:8080/v1/vad",
json={"model": "silero-vad", "audio": samples},
)
print(response.json())
Example response
{
"segments": [
{
"start": 0.5,
"end": 2.3
},
{
"start": 3.1,
"end": 5.8
}
]
}
Model Configuration
Create a YAML configuration file for the VAD model:
name: silero-vad
backend: silero-vad
Detection parameters can be overridden via model options (key:value entries):
name: silero-vad
backend: silero-vad
options:
- threshold:0.55
- min_silence_duration_ms:50
- speech_pad_ms:450
Supported options:
| Option | Type | Default | Description |
|---|---|---|---|
threshold |
float | 0.5 |
Speech probability threshold |
min_silence_duration_ms |
int | 100 |
Minimum silence before ending a speech segment |
speech_pad_ms |
int | 30 |
Padding added around each speech segment |
Thresholds must be greater than 0 and less than 1. Durations must be nonnegative integers. Malformed values, negative durations, and NaN thresholds are ignored; the default or last valid value remains in use.
Reload the model (or restart LocalAI) after changing these options.
parakeet-cpp backend
The parakeet-cpp backend serves the same endpoint. It runs one of two detectors:
-
Silero VAD from a GGUF file (gallery entry
parakeet-cpp-silero-vad-f16, 1.3 MB). One probability per 32 ms. -
The VAD head of a full Moondream Ultra or Redux model (gallery entries
parakeet-cpp-vad-moondream-ultra-q8_0andparakeet-cpp-vad-moondream-redux-packed). One probability per 80 ms. The packed Redux file runs on CPU only. -
A VAD-only slice of that head (gallery entries
parakeet-cpp-vad-moondream-redux, 9.9 MB, andparakeet-cpp-vad-moondream-ultra, 6.0 MB). The slice is cut out of the full model without retraining, so the segments are byte-identical to the full model's head, and the speed is the same. Compared with loading the whole model (213 MB to 1.4 GB), the file is 6 to 10 MB, loads in a few milliseconds instead of 0.1 to 0.7 s, and needs about 245 MiB of peak memory for a 33 s clip instead of 0.6 to 1.6 GiB. A slice cannot transcribe, and it needs a parakeet.cpp build with VAD-only GGUF support (pin e53a253 or newer). -
A bundle GGUF (gallery entries
parakeet-cpp-bundle-small,parakeet-cpp-bundle-standardandparakeet-cpp-bundle-moondream-redux). The bundle holds a Silero component next to the ASR model, and the backend uses it for this endpoint with no option. See [Bundle GGUF files]({{%relref "features/audio-to-text" %}}#bundle-gguf-files-several-models-in-one-file).
The entry parakeet-cpp-vad installs Silero. The detectors differ and are not variants of one model, so install the entry of the VAD head by name if you want it (parakeet-cpp-vad-moondream-redux or parakeet-cpp-vad-moondream-ultra for the small files). The request is the same as above: audio is 16 kHz mono float32 PCM, and the response lists segments with start and end in seconds. An ASR model that has no VAD head fails the request with model has no VAD head.
name: parakeet-vad
backend: parakeet-cpp
known_usecases:
- vad
parameters:
model: parakeet-cpp/silero-vad-f16.gguf
options:
- vad_threshold:0.5
- vad_min_pause:0.1
All options are optional. An unset value keeps the default of the detector in use (the library defaults differ between Silero and the head):
| Option | Unit | Silero default | Head default | Description |
|---|---|---|---|---|
vad_threshold |
0 to 1 | 0.5 |
0.5 |
Speech probability threshold |
vad_min_pause |
seconds | 0.1 |
0.2 |
A silence this long separates two segments; shorter gaps merge |
vad_min_speech |
seconds | 0.25 |
0.1 |
Shorter speech runs are dropped |
vad_speech_pad |
seconds | 0.03 |
0 |
Padding added around each segment |
Option names differ from the Silero backend above (min_silence_duration_ms and speech_pad_ms are in milliseconds there). The same options tune transcription with vad:true or vad_model; see [audio to text]({{%relref "features/audio-to-text" %}}). Requests on one loaded model run one at a time.
Detection Parameters
The Silero VAD backend uses the following internal defaults (overridable via options above):
- Sample rate: 16kHz
- Threshold: 0.5
- Min silence duration: 100ms
- Speech pad duration: 30ms
Error Responses
| Status Code | Description |
|---|---|
| 400 | Missing or invalid model or audio field |
| 500 | Backend error during VAD processing |