Files
LocalAI/docs/content/features/voice-activity-detection.md
T
372b1f8983 feat(silero-vad): allow threshold/silence/pad via model options (#12430)
* feat(silero-vad): allow threshold/silence/pad via model options

Signed-off-by: anton ziderer <Antonziderer@mail.ru>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(silero-vad): ignore NaN thresholds

NaN passes the detector validation but prevents speech comparisons from succeeding.
Ignore it like malformed input and document the option validation.
Convert the option tests to Ginkgo and cover invalid overrides.

Assisted-by: Codex:gpt-6
Signed-off-by: anton ziderer <Antonziderer@mail.ru>

---------

Signed-off-by: anton ziderer <Antonziderer@mail.ru>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-10-02 23:59:49 +02:00

3.9 KiB

+++ disableToc = false title = "Voice Activity Detection (VAD)" weight = 35 url = "/features/voice-activity-detection/" +++

Voice Activity Detection (VAD) identifies segments of speech in audio data. LocalAI provides a /v1/vad endpoint powered by the Silero VAD backend.

The [audio.cpp backend]({{%relref "features/audio-cpp" %}}) also serves this endpoint, and ships the silero_vad and marblenet_vad assets inside its own package, so VAD works there with nothing to download (model: bundled:silero_vad plus the family:silero_vad option).

API

  • Method: POST
  • Endpoints: /v1/vad, /vad

Request

The request body is JSON with the following fields:

Parameter Type Required Description
model string Yes Model name (e.g. silero-vad)
audio float32[] Yes Array of audio samples (16kHz PCM float)

Response

Returns a JSON object with detected speech segments:

Field Type Description
segments array List of detected speech segments
segments[].start float Start time in seconds
segments[].end float End time in seconds

Usage

Example request

The /v1/vad endpoint expects the audio field to be an array of raw 16kHz mono PCM samples as float32 values, so the request body is usually built from a real audio file rather than typed by hand.

First convert any audio file to 16kHz mono with ffmpeg:

ffmpeg -i input.mp3 -ar 16000 -ac 1 -f wav speech.wav

Then load the samples and POST them (this snippet needs pip install soundfile numpy requests):

import soundfile as sf
import numpy as np
import requests

audio, sample_rate = sf.read("speech.wav")
if audio.ndim > 1:
    audio = audio.mean(axis=1)  # downmix to mono
samples = audio.astype(np.float32).tolist()

response = requests.post(
    "http://localhost:8080/v1/vad",
    json={"model": "silero-vad", "audio": samples},
)
print(response.json())

Example response

{
  "segments": [
    {
      "start": 0.5,
      "end": 2.3
    },
    {
      "start": 3.1,
      "end": 5.8
    }
  ]
}

Model Configuration

Create a YAML configuration file for the VAD model:

name: silero-vad
backend: silero-vad

Detection parameters can be overridden via model options (key:value entries):

name: silero-vad
backend: silero-vad
options:
  - threshold:0.55
  - min_silence_duration_ms:50
  - speech_pad_ms:450

Supported options:

Option Type Default Description
threshold float 0.5 Speech probability threshold
min_silence_duration_ms int 100 Minimum silence before ending a speech segment
speech_pad_ms int 30 Padding added around each speech segment

Thresholds must be greater than 0 and less than 1. Durations must be nonnegative integers. Malformed values, negative durations, and NaN thresholds are ignored; the default or last valid value remains in use.

Reload the model (or restart LocalAI) after changing these options.

Detection Parameters

The Silero VAD backend uses the following internal defaults (overridable via options above):

  • Sample rate: 16kHz
  • Threshold: 0.5
  • Min silence duration: 100ms
  • Speech pad duration: 30ms

Error Responses

Status Code Description
400 Missing or invalid model or audio field
500 Backend error during VAD processing