* fix(distributed): stage sound detection audio Sound detection passes frontend temporary paths directly to remote workers, unlike transcription. Stage the WAV before classification so CED can read it without a shared temporary directory. Preserve the original request for retries and propagate staging errors without calling the backend. Cover staging, request preservation, and error handling with regression tests. Assisted-by: Codex:GPT-6 golangci-lint * test(distributed): verify routed sound staging Call sound detection through the client returned by SmartRouter.Route. This checks interface dispatch through both routing wrappers, rather than constructing FileStagingClient directly. The test fails without the sound-staging override and passes with it. Assisted-by: Codex:GPT-6 golangci-lint --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2.3 KiB
+++ disableToc = false title = "Sound Classification" weight = 32 url = "/features/audio-classification/" +++
Sound-event classification (audio tagging) answers the question "what am I hearing?" - given an audio clip, it returns a list of scored AudioSet labels (e.g. Baby cry, infant cry, Glass breaking, Dog bark, Alarm).
LocalAI exposes this through the /v1/audio/classification endpoint, modelled after /v1/audio/transcriptions. The reference backend is ced.cpp (CED, a 527-class AudioSet tagger), a small ViT over a log-mel spectrogram ported to ggml with full PyTorch parity. Apache-2.0 weights are redistributable as GGUF.
Because classification is exposed as a regular OpenAI-style endpoint, any HTTP client works - there is no Python dependency on the consumer side.
In distributed mode, LocalAI stages uploaded audio and realtime sound-detection windows on the selected worker before classification. The API server and worker do not need a shared temporary directory.
Endpoint
POST /v1/audio/classification
Content-Type: multipart/form-data
| Field | Type | Description |
|---|---|---|
file |
file (required) | audio file in any format ffmpeg accepts |
model |
string (required) | name of the sound-classification-capable model (e.g. ced-base-f16) |
top_k |
int | number of top tags to return (0 = backend default) |
threshold |
float | drop tags scoring below this value |
Response
{
"model": "ced-base-f16",
"detections": [
{"index": 23, "label": "Baby cry, infant cry", "score": 0.87},
{"index": 22, "label": "Crying, sobbing", "score": 0.41}
]
}
Detections are returned in score-descending order. Scores are per-class probabilities (multi-label, independent), so they do not sum to 1.
Example
First install a classification model from the gallery (the example below uses ced-base-f16):
local-ai run ced-base-f16
curl http://localhost:8080/v1/audio/classification \
-H "Content-Type: multipart/form-data" \
-F file="@/path/to/clip.wav" \
-F model="ced-base-f16" \
-F top_k=10
See also
- [Audio to Text]({{% relref "audio-to-text" %}}) - speech transcription
- [Speaker Diarization]({{% relref "audio-diarization" %}}) - who spoke when