Files
LocalAI/docs/content/features/audio-classification.md
T
localai-org-maint-botandEttore Di Giacinto fb8b7a359a fix(distributed): stage sound detection audio (#11907)
* fix(distributed): stage sound detection audio

Sound detection passes frontend temporary paths directly to remote
workers, unlike transcription. Stage the WAV before classification so
CED can read it without a shared temporary directory.

Preserve the original request for retries and propagate staging errors
without calling the backend. Cover staging, request preservation, and
error handling with regression tests.

Assisted-by: Codex:GPT-6 golangci-lint

* test(distributed): verify routed sound staging

Call sound detection through the client returned by SmartRouter.Route.
This checks interface dispatch through both routing wrappers, rather
than constructing FileStagingClient directly.

The test fails without the sound-staging override and passes with it.

Assisted-by: Codex:GPT-6 golangci-lint

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-09-07 17:12:37 +02:00

2.3 KiB

+++ disableToc = false title = "Sound Classification" weight = 32 url = "/features/audio-classification/" +++

Sound-event classification (audio tagging) answers the question "what am I hearing?" - given an audio clip, it returns a list of scored AudioSet labels (e.g. Baby cry, infant cry, Glass breaking, Dog bark, Alarm).

LocalAI exposes this through the /v1/audio/classification endpoint, modelled after /v1/audio/transcriptions. The reference backend is ced.cpp (CED, a 527-class AudioSet tagger), a small ViT over a log-mel spectrogram ported to ggml with full PyTorch parity. Apache-2.0 weights are redistributable as GGUF.

Because classification is exposed as a regular OpenAI-style endpoint, any HTTP client works - there is no Python dependency on the consumer side.

In distributed mode, LocalAI stages uploaded audio and realtime sound-detection windows on the selected worker before classification. The API server and worker do not need a shared temporary directory.

Endpoint

POST /v1/audio/classification
Content-Type: multipart/form-data
Field Type Description
file file (required) audio file in any format ffmpeg accepts
model string (required) name of the sound-classification-capable model (e.g. ced-base-f16)
top_k int number of top tags to return (0 = backend default)
threshold float drop tags scoring below this value

Response

{
  "model": "ced-base-f16",
  "detections": [
    {"index": 23, "label": "Baby cry, infant cry", "score": 0.87},
    {"index": 22, "label": "Crying, sobbing", "score": 0.41}
  ]
}

Detections are returned in score-descending order. Scores are per-class probabilities (multi-label, independent), so they do not sum to 1.

Example

First install a classification model from the gallery (the example below uses ced-base-f16):

local-ai run ced-base-f16
curl http://localhost:8080/v1/audio/classification \
  -H "Content-Type: multipart/form-data" \
  -F file="@/path/to/clip.wav" \
  -F model="ced-base-f16" \
  -F top_k=10

See also

  • [Audio to Text]({{% relref "audio-to-text" %}}) - speech transcription
  • [Speaker Diarization]({{% relref "audio-diarization" %}}) - who spoke when