vad-rs

Detects the start and stop of speech in an audio stream using a small neural model.

Silero VAD v5 inference in Rust via candle (CPU tensor ops). Accepts 24kHz PCM, resamples to 16kHz, and emits SpeechStart/SpeechEnd events with configurable thresholds and redemption-frame hysteresis. Weights converted from the original ONNX model to safetensors.

Runs
browser, WASM, native
Model
snakers4/silero-vad, 1.2MB, MIT
Source
idle-intelligence/vad-rs
HuggingFace
idle-intelligence/silero-vad-v5-safetensors
Status
maintained

← Home