silero-vad: transcribe.cpp GGUF

GGUF conversions of snakers4/silero-vad for use with transcribe.cpp.

Ported from upstream commit 5cd7945, pinned 2026-10-09. Validated against the silero-vad 6.2.3 (TorchScript, CPU, F32) reference at transcribe.cpp commit 2674b95d on 2026-10-10.

Voice activity detection by the Silero Team. Not a transcription model: the VAD role produces per-frame speech probabilities (offline or streaming) and offline speech segments. CPU is the default backend for reference fidelity. transcribe.cpp also loads whisper.cpp's ggml-silero-v5.1.2.bin / v6.2.0.bin.

Downloads

Quantization Download Size
F32 silero-vad-v6.2-F32.gguf 1 MB

Usage

Build transcribe.cpp and run on a mono WAV at the model's sample rate:

cmake -B build && cmake --build build --target transcribe-cli
build/bin/transcribe-cli -m silero-vad-v6.2-F32.gguf input.wav

From the C API, use the VAD role (include/transcribe/vad.h): offline segments with transcribe_vad_run, live per-frame probabilities with transcribe_vad_stream_feed.

See the model page and VAD usage.

License

Inherited from the base model: MIT. See the upstream model card for full terms.

Downloads last month
9
GGUF
Model size
310k params
Architecture
silero_vad
Hardware compatibility
Log In to add your hardware

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support