Parakeet Ultra GGUF

GGUF weights of Moondream Parakeet Ultra for s2s.cpp, a C++17/GGML voice assistant. Parakeet Ultra is NVIDIA Parakeet TDT 0.6B v3 post trained by Moondream: same FastConformer encoder and TDT transducer decoder, same config and tokenizer, 25 European languages detected on their own, punctuation and capitalization included. Moondream reports a lower word error rate than v3 in every group they test: English, multilingual, business speech and background noise. Runs on CPU, CUDA, Metal, Vulkan, SYCL.

s2s.cpp loads this model when it is present and Parakeet TDT 0.6B v3 otherwise.

Files

variant size use case
F32 2.5 GB reference, source of the quants
Q8_0 668 MB recommended default
Q6_K 537 MB
Q5_K_M 470 MB
Q4_K_M 407 MB lowest VRAM, see parity

The server loads the best quant it finds, up to Q8_0.

Quick start

git clone https://github.com/ServeurpersoCom/s2s.cpp.git
cd s2s.cpp
git submodule update --init
./buildcuda.sh
hf download Serveurperso/Parakeet-Ultra-GGUF parakeet-ultra-Q8_0.gguf --local-dir models
./build/parakeet-transcribe --model models/parakeet-ultra-Q8_0.gguf --file qwentts.cpp/examples/freeman.wav

./models.sh then fetches the rest of the pipeline and ./server.sh starts the voice assistant on http://localhost:8088.

Quantization policy

The quantizer of s2s.cpp mirrors llama-quantize: every 2D weight takes the K-quant of the variant, and in the M variants the attention values and the second feed forward linear are bumped to Q6_K on the first few, the last few and every third layer, the two places a low bit weight hurts the transcript first. Three groups stay F32 in every variant:

tensor why
subsampling kernels, depthwise conv kernels, sinusoid table read by the f32 convolution path, precision worth more than the bytes
relative attention biases added to the queries, no quantized add
1D tensors (norms, biases) as in llama.cpp

The conversion folds each conformer BatchNorm into its depthwise kernel, so the graph has none. The upstream checkpoint also carries a small voice activity head (vad_head.*) that the transducer never reads, left out of the GGUF.

Parity

test-parakeet holds each stage against transformers ParakeetForTDT loaded with the Parakeet Ultra weights, on the example recording: mel, subsampling, first block, encoder states, projection, prediction network, joint, and the transcript.

variant encoder states cossim transcript
F32 0.999999 identical
Q8_0 0.999777 identical
Q4_K_M 0.988200 0.92 sequence match

CUDA figures. F32 and Q8_0 give the same transcript on Vulkan and CPU. Q4_K_M drifts like v3 in the encoder, but the drift flips hesitation words ("gonna" read "going to"): 0.95 on Vulkan, identical on CPU. Stay on Q8_0.

License

Upstream model: Parakeet Ultra by Moondream, CC BY 4.0, derived from Parakeet TDT 0.6B v3 by NVIDIA, CC BY 4.0

GGUF tooling: s2s.cpp, MIT

Downloads last month
174
GGUF
Model size
0.6B params
Architecture
parakeet-tdt
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Serveurperso/Parakeet-Ultra-GGUF

Quantized
(22)
this model