Parakeet Ultra GGUF
GGUF weights of Moondream Parakeet Ultra for s2s.cpp, a C++17/GGML voice assistant. Parakeet Ultra is NVIDIA Parakeet TDT 0.6B v3 post trained by Moondream: same FastConformer encoder and TDT transducer decoder, same config and tokenizer, 25 European languages detected on their own, punctuation and capitalization included. Moondream reports a lower word error rate than v3 in every group they test: English, multilingual, business speech and background noise. Runs on CPU, CUDA, Metal, Vulkan, SYCL.
s2s.cpp loads this model when it is present and Parakeet TDT 0.6B v3 otherwise.
Files
| variant | size | use case |
|---|---|---|
| F32 | 2.5 GB | reference, source of the quants |
| Q8_0 | 668 MB | recommended default |
| Q6_K | 537 MB | |
| Q5_K_M | 470 MB | |
| Q4_K_M | 407 MB | lowest VRAM, see parity |
The server loads the best quant it finds, up to Q8_0.
Quick start
git clone https://github.com/ServeurpersoCom/s2s.cpp.git
cd s2s.cpp
git submodule update --init
./buildcuda.sh
hf download Serveurperso/Parakeet-Ultra-GGUF parakeet-ultra-Q8_0.gguf --local-dir models
./build/parakeet-transcribe --model models/parakeet-ultra-Q8_0.gguf --file qwentts.cpp/examples/freeman.wav
./models.sh then fetches the rest of the pipeline and ./server.sh starts the voice assistant on http://localhost:8088.
Quantization policy
The quantizer of s2s.cpp mirrors llama-quantize: every 2D weight takes the K-quant of the variant, and in the M variants the attention values and the second feed forward linear are bumped to Q6_K on the first few, the last few and every third layer, the two places a low bit weight hurts the transcript first. Three groups stay F32 in every variant:
| tensor | why |
|---|---|
| subsampling kernels, depthwise conv kernels, sinusoid table | read by the f32 convolution path, precision worth more than the bytes |
| relative attention biases | added to the queries, no quantized add |
| 1D tensors (norms, biases) | as in llama.cpp |
The conversion folds each conformer BatchNorm into its depthwise kernel, so the graph has none. The upstream checkpoint also carries a small voice activity head (vad_head.*) that the transducer never reads, left out of the GGUF.
Parity
test-parakeet holds each stage against transformers ParakeetForTDT loaded with the Parakeet Ultra weights, on the example recording: mel, subsampling, first block, encoder states, projection, prediction network, joint, and the transcript.
| variant | encoder states cossim | transcript |
|---|---|---|
| F32 | 0.999999 | identical |
| Q8_0 | 0.999777 | identical |
| Q4_K_M | 0.988200 | 0.92 sequence match |
CUDA figures. F32 and Q8_0 give the same transcript on Vulkan and CPU. Q4_K_M drifts like v3 in the encoder, but the drift flips hesitation words ("gonna" read "going to"): 0.95 on Vulkan, identical on CPU. Stay on Q8_0.
License
Upstream model: Parakeet Ultra by Moondream, CC BY 4.0, derived from Parakeet TDT 0.6B v3 by NVIDIA, CC BY 4.0
GGUF tooling: s2s.cpp, MIT
- Downloads last month
- 174
4-bit
5-bit
6-bit
8-bit
32-bit
Model tree for Serveurperso/Parakeet-Ultra-GGUF
Base model
moondream/parakeet-ultra