Instructions to use eulogik/polywhisper with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use eulogik/polywhisper with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="eulogik/polywhisper")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("eulogik/polywhisper") model = AutoModelForSpeechSeq2Seq.from_pretrained("eulogik/polywhisper", device_map="auto") - PEFT
How to use eulogik/polywhisper with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Download README.md from eulogik/polywhisper: direct link, hf CLI and curl.
- Browser
- Download file 13.7 kB
-
https://huggingface.co/eulogik/polywhisper/resolve/main/README.md
- Command line
-
hf download hf://eulogik/polywhisper/README.md
-
curl -L -o README.md https://huggingface.co/eulogik/polywhisper/resolve/main/README.md
language:
- hi
- ta
- te
- bn
- mr
license: mit
library_name: transformers
pipeline_tag: automatic-speech-recognition
base_model: openai/whisper-small
tags:
- polywhisper
- indic-asr
- hindi-asr
- tamil-speech-recognition
- telugu-stt
- bengali-asr
- marathi-speech-to-text
- speech-recognition
- multilingual
- lora
- whisper
- hindi
- tamil
- telugu
- bengali
- marathi
- indic-languages
- indian-languages
- automatic-speech-recognition
- speech-to-text
- low-resource-asr
- fleurs
- indicvoices
- onnx
- quantized
- efficient-asr
- edge-asr
- peft
datasets:
- ai4bharat/indicvoices-st
- google/fleurs
model-index:
- name: PolyWhisper v9 (Whisper-Small + Per-Language LoRA)
results:
- task:
type: automatic-speech-recognition
name: Hindi Speech Recognition
dataset:
name: FLEURS Hindi (hi_in)
type: google/fleurs
metrics:
- type: wer
value: 46.3
name: WER (beam=1, normalized)
- task:
type: automatic-speech-recognition
name: Tamil Speech Recognition
dataset:
name: FLEURS Tamil (ta_in)
type: google/fleurs
metrics:
- type: wer
value: 70.1
name: WER (beam=1, normalized)
- task:
type: automatic-speech-recognition
name: Telugu Speech Recognition
dataset:
name: FLEURS Telugu (te_in)
type: google/fleurs
metrics:
- type: wer
value: 100.1
name: WER (beam=1, normalized)
- task:
type: automatic-speech-recognition
name: Bengali Speech Recognition
dataset:
name: FLEURS Bengali (bn_in)
type: google/fleurs
metrics:
- type: wer
value: 130.2
name: WER (beam=1, normalized)
- task:
type: automatic-speech-recognition
name: Marathi Speech Recognition
dataset:
name: FLEURS Marathi (mr_in)
type: google/fleurs
metrics:
- type: wer
value: 116.9
name: WER (beam=1, normalized)
ποΈ PolyWhisper v9 β Efficient Multilingual Indic ASR
by Eulogik β Frontier Edge AI Β· Vernacular Intelligence Β· eulogik.com
TL;DR: PolyWhisper v9 is a research-ready automatic speech recognition (ASR) system for Hindi, Tamil, Telugu, Bengali, and Marathi. It pairs a frozen OpenAI Whisper-Small backbone (244M params) with tiny per-language LoRA adapters (~14MB each). Bengali WER drops β28.2% and Marathi β75.4% versus the no-augmentation baseline β at roughly 1% of the storage cost of full fine-tuning.
β¨ Why PolyWhisper?
| Full fine-tune (per language) | PolyWhisper v9 | |
|---|---|---|
| Storage per language | ~1.5 GB | ~14 MB (100Γ smaller) |
| Backbone | retrained each time | frozen once, shared by all 5 |
| Bengali (bn) FLEURS WER | 181.3 (baseline) | 130.2 (β28.2%) |
| Marathi (mr) FLEURS WER | 474.9 (baseline) | 116.9 (β75.4%) |
| Telugu (te) FLEURS WER | 103.0 (baseline) | 100.1 (β2.8%) |
| Hindi (hi) FLEURS WER | 43.0 (baseline) | 46.3 |
| Tamil (ta) FLEURS WER | 68.2 (baseline) | 70.1 |
| CPU deployment | heavy | ONNX INT8, no GPU needed |
WER = word error rate (lower is better). FLEURS test set, beam=1, punctuation-normalized scoring.
π Benchmarks (FLEURS, beam=1, normalized WER)
| Language | Code | Script | v7 (no augment) | v9 final | Ξ vs v7 |
|---|---|---|---|---|---|
| Hindi | hi |
Devanagari | 43.0 | 46.3 | +7.7% |
| Tamil | ta |
Tamil | 68.2 | 70.1 | +2.8% |
| Telugu | te |
Telugu | 103.0 | 100.1 | β β2.8% |
| Bengali | bn |
Bengali | 181.3 | 130.2 | β β28.2% |
| Marathi | mr |
Devanagari | 474.9 | 116.9 | β β75.4% |
π§ͺ The v9 finding: augment per language, not globally
Training with SpecAugment + speed perturbation on all languages damaged Hindi/Tamil (token-loop degeneration) while massively helping Bengali/Marathi. The v9 recipe augments only bn/mr and trains hi/ta/te clean:
| Language | Augmentation | Result |
|---|---|---|
| Hindi, Tamil, Telugu | none (clean) | avoids global-augment damage; stays near the no-augment baseline |
| Bengali, Marathi | SpecAugment + 0.9Γ/1.1Γ speed perturb | large gains on hard languages |
π― Decoding: per-language beam widths (measured, full FLEURS test)
Beam-5 + repetition penalty 1.3 helps every language except Telugu, where beam search collapses into repeated-token loops (0/472 perfect samples, 326/472 over 100% WER). The library/CLI defaults encode this (num_beams=None β per-language optimal):
| Language | beam-1 | beam-5 + rep 1.3 | Shipped default |
|---|---|---|---|
| Hindi | 46.3 | 45.0 (β2.8%) | beam-5 |
| Tamil | 70.1 | 68.6 (β2.2%) | beam-5 |
| Telugu | 100.1 | 120.5 (+20.4% β οΈ) | beam-1 |
| Bengali | 130.2 | 126.4 (β2.9%) | beam-5 |
| Marathi | 116.9 | 91.5 (β21.7%) | beam-5 |
π¦ Which adapter should I use?
| Language | Adapter file | Backbone | WER |
|---|---|---|---|
Hindi (hi) |
polywhisper_output_hi/adapters_v3/hi_best_clean.pt |
openai/whisper-small |
46.3 |
Tamil (ta) |
polywhisper_output_ta/adapters_v3/ta_best_clean.pt |
openai/whisper-small |
70.1 |
Telugu (te) |
polywhisper_output_gpu0/adapters_v3/te_best_prod.pt |
openai/whisper-small |
100.1 |
Bengali (bn) |
polywhisper_output_gpu0/adapters_v3/bn_best_prod.pt |
openai/whisper-small |
130.2 |
Marathi (mr) |
polywhisper_output_gpu1/adapters_v3/mr_best_prod.pt |
openai/whisper-small |
116.9 |
All adapters are rank-16 LoRA (decoder + encoder attention), ~14MB each. Backbone weights are not included β they load from openai/whisper-small at runtime. The _prod suffix is the v9 production-run tag, not an augmentation marker: Telugu was trained clean in the selective v9 recipe.
π Quickstart
pip install -e .
# Hindi speech to text
polywhisper transcribe audio.wav --lang hi
# Tamil with JSON output
polywhisper transcribe audio.wav --lang ta --format json
# Auto-detect language, SRT subtitles
polywhisper transcribe audio.wav --format srt > subs.srt
# Batch a folder
polywhisper batch ./audio_folder/ --lang bn --output results.json
from polywhisper import transcribe
result = transcribe("audio.wav", lang="mr")
print(result.text)
print(result.segments) # timestamped segments
π₯οΈ CPU-only inference (ONNX Runtime)
Export INT8-quantized ONNX graphs (no PyTorch, no GPU needed at inference):
polywhisper export --lang hi --variant prod --int8
Pre-exported v9 graphs live under export/onnx/ on the Hub β per language, fp32 + INT8:
| Lang | Encoder (fp32 / INT8) | Decoder (fp32 / INT8) |
|---|---|---|
| hi | 358MB / 97MB | 784MB / 204MB |
| ta | 358MB / 97MB | 784MB / 204MB |
| te | 358MB / 97MB | 784MB / 204MB |
| bn | 358MB / 97MB | 784MB / 204MB |
| mr | 358MB / 97MB | 784MB / 204MB |
Files are named {lang}_{lang}_best_prod_{encoder,decoder}{,_int8}.onnx. INT8 is ~4Γ smaller.
Verification: fp32 ONNX vs PyTorch max diff < 1e-3 on all five languages (encoder + decoder). End-to-end greedy spot-checks (FLEURS audio, beam=1):
| Lang | torch WER | ONNX INT8 WER |
|---|---|---|
| hi (10 samples) | 43.4% | 48.3% |
| ta (5 samples) | 100.0% | 100.0% |
| te (5 samples) | 100.0% | 101.6% |
| bn (5 samples) | 104.9% | 118.7% |
| mr (5 samples) | 82.9% | 89.4% |
Spot-checks are tiny (5β10 utterances) so single-sentence flips move the numbers; fp32 ONNX is at parity with torch. INT8 trades a few points for 4Γ smaller files.
ποΈ Training recipe (reproducible)
- Data: IndicVoices-ST (~19β20k clips/language) Β· Eval: FLEURS
- Backbone:
openai/whisper-small, frozen Β· Adapters: LoRA rank-16, encoder + decoder attention - Schedule: 3β5 epochs/language, batch 4, AdamW, cosine LR (peak 1e-4), 2Γ NVIDIA T4
- Augmentation (v9): SpecAugment + speed perturb for
bn/mronly;hi/ta/teclean - Selection: WER-gated checkpoints (
*_best_*.pt) on FLEURS dev slices - Code:
train_v3.pyΒ· orchestratorkaggle_train_resumable.pyΒ· scoringnormalize_ortho.py
β FAQ
What is PolyWhisper? PolyWhisper is an open-source Indic ASR toolkit: one frozen Whisper-Small backbone plus five small per-language LoRA adapters covering Hindi, Tamil, Telugu, Bengali, and Marathi.
How is it different from fine-tuning Whisper?
Full fine-tuning rewrites 244Mβ1.5B weights per language. PolyWhisper freezes the backbone and trains ~3.5M LoRA parameters per language (14MB), so five languages ship for the storage cost of a rounding error.
Which languages are usable? All five ship working adapters. Hindi (46.3 WER) and Tamil (70.1) are strongest; Telugu, Bengali, and Marathi remain high-WER research adapters, useful for assistive/search/subtitle-draft workflows rather than verbatim transcription.
Can I run it on CPU? Yes β export to ONNX INT8 and run with ONNX Runtime, no GPU required.
Can I run it on a Mac?
Yes β PyTorch MPS is supported (Device: mps), plus CPU via ONNX.
What data was it trained/evaluated on? Trained on IndicVoices-ST conversational speech, evaluated on FLEURS read speech with punctuation-normalized, script-aware scoring.
β οΈ Limitations
- Absolute WER on Telugu/Bengali/Marathi is still high β usable for assistive/search/subtitle-draft workflows, not verbatim legal/medical transcription.
- Evaluated on read speech (FLEURS); spontaneous conversational accuracy will differ.
- Beam=1 numbers in the benchmark table above (paper parity); shipped defaults use beam-5 + repetition penalty 1.3 except Telugu (beam-1), see decoding table.
π License & citation
MIT. Whisper weights Β© OpenAI. Training data: IndicVoices-ST (CC-BY) Β· Eval: FLEURS (CC-BY).
@misc{polywhisper2026,
title = {PolyWhisper: Efficient Multilingual Indic ASR via Frozen Backbone and Per-Language LoRA Adapters},
author = {Kishore, Gautam},
year = {2026},
publisher = {HuggingFace},
url = {https://huggingface.co/eulogik/polywhisper}
}
π Links
- π Eulogik: eulogik.com
- π€ Model: huggingface.co/eulogik/polywhisper
- π» Code: github.com/eulogik/PolyWhisper
- π£οΈ Train data: ai4bharat/indicvoices-st
- π§ͺ Eval data: google/fleurs




