Automatic Speech Recognition
Transformers
ONNX
PEFT
whisper
polywhisper
indic-asr
hindi-asr
tamil-speech-recognition
telugu-stt
bengali-asr
marathi-speech-to-text
speech-recognition
multilingual
lora
hindi
tamil
telugu
bengali
marathi
indic-languages
indian-languages
speech-to-text
low-resource-asr
fleurs
indicvoices
quantized
efficient-asr
edge-asr
Eval Results (legacy)
Instructions to use eulogik/polywhisper with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use eulogik/polywhisper with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="eulogik/polywhisper")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForSpeechSeq2Seq processor = AutoProcessor.from_pretrained("eulogik/polywhisper") model = AutoModelForSpeechSeq2Seq.from_pretrained("eulogik/polywhisper", device_map="auto") - PEFT
How to use eulogik/polywhisper with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
docs: per-language optimal decoding table (beam-5 except te=1)
Browse files
README.md
CHANGED
|
@@ -145,6 +145,18 @@ Training with SpecAugment + speed perturbation on **all** languages damaged Hind
|
|
| 145 |
| Hindi, Tamil | none (clean) | matches no-augment baseline |
|
| 146 |
| Telugu, Bengali, Marathi | SpecAugment + 0.9Γ/1.1Γ speed perturb | large gains on hard languages |
|
| 147 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 148 |
## π¦ Which adapter should I use?
|
| 149 |
|
| 150 |
| Language | Adapter file | Backbone | WER |
|
|
@@ -250,7 +262,7 @@ Trained on IndicVoices-ST conversational speech, evaluated on FLEURS read speech
|
|
| 250 |
|
| 251 |
- Absolute WER on Telugu/Bengali/Marathi is still high β usable for assistive/search/subtitle-draft workflows, not verbatim legal/medical transcription.
|
| 252 |
- Evaluated on read speech (FLEURS); spontaneous conversational accuracy will differ.
|
| 253 |
-
- Beam=1 numbers above; beam
|
| 254 |
|
| 255 |
## π License & citation
|
| 256 |
|
|
|
|
| 145 |
| Hindi, Tamil | none (clean) | matches no-augment baseline |
|
| 146 |
| Telugu, Bengali, Marathi | SpecAugment + 0.9Γ/1.1Γ speed perturb | large gains on hard languages |
|
| 147 |
|
| 148 |
+
### π― Decoding: per-language beam widths (measured, full FLEURS test)
|
| 149 |
+
|
| 150 |
+
Beam-5 + repetition penalty 1.3 helps every language **except Telugu**, where beam search collapses into repeated-token loops (0/472 perfect samples, 326/472 over 100% WER). The library/CLI defaults encode this (`num_beams=None` β per-language optimal):
|
| 151 |
+
|
| 152 |
+
| Language | beam-1 | beam-5 + rep 1.3 | Shipped default |
|
| 153 |
+
|---|---|---|---|
|
| 154 |
+
| Hindi | 46.3 | **45.0** (β2.8%) | beam-5 |
|
| 155 |
+
| Tamil | 70.1 | **68.6** (β2.2%) | beam-5 |
|
| 156 |
+
| Telugu | **100.1** | 120.5 (+20.4% β οΈ) | **beam-1** |
|
| 157 |
+
| Bengali | 130.2 | **126.4** (β2.9%) | beam-5 |
|
| 158 |
+
| Marathi | 96.7 | **91.5** (β5.4%) | beam-5 |
|
| 159 |
+
|
| 160 |
## π¦ Which adapter should I use?
|
| 161 |
|
| 162 |
| Language | Adapter file | Backbone | WER |
|
|
|
|
| 262 |
|
| 263 |
- Absolute WER on Telugu/Bengali/Marathi is still high β usable for assistive/search/subtitle-draft workflows, not verbatim legal/medical transcription.
|
| 264 |
- Evaluated on read speech (FLEURS); spontaneous conversational accuracy will differ.
|
| 265 |
+
- Beam=1 numbers in the benchmark table above (paper parity); shipped defaults use beam-5 + repetition penalty 1.3 except Telugu (beam-1), see decoding table.
|
| 266 |
|
| 267 |
## π License & citation
|
| 268 |
|