GautamKishore commited on
Commit
19afd17
Β·
verified Β·
1 Parent(s): c84f131

docs: per-language optimal decoding table (beam-5 except te=1)

Browse files
Files changed (1) hide show
  1. README.md +13 -1
README.md CHANGED
@@ -145,6 +145,18 @@ Training with SpecAugment + speed perturbation on **all** languages damaged Hind
145
  | Hindi, Tamil | none (clean) | matches no-augment baseline |
146
  | Telugu, Bengali, Marathi | SpecAugment + 0.9Γ—/1.1Γ— speed perturb | large gains on hard languages |
147
 
 
 
 
 
 
 
 
 
 
 
 
 
148
  ## πŸ“¦ Which adapter should I use?
149
 
150
  | Language | Adapter file | Backbone | WER |
@@ -250,7 +262,7 @@ Trained on IndicVoices-ST conversational speech, evaluated on FLEURS read speech
250
 
251
  - Absolute WER on Telugu/Bengali/Marathi is still high β€” usable for assistive/search/subtitle-draft workflows, not verbatim legal/medical transcription.
252
  - Evaluated on read speech (FLEURS); spontaneous conversational accuracy will differ.
253
- - Beam=1 numbers above; beam=5 decoding improves results at higher latency.
254
 
255
  ## πŸ“„ License & citation
256
 
 
145
  | Hindi, Tamil | none (clean) | matches no-augment baseline |
146
  | Telugu, Bengali, Marathi | SpecAugment + 0.9Γ—/1.1Γ— speed perturb | large gains on hard languages |
147
 
148
+ ### 🎯 Decoding: per-language beam widths (measured, full FLEURS test)
149
+
150
+ Beam-5 + repetition penalty 1.3 helps every language **except Telugu**, where beam search collapses into repeated-token loops (0/472 perfect samples, 326/472 over 100% WER). The library/CLI defaults encode this (`num_beams=None` β†’ per-language optimal):
151
+
152
+ | Language | beam-1 | beam-5 + rep 1.3 | Shipped default |
153
+ |---|---|---|---|
154
+ | Hindi | 46.3 | **45.0** (βˆ’2.8%) | beam-5 |
155
+ | Tamil | 70.1 | **68.6** (βˆ’2.2%) | beam-5 |
156
+ | Telugu | **100.1** | 120.5 (+20.4% ⚠️) | **beam-1** |
157
+ | Bengali | 130.2 | **126.4** (βˆ’2.9%) | beam-5 |
158
+ | Marathi | 96.7 | **91.5** (βˆ’5.4%) | beam-5 |
159
+
160
  ## πŸ“¦ Which adapter should I use?
161
 
162
  | Language | Adapter file | Backbone | WER |
 
262
 
263
  - Absolute WER on Telugu/Bengali/Marathi is still high β€” usable for assistive/search/subtitle-draft workflows, not verbatim legal/medical transcription.
264
  - Evaluated on read speech (FLEURS); spontaneous conversational accuracy will differ.
265
+ - Beam=1 numbers in the benchmark table above (paper parity); shipped defaults use beam-5 + repetition penalty 1.3 except Telugu (beam-1), see decoding table.
266
 
267
  ## πŸ“„ License & citation
268