Accuracy tiers
Transcription accuracy depends on how much training data exists for each language. We classify languages into three tiers:Word Error Rate (WER) measures the percentage of words transcribed incorrectly. A 5% WER means roughly 1 in 20 words may be wrong. These are approximate ranges — actual accuracy depends on audio quality, speaker clarity, background noise, and domain-specific vocabulary.
Tips for better accuracy
- Provide a
languagehint — Auto-detection works well for high-tier languages but can misidentify lower-resource languages. Passing thelanguageparameter eliminates detection errors entirely. - Clean audio matters — Background noise, music, and overlapping speakers reduce accuracy across all languages.
- Medium and low-tier languages benefit the most from a language hint. Without one, the detector may fall back to a related high-resource language.
Complete language list
High accuracy (under 8% WER)
These languages have extensive training data and deliver production-quality transcription.Medium accuracy (8–15% WER)
These languages produce good results suitable for most workflows. A language hint is recommended.Low accuracy (15–40%+ WER)
These languages have limited training data. Transcription captures the general content but expect frequent errors. Always provide alanguage hint for best results.
Language detection
When you omit thelanguage parameter, FastDrop automatically detects the spoken language from the audio. Detection accuracy follows the same tier pattern:
- High-tier languages are detected reliably (95%+ accuracy)
- Medium-tier languages are detected correctly most of the time (85–95%)
- Low-tier languages may be misidentified as a related higher-resource language (60–85%)
language parameter to bypass detection entirely.
Translation coverage
Translation to English is available for all 99 languages. Whentranslate is true (the default) and the source language is not English, FastDrop generates a full English translation alongside the source transcription.
Translation quality is generally high across all languages, as it operates on the transcribed text rather than raw audio. Even for low-accuracy transcription languages, the translation of whatever text is captured tends to be reliable.