Zero STT delivers state-of-the-art transcription across 216+ languages with native multilingual understanding, streaming performance under 500ms, and the industry's best composite accuracy on OpenASR.
3.10%
Composite WER
#1
OpenASR Leaderboard
216+
Languages Supported
240+
Concurrent Streams/GPU
Real Speech
Most speech recognition models are evaluated on clean datasets. Production audio isn't clean.
Zero STT was built for theseconversations.
Zero STT tops the OpenASR leaderboardat a 3.10% composite WER - a 45% relativereduction versus the best published baseline.
| Benchmark | Zero STT | NVIDIA Canary-Qwen 2.5BCanary 2.5B | Best Published BaselineBaseline | Relative GainGain |
|---|---|---|---|---|
| LibriSpeech Clean | 0.71% | 1.58% | 1.42% | 50% |
| SPGISpeech | 1.10% | 1.90% | 1.90% | 42% |
| TedLium | 1.43% | 2.71% | 2.71% | 47% |
| LibriSpeech Other | 2.17% | 3.12% | 2.87% | 24% |
| AMI | 4.19% | 9.65% | 9.12% | 54% |
| VoxPopuli | 4.34% | 5.63% | 5.63% | 23% |
| GigaSpeech | 4.99% | 9.43% | 9.43% | 47% |
| Earnings22 | 5.83% | 10.02% | 9.53% | 39% |
| Composite | 3.10% | 5.63% | n/a | 45% |
Sub-500ms first token for live conversations.
Separate speakers automatically.
Classify customer intent while transcribing.
Detect satisfaction, frustration, urgency.
Track emotional changes throughout conversations.
Punctuation, capitalization, timestamps, keyword detection, translation, transliteration and profanity masking in one API.