corrected measurement
Voice AI leaderboard reviewed subset
Provider columns return only after a complete corrected run and digest-bound adversarial review. Paid-tier and incomplete providers remain withdrawn.
speech to text
Corrected corpus accuracy and protocol-separated latency
Whisper normalization is applied symmetrically. Accuracy is corpus edit rate. Latency is serial request-to-completion wall clock from five repeats per clip in a pinned runner region.
| Provider | Protocol | Median | IQR | de WER | en WER | es WER | fr WER | hi WER | ja CER |
|---|---|---|---|---|---|---|---|---|---|
| smallest-stt | sync_batch | 537.1 ms | 876.1 ms | 6.7% | 3.8% | 1.3% | 8.1% | 3.9% | 16.8% |
| deepgram | sync_batch | 628.8 ms | 339.7 ms | 7.1% | 3.4% | 3.0% | 7.3% | 14.4% | 6.1% |
text to speech
Corrected TTS timing
Streaming reports first non-empty audio byte. Batch providers report full-synthesis wall clock and are not presented as conversational latency.
| Provider | Protocol | Metric | Median | IQR | Runs |
|---|---|---|---|---|---|
| deepgram-aura | streaming_http | time to first non-empty audio byte | 316.0 ms | 84.0 ms | 75 |
| rime | batch | full-synthesis wall clock (batch) | 834.8 ms | 1629.2 ms | 90 |
| smallest-tts | batch | full-synthesis wall clock (batch) | 1175.6 ms | 585.8 ms | 90 |