corrected measurement

Voice AI leaderboard reviewed subset

Provider columns return only after a complete corrected run and digest-bound adversarial review. Paid-tier and incomplete providers remain withdrawn.

speech to text

Corrected corpus accuracy and protocol-separated latency

Whisper normalization is applied symmetrically. Accuracy is corpus edit rate. Latency is serial request-to-completion wall clock from five repeats per clip in a pinned runner region.

ProviderProtocolMedianIQRde WERen WERes WERfr WERhi WERja CER
smallest-sttsync_batch537.1 ms876.1 ms6.7%3.8%1.3%8.1%3.9%16.8%
deepgramsync_batch628.8 ms339.7 ms7.1%3.4%3.0%7.3%14.4%6.1%

text to speech

Corrected TTS timing

Streaming reports first non-empty audio byte. Batch providers report full-synthesis wall clock and are not presented as conversational latency.

ProviderProtocolMetricMedianIQRRuns
deepgram-aurastreaming_httptime to first non-empty audio byte316.0 ms84.0 ms75
rimebatchfull-synthesis wall clock (batch)834.8 ms1629.2 ms90
smallest-ttsbatchfull-synthesis wall clock (batch)1175.6 ms585.8 ms90