methodology

Measurement correction log

Public numbers are claims. When the method fails review, the numbers come down first.

15 September 2026

Accuracy and latency rankings withdrawn

Defect 1 - text normalization. The 13-14 September STT board compared lowercase whitespace tokens without the Whisper paper normalizer. FLEURS references are bare lowercase text while providers often return casing, punctuation, contractions or formatted numbers. Content-equivalent output could be penalized. The published overall WER spread was 9.4%-17.2%; those values and derived scores are withdrawn, not silently replaced.

Defect 2 - aggregation. The board averaged per-clip WER, over-weighting short utterances. The rebuild uses corpus WER: total token edits divided by total reference tokens, after identical normalization of reference and hypothesis.

Defect 3 - latency instrumentation. Four async adapters polled at fixed four-second intervals while calls ran under six-way concurrency on an unspecified GitHub runner. Published STT p50s included polling and contention artifacts: between two runs, Groq moved 418ms to 755ms and Deepgram 460ms to 546ms while their WERs were unchanged. All STT and LLM latency rankings are withdrawn.

Defect 4 - TTS metric. Full synthesis wall clock was labeled as conversational latency. Streaming TTS must report time to first audio byte. Batch-only providers must be labeled "full-synthesis wall clock (batch)" and cannot imply conversational responsiveness. All TTS latency rankings are withdrawn pending this split.

Rebuild gate. Serial execution; adaptive polling from 250ms with backoff; pinned region recorded in JSON; at least five repeats per clip; median and interquartile range; sync and async methods separated; TTFB for streaming TTS; and an explicit adversarial review whose job is to break the measurement before publication.

corpus

Inputs and receipts

The source corpus remains Google FLEURS test audio (CC-BY-4.0), 15 clips each in English, Spanish, Hindi, French, German and Japanese. Historical raw results remain in the repository for audit but are not endorsed rankings.