How the experiment was set up, how it was scored, and what was left out.
Not audio quality, not naturalness. If the text says 05/20/2023, does the audio say "May twentieth, twenty twenty-three"? For a voice agent reading back a booking, that one behaviour decides whether the call works.
27 types, 40 each. A generator produces each sentence and its correct spoken form from the same value, so the right answer is correct by construction. Every sentence passed four checks before entering the set, and the whole set regenerates from a seed.
Each sentence went through seven open models as hosted on Replicate. Each clip went through two speech recognizers: OpenAI whisper-1 as the scorer and Deepgram nova-3 as a cross-check. Both sides were converted into one comparable form, so "May 20th" in a transcript doesn't count against a model that said "May twentieth".
There isn't one fixed wording a model has to match. The scorer compares possible readings rather than exact text, and accepts any valid one:
So a date is only marked wrong when what the model said isn't a valid reading at all, not when it picked a different valid one. The scoring code and its tests are in the repo.
15,120 records. 47 were lost to recognizer rate limits and 22 were screened out as empty or invented transcripts: 69 in total, 0.5%, excluded rather than guessed. The two recognizers ranked the seven models identically, but disagreed on 11% of individual clips (832 of 7,497). Only a human ear settles those, so we publish the disagreement set for you to check.
The data, code and results are published on GitHub, along with a free script that recalculates the numbers.