Dograh
Methodology · EXP-001

How we measured it

How the experiment was set up, how it was scored, and what was left out.

01 · Question

Does the model say the right words?

Not audio quality, not naturalness. If the text says 05/20/2023, does the audio say "May twentieth, twenty twenty-three"? For a voice agent reading back a booking, that one behaviour decides whether the call works.

02 · Data

1,080 sentences, generated rather than written

27 types, 40 each. A generator produces each sentence and its correct spoken form from the same value, so the right answer is correct by construction. Every sentence passed four checks before entering the set, and the whole set regenerates from a seed.

03 · Method

Seven models, two listeners

Each sentence went through seven open models as hosted on Replicate. Each clip went through two speech recognizers: OpenAI whisper-1 as the scorer and Deepgram nova-3 as a cross-check. Both sides were converted into one comparable form, so "May 20th" in a transcript doesn't count against a model that said "May twentieth".

04 · Scoring

What counts as correct

There isn't one fixed wording a model has to match. The scorer compares possible readings rather than exact text, and accepts any valid one:

  • Numbers said as words, following normal English number grammar, so "one thousand two hundred thirty four" counts.
  • Symbols that can be read more than one way. A dash can be "dash" or "hyphen", left silent in a serial number, "to" in a sports score, or "minus" in a sum.
  • Abbreviations either spelled out or expanded, so "kg" and "kilograms" both count.
  • Roman numerals, fractions and ordinals in the ways people normally say them.
  • A different word order within the tricky part, so "July third" and "third July" both count.

So a date is only marked wrong when what the model said isn't a valid reading at all, not when it picked a different valid one. The scoring code and its tests are in the repo.

05 · Ledger

What we threw away

15,120 records. 47 were lost to recognizer rate limits and 22 were screened out as empty or invented transcripts: 69 in total, 0.5%, excluded rather than guessed. The two recognizers ranked the seven models identically, but disagreed on 11% of individual clips (832 of 7,497). Only a human ear settles those, so we publish the disagreement set for you to check.

06 · Limits

What this does not tell you

  • It is not general accuracy: every sentence is chosen to be hard.
  • Models as hosted on Replicate, not bare checkpoints.
  • English only.
  • Saying "hashtag" aloud is our convention, not a universal one.
  • Compute figures come from 10 clips per model.
  • Assume the speech recognizers (ASR) have an error rate of about 5%, so a small share of verdicts may reflect the listener mishearing rather than the model.
07 · Data

Data and code

The data, code and results are published on GitHub, along with a free script that recalculates the numbers.