We gave seven open text-to-speech models 1,080 sentences a voice agent might read back, then checked what they actually said.
Dograh (2026). TTS verbalization benchmark: do open text-to-speech models say dates, prices and phone numbers correctly? Dograh Experiments. https://www.dograh.com/hub/experiments/tts-verbalization-benchmark
@misc{dograh2026tts,
author = {{Dograh}},
title = {TTS verbalization benchmark: do open text-to-speech models say dates, prices and phone numbers correctly?},
year = {2026},
howpublished = {Dograh Experiments},
url = {https://www.dograh.com/hub/experiments/tts-verbalization-benchmark},
note = {Data and code: https://github.com/dograh-hq/experiments/tree/main/benchmarks/tts-verbalization}
}This experiment tested whether seven open text-to-speech models say tricky text correctly: dates, prices, phone numbers, hashtags and 23 other types. Each model read the same 1,080 sentences, and two independent speech recognizers wrote down what it actually said.
Each industry needs different kinds of text read correctly. Pick yours to rank the models on just those types.
Switch to Table to see each model’s likely range and cost.
The four hardest types of text. Each number is the best score any of the seven models reached, so every model did this badly or worse.
One model can be perfect at one type and score zero on another, so choose by the types your agent reads, not the average.
| Models → | |||||||
|---|---|---|---|---|---|---|---|
| Type of text ↓ | kokoro-82m | orpheus-3b | qwen3-tts | chatterbox-turbo | chatterbox | chatterbox-multilingual | parler-tts |
| Initialism or Acronym | |||||||
| Musical Notation | |||||||
| Stock Ticker | |||||||
| Currency | |||||||
| Ordinal | |||||||
| Vehicle or Product Code | |||||||
| Decimal | |||||||
| Sports score | |||||||
| Roman Numeral | |||||||
| Abbreviation | |||||||
| Cardinal | |||||||
| Legal Reference | |||||||
| License Plate or Serial Numbers | |||||||
| Time | |||||||
| Unit | |||||||
| Mathematical Expression | |||||||
| Geographic Coordinates | |||||||
| Chemical Formula | |||||||
| Phone Number | |||||||
| Address | |||||||
| ISBN | |||||||
| URL or Email | |||||||
| Biological Classification | |||||||
| Fractions | |||||||
| Version Numbers | |||||||
| Date | |||||||
| Hashtag or Mention | |||||||
Her birthday falls on 2009-02-26 this year.
"her birthday falls on two thousand nine zero two twenty six this year"
kokoro-82m is the most accurate and the cheapest to run: about 18× less compute per clip than orpheus-3b, the runner-up.
The spread inside one model is wider than the spread between models: kokoro-82m scores 1.00 on musical notation and 0.00 on hashtags in the same run. So look at the type your product actually reads.
Top 3 of 7 on ID & serial numbers
Top 3 of 7 on Dates
Want a different type? Use the industry picker or the heatmap above.
Pick a type of text to see a real sentence from the test, how it should be read, and what each of the seven models said.
Real sentences from the run. What each model said is shown as the listener heard it, with numbers written out as words. Both listeners agreed on every verdict shown.
Here's one real sentence from the test, followed all the way through.
A single speech recognizer deciding what a model "said" is a weak foundation, so each clip was sent to two from different companies.
Four results that need a closer look before you rely on them.
The models split "#GreenBreak" into "Green Break" correctly, but none of them ever says the word "hashtag", and the test's correct reading includes it. That's a judgement call: plenty of people don't say "hashtag" out loud. If your product doesn't, read the other 26 types.
Its speech is garbled overall, not just on the hard part. Read its score as the floor that proves the test can tell good from bad.
Example: "Her birthday falls on 2009-02-26 this year." came out as "Herbal tea flowers enswore so, very gee. Thok!"
Deepgram nova-3 "repairs" some mistakes on its own, for example turning "mph" into "miles per hour", which would turn a model's failure into a pass. So it's used as a cross-check, and OpenAI whisper-1 is the scorer.
It is not general accuracy: every sentence is chosen to be hard. Models were tested as hosted on Replicate, in English only. Compute figures come from a 10-clip timing run per model, and 69 of 15,120 records (0.5%) never produced a score. Assume the speech recognizers (ASR) have an error rate of about 5%, so a small share of verdicts may reflect the listener mishearing rather than the model.
The test sentences, the code and the full results are published on GitHub, along with a free script that recalculates the numbers.
Dograh (2026). TTS verbalization benchmark: do open text-to-speech models say dates, prices and phone numbers correctly? Dograh Experiments. https://www.dograh.com/hub/experiments/tts-verbalization-benchmark
@misc{dograh2026tts,
author = {{Dograh}},
title = {TTS verbalization benchmark: do open text-to-speech models say dates, prices and phone numbers correctly?},
year = {2026},
howpublished = {Dograh Experiments},
url = {https://www.dograh.com/hub/experiments/tts-verbalization-benchmark},
note = {Data and code: https://github.com/dograh-hq/experiments/tree/main/benchmarks/tts-verbalization}
}