Dograh
Experiments /  TTS verbalization benchmark
EXP-001Benchmark✓ Run Sep 29, 2026 · 78 of 78 checks pass

Nobody's text-to-speech can say a date with confidence

We gave seven open text-to-speech models 1,080 sentences a voice agent might read back, then checked what they actually said.

7models
1,080sentences
27types of tricky text
2independent listeners
Cite as text
Dograh (2026). TTS verbalization benchmark: do open text-to-speech models say dates, prices and phone numbers correctly? Dograh Experiments. https://www.dograh.com/hub/experiments/tts-verbalization-benchmark
BibTeX (for academic papers)
@misc{dograh2026tts,
  author = {{Dograh}},
  title  = {TTS verbalization benchmark: do open text-to-speech models say dates, prices and phone numbers correctly?},
  year   = {2026},
  howpublished = {Dograh Experiments},
  url    = {https://www.dograh.com/hub/experiments/tts-verbalization-benchmark},
  note   = {Data and code: https://github.com/dograh-hq/experiments/tree/main/benchmarks/tts-verbalization}
}
Summary

This experiment tested whether seven open text-to-speech models say tricky text correctly: dates, prices, phone numbers, hashtags and 23 other types. Each model read the same 1,080 sentences, and two independent speech recognizers wrote down what it actually said.

Key findings
The smallest model is the most accurate.kokoro-82m (82 million parameters) scores 0.61, ahead of orpheus-3b (3 billion) at 0.54, and uses roughly 18x less compute.
No model can say a hashtag.0 correct out of 279.
Dates are broken everywhere.The best score any model reached is 0.23: three out of four dates come out wrong.
The ranking is reliable.Two independent listeners rank the models identically, even though they disagree on 11% of individual clips.
What we found

Even the best model gets only six in ten right

Which model fits your industry?

Each industry needs different kinds of text read correctly. Pick yours to rank the models on just those types.

Overall score across all 1,080 sentences and 27 types.

Switch to Table to see each model’s likely range and cost.

Where it breaks

The four hardest types of text. Each number is the best score any of the seven models reached, so every model did this badly or worse.

0.00
Hashtags
No model got a single one right
0.23
Dates
Even the best model: about 2 in 10 right
0.30
Version numbers
Even the best model: about 3 in 10 right
0.40
Fractions
Even the best model: about 4 in 10 right

Every model, every type

One model can be perfect at one type and score zero on another, so choose by the types your agent reads, not the average.

How to read this: each row is a type of text and each column is a model, so each square is one model on one type. The brighter the orange, the more often that model said it right. Dark squares mean it almost never did. Click any square for the details.
Models →
Type of text ↓kokoro-82morpheus-3bqwen3-ttschatterbox-turbochatterboxchatterbox-multilingualparler-tts
Initialism or Acronym
Musical Notation
Stock Ticker
Currency
Ordinal
Vehicle or Product Code
Decimal
Sports score
Roman Numeral
Abbreviation
Cardinal
Legal Reference
License Plate or Serial Numbers
Time
Unit
Mathematical Expression
Geographic Coordinates
Chemical Formula
Phone Number
Address
ISBN
URL or Email
Biological Classification
Fractions
Version Numbers
Date
Hashtag or Mention
Never rightAlways right
Date
Right on about 4 of 40 date sentences
6 / 7rank on this type
-0.51vs its overall 0.61
0.03–0.2095% range
All 7 models on date
chatterbox
0.23
orpheus-3b
0.20
qwen3-tts
0.17
chatterbox-multilingual
0.17
chatterbox-turbo
0.15
kokoro-82m
0.10
parler-tts
0.00
A real miss

Her birthday falls on 2009-02-26 this year.

"her birthday falls on two thousand nine zero two twenty six this year"

Type of textkokoro-82morpheus-3bqwen3-ttschatterbox-turbochatterboxchatterbox-multilingualparler-ttsBest
Initialism or Acronym1.000.890.800.970.970.820.151.00
Musical Notation1.000.400.230.380.400.330.001.00
Stock Ticker1.000.820.920.850.820.850.031.00
Currency0.970.880.620.680.470.030.000.97
Ordinal0.900.970.820.930.880.700.000.97
Vehicle or Product Code0.950.720.510.780.880.280.000.95
Decimal0.900.500.820.620.330.000.000.90
Sports score0.530.780.900.820.450.380.000.90
Roman Numeral0.880.550.310.380.570.150.000.88
Abbreviation0.850.690.590.800.600.620.570.85
Cardinal0.800.820.500.620.420.120.000.82
Legal Reference0.780.640.500.720.600.250.000.78
License Plate or Serial Numbers0.720.550.550.450.780.170.000.78
Time0.720.640.380.450.330.620.000.72
Unit0.620.650.720.530.420.380.000.72
Mathematical Expression0.680.400.700.400.000.000.000.70
Geographic Coordinates0.000.530.670.230.000.000.000.67
Chemical Formula0.620.470.610.450.620.500.000.62
Phone Number0.420.600.410.380.600.030.000.60
Address0.550.240.170.250.170.050.000.55
ISBN0.150.530.260.280.400.000.000.53
URL or Email0.420.400.160.500.150.000.000.50
Biological Classification0.450.460.330.400.450.200.000.46
Fractions0.100.100.400.030.150.030.000.40
Version Numbers0.300.280.230.280.200.200.000.30
Date0.100.200.170.150.230.170.000.23
Hashtag or Mention0.000.000.000.000.000.000.000.00

Share of sentences said correctly (0 = never, 1 = always). Orange = best model for that type.

The smallest model wins on both counts

kokoro-82m is the most accurate and the cheapest to run: about 18× less compute per clip than orpheus-3b, the runner-up.

← Compute per clip (shorter is cheaper)ModelCorrect (longer is better) →
0.31s
kokoro-82m
0.61
5.50s
orpheus-3b
0.54
4.99s
qwen3-tts
0.49
3.48s
chatterbox
0.44
3.06s
parler-tts
0.03
Reading the results

An average across 27 types is a headline, not a decision

The spread inside one model is wider than the spread between models: kokoro-82m scores 1.00 on musical notation and 0.00 on hashtags in the same run. So look at the type your product actually reads.

If your product reads order numbers
look at ID & serial numbers
chatterbox
0.78
kokoro-82m
0.72
orpheus-3b
0.55

Top 3 of 7 on ID & serial numbers

If it books appointments
look at Dates
chatterbox
0.23
orpheus-3b
0.20
qwen3-tts
0.17

Top 3 of 7 on Dates

Want a different type? Use the industry picker or the heatmap above.

Hear it fail

Listen to what each model actually said

Pick a type of text to see a real sentence from the test, how it should be read, and what each of the seven models said.

0 of 7 models said it right
The text
Her birthday falls on 2009-02-26 this year.
Should be said as
Her birthday falls on February twenty sixth two thousand nine this year.
kokoro-82m"her birthday falls on two thousand nine zero two twenty six this year"Wrong
orpheus-3b"her birthday falls on two thousand and nine oh two twenty six this year"Wrong
qwen3-tts"her birthday falls on two thousand and nine dot r six this year"Wrong
chatterbox-turbo"her birthday falls on two thousand and nine two twenty six this year"Wrong
chatterbox"her birthday falls on two thousand nine two two six this year"Wrong
chatterbox-multilingual"her birthday falls on two thousand and nine nif-two twenty six this year"Wrong
parler-tts"her birthday fais and voslo vargy vogue"Wrong

Real sentences from the run. What each model said is shown as the listener heard it, with numbers written out as words. Both listeners agreed on every verdict shown.

How it worked

One sentence, four steps, repeated 7,560 times

Here's one real sentence from the test, followed all the way through.

01 · THE TEXT

We write a tricky sentence

1,080
sentences across 27 types, each created by code together with its correct spoken form
Written
Her birthday falls on 2009-02-26 this year.
Should be said as
Her birthday falls on February twenty sixth two thousand nine this year.
02 · THE VOICE

A model reads it aloud

7,560
audio clips: every sentence read by all 7 open text-to-speech models
reads the sentence aloud
03 · THE LISTENER

A recognizer writes down what it heard

15,120
transcription attempts: each clip sent to 2 speech recognizers from different companies
It heard
"her birthday falls on two thousand nine zero two twenty six this year"
04 · THE VERDICT

We compare it to the right answer

0.61
is the best score: the share of clips a model got right
Needed
"february twenty sixth two thousand nine"
Wrong
Why you can trust it

Two independent listeners. The same ranking.

A single speech recognizer deciding what a model "said" is a weak foundation, so each clip was sent to two from different companies.

Same ranking from both listeners

Every record, accounted for

Records (2 listeners × 7,560 clips)15,120
Scored15,051
Excluded · listener rate limits47
Excluded · empty or invented transcripts22
Clips both listeners scored7,497
Clips the listeners disagreed on832
But the two listeners disagreed on 11% of individual clips (832 of 7,497). Only a human ear settles those, so we publish the disagreement set for you to check.

Behind the numbers

Four results that need a closer look before you rely on them.

Why do hashtags score zero?

The models split "#GreenBreak" into "Green Break" correctly, but none of them ever says the word "hashtag", and the test's correct reading includes it. That's a judgement call: plenty of people don't say "hashtag" out loud. If your product doesn't, read the other 26 types.

Why does parler-tts score 0.03?

Its speech is garbled overall, not just on the hard part. Read its score as the floor that proves the test can tell good from bad.

Example: "Her birthday falls on 2009-02-26 this year." came out as "Herbal tea flowers enswore so, very gee. Thok!"

Why is only one listener the official scorer?

Deepgram nova-3 "repairs" some mistakes on its own, for example turning "mph" into "miles per hour", which would turn a model's failure into a pass. So it's used as a cross-check, and OpenAI whisper-1 is the scorer.

What does this not tell you?

It is not general accuracy: every sentence is chosen to be hard. Models were tested as hosted on Replicate, in English only. Compute figures come from a 10-clip timing run per model, and 69 of 15,120 records (0.5%) never produced a score. Assume the speech recognizers (ASR) have an error rate of about 5%, so a small share of verdicts may reflect the listener mishearing rather than the model.

Data and code

The data and code behind this experiment

The test sentences, the code and the full results are published on GitHub, along with a free script that recalculates the numbers.

Cite as text
Dograh (2026). TTS verbalization benchmark: do open text-to-speech models say dates, prices and phone numbers correctly? Dograh Experiments. https://www.dograh.com/hub/experiments/tts-verbalization-benchmark
BibTeX (for academic papers)
@misc{dograh2026tts,
  author = {{Dograh}},
  title  = {TTS verbalization benchmark: do open text-to-speech models say dates, prices and phone numbers correctly?},
  year   = {2026},
  howpublished = {Dograh Experiments},
  url    = {https://www.dograh.com/hub/experiments/tts-verbalization-benchmark},
  note   = {Data and code: https://github.com/dograh-hq/experiments/tree/main/benchmarks/tts-verbalization}
}