Benchmarks
Converse-STT Benchmark
In collaboration withThe Conversational Speech-to-Text Benchmark (Converse-STT) measures how frontier AI models can accurately transcribe real-world, two-person American English conversations.
Models
Datasets
Higher conversational error
Lowest word error rate
Model
Score
- 01
- 02
- 03
- 04
- 05
- 06
- 07
- 08
- 09
- 10
- 11
- 12
- 13
- 14
- 15
Key Takeaways
Frequently Asked Questions
The Conversational Speech-to-Text Benchmark (Converse-STT) measures how frontier AI models transcribe real-world, two-person American English conversations. In collaboration with Cekura, 15 speech-to-text models were scored on Ocular AI’s Hi-Fi recordings and on public Pipecat audio. A lower word error rate is a more accurate transcript.
On the Ocular AI Hi-Fi dataset, Reson8 leads at 2.93%. On public Pipecat audio, AssemblyAI Universal 3.5 Pro leads at 1.93% and ranks ninth on the Hi-Fi conversations, at 4.23%. Twelve of the 15 models had a higher error rate on the conversations, and four more than doubled.
Ocular AI supplied four American English conversations of about 15 minutes each, roughly 60 minutes across eight speaker tracks. They were recorded at 48 kHz on separate tracks, with time-aligned verbatim transcripts, and delivered as 24-bit mono FLAC. The recordings keep the pauses, fillers, repetitions, and false starts of ordinary talk.
The score is word error rate: incorrect, missing, and extra words divided by the number of words in the reference. Breath and laughter tags are removed before scoring, so a model is not penalized for leaving them out. Under each bar, the other dataset’s score is shown, and the ranks run from that model’s place on Pipecat to its place on the Hi-Fi conversations.
Pipecat is a public set of short voice-agent clips. Scoring the same models there shows whether a result on public audio carries over to a two-person conversation. The list above can be switched between the two datasets.
A public sample of the Hi-Fi conversations is on Hugging Face. It includes recordings and reference transcripts. It is a sample of a larger collection and does not include the score files for this table.
Ocular AI licenses Hi-Fi conversational speech. Contact us to talk about a collection.