Word error rate · lower is better
2.93%
Reson8 (Reson8)
3.09%Cartesia Ink 2 (Cartesia)
3.46%
Smallest Pulse (Smallest)
3.53%GPT Realtime Whisper (OpenAI)
3.80%
Google Chirp 2 (Google)
AudioIn collaboration with
Converse-STT
Conversational Speech-to-Text Benchmark
In collaboration with Cekura, we scored 15 speech-to-text models on Ocular AI Real World Conversational Data and on public Pipecat audio. Twelve had a higher error rate on the conversational recordings, and the leading model changed with the dataset.
Category
Speech-to-textLanguages
English