Benchmarks

Converse-STT Benchmark

In collaboration withCekura

The Conversational Speech-to-Text Benchmark (Converse-STT) measures how frontier AI models can accurately transcribe real-world, two-person American English conversations.

15

Models

2

Datasets

12

Higher conversational error

2.93%

Lowest word error rate

Model

Score

  1. 01

    Reson8 (Reson8)

    2.93%

    Pipecat 2.10% · rank 3 → 1

  2. 02

    Cartesia Ink 2 (Cartesia)

    3.09%

    Pipecat 2.39% · rank 5 → 2

  3. 03

    Smallest Pulse (Smallest)

    3.46%

    Pipecat 3.66% · rank 11 → 3

  4. 04

    GPT Realtime Whisper (OpenAI)

    3.53%

    Pipecat 2.16% · rank 4 → 4

  5. 05

    Google Chirp 2 (Google)

    3.80%

    Pipecat 2.59% · rank 7 → 5

  6. 06

    Speechmatics Linden (Speechmatics)

    3.93%

    Pipecat 2.46% · rank 6 → 6

  7. 07

    Deepgram Flux English (Deepgram)

    4.18%

    Pipecat 3.90% · rank 13 → 7

  8. 08

    Google Chirp 3 (Google)

    4.19%

    Pipecat 2.00% · rank 2 → 8

  9. 09

    Pipecat 1.93% · rank 1 → 9

  10. 10

    Deepgram Nova-3 (Deepgram)

    4.39%

    Pipecat 3.46% · rank 10 → 10

  11. 11

    Pipecat 3.35% · rank 9 → 11

  12. 12

    Gradium (Gradium)

    4.45%

    Pipecat 6.48% · rank 15 → 12

  13. 13

    Pipecat 4.96% · rank 14 → 13

  14. 14

    Inworld STT-1 (Inworld)

    10.60%

    Pipecat 2.72% · rank 8 → 14

  15. 15

    GPT-4o Transcribe (OpenAI)

    12.45%

    Pipecat 3.90% · rank 12 → 15

Key Takeaways

  • 01

    Frontier AI models struggle to transcribe real-world, unscripted data. Accuracy that holds on public clips often drops once the speech is a real two-person conversation.

  • 02

    Twelve of 15 models had a higher word error rate on conversation. On Ocular’s two-person recordings, 12 models scored worse than on public Pipecat audio, and four more than doubled.

  • 03

    The leading model changes with the dataset. AssemblyAI Universal 3.5 Pro leads Pipecat at 1.93% and ranks ninth on the conversational recordings at 4.23%. Reson8 moves from third at 2.10% to first at 2.93%.

  • 04

    Three models were more accurate on conversation. Smallest Pulse, Deepgram Flux Multilingual, and Gradium each had a lower word error rate on the Hi-Fi recordings than on Pipecat.

  • 05

    The recordings are real two-person conversation. Four American English conversations, about 60 minutes across eight speaker tracks, recorded at 48 kHz as 24-bit mono FLAC with verbatim, time-aligned transcripts.

  • 06

    Both scores belong on the leaderboard. Breath and laughter tags are removed before scoring, and each model still carries its Pipecat result next to its conversational result.

Frequently Asked Questions

The Conversational Speech-to-Text Benchmark (Converse-STT) measures how frontier AI models transcribe real-world, two-person American English conversations. In collaboration with Cekura, 15 speech-to-text models were scored on Ocular AI’s Hi-Fi recordings and on public Pipecat audio. A lower word error rate is a more accurate transcript.

On the Ocular AI Hi-Fi dataset, Reson8 leads at 2.93%. On public Pipecat audio, AssemblyAI Universal 3.5 Pro leads at 1.93% and ranks ninth on the Hi-Fi conversations, at 4.23%. Twelve of the 15 models had a higher error rate on the conversations, and four more than doubled.

Ocular AI supplied four American English conversations of about 15 minutes each, roughly 60 minutes across eight speaker tracks. They were recorded at 48 kHz on separate tracks, with time-aligned verbatim transcripts, and delivered as 24-bit mono FLAC. The recordings keep the pauses, fillers, repetitions, and false starts of ordinary talk.

The score is word error rate: incorrect, missing, and extra words divided by the number of words in the reference. Breath and laughter tags are removed before scoring, so a model is not penalized for leaving them out. Under each bar, the other dataset’s score is shown, and the ranks run from that model’s place on Pipecat to its place on the Hi-Fi conversations.

Pipecat is a public set of short voice-agent clips. Scoring the same models there shows whether a result on public audio carries over to a two-person conversation. The list above can be switched between the two datasets.

A public sample of the Hi-Fi conversations is on Hugging Face. It includes recordings and reference transcripts. It is a sample of a larger collection and does not include the score files for this table.

Ocular AI licenses Hi-Fi conversational speech. Contact us to talk about a collection.

Ready to bring AI into the real world?
Converse-STT Benchmark | Ocular AI