September 21, 2026ReleaseResearchBenchmark

Introducing The Converse-STT Benchmark

The Conversational Speech-to-Text Benchmark (Converse-STT) measures how frontier AI models behave on real-world, conversational data. It tests transcription on two-person, real-world American English conversations and on voice-agent clips from Pipecat’s STT Benchmark dataset.

Louis Murerwa
Louis Murerwa
  1. Reson8 (Reson8)

    2.93%
  2. Cartesia Ink 2 (Cartesia)

    3.09%
  3. Smallest Pulse (Smallest)

    3.46%
  4. GPT Realtime Whisper (OpenAI)

    3.53%
  5. Google Chirp 2 (Google)

    3.80%
  6. Speechmatics Linden (Speechmatics)

    3.93%
  7. Deepgram Flux English (Deepgram)

    4.18%
  8. Google Chirp 3 (Google)

    4.19%
  9. Deepgram Nova-3 (Deepgram)

    4.39%
  10. Gradium (Gradium)

    4.45%
  11. Inworld STT-1 (Inworld)

    10.60%
  12. GPT-4o Transcribe (OpenAI)

    12.45%

We are introducing the Conversational Speech-to-Text Benchmark (Converse-STT), created in collaboration with Cekura. Ocular AI benchmarks measure how models behave on real-world data. Converse-STT applies that test to speech: how accurately frontier models transcribe two-person American English conversations.

Cekura was critical in running the model evaluations on these conversations. They ran the 15 AI model evaluations on a subset of full-duplex, studio-grade conversational data, created by American English experts hired through Ocular AI’s Workbolt platform, realistically testing how AI models perform on real-world human-generated data. The results reflect the transcription accuracy frontier models must reach to be useful on real human conversation. See the Converse-STT leaderboard.

On real-world conversation, 12 models had a higher word error rate than on public Pipecat audio, and four more than doubled. Three had a lower word error rate. The leading model changed once the speech was a real exchange between two people. That shift is what the benchmark is for: a public score, read against behavior on real-world data.

The comparison below shows every model on Ocular AI’s real-world conversational data and on public Pipecat audio. The Pipecat set is 1,000 public voice-agent clips from Pipecat’s STT benchmark dataset. Its reference transcripts are generated with Gemini and described by Pipecat as human-reviewed, so those clips serve as the public comparison. The real-world, human-generated data in this benchmark is the two-person conversations.

Converse-STT Benchmark

80% were less accurate on real-world conversation.

12 of 15 tested models had higher WER on Ocular AI’s real-world, two-person conversations than on public Pipecat audio. Four more than doubled; three had lower WER.

Higher WER on OcularLower WER on Ocular

Ocular ÷ Pipecat WER

3.90×
Pipecat 2.72% → Ocular 10.60%
More than double WER
3.19×
Pipecat 3.90% → Ocular 12.45%
More than double WER
2.19×
Pipecat 1.93% → Ocular 4.23%
More than double WER
2.10×
Pipecat 2.00% → Ocular 4.19%
More than double WER
1.63×
Pipecat 2.16% → Ocular 3.53%
Higher WER on Ocular
1.60×
Pipecat 2.46% → Ocular 3.93%
Higher WER on Ocular
1.47×
Pipecat 2.59% → Ocular 3.80%
Higher WER on Ocular
1.40×
Pipecat 2.10% → Ocular 2.93%
Higher WER on Ocular
1.33×
Pipecat 3.35% → Ocular 4.45%
Higher WER on Ocular
1.29×
Pipecat 2.39% → Ocular 3.09%
Higher WER on Ocular
1.27×
Pipecat 3.46% → Ocular 4.39%
Higher WER on Ocular
1.07×
Pipecat 3.90% → Ocular 4.18%
Higher WER on Ocular
0.95×
Pipecat 3.66% → Ocular 3.46%
Lower WER on Ocular
0.91×
Pipecat 4.96% → Ocular 4.53%
Lower WER on Ocular
0.69×
Pipecat 6.48% → Ocular 4.45%
Lower WER on Ocular

Each multiplier is Ocular WER ÷ Pipecat WER, calculated from rounded reported values. Below 1× means lower WER. Compare bars use 0–4×; both dataset views use the same 0–13% WER scale. Cards keep the same ratio-sorted order and background (higher/lower Ocular WER) while switching views. All 15 models are included.

The comparison includes models from AssemblyAI, Google, Reson8, OpenAI, Cartesia, Speechmatics, Inworld, Deepgram, Smallest, and Gradium. All 15 entries are shown, including the models that were more accurate on real two-person conversation.

What Converse-STT Measures

Converse-STT measures how a speech model behaves when the audio is real conversation. The score is word error rate. A lower score is a more accurate transcript of speech a product would actually hear, including repetitions, hesitations, false starts, and turn-taking. Every model carries both results on the leaderboard: one on the conversational recordings, one on the public Pipecat clips.

The Real-World, Human Dataset

This dataset was generated on Ocular AI’s platform. American English experts were hired through Workbolt and tasked to have a topic-based, free, full-duplex conversation: two people talking naturally, on a shared topic, at the same time. The audio is 48 kHz American English, recorded on separate speaker tracks with synchronization offsets and time-aligned, verbatim reference transcripts.

The package contains four conversations, each approximately 15 minutes long—about 60 minutes of conversation across eight speaker tracks. Each track is delivered as 24-bit mono FLAC with a paired JSON transcript containing segment and word timings.[1] We used these conversations in the evaluation, including a segmented-turn analysis.

The talk keeps the marks of a natural conversation: pauses, fillers, repetitions, and false starts. That is the speech a model has to handle to be useful in a product. The references are human-generated transcriptions and annotations. They preserve these spoken details and mark paralinguistic events with tags such as <laugh/> and <breath/>, also written by people.

Use the audio toggle below to hear an excerpt from one of these conversations next to a Pipecat voice-agent clip. Listen for the pauses, hesitations, and restarts in the conversational excerpt. These examples let you hear the supplied recordings; they are not a matched test or a summary of either dataset.

Listen to the reference

Real speech. Details intact.

Conversation 01 · Speaker B · American English
0:00 / 0:00
Verbatim reference
Okay. Okay. So the strategy I use is I got these, uh, <breath/> noise-canceling headphones, right? And I put them on, and I- I’ll listen to, like, the Rocky soundtrack, <breath/>
Uncut speaker-track excerpt · 00:08.320–00:31.234. The pause between the two “Okay”s is preserved. Breath tags are annotations, not spoken words; they are removed before WER scoring.

What The First Results Show

The transcripts were normalized before scoring. In the reviewed pipeline, non-speech annotation tags such as breaths and laughter were removed before word error rate was calculated. Models were therefore not penalized for omitting literal breath tags. This is distinct from removing all spoken fillers, repetitions, or false starts.[2]

The leading model changes with the dataset. AssemblyAI Universal 3.5 Pro leads the public Pipecat dataset at 1.93% WER, but appears ninth on Ocular AI’s conversational dataset at 4.23%. Reson8 moves from third on the public dataset at 2.10% to first on the conversational dataset at 2.93%.

The leading model changes with the dataset

Real-world speech. A different leader.

AssemblyAI Universal 3.5 Pro
Pipecat STT Dataset
#1
1.93% WER
Conversational · Ocular
#9
4.23% WER
Reson8
Pipecat STT Dataset
#3
2.10% WER
Conversational · Ocular
#1
2.93% WER
Rank among 15 tested models on each dataset. Lower WER is better. Rank is relative to the other models: moving up does not necessarily mean a lower absolute WER.

Three models perform better on the conversational dataset. Smallest Pulse records 3.46% WER, compared with 3.66% on the public dataset; Deepgram Flux Multilingual records 4.53%, compared with 4.96%; and Gradium records 4.45%, compared with 6.48%. These three run against the broader trend.

Why Do Models Fail?

Real conversation does not hold still. Two people revise themselves as they go, and inside a few seconds a turn can restart, repeat, stall, and repair. The human transcript keeps that motion, so every one of those words is in the reference a model is scored against.

The models that lead on public clips fail in much the same way once the speech is that exchange. The excerpt above keeps both “Okay”s and the pause between them, the filler “uh,” and a repaired start, “I- I’ll,” before the speaker reaches the point. A transcript that smooths that into one fluent sentence drops or rewrites words the reference still contains. The Pipecat comparison clip is a short, stable question: “How much juice is in one lime?” A low error rate there describes that question. Carrying every spoken word through a turn that changes while it is being said is a different task, and a cleaner clip will not close the gap. Models need to follow speech that stays in motion and hold every word of it through the whole turn.

Applicability

An application built on these models hears the transcript, then acts on it. A voice agent, a meeting note, a support summary, or a search index starts from the words the model wrote down. On a short public clip, AssemblyAI Universal 3.5 Pro reaches 1.93% word error rate, about two errors in a hundred words. On Ocular AI’s two-person conversations, that same model is at 4.23%, about four errors in a hundred words, and ninth of 15. Reson8, third on the public clips at 2.10%, is the model an application would actually be using on this speech, at 2.93%. The accuracy that ships is the conversational number.

That gap is what the rest of the product inherits. A word dropped or rewritten in the transcript is a word the agent, the summary, and every later step receive already wrong, and no amount of downstream reasoning recovers it. A model that doubles its error rate on this speech is the difference between a demo that holds and a product that misquotes its users. An application whose callers restart, repeat, and repair should be chosen on the score from that speech.

Key Takeaways

Six points from measuring frontier models on real-world conversation.

Key Takeaways

  • 01

    Frontier AI models struggle on real-world, unscripted data. Accuracy that holds on public clips often drops once the speech is a real two-person conversation, which is the behavior a product will see.

  • 02

    On real-world conversation, 12 of 15 models had a higher word error rate. They scored worse on Ocular AI’s two-person recordings than on public Pipecat audio, and four of them more than doubled.

  • 03

    The applicable leader changes with the data. AssemblyAI Universal 3.5 Pro leads Pipecat at 1.93% and ranks ninth on the conversational recordings at 4.23%. Reson8 moves from third at 2.10% to first at 2.93%.

  • 04

    Real-world data is where some models improve. Smallest Pulse, Deepgram Flux Multilingual, and Gradium each had a lower word error rate on the Hi-Fi recordings than on Pipecat.

  • 05

    The test is real human speech. Four American English conversations, about 60 minutes across eight speaker tracks, recorded at 48 kHz as 24-bit mono FLAC with verbatim, time-aligned transcripts.

  • 06

    A public score and a real-world score both belong on the leaderboard. Breath and laughter tags are removed before scoring, and each model still carries its Pipecat result next to its conversational result.

Conclusion

Converse-STT provides a benchmark for evaluating frontier speech models on real-world, two-person conversation and shows that accuracy on public voice-agent clips often drops once the speech is a real exchange between two people.

Converse-STT evaluates whether a model can transcribe full-duplex American English as people actually speak it: two voices, overlap, pauses, fillers, repetitions, and false starts, against human-generated references.

The benchmark consists of four conversations, about 60 minutes across eight speaker tracks, and a public comparison on 1,000 voice-agent clips from Pipecat’s STT benchmark dataset. It measures word error rate on both, and each model’s two scores sit side by side on the leaderboard.

Our evaluation revealed a clear gap in transfer: the models that lead on public audio are not the models that lead on real conversation.

The analysis indicates that a single public score leaves out the speech a product will hear, the reference convention behind the number, and how stable a model is across both datasets.

These recordings are American English, studio-captured, and human-transcribed, so the findings should be understood within that scope. Other languages, accents, and noisier environments are a further test. The public sample on Hugging Face contains recordings and reference transcripts from this collection; the score files behind the table stay with the evaluation.

We hope Converse-STT serves as a practical guide for choosing a speech model for real conversation and as a concrete target for the transcription accuracy frontier models need on human-generated speech.

See the Converse-STT leaderboard.[3]

Evaluate Your Model On Real-World Conversation

Building a speech model or choosing one for a product? The useful question is how it behaves on the speech your users actually produce. Ocular AI benchmarks are built for that test.

Ocular AI helps teams source real conversational audio and time-aligned annotations for evaluation. Share the languages, accents, recording conditions, and transcription requirements your application needs, and we can help scope a dataset around them.

Start with samples. Tell us about your model or product through the form, and our team will follow up with relevant sample clips and next steps.

Request conversational data samples →

Author

Louis Murerwa

Louis Murerwa

Co-founder & CTO

Ready to bring AI into the real world?