Introducing Converse: Benchmarks for Conversational Voice AI
Converse is a benchmark family for frontier conversational voice AI, building toward the Human Voice Turing Test. It measures how well AI understands speech, how naturally it speaks, how well it takes part in real conversation, and whether it achieves the outcomes people ask for, starting with Converse-STT.


Listen closely to two people talking and it is rarely a tidy exchange of turns. They overlap, murmur "mm-hmm" while the other speaks, stop mid-word to rephrase, and hand the floor back and forth in a fraction of a second. Voice AI is starting to work this way too. How we measure it hasn't caught up.
Converse is a benchmark family for conversational voice AI, built to close that gap. Its long-term target is the Human Voice Turing Test: a conversation in which a person can't tell whether they are talking to a person or to AI. Measured against real human conversation, Converse asks how well a voice model understands speech, how naturally it speaks, how well it takes part, and whether it achieves what the person wanted. The first benchmark, Converse-STT, launches alongside this post. Explore the Converse benchmarks.
From Cascades to Conversation
For most of the past decade, voice AI has been built as a cascade. A speech-to-text model transcribes what the user said, a language model decides what to say back, and a text-to-speech model reads it aloud. The design made sense: each part could be built, improved, and swapped out independently.
It also sets the limits of the conversation. Each handoff passes along text, so tone, hesitation, and timing are lost at the first step. Each stage adds latency, which produces the familiar listen, pause, think, speak rhythm. And because the system only acts once a turn is finished, it can't backchannel, can't be interrupted gracefully, and can't tell a pause for thought from the end of a sentence.
Over the past year, the frontier has started to move past that design. Full-duplex models fold listening, reasoning, and speaking into a single system that runs continuously. NVIDIA's PersonaPlex listens and speaks at the same time, handling interruptions and backchannels.[1] Thinking Machines Lab's interaction models interleave input and output in 200-millisecond micro-turns.[2] OpenAI's GPT-Live decides many times per second whether to speak, listen, pause, interrupt, or call a tool.[3] Google's Gemini 3.8 Live models reason and speak while tools run in the background.[4]

The transition is under way, not finished. Most voice products in production are still cascades, and they will be for some time. That leaves evaluation with two jobs at once: measure the components cascades depend on, on the speech real users produce, and measure the conversational behavior full-duplex models will be judged by. Converse is designed to do both, with real human speech as the reference.
Benchmarks Built for the Cascade
Most of the benchmarks we have grew up alongside the cascade, so most score a single component. The newest have started to test interaction.
Speech recognition. LibriSpeech, about 1,000 hours of read audiobooks, is effectively saturated, with leading systems below 2% word error rate.[5] Switchboard and Fisher contain real conversation,[6][7] and Fisher still trains dialogue models such as Moshi,[8] but they are 8 kHz telephone recordings from decades ago and likely sit in many training sets. AMI and CHiME-6 test meetings and far-field audio, which is mostly acoustic robustness.[9][10] The Open ASR Leaderboard averages public test sets,[11] and Pipecat's benchmark covers short voice-agent utterances.[12]
Speech synthesis. Text-to-speech is scored on isolated utterances: mean opinion score (MOS) or predictors like UTMOS,[13] plus intelligibility and speaker similarity in suites like Seed-TTS-eval.[14] None ask whether a line was delivered right given what came before.
Spoken dialogue. VoiceBench and Big Bench Audio test what an assistant says in response to a spoken prompt, largely as a single turn.[15][16] Full-Duplex-Bench and Talking Turns have begun to measure pauses, backchannels, and interruptions.[17][18]
What's missing across all of them, old and new, is the same thing: a trustworthy human reference. Most test speech is read or scripted. The conversational corpora that exist are decades-old narrowband recordings, often with both speakers mixed, so overlapping speech can't be attributed to either one. And test sets that have been public for years are at risk of contamination. Without separate, precisely timed recordings of real people talking, there is nothing reliable to compare a model's words or timing against.
What Converse Measures
No model is close to passing the Human Voice Turing Test, and no single score could say how close. So Converse breaks "Can AI converse like a human?" into four questions. The first two score individual skills, hearing and speaking. The last two score the whole system inside a conversation: how it behaves in the exchange, and whether the exchange ends where the person wanted. All four are scored against real human speech, recorded and transcribed for evaluation. The shape of that data follows the question: Converse-STT uses two-person conversations, while other benchmarks may call for group conversations, a caller and an agent, noisy environments, or domain-specific speech.
Understanding
Can a model accurately understand and transcribe natural conversation?
Every voice system has to get the words right first, whether it's a cascade's speech-to-text stage or a full-duplex model's listening. Converse measures this with word error rate against a human verbatim reference that keeps every restart, filler, and repair.
Word error rate
- Substituted words
- Deleted words
- Inserted words
- Words in the human reference
Separate tracks let errors be attributed per speaker, including where the speakers overlap. To show what that catches, we ran Whisper large-v3, a widely used open model, on one full 15-minute conversation from the public Converse sample. Its transcript reads fluently but has 172 errors in about 3,000 words. Most are deletions of exactly what a verbatim reference keeps: "um," "you know," cut-off restarts, and backchannels. Some are hallucinations, like "Thank you." on silence and a run of repeated "Yes." Use "Next error" to step through them.
One full conversation, scored
15 minutes of real conversation.
Loading transcript…
Word error rate weights every word equally: dropping "uh" costs the same as dropping the corrected date in "Thursday, sorry, Tuesday." Future versions will add measures for self-corrections, names, and numbers.
Speaking
Can a model generate natural, expressive, context-appropriate speech?
Getting the words right is only half of the exchange. "That's great" needs different delivery after good news than after a disappointed correction. Converse will evaluate speech in context, judging whether prosody, pacing, and emotion fit the preceding conversation, compared with how the real person responded at the same moment.
Conversing
Can a model listen and speak with appropriate timing, overlap, and interruption handling?
This is where full-duplex models are meant to pull ahead of cascades, and where the human reference matters most. Across ten languages, gaps between turns cluster around 200 milliseconds.[19] Separate-channel recordings let us compute the same statistics for people and models: floor-transfer offsets, backchannel timing and rate, how quickly a model yields when interrupted, and whether it waits through a mid-turn pause.
The excerpt below comes from the same kind of recording. In 22 seconds, two speakers overlap seven times, backchannel without taking the floor, start at the same moment and negotiate who continues, and repair their own sentences. A model that treats every pause as the end of a turn, or every overlap as an interruption, would get most of it wrong.
Listen to a real conversation
Two speakers. One timeline.
A<breath/><breath/>
B
B
B<breath/><breath/>
A<breath/>
A<breath/>
Achieving Outcomes
Can a model turn a spoken conversation into the right outcome?
Voice agents are increasingly judged by what they do, not just what they say. Agent benchmarks already measure this in text. τ-bench simulates customer-service conversations and scores whether the database ends in the correct state, with a pass^k metric for how consistently an agent succeeds across repeated trials.[20] τ²-bench extends it to settings where the user also has to take actions, such as troubleshooting their own phone.[21] In both, the user's request arrives as clean text.
Spoken, it rarely does. In "move it to Thursday, sorry, Tuesday, and send it to Maya's," the correction and the name decide whether the task succeeds. Testing that takes more than a list of prompts. It takes environments: simulated, data-rich worlds in which a voice model has to reach a goal through conversation.
Each Converse environment will have four parts:
- A world. Databases with realistic depth, such as accounts, orders, calendars, and contacts, plus the policies that govern them. The agent can only change the world through tools.
- A goal. A target state the conversation should produce, like a delivery moved to the right date and address, specified so it can be checked without judging the transcript.
- A simulated user who speaks like a person. Built from real Converse recordings, the user hesitates, corrects themselves, trails off, confirms with "mm-hmm," and interrupts when the agent gets something wrong.
- A verifier. The episode is scored on the final state of the world. Repeated runs give a reliability score in the style of pass^k.
That verifier is also a reward. The same environment that measures a model can train it with reinforcement learning, one way voice models can learn to act on speech rather than on clean transcripts. The episode below shows the loop, and what one missed correction costs.
Achieving an outcome in a simulated world
One spoken correction decides the reward.
Press play to run the episode, or step through it one action at a time.
The Data Underneath
Each of those measurements depends on the reference data being right, so every Converse benchmark shares one design, built to fix the gaps above.
- Real speech, shaped to the question. Unscripted recordings by experts from Workbolt, Ocular AI's expert network, with each speaker on a separate track. For Converse-STT, that means two-person conversations recorded at 48 kHz, 24-bit.
- Verbatim, time-aligned references. Human transcripts keep fillers, repetitions, and false starts, with word timestamps and tags like
<laugh/>and<breath/>. - Documented normalization. Rules are published with each benchmark. Non-speech tags are removed before scoring; spoken disfluencies stay.
- Private test sets, public samples. Full test sets are held out to limit contamination, with a public sample on Hugging Face.
- Public baselines alongside. Every model is also scored on public audio, and the gap is often the most informative number.
- Independent evaluation. For Converse-STT, Cekura ran all 15 model evaluations.
Built on Ocular's Expert Network and Data Foundry
Data like this is hard to produce at the quality a benchmark needs, and harder to keep producing as the benchmarks grow. Converse runs on the same two-part stack Ocular AI uses to build training data for frontier labs.
The Expert Network. Workbolt, our expert network, supplies the people each benchmark depends on, identity-verified and credential-checked. For Converse that means native speakers across languages and accents, recorded in whatever configuration a benchmark needs, linguists and trained transcribers who write verbatim references, and, as the family grows, raters who judge speech in context and domain professionals, such as customer-support leads, who design the goals, policies, and tasks inside outcome environments.
The Data Foundry. Our Data Foundry turns those contributions into benchmark-grade data. Capture tooling records each speaker on a separate studio-grade track with synchronization offsets. Structured task interfaces guide transcription and annotation. Automated checks for clipping, noise, completeness, and alignment run before anything reaches a reviewer, and every file then passes human review against a rubric. Each recording stays traceable to a contributor, a consent record, and its review passes.
How Converse is produced
Experts in, benchmarks out.
Because we recruit the speakers and run the capture, a benchmark is not limited to whatever data already exists. We design the data around what we want to measure: two people or a group, casual talk or domain calls, a quiet studio or a noisy street. The same stack that produced the Converse-STT recordings can build what comes next: preference judgments for speech in context, timing annotations for conversation, and the simulated worlds, spoken users, and verifiers that outcome environments require.
First Results: Converse-STT
We started with understanding because every voice system, cascade or full-duplex, depends on it. Converse-STT scores 15 speech-to-text models from AssemblyAI, Google, Reson8, OpenAI, Cartesia, Speechmatics, Inworld, Deepgram, Smallest, and Gradium on about 60 minutes of real conversation and on 1,000 public clips from Pipecat's STT benchmark.
- 12 of 15 models had a higher word error rate on real conversation, and four more than doubled.
- The leader changed with the data. AssemblyAI Universal 3.5 Pro leads on Pipecat at 1.93% but ranks ninth on real conversation. Reson8 moves from third to first at 2.93%.
- Some models did better on real speech, including Smallest Pulse, Deepgram Flux Multilingual, and Gradium.
A public score alone would have pointed a product team to the wrong model. That is the case for Converse in a single result: the speech a model is tested on changes which model looks best. Full methodology is in the Converse-STT post, and scores are on the leaderboard.
Where Converse Goes Next
Converse-STT covers understanding. The next Converse benchmarks will follow the same path as the models, from components to the full conversation, and explore these research directions:
- Speech in context. Whether a model's prosody, pacing, and emotion fit the conversation so far, compared with how real people responded at the same moment.
- End-to-end voice models. Scoring systems that hear and respond in a single model, where the pieces can't be evaluated separately.
- Conversational timing. Turn-taking, backchannels, interruptions, and pauses, measured against the timing statistics of real human conversation.
- Achieving outcomes. Simulated, data-rich environments with spoken users, where success is the final state of the world.
- Coverage and meaning. More languages, accents, and noisier recording conditions, and scores that separate harmless errors from ones that break the task.
Every benchmark will be scored against real human recordings, in whatever shape its question calls for.
The destination is the same one the models are heading toward: a conversation that feels human. Converse is how we will know how close they are.
Get Involved
Building a speech or voice model? Add it to a Converse benchmark. Want to collaborate on the research? Work with us or email research@useocular.com.
Authors

Michael Moyo
Co-founder & CEO

Louis Murerwa
Co-founder & CTO


