Building toward the Human Voice Turing Test

Converse Benchmarks

A benchmark family for frontier voice AI. Measuring how accurately AI understands speech, how naturally it speaks, and how effectively it participates in real conversation.

Work with us

Word error rate · lower is better

Show 10 more

15 Models · 2 Datasets · 60 Min of Conversation

Converse-STT

Conversational Speech-to-Text Benchmark

Can frontier models accurately understand and transcribe natural conversation?

Evaluates frontier AI and speech-to-text models on Real World Conversational Data and on agent-generated, publicly available Pipecat audio. It measures word error rate (WER) on each dataset and shows how each model's score shifts between them.

Category

Speech-to-Text

Languages

English
In collaboration withCekura

Can AI converse like a human?

We believe voice will be the natural interface between people and AI, but speech is uniquely hard to get right. Unlike code or math, there is rarely one right answer: a good conversation depends on the setting, and every utterance carries emotion, tone, pace, accent, and language on top of its words.

Strong scores on clean clips don't predict how a model holds up in a live conversation. Ocular Converse measures voice AI on how people actually speak, the conditions they speak in, and how a conversation unfolds.

Core Research Areas

Four capabilities of human conversation

Understanding

Can frontier models accurately understand and transcribe natural conversation?

People restart sentences, hesitate, overlap, and correct themselves mid-thought. A model has to capture what was actually said, across speakers, accents, and settings, not just clean read speech.

Measured today by Converse-STT

Speaking

Can the model generate natural, expressive, context-appropriate speech?

A line that sounds smooth on its own can sound wrong in context. “That’s great” needs different delivery after good news than after a disappointed correction. Emphasis, pacing, emotion, and pronunciation all have to fit what came before.

Conversing

Can the model listen and speak with appropriate timing, overlap, and interruption handling?

Excellent components don’t make a good conversation partner. A model has to know when to speak, wait, acknowledge, stop, and recover from a misunderstanding, in real time.

Acting

Can the model understand spoken instructions and carry them out?

Some words matter more than others. In “Send it Thursday, sorry, Tuesday, to Maya,” the correction, the date, and the name decide whether the task succeeds.

The goal is scores that predict what people experience when they talk to a model: the bar a Human Voice Turing Test sets.

Methodology

How we benchmark

Real conversation

Unscripted, two-person conversations recorded by experts hired through Workbolt, with a separate studio-grade track for each speaker.

Public baselines alongside

Every model is also scored on public audio, and both results are shown side by side so the gap is visible.

Open samples

A public sample of the recordings and reference transcripts is published on Hugging Face so results can be inspected.

Slices, not silos

Languages, accents, and recording conditions are added as test slices within a benchmark, keeping results comparable.

Frequently Asked Questions

Ocular Converse is a benchmark family for frontier voice AI. It measures how accurately AI understands speech, how naturally it speaks, and how well it takes part in real conversation. Every benchmark in the family works toward one question: can AI converse like a human? The long-term goal is a Human Voice Turing Test, where people talking to a model can't tell it apart from another person.

Converse-STT is the first benchmark in the family. In collaboration with Cekura, it scores 15 speech-to-text models on how accurately they transcribe natural, two-person American English conversation. Each benchmark in the family takes its name from the capability it measures.

The conversations are recorded by experts hired through Ocular AI's Workbolt platform. They are unscripted, two-person, full-duplex recordings, captured at studio quality on a separate track for each speaker, with verbatim, time-aligned transcripts.

They are test slices within a benchmark, not separate benchmarks. Converse-STT currently covers American English; new languages, accents, and recording conditions are added as slices so results stay comparable across the family.

Yes. Contact us to evaluate a model on Ocular Converse or to license Hi-Fi conversational speech.

Ready to bring AI into the real world?