$2M Pre-Seed Round from Drive Capital, Y Combinator, and Other Investors.
Today, Ocular AI is excited to announce our $2 million pre-seed round, led by Drive Capital, with participation from Y Combinator, Alumni Ventures, 1745 Ventures, Orange Collective, MyAsia VC, and a group of angel investors.



Silicon Valley, CA: Today, Ocular AI is excited to announce our $2 million pre-seed round, led by Drive Capital, with participation from Y Combinator, Alumni Ventures, 1745 Ventures (formerly known as Bertelsmann Digital Media Investments, BDMI), Orange Collective, MyAsia VC, and a group of angel investors.
Since launch, our mission has been simple: encode human expertise into frontier AI models, and bring frontier AI into the real world. We started with voice, the most natural interface humans have, and the one where the gap between a benchmark score and a real conversation is still audible to anyone who tries it.
Today, some of the top frontier AI labs and Fortune 100 enterprises train and evaluate on Ocular AI's datasets, evaluations, and benchmarks. We work hand in hand with the leading and fastest-growing labs building voice-native models, operating as an extension of their teams: capturing conversation at studio quality, building the evaluations that show where their models break, and shipping the data that fixes it. Revenue is in the seven figures and growing quickly, and the Expert Network behind it now counts thousands of vetted domain experts.
This round is just the beginning.
The most natural interface is the hardest to get right
AI only becomes meaningful when you can naturally interact with it. Voice is the most human interface we have, and it carries far more information bandwidth than text alone can capture: tone, timing, hesitation, emphasis, the "mm-hmm" that keeps a speaker going. A transcript keeps the words and throws the rest away. If the goal is to encode human expertise into models, speech is where that expertise is most visible and least written down.
Over the past year it became the frontier's priority. NVIDIA, Thinking Machines Lab, OpenAI, and Google have each shipped or previewed audio-native models that are also full duplex: they listen and speak at the same time.[1][2][3][4] Until now, most voice AI has been a cascade: a speech-to-text model transcribes what you said, a language model reads the transcript and writes a reply, and a text-to-speech model reads that reply aloud. Each stage only sees text, so everything that isn't a word, the tone, the pause, the fact that you started talking again, is lost at the first step. A voice-native model skips the transcript. One model hears tone and timing directly, decides when to speak and when to listen, and handles turn-taking and interruptions in real time. We're moving from AI you type at to AI you talk with. Yet these models still stumble on underrepresented accents, miss the backchannels and barge-ins that make conversation feel natural, and freeze when a real human interrupts.
Voice is also not the end of it. We're seeing labs build models that can talk and see at the same time, the way humans do: models that watch a face, read a gesture, and hear a tone in one stream and respond in real time. We're working with some of the biggest players building these models to generate high-fidelity audiovisual datasets on our infrastructure.
Frontier data is the missing layer
The bottleneck is not data in general. There is more speech on the internet than any model could consume. The bottleneck is frontier data: data that captures what models can't already do, at a fidelity they can learn from. Working with the top labs, the same two problems come up every time. Collecting that data is hard: models have to learn from real conversation, with the overlap and the pause before an answer, captured with each speaker on their own clean channel, and none of that is on the open web. Evaluating models is just as hard, because the benchmarks that exist rarely use real-world speech. As we showed in Beyond Fisher, almost every open full-duplex model still leans on one 8 kHz telephone corpus recorded in 2004.
Compute is commoditizing and architectures diffuse quickly through open research. The data a lab trains and evaluates on is becoming the thing that sets it apart, and it is the layer the field has chronically underinvested in.
Our approach: experts and research, together
Our core belief is that data complex enough to teach today's frontier models takes two things together: domain experts, and rigorous research to understand where models fail in the real world. Experts alone produce more data. Research alone produces findings nobody can train on. Scaling AI to the next frontier takes both: evaluations that show exactly where a model breaks, experts who know what right looks like, and the systems to turn the two into frontier datasets that push the model past its failures. We believe training data deserves the same level of rigor that the frontier AI labs apply to their work on training algorithms and model architectures, and we treat it that way.
Experts know what right looks like. Research shows where models fall short. Frontier data needs both.
That is what we've built: an Expert Network of thousands of vetted domain experts, and the data infrastructure to capture and generate high-fidelity datasets at the scale and quality that push models to the next frontier.
What's next: building an Applied AI Data Research Lab
We've built the infrastructure and the Expert Network to generate high-fidelity datasets at scale. We're now doubling down on evaluation suites and research, starting with Converse, our benchmark family for frontier voice AI, built toward one question: can AI converse like a human? Converse measures models across four areas. Understanding: can the model accurately transcribe natural conversation, with its restarts, overlaps, and accents? Speaking: can it generate speech whose emphasis, pacing, and emotion fit the context? Conversing: can it listen and speak with the right timing, and handle interruptions in real time? Achieving outcomes: can it turn a spoken conversation into the right result? The first benchmark, Converse-STT, is live today. It measures how accurately frontier speech-to-text models transcribe real-world conversational data, with all the restarts, overlaps, and hesitations that clean read speech leaves out.
The same infrastructure is not limited to voice. The Expert Network and the pipeline behind it can power data for any domain where expertise is the bottleneck, including professional domains such as medicine, law, finance, and software engineering, and we're beginning to expand into them as we go.
What we're building toward
If you're building voice or multimodal models, talk to us. If you're an expert who wants to shape the next generation of AI, join the network.
Join us in pushing the frontier
We're a small team in San Francisco, with engineers from Dartmouth College, IITs, Microsoft, Google, and Atlassian, and the next stage of this work needs more people who care about getting the data right. If you want to build the data layer that brings frontier AI into the real world, see our open roles or write to us at careers@useocular.com.

Authors

Michael Moyo
CEO & Co-founder

Louis Murerwa
CTO & Co-founder


