September 28, 2026ProductRelease

Introducing High-Fidelity, Full-Duplex Audiovisual Conversational Datasets

Speech models went full-duplex in 2024, and video conversational models are going full-duplex now. This full-duplex audiovisual dataset records two-person video conversations, with each speaker's audio and video as separate streams on one shared clock: the data shape those models need.

Louis MurerwaMichael Moyo
Louis Murerwa and Michael Moyo
Introducing High-Fidelity, Full-Duplex Audiovisual Conversational Datasets

Today we are introducing our full-duplex audiovisual dataset: two-person video conversations in which each speaker's audio and video are recorded as separate streams on one shared clock.

It is built for the next generation of Video Conversational Models, that listen and speak at the same time.

Models Are Going Full-Duplex

In 2024, Kyutai introduced Moshi, the first real-time full-duplex spoken language model.[1] Moshi's breakthrough was to stop treating conversation as a sequence of turns and instead generate speech "while modeling separately its own speech and that of the user into parallel streams." This allowed for "the removal of explicit speaker turns, and the modeling of arbitrary conversational dynamics."

Two streams, both always running, each one belonging to one side of the conversation.

In 2026 the same idea has reached Video Conversational Models.

Full-duplex reaches video

Five systems, in their authors' own words.

SystemDateIn its authors' words
INFP[2]December 2024A conversational agent "should not be predetermined a role, but rather be able to freely switch between listening and speaking states"
LPM 1.0[3]April 2026"single-person full-duplex audio-visual conversational performance", in which "characters simultaneously speak, listen, react, and emote"
TML-Interaction-Small[7]May 2026Interaction models "continuously take in audio, video, and text, and think, respond, and act in real time"
Wan-Streamer[4]June 2026"real-time, low-latency, full-duplex audio-visual interaction", modelling "language, audio, and video as both input and output within a single Transformer"
Tavus Phoenix-4.5[5]September 2026The animator "listens to two audio streams, the PAL’s own speech and the user’s"

The direction is the same in research as it is in product innovation: an avatar evolving from a face that speaks when given audio into an active participant with the ability to listen, react and speak at the same time. It is one side of a conversation, and it has to behave like one the whole time, including when it is silent.

"In a real conversation, we keep reacting while we listen — our expression changes, our head and posture shift, and small movements signal that we’re still present in the exchange."

— Tavus, Phoenix-4.5[5]

The same recording serves a model that watches rather than one that is watched. In May 2026, Thinking Machines Lab introduced interaction models that keep taking in audio and video while they respond, in 200 ms slices instead of turns.[7] An avatar has to be seen listening. An interaction model has to see and hear the person it is listening to, including when both sides are talking. Both need each speaker's audio and video kept as separate streams on one clock.

Full-Duplex Video Models Require Full-Duplex Data

A model that handles two streams at once has to be trained on two streams at once, and that poses a requirement on the shape of the data, before any question of quantity or quality.

Video data comes in three shapes.

The shape of the data

Three shapes of conversation video, and what a model can learn from each.

ShapeWhat it containsWhat a model can learn from it
Single speakerOne person facing a camera: one voice, one faceHow to speak
Edited conversationTwo people and one output: both voices mixed together, the camera on one person at a timeHow to speak, and fragments of how to listen
Full duplexTwo people, each recorded separately: each voice and each face as its own stream, all on one clockSpeaking, listening, overlap, and the timing between two people

Almost all public data is in the first two shapes.

Single speaker. The datasets that taught avatars to talk are one person facing a camera: VoxCeleb2, HDTF, CelebV-HQ, MEAD. HeyGen describes the audio stage of Avatar V as trained on "a broad corpus of talking-head video spanning diverse speakers, languages, and speaking styles."[6]

Edited conversation. Interviews, podcasts and films contain two people, but the edit follows the speaker. The authors of LPM 1.0 measured what is left for the listener:[3] "only approximately 10% of all conversational segments are framed on the listener". Their conclusion: "This inherent imbalance makes listener-centric data both scarce and valuable for training models that generate realistic listening behavior."

Full duplex. Here the public record is thin. The INFP paper surveys it:[2] ViCo "is a small-scale dataset with 1.6 hours and 96 IDs, and it lacks multi-turn conversation scenarios", and for ViCo-X "the total duration is only 0.4 hours."

The evolution to full duplex started in speech, where the need for two-channel audio was the bottleneck. We wrote about that in Beyond Fisher. Video is at the same point now, with one more modality to keep in sync, and the full-duplex audiovisual dataset is our answer for video, as the Full-Duplex Hi-Fi corpus was for speech.

The Full-Duplex Audiovisual Dataset

The full-duplex audiovisual dataset is built to address the bottleneck in full-duplex video data. Each conversation is recorded with two people live on a video call, talking for about 15 minutes. Each person is recorded as one audio stream and one video stream, and all four streams sit on a shared timeline.

Below we share a sample, shown the way it appears in the delivered set. The two videos carry five minutes of Conversation 1, one stream per speaker, played together. Speaker A does most of the talking for the first two minutes. At about 2:15 the roles change and Speaker B takes over. Both people are on camera and on their own microphone throughout, whether they are speaking or listening.

Sample

One conversation, four synchronized streams.

A live video call, with each person on their own camera and microphone.

Speaker A · Male, Asian · video, H.264

Speaker B · Female, two or more races · video, H.264

0:00 / 0:00

Preview: five minutes of the conversation, from 1:00 to 6:00, with faces blurred and the picture scaled down for the web. The delivered files are not blurred: one video file per speaker, H.264 at 720p.

Speaker A · audio, FLAC, 48 kHz, 24-bit, mono

0:00 / 0:00

Speaker B · audio, FLAC, 48 kHz, 24-bit, mono

0:00 / 0:00

Alignment & verification

Turn-taking

Both speakers' voice activity on the shared timeline: who is talking when, and where they overlap. This window is the first minute and a half of the preview.

Spectrogram

Frequency against time for each speaker's full 15-minute recording, up to 24 kHz.

Alignment

Each track's offset from an independent reference, one point per measurement window across the full recording. Grey spans are windows where no reference was available.

Lip sync

Voice activity from each speaker's audio file and facial motion from their video file, on one time axis. This window is the first minute of the preview, where Speaker B is mostly listening.

Request the sample set

Verification

Four figures ship with every conversation.

FigureWhat it shows
Turn-takingBoth speakers' voice activity as two lanes on the shared timeline: who is speaking, when, and where they overlap
SpectrogramEnergy across frequency and time for each speaker, up to about 24 kHz. Confirms full-band audio with no clipping, dropouts or resampling artefacts.
AlignmentEach track's offset from an independent reference, in milliseconds, measured across the whole recording
Lip syncVoice activity and facial motion for each speaker, overlaid on one time axis

Each file's measured values are in its metadata.json. After the automated checks, a reviewer watches and listens to every conversation on synchronised side-by-side playback.

Pushing the Frontier with the Full-Duplex Audiovisual Dataset

The full-duplex audiovisual dataset is designed to push the frontier of video conversational models. The capabilities frontier companies are pushing to improve are expressiveness, speed, identity preservation, whole-frame generation and listening. For each one, this is what our dataset offers.

Expressiveness. An expressive model is more present in a conversation: its face and body say what words do not. Our dataset contains the real expression of both people, in conversations spoken in their own words.

Speed. A conversation breaks when the reply comes late. Our dataset contains natural backchannels and turn changes on a shared timeline, so a model can learn realistic turn-taking: when people respond, and how fast.

Identity preservation. A person has to stay recognisably themselves as they move. Our dataset records each person for about 15 minutes without a cut, so a model sees one face through many expressions.

Whole-frame generation. A model that generates the head, shoulders and torso has to learn from video that shows them. Our dataset records the full camera frame at 720p on each speaker's own device, not a region around the face.

Listening. A model that goes still when it stops talking has left the conversation. Our dataset records each person on separate streams, so a model can see and hear how people listen, from a nod to an "mm".

Both sides at once. Moshi models two audio streams. A full-duplex video model has four: a voice and a face for each side. Our dataset provides all four on one clock.

Interaction. An interaction model answers while the other person is still going: an interjection, a silence, a nod, a correction. Our dataset leaves both people on camera and on their own microphone for the whole conversation, so those moments stay in the file.

How It Compares

Dataset comparison

The full-duplex audiovisual dataset against the datasets the field trains on.

DatasetShapeSizeAudioVideoSettingLicence
Full-duplex audiovisual datasetFull duplexOver 1,000 hours48 kHz, 24-bit, lossless720p or higherLive video callOn request
DyConvTwo peopleOver 200 hoursNot statedNot statedCollected from the internetNot stated
ViCoTwo people1.6 hoursNot statedNot statedNot statedNot stated
CelebV-HQSingle speaker35,666 clipsWeb video512×512 or largerWeb videoNot stated
HDTFSingle speakerAbout 16 hoursWeb video720p to 1080pWeb videoNot stated
VoxCeleb2Single speakerOver a million utterances, over 6,000 speakersWeb videoWeb videoInterviewsNot stated

Comparator values are taken from each dataset's paper. ViCo and DyConv values are as reported in the INFP paper.

Train Your Model on Full-Duplex Video Data

Building a video conversational model, an interactive avatar, or an interaction model? The useful question is how it behaves when the other person is talking. The full-duplex audiovisual dataset is built for that test. The full specifications, verification figures, and dataset comparison are on the dataset page.

Ocular AI helps teams source real two-person video conversation, with each speaker's audio and video as separate streams on a shared timeline. Share the conversations, formats and annotation layers your model needs, and we can help scope a dataset around them.

Start with samples. Tell us about your model or product through the form, and our team will follow up with sample conversations and next steps. Each one comes with per-speaker audio and video, a metadata file, and its four verification figures.

Request full-duplex audiovisual samples →

Get in touch here to access our off-the-shelf full-duplex conversational, Hi-Fi multilingual, and multi-accent English ASR datasets, or reach out to us directly at research@useocular.com.

Authors

Louis Murerwa

Louis Murerwa

Co-founder & CTO

Michael Moyo

Michael Moyo

Co-founder & CEO

Ready to bring AI into the real world?