Introducing High-Fidelity, Full-Duplex Audiovisual Conversational Datasets
Speech models went full-duplex in 2024, and video conversational models are going full-duplex now. This full-duplex audiovisual dataset records two-person video conversations, with each speaker's audio and video as separate streams on one shared clock: the data shape those models need.



Today we are introducing our full-duplex audiovisual dataset: two-person video conversations in which each speaker's audio and video are recorded as separate streams on one shared clock.
It is built for the next generation of Video Conversational Models, that listen and speak at the same time.
Models Are Going Full-Duplex
In 2024, Kyutai introduced Moshi, the first real-time full-duplex spoken language model.[1] Moshi's breakthrough was to stop treating conversation as a sequence of turns and instead generate speech "while modeling separately its own speech and that of the user into parallel streams." This allowed for "the removal of explicit speaker turns, and the modeling of arbitrary conversational dynamics."
Two streams, both always running, each one belonging to one side of the conversation.
In 2026 the same idea has reached Video Conversational Models.
Full-duplex reaches video
Five systems, in their authors' own words.
| System | Date | In its authors' words |
|---|---|---|
| INFP[2] | December 2024 | A conversational agent "should not be predetermined a role, but rather be able to freely switch between listening and speaking states" |
| LPM 1.0[3] | April 2026 | "single-person full-duplex audio-visual conversational performance", in which "characters simultaneously speak, listen, react, and emote" |
| TML-Interaction-Small[7] | May 2026 | Interaction models "continuously take in audio, video, and text, and think, respond, and act in real time" |
| Wan-Streamer[4] | June 2026 | "real-time, low-latency, full-duplex audio-visual interaction", modelling "language, audio, and video as both input and output within a single Transformer" |
| Tavus Phoenix-4.5[5] | September 2026 | The animator "listens to two audio streams, the PAL’s own speech and the user’s" |
The direction is the same in research as it is in product innovation: an avatar evolving from a face that speaks when given audio into an active participant with the ability to listen, react and speak at the same time. It is one side of a conversation, and it has to behave like one the whole time, including when it is silent.
"In a real conversation, we keep reacting while we listen — our expression changes, our head and posture shift, and small movements signal that we’re still present in the exchange."
— Tavus, Phoenix-4.5[5]
The same recording serves a model that watches rather than one that is watched. In May 2026, Thinking Machines Lab introduced interaction models that keep taking in audio and video while they respond, in 200 ms slices instead of turns.[7] An avatar has to be seen listening. An interaction model has to see and hear the person it is listening to, including when both sides are talking. Both need each speaker's audio and video kept as separate streams on one clock.
Full-Duplex Video Models Require Full-Duplex Data
A model that handles two streams at once has to be trained on two streams at once, and that poses a requirement on the shape of the data, before any question of quantity or quality.
Video data comes in three shapes.
The shape of the data
Three shapes of conversation video, and what a model can learn from each.
| Shape | What it contains | What a model can learn from it |
|---|---|---|
| Single speaker | One person facing a camera: one voice, one face | How to speak |
| Edited conversation | Two people and one output: both voices mixed together, the camera on one person at a time | How to speak, and fragments of how to listen |
| Full duplex | Two people, each recorded separately: each voice and each face as its own stream, all on one clock | Speaking, listening, overlap, and the timing between two people |
Almost all public data is in the first two shapes.
Single speaker. The datasets that taught avatars to talk are one person facing a camera: VoxCeleb2, HDTF, CelebV-HQ, MEAD. HeyGen describes the audio stage of Avatar V as trained on "a broad corpus of talking-head video spanning diverse speakers, languages, and speaking styles."[6]
Edited conversation. Interviews, podcasts and films contain two people, but the edit follows the speaker. The authors of LPM 1.0 measured what is left for the listener:[3] "only approximately 10% of all conversational segments are framed on the listener". Their conclusion: "This inherent imbalance makes listener-centric data both scarce and valuable for training models that generate realistic listening behavior."
Full duplex. Here the public record is thin. The INFP paper surveys it:[2] ViCo "is a small-scale dataset with 1.6 hours and 96 IDs, and it lacks multi-turn conversation scenarios", and for ViCo-X "the total duration is only 0.4 hours."
The evolution to full duplex started in speech, where the need for two-channel audio was the bottleneck. We wrote about that in Beyond Fisher. Video is at the same point now, with one more modality to keep in sync, and the full-duplex audiovisual dataset is our answer for video, as the Full-Duplex Hi-Fi corpus was for speech.
The Full-Duplex Audiovisual Dataset
The full-duplex audiovisual dataset is built to address the bottleneck in full-duplex video data. Each conversation is recorded with two people live on a video call, talking for about 15 minutes. Each person is recorded as one audio stream and one video stream, and all four streams sit on a shared timeline.
Below we share a sample, shown the way it appears in the delivered set. The two videos carry five minutes of Conversation 1, one stream per speaker, played together. Speaker A does most of the talking for the first two minutes. At about 2:15 the roles change and Speaker B takes over. Both people are on camera and on their own microphone throughout, whether they are speaking or listening.
Sample
One conversation, four synchronized streams.
A live video call, with each person on their own camera and microphone.
Speaker A · Male, Asian · video, H.264
Speaker B · Female, two or more races · video, H.264
Preview: five minutes of the conversation, from 1:00 to 6:00, with faces blurred and the picture scaled down for the web. The delivered files are not blurred: one video file per speaker, H.264 at 720p.
Speaker A · audio, FLAC, 48 kHz, 24-bit, mono
Speaker B · audio, FLAC, 48 kHz, 24-bit, mono
Alignment & verification
Turn-taking
Both speakers' voice activity on the shared timeline: who is talking when, and where they overlap. This window is the first minute and a half of the preview.
Spectrogram
Frequency against time for each speaker's full 15-minute recording, up to 24 kHz.
Alignment
Each track's offset from an independent reference, one point per measurement window across the full recording. Grey spans are windows where no reference was available.
Lip sync
Voice activity from each speaker's audio file and facial motion from their video file, on one time axis. This window is the first minute of the preview, where Speaker B is mostly listening.
Verification
Four figures ship with every conversation.
| Figure | What it shows |
|---|---|
| Turn-taking | Both speakers' voice activity as two lanes on the shared timeline: who is speaking, when, and where they overlap |
| Spectrogram | Energy across frequency and time for each speaker, up to about 24 kHz. Confirms full-band audio with no clipping, dropouts or resampling artefacts. |
| Alignment | Each track's offset from an independent reference, in milliseconds, measured across the whole recording |
| Lip sync | Voice activity and facial motion for each speaker, overlaid on one time axis |
Each file's measured values are in its metadata.json. After the automated checks, a reviewer watches and listens to every conversation on synchronised side-by-side playback.
Pushing the Frontier with the Full-Duplex Audiovisual Dataset
The full-duplex audiovisual dataset is designed to push the frontier of video conversational models. The capabilities frontier companies are pushing to improve are expressiveness, speed, identity preservation, whole-frame generation and listening. For each one, this is what our dataset offers.
Expressiveness. An expressive model is more present in a conversation: its face and body say what words do not. Our dataset contains the real expression of both people, in conversations spoken in their own words.
Speed. A conversation breaks when the reply comes late. Our dataset contains natural backchannels and turn changes on a shared timeline, so a model can learn realistic turn-taking: when people respond, and how fast.
Identity preservation. A person has to stay recognisably themselves as they move. Our dataset records each person for about 15 minutes without a cut, so a model sees one face through many expressions.
Whole-frame generation. A model that generates the head, shoulders and torso has to learn from video that shows them. Our dataset records the full camera frame at 720p on each speaker's own device, not a region around the face.
Listening. A model that goes still when it stops talking has left the conversation. Our dataset records each person on separate streams, so a model can see and hear how people listen, from a nod to an "mm".
Both sides at once. Moshi models two audio streams. A full-duplex video model has four: a voice and a face for each side. Our dataset provides all four on one clock.
Interaction. An interaction model answers while the other person is still going: an interjection, a silence, a nod, a correction. Our dataset leaves both people on camera and on their own microphone for the whole conversation, so those moments stay in the file.
How It Compares
Dataset comparison
The full-duplex audiovisual dataset against the datasets the field trains on.
| Dataset | Shape | Size | Audio | Video | Setting | Licence |
|---|---|---|---|---|---|---|
| Full-duplex audiovisual dataset | Full duplex | Over 1,000 hours | 48 kHz, 24-bit, lossless | 720p or higher | Live video call | On request |
| DyConv | Two people | Over 200 hours | Not stated | Not stated | Collected from the internet | Not stated |
| ViCo | Two people | 1.6 hours | Not stated | Not stated | Not stated | Not stated |
| CelebV-HQ | Single speaker | 35,666 clips | Web video | 512×512 or larger | Web video | Not stated |
| HDTF | Single speaker | About 16 hours | Web video | 720p to 1080p | Web video | Not stated |
| VoxCeleb2 | Single speaker | Over a million utterances, over 6,000 speakers | Web video | Web video | Interviews | Not stated |
Comparator values are taken from each dataset's paper. ViCo and DyConv values are as reported in the INFP paper.
Train Your Model on Full-Duplex Video Data
Building a video conversational model, an interactive avatar, or an interaction model? The useful question is how it behaves when the other person is talking. The full-duplex audiovisual dataset is built for that test. The full specifications, verification figures, and dataset comparison are on the dataset page.
Ocular AI helps teams source real two-person video conversation, with each speaker's audio and video as separate streams on a shared timeline. Share the conversations, formats and annotation layers your model needs, and we can help scope a dataset around them.
Start with samples. Tell us about your model or product through the form, and our team will follow up with sample conversations and next steps. Each one comes with per-speaker audio and video, a metadata file, and its four verification figures.
Request full-duplex audiovisual samples →
Get in touch here to access our off-the-shelf full-duplex conversational, Hi-Fi multilingual, and multi-accent English ASR datasets, or reach out to us directly at research@useocular.com.
Authors

Louis Murerwa
Co-founder & CTO

Michael Moyo
Co-founder & CEO


