Full-Duplex Audiovisual Datasets
Hi-Fi, Studio-Grade TierTwo-person video conversations in which each speaker's audio and video are recorded as separate streams on one shared clock. Built for the next generation of video conversational models, interactive avatars, and interaction models — the ones that listen and speak at the same time.
Overview
In 2024 speech models went full-duplex: two audio streams, both always running, each belonging to one side of the conversation. In 2026 the same idea has reached video. A model that handles two streams at once has to be trained on two streams at once — and almost all public video is one person facing a camera, or an edit that follows whoever is speaking. This dataset is the third shape: two people recorded separately on a live video call, each voice and each face as its own stream, all four on one clock — so a model can learn speaking, listening, overlap, and the timing between two people.
Key highlights
Sample
One conversation, as delivered
Five minutes of Conversation 1, one video per speaker, played together. Speaker A does most of the talking for the first two minutes; at about 2:15 the roles change and Speaker B takes over. Both people are on camera and on their own microphone throughout, whether they are speaking or listening.
Sample
One conversation, four synchronized streams.
A live video call, with each person on their own camera and microphone.
Speaker A · Male, Asian · video, H.264
Speaker B · Female, two or more races · video, H.264
Preview: five minutes of the conversation, from 1:00 to 6:00, with faces blurred and the picture scaled down for the web. The delivered files are not blurred: one video file per speaker, H.264 at 720p.
Speaker A · audio, FLAC, 48 kHz, 24-bit, mono
Speaker B · audio, FLAC, 48 kHz, 24-bit, mono
Alignment & verification
Turn-taking
Both speakers' voice activity on the shared timeline: who is talking when, and where they overlap. This window is the first minute and a half of the preview.
Spectrogram
Frequency against time for each speaker's full 15-minute recording, up to 24 kHz.
Alignment
Each track's offset from an independent reference, one point per measurement window across the full recording. Grey spans are windows where no reference was available.
Lip sync
Voice activity from each speaker's audio file and facial motion from their video file, on one time axis. This window is the first minute of the preview, where Speaker B is mostly listening.
Technical specifications
Shape
Each conversation is two people live on a video call, talking for about 15 minutes. Each person is recorded as one audio stream and one video stream, and all four streams sit on a shared timeline. Nobody is cut away from: both people stay on camera and on their own microphone for the whole conversation, whether they are speaking or listening.
Capture specs
Audio is captured per speaker at 48 kHz / 24-bit and delivered as lossless FLAC — the same Hi-Fi spec as our speech catalogue. Video is captured per speaker at 720p or higher and delivered as H.264, recording the full camera frame on the speaker's own device rather than a region around the face, so head, shoulders, and torso are all in the picture.
Verification & metadata
Four figures ship with every conversation — turn-taking, spectrogram, alignment, and lip sync — and each file's measured values are in its metadata.json. After the automated checks, a reviewer watches and listens to every conversation on synchronised side-by-side playback. Speaker demographics ship with each conversation; transcripts, diarization, and custom annotation layers are available on request.
What a model can learn from it
- Expressiveness — real facial and body expression from both people, in conversations spoken in their own words
- Speed — natural backchannels and turn changes on a shared timeline, so a model learns when people respond and how fast
- Identity preservation — about 15 minutes of one face without a cut, through many expressions
- Whole-frame generation — the full camera frame at 720p, not a region around the face
- Listening — each person on separate streams, so a model can see and hear how people listen, from a nod to an “mm”
- Both sides at once — a voice and a face for each side, all four streams on one clock
- Interaction — both people on camera and microphone throughout, so interjections, silences, and corrections stay in the file
Verification
Four figures ship with every conversation.
| Turn-taking | Both speakers' voice activity as two lanes on the shared timeline: who is speaking, when, and where they overlap. |
|---|---|
| Spectrogram | Energy across frequency and time for each speaker, up to about 24 kHz. Confirms full-band audio with no clipping, dropouts, or resampling artefacts. |
| Alignment | Each track's offset from an independent reference, in milliseconds, measured across the whole recording. |
| Lip sync | Voice activity and facial motion for each speaker, overlaid on one time axis. |
Each file's measured values are in its metadata.json. After the automated checks, a reviewer watches and listens to every conversation on synchronised side-by-side playback.
Comparisons
The datasets the field trains on vs. Ocular Full-Duplex AV
The datasets that taught avatars to talk — VoxCeleb2, HDTF, CelebV-HQ — are one person facing a camera. The two-person references, ViCo and DyConv, are small, or collected from edited web video where the camera follows whoever is speaking. This dataset is the inverse on every axis that matters for a full-duplex model: two people recorded separately, on a live call, at Hi-Fi audio and 720p+ video, at over 1,000 hours, with commercial licensing.
Dataset comparison
The full-duplex audiovisual dataset against the datasets the field trains on.
| Category | Attribute | Ocular Full-Duplex AVCommercial | DyConvINFP, ByteDance | ViCoAcademic | CelebV-HQAcademic | HDTFAcademic | VoxCeleb2Academic |
|---|---|---|---|---|---|---|---|
| Shape | ShapeHow many people, and how they are recorded | Full duplex — each speaker on separate audio + video streams | Two people | Two people | Single speaker | Single speaker | Single speaker |
| SizeTotal duration or clip count | Over 1,000 hours | Over 200 hours | 1.6 hours | 35,666 clips | About 16 hours | Over a million utterances, 6,000+ speakers | |
| Format | AudioCapture quality of the speech track | 48 kHz, 24-bit, lossless | Not stated | Not stated | Web video | Web video | Web video |
| VideoResolution of the picture | 720p or higher | Not stated | Not stated | 512×512 or larger | 720p to 1080p | Web video | |
| Provenance | SettingWhere the recordings come from | Live video call | Collected from the internet | Not stated | Web video | Web video | Interviews |
| LicenceCommercial usability | Commercial — on request | Not stated | Not stated | Not stated | Not stated | Not stated |
Ocular values shown in bold are the per-row reference. Comparator values are taken from each dataset's paper; ViCo and DyConv values are as reported in the INFP paper.
Dataset information
Every property a buyer asks about before booking a sample, in one datasheet.
| Name | Full-Duplex Audiovisual Conversational Dataset |
|---|---|
| Tier | Hi-Fi, Studio-Grade |
| Data modalities |
|
| Shape | Full duplex — two people, each recorded separately, each voice and each face as its own stream on one shared clock |
| Streams per conversation |
|
| Size | Over 1,000 hours of conversation |
| Conversation length | About 15 minutes per conversation, recorded without a cut |
| Setting | Live video call — each person on their own camera and microphone, on their own device |
| Audio | 48 kHz / 24-bit per speaker, lossless FLAC |
| Video | 720p or higher per speaker, H.264 — full camera frame (head, shoulders, and torso), not a crop around the face |
| Synchronisation | All four streams on one shared clock; each track's offset measured against an independent reference across the full recording |
| Verification |
|
| Speaker metadata | Gender and race / ethnicity per speaker |
| Annotations |
|
| Licence | Commercial — terms scoped on request |
Properties listed here apply to every conversation in the dataset. Additional languages, scenarios, or annotation layers are scoped per engagement.
Building a video conversational model, an interactive avatar, or an interaction model? The useful question is how it behaves when the other person is talking. This dataset is built for that test. Share the conversations, formats, and annotation layers your model needs, and we can help scope a dataset around them — each sample comes with per-speaker audio and video, a metadata file, and its four verification figures.
Request samples