Speaker A
Speaker B

Full-Duplex Audiovisual Datasets

Hi-Fi, Studio-Grade Tier

Two-person video conversations in which each speaker's audio and video are recorded as separate streams on one shared clock. Built for the next generation of video conversational models, interactive avatars, and interaction models — the ones that listen and speak at the same time.

Overview

In 2024 speech models went full-duplex: two audio streams, both always running, each belonging to one side of the conversation. In 2026 the same idea has reached video. A model that handles two streams at once has to be trained on two streams at once — and almost all public video is one person facing a camera, or an edit that follows whoever is speaking. This dataset is the third shape: two people recorded separately on a live video call, each voice and each face as its own stream, all four on one clock — so a model can learn speaking, listening, overlap, and the timing between two people.

Key highlights

  • 01

    Full-duplex shape: two people, each recorded separately — each voice and each face as its own stream, all four on one shared clock.

  • 02

    Recorded as a live video call, about 15 minutes per conversation, with no cuts — one face through many expressions.

  • 03

    Both people stay on camera and on their own microphone the whole time, whether speaking or listening, so backchannels, nods, and overlap stay in the file.

  • 04

    Audio per speaker at 48 kHz / 24-bit, delivered as lossless FLAC — the same Hi-Fi capture spec as our speech catalogue.

  • 05

    Video per speaker at 720p or higher, H.264, recording the full camera frame on the speaker's own device — head, shoulders, and torso, not a crop around the face.

  • 06

    Over 1,000 hours of conversation — against 1.6 hours for ViCo and roughly 200 hours for DyConv, the field's full-duplex references.

  • 07

    Four verification figures ship with every conversation: turn-taking, spectrogram, alignment, and lip sync.

  • 08

    Every file's measured values live in its metadata.json; after the automated checks, a reviewer watches and listens to every conversation on synchronised side-by-side playback.

  • 09

    Speaker demographics — gender and race / ethnicity — shipped with each conversation; transcripts, diarization, and custom annotation layers on request.

  • 10

    Commercial licensing with signed consent from paid participants; terms scoped on request.

Sample

One conversation, as delivered

Five minutes of Conversation 1, one video per speaker, played together. Speaker A does most of the talking for the first two minutes; at about 2:15 the roles change and Speaker B takes over. Both people are on camera and on their own microphone throughout, whether they are speaking or listening.

Sample

One conversation, four synchronized streams.

A live video call, with each person on their own camera and microphone.

Speaker A · Male, Asian · video, H.264

Speaker B · Female, two or more races · video, H.264

0:00 / 0:00

Preview: five minutes of the conversation, from 1:00 to 6:00, with faces blurred and the picture scaled down for the web. The delivered files are not blurred: one video file per speaker, H.264 at 720p.

Speaker A · audio, FLAC, 48 kHz, 24-bit, mono

0:00 / 0:00

Speaker B · audio, FLAC, 48 kHz, 24-bit, mono

0:00 / 0:00

Alignment & verification

Turn-taking

Both speakers' voice activity on the shared timeline: who is talking when, and where they overlap. This window is the first minute and a half of the preview.

Spectrogram

Frequency against time for each speaker's full 15-minute recording, up to 24 kHz.

Alignment

Each track's offset from an independent reference, one point per measurement window across the full recording. Grey spans are windows where no reference was available.

Lip sync

Voice activity from each speaker's audio file and facial motion from their video file, on one time axis. This window is the first minute of the preview, where Speaker B is mostly listening.

Request the sample set

Technical specifications

Shape

Each conversation is two people live on a video call, talking for about 15 minutes. Each person is recorded as one audio stream and one video stream, and all four streams sit on a shared timeline. Nobody is cut away from: both people stay on camera and on their own microphone for the whole conversation, whether they are speaking or listening.

Capture specs

Audio is captured per speaker at 48 kHz / 24-bit and delivered as lossless FLAC — the same Hi-Fi spec as our speech catalogue. Video is captured per speaker at 720p or higher and delivered as H.264, recording the full camera frame on the speaker's own device rather than a region around the face, so head, shoulders, and torso are all in the picture.

Verification & metadata

Four figures ship with every conversation — turn-taking, spectrogram, alignment, and lip sync — and each file's measured values are in its metadata.json. After the automated checks, a reviewer watches and listens to every conversation on synchronised side-by-side playback. Speaker demographics ship with each conversation; transcripts, diarization, and custom annotation layers are available on request.

What a model can learn from it

  • Expressiveness — real facial and body expression from both people, in conversations spoken in their own words
  • Speed — natural backchannels and turn changes on a shared timeline, so a model learns when people respond and how fast
  • Identity preservation — about 15 minutes of one face without a cut, through many expressions
  • Whole-frame generation — the full camera frame at 720p, not a region around the face
  • Listening — each person on separate streams, so a model can see and hear how people listen, from a nod to an “mm”
  • Both sides at once — a voice and a face for each side, all four streams on one clock
  • Interaction — both people on camera and microphone throughout, so interjections, silences, and corrections stay in the file

Verification

Four figures ship with every conversation.

Turn-takingBoth speakers' voice activity as two lanes on the shared timeline: who is speaking, when, and where they overlap.
SpectrogramEnergy across frequency and time for each speaker, up to about 24 kHz. Confirms full-band audio with no clipping, dropouts, or resampling artefacts.
AlignmentEach track's offset from an independent reference, in milliseconds, measured across the whole recording.
Lip syncVoice activity and facial motion for each speaker, overlaid on one time axis.

Each file's measured values are in its metadata.json. After the automated checks, a reviewer watches and listens to every conversation on synchronised side-by-side playback.

Comparisons

The datasets the field trains on vs. Ocular Full-Duplex AV

The datasets that taught avatars to talk — VoxCeleb2, HDTF, CelebV-HQ — are one person facing a camera. The two-person references, ViCo and DyConv, are small, or collected from edited web video where the camera follows whoever is speaking. This dataset is the inverse on every axis that matters for a full-duplex model: two people recorded separately, on a live call, at Hi-Fi audio and 720p+ video, at over 1,000 hours, with commercial licensing.

Dataset comparison

The full-duplex audiovisual dataset against the datasets the field trains on.

CategoryAttributeOcular Full-Duplex AVCommercialDyConvINFP, ByteDanceViCoAcademicCelebV-HQAcademicHDTFAcademicVoxCeleb2Academic
ShapeShapeHow many people, and how they are recordedFull duplex — each speaker on separate audio + video streamsTwo peopleTwo peopleSingle speakerSingle speakerSingle speaker
SizeTotal duration or clip countOver 1,000 hoursOver 200 hours1.6 hours35,666 clipsAbout 16 hoursOver a million utterances, 6,000+ speakers
FormatAudioCapture quality of the speech track48 kHz, 24-bit, losslessNot statedNot statedWeb videoWeb videoWeb video
VideoResolution of the picture720p or higherNot statedNot stated512×512 or larger720p to 1080pWeb video
ProvenanceSettingWhere the recordings come fromLive video callCollected from the internetNot statedWeb videoWeb videoInterviews
LicenceCommercial usabilityCommercial — on requestNot statedNot statedNot statedNot statedNot stated

Ocular values shown in bold are the per-row reference. Comparator values are taken from each dataset's paper; ViCo and DyConv values are as reported in the INFP paper.

Dataset information

Every property a buyer asks about before booking a sample, in one datasheet.

NameFull-Duplex Audiovisual Conversational Dataset
TierHi-Fi, Studio-Grade
Data modalities
  • Audio
  • Video
  • Audiovisual
ShapeFull duplex — two people, each recorded separately, each voice and each face as its own stream on one shared clock
Streams per conversation
  • Speaker A — video (MP4, H.264)
  • Speaker A — audio (FLAC, mono)
  • Speaker B — video (MP4, H.264)
  • Speaker B — audio (FLAC, mono)
  • Shared timeline + metadata.json
SizeOver 1,000 hours of conversation
Conversation lengthAbout 15 minutes per conversation, recorded without a cut
SettingLive video call — each person on their own camera and microphone, on their own device
Audio48 kHz / 24-bit per speaker, lossless FLAC
Video720p or higher per speaker, H.264 — full camera frame (head, shoulders, and torso), not a crop around the face
SynchronisationAll four streams on one shared clock; each track's offset measured against an independent reference across the full recording
Verification
  • Turn-taking figure
  • Spectrogram per speaker
  • Alignment per speaker
  • Lip sync per speaker
  • Measured values in metadata.json
  • Human review on synchronised side-by-side playback
Speaker metadataGender and race / ethnicity per speaker
Annotations
  • Word-level transcripts (on request)
  • Diarization and speaker turns (on request)
  • Custom annotation layers (on request)
LicenceCommercial — terms scoped on request

Properties listed here apply to every conversation in the dataset. Additional languages, scenarios, or annotation layers are scoped per engagement.

Building a video conversational model, an interactive avatar, or an interaction model? The useful question is how it behaves when the other person is talking. This dataset is built for that test. Share the conversations, formats, and annotation layers your model needs, and we can help scope a dataset around them — each sample comes with per-speaker audio and video, a metadata file, and its four verification figures.

Request samples

Tell us about your model or product and we'll follow up with sample conversations, pricing, and next steps.

What are you interested in?
How do you plan to use the data?
Ready to bring AI into the real world?