Egocentric First-Person Audio Dataset
Egocentric & First-Person Data
Tags and Keywords

£150,000
About
A large-scale first-person perspective audio-video corpus capturing real-world human activities, conversations, and ambient acoustic conditions from a wearable point of view. Designed for embodied AI, multimodal model training, and acoustic scene understanding — one of the few egocentric audio-video datasets available at scale.
Data Product Features
- First-person perspective audio and video from wearable recording setups
- Captures real-world human activities and daily interactions
- Ambient acoustic conditions across diverse environments
- Suitable for multimodal pairing with sensor or IMU data
- Large-scale collection with multilingual speaker coverage
Distribution
Format: Audio + Video + Metadata
Size: Large-scale
Records: Large-scale collection
Data Volume
Large-scale egocentric audio-video corpus — exact volume arranged per project scope
Usage
- Embodied AI model training with first-person audio-visual context
- Acoustic scene understanding and classification
- Multimodal model training paired with IMU or sensor data
- Wearable AI and ambient intelligence research
Coverage
Geographic Coverage: Global
Time Range: Historical
Languages: Multilingual
License
CC0 — No Rights Reserved
AI Training Rights
Licensee is granted a non-exclusive, worldwide, and perpetual right to use this data product to train, fine-tune, and evaluate machine learning models. The data product itself may not be redistributed or shared outside licensed usage. Licensee must comply with all applicable laws, including data protection and privacy regulations.
Who Can Use It
- Robotics Researchers: For training embodied agents with first-person audio-visual perception
- AI/ML Engineers: For multimodal and acoustic scene models
- Wearable Tech Companies: For ambient sound and vision understanding applications
Data Dictionary
- clip_id (string) — Unique identifier for each audio-video clip
- duration_seconds (float) — Length of clip in seconds
- activity_label (string) — Annotated human activity during recording
- environment (string) — Indoor, outdoor, or mixed environment
- language (string) — Language of speech captured, if any
- recording_device (string) — Wearable device type used
- modality (string) — Audio, Video, or Both
- metadata (object) — Additional contextual annotations
Loading...
£150,000
Download Dataset in VIDEO Format
Recommended Datasets
Loading recommendations...
