220K+ Hours Narrated Egocentric Video Dataset
Computer Vision & Video Datasets
Tags and Keywords

£371,000
About
220K+ Hours Narrated Egocentric Video Dataset
Description
The 220K+ Hours Narrated Egocentric Video Dataset is a large-scale collection of first-person, egocentric video recordings with spoken narration embedded directly within the video. The dataset captures real-world visual experiences together with accompanying spoken descriptions, providing rich audio-visual context for understanding activities, objects, environments, and events from a first-person perspective.
Designed for Computer Vision, Multimodal AI, Video Understanding, Embodied AI, Robotics, and Human Activity Recognition, the dataset supports the development and evaluation of AI models that learn to connect visual observations with spoken information. The combination of egocentric video and embedded narration makes the dataset particularly valuable for video-language learning, activity recognition, scene understanding, object interaction analysis, multimodal model training, and real-world AI systems.
Note: The listed price applies to the specified initial batch of 50,000 hours of narrated egocentric video. Pricing for larger batches or the complete dataset library varies depending on total video duration, video quality and resolution, narration complexity, number of attributes, metadata availability, scene and task diversity, camera perspective, file formats, licensing terms, and customization needs. Final pricing will be determined based on the specific dataset requirements.
Data Product Features
| Feature | Description |
|---|---|
| Egocentric Video | First-person video captured from the viewpoint of the individual, representing real-world visual experiences. |
| Embedded Spoken Narration | Spoken narration is recorded directly within the video, providing auditory context about activities, events, objects, or surroundings. |
| First-Person Perspective | Provides a natural viewpoint for understanding human activities and interactions from the user's visual perspective. |
| Audio-Visual Content | Combines visual video content with embedded spoken audio for multimodal AI development. |
| Human Activities | Captures a broad range of activities, actions, and real-world tasks performed or observed from an egocentric viewpoint. |
| Object Interactions | Provides visual context for interactions between people, objects, and surrounding environments. |
| Scene Understanding | Supports analysis of real-world environments, scenes, objects, and contextual information. |
| Temporal Information | Continuous video sequences support temporal event understanding, action recognition, and activity modeling. |
| Multimodal Context | Enables models to learn relationships between visual observations and spoken narration. |
Distribution
The dataset is distributed as video files containing both visual content and embedded spoken narration.
- Video Format: MP4 and MOV formats
- Audio: Spoken narration embedded within the video
- Data Volume: 220K+ hours of narrated egocentric video
- Perspective: Egocentric / first-person
- Data Size: May vary depending on video resolution, frame rate, codec, compression, audio quality, metadata, and delivery configuration.
Usage
This data product is ideal for a variety of applications:
- Computer Vision: Develop models for visual recognition, scene understanding, and real-world perception.
- Video Understanding: Analyze activities, actions, events, objects, and temporal relationships in video.
- Multimodal AI: Train models that jointly process visual and spoken information.
- Video-Language Learning: Learn relationships between visual events and spoken descriptions within video.
- Human Activity Recognition: Detect and classify human activities from an egocentric viewpoint.
- Action Recognition: Identify actions and events across continuous first-person video sequences.
- Embodied AI: Support AI systems that learn from human-centric visual and auditory experiences.
- Robotics: Develop perception, navigation, activity-understanding, and interaction models.
- Object Interaction Analysis: Understand how people interact with objects in real-world environments.
- Scene Understanding: Analyze environments and contextual information from a first-person perspective.
- Multimodal Model Training: Train, fine-tune, and evaluate audio-visual and video-language models.
- Video Analytics: Support intelligent analysis of real-world first-person video content.
Coverage
- Geographic Coverage: Global.
- Environment Coverage: Real-world environments represented within the underlying recordings, including residential, commercial, workplace, indoor, outdoor, and other settings where available.
- Activity Coverage: Diverse human activities, actions, tasks, object interactions, and events.
- Audio Coverage: Spoken narration embedded within the video recordings.
License
CC BY 4.0 (Creative Commons Attribution 4.0 International)
AI Training Rights
InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.
Data Dictionary
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
file_format | STRING | Video file format | MP4, MOV, AVI, etc. |
duration_sec | FLOAT | Duration of the video in seconds | Non-negative numeric value |
duration_hours | FLOAT | Duration of the video in hours | Derived from duration |
activity | STRING | Activity or action represented in the video | Dataset-specific values |
object_category | STRING | Objects appearing in or interacted with during the video | Dataset-specific categories |
environment_type | STRING | General environment represented in the video | Residential, commercial, workplace, indoor, outdoor, etc. |
scene_type | STRING | Scene or setting represented in the video | Dataset-specific values |
camera_perspective | STRING | Perspective from which the video was captured | Egocentric / First-Person |
resolution | STRING | Video resolution | Dataset-specific |
frame_rate | FLOAT | Video frame rate | Frames per second |
audio_present | BOOLEAN | Indicates whether spoken audio is embedded in the video | Yes / No |
audio_type | STRING | Type of audio content present in the video | Spoken narration, where applicable |
Considerations
This dataset is provided for research and educational purposes only. It contains only sample data.
Additional Notes
- The dataset contains 220K+ hours of narrated egocentric video with spoken narration embedded within the video files.
- The dataset is suitable for Computer Vision, Multimodal AI, Video Understanding, Embodied AI, Robotics, and Human Activity Recognition.
- Dataset volume and technical specifications may vary based on video quality, resolution, frame rate, codec, compression, audio characteristics, metadata, and delivery requirements.
Loading...
£371,000
Download Dataset in VIDEO Format
Recommended Datasets
Loading recommendations...
