220K+ Hours Narrated Egocentric Video Dataset

Computer Vision & Video Datasets

Tags and Keywords

Egocentric,

Narrated,

Firstperson,

Computervision,

Embodiedai,

Robotics,

Multimodal,

Video

220K+ Hours Narrated Egocentric Video Dataset Dataset on Opendatabay data marketplace

£371,000

About

220K+ Hours Narrated Egocentric Video Dataset

Description

The 220K+ Hours Narrated Egocentric Video Dataset is a large-scale collection of first-person, egocentric video recordings with spoken narration embedded directly within the video. The dataset captures real-world visual experiences together with accompanying spoken descriptions, providing rich audio-visual context for understanding activities, objects, environments, and events from a first-person perspective.
Designed for Computer Vision, Multimodal AI, Video Understanding, Embodied AI, Robotics, and Human Activity Recognition, the dataset supports the development and evaluation of AI models that learn to connect visual observations with spoken information. The combination of egocentric video and embedded narration makes the dataset particularly valuable for video-language learning, activity recognition, scene understanding, object interaction analysis, multimodal model training, and real-world AI systems.
Note: The listed price applies to the specified initial batch of 50,000 hours of narrated egocentric video. Pricing for larger batches or the complete dataset library varies depending on total video duration, video quality and resolution, narration complexity, number of attributes, metadata availability, scene and task diversity, camera perspective, file formats, licensing terms, and customization needs. Final pricing will be determined based on the specific dataset requirements.

Data Product Features

FeatureDescription
Egocentric VideoFirst-person video captured from the viewpoint of the individual, representing real-world visual experiences.
Embedded Spoken NarrationSpoken narration is recorded directly within the video, providing auditory context about activities, events, objects, or surroundings.
First-Person PerspectiveProvides a natural viewpoint for understanding human activities and interactions from the user's visual perspective.
Audio-Visual ContentCombines visual video content with embedded spoken audio for multimodal AI development.
Human ActivitiesCaptures a broad range of activities, actions, and real-world tasks performed or observed from an egocentric viewpoint.
Object InteractionsProvides visual context for interactions between people, objects, and surrounding environments.
Scene UnderstandingSupports analysis of real-world environments, scenes, objects, and contextual information.
Temporal InformationContinuous video sequences support temporal event understanding, action recognition, and activity modeling.
Multimodal ContextEnables models to learn relationships between visual observations and spoken narration.

Distribution

The dataset is distributed as video files containing both visual content and embedded spoken narration.
  • Video Format: MP4 and MOV formats
  • Audio: Spoken narration embedded within the video
  • Data Volume: 220K+ hours of narrated egocentric video
  • Perspective: Egocentric / first-person
  • Data Size: May vary depending on video resolution, frame rate, codec, compression, audio quality, metadata, and delivery configuration.

Usage

This data product is ideal for a variety of applications:
  • Computer Vision: Develop models for visual recognition, scene understanding, and real-world perception.
  • Video Understanding: Analyze activities, actions, events, objects, and temporal relationships in video.
  • Multimodal AI: Train models that jointly process visual and spoken information.
  • Video-Language Learning: Learn relationships between visual events and spoken descriptions within video.
  • Human Activity Recognition: Detect and classify human activities from an egocentric viewpoint.
  • Action Recognition: Identify actions and events across continuous first-person video sequences.
  • Embodied AI: Support AI systems that learn from human-centric visual and auditory experiences.
  • Robotics: Develop perception, navigation, activity-understanding, and interaction models.
  • Object Interaction Analysis: Understand how people interact with objects in real-world environments.
  • Scene Understanding: Analyze environments and contextual information from a first-person perspective.
  • Multimodal Model Training: Train, fine-tune, and evaluate audio-visual and video-language models.
  • Video Analytics: Support intelligent analysis of real-world first-person video content.

Coverage

  • Geographic Coverage: Global.
  • Environment Coverage: Real-world environments represented within the underlying recordings, including residential, commercial, workplace, indoor, outdoor, and other settings where available.
  • Activity Coverage: Diverse human activities, actions, tasks, object interactions, and events.
  • Audio Coverage: Spoken narration embedded within the video recordings.

License

CC BY 4.0 (Creative Commons Attribution 4.0 International)

AI Training Rights

InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.

Data Dictionary

Column NameData TypeDescriptionPossible Values/Notes
file_formatSTRINGVideo file formatMP4, MOV, AVI, etc.
duration_secFLOATDuration of the video in secondsNon-negative numeric value
duration_hoursFLOATDuration of the video in hoursDerived from duration
activitySTRINGActivity or action represented in the videoDataset-specific values
object_categorySTRINGObjects appearing in or interacted with during the videoDataset-specific categories
environment_typeSTRINGGeneral environment represented in the videoResidential, commercial, workplace, indoor, outdoor, etc.
scene_typeSTRINGScene or setting represented in the videoDataset-specific values
camera_perspectiveSTRINGPerspective from which the video was capturedEgocentric / First-Person
resolutionSTRINGVideo resolutionDataset-specific
frame_rateFLOATVideo frame rateFrames per second
audio_presentBOOLEANIndicates whether spoken audio is embedded in the videoYes / No
audio_typeSTRINGType of audio content present in the videoSpoken narration, where applicable

Considerations

This dataset is provided for research and educational purposes only. It contains only sample data.

Additional Notes

  • The dataset contains 220K+ hours of narrated egocentric video with spoken narration embedded within the video files.
  • The dataset is suitable for Computer Vision, Multimodal AI, Video Understanding, Embodied AI, Robotics, and Human Activity Recognition.
  • Dataset volume and technical specifications may vary based on video quality, resolution, frame rate, codec, compression, audio characteristics, metadata, and delivery requirements.

Listing Stats

VIEWS

3

DELIVERY

CUSTOM, S3

LISTED

15/09/2026

UPDATED

16/09/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£371,000

Download Dataset in VIDEO Format