229K+ Hours Annotated Egocentric Dataset
Generative AI & Computer Vision
Tags and Keywords

£148,500
About
229K+ Hours Annotated Egocentric Dataset
Description
The 229K+ Hours Annotated Egocentric Dataset is a large-scale, structured collection of first-person (egocentric) visual data captured from wearable cameras and enriched with detailed annotations. The dataset is designed to support the development of advanced computer vision, artificial intelligence (AI), machine learning (ML), robotics, and augmented reality (AR) systems that require a human-centric understanding of real-world environments.
The dataset enables AI models to learn from natural human experiences and interactions, making it valuable for applications such as activity recognition, object detection, scene understanding, visual question answering, autonomous agents, assistive technologies, and embodied AI research.
Note: The listed price applies to the specified initial batch of 10,000 hours of annotated egocentric video. Pricing for larger batches or the complete dataset library varies depending on total video duration, video quality and resolution, annotation complexity, number of attributes, metadata availability, scene and task diversity, camera perspective, file formats, licensing terms, and customization needs. Final pricing will be determined based on the specific dataset requirements.
Data Product Features
| Feature | Description |
|---|---|
| Activity Label | Human activity being performed |
| Action Label | Fine-grained action performed by the wearer |
| Object Label | Identified object within the scene |
| Object Category | Classification category of detected objects |
| Bounding Boxes | Object localization coordinates |
| Hand Interaction | Hand-object interaction annotations |
| Human Pose | Body or hand keypoints |
| Scene Category | Indoor, home, kitchen, etc. |
| Environment Context | Environmental conditions and surroundings |
| Occlusion Status | Visibility status of objects |
| Camera Motion | Wearer movement information |
| Frame Resolution | video resolution |
| Duration | Recording duration |
Distribution
- File Formats: MP4
Data Volume
- Data Type: Annotated Egocentric Videos
- Records: Thousands to millions of annotated frames
- Media: Videos
- Volume: 229K+ Hours
- Dataset Size: The dataset size may vary depending on video resolution, duration, frame rate, file formats, compression quality, metadata availability, annotations, and dataset version.
Usage
This data product is ideal for numerous AI and computer vision applications.
- Activity Recognition: Train models to recognize daily human activities.
- Action Recognition: Learn fine-grained human actions from first-person perspectives.
- Object Detection: Detect and classify objects appearing in egocentric scenes.
- Object Tracking: Track objects across video sequences.
- Scene Understanding: Improve environmental awareness for AI systems.
- Hand-Object Interaction Modeling: Understand manipulation and interaction with objects.
- Embodied AI: Train intelligent agents operating in real-world environments.
- Robotics: Enhance robotic perception and task execution.
- Augmented Reality (AR): Build context-aware AR applications.
- Virtual Reality (VR): Improve immersive first-person experiences.
- Assistive Technologies: Develop AI assistants for navigation and accessibility.
- Computer Vision Research: Benchmark detection, segmentation, and recognition algorithms.
- Foundation Model Training: Fine-tune multimodal vision-language models.
- Autonomous Systems: Support perception modules for intelligent systems.
Coverage
Geographic Coverage
Global, covering diverse indoor and outdoor environments where applicable.
Time Range
Multiple recording sessions collected over various time periods
Demographics
May include diverse participants across:
- Various age groups
- Different genders
- Multiple occupations
- Daily living activities
- Household environments
- Workplace environment
License
CC BY 4.0 (Creative Commons Attribution 4.0 International)
AI Training Rights
InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.
Data Dictionary
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
| activity | String | High-level activity | Cooking, Walking, Shopping, Working, etc. |
| action | String | Fine-grained action | Pick, Place, Open, Close, Cut, Hold, Pour, etc. |
| object_name | String | Name of detected object | Cup, Laptop, Phone, Bottle, Chair, etc. |
| object_category | String | Object class | Electronics, Furniture, Food, Tools, etc. |
| bounding_box | Array | Object coordinates | x, y, width, height |
| hand_interaction | Boolean/String | Hand interaction status | Left, Right, Both, None |
| pose_keypoints | JSON | Human pose annotations | Keypoint coordinates |
| scene_type | String | Environment category | Kitchen, , Home, Store, etc. |
| occlusion | String | Visibility level | None, Partial, Full |
| camera_motion | String | Camera movement | Static, Moving |
| file_name | String | Associated media file | Videos |
Considerations
This dataset is provided for research and educational purposes only. It contains only sample data.
Additional Notes
- Optimized for AI, machine learning, computer vision, embodied AI, and robotics workflows.
- Supports object detection, segmentation, action recognition, activity understanding, and multimodal learning tasks.
- Well-suited for foundation model pretraining, benchmarking, transfer learning, and commercial AI product development.
Loading...
£148,500
Download Dataset in VIDEO Format
Recommended Datasets
Loading recommendations...
