220K+ Commercial and Residential Egocentric Dataset
Computer Vision & Video Datasets
Tags and Keywords

£185,500
About
220K+ Hours Commercial & Residential Egocentric Video Dataset
Description
The 220K+ Hours Commercial & Residential Egocentric Video Dataset is a large-scale collection of first-person (egocentric) video data captured across commercial and residential environments. With 220K+ hours of egocentric video, the dataset provides extensive visual information for understanding real-world environments, human activities, object interactions, navigation, and spatial context from a first-person perspective.
The dataset is designed to support Computer Vision, Artificial Intelligence, Machine Learning, Embodied AI, Robotics, Activity Recognition, Object Detection, Scene Understanding, Spatial Intelligence, and Multimodal AI applications. The combination of diverse indoor environments and first-person perspectives makes the dataset valuable for developing and evaluating AI systems that need to interpret real-world activities and environments from the viewpoint of a person.
Commercial environments may include workplaces, offices, retail spaces, facilities, and other business settings, while residential environments may include homes and other living spaces. The dataset can support large-scale AI model training, fine-tuning, evaluation, video understanding, behavior analysis, and real-world environment modeling.
Note: The listed price applies to the specified initial batch of 25,000 hours of egocentric video. Pricing for larger batches or the complete dataset library varies depending on total video duration, video quality and resolution, metadata availability, annotation requirements, number of attributes, scene and task diversity, camera perspective, file formats, licensing terms, and customization needs. Final pricing will be determined based on the specific dataset requirements.
Data Product Features
| Feature | Description |
|---|---|
| Egocentric Video | First-person video captured from a human-centered viewpoint for real-world visual understanding. |
| Commercial Environments | Video representing workplaces, offices, retail environments, facilities, and other commercial settings where available. |
| Residential Environments | Video representing homes and residential environments. |
| Human Activities | Visual representation of activities and interactions occurring within the recorded environments. |
| Object Interactions | First-person views of interactions with objects and environmental elements. |
| Scene Information | Visual information describing indoor scenes, spaces, rooms, and surrounding environments. |
| Spatial Context | First-person spatial information useful for navigation and environment understanding. |
| Temporal Information | Continuous video sequences capturing activities and events over time. |
| Environmental Context | Contextual information about commercial and residential environments. |
Distribution
The dataset is distributed as digital egocentric video files, with video sequences organized according to the available environment, recording, and metadata structure.
-
Format: Video files such as MP4 and MOV formats.
-
Data Type: Egocentric/first-person video.
-
Data Volume: 220K+ hours of video.
-
Environment Types: Commercial and Residential.
-
Delivery: Video files may be organized by environment, sequence, recording session, or other applicable metadata.
-
Dataset Size: The overall file size may vary depending on video resolution, frame rate, codec, compression, metadata, and delivery configuration.
Usage
This data product is ideal for a variety of Computer Vision, AI, Robotics, and Video Understanding applications:
- Egocentric Video Understanding: Train AI models to interpret real-world scenes and activities from a first-person perspective.
- Computer Vision: Develop models for scene understanding, object detection, recognition, and visual perception.
- Activity Recognition: Identify and classify human activities and interactions in commercial and residential environments.
- Object Detection & Recognition: Train models to identify objects and environmental elements in first-person video.
- Embodied AI: Support AI systems that understand and interact with physical environments.
- Robotics: Develop perception systems for robots operating in indoor and human-centered environments.
- Spatial Intelligence: Train models to understand spatial relationships, environments, movement, and navigation.
- Navigation: Support autonomous navigation and environment-aware AI systems.
- Video Analytics: Analyze activities, events, objects, and environmental changes over time.
- Multimodal AI: Combine video and available metadata to develop multimodal AI systems.
- Human-Object Interaction: Study interactions between people, objects, and environments.
- Scene Understanding: Develop models for recognizing and interpreting residential and commercial scenes.
- AI Model Training: Train, fine-tune, and evaluate large-scale video and vision models.
- Research & Benchmarking: Support academic and commercial research in computer vision, video understanding, and embodied intelligence.
Coverage
- Geographic Coverage: Global
- Time Range: Multi-period video collection; exact recording dates may vary by source and collection.
- Environment Coverage: Commercial and residential environments.
- Scene Coverage: Indoor spaces, workplaces, offices, homes, retail environments, facilities, and other applicable environments.
- Activity Coverage: Everyday activities, human-object interactions, movement, navigation, and environment-related activities where represented in the source data.
License
CC BY 4.0 (Creative Commons Attribution 4.0 International)
AI Training Rights
InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.
Data Dictionary
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
file_name | STRING | Name of the video file. | Video filename |
file_format | STRING | Format of the video file. | MP4 and MOV |
duration_sec | FLOAT | Duration of the video in seconds. | Positive numeric value |
duration_hours | FLOAT | Duration of the video in hours. | Positive numeric value |
environment_type | STRING | Classification of the environment represented in the video. | Commercial, Residential |
scene_type | STRING | Type of scene or location represented in the video. | Office, Home, Retail, Facility, etc. |
resolution | STRING | Video resolution. | 720p, 1080p, 4K, etc. |
frame_rate | FLOAT | Number of frames captured per second. | FPS value |
codec | STRING | Video encoding codec. | H.264, H.265, etc. |
camera_perspective | STRING | Perspective from which the video was captured. | Egocentric, First-Person |
Scene_Category | String | Environment classification | Indoor, Outdoor, Office, Kitchen, etc. |
Activity_Label | String | Human activity | Walking, Cooking, Washing, etc. |
Object_Labels | String/Array | Objects present in the scene | Person, Laptop, Cup, Door, Phone, etc. |
Interaction_Label | String | Human-object interaction | Holding, Picking, Opening, Closing, Carrying |
Considerations
This dataset is provided for research and educational purposes only. It contains only sample data.
Additional Notes
- The dataset contains 220K+ hours of commercial and residential egocentric video, providing substantial volume for large-scale AI and computer vision development.
- The first-person perspective makes the dataset particularly relevant to egocentric vision, embodied AI, robotics, activity recognition, spatial intelligence, and real-world video understanding.
- Commercial and residential environments provide varied visual contexts for training models to understand indoor scenes, human activities, objects, interactions, and spatial relationships.
- Dataset size may vary depending on video resolution, frame rate, codec, compression, metadata availability, and the selected delivery batch.
- The listed dataset volume represents 220K+ hours of video; the corresponding storage size may vary based on the technical characteristics of the video files.
Loading...
£185,500
Download Dataset in VIDEO Format
Recommended Datasets
Loading recommendations...
