Wreck-7K: 7,000-Hour Real-World Human–Object Interaction Video
Egocentric & First-Person Data
Tags and Keywords

£55,000
About
Wreck-7K: 7,000-Hour Real-World Human–Object Interaction Video Archive
Wreck-7K is a proprietary video dataset containing approximately 7,000 hours of third-person, or exocentric, footage captured inside an operating physical entertainment facility in Las Vegas, Nevada.
The archive documents natural human–object interaction across thousands of real customer sessions. Recorded activities include object handling, carrying, tool selection, striking, impact, deformation, breakage, and debris movement.
Unlike datasets collected through tightly scripted laboratory demonstrations, Wreck Room Exo-7K captures natural variation in human movement, participant behavior, object selection, action intensity, scene composition, occlusion, debris patterns, and object-state transitions.
The dataset is designed for organizations developing video foundation models, computer-vision systems, vision-language models, multimodal models, action-recognition systems, physical-world models, and embodied-AI applications.
Data Product Features
The primary data product consists of raw video assets accompanied by a machine-readable file manifest.
Key features may include:
- Third-Person Video: Exocentric recordings showing participants, objects, tools, and the surrounding environment.
- Natural Human Behavior: Unscripted activity recorded during real customer sessions rather than staged laboratory demonstrations.
- Human–Object Interaction: Repeated examples of grasping, carrying, positioning, striking, sweeping, shoveling, sorting, and collecting objects.
- Tool Use: Participants interact with a variety of handheld tools and physical objects.
- Impact and Breakage Events: Video includes object impacts, deformation, fragmentation, breakage, and changes in object condition.
- Object-State Transitions: Objects and scenes change from intact or organized states to damaged, fragmented, displaced, or collected states.
- Scene Transformation: The environment changes throughout each session as objects are moved, broken, scattered, and removed.
- Multi-Person Activity: Some sessions include multiple participants interacting within the same environment.
- Variable Occlusion: Participants, objects, tools, and debris may partially or fully occlude one another during activity.
- Long-Form Temporal Context: Continuous video captures action sequences and scene evolution rather than only isolated clips.
- Real-World Variation: The archive includes variation in movement style, participant height, clothing, object type, tool choice, action speed, and interaction intensity.
- Commercial Licensing: Qualified footage is available for commercial machine-learning training, evaluation, and model development.
Distribution
The dataset is delivered as raw video files with an accompanying CSV or JSON manifest.
-
Primary Format: Video files in their original or standardized delivery format
-
Video Codec: H.264
-
Resolution: 1920X1080
-
Frame Rate: 30 FPS
-
Audio: Removed
-
Manifest Format: CSV and/or JSON
-
Delivery Method: Secure cloud transfer, encrypted storage device, or another mutually agreed delivery method
-
Folder Structure: Organized by session, camera, date, or asset identifier
-
Annotations: Raw video is the standard deliverable. Additional segmentation, tagging, transcription, object detection, face blurring, audio removal, or metadata enrichment can be quoted separately.
-
Data Volume: Approximately 7,000 hours of third-person video
-
Number of Sessions: Thousands of individual customer sessions
-
Number of Video Assets: Dependent on the final licensed subset and delivery configuration
-
Manifest Records: One record per delivered video asset
-
Estimated Storage Size: 12TB
-
Ongoing Availability: Additional footage may be added through continued capture operations
Buyers may license the full archive or request a curated subset based on duration, activity type, object category, technical quality, date range, or privacy requirements.
Usage
This data product is ideal for a variety of applications:
- Video Foundation Models: Large-scale pretraining and evaluation using real-world human activity and physical interactions.
- Vision-Language Models: Training systems to connect video content with descriptions of actions, objects, events, and scene changes.
- Action Recognition: Identification and classification of carrying, striking, sweeping, shoveling, sorting, cleanup, and other physical actions.
- Temporal Action Segmentation: Detecting when actions begin, change, and end within long-form video.
- Human–Object Interaction Recognition: Modeling relationships among people, tools, objects, and the surrounding environment.
- Object-State Recognition: Identifying transitions such as intact to damaged, organized to scattered, or debris to collected.
- Physical-World Understanding: Learning how objects move, deform, break, scatter, and interact after physical contact.
- World-Model Development: Training or evaluating models that predict future visual states following human actions and object interactions.
- Scene Understanding: Analyzing changing layouts, clutter, irregular objects, occlusion, debris, and multi-person activity.
- Tool-Use Recognition: Detecting tools, tool selection, tool handling, and resulting interactions with target objects.
- Event Detection: Identifying impacts, breakage events, debris movement, cleanup actions, and other meaningful moments.
- Synthetic Video Evaluation: Comparing generated video, simulated physics, and synthetic activity against real-world examples.
- Robotics Visual Pretraining: Supporting affordance learning, activity understanding, visual representation learning, and human-to-robot transfer research.
- Safety and Behavior Analysis: Studying movement patterns, proximity, tool use, high-motion activity, and interactions in cluttered environments.
This dataset does not include robot joint states, force measurements, depth sensing, motion-capture trajectories, or direct robot-control actions. It is therefore best suited to visual learning, physical-world understanding, action recognition, model pretraining, and evaluation rather than direct robot-control supervision.
Coverage
- Geographic Coverage: 60% tourists in Las Vegas, Nevada, United States
- Capture Environment: Indoor operating physical entertainment facility
- Time Range: 2018 – Ongoing
- Session Coverage: 4000+ real customer sessions
- Participant Coverage: Thousands of participants across a wide range of body types, movement styles, clothing, group sizes, and activity patterns
- Age Coverage: Primarily adults. Exact demographic labels are not included unless separately documented and made available.
- Gender Coverage: Participants of multiple genders appear in the footage. Gender labels are not included as structured metadata.
- Environmental Coverage: Variable object configurations, debris levels, lighting conditions, participant positions, action intensity, and scene states within the same operating facility
The dataset is geographically concentrated in one facility and should not be represented as globally distributed footage.
License
Proprietary
The dataset is available under a negotiated commercial data-license agreement. Licensing options may include:
- Evaluation access
- Curated hourly subsets
- Full non-exclusive commercial licensing
- Field-limited licensing
- Internal research and model-development licensing
- Ongoing data delivery
- Custom privacy processing
- Custom annotation and enrichment
Only footage with verified rights applicable to the agreed use will be included in the licensed delivery. Supporting provenance and participant-rights documentation may be made available for diligence under NDA.
AI Training Rights
Subject to the final executed license agreement, the licensee may be granted a non-exclusive, worldwide, and perpetual right to:
- Use the Data Product to train, fine-tune, test, validate, and evaluate machine-learning models, including video models, computer-vision models, vision-language models, multimodal foundation models, and embodied-AI systems.
- Incorporate information learned from the Data Product into commercial models and commercialize resulting model outputs.
- Create permitted derivative works, including model weights, embeddings, representations, labels, and analytical outputs.
- Use the Data Product for internal research, model development, benchmarking, and evaluation.
Restrictions:
- The raw Data Product may not be sold, redistributed, sublicensed, published, or shared outside the licensed organization except as expressly authorized.
- The Data Product may not be used to identify, contact, track, or profile individual participants.
- The licensee must comply with all applicable privacy, data-protection, intellectual-property, and machine-learning regulations.
- The licensee must maintain reasonable technical and organizational safeguards protecting the Data Product.
- Any permitted use remains subject to the final commercial license agreement.
Who Can Use It
- AI and Machine-Learning Companies: For training video, vision-language, multimodal, physical-world, and foundation models.
- Computer-Vision Teams: For action recognition, human–object interaction detection, object-state recognition, and scene understanding.
- Robotics Companies: For visual pretraining, affordance learning, task understanding, and human-demonstration research.
- Research Organizations: For academic or commercial research involving video understanding, physical interaction, human behavior, and temporal modeling.
- Simulation and Synthetic-Data Companies: For evaluating generated motion, object interactions, impact dynamics, breakage, and scene transformation.
- Video-Generation Companies: For training and evaluating models that generate or predict realistic human actions and physical-world events.
- Safety and Behavior Researchers: For analyzing tool use, movement, proximity, impacts, and activity in cluttered environments.
- Data Aggregators and Licensing Platforms: For authorized licensing and delivery to qualified enterprise customers, subject to contractual restrictions.
Data Dictionary
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
asset_id | String | Unique identifier assigned to each delivered media asset | Alphanumeric identifier generated during delivery preparation |
relative_file_path | String | File location within the delivered dataset folder structure | Internal absolute file paths are removed before delivery |
top_level_batch | String | Source export folder or recording batch associated with the asset | Example: Dav Files 8 5 22; may be standardized into a batch identifier |
camera_location | String | Physical camera location parsed from the filename or source folder | Room 1, Room 2, Room 3, Room 4, Warehouse |
camera_channel | Integer/Null | Recorder channel inferred from the folder suffix or filename structure | Example: 4; older folders may not contain a channel value |
stream_type | String | Video stream variant associated with the asset | Usually main; limited files may contain sub1 or sub |
capture_timestamp_raw | String | Original 14-digit recording timestamp parsed from the filename | Format: YYYYMMDDHHMMSS |
capture_date | Date | Calendar date parsed from the original recording timestamp | Format: YYYY-MM-DD |
capture_time | Time | Time of day parsed from the original recording timestamp | Format: HH:MM:SS |
file_extension | String | Original file extension | .dav, .jpg, .mp4, .asf |
media_type | String | General type of media represented by the asset | video, still_image |
file_size_bytes | Integer | Size of the media file in bytes | Positive integer; zero-byte files are separately flagged |
source_last_modified_at | Datetime | Filesystem modification timestamp from the source archive | May reflect copying or export activity rather than the original capture time |
has_companion_still | Boolean | Indicates whether a still image with the same base filename exists | true, false |
has_companion_video | Boolean | Indicates whether a corresponding video exists for a still-image asset | true, false |
image_width_px | Integer/Null | Width of a still image in pixels | Sampled JPG files were 2688 pixels wide |
image_height_px | Integer/Null | Height of a still image in pixels | Sampled JPG files were 1520 pixels high |
zero_byte_flag | Boolean | Indicates that the source file exists but contains zero bytes of data | true, false; zero-byte files are excluded from licensed delivery |
delivery_status | String | Indicates whether the asset passed technical and licensing review | approved, review_required, excluded |
rights_status | String | Indicates the applicable licensing or rights-verification status | Example: verified_for_delivery |
privacy_processing | String | Identifies privacy processing applied before delivery | none, audio_removed, face_blurred |
notes | String | Additional technical, content, or quality-control notes | Free-text field |
Additional Notes
The standard product is a large-scale raw-video archive rather than a fully human-annotated benchmark dataset.
Activity labels, object categories, participant counts, and other enriched metadata may not be available for every asset unless specifically included in the licensed scope. Buyers requiring a targeted subset should provide their desired activities, object types, video quality requirements, duration, privacy requirements, and intended use.
Available optional services may include:
- Curated subset creation
- Scene and action segmentation
- Object and activity tagging
- Automated object detection
- Tracking or segmentation masks
- Speech transcription
- Face blurring
- Audio removal
- Clip extraction
- Quality filtering
- Custom manifest development
- Ongoing directed data capture
Wreck Room also offers custom first-person, third-person, and synchronized multi-camera capture. Customers may define the requested objects, actions, tools, camera placements, participant instructions, metadata, privacy processing, and delivery requirements.
Listing Stats
VIEWS
21
DELIVERY
CUSTOM, S3
LISTED
11/07/2026
UPDATED
12/07/2026
REGION
NORTH AMERICA
TRUST
5 / 5
Loading...
£55,000
Download Dataset in JSON Format
Recommended Datasets
Loading recommendations...
