Commercial Source Code & Engineering Data
Source Code & Programming Data
Tags and Keywords

£200,000
About
A large-scale, de-identified corpus of commercial source code and engineering artefacts sourced from 240+ private repositories. Covers full git histories, merged pull requests with developer-written change descriptions, and a rich quality/profiling layer — purpose-built for code LLM training, SWE-bench-style evaluation, and diff-to-intent research.
Data Product Features
- 240+ private commercial repositories (190+ web/backend, 50+ iOS/Android)
- 300,000+ commits with complete git history
- ~42M first-party lines of code, post de-identification
- 44,000+ merged MRs with developer-written change descriptions (36,000+ substantive, ~2.3M tokens)
- 35-field repository profiler + 102-field quality layer
- 9,000+ mined SWE-bench-style task candidates
- Full de-identification across commit metadata, file contents, and file paths
Distribution
Format: Git / NDJSON / CSV
Size: ~463M tokens
Records: 240+ repos · 44,000+ MRs
Data Volume
300,000+ commits, 44,000+ merged MRs, ~42M lines of code
Usage
- Code LLM pre-training and fine-tuning on real commercial codebases
- Diff-to-intent model training using MR descriptions as supervision signal
- SWE-bench-style evaluation dataset mining
- Repository-level code quality analysis and profiling
Coverage
Geographic Coverage: Global (India-origin commercial repos)
Time Range: 2013–2026
Languages: JavaScript, TypeScript, PHP, Python, Swift, Kotlin
License
CC0 — No Rights Reserved
AI Training Rights
Licensee is granted a non-exclusive, worldwide, and perpetual right to use this data product to train, fine-tune, and evaluate machine learning models. The data product itself may not be redistributed or shared outside licensed usage. Licensee must comply with all applicable laws, including data protection and privacy regulations.
Who Can Use It
- AI/ML Engineers: For training code generation and completion models
- Researchers: For studying software engineering practices at scale
- Enterprises: For building internal developer productivity tools
Data Dictionary
- repo_id (string) — De-identified repository identifier
- commit_hash (string) — Git commit SHA, 40-char hex
- commit_message (string) — Developer-written commit message, free text
- diff (string) — Code diff per commit, unified diff format
- mr_description (string) — Merged MR change description, free text
- language (string) — Primary programming language: JS, TS, PHP, Python, Swift, Kotlin
- quality_score (float) — 102-field quality layer score, range 0.0–1.0
Listing Stats
VIEWS
11
DELIVERY
SPECIFIED IN DESCRIPTION
LISTED
08/09/2026
UPDATED
12/09/2026
REGION
GLOBAL
TRUST
5 / 5
Loading...
£200,000
Download Dataset in CODE Format
Recommended Datasets
Loading recommendations...
