Commercial Source Code & Engineering Data

Source Code & Programming Data

Tags and Keywords

Sourcecode

Githistory

Llmtraining

Codegeneration

Pullrequests

Softwareengineering

De-identified

Commercialrepositories

Commercial Source Code & Engineering Data Dataset on Opendatabay data marketplace

£200,000

About

A large-scale, de-identified corpus of commercial source code and engineering artefacts sourced from 240+ private repositories. Covers full git histories, merged pull requests with developer-written change descriptions, and a rich quality/profiling layer — purpose-built for code LLM training, SWE-bench-style evaluation, and diff-to-intent research.
Data Product Features
  • 240+ private commercial repositories (190+ web/backend, 50+ iOS/Android)
  • 300,000+ commits with complete git history
  • ~42M first-party lines of code, post de-identification
  • 44,000+ merged MRs with developer-written change descriptions (36,000+ substantive, ~2.3M tokens)
  • 35-field repository profiler + 102-field quality layer
  • 9,000+ mined SWE-bench-style task candidates
  • Full de-identification across commit metadata, file contents, and file paths
Distribution Format: Git / NDJSON / CSV Size: ~463M tokens Records: 240+ repos · 44,000+ MRs
Data Volume 300,000+ commits, 44,000+ merged MRs, ~42M lines of code
Usage
  • Code LLM pre-training and fine-tuning on real commercial codebases
  • Diff-to-intent model training using MR descriptions as supervision signal
  • SWE-bench-style evaluation dataset mining
  • Repository-level code quality analysis and profiling
Coverage Geographic Coverage: Global (India-origin commercial repos) Time Range: 2013–2026 Languages: JavaScript, TypeScript, PHP, Python, Swift, Kotlin
License CC0 — No Rights Reserved
AI Training Rights Licensee is granted a non-exclusive, worldwide, and perpetual right to use this data product to train, fine-tune, and evaluate machine learning models. The data product itself may not be redistributed or shared outside licensed usage. Licensee must comply with all applicable laws, including data protection and privacy regulations.
Who Can Use It
  • AI/ML Engineers: For training code generation and completion models
  • Researchers: For studying software engineering practices at scale
  • Enterprises: For building internal developer productivity tools
Data Dictionary
  • repo_id (string) — De-identified repository identifier
  • commit_hash (string) — Git commit SHA, 40-char hex
  • commit_message (string) — Developer-written commit message, free text
  • diff (string) — Code diff per commit, unified diff format
  • mr_description (string) — Merged MR change description, free text
  • language (string) — Primary programming language: JS, TS, PHP, Python, Swift, Kotlin
  • quality_score (float) — 102-field quality layer score, range 0.0–1.0

Listing Stats

VIEWS

11

DELIVERY

SPECIFIED IN DESCRIPTION

LISTED

08/09/2026

UPDATED

12/09/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£200,000

Download Dataset in CODE Format