Data-Analysis-Platform
Data Science and Analytics
Tags and Keywords

£50,000
About
In context, this represents the kind of real-world, multi-service enterprise engineering history covering the full lifecycle of a data platform build-out — from initial scaffolding through infrastructure provisioning, API development, governance tooling, and testing.
Useful for research into software engineering productivity and patterns, repository-scale benchmarking, or as training/reference material for tools that need to reason about realistic, multi-repository enterprise development history rather than isolated code snippets.
Data Product Features
Multi-domain technical coverage — the codebase spans Terraform/infrastructure-as-code, Python backend services, IAM/Cognito configuration, SQL data migrations, and Postman/Newman API test suites, giving breadth across the typical layers of a real platform rather than a single language or service type.Quantifiable scale — roughly 108,000 lines of code, 1,474 tracked files, and ~1.1M tokens of text content, already computed and ready to report as listing specifications.
Distribution
The data product is delivered as a native Git repository — full version-controlled history, not a flat export. It can be cloned and browsed with standard git tooling (git log, git blame, git diff all work normally), or, if a marketplace requires a static download instead, exported as commit-level and file-level metric tables in CSV/JSON.
- **Data Volume: Working tree (actual files, no git metadata): 6.0 MB Git object database (.git, full commit history): 4.5 MB Combined: ~10.5 MB total 5,683 commits, 1,474 tracked files, ~108,761 lines of code, ~317,744 words, ~4.41M characters (~1.1M tokens estimated)
Usage
This data product is ideal for a variety of applications:
Application: Software engineering research — studying real commit cadence, sequencing, and team behavior across a multi-service platform build, rather than relying on synthetic or single-repo samples.
Application: LLM and AI training data — fine-tuning or evaluating code-and-commit-understanding models on authentic enterprise development history, including realistic commit messages and diff patterns.
Application: Repository sizing and cost estimation benchmarking — using the measured LOC, file counts, and token volumes as reference points for scoping similar platform engagements.
Application: DevOps and CI/CD pattern analysis — examining how infrastructure-as-code (Terraform), IAM configuration, and application code evolved together across interdependent services.
Application: Codebase migration and monorepo case studies — as a worked example of consolidating multiple independently-versioned repositories into a single repository while preserving full commit history.
Application: Developer productivity and engineering-velocity modeling — commit frequency, file-churn, and cross-service coordination patterns over a defined ~3-month delivery window.
Application: Training material for engineering teams — a realistic reference codebase (data platform: API services, IAM, governance, curation, infrastructure) for onboarding, tooling demos, or internal workshops.
Coverage
Geographic Coverage: Development-side coverage is India — every commit in the repository carries an IST (UTC+05:30) author/committer timestamp, indicating an India-based engineering team. The platform itself was built for a global healthcare organization's cloud infrastructure (AWS-hosted), so the deployment target is global/US-centric even though the development activity captured here is India-based.
Time Range: Earliest recorded commit: 2023-02-13 (data, infrastructure, client, core-api, curation). Latest recorded commit: 2025-01-28 (dm-producer-mvp). So the dataset spans roughly February 2023 – January 2025 (~23 months) in total, though individual projects have different active windows within that span
License
CC0
AI Training Rights
Licensee is granted a non-exclusive, worldwide, and perpetual right to:
- Use the Data Product to train, fine-tune, and evaluate machine learning models, including large language models.
- Incorporate Data Product content into models and commercialize resulting model outputs.
- Create derivative works (model weights, embeddings, etc.) for any lawful purpose.
Restrictions:
- The Data Product itself may not be sold, redistributed, or shared outside of licensed usage.
- Licensee must comply with all applicable laws, including data protection and privacy regulations.
Who Can Use It
Data Scientists / ML Engineers: For training or evaluating code-and-commit-understanding models — fine-tuning LLMs on realistic commit sequencing, diff patterns, and multi-service coordination rather than synthetic code samples.
Software Engineering Researchers: For academic or applied studies on developer productivity, commit cadence, monorepo consolidation patterns, or how infrastructure-as-code and application code co-evolve across interdependent services over a real ~2-year delivery window.
DevOps / Platform Engineering Teams: For benchmarking — using the measured LOC, file counts, commit volume, and project structure as reference points when scoping, estimating, or comparing similar data-platform engagements.
Engineering Managers / Consultancies: For internal training material — a realistic multi-service reference codebase (API services, IAM, data governance, curation, Terraform infrastructure) useful for onboarding, tooling demonstrations, or workshops without needing a live production system.
AI/Code-Tooling Businesses: For product development and testing — evaluating code search, static analysis, repository-migration tooling, or git-history-analysis products against a real, non-trivial monorepo rather than toy examples.
Data Dictionary
The primary deliverable is a git repository, not inherently tabular — but for buyers who want a tabular view, two derived summary tables can be exported (commit-level and file-level). Here's the dictionary for both:
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
| repo_name | string | Project directory the commit belongs to | One of: client, core-api, curation, data, dm-central-governance, dm-domain-mvp, dm-producer-mvp, health-data-platform-iam-role, infrastructure |
| commit_hash | string (40 chars) | Git SHA-1 commit hash within the monorepo | Hexadecimal string |
| author_name | string | Commit author display name | Free text; individual contributor names present in git history |
| author_email | string | Commit author email address | Free text; sanitized to remove client-domain addresses |
| commit_date | datetime (ISO 8601) | Author/committer timestamp | Range 2023-02-13 to 2025-01-28; all values carry +05:30 (IST) offset |
| commit_message | string | Commit message (subject line) | Free text, typically <100 chars |
| files_changed | integer | Count of files touched by this commit | 0 (empty commit) upward; some commits are empty by design (see Notes) |
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
| repo_name | string | Project directory the file belongs to | Same 9 values as above |
| file_path | string | Path relative to the project directory | e.g. src/dataset/app.py |
| file_extension | string | File extension, lowercase | py, json, tf, tfvars, yml, yaml, sh, js, jsx, md, sql, hcl, txt, etc. |
| line_count | integer | Newline-delimited line count | ≥0 |
| word_count | integer | Whitespace-delimited word count | ≥0 |
| char_count | integer | Character count (UTF-8 decoded) | ≥0; binary files excluded (87 files) |
Binary/asset files (fonts, images, zip archives, compiled .pyc) are excluded from the files table and from all LOC/word/character counts throughout this listing, since line/word counts aren't meaningful for them
Listing Stats
VIEWS
15
DELIVERY
INSTANT DOWNLOAD
LISTED
10/08/2026
UPDATED
02/09/2026
REGION
GLOBAL
TRUST
5 / 5
Loading...
£50,000
Download Dataset in ZIP Format
Recommended Datasets
Loading recommendations...
