Data-Analysis-Platform

Data Science and Analytics

Tags and Keywords

Healthcare

Data

Code

Llm

Data-Analysis-Platform Dataset on Opendatabay data marketplace

£50,000

About

In context, this represents the kind of real-world, multi-service enterprise engineering history covering the full lifecycle of a data platform build-out — from initial scaffolding through infrastructure provisioning, API development, governance tooling, and testing.
Useful for research into software engineering productivity and patterns, repository-scale benchmarking, or as training/reference material for tools that need to reason about realistic, multi-repository enterprise development history rather than isolated code snippets.

Data Product Features

Multi-domain technical coverage — the codebase spans Terraform/infrastructure-as-code, Python backend services, IAM/Cognito configuration, SQL data migrations, and Postman/Newman API test suites, giving breadth across the typical layers of a real platform rather than a single language or service type.Quantifiable scale — roughly 108,000 lines of code, 1,474 tracked files, and ~1.1M tokens of text content, already computed and ready to report as listing specifications.

Distribution

The data product is delivered as a native Git repository — full version-controlled history, not a flat export. It can be cloned and browsed with standard git tooling (git log, git blame, git diff all work normally), or, if a marketplace requires a static download instead, exported as commit-level and file-level metric tables in CSV/JSON.
  • **Data Volume: Working tree (actual files, no git metadata): 6.0 MB Git object database (.git, full commit history): 4.5 MB Combined: ~10.5 MB total 5,683 commits, 1,474 tracked files, ~108,761 lines of code, ~317,744 words, ~4.41M characters (~1.1M tokens estimated)

Usage

This data product is ideal for a variety of applications:
Application: Software engineering research — studying real commit cadence, sequencing, and team behavior across a multi-service platform build, rather than relying on synthetic or single-repo samples. Application: LLM and AI training data — fine-tuning or evaluating code-and-commit-understanding models on authentic enterprise development history, including realistic commit messages and diff patterns. Application: Repository sizing and cost estimation benchmarking — using the measured LOC, file counts, and token volumes as reference points for scoping similar platform engagements. Application: DevOps and CI/CD pattern analysis — examining how infrastructure-as-code (Terraform), IAM configuration, and application code evolved together across interdependent services. Application: Codebase migration and monorepo case studies — as a worked example of consolidating multiple independently-versioned repositories into a single repository while preserving full commit history. Application: Developer productivity and engineering-velocity modeling — commit frequency, file-churn, and cross-service coordination patterns over a defined ~3-month delivery window. Application: Training material for engineering teams — a realistic reference codebase (data platform: API services, IAM, governance, curation, infrastructure) for onboarding, tooling demos, or internal workshops.

Coverage

Geographic Coverage: Development-side coverage is India — every commit in the repository carries an IST (UTC+05:30) author/committer timestamp, indicating an India-based engineering team. The platform itself was built for a global healthcare organization's cloud infrastructure (AWS-hosted), so the deployment target is global/US-centric even though the development activity captured here is India-based. Time Range: Earliest recorded commit: 2023-02-13 (data, infrastructure, client, core-api, curation). Latest recorded commit: 2025-01-28 (dm-producer-mvp). So the dataset spans roughly February 2023 – January 2025 (~23 months) in total, though individual projects have different active windows within that span

License

CC0

AI Training Rights

Licensee is granted a non-exclusive, worldwide, and perpetual right to:
  • Use the Data Product to train, fine-tune, and evaluate machine learning models, including large language models.
  • Incorporate Data Product content into models and commercialize resulting model outputs.
  • Create derivative works (model weights, embeddings, etc.) for any lawful purpose.
Restrictions:
  • The Data Product itself may not be sold, redistributed, or shared outside of licensed usage.
  • Licensee must comply with all applicable laws, including data protection and privacy regulations.

Who Can Use It

Data Scientists / ML Engineers: For training or evaluating code-and-commit-understanding models — fine-tuning LLMs on realistic commit sequencing, diff patterns, and multi-service coordination rather than synthetic code samples. Software Engineering Researchers: For academic or applied studies on developer productivity, commit cadence, monorepo consolidation patterns, or how infrastructure-as-code and application code co-evolve across interdependent services over a real ~2-year delivery window. DevOps / Platform Engineering Teams: For benchmarking — using the measured LOC, file counts, commit volume, and project structure as reference points when scoping, estimating, or comparing similar data-platform engagements. Engineering Managers / Consultancies: For internal training material — a realistic multi-service reference codebase (API services, IAM, data governance, curation, Terraform infrastructure) useful for onboarding, tooling demonstrations, or workshops without needing a live production system. AI/Code-Tooling Businesses: For product development and testing — evaluating code search, static analysis, repository-migration tooling, or git-history-analysis products against a real, non-trivial monorepo rather than toy examples.

Data Dictionary

The primary deliverable is a git repository, not inherently tabular — but for buyers who want a tabular view, two derived summary tables can be exported (commit-level and file-level). Here's the dictionary for both:
Column NameData TypeDescriptionPossible Values/Notes
repo_namestringProject directory the commit belongs toOne of: client, core-api, curation, data, dm-central-governance, dm-domain-mvp, dm-producer-mvp, health-data-platform-iam-role, infrastructure
commit_hashstring (40 chars)Git SHA-1 commit hash within the monorepoHexadecimal string
author_namestringCommit author display nameFree text; individual contributor names present in git history
author_emailstringCommit author email addressFree text; sanitized to remove client-domain addresses
commit_datedatetime (ISO 8601)Author/committer timestampRange 2023-02-13 to 2025-01-28; all values carry +05:30 (IST) offset
commit_messagestringCommit message (subject line)Free text, typically <100 chars
files_changedintegerCount of files touched by this commit0 (empty commit) upward; some commits are empty by design (see Notes)
Column NameData TypeDescriptionPossible Values/Notes
repo_namestringProject directory the file belongs toSame 9 values as above
file_pathstringPath relative to the project directorye.g. src/dataset/app.py
file_extensionstringFile extension, lowercasepy, json, tf, tfvars, yml, yaml, sh, js, jsx, md, sql, hcl, txt, etc.
line_countintegerNewline-delimited line count≥0
word_countintegerWhitespace-delimited word count≥0
char_countintegerCharacter count (UTF-8 decoded)≥0; binary files excluded (87 files)

Binary/asset files (fonts, images, zip archives, compiled .pyc) are excluded from the files table and from all LOC/word/character counts throughout this listing, since line/word counts aren't meaningful for them

Listing Stats

VIEWS

15

DELIVERY

INSTANT DOWNLOAD

LISTED

10/08/2026

UPDATED

02/09/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£50,000

Download Dataset in ZIP Format