200+ Legacy Codebases Dataset

Foundation Model Datasets

Tags and Keywords

Legacy

Codebase

Software

Programming

Refactoring

Modernization

Llm

Ai

200+ Legacy Codebases Dataset Dataset on Opendatabay data marketplace

£742,600

About

200+ Legacy Codebases Dataset

Description

The 200+ Legacy Codebases Dataset is a comprehensive collection of 200+ legacy software codebases spanning 11 industries, comprising 90.85M+ lines of source code and 455K+ source code files. Designed for AI/ML, foundation model training, software engineering, code intelligence, and application modernization, this dataset provides diverse real-world legacy code for developing and evaluating intelligent software development tools.
The dataset supports a wide range of AI applications, including code generation, code completion, code summarization, code migration, automated refactoring, software maintenance, technical debt analysis, bug detection, program comprehension, and legacy application modernization. It is an ideal resource for AI researchers, software engineering teams, and organizations building next-generation coding assistants and foundation models.
Note: The listed price applies to the specified initial batch of 2 million lines of code. Pricing for larger batches or the complete dataset library varies depending on the number of code repositories, total lines of code, programming language coverage, code complexity, metadata availability, annotation requirements, file formats, licensing terms, and customization needs. Final pricing will be determined based on the specific dataset requirements.

Data Product Features

FeatureDescription
IndustryIndustry associated with the Codebases
Project NameLegacy application or project name
Programming LanguageProgramming language(s) used
Source CodeComplete source code files
File CountNumber of files within each repository
Lines of CodeTotal lines of source code
File ExtensionSource code file extensions

Distribution

  • Format: Source code files (e.g., .java, .cs, .cpp, .c, .py, .js, .ts, .php, .json, configuration files, and related project files where applicable)
  • Structure: Organized repository folders containing source code, project files.

Data Volume

  • Codebases: 200+
  • Industries: 11
  • Lines of Code: 90.85M+
  • Files: 455K+
  • Dataset Size: Dataset size may vary depending on repository size, programming language coverage, metadata availability, and dataset version.

Usage

This data product is ideal for a variety of applications:
  • Foundation Model Training: Train large language models on enterprise-scale legacy source code.
  • Code Modernization: Develop AI systems for migrating and modernizing legacy software.
  • Code Generation: Build models capable of generating maintainable source code.
  • Code Completion: Enhance intelligent developer assistants and IDE autocomplete tools.
  • Automated Refactoring: Train AI models to improve code quality and maintainability.
  • Code Summarization: Generate documentation and natural language explanations from legacy code.
  • Program Comprehension: Build models for understanding complex software systems.
  • Bug Detection: Develop AI systems for identifying defects and code vulnerabilities.
  • Technical Debt Analysis: Analyze software complexity and modernization opportunities.
  • Software Engineering Research: Benchmark AI models for legacy software maintenance and evolution.

Coverage

Geographic Coverage

Global.

License

CC BY 4.0 (Creative Commons Attribution 4.0 International)

AI Training Rights

InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.

Data Dictionary

Column NameData TypeDescriptionPossible Values/Notes
project_nameStringName of the software projectOptional
industryStringIndustry categoryFinance, Healthcare, Retail, Manufacturing, etc.
programming_languageStringProgramming language(s) usedJava, C#, C++, Python, JavaScript, etc.
source_fileStringSource code filenameVarious file types
file_extensionStringSource code file extension.java, .cs, .cpp, .py, .js, etc.
lines_of_codeIntegerNumber of lines of codePositive integer
file_countIntegerNumber of files in the repositoryPositive integer

Considerations

This dataset is provided for research and educational purposes only. It contains only sample data.

Additional Notes

  • This dataset contains 200+ legacy software codebases spanning 11 industries, comprising 90.85M+ lines of code and 455K+ source code files.

Listing Stats

VIEWS

6

DELIVERY

CUSTOM, S3

LISTED

04/08/2026

UPDATED

08/08/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£742,600

Download Dataset in CODE Format