200+ Legacy Codebases Dataset
Foundation Model Datasets
Tags and Keywords

£742,600
About
200+ Legacy Codebases Dataset
Description
The 200+ Legacy Codebases Dataset is a comprehensive collection of 200+ legacy software codebases spanning 11 industries, comprising 90.85M+ lines of source code and 455K+ source code files. Designed for AI/ML, foundation model training, software engineering, code intelligence, and application modernization, this dataset provides diverse real-world legacy code for developing and evaluating intelligent software development tools.
The dataset supports a wide range of AI applications, including code generation, code completion, code summarization, code migration, automated refactoring, software maintenance, technical debt analysis, bug detection, program comprehension, and legacy application modernization. It is an ideal resource for AI researchers, software engineering teams, and organizations building next-generation coding assistants and foundation models.
Note: The listed price applies to the specified initial batch of 2 million lines of code. Pricing for larger batches or the complete dataset library varies depending on the number of code repositories, total lines of code, programming language coverage, code complexity, metadata availability, annotation requirements, file formats, licensing terms, and customization needs. Final pricing will be determined based on the specific dataset requirements.
Data Product Features
| Feature | Description |
|---|---|
| Industry | Industry associated with the Codebases |
| Project Name | Legacy application or project name |
| Programming Language | Programming language(s) used |
| Source Code | Complete source code files |
| File Count | Number of files within each repository |
| Lines of Code | Total lines of source code |
| File Extension | Source code file extensions |
Distribution
- Format: Source code files (e.g.,
.java,.cs,.cpp,.c,.py,.js,.ts,.php,.json, configuration files, and related project files where applicable) - Structure: Organized repository folders containing source code, project files.
Data Volume
- Codebases: 200+
- Industries: 11
- Lines of Code: 90.85M+
- Files: 455K+
- Dataset Size: Dataset size may vary depending on repository size, programming language coverage, metadata availability, and dataset version.
Usage
This data product is ideal for a variety of applications:
- Foundation Model Training: Train large language models on enterprise-scale legacy source code.
- Code Modernization: Develop AI systems for migrating and modernizing legacy software.
- Code Generation: Build models capable of generating maintainable source code.
- Code Completion: Enhance intelligent developer assistants and IDE autocomplete tools.
- Automated Refactoring: Train AI models to improve code quality and maintainability.
- Code Summarization: Generate documentation and natural language explanations from legacy code.
- Program Comprehension: Build models for understanding complex software systems.
- Bug Detection: Develop AI systems for identifying defects and code vulnerabilities.
- Technical Debt Analysis: Analyze software complexity and modernization opportunities.
- Software Engineering Research: Benchmark AI models for legacy software maintenance and evolution.
Coverage
Geographic Coverage
Global.
License
CC BY 4.0 (Creative Commons Attribution 4.0 International)
AI Training Rights
InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.
Data Dictionary
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
| project_name | String | Name of the software project | Optional |
| industry | String | Industry category | Finance, Healthcare, Retail, Manufacturing, etc. |
| programming_language | String | Programming language(s) used | Java, C#, C++, Python, JavaScript, etc. |
| source_file | String | Source code filename | Various file types |
| file_extension | String | Source code file extension | .java, .cs, .cpp, .py, .js, etc. |
| lines_of_code | Integer | Number of lines of code | Positive integer |
| file_count | Integer | Number of files in the repository | Positive integer |
Considerations
This dataset is provided for research and educational purposes only. It contains only sample data.
Additional Notes
- This dataset contains 200+ legacy software codebases spanning 11 industries, comprising 90.85M+ lines of code and 455K+ source code files.
Loading...
£742,600
Download Dataset in CODE Format
Recommended Datasets
Loading recommendations...
