12K+ DSA Codebases & 64K+ Programming Solutions Dataset
Foundation Model Datasets
Tags and Keywords

£286,900
About
12K+ DSA Codebases & 64K+ Programming Solutions Dataset
Description
The 12K+ DSA Codebases & 64K+ Programming Solutions Dataset is a comprehensive multi-language source code dataset comprising 1,200+ Data Structures and Algorithms (DSA) codebases, 3.86M+ lines of source code, and 64K+ programming solutions across 9 programming languages. Designed for AI/ML, foundation model training, code generation, software engineering, and developer tools, this dataset provides high-quality code samples for building and evaluating intelligent coding systems.
The dataset supports a wide range of applications, including large language model (LLM) training, code generation, code completion, program synthesis, code summarization, automated code review, bug detection, and algorithmic reasoning. It is an ideal resource for AI developers, educational platforms, and software organizations building next-generation coding assistants and foundation models.
Note: Pricing varies depending on several factors, including the number of code repositories, total lines of code, programming language coverage, metadata availability, annotation requirements, and customization needs. The final price will be determined based on the specific dataset requirements.
Data Product Features
| Feature | Description |
|---|---|
| Programming Language | Language used in the source code |
| Algorithm Category | Data structure or algorithm classification |
| Problem Title | Name of the coding problem |
| Source Code | Complete implementation of the solution |
| Lines of Code | Number of lines in each source file |
| File Extension | Programming language file extension |
Distribution
The dataset is distributed as organized source code files suitable for AI training, software engineering research, and foundation model development.
- Format: Source code files (e.g.,
.cpp,.py,.java,.js,.cs,.c.).
Data Volume
- Codebases: 12,000+
- Programming Solutions: 64,000+
- Lines of Code: 3.86M+
- Programming Languages: 9
- Dataset Size: Dataset size may vary depending on source code volume, file formats, metadata availability, and dataset version.
Usage
This data product is ideal for a variety of applications:
- Foundation Model Training: Train large language models on high-quality multi-language source code.
- Code Generation: Develop AI models capable of generating correct and efficient source code.
- Code Completion: Build intelligent programming assistants and IDE autocomplete systems.
- Program Synthesis: Generate executable code from programming tasks and algorithmic descriptions.
- Algorithmic Reasoning: Train models to understand and solve data structures and algorithms problems.
- Code Summarization: Generate natural language summaries of source code.
- Bug Detection: Build AI systems for identifying coding errors and software vulnerabilities.
- Software Engineering Research: Analyze coding patterns, algorithm implementations, and software quality.
- Educational Platforms: Support coding education, competitive programming, and technical interview preparation.
- Model Benchmarking: Evaluate AI coding assistants and code-focused foundation models.
Coverage
Geographic Coverage
Global.
License
CC BY 4.0 (Creative Commons Attribution 4.0 International)
AI Training Rights
InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.
Data Dictionary
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
| programming_language | String | Programming language used | C++, Java, Python, JavaScript, PHP, C#, C, CPP, etc. |
| algorithm_category | String | DSA topic or algorithm category | Arrays, Trees, Graphs, DP, Sorting, etc. |
| problem_title | String | Name of the coding problem | Optional |
| source_code | Text | Complete source code implementation | Source code |
| lines_of_code | Integer | Number of lines in the source file | Positive integer |
| file_extension | String | Source code file extension | .cpp, .py, .java, .js, .cs, etc. |
Considerations
This dataset is provided for research and educational purposes only. It contains only sample data.
Additional Notes
- This dataset contains multi-language Data Structures and Algorithms (DSA) codebases and programming solutions for AI training and software engineering applications.
- It is optimized for foundation model training, code generation, code completion, program synthesis, algorithmic reasoning, software engineering research, and AI-powered developer tools.
Loading...
£286,900
Download Dataset in CODE Format
Recommended Datasets
Loading recommendations...
