12K+ DSA Codebases & 64K+ Programming Solutions Dataset

Foundation Model Datasets

Tags and Keywords

Dsa

Algorithms

Codebases

Programming

Multilingual

Coding

Software

Llm

12K+ DSA Codebases & 64K+ Programming Solutions Dataset Dataset on Opendatabay data marketplace

£286,900

About

12K+ DSA Codebases & 64K+ Programming Solutions Dataset

Description

The 12K+ DSA Codebases & 64K+ Programming Solutions Dataset is a comprehensive multi-language source code dataset comprising 1,200+ Data Structures and Algorithms (DSA) codebases, 3.86M+ lines of source code, and 64K+ programming solutions across 9 programming languages. Designed for AI/ML, foundation model training, code generation, software engineering, and developer tools, this dataset provides high-quality code samples for building and evaluating intelligent coding systems.
The dataset supports a wide range of applications, including large language model (LLM) training, code generation, code completion, program synthesis, code summarization, automated code review, bug detection, and algorithmic reasoning. It is an ideal resource for AI developers, educational platforms, and software organizations building next-generation coding assistants and foundation models.
Note: Pricing varies depending on several factors, including the number of code repositories, total lines of code, programming language coverage, metadata availability, annotation requirements, and customization needs. The final price will be determined based on the specific dataset requirements.

Data Product Features

FeatureDescription
Programming LanguageLanguage used in the source code
Algorithm CategoryData structure or algorithm classification
Problem TitleName of the coding problem
Source CodeComplete implementation of the solution
Lines of CodeNumber of lines in each source file
File ExtensionProgramming language file extension

Distribution

The dataset is distributed as organized source code files suitable for AI training, software engineering research, and foundation model development.
  • Format: Source code files (e.g., .cpp, .py, .java, .js, .cs, .c.).

Data Volume

  • Codebases: 12,000+
  • Programming Solutions: 64,000+
  • Lines of Code: 3.86M+
  • Programming Languages: 9
  • Dataset Size: Dataset size may vary depending on source code volume, file formats, metadata availability, and dataset version.

Usage

This data product is ideal for a variety of applications:
  • Foundation Model Training: Train large language models on high-quality multi-language source code.
  • Code Generation: Develop AI models capable of generating correct and efficient source code.
  • Code Completion: Build intelligent programming assistants and IDE autocomplete systems.
  • Program Synthesis: Generate executable code from programming tasks and algorithmic descriptions.
  • Algorithmic Reasoning: Train models to understand and solve data structures and algorithms problems.
  • Code Summarization: Generate natural language summaries of source code.
  • Bug Detection: Build AI systems for identifying coding errors and software vulnerabilities.
  • Software Engineering Research: Analyze coding patterns, algorithm implementations, and software quality.
  • Educational Platforms: Support coding education, competitive programming, and technical interview preparation.
  • Model Benchmarking: Evaluate AI coding assistants and code-focused foundation models.

Coverage

Geographic Coverage

Global.

License

CC BY 4.0 (Creative Commons Attribution 4.0 International)

AI Training Rights

InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.

Data Dictionary

Column NameData TypeDescriptionPossible Values/Notes
programming_languageStringProgramming language usedC++, Java, Python, JavaScript, PHP, C#, C, CPP, etc.
algorithm_categoryStringDSA topic or algorithm categoryArrays, Trees, Graphs, DP, Sorting, etc.
problem_titleStringName of the coding problemOptional
source_codeTextComplete source code implementationSource code
lines_of_codeIntegerNumber of lines in the source filePositive integer
file_extensionStringSource code file extension.cpp, .py, .java, .js, .cs, etc.

Considerations

This dataset is provided for research and educational purposes only. It contains only sample data.

Additional Notes

  • This dataset contains multi-language Data Structures and Algorithms (DSA) codebases and programming solutions for AI training and software engineering applications.
  • It is optimized for foundation model training, code generation, code completion, program synthesis, algorithmic reasoning, software engineering research, and AI-powered developer tools.

Listing Stats

VIEWS

5

DELIVERY

CUSTOM, S3

LISTED

04/08/2026

UPDATED

07/08/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£286,900

Download Dataset in CODE Format