7.3B+ Words Multi-Domain Academic Thesis Dataset
Natural Language Processing
Tags and Keywords

£741,200
About
7.3B+ Words Multi-Domain Academic Thesis Dataset
Description
The 7.3B+ Words Multi-Domain Academic Thesis Text Dataset (10+ Disciplines) is a large-scale collection of academic thesis documents spanning more than ten specialized research domains. Designed for artificial intelligence, natural language processing (NLP), large language model (LLM) development, and academic research, the dataset contains over 7.3 billion words of scholarly text covering diverse scientific, engineering, medical, business, social science, and humanities disciplines.
The dataset provides rich, structured academic content suitable for training, fine-tuning, and evaluating language models, information retrieval systems, document understanding models, text summarization, semantic search, and other AI applications. Its broad disciplinary coverage enables the development of robust AI systems capable of understanding technical and scholarly language across multiple fields.
Note: The listed price applies to the specified initial batch of 500 million words. Pricing for larger batches or the complete dataset library varies depending on the number of thesis documents, total word count, subject coverage, language, metadata availability, document format, annotation requirements, licensing terms, and customization needs. Final pricing will be determined based on the specific dataset requirements.
Data Product Features
The dataset may include the following metadata fields (availability depends on the licensed version):
| Feature | Description |
|---|---|
| Academic Domain | Primary academic field or research discipline to which the thesis belongs (e.g., Computer Science, Medicine, Engineering, Social Sciences). |
| Language | English or other supported languages |
| Abstract | Summary of the thesis. |
| Chapters | Structured organization of the thesis, including chapter titles, sections, and subsections where available |
| Full Text | Complete thesis content in its original format for research, analysis, and AI applications. |
| File Format |
Distribution
-
Format: PDF
-
Data Volume:
- Corpus Size: 7.3B+ Words
- Domains Covered: 10+ Academic Disciplines
- Dataset Size: The dataset size may vary depending on the number of thesis, page count, word count, text length, file formats, metadata availability, annotations, and dataset version.
Usage
This data product is ideal for a variety of applications:
- LLM Pretraining: Train large language models using high-quality academic text.
- LLM Fine-Tuning: Adapt foundation models for academic and scientific language understanding.
- Natural Language Processing: Develop text classification, entity recognition, semantic analysis, and information extraction models.
- Text Summarization: Train systems for automatic thesis and research paper summarization.
- Semantic Search: Build intelligent academic search and retrieval systems.
- Question Answering: Develop AI systems capable of answering research-related questions.
- Knowledge Graph Construction: Extract entities and relationships from scholarly documents.
- Academic Research: Conduct bibliometric, linguistic, and computational research.
Coverage
The dataset provides comprehensive scholarly coverage across multiple research disciplines.
- Geographic Coverage: Global
- Academic Domains: 10+ disciplines including engineering, computer science, medicine, business, social sciences, education, law, economics, natural sciences, and humanities.
License
CC BY 4.0 (Creative Commons Attribution 4.0 International)
AI Training Rights
InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.
Data Dictionary
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
| Thesis_Title | String | Title of the thesis document | Free text |
| Academic_Domain | String | Research discipline | Computer Science, Medicine, Engineering, Business, Law, Education, Economics, Social Sciences, Humanities, Natural Sciences, etc. |
| Language | String | Language of the thesis | English or other supported languages |
| Abstract | Text | Thesis summary | Free text |
| Keywords | String | Research keywords | Multiple values |
| Chapters | Integer | Number of thesis chapters | Positive integer |
| Word_Count | Integer | Number of words in the document | Positive integer |
| File_Format | String | Original file format | |
| Full_Text | Text | Complete thesis content | Long-form text |
Considerations
This dataset is provided for research and educational purposes only. It contains only sample data.
Loading...
£741,200
Download Dataset in TEXT Format
Recommended Datasets
Loading recommendations...
