7.3B+ Words Multi-Domain Academic Thesis Dataset

Natural Language Processing

Tags and Keywords

Thesis

Academic

Corpus

Multidomain

Research

Nlp

Llm

Education

7.3B+ Words Multi-Domain Academic Thesis Dataset Dataset on Opendatabay data marketplace

£741,200

About

7.3B+ Words Multi-Domain Academic Thesis Dataset

Description

The 7.3B+ Words Multi-Domain Academic Thesis Text Dataset (10+ Disciplines) is a large-scale collection of academic thesis documents spanning more than ten specialized research domains. Designed for artificial intelligence, natural language processing (NLP), large language model (LLM) development, and academic research, the dataset contains over 7.3 billion words of scholarly text covering diverse scientific, engineering, medical, business, social science, and humanities disciplines.
The dataset provides rich, structured academic content suitable for training, fine-tuning, and evaluating language models, information retrieval systems, document understanding models, text summarization, semantic search, and other AI applications. Its broad disciplinary coverage enables the development of robust AI systems capable of understanding technical and scholarly language across multiple fields.
Note: The listed price applies to the specified initial batch of 500 million words. Pricing for larger batches or the complete dataset library varies depending on the number of thesis documents, total word count, subject coverage, language, metadata availability, document format, annotation requirements, licensing terms, and customization needs. Final pricing will be determined based on the specific dataset requirements.

Data Product Features

The dataset may include the following metadata fields (availability depends on the licensed version):
FeatureDescription
Academic DomainPrimary academic field or research discipline to which the thesis belongs (e.g., Computer Science, Medicine, Engineering, Social Sciences).
LanguageEnglish or other supported languages
AbstractSummary of the thesis.
ChaptersStructured organization of the thesis, including chapter titles, sections, and subsections where available
Full TextComplete thesis content in its original format for research, analysis, and AI applications.
File FormatPDF

Distribution

  • Format: PDF
  • Data Volume:
    • Corpus Size: 7.3B+ Words
    • Domains Covered: 10+ Academic Disciplines
    • Dataset Size: The dataset size may vary depending on the number of thesis, page count, word count, text length, file formats, metadata availability, annotations, and dataset version.

Usage

This data product is ideal for a variety of applications:
  • LLM Pretraining: Train large language models using high-quality academic text.
  • LLM Fine-Tuning: Adapt foundation models for academic and scientific language understanding.
  • Natural Language Processing: Develop text classification, entity recognition, semantic analysis, and information extraction models.
  • Text Summarization: Train systems for automatic thesis and research paper summarization.
  • Semantic Search: Build intelligent academic search and retrieval systems.
  • Question Answering: Develop AI systems capable of answering research-related questions.
  • Knowledge Graph Construction: Extract entities and relationships from scholarly documents.
  • Academic Research: Conduct bibliometric, linguistic, and computational research.

Coverage

The dataset provides comprehensive scholarly coverage across multiple research disciplines.
  • Geographic Coverage: Global
  • Academic Domains: 10+ disciplines including engineering, computer science, medicine, business, social sciences, education, law, economics, natural sciences, and humanities.

License

CC BY 4.0 (Creative Commons Attribution 4.0 International)

AI Training Rights

InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.

Data Dictionary

Column NameData TypeDescriptionPossible Values/Notes
Thesis_TitleStringTitle of the thesis documentFree text
Academic_DomainStringResearch disciplineComputer Science, Medicine, Engineering, Business, Law, Education, Economics, Social Sciences, Humanities, Natural Sciences, etc.
LanguageStringLanguage of the thesisEnglish or other supported languages
AbstractTextThesis summaryFree text
KeywordsStringResearch keywordsMultiple values
ChaptersIntegerNumber of thesis chaptersPositive integer
Word_CountIntegerNumber of words in the documentPositive integer
File_FormatStringOriginal file formatPDF
Full_TextTextComplete thesis contentLong-form text

Considerations


This dataset is provided for research and educational purposes only. It contains only sample data.

Listing Stats

VIEWS

3

DELIVERY

CUSTOM, S3

LISTED

24/07/2026

UPDATED

08/08/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£741,200

Download Dataset in TEXT Format