27M+ Words covering 400+ Nepali Non-STEM Textbooks Dataset
Natural Language Processing
Tags and Keywords

£40,900
About
27M+ Words covering 400+ Nepali Non-STEM Textbooks Dataset
Description
The 27M+ Words covering 400+ Nepali Non-STEM Textbooks Dataset is a large-scale educational text corpus comprising more than 400+ non-STEM textbooks collected from diverse academic disciplines. Covering subjects such as history, literature, social sciences, languages, economics, philosophy, arts, law, business, and humanities, this dataset contains over 27M+ words of structured textual content.
Designed for Artificial Intelligence (AI), Natural Language Processing (NLP), Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), educational technology, and academic research, the dataset provides rich educational language resources suitable for training, fine-tuning, evaluating, and benchmarking language models. The collection supports multilingual and cross-domain learning, making it valuable for developing educational AI systems, semantic search engines, text analytics tools, and domain-specific language models.
Note: Pricing varies depending on several factors, including the number of textbooks, total word count, subject coverage, language, metadata availability, annotation requirements, and customization needs. The final price will be determined based on the specific dataset requirements.
Data Product Features
| Feature | Description |
|---|---|
| Title | Name of the textbook |
| Subject Category | Humanities, Social Sciences, Languages, Business, Arts, Law, etc. |
| Language | Language of the textbook content |
| Chapter Information | Chapter titles and structure |
| Text Content | Full extracted textual content |
| Word Count | Number of words per textbook |
| Page Count | Number of pages |
Distribution
- Formats: PDF
Data Volume:
- Books: 400+ Non-STEM textbooks
- Text Volume: 27M+ words
- Dataset Size: The dataset size may vary depending on the number of textbooks, page count, text volume, file formats, metadata availability, annotations, and dataset version.
Usage
This data product is ideal for a variety of applications:
- LLM Training: Train and fine-tune educational and domain-specific language models.
- NLP Research: Support text classification, summarization, semantic analysis, and topic modeling.
- RAG Systems: Build educational retrieval and question-answering systems.
- Educational AI: Develop tutoring platforms, intelligent assistants, and e-learning tools.
- OCR Improvement: Enhance text extraction systems for educational documents.
- Search Engines: Create semantic search and recommendation systems for academic content.
- Knowledge Graphs: Generate structured educational knowledge bases.
Coverage
The dataset provides comprehensive coverage of Nepali-language Non-STEM content across multiple disciplines.
- Geographic Coverage: Global
- Domains: Education, Humanities, Social Sciences, Languages, Literature, History, Economics, Business, Law, Arts
License
CC BY 4.0 (Creative Commons Attribution 4.0 International)
AI Training Rights
InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.
Data Dictionary
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
| Document_Title | String | Title or name of the document | Free text |
| Subject | String | Non-STEM discipline | History, Literature, Economics, Arts, etc. |
| Language | String | Language of the document | Nepali |
| Content | Text | Complete document text | Nepali text |
| Keywords | String | Document keywords | Comma-separated values |
| Word_Count | Integer | Total number of words in the document | Positive integer |
Considerations
This dataset is provided for research and educational purposes only. It contains only sample data.
Loading...
£40,900
Download Dataset in TEXT Format
Recommended Datasets
Loading recommendations...
