7M+ Words covering 100+ Marathi STEM Textbooks Dataset

Natural Language Processing

Tags and Keywords

Stem

Science

Engineering

Mathematics

Technology

Education

Nlp

Marathi

7M+ Words covering 100+ Marathi STEM Textbooks Dataset Dataset on Opendatabay data marketplace

£11,600

About

7M+ Words covering 100+ Marathi STEM Textbooks Dataset

Description

The 7M+ Words Marathi STEM Text Dataset is a large-scale collection of Marathi-language Science, Technology, Engineering, and Mathematics (STEM) text designed for artificial intelligence, natural language processing (NLP), and machine learning applications. The dataset contains diverse STEM content sourced from educational and technical materials, with interwoven visuals for improved contextual understanding, making it suitable for language model training, knowledge extraction, semantic search, document understanding, and AI-driven educational technologies.
Designed for researchers, AI developers, educational institutions, and enterprises, this dataset provides high-quality STEM text to support advanced language models, intelligent search systems, domain-specific AI applications, and scientific text analytics.
Note: Pricing varies depending on several factors, including the number of textbooks, total word count, subject coverage, language, metadata availability, annotation requirements, and customization needs. The final price will be determined based on the specific dataset requirements.

Data Product Features

FeatureDescription
SubjectSTEM discipline such as Science, Technology, Engineering, or Mathematics.
LanguageLanguage in which the document is written (Marathi).
Document TypeType of content, such as textbook, article, manual, guide, lecture notes, or reference material.
ContentComplete textual content of the document.
KeywordsKeywords describing the document's primary topics.

Distribution

  • Format: PDF
  • Data Volume: 7M+ Words across Marathi STEM Textbook

Usage

This data product is ideal for a variety of applications:
  • Large Language Models (LLMs): Train and fine-tune domain-specific language models.
  • Natural Language Processing: Develop text classification, summarization, and information extraction systems.
  • Semantic Search: Build intelligent search and retrieval applications for STEM content.
  • Educational AI: Create AI-powered tutoring, learning assistants, and question-answering systems.
  • Knowledge Graphs: Extract structured knowledge from scientific and technical text.
  • Document Intelligence: Automate indexing, categorization, and document analysis.
  • Research Analytics: Support scientific literature mining and educational research.
  • Retrieval-Augmented Generation (RAG): Enhance AI systems with reliable STEM knowledge.

Coverage

The dataset provides comprehensive coverage of Marathi-language STEM content across multiple disciplines.
  • Geographic Coverage: Global
  • Domains: Science, Technology, Engineering, Mathematics, Education, Research, and Technical Documentation

License

CC BY 4.0 (Creative Commons Attribution 4.0 International)

AI Training Rights

InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.

Data Dictionary

Column NameData TypeDescriptionPossible Values/Notes
Document_TitleStringTitle or name of the documentFree text
SubjectStringSTEM disciplineScience, Technology, Engineering, Mathematics
LanguageStringLanguage of the documentMarathi
Document_TypeStringType of documentTextbook, Article, Manual, Guide, Lecture Notes, Reference
ContentTextComplete document textMarathi text
KeywordsStringDocument keywordsComma-separated values
Word_CountIntegerTotal number of words in the documentPositive integer

Considerations

This dataset is provided for research and educational purposes only. It contains only sample data.

Listing Stats

VIEWS

1

DELIVERY

CUSTOM, S3

LISTED

25/07/2026

UPDATED

05/08/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£11,600

Download Dataset in TEXT Format