Multilingual Books Corpus — 10 Indian Languages

Natural Language Processing

Tags and Keywords

Multilingual

Books

Indianlanguages

Nlp

Llmtraining

Textcorpus

Indiclanguages

Full-text

Multilingual Books Corpus — 10 Indian Languages Dataset on Opendatabay data marketplace

£400,000

About

A large-scale full-text books corpus spanning 10 Indian languages — Bengali, English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, and Tamil. Contains 76,087 books totalling 29.3 million pages and approximately 11.7 billion tokens. One of the largest curated multilingual book corpora available for Indian language LLM training and NLP research.
Data Product Features
  • 76,087 full-text books across 10 Indian languages
  • 29.3 million pages — approximately 11.7 billion tokens
  • English largest single language at 14,734 titles
  • Covers Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and English
  • Delivered in PDF, ePub, and plain text formats
  • Historical archive spanning diverse literary and non-fiction genres
Distribution
Format: PDF / ePub / Text Size: ~11.7 billion tokens Records: 76,087 books
Data Volume
76,087 full-text books, 29.3 million pages, approximately 11.7 billion tokens across 10 Indian languages.
Usage
  • LLM pre-training and fine-tuning for Indian language models
  • Multilingual NLP research and benchmarking
  • Cross-lingual transfer learning
  • Text summarisation and document understanding model training
  • Building Indic language AI assistants and translation systems
Coverage
Geographic Coverage: India Time Range: Historical archive Languages: Bengali, English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil
License
CC0 — No Rights Reserved
AI Training Rights
Licensee is granted a non-exclusive, worldwide, and perpetual right to use this data product to train, fine-tune, and evaluate machine learning models. The data product itself may not be redistributed or shared outside licensed usage. Licensee must comply with all applicable laws, including data protection and privacy regulations.
Who Can Use It
  • AI/ML Engineers: For training and fine-tuning Indian language LLMs
  • NLP Researchers: For multilingual benchmarking and cross-lingual studies
  • Enterprises: For building Indic language AI products and assistants
  • Academic Institutions: For computational linguistics and language preservation research
Data Dictionary
  • book_id (string) — Unique identifier for each book
  • title (string) — Title of the book
  • language (string) — Primary language of the book
  • genre (string) — Genre or subject category
  • page_count (integer) — Total number of pages
  • token_count (integer) — Approximate token count for the book
  • file_format (string) — PDF, ePub, or plain text
  • publication_year (integer) — Year of original publication where available
  • author (string) — Author name where available
  • source (string) — Anonymised source or collection identifier

Listing Stats

VIEWS

8

DELIVERY

CUSTOM, S3

LISTED

11/09/2026

UPDATED

12/09/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£400,000

Download Dataset in TEXT Format