Multilingual Books Corpus — 10 Indian Languages
Natural Language Processing
Tags and Keywords

£400,000
About
A large-scale full-text books corpus spanning 10 Indian languages — Bengali, English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, and Tamil. Contains 76,087 books totalling 29.3 million pages and approximately 11.7 billion tokens. One of the largest curated multilingual book corpora available for Indian language LLM training and NLP research.
Data Product Features
- 76,087 full-text books across 10 Indian languages
- 29.3 million pages — approximately 11.7 billion tokens
- English largest single language at 14,734 titles
- Covers Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and English
- Delivered in PDF, ePub, and plain text formats
- Historical archive spanning diverse literary and non-fiction genres
Distribution
Format: PDF / ePub / Text
Size: ~11.7 billion tokens
Records: 76,087 books
Data Volume
76,087 full-text books, 29.3 million pages, approximately 11.7 billion tokens across 10 Indian languages.
Usage
- LLM pre-training and fine-tuning for Indian language models
- Multilingual NLP research and benchmarking
- Cross-lingual transfer learning
- Text summarisation and document understanding model training
- Building Indic language AI assistants and translation systems
Coverage
Geographic Coverage: India
Time Range: Historical archive
Languages: Bengali, English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil
License
CC0 — No Rights Reserved
AI Training Rights
Licensee is granted a non-exclusive, worldwide, and perpetual right to use this data product to train, fine-tune, and evaluate machine learning models. The data product itself may not be redistributed or shared outside licensed usage. Licensee must comply with all applicable laws, including data protection and privacy regulations.
Who Can Use It
- AI/ML Engineers: For training and fine-tuning Indian language LLMs
- NLP Researchers: For multilingual benchmarking and cross-lingual studies
- Enterprises: For building Indic language AI products and assistants
- Academic Institutions: For computational linguistics and language preservation research
Data Dictionary
- book_id (string) — Unique identifier for each book
- title (string) — Title of the book
- language (string) — Primary language of the book
- genre (string) — Genre or subject category
- page_count (integer) — Total number of pages
- token_count (integer) — Approximate token count for the book
- file_format (string) — PDF, ePub, or plain text
- publication_year (integer) — Year of original publication where available
- author (string) — Author name where available
- source (string) — Anonymised source or collection identifier
Loading...
£400,000
Download Dataset in TEXT Format
Recommended Datasets
Loading recommendations...
