Bilingual Multi-Genre Foundational Text Corpus

LLM Fine-Tuning Data

Tags and Keywords

Text-corpus

Bilingual-indic

Llm-pretraining

Fiction-data

Poetry-matrix

Couplet/doha-collection

Academic-texts

Human-curated/nonfiction/fiction/verses/stories/novels/articles

Bilingual Multi-Genre Foundational Text Corpus Dataset on Opendatabay data marketplace

$16,400

About

📂 Data Product Executive Overview

This massive, high-density bilingual foundational text corpus is officially curated, digitized, and structured under Prakhar Goonj Publications—a globally recognized, D&B U-N-S Certified international publishing house (D-U-N-S No. 64-125-5366) based in Delhi, India. We are a Tier-1 registered member of the Federation of Indian Publishers (FIP) and officially empanelled with the Ministry of Information & Broadcasting (RNI, Government of India).

🏛️ Dataset Scope & Domain Density (Verified Sourcing Core)

This foundational database consists of massive linguistic text arrays built for pre-training and fine-tuning Large Language Models (LLMs) in deep native Indic token structures and multilingual fluency. Stripped entirely of isolated legal, medical, astronomical, and local low-resource provincial dialects, this corpus focuses strictly on mainstream bilingual vocabulary, creative prose meters, and multi-disciplinary structures.
  • Total Clean Dataset Size: 616 Expert-Authored Published Book Titles
  • Total Volume Infrastructure: ~89,145 Fully Proofread, Human-Authored Digitized Pages (Approx)
  • Linguistic Footprint: ~31 Million Exact Words / ~52 Million Tokens (Approx)
  • Language Array Distribution: Balanced Bilingual Matrices (~65% Pure Indic Hindi & ~35% Academic English)
  • Commercial Asset Valuation Pricing: £12,500 GBP (Lumpsum Non-Exclusive License Payout)

📚 Some of the Core Featured Multi-Genre Categories Included:

  1. Extensive Poetry, Doha, & Gazal Repositories: High-context literary frameworks detailing structured meters, rhyming variables, and classical couplets (featuring an extensive collection of 5,233 standard couplets by Dr. Om Joshi).
  2. Contemporary Fiction, Novels, & Short Stories: Long-form descriptive prose, creative storytelling blocks, and rich structural dialogue patterns natively tracking diverse human interactions.
  3. Academic, STEM, & School Curriculums: Comprehensive higher-education modules, teaching methodologies, and CBSE/NCERT-aligned school textbooks optimized for training logical AI prompt alignment.
  4. Memoirs, Biographies, & Social Studies: Detailed historical retrospectives, political essays, disability awareness treatises, and cultural deep dives spanning localized socio-economic landscapes.

🛡️ Data Provenance & Compliance Governance

  • Collection Method: 100% human-authored, peer-reviewed, and professionally proofread published manuscripts. Zero public web scraping or unverified crowd-sourced data dumps.
  • Licensing Model: Available under a flexible, 100% Non-Exclusive Commercial License for AI Model Training and evaluation purposes. Original text remains secure under publisher custody (Custom Delivery Model) until full transaction clearance.

Listing Stats

VIEWS

9

DELIVERY

CUSTOM, S3

LISTED

01/10/2026

UPDATED

01/10/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

$16,400

Download Dataset in TEXT Format