Mahatma Gandhi Multimodal Historical Archive
Synthetic Images & Vision Datasets
Tags and Keywords

£125,000
About
Mahatma Gandhi Multimodal Historical Archive
A large, human-curated digital archive of historical material relating to Mahatma Gandhi, the Indian independence movement, nonviolence, satyagraha, and the people, places and events surrounding Gandhi's life and legacy.
The collection contains approximately 21,500 catalogued assets spanning historical photographs, writings and correspondence, audio recordings, and video material, accompanied by structured archival metadata.
Unlike datasets assembled through large-scale web scraping, this collection originates from a purpose-built historical archive and preserves source, creator, rights, contextual and descriptive metadata alongside the underlying media.
The archive is designed for organisations seeking historically grounded, provenance-rich primary-source material for multimodal AI development, computer vision, document intelligence, retrieval systems, cultural heritage technology and research.
Data Product Features
The archive currently contains approximately:
- 16,636 image records, including historical photographs, artwork, cartoons and colourised archival material.
- 4,362 writing and document records, including letters, articles, diaries, books, notes and related primary-source documents.
- 304 audio records, representing approximately 205 hours of catalogued audio.
- 199 video records, representing approximately 55 hours of catalogued video.
- 21 structured metadata fields describing identity, media type, date, location, people, subjects, provenance, creator, rights status and other archival attributes.
- A substantial collection of original and colourised historical-image relationships suitable for image restoration, colourisation and multimodal vision research.
- Human-curated subject and contextual metadata covering major people, places, historical events and themes associated with Gandhi.
The underlying material includes photographs, correspondence, speeches, interviews, testimonials, documentary footage, writings and related historical records.
Distribution
The archive is delivered as a multimodal collection combining media assets with structured metadata.
Primary media formats include:
- JPEG / JPG archival images and document scans
- PDF documents
- MP3 audio
- MP4 video
- XLSX structured archival metadata
- ZIP/archive packages where appropriate
Data Volume:
- Catalog records: 21,501
- Unique catalogued asset IDs: approximately 21,400
- Structured metadata fields: 21
- Total archive size: approximately 159 GB
For enterprise licences, the full archive is delivered through an agreed secure cloud or managed transfer mechanism rather than a browser download.
A representative sample package is available for technical and commercial evaluation.
Usage
This data product is suitable for a range of advanced AI, research and cultural-heritage applications:
- Multimodal Foundation Model Training: Train or evaluate models that jointly reason across historical images, documents, audio, video and metadata.
- Vision-Language Models: Develop models capable of captioning, retrieving, classifying and reasoning over historical imagery and archival material.
- Historical Image Understanding: Train computer-vision systems for archival image classification, semantic search, entity recognition and visual retrieval.
- Image Restoration and Colourisation: Use relationships between archival originals and colourised derivatives for restoration, enhancement and historical-image transformation research.
- Document AI and OCR: Develop systems for transcription, document classification, handwriting or print recognition, archival search and document understanding.
- RAG and Knowledge Systems: Build historically grounded retrieval systems that connect primary-source material with people, places, dates, topics and archival provenance.
- Historical Audio Intelligence: Develop speech recognition, retrieval, classification and contextualisation systems using archival speeches, interviews and recordings.
- Historical Video Understanding: Support video retrieval, multimodal search, scene understanding, historical-event recognition and documentary research.
- Digital Humanities: Support computational history, archival research, entity analysis, historical network analysis and primary-source scholarship.
- Cultural Heritage Technology: Develop searchable archives, museum experiences, educational applications and historically grounded digital experiences.
Coverage
Geographic Coverage: Global, with particularly significant material relating to India, South Africa and the United Kingdom.
Temporal Coverage: Catalogued material spans approximately 1857–2009, with the core historical collection concentrated around Gandhi's lifetime and the Indian independence period.
Subject Coverage: Mahatma Gandhi, Indian independence, satyagraha, nonviolence, ashram life, political leaders, correspondence, the Salt March, South Africa, Round Table Conferences, prayer meetings, imprisonment, spinning and khadi, independence, Gandhi's assassination and funeral, and related historical themes.
Notable People Represented: Mahatma Gandhi, Kasturba Gandhi, Jawaharlal Nehru, Vallabhbhai Patel, Vinoba Bhave, Khan Abdul Ghaffar Khan, Mirabehn, Sarojini Naidu, Rabindranath Tagore, Subhas Chandra Bose, Muhammad Ali Jinnah and other historical figures.
License
CUSTOM COMMERCIAL DATA LICENCE
This archive is not released under CC0 or an open-data licence.
Access and permitted use are governed by a custom licence defining the specific archive materials included, permitted use cases, territory, term, technical delivery, attribution requirements and any asset-specific restrictions.
Different archive materials may carry different attribution, copyright or usage conditions. The rights applicable to the licensed subset will be identified as part of the final licence and delivery manifest.
Non-AI publication, broadcasting, exhibition, redistribution or other uses require appropriate written licensing terms and are not automatically included.
AI Training Rights
Commercial AI and machine-learning rights are available subject to the final written licence applicable to the specific transaction and licensed archive subset.
Where expressly granted, permitted uses may include:
- Training, fine-tuning and evaluating machine-learning and multimodal AI models.
- Training vision-language models, computer-vision models, OCR/document-intelligence systems, speech models and multimodal retrieval systems.
- Creating model weights, embeddings and other machine-learning representations derived from the licensed data.
- Commercial deployment of resulting models and applications where expressly authorised by the applicable licence.
Unless expressly authorised in writing:
- The underlying archive data may not be resold, redistributed, published, sublicensed or made available to third parties.
- The archive may not be used outside the permitted purposes specified in the licence.
- Copyright notices, provenance information, attribution requirements and embedded rights metadata may not be removed or misrepresented.
- Publication, broadcasting, merchandising, standalone media distribution and other non-AI uses require separate permission where applicable.
Exact AI rights, licensed assets and permitted uses are confirmed in the applicable order and custom licence.
Who Can Use It
Foundation Model Developers: For multimodal pre-training, fine-tuning, retrieval, evaluation and historically grounded model development.
Computer Vision and VLM Teams: For historical image understanding, image-text alignment, retrieval, classification, restoration and colourisation.
LLM and RAG Developers: For primary-source retrieval, archival reasoning, knowledge-grounding and citation-oriented historical systems.
Document AI Teams: For OCR, document classification, archival search and document understanding.
Speech and Video AI Teams: For historical audio/video retrieval, transcription and multimodal understanding.
Universities and Research Institutions: For approved academic, historical, digital-humanities and computational-research applications.
Museums, Archives and Cultural Institutions: For separately licensed digital heritage, education, research and discovery applications.
Data Dictionary
| Column Name | Data Type | Description | Possible Values / Notes |
|---|---|---|---|
| asset_id | String | Unique catalogue identifier for an archive asset | Primary identifier; some derivative assets have related IDs |
| asset_type | Categorical String | High-level asset category | Image, Writing, Audio, Video |
| file_format | String | Digital file format | jpg, pdf, mp3, mp4 |
| title | String | Archival title or identifying label | Human-curated text |
| description | String | Description or contextual information about the asset | Free text; availability varies |
| language | String | Language associated with the asset | Primarily relevant to textual/audio material; may be blank |
| media_type | Categorical String | More specific archival media classification | Photograph, Letter, Article, Diary, Artwork, Audio Speech, Footage, Documentary, etc. |
| date_iso | String / Date | Date associated with the asset | May contain year, year-month or full date; some records undated |
| location | String | Geographic location associated with the asset | City, region, country or historical place name |
| format_type | String | Additional archival format classification | Availability varies |
| recipient | String | Recipient of correspondence where applicable | Primarily used for letters |
| persons | String / Entity List | People represented, mentioned or associated with the asset | Historical-person entities; normalization may be required |
| duration | Duration | Duration of audio or video material | Populated primarily for Audio and Video |
| keywords | String / Tag List | Human-curated topical or descriptive keywords | Multiple subjects may be represented |
| page_count | Integer | Number of pages associated with written material | Primarily applicable to Writing assets |
| resolution | String | Image/video resolution field | Currently not consistently populated |
| colorized_by | String | Identifies colourisation provenance where applicable | Includes GandhiServe for relevant derivative images |
| original_creator | String | Creator, photographer, author or originator where known | May include named creator or unknown |
| source | String | Archival source or collection provenance | Human-curated provenance information |
| copyright_status | String | Copyright, credit or rights-status information associated with the asset | Asset-specific free-text rights information |
| license_type | Categorical String | High-level licensing classification recorded in the catalogue | Includes Credit required, All Rights Reserved, and territory-specific variants |
Data Quality and Preparation Notes
This is an archival source collection rather than a fully homogenised synthetic benchmark dataset.
Some metadata fields are incomplete where the information is historically unavailable or not applicable to that media type. Names, places and historical entities may also contain archival naming variants.
Many historical writings are currently represented as page images or PDF scans rather than uniformly production-grade OCR text. Buyers requiring OCR, normalized entities, image embeddings, transcripts, scene segmentation or other enrichment layers can discuss an enriched delivery specification.
The commercial delivery package can include a manifest, metadata export, rights information, checksums and documentation appropriate to the agreed licence and use case.
Loading...
£125,000
Download Dataset in MULTIMODAL Format
Recommended Datasets
Loading recommendations...
