Multilingual Speech & Audio Dataset

Audio, Speech & Acoustic Datasets

Tags and Keywords

Speech

Audio

Customerservice

Callcentre

Multilingual

Asr

Conversationalai

Voiceai

Multilingual Speech & Audio Dataset Dataset on Opendatabay data marketplace

£500,000

About

A large-scale two-way conversational audio corpus of real customer–agent interactions across industries. Covers extensive Indian language diversity including regional and low-resource languages, plus South Asian and Middle Eastern languages. Transcripts available alongside audio. Purpose-built for ASR model training, voice AI, and multilingual speech research.
Data Product Features
  • Real customer–agent conversations across multiple industries
  • Extensive Indian language coverage including regional and low-resource languages
  • South Asian and Middle Eastern language coverage
  • Transcripts available alongside audio files
  • Two-way (full duplex) conversational format
Distribution Format: WAV / MP3 / JSON Size: Multi-million words Records: Multi-million utterances
Data Volume Multi-million words of conversational speech across Indian and international languages
Usage
  • ASR (Automatic Speech Recognition) model training for Indian languages
  • Voice AI and conversational agent development
  • Multilingual speech synthesis and TTS model training
  • Low-resource language research
Coverage Geographic Coverage: India and South Asia Time Range: Historical — expandable Languages: Indian regional languages, South Asian and Middle Eastern languages
License CC0 — No Rights Reserved
AI Training Rights Licensee is granted a non-exclusive, worldwide, and perpetual right to use this data product to train, fine-tune, and evaluate machine learning models. The data product itself may not be redistributed or shared outside licensed usage. Licensee must comply with all applicable laws, including data protection and privacy regulations.
Who Can Use It
  • AI/ML Engineers: For training multilingual ASR and speech models
  • Voice AI Companies: For building Indian language voice assistants
  • Researchers: For low-resource language speech research
Data Dictionary
  • audio_id (string) — Unique identifier for each audio file
  • language (string) — Language of the interaction
  • duration_seconds (float) — Length of audio clip in seconds
  • speaker_role (string) — Customer or Agent
  • transcript (string) — Text transcript of the audio
  • industry (string) — Industry vertical of the interaction
  • channel (string) — WAV or MP3
  • sampling_rate (integer) — Audio sampling rate in Hz

Listing Stats

VIEWS

8

DELIVERY

CUSTOM, S3

LISTED

06/09/2026

UPDATED

12/09/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£500,000

Download Dataset in AUDIO Format