14K+ Hours Malayalam Dual-Channel Call Center Dataset

Audio, Speech & Acoustic Datasets

Tags and Keywords

Malayalam

Speech

Dataset

Audio

Language

Dual-channel

Stereo

Wav

14K+ Hours Malayalam Dual-Channel Call Center Dataset Dataset on Opendatabay data marketplace

£112,100

About

Malayalam Dual-Channel Call Center Dataset

Description

The Malayalam Dual-Channel Call Center Dataset is a collection of Malayalam call center audio recordings provided in Dual-channel of 14,980 hours. The dataset is designed to support research and development in speech processing, automatic speech recognition (ASR), natural language processing (NLP), conversational AI, and machine learning.
The dataset can be used for developing and evaluating AI systems that process Malayalam speech in customer service and call center environments. It is suitable for academic research, commercial AI applications, speech analytics, and voice technology development.
Note: Pricing varies depending on several factors, including the language, total audio hours, metadata availability, number of attributes, annotation requirements, and customization needs. The final price will be determined based on the specific dataset requirements.

Data Product Features

The dataset includes the following key components:
FeatureDescription
Audio FilesMalayalam call center speech recordings in WAV, OBB, MP3 format
Audio FormatWAV , OBB, MP3
Channel TypeDual-Channel
LanguageMalayalam
LicenseCC BY 4.0 (Creative Commons Attribution 4.0 International)

Distribution

The dataset is distributed as a compressed archive containing audio files.
  -Format: WAV, MP3, OBB
  -Structure: Audio files

Data Volume

-Records: Call Center recordings
-Volume: 14,980 hours
-Channel Configuration: Dual-Channel 

Usage

This data product is suitable for a wide range of AI and speech technology applications.
-Automatic Speech Recognition (ASR): Train and evaluate speech recognition systems.
-Speech-to-Text: Develop transcription models.
-Conversational AI: Build intelligent virtual assistants and customer support systems.
-Speech Analytics: Analyze customer interactions and speech patterns.
-Machine Learning: Train and benchmark speech-based AI models.
-Natural Language Processing (NLP): Support spoken language understanding and downstream NLP tasks.
-Voice Technology: Develop and evaluate voice-enabled applications.
-Academic Research: Conduct research in speech processing, AI, and computational linguistics.

Coverage

The dataset covers Malayalam-language call center recordings.
-Geographic Coverage: Malayalam-speaking regions
-Time Range: Not specified

License

CC BY 4.0 (Creative Commons Attribution 4.0 International)

AI Training Rights

InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.

Data Dictionary

Column NameData TypeDescriptionPossible Values/Notes
file_nameStringName of the audio recordingMalayalam Call Center Dataset
languageStringLanguage of the recordingMalayalam
audio_formatStringAudio file formatWAV , OBB, MP3
channel_typeStringAudio channel configurationDual-Channel
file_sizeIntegerSize of the audio filekilobyte
licenseStringDataset licenseCC BY 4.0

Considerations


This dataset is provided for research and educational purposes only. It contains only sample data.

Listing Stats

VIEWS

1

DELIVERY

CUSTOM, S3

LISTED

12/07/2026

UPDATED

24/07/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£112,100

Download Dataset in AUDIO Format