20k B2C Call Center Audio & Verbatim Transcripts Dataset for Voice AI
Audio, Speech & Acoustic Datasets
Tags and Keywords

£373,000
About
This data product provides a highly authentic, ground-truth dataset of real human-to-human customer service interactions. Unlike synthetic speech corpora, this dataset captures real-world acoustic environments, unscripted dialogues, genuine emotional stress, and overlapping speech. It is an essential resource for AI teams building robust Automatic Speech Recognition (ASR) systems and conversational AI agents.
The uploaded sample file is a preview catalog only and is not the full dataset. Final pricing depends on selected languages, number of audio hours, enrichment files, delivery format, and licensing scope. Average price per hour is $25.
Data Product Features
The dataset pairs high-fidelity audio recordings with word-for-word text transcripts. Key features include:
- Speaker Separation: Stereo audio formatting isolates the "Agent" and the "Customer" on separate channels.
- Rich Metadata: Includes device type, caller location, call duration, hold times, reason of call, company name etc.
- Outcome Labels: Features user-reported reasons for calling and AI-generated summaries of the conversation, providing context for intent classification.
- PII Redaction: All audio and text files are strictly processed to remove Personally Identifiable Information (PII) to ensure commercial compliance.
Distribution
- Format: Transcripts and metadata are delivered in CSV/JSON. Corresponding audio files are provided in high-quality OGG or FLAC formats (48kHz, 192 kbps).
- Data Volume: The full underlying database contains over 25,000 hours of audio. (This specific listing acts as a sample/entry point).
- Delivery: Custom delivery via secure S3 Bucket.
Usage
This data product is ideal for a variety of AI/ML applications:
- Application: Training and benchmarking Automatic Speech Recognition (ASR) models in noisy, real-world telephony environments.
- Application: Training Speaker Diarization models to accurately distinguish between multiple speakers.
- Application: Fine-tuning LLMs and Conversational AI agents on natural human complaint resolution and negotiation tactics.
- Application: Developing Voice Emotion & Sentiment models by analyzing real customer frustration and stress tones.
Coverage
- Geographic Coverage: USA.
- Time Range: 2025 - Present.
- Demographics: B2C consumers interacting with support teams across 140,000+ brands (Retail, Finance, Travel, E-commerce, etc.).
License
Proprietary
AI Training Rights
Licensee is granted a non-exclusive, worldwide, and perpetual right to:
- Use the Data Product to train, fine-tune, and evaluate machine learning models, including large language models.
- Incorporate Data Product content into models and commercialize resulting model outputs. Create derivative works (model weights, embeddings, etc.) for any lawful purpose.
Restrictions:
- The Data Product itself may not be sold, redistributed, or shared outside of licensed usage.
- Licensee must comply with all applicable laws, including data protection and privacy regulations.
Who Can Use It
- Data Scientists & ML Engineers: For training Voice AI, ASR, and NLP models.
- CX & Product Teams: For analyzing customer support interactions and automating QA processes.
- Researchers: For studying human-computer interaction and natural language processing.
Data Dictionary
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
| Call id | String | Unique identifier for each recorded support call. | e.g., CA131dad885... |
| Company Name | String | The name of the brand or company the customer is calling. | e.g., Groupon, DoorDash |
| Device | String | The type of device the customer used to make the call. | mobile, desktop, etc. |
| OS | String | The operating system of the customer's device. | ios, android, windows |
| Date of call | Datetime | The exact date and time when the call took place. | Format: MM/DD/YYYY HH:MM |
| City of the user | String | The city where the caller is located. | |
| State of the user | String | The state or province where the caller is located. | |
| Country of the user | String | The country where the caller is located. | |
| Total call duration(in seconds) | Number | Total length of the call from start to finish, including wait times. | |
| Call duration with company (in seconds) | Number | Actual duration of the active conversation with the agent, excluding hold time. | |
| Reason of the Call | String | Specific issue or reason for the call. Highly valuable as a label for intent detection models. | e.g., Missing order, Refund request |
| Reason Group | String | High-level categorization of the call reason for broader trend analysis. | e.g., charge, delivery, product_service quality |
| Representative waiting time (in seconds) | Number | Time spent on hold waiting for an agent to answer. | |
| Transcription | String | Full, verbatim word-for-word text transcription of the dialogue with speaker diarization. | Contains speaker tags like [Speaker 0], [Speaker 1] |
| Summary of the call (AI generated) | String | AI-generated text summary explaining the main issue and resolution. | Ideal for training text summarization models |
Note for buyers: The instant download provided is a metadata and transcript sample. For access to the corresponding raw OGG/FLAC audio files, or to purchase larger historical subsets, please contact us directly after reviewing the sample schema.
Loading...
£373,000
Download Dataset in Other Format
Recommended Datasets
Loading recommendations...
