Audio Data: English (US) Conversational Speech Audio Dataset
Audio, Speech & Acoustic Datasets
Tags and Keywords

£18,600
About
This audio data product provides a 1,000-hour curated package of real, unscripted human-to-human customer service interactions. If your team is looking for a high-quality audio dataset, speech recognition models require real-world acoustic environments to achieve high accuracy. Unlike synthetic or read-speech datasets, this authentic audio dataset captures spontaneous conversations, natural pauses, overlapping speech, and high-emotion scenarios.
Data Product Features
This specific package contains 1,000 hours of US English audio, paired with verbatim text transcripts.
- Audio Specs: Dual-channel (stereo) separation, isolating the Agent and the Customer on independent tracks. Available in native or upsampled sample rates based on your requirements (FLAC/OGG).
- High-Stakes Emotional Register: Contains conflict-resolution and support interactions, providing the rare edge-case conversational data where baseline ASR models typically fail.
- Rich Metadata: Includes device type, call duration, user-reported intent, and AI-generated call summaries.
- PII Redaction: 100% of the audio data and text transcripts are strictly processed to remove Personally Identifiable Information (PII) to ensure full commercial compliance.
Distribution
- Data Volume: 1,000 hours of audio + corresponding metadata/transcripts.
- Format: Audio files delivered in FLAC or OGG. Transcripts and labels delivered via CSV/JSON.
- Note: We hold over 30,000+ hours in our internal database. This 1,000-hour dataset is packaged specifically for rapid model testing, fine-tuning, and deployment. Scalable up to enterprise volumes upon request
Usage
This audio dataset is ideal for AI/ML teams:
- Application: Fine-tuning Automatic Speech Recognition (ASR) models on complex, real-world telephony speech.
- Application: Training Speaker Diarization algorithms using our cleanly separated stereo tracks.
- Application: Developing Voice AI Agents and Customer Support Copilots that need to understand emotional nuances, interruptions, and code-switching under stress.
Coverage
- Geographic Coverage: United States
- Language: English (US)
License
Proprietary
AI Training Rights
Licensee is granted a non-exclusive, worldwide, and perpetual right to:
- Use the Data Product to train, fine-tune, and evaluate machine learning models, including large language models.
- Incorporate Data Product content into models and commercialize resulting model outputs.
- Create derivative works (model weights, embeddings, etc.) for any lawful purpose.
Restrictions:
- The Data Product itself may not be sold, redistributed, or shared outside of licensed usage.
- Licensee must comply with all applicable laws, including data protection and privacy regulations.
Who Can Use It
- Data Scientists & ML Engineers: For training Voice AI, ASR, and NLP models.
- CX & Product Teams: For analyzing customer support interactions and automating QA processes.
- Researchers: For studying human-computer interaction and natural language processing.
Data Dictionary
Provide a data dictionary that defines each column or key in the data product, including data types, possible values, and any relevant notes.
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
| Call id | String | Unique identifier for each recorded support call. | e.g., CA131dad885... |
| Company Name | String | The name of the brand or company the customer is calling. | e.g., Groupon, DoorDash |
| Device | String | The type of device the customer used to make the call. | mobile, desktop, etc. |
| OS | String | The operating system of the customer's device. | ios, android, windows |
| Date of call | Datetime | The exact date and time when the call took place. | Format: MM/DD/YYYY HH:MM |
| City of the user | String | The city where the caller is located. | |
| State of the user | String | The state or province where the caller is located. | |
| Country of the user | String | The country where the caller is located. | |
| Total call duration(in seconds) | Number | Total length of the call from start to finish, including wait times. | |
| Call duration with company (in seconds) | Number | Actual duration of the active conversation with the agent, excluding hold time. | |
| Reason of the Call | String | Specific issue or reason for the call. Highly valuable as a label for intent detection models. | e.g., Missing order, Refund request |
| Reason Group | String | High-level categorization of the call reason for broader trend analysis. | e.g., charge, delivery, product_service quality |
| Representative waiting time (in seconds) | Number | Time spent on hold waiting for an agent to answer. | |
| Transcription | String | Full, verbatim word-for-word text transcription of the dialogue with speaker diarization. | Contains speaker tags like [Speaker 0], [Speaker 1] |
| Summary of the call (AI generated) | String | AI-generated text summary explaining the main issue and resolution. | Ideal for training text summarization models |
Note for buyers: The instant download provided is a metadata and transcript sample. For access to the corresponding raw OGG/FLAC audio files, or to purchase larger historical subsets, please contact us directly after reviewing the sample schema.
Listing Stats
VIEWS
6
DELIVERY
CUSTOM, S3
LISTED
30/07/2026
UPDATED
30/07/2026
REGION
NORTH AMERICA
TRUST
5 / 5
Loading...
£18,600
Download Dataset in Unknown Format
Recommended Datasets
Loading recommendations...
