Audio Data: English (US) Conversational Speech Audio Dataset

Audio, Speech & Acoustic Datasets

Tags and Keywords

Audio

Data

Dataset

Speech

Conversational

Recognition

Asr

Voice

Audio Data: English (US) Conversational Speech Audio Dataset Dataset on Opendatabay data marketplace

£18,600

About

This audio data product provides a 1,000-hour curated package of real, unscripted human-to-human customer service interactions. If your team is looking for a high-quality audio dataset, speech recognition models require real-world acoustic environments to achieve high accuracy. Unlike synthetic or read-speech datasets, this authentic audio dataset captures spontaneous conversations, natural pauses, overlapping speech, and high-emotion scenarios.

Data Product Features

This specific package contains 1,000 hours of US English audio, paired with verbatim text transcripts.
  • Audio Specs: Dual-channel (stereo) separation, isolating the Agent and the Customer on independent tracks. Available in native or upsampled sample rates based on your requirements (FLAC/OGG).
  • High-Stakes Emotional Register: Contains conflict-resolution and support interactions, providing the rare edge-case conversational data where baseline ASR models typically fail.
  • Rich Metadata: Includes device type, call duration, user-reported intent, and AI-generated call summaries.
  • PII Redaction: 100% of the audio data and text transcripts are strictly processed to remove Personally Identifiable Information (PII) to ensure full commercial compliance.

Distribution

  • Data Volume: 1,000 hours of audio + corresponding metadata/transcripts.
  • Format: Audio files delivered in FLAC or OGG. Transcripts and labels delivered via CSV/JSON.
  • Note: We hold over 30,000+ hours in our internal database. This 1,000-hour dataset is packaged specifically for rapid model testing, fine-tuning, and deployment. Scalable up to enterprise volumes upon request

Usage

This audio dataset is ideal for AI/ML teams:
  • Application: Fine-tuning Automatic Speech Recognition (ASR) models on complex, real-world telephony speech.
  • Application: Training Speaker Diarization algorithms using our cleanly separated stereo tracks.
  • Application: Developing Voice AI Agents and Customer Support Copilots that need to understand emotional nuances, interruptions, and code-switching under stress.

Coverage

  • Geographic Coverage: United States
  • Language: English (US)

License

Proprietary

AI Training Rights

Licensee is granted a non-exclusive, worldwide, and perpetual right to:
  • Use the Data Product to train, fine-tune, and evaluate machine learning models, including large language models.
  • Incorporate Data Product content into models and commercialize resulting model outputs.
  • Create derivative works (model weights, embeddings, etc.) for any lawful purpose.
Restrictions:
  • The Data Product itself may not be sold, redistributed, or shared outside of licensed usage.
  • Licensee must comply with all applicable laws, including data protection and privacy regulations.

Who Can Use It

  • Data Scientists & ML Engineers: For training Voice AI, ASR, and NLP models.
  • CX & Product Teams: For analyzing customer support interactions and automating QA processes.
  • Researchers: For studying human-computer interaction and natural language processing.

Data Dictionary

Provide a data dictionary that defines each column or key in the data product, including data types, possible values, and any relevant notes.
Column NameData TypeDescriptionPossible Values/Notes
Call idStringUnique identifier for each recorded support call.e.g., CA131dad885...
Company NameStringThe name of the brand or company the customer is calling.e.g., Groupon, DoorDash
DeviceStringThe type of device the customer used to make the call.mobile, desktop, etc.
OSStringThe operating system of the customer's device.ios, android, windows
Date of callDatetimeThe exact date and time when the call took place.Format: MM/DD/YYYY HH:MM
City of the userStringThe city where the caller is located.
State of the userStringThe state or province where the caller is located.
Country of the userStringThe country where the caller is located.
Total call duration(in seconds)NumberTotal length of the call from start to finish, including wait times.
Call duration with company (in seconds)NumberActual duration of the active conversation with the agent, excluding hold time.
Reason of the CallStringSpecific issue or reason for the call. Highly valuable as a label for intent detection models.e.g., Missing order, Refund request
Reason GroupStringHigh-level categorization of the call reason for broader trend analysis.e.g., charge, delivery, product_service quality
Representative waiting time (in seconds)NumberTime spent on hold waiting for an agent to answer.
TranscriptionStringFull, verbatim word-for-word text transcription of the dialogue with speaker diarization.Contains speaker tags like [Speaker 0], [Speaker 1]
Summary of the call (AI generated)StringAI-generated text summary explaining the main issue and resolution.Ideal for training text summarization models

Note for buyers: The instant download provided is a metadata and transcript sample. For access to the corresponding raw OGG/FLAC audio files, or to purchase larger historical subsets, please contact us directly after reviewing the sample schema.

Listing Stats

VIEWS

6

DELIVERY

CUSTOM, S3

LISTED

30/07/2026

UPDATED

30/07/2026

REGION

NORTH AMERICA

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£18,600

Download Dataset in Unknown Format