59,000 Hours of Real Business Coaching & Sales Calls for AI Training

Audio, Speech & Acoustic Datasets

Tags and Keywords

Conversational

Conversationalai

Speech

Sales

Coaching

Training

Emotion

Audio

59,000 Hours of Real Business Coaching & Sales Calls for AI Training Dataset on Opendatabay data marketplace

$1,972,400

About

59,000 hours of real, unscripted conversations from a business coaching company, recorded on Zoom from about 2015 to 2026. That is nearly seven years of nonstop talk.
Every call ships with its audio plus three files: a full JSON record, a readable Markdown transcript, and an HTML viewer.
The library holds two kinds of calls.
Most are coaching calls, where expert coaches work live with entrepreneurs on real problems in their businesses, usually in group sessions where students take turns in the "hot seat."
The rest are sales calls, where our sales team talks with entrepreneurs deciding whether to join a coaching program.
Together they show, turn by turn, how skilled people help someone change and help someone decide.
AI agents are now being built to do both, and real examples are very hard to find. None of this material has ever been posted on the open web, so it is not in public training data. No video is included.

Data Product Features

  • Real experts, real stakes: No actors and no scripts. People bring real businesses, real money worries and real decisions.
  • Group session structure: Hot seats, Q&A and teaching blocks are marked with start and end times and the student in focus.
  • Speaker roles: Every speaker is labeled coach, student or other participant. Pseudonyms are unique to each call.
  • Stages, summary and themes: Each call has a summary, a main theme and supporting themes, and is split into titled stages with times and a short summary of each.
  • Tone on every segment: Each transcript segment has a tone label (such as encouraging, frustrated, curious, empathetic or relieved) and a strength of low, medium or high.
  • Conversation flow: Each segment is labeled with what it does in the conversation, such as check-in, exploration, teaching, problem statement, reframe, commitment or close.
  • Coaching techniques and their effect: Each technique the coach uses (such as a diagnostic question, normalizing, giving an example, a reframe or celebrating a win) is labeled. The data shows whether the student's emotional score then went up, stayed flat, went down, or drew no response.
  • Key moments and turning points: Insights, breakthroughs and resistance are marked, along with coach redirects, topic shifts and student pushback, each with a short note.
  • Emotional arcs: Each speaker's emotional path through the call is scored from −3 to +3. Notes explain each emotion shift and the moment that triggered it.
  • Voice measurements: Speaking speed, average pitch, pitch variation, relative loudness and the pause before each segment.
  • Quality data: Every call records its signal-to-noise ratio, clipping, speech recognition confidence and the privacy checks it went through.
  • Living dataset: More hours are added as we process our archive.

Distribution

  • Format: Each call ships as four files: audio (mono, from the original Zoom recording, AAC at 32 kHz); a JSON file with every label; a Markdown transcript with tags on every segment; and a self-contained HTML viewer with a table of contents and an emotional-arc chart.
  • Structure: One folder per call, named by a random call ID (e.g. CC-WYKCVAP5), grouped by year.
  • Size: About 3.4 TB, almost all of it audio. The three text files add roughly 25 GB.
  • Delivery: S3 bucket.
  • Data Volume: About 59,000 hours; about 50,000 calls; roughly 13 million transcript segments (about 220 per hour); 11 fields per segment, plus call-level data.

Usage

This data product is ideal for a variety of applications:
  • AI coaching and support agents: Train agents to ask good questions, notice how someone feels, and respond with care.
  • Technique-to-outcome modeling: Learn which coaching techniques help people move forward, using the before-and-after emotion scores.
  • AI sales assistants: Train and test sales agents on real conversations where people voice doubts, push back and decide.
  • Sales call scoring: Build tools that review sales calls and coach reps.
  • Emotion recognition: Recognize feelings from both words and voice in real, multi-person conversations.
  • Speech recognition and speaker separation: Improve accuracy on real Zoom calls with many speakers, microphones and accents.
  • Expressive voice models: Help synthetic voices sound warm and natural.
  • Evaluation sets: Test how well models follow long, real conversations.

Coverage

  • Geographic Coverage: Global; most participants are in the United States, Canada, the United Kingdom and Australia.
  • Time Range: About 2015 – 2026.
  • Demographics: Adult entrepreneurs and small-business owners, from first-time founders to experienced owners, talking with expert business coaches and a sales team. Language: English.

License

Proprietary

AI Training Rights

Licensee is granted a non-exclusive, worldwide, and perpetual right to:
  • Use the Data Product to train, fine-tune, and evaluate machine learning models, including large language models.
  • Incorporate Data Product content into models and commercialize resulting model outputs.
  • Create derivative works (model weights, embeddings, etc.) for any lawful purpose.
Restrictions:
  • The Data Product itself may not be sold, redistributed, or shared outside of licensed usage.
  • Licensee may not attempt to identify any person in the data.
  • Licensee must comply with all applicable laws, including data protection and privacy regulations.
Exclusive or time-limited licenses are available on request.

Who Can Use It

  • AI labs and model builders: For training and fine-tuning conversational and voice models.
  • Data scientists and ML engineers: For building emotion, speech and dialogue models.
  • Sales technology companies: For AI sales assistants and call-scoring tools.
  • Coaching, education and support platforms: For AI coaches, tutors and support agents.
  • Researchers: For studies of conversation, emotion, group dynamics and how people change.

Data Dictionary

Column NameData TypeDescriptionPossible Values/Notes
call_idstringRandom ID for the calle.g. CC-WYKCVAP5
collectionstringWhich collection the call belongs toe.g. coaching_calls
call_typestringKind of sessione.g. coaching_session
durationnumberCall length in secondse.g. 4289.34
recording_period.year, .quarterintegerWhen the call was recordedYear and quarter only, for privacy
languagestringLanguage codeen
audio_quality.snr_db, .clipping, .channelsnumberSignal-to-noise ratio, clipping share and channel count
source_formatobjectOriginal recording formatContainer, codec, bitrate, sample rate, channels
asr_engine, asr_mean_confidencestring, numberSpeech recognizer used and its average confidenceConfidence from 0 to 1
participants[]arrayEveryone who spoke: placeholder, pseudonym, role, how the role was assigned, and confidenceRoles: coach, student, participant
primary_theme, secondary_themesstring, arrayMain and supporting topicse.g. mindset_confidence, offer_design, launch_planning
summarystringShort summary of the whole call
stages[]arrayTitled stages of the call, with start and end times, segment range and summary
blocks[]arraySession formats within the call, with the student in focushot_seat, qa, teaching_block
segments[].idintegerOrder of the segment in the call
segments[].start, .endnumberStart and end time, in seconds
segments[].speakerstringPlaceholder for the speakere.g. [Coach A], [Student S-9831]
segments[].textstringWhat was said, with identifying details replacede.g. [Person 3], [Business 2], [Product 1], [Location 1]
segments[].tones[]arrayTone and strengthTone e.g. encouraging, frustrated, curious, empathetic, realization, relieved; strength: low, medium, high
segments[].flow_rolestringWhat the segment does in the conversatione.g. check_in, exploration, teaching, problem_statement, reframe, commitment, close
segments[].behaviorstringCoaching technique used, if anye.g. diagnostic_question, normalize, give_example, reframe, celebrate_win, reflect_back, commitment_confirm
segments[].momentstringKey moment, if anye.g. insight, breakthrough, resistance
segments[].overlapbooleanWhether speakers talked over each other
segments[].acousticsobjectWords per minute, average pitch, pitch variation, relative loudness, pause beforeCan be empty for very short segments
emotion_notes[]arrayEmotion shifts: speaker, from tone, to tone, the segment that triggered it, and a note
transitions[]arrayTurning points in the conversation, with a notee.g. coach_redirect, topic_shift, breakthrough, student_pushback
emotional_arcs[]arrayEach speaker's emotional path through the call, point by pointScore from −3 to +3
behavior_effects[]arrayHow each technique landed: the student, their emotional score before and after, and the directionDirection: up, flat, down, no_response
redaction_summaryobjectCount of each kind of detail replacede.g. person, business, product, location, date, amount, contact
qcobjectPrivacy and consistency checks run on the call, and their results
human_checkedbooleanWhether a person has reviewed the call
consent_status, consent_terms_versionstringConsent record for the call
schema_version, taxonomy_version, pipeline_versionstringVersions of the file format, label set and pipeline

  • How the labels are made: Our own automated pipeline makes every label, running entirely on our own computers. Transcripts come from NVIDIA Parakeet speech recognition, with Whisper as a backup. The labels are made by AI models, not typed by hand, and each call records whether a person has reviewed it. Buyers can check label quality in the sample before licensing.
  • Privacy: No video is included. Names of people, businesses, products and places, plus dates, amounts, links and contact details, are replaced with numbered placeholders like [Person 3] or [Business 2]. Spoken names are covered in the audio before delivery. Calls are dated by year and quarter only, and pseudonyms are unique to each call, so no one can be followed from call to call. Every call goes through automated privacy checks, including an AI privacy review.
  • Consent: Every coaching call is covered by signed agreements from everyone on it. Every sales call starts with a notice that the call is recorded and may be shared or used for training. A consent summary is available on request, and full terms can be reviewed under NDA.
  • Slices: License the whole library, or just part of it: coaching only, sales only, certain years, or text only.
  • Sample: 5 full calls (3 coaching, 2 sales), each with all four files.

Listing Stats

VIEWS

6

DELIVERY

CUSTOM, S3

LISTED

30/09/2026

UPDATED

30/09/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

$1,972,400

Download Dataset in Unknown Format