AI Prompt & Conversation Dataset

Natural Language Processing

Tags and Keywords

Ai

Prompts

Conversations

Llm

Chatbot

Nlp

Training

Dialogue

AI Prompt & Conversation Dataset Dataset on Opendatabay data marketplace

£387,600

About

AI Prompt & Conversation Dataset

Description

The AI Prompt & Conversation Dataset is a structured collection of AI prompts and conversational interactions contains 259K real-world sessions capturing complex, multi-turn human–AI workflows comprising 13.4M+ messages, 2.75B+ tokens, and 607K+ linked multimodal artifacts, including code, PDFs, and 370K+ images, designed for artificial intelligence, machine learning, natural language processing (NLP), and large language model (LLM) development. The dataset captures prompt-and-response interactions that can support conversational AI, chatbot development, prompt engineering, LLM fine-tuning, model evaluation, dialogue understanding, and AI research.
The dataset provides valuable conversational language data for analyzing how users interact with AI systems, understanding dialogue patterns, developing instruction-following models, and evaluating the quality and relevance of AI-generated responses.
Note: Pricing varies depending on several factors, including the number of conversation sessions, number of prompts and responses, data volume, conversation depth, source platform, metadata availability, attachment information, file format, licensing terms, and customization needs. The final price will be determined based on the specific dataset requirements and selected package.

Data Product Features

FeatureDescription
Conversation SessionsGroups related prompts and responses into individual conversational sessions, where available.
Session IDIdentifier used to associate multiple interactions with the same conversation session.
User PromptsUser-provided instructions, questions, requests, or queries submitted to an AI system.
AI ResponsesAI-generated responses corresponding to user prompts.
Conversation TurnsIndividual exchanges between the user and AI within a session.
Prompt-Response PairsPaired user prompts and corresponding AI responses for conversational AI analysis and model development.
Conversation ContextPrevious conversational information that provides context for subsequent interactions, where available.
Source PlatformPlatform or AI system from which the prompts and conversation sessions were collected or sourced.
Attachment NameName of the file or attachment associated with a conversation or prompt, where applicable.
TimestampsDate and/or time information associated with conversations, where available.
Conversation MetadataAdditional attributes associated with sessions or interactions, where available.

Distribution

  • Data Volume: Dataset size varies based on the number of conversation sessions, prompts, responses, and interaction records included.
  • Format: Structured tabular data; exact file format may vary by selected package.
  • Data Type: AI / NLP / Conversational Data
  • Structure: Conversation-level and/or interaction-level records containing prompts, responses, session information, and associated metadata.
  • Record Structure: Records may represent individual prompt-response interactions or complete conversation sessions.
  • Session Structure: Multiple conversation turns may be associated with a single session where session identifiers are available.
  • Dataset Size: Dataset size may vary depending on the selected package, number of sessions, number of interactions, metadata, and specific requirements.

Usage

This data product is ideal for a variety of applications:
  • LLM Fine-Tuning: Training and fine-tuning language models using prompt-response and conversational data.
  • Conversational AI: Developing AI assistants, chatbots, and dialogue systems.
  • Prompt Engineering: Studying prompt structures and their relationship to AI-generated responses.
  • NLP Research: Analyzing natural language, dialogue patterns, intent, and conversational context.
  • AI Model Evaluation: Evaluating response quality, instruction following, relevance, and conversational consistency.
  • Dialogue Understanding: Developing models that understand multi-turn conversations and contextual interactions.
  • Chatbot Development: Building and improving conversational interfaces and virtual assistants.
  • AI Research: Supporting research into human-AI interaction, language models, and conversational systems.

Coverage

  • Geographic Coverage: Global
  • Language Coverage: Dataset-specific; may include one or multiple languages depending on the source conversations.
  • Conversation Coverage: User prompts, AI responses, multi-turn interactions, and conversational context where available.
  • Domain Coverage: Software & Engineering, Writing & Knowledge Work, Business & Career, Design & Media, Personal & Everyday, Other / unmapped.

License

CC BY 4.0 (Creative Commons Attribution 4.0 International)

AI Training Rights

InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.

Data Dictionary

Column NameData TypeDescriptionPossible Values/Notes
session_idSTRINGUnique identifier for a conversation session.Session-specific identifier
conversation_idSTRINGIdentifier for a complete conversation, where available.Dataset-specific
user_promptTEXTPrompt, question, instruction, or request submitted by the user.Free-text
ai_responseTEXTResponse generated by the AI system for the corresponding prompt.Free-text
prompt_response_pairTEXT/JSONCombined or structured representation of a prompt and its corresponding response.Dataset-specific
conversation_contextTEXTPrevious conversational context associated with the interaction.Free-text; may be unavailable
timestampDATETIMEDate and time associated with the interaction, where available.Dataset-specific
languageSTRINGLanguage used in the conversation.Multilingual
source_platformSTRINGIdentifies the platform or AI system from which the conversation or prompt was sourced.Claude,ChatGPT, Gemini / AI Studio, DeepSeek, Perplexity,Lovable, Web app, Emergent, Grok, Replit, MCP server — Cursor, Antigravity,NotebookLM,Copilot
attachment_nameSTRINGName of the file or attachment associated with the conversation or interaction.Filename such asImages-PNG, WebP, JPEG, Text & markup-txt, md, json, xml, yaml,Source code-18 languages, Office — Word, PowerPoint, sheets, CSV, PDF; may be blank/null when no attachment is present.
prompt_typeSTRINGClassification or category of the user prompt, where available.Dataset-specific
topicSTRINGTopic or subject associated with the conversation.Dataset-specific
metadataJSON/STRINGAdditional information associated with the conversation or interaction.Dataset-specific

Considerations

This dataset is provided for research and educational purposes only. It contains only sample data.

Listing Stats

VIEWS

17

DELIVERY

CUSTOM, S3

LISTED

17/09/2026

UPDATED

18/09/2026

REGION

GLOBAL

Universal Data Trust Rating UDTRTRUST

5 / 5

Loading...

£387,600

Download Dataset in TEXT Format