AI Prompt & Conversation Dataset
Natural Language Processing
Tags and Keywords

£387,600
About
AI Prompt & Conversation Dataset
Description
The AI Prompt & Conversation Dataset is a structured collection of AI prompts and conversational interactions contains 259K real-world sessions capturing complex, multi-turn human–AI workflows comprising 13.4M+ messages, 2.75B+ tokens, and 607K+ linked multimodal artifacts, including code, PDFs, and 370K+ images, designed for artificial intelligence, machine learning, natural language processing (NLP), and large language model (LLM) development. The dataset captures prompt-and-response interactions that can support conversational AI, chatbot development, prompt engineering, LLM fine-tuning, model evaluation, dialogue understanding, and AI research.
The dataset provides valuable conversational language data for analyzing how users interact with AI systems, understanding dialogue patterns, developing instruction-following models, and evaluating the quality and relevance of AI-generated responses.
Note: Pricing varies depending on several factors, including the number of conversation sessions, number of prompts and responses, data volume, conversation depth, source platform, metadata availability, attachment information, file format, licensing terms, and customization needs. The final price will be determined based on the specific dataset requirements and selected package.
Data Product Features
| Feature | Description |
|---|---|
| Conversation Sessions | Groups related prompts and responses into individual conversational sessions, where available. |
| Session ID | Identifier used to associate multiple interactions with the same conversation session. |
| User Prompts | User-provided instructions, questions, requests, or queries submitted to an AI system. |
| AI Responses | AI-generated responses corresponding to user prompts. |
| Conversation Turns | Individual exchanges between the user and AI within a session. |
| Prompt-Response Pairs | Paired user prompts and corresponding AI responses for conversational AI analysis and model development. |
| Conversation Context | Previous conversational information that provides context for subsequent interactions, where available. |
| Source Platform | Platform or AI system from which the prompts and conversation sessions were collected or sourced. |
| Attachment Name | Name of the file or attachment associated with a conversation or prompt, where applicable. |
| Timestamps | Date and/or time information associated with conversations, where available. |
| Conversation Metadata | Additional attributes associated with sessions or interactions, where available. |
Distribution
- Data Volume: Dataset size varies based on the number of conversation sessions, prompts, responses, and interaction records included.
- Format: Structured tabular data; exact file format may vary by selected package.
- Data Type: AI / NLP / Conversational Data
- Structure: Conversation-level and/or interaction-level records containing prompts, responses, session information, and associated metadata.
- Record Structure: Records may represent individual prompt-response interactions or complete conversation sessions.
- Session Structure: Multiple conversation turns may be associated with a single session where session identifiers are available.
- Dataset Size: Dataset size may vary depending on the selected package, number of sessions, number of interactions, metadata, and specific requirements.
Usage
This data product is ideal for a variety of applications:
- LLM Fine-Tuning: Training and fine-tuning language models using prompt-response and conversational data.
- Conversational AI: Developing AI assistants, chatbots, and dialogue systems.
- Prompt Engineering: Studying prompt structures and their relationship to AI-generated responses.
- NLP Research: Analyzing natural language, dialogue patterns, intent, and conversational context.
- AI Model Evaluation: Evaluating response quality, instruction following, relevance, and conversational consistency.
- Dialogue Understanding: Developing models that understand multi-turn conversations and contextual interactions.
- Chatbot Development: Building and improving conversational interfaces and virtual assistants.
- AI Research: Supporting research into human-AI interaction, language models, and conversational systems.
Coverage
-
Geographic Coverage: Global
-
Language Coverage: Dataset-specific; may include one or multiple languages depending on the source conversations.
-
Conversation Coverage: User prompts, AI responses, multi-turn interactions, and conversational context where available.
-
Domain Coverage: Software & Engineering, Writing & Knowledge Work, Business & Career, Design & Media, Personal & Everyday, Other / unmapped.
License
CC BY 4.0 (Creative Commons Attribution 4.0 International)
AI Training Rights
InfoBay.AI ensures that all datasets are sourced, curated, and managed with proper ownership verification, licensing documentation, and data provenance records. We hold the necessary rights to license and sublicense the datasets we provide through formal agreements with our data vendors, which grant us the required permissions for commercial licensing and AI training use cases. To ensure transparency and compliance, we maintain relevant documentation and have previously shared redacted agreements for selected datasets as evidence of our data rights and licensing authority.
Data Dictionary
| Column Name | Data Type | Description | Possible Values/Notes |
|---|---|---|---|
session_id | STRING | Unique identifier for a conversation session. | Session-specific identifier |
conversation_id | STRING | Identifier for a complete conversation, where available. | Dataset-specific |
user_prompt | TEXT | Prompt, question, instruction, or request submitted by the user. | Free-text |
ai_response | TEXT | Response generated by the AI system for the corresponding prompt. | Free-text |
prompt_response_pair | TEXT/JSON | Combined or structured representation of a prompt and its corresponding response. | Dataset-specific |
conversation_context | TEXT | Previous conversational context associated with the interaction. | Free-text; may be unavailable |
timestamp | DATETIME | Date and time associated with the interaction, where available. | Dataset-specific |
language | STRING | Language used in the conversation. | Multilingual |
source_platform | STRING | Identifies the platform or AI system from which the conversation or prompt was sourced. | Claude,ChatGPT, Gemini / AI Studio, DeepSeek, Perplexity,Lovable, Web app, Emergent, Grok, Replit, MCP server — Cursor, Antigravity,NotebookLM,Copilot |
attachment_name | STRING | Name of the file or attachment associated with the conversation or interaction. | Filename such asImages-PNG, WebP, JPEG, Text & markup-txt, md, json, xml, yaml,Source code-18 languages, Office — Word, PowerPoint, sheets, CSV, PDF; may be blank/null when no attachment is present. |
prompt_type | STRING | Classification or category of the user prompt, where available. | Dataset-specific |
topic | STRING | Topic or subject associated with the conversation. | Dataset-specific |
metadata | JSON/STRING | Additional information associated with the conversation or interaction. | Dataset-specific |
Considerations
This dataset is provided for research and educational purposes only. It contains only sample data.
Loading...
£387,600
Download Dataset in TEXT Format
Recommended Datasets
Loading recommendations...
