High Quality Arabic Corpus
Foundation Model Datasets
Tags and Keywords

£12,000
About
13M docs, 15B Tokens, 4+ FineWeb-edu score collection of high-quality Arabic text data with their metadata.
Buy the Multilingual Pack for £28,000 instead of £36,000 and save £8,000.
Save over $20,000 in GPU costs with our ready-to-use dataset.
Dataset License
Creation
The dataset was created by filtering all
English common crawl data for high-quality text using the FineWeb-Edu classifier with education score of 4 or higher over 5.
The data is source from the v1.0.0 of the HuggingFaceFW/fineweb-edu dataset which corresponds to CC-MAIN-2024-10 from common crawl.
The data was also fully deduplicated and labeled for Topic and Format using the WebOrganizer Classifiers, and then we only keep documents with a specific format (list below).
All documents were then translated from English to Arabic using the Qwen3-235B-A22B LLM model, while also removing any webscraping artifacts and reformating the output text using markdown (added headings, lists, or other formatting elements to improve readability), ensuring the text is high quality and clean.
The LLM was also used to generate a title if the document did not have one.Data Statistics
- Total Documents: 13,287,694
- Total Tokens: 14.9B GPT-4o Tokens (14,895,034,936 Tokens)
- Total Size: ~71GB
- Total GPU Hours Needed: 15,000 H100 Hours per language
Data Fields
id: (str) Unique identifier for the document.title: (str) Title of the document.text: (str) The main content of the document, translated to Arabic.metadata: (dict) Additional metadata about the document, including:url: (str) The original URL of the document.dump: (str) The common crawl dump from which the document was extracted.date: (str) The date when the document was scraped.file_path: (str) The path to the original file in the common crawl dataset.language: (str) The language of the original document (always "English"en).language_score: (float) The language quality score of the document, ranging from 0 to 1.minhash_cluster_size: (int) The size of the deduplication cluster the document belongs to.fw_edu_int_score: (int) The rounded FineWeb-Edu classifier score for the document, indicating its educational quality (0-5).fw_edu_score: (float) The FineWeb-Edu classifier score for the document, indicating its educational quality (0-5).wo_format_label: (str) The format label assigned by the WebOrganizer classifier, indicating the type of content. Check the WebOrganizer Classifiers for more details.wo_format_score: (float) The confidence score for the format label assigned by the WebOrganizer classifier.wo_topic_label: (str) The topic label assigned by the WebOrganizer classifier, indicating the main subject of the content. Check the WebOrganizer Classifiers for more details.wo_topic_score: (float) The confidence score for the topic label assigned by the WebOrganizer classifier.wo_format_output: (list[dict]) The full output of the WebOrganizer classifier for the format label, including the label and score of all formats.wo_topic_output: (list[dict]) The full output of the WebOrganizer classifier for the topic label, including the label and score of all topics.length: (int) The length of the document in characters.token_count: (int) The number of tokens in the document, calculated using the GPT-4o tokenizer.orig_text: (str) The original text of the document before translation.orig_len: (int) The length of the original text in characters.orig_token_count: (int) The number of tokens in the original text, using the gpt2 tokenizer.
Data Formats
The dataset contains documents in the following formats, filtered from all formats available in the WebOrganizer classifier:
- Academic Writing
- Nonfiction Writing
- Personal Blog
- Q&A Forum
- Structured Data
- Creative Writing
- Documentation
- Tutorial
- Knowledge Article
Topics
The dataset contains documents on the following topics:
- Adult
- Art & Design
- Software Dev.
- Crime & Law
- Education & Jobs
- Hardware
- Entertainment
- Social Life
- Fashion & Beauty
- Finance & Business
- Food & Dining
- Games
- Health
- History
- Home & Hobbies
- Industrial
- Literature
- Politics
- Religion
- Science & Tech.
- Software
- Sports & Fitness
- Transportation
- Travel
Deduplication
The dataset has been fully deduplicated using the MinHash algorithm with the following parameters:
- Num Buckets: 16
- Hashes per Bucket: 8
- Ngrams: 13
Listing Stats
VIEWS
80
DELIVERY
INSTANT DOWNLOAD
LISTED
17/10/2025
UPDATED
21/10/2025
REGION
GLOBAL
TRUST
5 / 5
Loading...
£12,000
Download Dataset in TEXT Format
Recommended Datasets
Loading recommendations...
