AI Compliance

Selling AI Training Data? Read This 10-Point EU AI Act Compliance Checklist First

Selling AI Training Data? Read This 10-Point EU AI Act Compliance Checklist First

A 10-point EU AI Act compliance checklist for data sellers, covering provenance documentation, IP rights, bias audits, PII handling, and liability management when selling AI training data.

Clearview AI, an American facial recognition company, has been fined more than €100 million combined by European regulators, including a €30.5 million fine from the Dutch Data Protection Authority, because of how it collected its AI training data. The facial recognition AI model the firm developed works like a search engine, helping their clients find a certain person by facial images their clients uploaded. After the engine finds the same image, it will give their client the person's social media link.

Clearview AI helped American police identify suspects; however, since the data fed into the AI model was scraped from social media platforms without users' consent, it also faced a $52 million class-action settlement in the US, and a separate £7.5 million fine from the UK's ICO.

With the rapid growth of the AI and LLMs industry, more and more businesses and individual developers have been buying data from AI training data providers. Nonetheless, as business scale extends under this trend, data providers need to be more cautious while listing and trading their data products under the law, especially the strict EU AI Act.

The 2026 EU AI Act Explained: Why Does it Matter to Data Sellers

In the article 6 questions AI teams should ask before buying training data under the 2026 EU AI Act we explained which provisions matter to AI training data buyers and why. That raises another question: if the EU AI Act primarily regulates AI providers and deployers, should data sellers care too?

The answer is yes. While the AI Act places most legal obligations on AI providers, those obligations inevitably influence what buyers expect from their data suppliers. To comply with the regulation, buyers demand datasets with clear provenance, appropriate licensing, and documented quality. Sellers who can provide that information will be far better positioned to win business.

10 Things Data Sellers Should Do Before Listing AI Training Data

Whether you are a creative individual or a data-collection business, the first thing you need to worry about isn't what kind of data is most popular or if AI companies would like to buy your data. Developing AI models needs various types of data, so any type of data will be useful to AI developers.

The first thing to consider is whether your dataset is legal for buyers to use. Under the EU AI Act, the training data AI companies use must meet certain quality criteria laid out in Article 10. That means it should be clean, relevant, and its provenance well-documented.

On top of that, AI companies providing general-purpose AI models also need to draw up and maintain up-to-date technical documentation under Article 53, covering the model's training and testing process and the results of its evaluation.

To keep this process transparent, AI companies are expected to publish a summary of the datasets they use. To build a trustworthy reputation for your data products, here are 10 things data sellers should do to ensure their products' quality and safety.

1. Document your data lineage before anyone asks

Sellers need to be able to show exactly where the data came from, how it was collected, who consented, and what rights were granted. If you cannot trace every record back to its source, you might cause legal issues for your customers. Under the EU AI Act, AI providers are required to document their training data sources, so buyers will expect this information upfront.

Opendatabay suggests sellers start with a smaller data product, the one where they're 100% sure of the provenance. It is always better to be safe than to cause problems.

2. Get your licensing house in order

Sellers need clear, written agreements proving they have the right to licence data for AI training. This includes contracts with original data owners, consent forms from participants, and model releases. "We scraped it" or "it was publicly available" won't pass marketplaces like Opendatabay's listing policy because those are vague, high-risk answers about how the data was sourced. The EU AI Act's emphasis on lawful data processing makes these answers even harder to defend.

3. Remove or de-identify personal information properly

GDPR and the EU AI Act both expect sellers to remove PII using recognised methods, such as pseudonymisation, anonymisation techniques like k-anonymity or differential privacy, or synthetic data generation, before listing. Simply deleting names is not enough under those regulations. Metadata, file paths, and embedded identifiers all need scrubbing.

Data sellers should also pay extra attention to sample files. These are the files that are publicly visible on product listings and shared with buyers, so make sure all personal information is fully removed before samples go live or are shared.

4. Be transparent about what your data does and does not cover

Under the EU AI Act, AI developers can only use relevant datasets in their AI models. This means data buyers will likely ask about demographics, regions, languages, and domain coverage. Sellers who can clearly describe their data's representation and limitations are the ones who win deals. Hiding gaps only creates liability when the buyer's model fails on an underrepresented population.

Selling biased data is fine, as long as buyers are aware of it. In some cases, messy, filtered, unfiltered, or biased data is exactly what a buyer is looking for. Describing your products as honestly as possible should always be the priority. Article 10 specifically requires datasets to be "sufficiently representative," so being upfront about limitations actually helps your buyers stay compliant.

5. Prepare for IP indemnification requests

Enterprise buyers increasingly require sellers to warrant that the data is rightfully owned and not infringing. Sellers should be ready to sign indemnification clauses standing behind their data. If you can't indemnify, you'll most likely lose the deal. With the EU AI Act holding AI providers accountable for the data they use, buyers have even more reason to demand these protections from their suppliers.

For example, if a data buyer on Opendatabay gets sued by the original data owner (or by an entity in the data whose consent was never given), that liability gets passed back to Opendatabay, and Opendatabay passes it back to the data seller. Always be prepared to stand behind your data.

(Check Opendatabay's Data Provider Agreement for further information.)

6. Consider offering samples and pilot-scale pricing

Providing samples allows buyers to verify data quality and suitability before committing, which supports their compliance obligations under Article 10. Realistic pricing at pilot scale also allows buyers to run compliance checks on a small batch before scaling up. Given the EU AI Act's requirements around data quality assessment, this approach makes it easier for buyers to justify the purchase internally.

7. Clean and standardise your data products before listing

Article 10 requires training data to be "relevant, sufficiently representative, and to the best extent possible, free of errors." Delivering your data in standard, clean formats reduces the risk of errors introduced during format conversion or cleaning, and it supports your buyer's compliance too.

8. Build a documented trust record

Buyers will increasingly require evidence that their data suppliers are trustworthy and auditable. The EU AI Act requires AI providers to maintain detailed technical documentation covering their training data, which means they need suppliers who can back up their claims with evidence. A documented trust record and third-party scoring give buyers the supplier-level assurance the regulation demands.

9. Understand which AI risk category your data serves

If your data is used in high-risk AI systems (medical devices, recruitment tools, law enforcement), buyers face stricter compliance obligations under the EU AI Act and will pass those requirements down to you. Know where your data fits in the risk framework so you can answer buyer questions confidently and price accordingly.

10. Keep your listings updated

Stale listings with outdated volumes, old samples, or missing fields signal neglect. Update your inventory regularly, add new samples, and respond to buyer enquiries promptly. The EU AI Act expects ongoing data governance, not just a one-time compliance check. The sellers who close deals are the ones who treat their marketplace presence like a shopfront, not a filing cabinet.

Opendatabay: An Experienced Partner to Support Your Business

Opendatabay was founded by Justinas Kairys, a seasoned software engineer turned founder and CEO who has built software for the UK government, the United Nations, and enterprises across Asia. After years of developing AI systems, he saw first-hand that high-quality training data is one of the most important ingredients for successful AI.

Being a developer himself, he spent countless hours searching for trustworthy data providers, only to find that closing a deal took even longer than finding the right data. He figured there should be a marketplace where finding, licensing, and buying datasets is as simple as shopping online. That idea became the foundation of Opendatabay.

To protect data sellers from liability and potential legal claims, Opendatabay requires every seller to demonstrate they have the legal right to licence the data they list. Our vetting process, built on Palantir Foundry, helps identify licensing and compliance issues before your data goes live.

But Opendatabay is more than a marketplace. We work alongside sellers to navigate the complexities of data licensing and commercialisation (whether you're listing your first dataset, turning underused company data into a new revenue stream, or bringing a valuable data product to market).

If that sounds relevant to you, we'd be happy to talk.

Frequently Asked Questions

What fines do I face for selling non-compliant AI training data under the EU AI Act?
Under Article 99 of the EU AI Act, you can face fines up to €35 million or 7% of annual global turnover for providing false or misleading information about data provenance, quality, or rights. Additionally, your buyers face liability when they deploy models trained on your non-compliant data and will pursue indemnification claims against you, which can be far more costly than regulatory fines.
What documentation do I need to sell AI training data to EU buyers?
Prepare: data lineage and provenance records showing where data originated and how it was collected, written agreements proving you own or have rights to the data, consent forms or GDPR legal basis documentation if data involves people, bias audit reports showing representativeness, de-identification methodology if PII was present, quality assurance processes, data samples, and clear licensing terms specifying commercial vs. general AI training use.
Do I need to remove personally identifiable information (PII) before selling training data?
Yes. You must either remove PII using recognised de-identification techniques (pseudonymisation, anonymisation like k-anonymity, or synthetic data generation), have valid GDPR consent, or have a documented legal basis. Simply deleting names is not sufficient—metadata, file paths, embedded identifiers, and information in data samples must also be scrubbed.