AI Compliance

6 Questions AI Teams Should Ask Before Buying Training Data Under the 2026 EU AI Act

6 Questions AI Teams Should Ask Before Buying Training Data Under the 2026 EU AI Act

Essential questions AI teams should consider when acquiring training data under the 2026 EU AI Act regulations.

In July 2026, a US court granted final approval to a record-breaking settlement. Anthropic agreed to pay a group of authors and publishers $1.5 billion to settle a copyright lawsuit. The lawsuit dates back to August 2024, when writers Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson sued Anthropic for downloading e-books from shadow libraries like LibGen to train Claude AI models. The judge in this case ruled that training AI models on legally acquired books counts as transformative fair use, but downloading pirated copies from shadow libraries does not, regardless of how the books were later used. Once the case was certified as a class action covering roughly 500,000 works, Anthropic faced theoretical damages that legal analysts calculated at around $72 billion if the case went to trial. To avoid that risk, Anthropic settled for $1.5 billion, the largest copyright recovery in US history.

This case shows just how tricky a situation AI developers can face if they ignore the importance of using legally obtained training data. And in 2026, the rules AI developers are racing to keep up with just changed.

What Actually Changed in 2026

On June 29, 2026, the Council of the EU gave final approval to the Digital Omnibus on AI, which was published in the Official Journal on July 24, 2026 and entered into force days later. The headline change is a 16-month deferral for standalone high-risk AI systems under Annex III, covering areas like hiring, credit scoring, biometric identification, and education, pushing the deadline from August 2, 2026 to December 2, 2027.

But this delay doesn't mean AI developers can relax across the board. Article 50's transparency obligations, which disclose AI-generated content and labelling chatbot interactions, are still due August 2, 2026, unchanged. Article 53 obligations for general-purpose AI (GPAI) providers have already been in force since August 2025. And a new set of prohibited practices, covering AI-generated non-consensual intimate imagery and similar abuse, kicks in from December 2, 2026. To clarify, although the pressure eased specifically on the data governance requirements under Article 10 below, almost everything else on the compliance calendar is still moving forward as originally planned.

Changes That Matter to AI Model and LLM Developers

The EU AI Act sets out clear rules for different scenarios, covering both individuals and businesses building AI models. Taking a quick glance at the EU AI Act, these are the articles that strongly impact developers when creating an AI model or application.

Article 10: Data and Data Governance

High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria.

The training, validation and testing data sets shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose. They shall have the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used.

This regulation doesn't explicitly ban scraped data, but it makes it very hard to use in practice because scraped data rarely comes with the origin records, consent history, or representativeness documentation Article 10 requires, so it's difficult to prove it meets the quality criteria above.

Article 53: Obligations for Providers of General-Purpose AI Models

AI model providers need to draw up and keep up-to-date the technical documentation of the model, including its training and testing process and the results of its evaluation, for the purpose of providing it, upon request, to the AI Office and the national competent authorities.

Article 53 (together with its accompanying annex) clearly lists what providers need to record in their documentation. Under this regulation, AI developers are required to keep detailed records of how they obtain training data and put in place a policy for complying with copyright law. In other words, they can't knowingly build on illegally sourced datasets.

Article 26: Obligations of Deployers of High-Risk AI Systems

AI model deployers shall monitor the operation of the high-risk AI system on the basis of the instructions for use and, where relevant, inform providers. Deployers must avoid using high-risk AI systems that may result in presenting risks. If they find out that AI systems contain data or usage that is against the law, they shall, without undue delay, inform the provider or distributor and the relevant market surveillance authority, and shall suspend the use of that system.

In this case, using illegal and high-risk data not only causes issues for developers, but also impacts businesses that use the model.

According to Article 99, the most serious violations of the EU AI Act (such as breaching Article 5's prohibited practices) can result in fines of up to €35,000,000 or, if the offender is an undertaking, up to 7% of its total worldwide annual turnover, whichever is higher.

Practical Steps for Companies to Face the Change

The question that AI developers might ask now is: how can we keep our AI models and AI projects from being fined? Here, Opendatabay offers 6 steps that AI developers should take to decrease the risk of violating the law.

Know your risk classification first

The EU AI Act sorts AI systems into risk tiers: unacceptable, high, limited, and minimal. High-risk systems (such as those in healthcare, hiring, and law enforcement) have strict data governance requirements. Buyers need to classify their system before they even start shopping for data. Without understanding which regulations apply to them, developers risk breaking the law unintentionally.

Ask for provenance documentation

Article 10 requires training data to have clear lineage, so data buyers should be asking suppliers: where did this data come from, how was it collected, do you have consent records, and is there a chain of custody? If a supplier can't answer these questions, that's a red flag. Remember, even though data buyers aren't the ones creating illegal datasets, they can still be fined for using them.

Check for bias and representativeness

The EU AI Act also requires high-risk AI training data to be relevant, representative, and as free of errors as possible. Thus, it is essential for buyers to ask data suppliers if data, for example, demographics, regions, and languages in the dataset are related to their AI model. A dataset that only covers one population will fail compliance for a product deployed across Europe since results could be biased.

Demand IP indemnification in writing

If a supplier provides data that turns out to be scraped, illegally acquired, or unlicensed, the buyer could face serious consequences. Make sure there are indemnification clauses in the contract that put liability on the supplier for rights and consent.

PII and GDPR are separate but connected

The EU AI Act does not replace GDPR. Buyers still need to confirm data is either PII-free, properly de-identified, or collected with valid consent. Ask suppliers for their de-identification methodology, for example, pseudonymisation, anonymisation techniques like k-anonymity or differential privacy, or synthetic data generation.

Keep an audit trail

The EU AI Act expects documentation of what data was used, when, and why. Buyers should be logging every dataset purchase, its provenance, its licence terms, and its intended use from day one. Deleting training data and covering your traces is what gets you fined or even banned from operating in the EU.

Opendatabay: Buy AI Training Data Without Concerns

Under the 2026 EU AI Act, finding a trustworthy and professional data seller to access high-quality AI training data can be exhausting. Give yourself more time to focus on your AI model's infrastructure by buying data from trusted AI data marketplaces like Opendatabay. We help clients save time by connecting AI teams with licensed, verified data sellers. While browsing licensed data products on our platform, you can make decisions without worrying about the legality of your training data.

To help you avoid the awkward situation of buying datasets from sellers who can't answer basic questions about where their data came from, Opendatabay ensures every supplier is able to answer critical questions clearly about their data products before they are even onboarded on the platform. Under Opendatabay's requirements, all data suppliers stand behind their data products listed on the platform and confirm they have full rights to list, sell, or licence them.

Cleaning and censoring datasets that contain personal or sensitive information can also be overwhelming. Opendatabay helps data suppliers with personal information scrubbing and anonymisation, partnering with Maya Data Privacy to double-check all data products listed on the platform and ensure they do not contain any personally sensitive data.

Developing AI models is hard enough without the headache of sourcing data. Opendatabay is built to make that part easier, giving you a fast, straightforward experience when searching for AI training data.

You can browse our data products or request a free sample to see if it's the right fit.

Frequently Asked Questions

What changed in the EU AI Act in 2026?
In June 2026, the EU approved the Digital Omnibus on AI with key changes: high-risk AI systems got a 16-month deadline extension (to December 2027), but transparency obligations and GPAI provider documentation requirements remain on their original timeline. New prohibited practices around AI-generated intimate imagery also take effect December 2, 2026.
What happens if I buy training data that turns out to be illegally sourced?
Even if you didn't create the illegal dataset, you can still face serious penalties under the EU AI Act. Article 99 allows fines up to €35 million or 7% of global annual turnover. This is why demanding provenance documentation and IP indemnification clauses in contracts with data suppliers is critical.
What training data documentation does Article 10 require?
Article 10 requires high-risk AI systems to use training data that is relevant, sufficiently representative, free of errors, and complete for the intended purpose. Data must have clear lineage with origin records, consent history, and representativeness documentation. Scraped data rarely meets these requirements because it lacks this crucial metadata.