AI Compliance

6 Questions AI Teams Should Ask Before Buying Training Data Under the 2026 EU AI Act
Essential questions AI teams should consider when acquiring training data under the 2026 EU AI Act regulations.
In July 2026, a US court granted final approval to a record-breaking settlement. Anthropic agreed to pay a group of authors and publishers $1.5 billion to settle a copyright lawsuit. The lawsuit dates back to August 2024, when writers Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson sued Anthropic for downloading e-books from shadow libraries like LibGen to train Claude AI models. The judge in this case ruled that training AI models on legally acquired books counts as transformative fair use, but downloading pirated copies from shadow libraries does not, regardless of how the books were later used. Once the case was certified as a class action covering roughly 500,000 works, Anthropic faced theoretical damages that legal analysts calculated at around $72 billion if the case went to trial. To avoid that risk, Anthropic settled for $1.5 billion, the largest copyright recovery in US history.
This case shows just how tricky a situation AI developers can face if they ignore the importance of using legally obtained training data. And in 2026, the rules AI developers are racing to keep up with just changed.
What Actually Changed in 2026
On June 29, 2026, the Council of the EU gave final approval to the Digital Omnibus on AI, which was published in the Official Journal on July 24, 2026 and entered into force days later. The headline change is a 16-month deferral for standalone high-risk AI systems under Annex III, covering areas like hiring, credit scoring, biometric identification, and education, pushing the deadline from August 2, 2026 to December 2, 2027.
But this delay doesn't mean AI developers can relax across the board. Article 50's transparency obligations, which disclose AI-generated content and labelling chatbot interactions, are still due August 2, 2026, unchanged. Article 53 obligations for general-purpose AI (GPAI) providers have already been in force since August 2025. And a new set of prohibited practices, covering AI-generated non-consensual intimate imagery and similar abuse, kicks in from December 2, 2026. To clarify, although the pressure eased specifically on the data governance requirements under Article 10 below, almost everything else on the compliance calendar is still moving forward as originally planned.
Changes That Matter to AI Model and LLM Developers
The EU AI Act sets out clear rules for different scenarios, covering both individuals and businesses building AI models. Taking a quick glance at the EU AI Act, these are the articles that strongly impact developers when creating an AI model or application.
Article 10: Data and Data Governance
High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria.
The training, validation and testing data sets shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose. They shall have the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used.
This regulation doesn't explicitly ban scraped data, but it makes it very hard to use in practice because scraped data rarely comes with the origin records, consent history, or representativeness documentation Article 10 requires, so it's difficult to prove it meets the quality criteria above.
Article 53: Obligations for Providers of General-Purpose AI Models
AI model providers need to draw up and keep up-to-date the technical documentation of the model, including its training and testing process and the results of its evaluation, for the purpose of providing it, upon request, to the AI Office and the national competent authorities.
Article 53 (together with its accompanying annex) clearly lists what providers need to record in their documentation. Under this regulation, AI developers are required to keep detailed records of how they obtain training data and put in place a policy for complying with copyright law. In other words, they can't knowingly build on illegally sourced datasets.
Article 26: Obligations of Deployers of High-Risk AI Systems
AI model deployers shall monitor the operation of the high-risk AI system on the basis of the instructions for use and, where relevant, inform providers. Deployers must avoid using high-risk AI systems that may result in presenting risks. If they find out that AI systems contain data or usage that is against the law, they shall, without undue delay, inform the provider or distributor and the relevant market surveillance authority, and shall suspend the use of that system.
In this case, using illegal and high-risk data not only causes issues for developers, but also impacts businesses that use the model.
According to Article 99, the most serious violations of the EU AI Act (such as breaching Article 5's prohibited practices) can result in fines of up to €35,000,000 or, if the offender is an undertaking, up to 7% of its total worldwide annual turnover, whichever is higher.
Practical Steps for Companies to Face the Change
The question that AI developers might ask now is: how can we keep our AI models and AI projects from being fined? Here, Opendatabay offers 6 steps that AI developers should take to decrease the risk of violating the law.
Know your risk classification first
The EU AI Act sorts AI systems into risk tiers: unacceptable, high, limited, and minimal. High-risk systems (such as those in healthcare, hiring, and law enforcement) have strict data governance requirements. Buyers need to classify their system before they even start shopping for data. Without understanding which regulations apply to them, developers risk breaking the law unintentionally.
Ask for provenance documentation
Article 10 requires training data to have clear lineage, so data buyers should be asking suppliers: where did this data come from, how was it collected, do you have consent records, and is there a chain of custody? If a supplier can't answer these questions, that's a red flag. Remember, even though data buyers aren't the ones creating illegal datasets, they can still be fined for using them.
Check for bias and representativeness
The EU AI Act also requires high-risk AI training data to be relevant, representative, and as free of errors as possible. Thus, it is essential for buyers to ask data suppliers if data, for example, demographics, regions, and languages in the dataset are related to their AI model. A dataset that only covers one population will fail compliance for a product deployed across Europe since results could be biased.
Demand IP indemnification in writing
If a supplier provides data that turns out to be scraped, illegally acquired, or unlicensed, the buyer could face serious consequences. Make sure there are indemnification clauses in the contract that put liability on the supplier for rights and consent.
PII and GDPR are separate but connected
The EU AI Act does not replace GDPR. Buyers still need to confirm data is either PII-free, properly de-identified, or collected with valid consent. Ask suppliers for their de-identification methodology, for example, pseudonymisation, anonymisation techniques like k-anonymity or differential privacy, or synthetic data generation.
Keep an audit trail
The EU AI Act expects documentation of what data was used, when, and why. Buyers should be logging every dataset purchase, its provenance, its licence terms, and its intended use from day one. Deleting training data and covering your traces is what gets you fined or even banned from operating in the EU.
Opendatabay: Buy AI Training Data Without Concerns
Under the 2026 EU AI Act, finding a trustworthy and professional data seller to access high-quality AI training data can be exhausting. Give yourself more time to focus on your AI model's infrastructure by buying data from trusted AI data marketplaces like Opendatabay. We help clients save time by connecting AI teams with licensed, verified data sellers. While browsing licensed data products on our platform, you can make decisions without worrying about the legality of your training data.
To help you avoid the awkward situation of buying datasets from sellers who can't answer basic questions about where their data came from, Opendatabay ensures every supplier is able to answer critical questions clearly about their data products before they are even onboarded on the platform. Under Opendatabay's requirements, all data suppliers stand behind their data products listed on the platform and confirm they have full rights to list, sell, or licence them.
Cleaning and censoring datasets that contain personal or sensitive information can also be overwhelming. Opendatabay helps data suppliers with personal information scrubbing and anonymisation, partnering with Maya Data Privacy to double-check all data products listed on the platform and ensure they do not contain any personally sensitive data.
Developing AI models is hard enough without the headache of sourcing data. Opendatabay is built to make that part easier, giving you a fast, straightforward experience when searching for AI training data.
You can browse our data products or request a free sample to see if it's the right fit.