Data Sourcing, Governance & Provenance

AI Center › Data Sourcing, Governance & Provenance

Global · The AI Lifecycle — Stage 2 of 14

Data Sourcing, Governance & Provenance

Stage 2 of 14 in The AI Lifecycle records where training, validation and testing data comes from, the legal rights under which it is collected and used, and the governance and provenance requirements that attach to it before any model is trained.

Coverage spans the EU AI Act's Article 10 data-governance requirements, GDPR principles and lawful bases as examined in EDPB Opinion 28/2024, the text-and-data-mining regime of Directive (EU) 2019/790, the Article 53(1)(d) public summary of training content, and the ISO/IEC standards for data quality and the data life cycle. Every statement on this page is mapped to a named authority — EUR-Lex legal texts, the EDPB, the European Commission and ISO — each with a link verified live, and applicability dates are stated as they stand under Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026 and in force from 27 July 2026.

What does Article 10 of the EU AI Act require for the data behind high-risk AI systems?

Article 10(1) of Regulation (EU) 2024/1689 requires high-risk AI systems that involve training AI models to be developed on the basis of training, validation and testing data sets that meet the quality criteria in paragraphs 2 to 5, "whenever such data sets are used". Article 10(2) places those data sets under "data governance and management practices appropriate for the intended purpose of the high-risk AI system".

Article 10 sits in Chapter III, Section 2 of the Regulation, among the requirements a high-risk AI system must satisfy before it is placed on the market or put into service. The governance duty is purpose-relative: the practices required are those appropriate for the system's intended purpose, and Article 10(2) states that they "shall concern in particular" eight matters, enumerated in the table on this page. For high-risk AI systems developed without techniques involving the training of AI models, Article 10(6) applies paragraphs 2 to 5 only to the testing data sets. Under Article 10(3), the required characteristics may be met at the level of individual data sets or at the level of a combination of them. Article 113 of Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744, defers Chapter III, Sections 1, 2 and 3 — which contain Article 10 — so that they apply from 2 December 2027 for systems that are high-risk under Article 6(2) and Annex III, and from 2 August 2028 for systems that are high-risk under Article 6(1) and Annex I, as set out in the applicability entry on this page.

Source: EU AI Act, Art. 10 — Regulation (EU) 2024/1689 ↗

Which practices must Article 10(2) data governance cover for training, validation and testing data?

Article 10(2)(a)–(h) enumerates eight matters that the data governance and management practices "shall concern in particular", running from design choices and the origin of data through bias examination and the identification of data gaps.

Point Data governance practice — Article 10(2), Regulation (EU) 2024/1689
(a) The relevant design choices
(b) Data collection processes and the origin of data, and in the case of personal data, the original purpose of the data collection
(c) Relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation
(d) The formulation of assumptions, in particular with respect to the information that the data are supposed to measure and represent
(e) An assessment of the availability, quantity and suitability of the data sets that are needed
(f) Examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations
(g) Appropriate measures to detect, prevent and mitigate possible biases identified according to point (f)
(h) The identification of relevant data gaps or shortcomings that prevent compliance with this Regulation, and how those gaps and shortcomings can be addressed

What quality standard does Article 10(3) set for training, validation and testing data sets?

Article 10(3) of Regulation (EU) 2024/1689 requires that training, validation and testing data sets be "relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose". In the enacted sentence, the qualifier "to the best extent possible" immediately precedes "free of errors and complete".

The provision continues: the data sets "shall have the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used". Its final sentence states that "those characteristics of the data sets may be met at the level of individual data sets or at the level of a combination thereof", so the standard can be satisfied by data sets in combination as well as individually. Article 10(3) operates alongside Article 10(4), which addresses the geographical, contextual, behavioural and functional setting of intended use, and both are measured against the system's intended purpose — a term defined in Article 3(12) of the Regulation.

Source: EU AI Act, Art. 10(3) — Regulation (EU) 2024/1689 ↗

What does Article 10(4) require about the setting in which a high-risk AI system will operate?

Article 10(4) of Regulation (EU) 2024/1689 requires that data sets take into account, "to the extent required by the intended purpose", the characteristics or elements particular to "the specific geographical, contextual, behavioural or functional setting within which the high-risk AI system is intended to be used".

The obligation is bounded twice by intended purpose: the relevant setting is the one within which the system "is intended to be used", and the degree to which data sets must reflect it is "to the extent required by the intended purpose". Four dimensions of setting are named — geographical, contextual, behavioural and functional. Article 10(4) works together with the statistical-properties sentence of Article 10(3), which refers, where applicable, to "the persons or groups of persons in relation to whom the high-risk AI system is intended to be used". Under Article 10(6), for high-risk AI systems developed without techniques involving the training of AI models, this requirement — like the rest of paragraphs 2 to 5 — applies only to the testing data sets.

Source: EU AI Act, Art. 10(4) — Regulation (EU) 2024/1689 ↗

When may special categories of personal data be processed to detect and correct bias in high-risk AI systems?

Article 10(5) of Regulation (EU) 2024/1689 permits providers to process special categories of personal data "exceptionally" and only "to the extent that it is strictly necessary" for bias detection and correction under Article 10(2), points (f) and (g), subject to appropriate safeguards for fundamental rights and freedoms and to all of the conditions in points (a) to (f) being met.

The processing remains additionally subject to Regulations (EU) 2016/679 and (EU) 2018/1725 and Directive (EU) 2016/680. The six cumulative conditions are: (a) bias detection and correction cannot be effectively fulfilled by processing other data, including synthetic or anonymised data; (b) the special categories of personal data are subject to technical limitations on re-use and to state-of-the-art security and privacy-preserving measures, "including pseudonymisation"; (c) measures ensure the data are secured and protected, with strict controls and documentation of access, so that only authorised persons have access under appropriate confidentiality obligations; (d) the data are not transmitted, transferred or otherwise accessed by other parties; (e) the data are deleted once the bias has been corrected or the retention period ends, whichever comes first; and (f) the records of processing activities include the reasons why processing special categories of personal data was strictly necessary and why that objective could not be achieved by processing other data. The categories themselves are defined in Article 9(1) GDPR, set out in the terms below.

Source: EU AI Act, Art. 10(5) — Regulation (EU) 2024/1689 ↗

Which GDPR principles and lawful bases govern personal data used in AI training?

Article 5(1) GDPR requires personal data — including personal data collected and used as training data — to be processed in accordance with the principles of lawfulness, fairness and transparency, purpose limitation, data minimisation, accuracy, storage limitation, and integrity and confidentiality, with the controller responsible under Article 5(2) for demonstrating compliance. Article 6(1) makes processing lawful "only if and to the extent that" at least one of six bases applies.

The six Article 6(1) bases are: (a) the data subject's consent for one or more specific purposes; (b) necessity for a contract with the data subject; (c) compliance with a legal obligation of the controller; (d) protection of vital interests; (e) performance of a task in the public interest or exercise of official authority; and (f) necessity "for the purposes of the legitimate interests pursued by the controller or by a third party, except where such interests are overridden by the interests or fundamental rights and freedoms of the data subject which require protection of personal data". Purpose limitation, Article 5(1)(b), requires collection "for specified, explicit and legitimate purposes" with no incompatible further processing; data minimisation, Article 5(1)(c), requires data to be "adequate, relevant and limited to what is necessary" for those purposes. EDPB Opinion 28/2024 examines how supervisory authorities assess the Article 6(1)(f) basis for AI model development and deployment, set out in the next entry.

Source: GDPR, Arts. 5–6 — Regulation (EU) 2016/679 ↗

How does EDPB Opinion 28/2024 address legitimate interest as a basis for AI model training?

Opinion 28/2024, adopted by the European Data Protection Board on 17 December 2024 under Article 64(2) GDPR at the request of the Irish supervisory authority, recalls a three-step test for relying on legitimate interest: identifying the legitimate interest pursued, a necessity test, and a balancing test.

The Opinion answers four questions: when an AI model can be considered anonymous; how controllers can demonstrate legitimate interest as a legal basis in the development phase; the same in the deployment phase; and the consequences of unlawful processing during development. On the first step, it recalls that an interest may be regarded as legitimate where three cumulative criteria are met: it is lawful, it is clearly and precisely articulated, and it is real and present rather than speculative. On the second step, necessity entails considering whether the processing will allow the pursuit of the interest and whether there is no less intrusive way of pursuing it, with attention to the amount of personal data processed. The third step assesses whether the interest is overridden by the interests or fundamental rights and freedoms of data subjects. The Opinion also notes that there is no hierarchy between GDPR legal bases, provides a non-exhaustive list of examples of mitigating measures for the development phase — including in relation to web scraping — and for the deployment phase, and states that supervisory authorities assess measures case by case.

Source: EDPB, Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models (December 2024) ↗

What copyright regime governs text and data mining of works used as AI training data in the EU?

Directive (EU) 2019/790 provides two text-and-data-mining exceptions: Article 3, for reproductions and extractions by research organisations and cultural heritage institutions carrying out text and data mining for scientific research on works to which they have lawful access, and Article 4, a general exception for lawfully accessible works that applies only where rightholders have not expressly reserved their rights.

Under Article 3(2), copies made under the research exception are stored with an appropriate level of security and may be retained for scientific research, including verification of research results. Under Article 4(2), reproductions and extractions may be retained for as long as necessary for the text and data mining. Article 4(3) sets the reservation condition: the exception applies "on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online". Article 4(4) states that Article 4 does not affect the application of Article 3. The AI Act connects this regime to model training: Article 53(1)(c) of Regulation (EU) 2024/1689 requires providers of general-purpose AI models to put in place "a policy to comply with Union law on copyright and related rights, and in particular to identify and comply with, including through state-of-the-art technologies, a reservation of rights expressed pursuant to Article 4(3) of Directive (EU) 2019/790", and Recital 106 states that this applies regardless of the jurisdiction in which the copyright-relevant acts underpinning training take place.

Source: CDSM Directive, Arts. 3–4 — Directive (EU) 2019/790 ↗

What must providers of general-purpose AI models publish about their training content?

Article 53(1)(d) of Regulation (EU) 2024/1689 requires every provider of a general-purpose AI model to "draw up and make publicly available a sufficiently detailed summary about the content used for training of the general-purpose AI model, according to a template provided by the AI Office". The European Commission published the Explanatory Notice and Template for the Public Summary of Training Content on 24 July 2025.

The Explanatory Notice presents its annexed document as a common minimal baseline for the information to be made public in the summary. The Commission's accompanying Questions & Answers states that its use is mandatory and describes three sections: general information about the provider and model, including training content types and data characteristics; a list of data sources, covering publicly available datasets, private datasets, scraped data, user data and synthetic data; and relevant data-processing aspects important for exercising rights under Union law. For content scraped from the internet, disclosure covers the crawlers used, the collection period, a description of the content scraped, and the top 10% of all domains scraped — for SMEs, the top 5% or 1,000 domains, whichever is lower. The obligation has applied since 2 August 2025 under Article 113, a date left unchanged by Regulation (EU) 2026/1744; for models placed on the market before that date, summaries are to be made available no later than 2 August 2027, with information gaps stated and justified where the information is unavailable or retrieval would impose a disproportionate burden. Regulation (EU) 2026/1744, the Digital Omnibus on AI, in force from 27 July 2026, defers the separate Article 6 high-risk dates set out in the applicability entry. Recital 107 describes the summary as generally comprehensive in scope rather than technically detailed.

Source: European Commission, Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models (July 2025) ↗

How do international standards frame data processing across the AI life cycle?

ISO/IEC 5338:2023 sets out life cycle processes for AI systems built with machine learning and with heuristic approaches, reworking the system and software life cycle processes of ISO/IEC/IEEE 15288 and ISO/IEC/IEEE 12207 and adding AI-specific processes drawn from ISO/IEC 22989 and ISO/IEC 23053 (paraphrase of the scope, Clause 1).

Published in December 2023 as a first edition by ISO/IEC JTC 1/SC 42, the document — per its scope, freshly worded here — supplies processes through which an AI system can be defined, controlled, managed, executed and improved over its life cycle stages, whether an organisation or project is developing the system or acquiring it; where a component of an AI system is conventional software or a conventional system, the corresponding ISO/IEC/IEEE 12207 and 15288 processes remain available to implement it. Data-centred framing comes from the companion document ISO/IEC 8183:2023, whose scope (Clause 1), paraphrased, marks out the stages of data processing across an AI system's life — acquisition, creation, development, deployment, maintenance and decommissioning — together with the actions belonging to each stage, prescribing no particular service, platform or tool and addressing organisations of every type and size. Both documents are available through the ISO Online Browsing Platform; neither had been cited as a harmonised standard under the AI Act in the Official Journal of the European Union as of 1 July 2026.

Source: ISO/IEC 5338:2023, Information technology — Artificial intelligence — AI system life cycle processes ↗

Which standards address data quality and data provenance for AI training data?

The ISO/IEC 5259 series, developed by ISO/IEC JTC 1/SC 42 under the title Artificial intelligence — Data quality for analytics and machine learning (ML), addresses data quality across six published parts, while Article 10(2)(b) of the EU AI Act anchors provenance in law by requiring data governance practices to cover data collection processes and the origin of data.

The six parts, with publication status verified on iso.org: Part 1, Overview, terminology, and examples (2024); Part 2, Data quality measures (2024); Part 3, Data quality management requirements and guidelines (2024); Part 4, Data quality process framework (2024); Part 5, Data quality governance framework (published February 2025); and Technical Report Part 6, Visualization framework for data quality (published May 2026). ISO/IEC 8183:2023 frames the data life cycle within which these quality activities sit, from acquisition through decommissioning (see terms below). None of the ISO/IEC documents named here had been cited as harmonised standards under Regulation (EU) 2024/1689 in the Official Journal of the European Union as of 1 July 2026; zero harmonised standards for the AI Act had been so cited as of that date.

Source: ISO/IEC 5259-1:2024, Artificial intelligence — Data quality for analytics and machine learning (ML) — Part 1: Overview, terminology, and examples ↗

When do the AI Act's data-related obligations apply?

Under Article 113 of Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744, the Regulation applies generally from 2 August 2026, but Article 10's data governance requirements sit in Chapter III, Section 2, which applies from 2 December 2027 for high-risk systems under Article 6(2) and Annex III and from 2 August 2028 for those under Article 6(1) and Annex I, while Chapter V, containing the Article 53(1)(d) training-content summary obligation, has applied since 2 August 2025.

The Regulation was published in the Official Journal on 12 July 2024 and entered into force on the twentieth day following publication. Article 113 staggers application: Chapters I and II have applied since 2 February 2025, except the inserted Article 5(1), points (ba) and (bb), and Article 5(1a) and (1b), which apply from 2 December 2026; Chapter III Section 4, Chapter V, Chapter VII, Chapter XII other than Article 101, and Article 78 have applied since 2 August 2025; Articles 102 to 110 apply from 27 July 2026; and general application runs from 2 August 2026, which is also the date for Chapter III Section 5 on standards, conformity assessment and registration, for Article 6(5), for Article 50 and for Chapter IX. Article 10 sits in Chapter III, Section 2, which Regulation (EU) 2026/1744 defers to 2 December 2027 for high-risk systems under Article 6(2) and Annex III and to 2 August 2028 for those under Article 6(1) and Annex I. That amending act — the Digital Omnibus on AI, procedure 2025/0359(COD) — was published as Regulation (EU) 2026/1744, OJ L 2026/1744 of 24 July 2026, and is in force from 27 July 2026. It also inserts Article 111(4), under which providers of AI systems generating synthetic audio, image, video or text placed on the market before 2 August 2026 must comply with Article 50(2) by 2 December 2026. The EUR-Lex procedure record for 2025/0359(COD) shows the Council approval of 29 June 2026 and the signature of 8 July 2026.

Source: EU AI Act, Art. 113 — Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744 ↗

Terms defined at this stage

special categories of personal data
Personal data revealing racial or ethnic origin, political opinions, religious or philosophical beliefs, or trade union membership, together with genetic data, biometric data processed for the purpose of uniquely identifying a natural person, data concerning health, and data concerning a natural person's sex life or sexual orientation. Article 9(1) GDPR prohibits processing these categories unless a condition in Article 9(2) applies; Article 10(5) of the EU AI Act adds six cumulative conditions where such data are processed for bias detection and correction in high-risk AI systems.
legitimate interests (lawful basis)
The sixth lawful basis for processing personal data under the GDPR: processing "necessary for the purposes of the legitimate interests pursued by the controller or by a third party, except where such interests are overridden by the interests or fundamental rights and freedoms of the data subject which require protection of personal data". EDPB Opinion 28/2024 recalls a three-step test for its use in AI model development and deployment: identification of the legitimate interest, a necessity test, and a balancing test.
purpose limitation
The GDPR principle that personal data be "collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes". Further processing for archiving purposes in the public interest, scientific or historical research purposes or statistical purposes is, in accordance with Article 89(1), not considered incompatible with the initial purposes.
data minimisation
The GDPR principle that personal data be "adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed". It applies to every processing operation on personal data, including the assembly of AI training corpora.
rights reservation (text-and-data-mining opt-out)
The mechanism in Article 4(3) of Directive (EU) 2019/790 by which the general text-and-data-mining exception applies only "on condition that the use of works and other subject matter referred to in that paragraph has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online". Article 53(1)(c) of the EU AI Act requires providers of general-purpose AI models to put in place a copyright policy that identifies and complies with such reservations, including through state-of-the-art technologies.
summary of training content
The public disclosure required of every provider of a general-purpose AI model: "a sufficiently detailed summary about the content used for training of the general-purpose AI model, according to a template provided by the AI Office". The European Commission published the Explanatory Notice and Template for the Public Summary of Training Content on 24 July 2025; the obligation has applied since 2 August 2025 under Article 113 of Regulation (EU) 2024/1689.
data life cycle (AI systems)
As framed by ISO/IEC 8183:2023, the ordered passage of data through an AI system's life, marked out as stages with actions attached to each — running from acquisition and creation through development, deployment, maintenance and decommissioning (paraphrase of Clause 1, Scope). The document prescribes no specific service, platform or tool and is addressed to organisations of every type and size that use data in developing and using AI systems.

Cite this page

1BusinessWorld AI Center, "Data Sourcing, Governance & Provenance — The AI Lifecycle." https://1businessworld.com/ai-center/data-sourcing-and-governance/ Version as of July 26, 2026.

The AI Center is informational only. It is provided by 1BusinessWorld strictly for general informational and educational purposes. Nothing in the AI Center constitutes, or should be construed as, legal, regulatory, compliance, technical, engineering, security, investment, financial, or other professional advice, or a recommendation, endorsement, solicitation, or offer regarding any technology, product, model, provider, framework, or course of action. 1BusinessWorld is not a law firm, regulatory authority, standards body, conformity-assessment or certification body, or investment adviser, and nothing in the AI Center creates any advisory, fiduciary, attorney-client, or other professional relationship with 1BusinessWorld. Although the AI Center references official materials published by legislatures, regulators, standards bodies, research organizations, and other named authorities, 1BusinessWorld makes no representation or warranty, express or implied, as to the accuracy, completeness, timeliness, or fitness for any purpose of any content, and, to the fullest extent permitted by law, disclaims all liability for any loss or damage of any kind arising directly or indirectly from the use of, or reliance on, any information presented. Laws, regulations, standards, technical practices, and AI capabilities change frequently and differ by jurisdiction; readers must verify all information against the current official text or source and consult qualified legal, compliance, technical, and other professional advisors before acting. Any decision relating to the development, deployment, procurement, or governance of AI systems is made solely at the reader's own risk. Last reviewed: July 26, 2026.