Data Preparation & Documentation

AI Center › Data Preparation & Documentation

Global · The AI Lifecycle — Stage 3 of 14

Data Preparation & Documentation

Stage 3 of 14 in The AI Lifecycle: the point at which collected data is prepared for model building — annotated, labelled, cleaned, updated, enriched and aggregated — and at which the data sets and the AI system itself are documented, logged and retained under defined obligations.

This stage records the data-preparation and documentation provisions of Regulation (EU) 2024/1689 (Articles 10–12, 18–19 and 53, with Annexes IV, XI and XII), the datasheet and model-card documentation frameworks of Gebru et al. and Mitchell et al., and the ISO/IEC 5259 data-quality series. Every statement is mapped to a named authority and a linked source verified against the cited text.

Which data-preparation operations does the EU AI Act name?

Article 10(2)(c) of Regulation (EU) 2024/1689 names "relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation" among the matters that data governance and management practices for high-risk AI systems must address.

Article 10(1) requires high-risk AI systems that involve training AI models with data to be developed on the basis of training, validation and testing data sets that meet the quality criteria of Article 10(2) to (5) whenever such data sets are used. Article 10(2) subjects those data sets to data governance and management practices appropriate for the intended purpose of the system; the full enumeration of the eight matters those practices shall concern in particular, points (a) to (h), is displayed at the data sourcing and governance stage of this series. Point (c) names the operations belonging to this stage: annotation, labelling, cleaning, updating, enrichment and aggregation. Point (d) adjoins them, covering the formulation of assumptions, in particular with respect to the information that the data are supposed to measure and represent. The documentation of these operations is fixed by Annex IV, point 2(d), which requires the technical documentation of a high-risk AI system to describe, where relevant, labelling procedures (e.g. for supervised learning) and data cleaning methodologies (e.g. outliers detection). Under Article 10(6), for high-risk AI systems developed without techniques involving the training of AI models, paragraphs 2 to 5 apply only to the testing data sets.

Source: EU AI Act, Art. 10(2) — Regulation (EU) 2024/1689 ↗

What is a datasheet for a dataset?

A datasheet for a dataset is a document that accompanies a dataset and answers a structured set of questions about how and why it was created, what it contains, how it was collected and prepared, and how it is distributed and maintained — a documentation practice proposed by Gebru et al. by analogy with the datasheets that accompany electronic components.

The proposal was first circulated on arXiv in March 2018 and published in Communications of the ACM in December 2021. It organises the datasheet as question sets keyed to stages of the dataset lifecycle: Motivation (why the dataset was created); Composition (what it contains); Collection Process (how the data were gathered); Preprocessing/Cleaning/Labeling (what preparation was applied); Uses (intended and potential applications); Distribution (how the dataset is shared); and Maintenance (how it is updated). The stated aims are improved communication between dataset creators and dataset consumers and greater transparency and accountability in the machine-learning community, with particular reference to high-stakes application contexts. The same term appears in Union law: Annex IV point 2(d) of Regulation (EU) 2024/1689 requires technical documentation of a high-risk AI system to state, where relevant, "the data requirements in terms of datasheets describing the training methodologies and techniques and the training data sets used".

Source: Gebru et al., "Datasheets for Datasets" (2021) — arXiv:1803.09010 ↗

How does a model card differ from a datasheet?

A datasheet documents a dataset; a model card documents a trained model. Mitchell et al. describe model cards as short documents accompanying trained machine-learning models that report benchmarked evaluation across a variety of conditions, including different cultural, demographic or phenotypic groups and intersectional groups relevant to the intended application domains.

The model-card framework, presented at the FAT* '19 conference (Atlanta, January 2019), sets out nine sections: Model Details; Intended Use; Factors; Metrics; Evaluation Data; Training Data; Quantitative Analyses; Ethical Considerations; and Caveats and Recommendations. The datasheet framework of Gebru et al. (arXiv:1803.09010) addresses the dataset lifecycle from motivation through maintenance, whereas the model card addresses a trained model's characteristics, evaluation conditions and disaggregated performance. Both originated as voluntary research-community documentation practices rather than legal instruments; in Regulation (EU) 2024/1689, dataset-level documentation appears as the data requirements framed "in terms of datasheets" in Annex IV point 2(d) for high-risk AI systems, and model-level documentation appears as the Annex XI technical documentation required of general-purpose AI model providers under Article 53(1)(a).

Source: Mitchell et al., "Model Cards for Model Reporting" (2019) — arXiv:1810.03993 ↗

What must the technical documentation of a high-risk AI system contain, and when must it exist?

Under Article 11(1), the technical documentation of a high-risk AI system shall be drawn up before the system is placed on the market or put into service and kept up to date; it must demonstrate compliance with the requirements of Chapter III, Section 2, give national competent authorities and notified bodies the information needed to assess that compliance in a clear and comprehensive form, and contain at minimum the elements set out in Annex IV.

Article 11(2) provides that where a high-risk AI system relates to a product covered by the Union harmonisation legislation listed in Section A of Annex I, a single set of technical documentation is drawn up containing both the Article 11(1) information and the information required under those legal acts. Article 11(3) empowers the Commission to adopt delegated acts under Article 97 to amend Annex IV in light of technical progress. Within Annex IV, point 2(d) carries the dataset-documentation requirement: where relevant, the data requirements in terms of datasheets describing the training methodologies and techniques and the training data sets used, including a general description of the data sets, information about their provenance, scope and main characteristics, how the data was obtained and selected, labelling procedures (e.g. for supervised learning) and data cleaning methodologies (e.g. outliers detection).

Source: EU AI Act, Art. 11 — Regulation (EU) 2024/1689 ↗

What are the main heads of information in Annex IV technical documentation?

Annex IV lists nine heads of information that the technical documentation referred to in Article 11(1) shall contain at least, as applicable to the relevant AI system.

Annex IV point Required information (summarised)
1 A general description of the AI system: intended purpose, provider and version; interaction with hardware, software or other AI systems; software and firmware versions; the forms of placing on the market or putting into service; target hardware; product illustrations; user interface and instructions for use for the deployer
2 A detailed description of the system's elements and development process: methods and steps, including third-party pre-trained systems or tools; design specifications; system architecture and computational resources; data requirements in terms of datasheets (point 2(d)); assessment of Article 14 human oversight measures; pre-determined changes; validation and testing procedures, metrics and test logs; cybersecurity measures
3 Detailed information about monitoring, functioning and control: performance capabilities and limitations; degrees of accuracy for specific persons or groups; foreseeable unintended outcomes and sources of risks; human oversight measures; specifications on input data, as appropriate
4 The appropriateness of the performance metrics for the specific AI system
5 The risk management system in accordance with Article 9
6 Relevant changes made by the provider to the system through its lifecycle
7 A list of harmonised standards applied, as published in the Official Journal of the European Union; where none, the solutions adopted to meet the Chapter III, Section 2 requirements, and other standards and technical specifications applied
8 A copy of the EU declaration of conformity referred to in Article 47
9 The system for evaluating performance in the post-market phase in accordance with Article 72, including the post-market monitoring plan referred to in Article 72(3)

What does the simplified technical documentation regime for SMEs provide?

Article 11(1) provides that SMEs, including start-ups, may provide the elements of the technical documentation specified in Annex IV in a simplified manner, and that the Commission shall establish a simplified technical documentation form targeted at the needs of small and microenterprises.

Under the same paragraph, where an SME, including a start-up, opts to provide the information required in Annex IV in a simplified manner, it shall use the form referred to in Article 11(1), and notified bodies shall accept that form for the purposes of the conformity assessment. The provision addresses the manner in which the Annex IV information may be provided; the elements to be covered remain those specified in Annex IV. The regime sits within the same article that requires technical documentation to be drawn up before the high-risk AI system is placed on the market or put into service and to be kept up to date.

Source: EU AI Act, Art. 11(1) — Regulation (EU) 2024/1689 ↗

How long must documentation be kept, and by whom?

Under Article 18(1), the provider keeps, for a period ending 10 years after the high-risk AI system has been placed on the market or put into service, at the disposal of the national competent authorities: the Article 11 technical documentation; the documentation concerning the Article 17 quality management system; documentation of changes approved by notified bodies and the decisions and other documents issued by notified bodies, where applicable; and the EU declaration of conformity referred to in Article 47.

Article 18(2) provides that each Member State determines the conditions under which that documentation remains at the disposal of the national competent authorities for the 10-year period where a provider or its authorised representative established on its territory goes bankrupt or ceases activity before the period ends. Article 18(3) provides that providers that are financial institutions subject to Union financial services law requirements on internal governance, arrangements or processes maintain the technical documentation as part of the documentation kept under the relevant Union financial services law.

Source: EU AI Act, Art. 18(1) — Regulation (EU) 2024/1689 ↗

Who keeps automatically generated logs, and for how long?

Article 12(1) requires high-risk AI systems to technically allow for the automatic recording of events (logs) over the lifetime of the system; Article 19(1) requires providers to keep the logs automatically generated by their high-risk AI systems, to the extent the logs are under their control, for a period appropriate to the intended purpose and of at least six months, unless applicable Union or national law — in particular Union law on the protection of personal data — provides otherwise.

Article 12(2) requires logging capabilities that enable the recording of events relevant for: (a) identifying situations that may result in the system presenting a risk within the meaning of Article 79(1) or in a substantial modification; (b) facilitating post-market monitoring under Article 72; and (c) monitoring the operation of high-risk AI systems by deployers under Article 26(5). For the remote biometric identification systems referred to in Annex III point 1(a), Article 12(3) sets minimum logging content: the period of each use (start and end date and time); the reference database against which input data has been checked; the input data for which the search led to a match; and the identification of the natural persons involved in verifying the results, as referred to in Article 14(5). Article 19(2) provides that financial-institution providers maintain the automatically generated logs as part of the documentation kept under the relevant financial services law.

Source: EU AI Act, Art. 12(1) — Regulation (EU) 2024/1689 ↗

What documentation do providers of general-purpose AI models owe?

Article 53(1)(a) requires providers of general-purpose AI models to draw up and keep up to date technical documentation of the model — including its training and testing process and the results of its evaluation — containing at minimum the Annex XI information, for provision on request to the AI Office and the national competent authorities; Article 53(1)(b) requires them to draw up, keep up to date and make available to providers of AI systems that intend to integrate the model information and documentation containing at minimum the Annex XII elements.

Annex XI Section 1 covers, for all general-purpose AI model providers: a general description (intended tasks and types of AI systems for integration, acceptable use policies, release date and distribution methods, architecture and number of parameters, input/output modality and format, licence) and a detailed description of the development process — including, for the training, testing and validation data, the type and provenance of data and curation methodologies, the number of data points, their scope and main characteristics, how the data was obtained and selected, and measures to detect unsuitable data sources and identifiable biases — plus computational resources used for training (e.g. number of floating point operations), training time and known or estimated energy consumption. Section 2 adds, for models with systemic risk, evaluation strategies and results, adversarial-testing measures (e.g. red teaming) and system-architecture description. Article 53(2) exempts providers of models released under a free and open-source licence with publicly available parameters, weights, architecture and usage information from points (a) and (b), except for general-purpose AI models with systemic risks. Article 53(1)(d) separately requires a publicly available, sufficiently detailed summary of the content used for training, drawn up according to a template provided by the AI Office.

Source: EU AI Act, Art. 53(1) — Regulation (EU) 2024/1689 ↗

What does the ISO/IEC 5259 series cover?

ISO/IEC 5259, Artificial intelligence — Data quality for analytics and machine learning (ML), is a six-part series on measuring, managing, governing and visualizing the quality of data used in analytics and machine learning; the ISO catalogue lists all six parts as published — parts 1, 3 and 4 in July 2024, part 2 in November 2024, part 5 in February 2025 and the part 6 technical report in May 2026.

Per the ISO catalogue: ISO/IEC 5259-1:2024 provides the series overview with terminology and examples (iso.org/standard/81088.html). ISO/IEC 5259-2:2024 sets out data quality measures (iso.org/standard/81860.html). ISO/IEC 5259-3:2024 states data quality management requirements and guidelines (iso.org/standard/81092.html). ISO/IEC 5259-4:2024 describes a data quality process framework (iso.org/standard/81093.html). ISO/IEC 5259-5:2025 describes a data quality governance framework by which governing bodies direct and oversee data-quality measures, management and related processes across the data life cycle (iso.org/standard/84150.html). ISO/IEC TR 5259-6:2026, a technical report published in May 2026, sets out a framework for presenting the results of data-quality measurement visually; its catalogue abstract frames the aim as supporting stakeholder assessment of those results through visualisation methods (iso.org/standard/86532.html). Regulation (EU) 2024/1689 addresses the quality of training, validation and testing data for high-risk AI systems separately, in Article 10.

Source: ISO/IEC 5259-1:2024, Artificial intelligence — Data quality for analytics and machine learning (ML) — Part 1: Overview, terminology, and examples ↗

When do these data and documentation obligations apply?

Under Article 113 of Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744, the requirements of Chapter III, Sections 1 to 3 — which carry the Articles 10, 11 and 12 data-governance, technical-documentation and logging duties and the Articles 18 and 19 retention duties for high-risk AI systems — apply from 2 December 2027 for systems that are high-risk under Article 6(2) and Annex III, and from 2 August 2028 for systems that are high-risk under Article 6(1) and Annex I; the Chapter V obligations for general-purpose AI models (Article 53) have applied since 2 August 2025, and Chapters I and II since 2 February 2025.

Article 6(5), under which the Commission issues guidelines on high-risk classification, is expressly carved out of that deferral and applies from 2 August 2026, as do Chapter III, Section 5 (Articles 40 to 49 on harmonised standards, conformity assessment, certificates and registration, including the Article 47 EU declaration of conformity listed at Annex IV point 8), Article 50 on transparency, and Chapter IX (Articles 72 to 94, including the Article 72 post-market monitoring listed at Annex IV point 9). The amending instrument is Regulation (EU) 2026/1744, the Digital Omnibus on AI (PE-CONS 30/26; procedure 2025/0359(COD), titled on the European Parliament Legislative Observatory "Simplification of the implementation of harmonised rules on artificial intelligence – Digital Omnibus on AI (Omnibus VII)") was adopted by the European Parliament on 16 June 2026 and by the Council on 29 June 2026, signed on 8 July 2026 and published in the Official Journal on 24 July 2026 (OJ L 2026/1744); it has been in force since 27 July 2026. It defers the Article 6(2)/Annex III high-risk obligations to 2 December 2027 and the Article 6(1)/Annex I obligations to 2 August 2028 (procedure file: oeil.europarl.europa.eu/oeil/en/procedure-file?reference=2025/0359(COD)). Articles 102 to 110, which amend other Union acts, apply from 27 July 2026, and under Article 111(2) high-risk AI systems placed on the market or put into service for use by public authorities before 2 August 2026 are to be brought into compliance by 2 August 2030. As of 1 July 2026, no harmonised standards in support of the AI Act had been cited in the Official Journal of the European Union; Annex IV point 7 provides for listing harmonised standards applied or, where none are applied, describing the solutions adopted to meet the requirements.

Source: EU AI Act, Art. 113 — Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744 ↗

Terms defined at this stage

data-preparation processing operations
Operations applied to bring training, validation and testing data sets into usable form, named in Article 10(2)(c) by example as "annotation, labelling, cleaning, updating, enrichment and aggregation"; for high-risk AI systems these operations fall within the required data governance and management practices.
input data
"Data provided to or directly acquired by an AI system on the basis of which the system produces an output" (Article 3(33)); Annex IV point 3 requires technical documentation to include specifications on input data, as appropriate.
data governance and management practices
The practices, appropriate for the intended purpose of a high-risk AI system, to which its training, validation and testing data sets shall be subject under Article 10(2), concerning in particular design choices; data collection and origin; data-preparation processing operations; formulation of assumptions; assessment of availability, quantity and suitability; examination and mitigation of possible biases; and identification of data gaps or shortcomings.
datasheet (for a dataset)
A document accompanying a dataset that answers structured questions organised by dataset-lifecycle stage — Motivation, Composition, Collection Process, Preprocessing/Cleaning/Labeling, Uses, Distribution and Maintenance — proposed by Gebru et al. by analogy with electronics datasheets; the term also appears in Annex IV point 2(d) of Regulation (EU) 2024/1689, which frames data requirements "in terms of datasheets".
automatically generated logs
Records of events that a high-risk AI system must technically allow to be recorded automatically over its lifetime under Article 12(1), and that providers keep under Article 19(1), to the extent the logs are under their control, for a period appropriate to the system's intended purpose and of at least six months, subject to applicable Union or national law.
simplified technical documentation form
The form that the Commission shall establish, targeted at the needs of small and microenterprises, through which SMEs including start-ups may provide the elements of technical documentation specified in Annex IV in a simplified manner under Article 11(1); notified bodies shall accept the form for the purposes of the conformity assessment.

Cite this page

1BusinessWorld AI Center, "Data Preparation & Documentation — The AI Lifecycle." https://1businessworld.com/ai-center/data-preparation-and-documentation/ Version as of July 26, 2026.

The AI Center is informational only. It is provided by 1BusinessWorld strictly for general informational and educational purposes. Nothing in the AI Center constitutes, or should be construed as, legal, regulatory, compliance, technical, engineering, security, investment, financial, or other professional advice, or a recommendation, endorsement, solicitation, or offer regarding any technology, product, model, provider, framework, or course of action. 1BusinessWorld is not a law firm, regulatory authority, standards body, conformity-assessment or certification body, or investment adviser, and nothing in the AI Center creates any advisory, fiduciary, attorney-client, or other professional relationship with 1BusinessWorld. Although the AI Center references official materials published by legislatures, regulators, standards bodies, research organizations, and other named authorities, 1BusinessWorld makes no representation or warranty, express or implied, as to the accuracy, completeness, timeliness, or fitness for any purpose of any content, and, to the fullest extent permitted by law, disclaims all liability for any loss or damage of any kind arising directly or indirectly from the use of, or reliance on, any information presented. Laws, regulations, standards, technical practices, and AI capabilities change frequently and differ by jurisdiction; readers must verify all information against the current official text or source and consult qualified legal, compliance, technical, and other professional advisors before acting. Any decision relating to the development, deployment, procurement, or governance of AI systems is made solely at the reader's own risk. Last reviewed: July 26, 2026.