art_10__para_2
2. Training, validation and testing data sets shall be subject to data governance and management practices appropriate for the intended purpose of the high-risk AI system. Those practices shall concern in particular: the relevant design choices; data collection processes and the origin of data, and in the case of personal data, the original purpose of the data collection; relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation; the formulation of assumptions, in particular with respect to the information that the data are supposed to measure and represent; an assessment of the availability, quantity and suitability of the data sets that are needed; examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations; appropriate measures to detect, prevent and mitigate possible biases identified according to point (f); the identification of relevant data gaps or shortcomings that prevent compliance with this Regulation, and how those gaps and shortcomings can be addressed.
← art_10__para_1 · All articles · (a) →
Source: EUR-Lex CELLAR · retrieved 2026-08-26 · Text as adopted (Official Journal); later amendments are not incorporated in this text.