An EU AI Act timeline for data teams, ordered by what cannot be recovered
Data work under the AI Act should be sequenced by what becomes impossible to recover later, not by which deadline is nearest.
Questions about the EU AI Act usually arrive as a schedule: what is due when. For a data team, the more useful sort is by irreversibility. Some obligations can be satisfied late with effort, because they depend on documents that can be written afterwards. Others depend on records that either exist or do not. A consent form used last month, a guideline version retired in spring, a recording log that was never kept — none of these can be reconstructed once the moment passes.
Sequence the work by that distinction and most of the schedule sorts itself. The application dates have also been the subject of proposed amendments, so check the current official schedule before committing any specific date to a plan. The ordering below is stable even when the dates move.
First: whatever you are collecting right now
Every active collection campaign is the most urgent item on the list, regardless of any deadline. The consent language being used this week, the guideline version in force this month, the metadata fields captured on today's sessions — these are the only items where "later" means "never."
The concrete work: audit the consent form and the recording log of every running campaign. If the form does not describe AI training in terms a speaker would understand, changing it now costs a reprint and a briefing update. Changing it after the campaign means the batch is unusable, or needs re-consent from people who have scattered. This is one week of work with an outsized payoff.
Second: the inventory, because everything else depends on it
Classification decisions need a model inventory with the datasets attached, and the data team is the only group that can produce the attachment. The output is a table: model, version, dataset, role in the pipeline (training, validation or test), where the data subjects are, and whether personal data is involved.
Until that table exists, no one — legal, procurement or engineering — can say which obligations attach to what. It is an afternoon of assembly that unblocks three departments at once, and it is also the map that decides which datasets get governance files first.
Third: the paperwork with the longest lead time
Contracts move at the speed of the counterparty, which is why they start early:
- Data processing agreements with every supplier that touches personal data, including annotation subcontractors.
- License amendments for any corpus whose current grant does not name machine learning training.
- Transfer mechanisms where data crosses jurisdictions, agreed with the other side rather than asserted.
- Consent forms and briefing scripts for campaigns that have not started recruiting yet.
Fourth: the documentation that can be written late
Governance files, data summaries and transparency disclosures consume time but not access — as long as the underlying records exist. That is why they sit lower in the order, and why the sequence inside this category matters: gather records first, write second. Writing before gathering produces documents that contradict the logs, which is the worst outcome for a document meant to serve as evidence.
Practically: collect the logs, archived guidelines and scanned consent forms into one place per dataset, then draft from what is actually there. Where a record is missing, the document says so rather than smoothing it over.
What can wait
Tooling. Lineage platforms, governance dashboards, automated metadata pipelines — all genuinely useful, none on the critical path. Buying a platform before the inventory exists automates a process that has not been defined. The same goes for template libraries for obligations that attach to systems you do not build: a team with two models needs two good files, not a framework.
There is one trap worth naming explicitly. The obligations that applied earliest were not the heaviest ones, and the data governance duties — the ones that depend on collection-time records — arrive later. A team that spends its preparation budget on the early, narrower obligations while collection continues without updated consent has optimized for the wrong end of the schedule.