1. The implementation phase first changes the boundaries of responsibility
The AI Act entered into force on 1 August 2024, but its obligations apply in stages according to risk and market role. Prohibited AI practices and AI literacy requirements began to apply on 2 February 2025. The main obligations for general-purpose AI models followed on 2 August 2025. For most system obligations and transparency requirements, 2 August 2026 is the critical date. Some high-risk systems connected to regulated product safety rules have a longer transition period, with certain requirements extending to 2 August 2027. Companies therefore need to confirm the applicable date against the type of system and the relevant transitional provision.[1]
These dates matter for more than the moment when regulators can begin enforcement. They connect decisions that used to be made separately across product, procurement, human resources, and management. Product managers need time in release plans for reclassification and regression testing. Procurement teams need current technical documentation before renewing a supplier. Human resources teams need a fundamental rights review before introducing recruitment or workforce analytics tools. Boards also need visibility into high-risk gaps that remain unresolved.
If the dates live only in a legal calendar while the rest of the company continues to release systems in the old way, the regulation has not entered the operating model. When each deadline is tied to an owner, a deliverable, and a location for evidence, the company can reconstruct the basis for a decision when a customer asks questions, a regulator carries out an inspection, or an incident occurs. The European Commission will continue to publish implementation material, and Member States will continue to develop national supervision and penalties. Regulatory versions, Commission guidance, national notices, and internal conclusions therefore belong in the same change record.[2]
The territorial scope also means that this is not an issue only for companies headquartered in the EU. A company outside the Union may fall within scope when it places a system on the EU market or when the system's output is used in the EU. Servers in the United States or Asia and a model hosted by a third party do not by themselves remove the link to EU regulation. Cross-border businesses need to look at where customers are located, where outputs are used, and who is affected, rather than focusing only on the location of the development team or data centre.
Once a company is within scope, its obligations depend on its role in the supply chain. A party that develops a system and places it on the market under its own name is a provider. A party that uses the system in its operations is a deployer. An importer brings a system from a third country into the EU market, while a distributor makes it available on that market. A provider outside the EU may also need an authorised representative. These roles carry different documentation, oversight, and market obligations, so role assessment is not a formal exercise in legal terminology. It is the starting point for every control that follows.[1]
One company may hold several roles at the same time. It may begin as the deployer of a general-purpose model, add its own data, workflow, and decision rules, and then sell the resulting service under its own brand. If it changes the intended purpose of the system or makes a substantial modification that affects compliance, some provider obligations that previously sat with the supplier may move to the company.
This is why a contract saying that the supplier is "responsible for compliance" cannot replace the company's own assessment. The supplier may be responsible for model-level documentation, version information, and vulnerability fixes, but the company still controls the business purpose, input data, user notices, human review, and final decision. Responsibility does not disappear because a contract contains a broad assurance, and it does not become clear merely because the underlying model comes from a well-known vendor.
A practical response is to create a role card for every system. It should record who defines the intended purpose, who selects the model version, who controls the input data, who approves release, who can suspend the service, and who ultimately uses the output. The card should be tied to the system version and contracting entity, then reviewed whenever the purpose, brand, model, or market scope changes. This turns an abstract legal role into operational ownership instead of forcing the company to search for an accountable person during an annual audit.
Once the role is clear, the company can decide which obligations the supplier must meet, which evidence it must retain itself, and which incidents require a joint response. The next question is no longer who is responsible. It is how much control the system needs in the setting where it is actually used.
2. Intended use and impact determine the level of control
The AI Act does not impose the same obligations on every system. It distinguishes prohibited practices, high-risk systems, systems subject to transparency duties, and systems that present minimal risk according to what they do and the impact they may have. The object of classification is not the model's name or the category on a vendor's product page. It is the function the model performs inside a specific business process.
The first boundary is the list of prohibited practices. The regulation prohibits certain forms of subliminal manipulation or deliberate deception that materially distort behaviour, as well as exploitation of vulnerabilities associated with age, disability, or particular social and economic circumstances when serious harm results. Social scoring, biometric categorisation that infers sensitive attributes, certain uses of emotion recognition in workplaces and schools, and prediction of individual criminal risk based only on profiling are also prohibited or tightly restricted. Real-time remote biometric identification in public spaces is subject to only narrow law-enforcement exceptions.[1]
The practical difficulty is rarely that a company openly purchases a product called "social scoring." More often, a risky function is presented under a softer name. Calling employee emotion recognition "workforce insight" or automated credit denial "customer segmentation" does not change the effect on the person concerned. A review should state in ordinary language whose data the system reads, what judgment it produces, what action follows, and whether the affected person can challenge it. If the company cannot answer one of those questions, it should not continue with an automated release.
A system that is not prohibited may still be high-risk because of its purpose and consequences. High-risk systems mainly arise through two routes. One route covers safety components of regulated products. The other covers the areas listed in Annex III, including biometrics, education, recruitment and worker management, critical infrastructure, essential services, law enforcement, migration and border management, and the administration of justice and democratic processes. Recruitment ranking, credit eligibility, insurance pricing, and allocation of medical resources need stronger controls because they can directly change a person's access to work, services, and opportunity.[1]
The same underlying model may produce different classifications in different workflows. A language model that helps a recruiter draft a job description is not equivalent to a system that automatically decides who receives an interview. A customer-service assistant that drafts a refund response is not equivalent to a system that closes an account. The classification record must describe the intended purpose, input data, affected groups, human decision points, and likely consequences. Without that information, the company has no defensible reason for choosing one level of control over another.
Some systems are not high-risk but still affect a user's understanding of the content or the party with which they are interacting. Transparency obligations reduce this information imbalance. A chatbot will usually need to disclose that it is a machine at the beginning of the interaction. Synthetic images, audio, and video may require labelling in relevant settings, with particular care for deepfakes and content concerning matters of public interest. A system that infers emotions in the workplace or categorises people by biometric characteristics may move into a much more restrictive category.[1]
Minimal risk does not mean no governance. An internal writing assistant should not receive confidential customer information. A coding assistant must not expose credentials or unreleased code. Employees need to know how to verify factual claims and links produced by a model. Risk classification is not an exemption label. It directs more resources to more consequential uses while preserving a minimum set of controls for every system.
3. Compliance depends on an evidence chain across the supply chain
When a system is classified as high-risk, the company must prove more than the existence of a policy. Risk management, data governance, technical testing, human oversight, and operational monitoring need to form a continuous chain. If each document exists in isolation but the company cannot show how a risk was identified, which control reduced it, and how the remaining risk is monitored after release, a large file archive will still fail to provide credible evidence.
Risk management should begin with specific harms. The team identifies possible effects on safety, discrimination, privacy, and fundamental rights, estimates their likelihood and severity, sets an acceptable range, applies mitigation, and then uses operating data to measure residual risk. Accuracy is only one metric. Error rates across groups, distribution shift, explainability, robustness, and cybersecurity also belong in testing. Comparable benchmarks should be used before and after model updates so the company can see whether a risk was reduced or merely moved somewhere else.
Data governance explains why the test results deserve trust. The company should record where training and validation data came from, whether the necessary rights exist, whether the data contains duplicate examples or incorrect labels, whether it represents the real user population, and who approved the cleaning rules. Refusing to collect sensitive attributes does not by itself remove discrimination risk because postcode, school, language, and employment history may act as proxies. If sensitive attributes are genuinely necessary for bias testing, they should be handled in a controlled environment with data minimisation, access restrictions, and a defined deletion date.
Technical documentation and operational logs connect design choices to actual use. Documentation should describe the intended purpose, system architecture, training methods, performance limits, test results, human oversight design, and security measures. Logs should record inputs, outputs, human intervention, anomalies, and version changes. It is not enough for all logs to remain in the supplier's environment. The deployer must at least be able to export records relevant to its own decisions and link them to the model version and operator involved.[1]
Human oversight must also be capable of changing the outcome. Reviewers need training, authority to pause or override a model output, and enough time to make an independent judgment. In recruitment, credit, healthcare, and public services, a model score cannot simply become the final decision. The interface should not make "accept model recommendation" the only practical default. When a reviewer disagrees with the model, the system should retain the reason. That record supports an appeal and helps the team identify persistent defects.
Security and robustness testing should cover prompt injection, data poisoning, model inversion, unauthorised access, adversarial examples, and service degradation. Generative features also require tests for hallucination, repetition of sensitive information, and jailbreak prompts. Every failed test needs a defined response such as delaying release, limiting the use case, or retesting after a fix. It cannot remain an "accepted issue" with no owner or deadline.
Where applicable, providers must also maintain a quality management system, complete conformity assessment, issue the EU declaration of conformity, apply CE marking, register the system, monitor it after release, and report serious incidents. CE marking applies to a particular version and intended purpose. It does not make every future update automatically compliant. Deployers therefore need to connect the supplier's declaration, the system version, and their own recorded use. Only then do risk identification, technical controls, release decisions, and operating monitoring form a complete evidence chain.
That evidence chain cannot stop at the company's own boundary. Providers of general-purpose AI models must prepare technical documentation, give downstream providers information about capabilities and limitations, maintain a policy for compliance with EU copyright law, and publish a sufficiently detailed summary of training content. Providers of models with systemic risk face additional duties for model evaluation, identification and mitigation of systemic risks, reporting of serious incidents, and adequate cybersecurity and model security.[3]
These obligations do not allow a buyer to transfer all responsibility to the model provider. A company first needs to establish whether it is buying an API, downloadable weights, a fine-tuning service, or a feature embedded in office software. The delivery model determines which logs the company can see, which versions it can control, how it can exit the service, and whether the supplier may use input data for training. A procurement questionnaire should request the model name and version, retention rules, retraining policy, processing locations, copyright complaint channel, incident notification period, log export method, and rollback arrangements.
Contracts should follow the actual division of responsibility. The supplier handles model-level documentation, vulnerability fixes, and notice of major version changes. The company handles the business purpose, user notices, human oversight, and final decision. Both parties share responsibility for incident classification, regulatory enquiries, and exit exercises. If the supplier refuses to provide necessary evidence, the company can restrict input data, disable automated execution, or reduce the risk of the use case. It cannot fill an external black box with an internal verbal assurance.
When a company performs extensive fine-tuning, changes the intended purpose, or incorporates the model into a high-risk system, it must reassess whether it has taken on provider obligations. Model cards and supplier white papers are useful inputs, but they cannot replace independent testing of the company's actual application. GPAI compliance is therefore not an extra procurement questionnaire. It brings supply-chain transparency into the company's own chain of accountability.
The evidence chain cannot remain in back-office documents either. It must reach the interface and the business process that users can see. A user should know at the first interaction that a chatbot is a machine. The notice can be brief, but it cannot be buried in terms that few people will find. For synthetic video, audio, and images, the production process should record the generation tool, model version, human edits, and publication channel. Visible labels and verifiable provenance should be used together where possible because metadata may disappear as content moves between platforms.
Consumers are not the only people who need notice. Employees need to know which content came from a model, which fields came from external data, and which decisions must be made by a person. An unexplained grey icon does not tell a user how much confidence to place in an output. For customer-supplied models or embedded features, the contract should also state who applies the label, who responds to complaints, and who retains the relevant records.
More importantly, transparency must connect to the protection of fundamental rights. Recruitment, credit, insurance, education, and public services affect a person's opportunities and dignity. Certain deployers of high-risk systems must complete a fundamental rights impact assessment before first use, identifying the affected groups, likely consequences, mitigation measures, and oversight arrangements. An assessment conducted only inside the development team can easily miss the problems experienced by real users, so the process should involve users, employee representatives, or other affected groups where appropriate.[1]
This work must also be coordinated with the GDPR. A company still needs a lawful basis for processing personal data, a defined purpose and retention period, a process for data subject requests, and a data protection impact assessment where required. For automated decisions that produce legal or similarly significant effects, the company also needs to consider Article 22 of the GDPR, human intervention, and appeal rights, alongside employment, anti-discrimination, consumer, and sector-specific law.
Common use cases show how these obligations connect. A recruitment screening system needs group-level testing, human review, and a route for candidates to challenge an outcome. A customer-service bot must identify itself as a machine and hand a case to a person before closing an account or changing a contract. Credit and insurance systems need more than an explanation of the model. They must show that a marketing segment was not quietly turned into a decision on eligibility. Notice, review, and appeal in the product experience are the user-facing end of the evidence chain.
4. Sustainable operations require governance in everyday workflows
A complete chain of accountability cannot depend on one compliance employee maintaining it by hand. A company may create a cross-functional AI governance committee, but the committee needs authority to approve and suspend systems. Product leaders explain the purpose and user impact. Engineering leaders explain testing and security. The data protection officer covers personal data. Procurement and legal teams cover suppliers and contracts. Information security leads incident response. Each system must also have an owner who can make a real decision during an incident, not simply a name on an organisation chart.
Existing workflows need corresponding controls. Project intake confirms the purpose and legal role. Procurement reviews supplier evidence. Release review checks the risk classification and test results. Operations monitor drift, complaints, and human override rates. Retirement includes data deletion and closure of access. This is not a separate compliance structure placed beside the business. It is the integration of AI Act requirements into the product, procurement, security, and audit processes the company already uses.
Employee policy should state which data must not be entered into external tools, which outputs require human verification, how errors are reported, and when use must stop. Training should use examples from the work people actually perform. Customer-service staff need to recognise hallucinations and know when to transfer a case. Developers need to prevent credential leakage. Recruiters need to understand model bias and appeals. Managers need to read risk reports. AI literacy is not a single training video. It is the ability of each role to make judgments proportionate to its responsibilities.[2]
Incident exercises test whether these arrangements work. A team can begin with an incorrect output or a data leak and then determine who assesses the impact, who preserves evidence, who suspends the system, and who notifies customers and regulators. If the exercise still requires an urgent search for contacts or a debate about which contract applies, the chain of accountability is not yet operational.
Each system should have its own evidence pack, named by system ID, version, and date so the material remains usable over time. The pack should include the role card, classification memorandum, data-flow map, supplier due diligence, contracts and change records, risk management report, test results, human oversight plan, user notices, logging policy, incident records, training records, and exit plan. It is not simply a document repository. It should answer why the system was approved, what changed, who made each decision, and how a problem was handled.
A control register can connect these materials. Each control should state its objective, evidence, owner, review frequency, and response to failure. For example, "complete regression testing before a major model update" should point to a versioned test report and require release to be blocked until the governance committee reviews a failure. When control objectives and evidence correspond one to one, the company can use the same compliance work for customer due diligence, internal audit, and comparison of future models.
For a company that does not yet have systematic AI governance, the first 90 days can launch an initial chain of accountability, but they cannot mark the end of compliance. During the first 30 days, the company should appoint accountable owners, pause unregistered AI procurement, inventory approved systems and employee-purchased tools, and complete an initial role and risk classification. Suspected prohibited practices, systems that make consequential decisions without meaningful human involvement, and systems with no clear data provenance should be suspended or restricted first.
Days 31 to 60 are for installing the most important controls. High-risk systems need baseline tests, human oversight, and a logging plan. Procurement teams should obtain model versions, data policies, and incident mechanisms from suppliers. Legal teams should update notification, audit, exit, and major-change clauses. The aim is not to make every document perfect at once. It is to give each serious gap an owner, a deadline, and a verifiable condition for closure.
Days 61 to 90 are for assembling evidence packs and running an exercise. The team can start with one incorrect output and walk through discovery, suspension, notification, repair, and review. The governance committee then closes critical gaps, while the board reviews unresolved risks, budgets, and business trade-offs. When a problem cannot yet be closed, the company should record why the risk is being accepted, how long that decision remains valid, and when it will be reassessed. A temporary exception should not become permanent by default.
After 90 days, the company still needs to review systems as versions, uses, and regulatory expectations change. A supplier may update the model, a business team may connect a drafting tool to a formal decision, or the affected user population may change. Any of these events can trigger a new classification. The value of the 90-day plan is that it starts a repeatable operating mechanism. It does not give the company permission to return to its old release process after one concentrated programme of work.
Management must remain responsible for that mechanism. The AI Act sets substantial administrative fines according to the type of violation. A breach of the prohibited-practices rules can carry a maximum fine of EUR 35 million or 7 percent of worldwide annual turnover for the preceding financial year, whichever is higher. Other obligations generally carry a top tier of EUR 15 million or 3 percent of worldwide turnover. Supplying incorrect, incomplete, or misleading information to authorities has a separate ceiling. The final penalty also takes account of factors such as company size, duration, intent or negligence, and cooperation.[1]
Yet a board that looks only at fines will underestimate the operating effect of the regulation. A system may have to be suspended, withdrawn, or modified. Customer contracts may trigger compensation and audit rights. Employees and consumers may bring claims under data protection, employment, or consumer law. The more immediate consequences may be slower product releases, difficulty replacing a supplier, and failed customer due diligence. A company may have a capable model and still be unable to expand in Europe because it cannot produce the evidence required to use it responsibly.
Board members do not need to read model code, but they should regularly see the number and risk distribution of AI systems, unresolved gaps, major version changes, supplier incidents, user complaints, human override rates, training coverage, and governance budgets. These indicators show whether the company can operate AI over time. A count of forms completed by the legal team does not.
Ultimately, as the EU AI Act moves into implementation, compliance becomes a basic capability for deploying and operating AI in the European market. An effective system does not end with a risk questionnaire or a set of technical documents. Role assessment determines responsibility, risk classification determines the strength of controls, and technical and organisational measures produce evidence that can be tested over time. Users who are affected must also be able to obtain an explanation, human review, and correction.
Companies do not need to wait until every implementation detail is settled before acting, and they should not treat compliance as a project that can be completed once. What they need is a governance mechanism that changes with model versions, business uses, and regulatory requirements. Companies that can continue to demonstrate clear responsibility, controlled risk, and effective response will be in a position to use AI in Europe over the long term, launch new products, and earn customer trust.