AI SDLC
The Record You Cannot Refactor
What an AI development lifecycle has to produce, and why the deadline that moved was not the one that mattered.
Executive brief
The AI Act's revised high-risk dates do not defer the creation of lifecycle evidence. Defensible AI systems need contemporaneous records of design choices, versions, evaluations, approvals, releases, monitoring, incidents, and accountability because that history cannot be reconstructed reliably after models, prompts, data, suppliers, and teams have changed.
Key takeaways
- Moving a legal application date does not recreate lifecycle evidence that was never captured.
- Risk decisions, design choices, versions, evaluations, approvals, releases, and incidents must be recorded when they happen.
- The traceable identity of an AI system includes models, prompts, retrieval data, embeddings, tools, policies, parameters, and evaluation assets—not code alone.
- Probabilistic systems turn testing into a continuing measurement programme with versioned regression evidence.
- Supplier-controlled model changes make lifecycle evidence and continuous successor evaluation operational-resilience requirements.
- One accountable owner must answer for the assembled system's behaviour, not merely for one component.
- The durable compliance asset is a contemporaneous record that survives changes to technology, suppliers, and personnel.
TL;DR
The approved Digital Omnibus moves the EU AI Act's high-risk application dates, but it does not return elapsed development time. Lifecycle evidence is created while a system is designed, tested, approved, released, and monitored. A system placed on the market in 2027 or 2028 will be judged through a record that must already exist by then. That contemporaneous record is the part of an AI system that cannot be reconstructed reliably later.
Key takeaways
- Article 50 remains on its original application path from 2 August 2026, while the approved Omnibus gives providers until 2 December 2026 to implement artificial-content transparency solutions.
- Article 9 requires risk management across the entire lifecycle and Article 12 requires automatic logging over the lifetime of the system, which together make the compliance artefact retrospective by construction.
- Harmonised standards were not available on the original schedule, so institutions must be ready to explain and defend the methods they use in the meantime.
- The change surface of an AI system extends well beyond code, and most organisations version-control only the part that looks like software.
- The release calendar of a model-dependent system is set partly by a supplier, which is an availability and evidence problem before it is a procurement one.
Sometime in the last eighteen months, most large regulated organisations acquired an AI inventory. Not a strategy, and not a platform, but a spreadsheet — usually maintained by a risk function, usually incomplete, and usually assembled after somebody senior asked how many AI systems the institution was running and received four different answers. The exercise tends to be revealing in a way that its sponsors did not intend. The systems are found. What is rarely found alongside them is any account of how they were built, what was tested before they went live, which model version produced which decision, or who agreed that the residual risk was acceptable. The systems exist. Their history does not.
That gap did not matter much while AI regulation was prospective. It is about to matter a great deal, and the reason is slightly counter-intuitive, because the most widely reported development of the past year has been a deferral.
What actually moved, and what did not
On 29 June 2026, the Council gave final approval to the Digital Omnibus on AI. The approved text sets 2 December 2027 for stand-alone high-risk systems under Annex III and 2 August 2028 for high-risk AI embedded in regulated products. The change was a response to an implementation environment in which standards and other support measures were not ready on the original schedule.
The shorthand that the EU "delayed the AI Act" is imprecise enough to be dangerous. The high-risk application dates moved; the rest of the framework did not move as one block. Prohibited practices have applied since February 2025, and obligations for providers of general-purpose AI models have applied since August 2025. Article 50 transparency duties remain on the path to apply from 2 August 2026. The approved Omnibus adds a three-month implementation period, ending 2 December 2026, for providers to put artificial-content transparency solutions in place.
Those transparency requirements are not confined to high-risk systems. They address interactions with people and the generation or manipulation of content, including disclosure and machine-readable marking duties. For customer-facing generative systems, this is an engineering concern in output paths, product behaviour, and retained evidence—not merely a sentence added to a banner.
For Swiss institutions the picture has an additional fold. Switzerland is not adopting a single horizontal AI act. The Federal Council decided in February 2025 to ratify the Council of Europe Framework Convention on Artificial Intelligence and to implement it through targeted legal amendments and non-binding measures, with a consultation draft due by the end of 2026. The absence of a Swiss AI act is therefore easy to mistake for the absence of AI obligations. A Swiss institution serving the EU market can still fall within the EU AI Act's scope, while FINMA already applies technology-neutral governance and risk-management expectations to supervised firms using AI. There is no single local statute to wait for; the relevant duties arrive through the frameworks the institution already has.
Why a deferred lifecycle obligation is not deferred work
The high-risk requirements are worth reading closely, because their structure is what makes the deferral less generous than it appears.
Article 9 requires a risk management system that is, in the Regulation's own words, a continuous iterative process planned and run throughout the entire lifecycle of the system, subject to regular systematic review and updating, and drawing on data gathered from post-market monitoring under Article 72. Article 12 requires that the system technically allow automatic recording of events over its lifetime, and the point is sharpened by what the Article excludes: manual recording does not satisfy it, because the logs have to serve risk identification, post-market monitoring, and verification of operational performance without depending on somebody remembering to write things down. Article 11 and Annex IV require technical documentation, drawn up before the system is placed on the market, covering the general logic of the system, the key design choices, and the main classification choices. Article 17 requires a documented quality management system spanning design, development, testing, data governance, change management and post-market monitoring. Article 14 requires human oversight arrangements under which a person can actually understand, intervene and override.
Read as a set, these are not disclosure obligations. They are obligations to have conducted the work in a particular way and to be able to demonstrate it. The technical documentation required by Article 11 has to describe key design choices, and a key design choice is a thing that happens once, in a meeting, in 2026. It can be recorded at the time or it can be reconstructed afterwards from memory and inference, and the second of those produces a document rather than a record. The distinction is not academic. Supervisors who have spent a career reading model validation files can recognise a narrative assembled after the fact, usually because it is too coherent — the risks identified are precisely the risks that materialised, the alternatives considered are precisely the alternatives that were rejected, and nothing in the file is inconvenient. Genuine contemporaneous records are messier and considerably more persuasive.
What follows is uncomfortable arithmetic. Systems being designed now, in the deferral window, will still be running in December 2027. They are the estate that will be assessed. Everything built between today and then is being built either with a lifecycle record or without one, and the choice is being made now whether or not anybody is framing it as a choice. Sixteen months of deferral applied to a lifecycle obligation does not defer sixteen months of work; it defers the moment at which the absence of that work becomes visible.
This is also where the analogy that most institutions already possess turns out to be useful. Banks have run model risk management for well over a decade, with inventories, independent validation, ongoing performance monitoring, tiering by materiality, and the principle that a model's owner is not its validator. The discipline is imperfect and frequently resented, but it exists, it is understood by supervisors, and it was built for exactly this problem: systems whose behaviour is statistical, whose inputs shift, and whose failures are gradual. The reason so little of it was applied to generative AI is not that it was judged unsuitable. It is that generative systems did not arrive through the channel that triggers model governance. They arrived as software features, procured as SaaS or built by product teams, and the second line was frequently not asked. Institutions that have started routing AI systems back through the model inventory have found the retrofit painful and have also found that they were roughly two years further along than they thought.
Nobody is going to hand you the method
The candid reason for the deferral is that the compliance infrastructure was not going to exist in time, and the clearest evidence of that is the state of the harmonised standards.
CEN-CENELEC's Joint Technical Committee 21 was asked to produce the European standards that would give providers a presumption of conformity with the high-risk requirements. The original delivery date was April 2025. That slipped to the end of 2025, then to late 2026, and by early 2026 the committee's own reporting indicated that full coverage of the standardisation request was not expected before 2027. In October 2025 the CEN and CENELEC technical boards adopted what they described as an exceptional package of measures to get the foundational work items out by the fourth quarter of 2026, permitting drafts with a positive Enquiry vote to be published without a separate formal vote and convening a small group of already-active experts to finish the six most delayed texts. Members raised concerns about the consequences of compressing a consensus process, and the disagreement was public.
The most instructive explanation offered for the delay was not resourcing or committee size, though both were real. It was that the Regulation asks for standardisation in areas where there is no established state of the art. That is a remarkable admission and it should change how an engineering leader reads the whole exercise. The Act requires demonstrable accuracy, robustness, bias management and human oversight for systems whose properties the field is still learning how to measure. The standards bodies are not late because they are slow. They are late because they were asked to codify a practice that is being invented while it is being required.
The operational conclusion is not that compliance is impossible. It is that the method will not be delivered to you, and that an institution's defensibility will rest on having designed a method, documented why it was chosen, applied it consistently, and retained the evidence. That is a heavier obligation than following a standard, and it is also more durable, because a method with a documented rationale survives a change of standard while a checklist inherited from a vendor does not.
The change surface is larger than the code
Underneath the regulatory question sits an engineering one that most organisations have not yet resolved, and it concerns what exactly is under change control.
A conventional application changes when somebody changes it. That single assumption is what makes the traditional lifecycle work: a release is an event, a version is a commit, a rollback is a redeploy, and a test that passed before the change is expected to pass after it unless the change caused it not to. An AI system violates this quietly. Its behaviour is determined by the model version, the system prompt, the retrieval corpus and the way that corpus was chunked, the embedding model used to index it, the tool definitions available to it, the guardrail policies wrapped around it, the evaluation set used to judge it, and the temperature and sampling parameters set somewhere in a configuration console. Most institutions version-control the application code, keep the prompts in a repository if the team happens to be disciplined, and manage everything else through a mixture of platform settings, a vector database that nobody snapshots, and a wiki page. When something changes in production behaviour, the first question — what changed — frequently cannot be answered, and the honest reason is that the change surface was never fully enumerated.
This is where an AI lifecycle stops being a rebranding of DevOps. The version identity of the system has to include artefacts that are not code and were never treated as configuration, and reproducing a decision from six months ago requires reconstituting all of them together. Article 12 asks for logs sufficient to trace the operation of the system, and a log that records the prompt and the response while omitting which retrieval index version was live, which guardrail policy was in force, and which model snapshot answered is not a trace. It is a transcript.
Testing stops being a gate and becomes a measurement programme
The second engineering break is in test.
Testing has historically been a gate: assertions pass or fail, coverage is a percentage, and a build either ships or does not. Probabilistic systems dissolve that contract. The same input can produce different outputs across invocations because of sampling, context effects, tool latency, or a supplier's change to the underlying weights, so a single fixed test case cannot establish anything on its own. Establishing that a system behaves reliably requires running the same task repeatedly and characterising the variation, which is expensive in a way that traditional test suites are not, and which produces a distribution and a threshold rather than a verdict. The practical shape of this is component-level evaluation — retrieval, planning, tool selection and generation assessed separately, so that a degradation can be located rather than merely detected — combined with deterministic checks wherever the property is genuinely deterministic, because an exact tool name or a required parameter does not need a probabilistic judge.
The evaluation set becomes an asset with a lifecycle of its own, and this is the part that most teams under-invest in. Every production failure ought to become a permanent regression case, because a system that passes its evaluations while failing in production is usually one whose evaluation set has never been updated by contact with reality. The set is also the closest thing an institution has to a specification of intended behaviour, which makes it a governance artefact rather than a test fixture, and it should be versioned, reviewed, and owned accordingly. I have seen an evaluation suite treated as engineering scratch material, deleted in a repository cleanup, and reconstructed from scratch four months later by people who no longer remembered why several of the original cases had been there. The cases that get lost in that process are disproportionately the ones that were added after an incident.
The release calendar belongs partly to somebody else
The third break is the one that most clearly distinguishes an AI system from anything an enterprise architecture function has governed before, and it has nothing to do with probability.
A model-dependent system has a supplier who can change or withdraw the dependency on its own schedule. Retirement dates and upgrade paths vary by model, deployment type, and hosting platform, so the same nominal model accessed through two providers can create two separate lifecycle exposures. An automatic upgrade may solve an availability problem while creating a behavioural one: the system continues to answer, but it may answer differently.
The failure mode is straightforward. A system is tuned against one model snapshot, the supplier changes or retires that dependency, and production quality falls even though no application code changed. A release process that treats deployment as the only moment of change will miss that event unless model identity and behaviour are monitored explicitly.
Two disciplines follow directly. First, production call sites should use identifiable model versions rather than unrecorded rolling aliases wherever the provider permits it; otherwise the system may be unable to show which model produced a decision. Second, evaluation suites should run against likely successor models as well as the incumbent, so replacement behaviour is understood before a retirement deadline. Both controls are easier to establish before the first forced migration.
Where these systems actually fail
The failures I have watched at close range have rarely been failures of a component. They have been failures at the seams between owners.
An AI system in a large institution typically has its model owned by a supplier, its data owned by a domain team with its own roadmap, its prompts owned by a product function, its infrastructure owned by a platform group, its guardrails owned by security, and its risk owned by a second line whose staff cannot read any of the artefacts involved. Each of those parties can answer for its own component. What frequently has no owner is the behaviour of the assembled system in production, and that is the only thing the customer, the supervisor and the complaint file are concerned with. The most valuable structural decision available is also the least fashionable: name a single accountable owner for the system's behaviour, distinct from the owners of its parts, and give that person the authority to stop it. Every governance instrument downstream of that decision works better, and most governance instruments introduced without it become documentation exercises.
The corollary is that the second line has to be able to read the evidence. A validation function that cannot interpret an evaluation report, cannot assess whether a test set is representative, and cannot form a view on whether a threshold is appropriate will fall back on process compliance, which is the failure mode that produces thick files and unexamined systems. This is a hiring and training problem that institutions consistently defer because it is slower to fix than a policy.
The questions the file has to answer
The checklist that belongs at the end of a piece like this is not a maturity model. It is the set of questions an institution should be able to answer from its own records, without a project to assemble the answer, for any AI system it is running. Each maps to an obligation that is either already live or arriving.
| Question | Why it is asked |
|---|---|
| Which model version, prompt, retrieval index, guardrail policy and tool set produced this specific output? | Traceability; Article 12 logging over the lifetime of the system |
| What were the key design choices, and what was rejected and why? | Article 11 and Annex IV technical documentation, drawn up before market placement |
| What was measured before go-live, against what set, and at what threshold? | Article 15 accuracy and robustness; the basis of any conformity argument |
| Who reviewed the residual risk and on what date? | Article 9 lifecycle risk management |
| How would a person have intervened in this decision, and is there evidence they could? | Article 14 human oversight as a design property |
| How is the system monitored in production, and what has it triggered? | Article 72 post-market monitoring; Article 73 incident reporting |
| Is the user told they are dealing with an AI system, and is generated content marked? | Article 50 begins applying 2 August 2026; the approved implementation period for artificial-content transparency solutions ends 2 December 2026 |
| What happens when the supplier deprecates the model? | Operational continuity and behavioural change control |
| Who is accountable for the behaviour of this system, as distinct from its components? | The question everything above ultimately resolves to |
An institution that can answer these from existing records is compliant in substance well ahead of the dates, whatever the standards eventually say. An institution that cannot is holding a portfolio of systems it can operate and cannot explain, and the deferral has given it sixteen months in which to grow that portfolio.
The asymmetry worth remembering
There is a reason to be sympathetic to organisations that read the deferral as relief. The regulation is demanding, the standards that would clarify it are late, the tooling is immature, and the engineering discipline it presupposes is being worked out in public by everybody at once. Nobody involved in this — not the Commission, not the standards committees, not the supervisors, and certainly not the practitioners — is operating from a settled understanding of what good looks like. Architectural humility is the correct posture, and anybody selling certainty about AI governance in 2026 is selling something else.
What is not uncertain is the direction of the obligation. Every instrument now in play, in the EU and through the sectoral route in Switzerland, asks the same underlying thing: show what the system did, show why it was built that way, and show who decided. That question will not soften, because it is the question supervision has always asked, arriving now at systems that were not built to answer it.
A system that turns out to be wrong can be corrected, replaced, retrained or withdrawn, and institutions do this routinely. What cannot be done later is the recording. Code can be refactored, models can be swapped, architectures can be migrated, and the estate that exists in December 2027 will differ substantially from the one being designed today. The record of how it came to exist is the one component that has to be written while the work is happening, and it is the component most likely to be postponed, because it is the only one that nothing visibly breaks without.
Part of the SDOP series on enterprise AI adoption, AI governance, software engineering, platform architecture and continuous modernisation.
References
- Council of the European Union: AI Act legislative timeline — Official timeline recording final Council approval on 29 June 2026, the revised high-risk dates, and the transparency implementation period.
- EUR-Lex: Regulation (EU) 2024/1689—the Artificial Intelligence Act — Primary legal text for lifecycle risk management, technical documentation, logging, quality management, human oversight, transparency, and post-market monitoring.
- European Commission: Standardisation of the AI Act — Official overview of the harmonised-standards programme and its relationship to the revised high-risk application dates.
- CEN-CENELEC: Update on accelerating AI standards development — Official update on exceptional measures adopted to accelerate foundational AI standards work.
- Swiss Federal Council: AI regulation and the Council of Europe Convention — Official Swiss decision to pursue Convention ratification through targeted legal amendments and complementary measures.
- FINMA Guidance 08/2024: Governance and risk management when using AI — Swiss supervisory expectations for identifying, assessing, managing, and monitoring AI risks under technology-neutral requirements.
- FINMA: Survey on AI use at Swiss financial institutions — Evidence on AI adoption, external-provider dependency, and developing governance structures in supervised Swiss institutions.
- Microsoft: Foundry Models lifecycle and retirement policy — A provider-specific example showing that model availability, upgrade, and retirement dates depend on the hosting platform and deployment type.
- AgentAssay: Regression testing for non-deterministic AI agent workflows — Research on token-efficient regression testing for non-deterministic agent workflows.
- Evaluation and Benchmarking of LLM Agents: A Survey — Survey of agent-evaluation methods, benchmarks, and open challenges.
Author
Géza Kuti is a senior Data and AI executive based in Bülach (ZH), Switzerland, focused on data strategy, enterprise architecture, AI governance, hybrid cloud, and regulated delivery.
Related expertise