AI SDLC
How to Build an AI-Ready Engineering Organization
Readiness is an evidentiary problem long before it is a tooling problem.
Executive brief
An AI-ready engineering organization is one that can still explain and defend its production changes when authorship is partly synthetic. It baselines outcomes before adoption becomes ambient, captures provenance and verification evidence at the moment of change, tiers autonomy by consequence, funds review capacity, tests vendor exit, and preserves the competence needed to hold machine-authored work accountable.
Key takeaways
- AI readiness is the ability to attribute, verify, reconstruct, and defend AI-assisted production changes—not an inventory of licences, seats, or training.
- Once AI becomes ambient, clean productivity counterfactuals disappear; outcome baselines must be captured before broad deployment.
- Every change should retain structured provenance, a named accountable human, and contemporaneous verification evidence.
- Autonomy belongs on a consequence-based tier enforced by the platform, with the strictest boundaries around critical and regulated functions.
- Verification capacity must be funded and measured because approval events alone do not prove substantive review.
- DORA portability requires tested exit capability, not a document describing an untested fallback.
- Preserving apprenticeship and comprehension is an operational-resilience control because tomorrow's reviewers are produced by doing today's work.
Readiness is an evidentiary problem long before it is a tooling problem.
TL;DR
- AI readiness is not a licence, seat, or training inventory; it is the ability to attribute, verify, reconstruct, and defend AI-assisted changes.
- Broad adoption is closing the window for clean before-and-after productivity baselines, so institutions should govern to defensible outcome metrics.
- AI increases delivery throughput while amplifying weaknesses in review, stability, integration, and accountability.
- Provenance must be captured when a change is made: the assistive system and version, degree of assistance, accountable human, and verification evidence.
- Autonomy should be tiered by consequence and enforced in the platform, with stricter boundaries for critical functions, regulatory reporting, client money, and sensitive data.
- Verification capacity, portability, and tested exit plans must be funded as operating capabilities rather than treated as documentation or residual work.
- AI-ready organisations deliberately preserve the apprenticeship and comprehension pathways that produce tomorrow's reviewers.
The question arrives in a predictable form. A steering committee wants to know whether the engineering organisation is ready for AI, and what comes back is an inventory: licences procured, seats allocated, a training curriculum, a pilot in the payments team that went well enough to justify the next tranche. That inventory is a competent answer to a procurement question, which is not the question that was asked, though it usually takes two or three quarters for the difference to become visible.
The difference becomes visible at the point where somebody has to explain a change. Not defend it — explain it. A production incident, a supervisory review of the change management framework, an internal audit sample, a client complaint that traces back to a calculation. The question is always some version of the same one: who made this change, on what authority, what was checked before it went in, and why did the checks not catch this. For thirty years the engineering organisations in regulated institutions have been able to answer that question because the answer was structurally cheap to produce. A named person wrote the change. A different named person approved it. Both were available and both remembered something. The four-eyes principle, segregation of duties, the change advisory board, the SOX-era control narrative — all of it rested on the assumption that authorship was human, singular, and attributable by default.
That assumption is the thing that is dissolving. The obligation to answer has not moved at all.
This is, I think, the useful frame for AI readiness, and it is a less comfortable frame than the tooling one because it does not resolve with a purchase order. An AI-ready engineering organisation is one that can still produce evidence about its own changes when a meaningful share of those changes were drafted by a system that does not remember, cannot be deposed, and will have been silently replaced by a newer version before the audit cycle closes. Everything else — platform quality, developer experience, training, model selection — is downstream of that, and mostly serves it.
A note on nomenclature, because the collision causes real confusion in steering papers: DORA in this piece means the Digital Operational Resilience Act, Regulation (EU) 2022/2554, unless it is explicitly the DORA research programme, meaning DevOps Research and Assessment, now run out of Google Cloud. Both are relevant here. They are unrelated. I have seen a board pack conflate them.
The Problem
The industry mislaid its control group
The most instructive thing published on AI and developer productivity in the last two years is not a finding. It is a methodological retreat.
In July 2025, METR published a randomised controlled trial of experienced open-source developers working on repositories they had contributed to for an average of five years. Sixteen developers, 246 real tasks, each randomly assigned to permit or forbid AI assistance. The developers forecast that AI would cut their completion time by roughly a quarter. Afterwards, having done the work, they estimated it had sped them up by about twenty per cent. The screen recordings and the clock showed they had been about nineteen per cent slower. METR was careful about scope — mature codebases, high implicit quality standards, early-2025 tooling — and explicitly declined to generalise. The finding that travelled was the slowdown. The finding that mattered was the gap between what the developers believed and what was measured, in a population selected for experience and for prior familiarity with the tools.
METR then tried to repeat the study with better tooling and a larger pool, and in February 2026 published the result, which was that they could not. Not that AI had helped or not helped, but that the experimental design had stopped working. Developers declined to enrol because they did not want to spend half their tasks without AI. Between thirty and fifty per cent of participants said they were withholding tasks from submission — specifically the tasks where they expected the largest uplift, because being randomised into the no-AI arm on those would have been painful. Developers running several agents concurrently could not report time-on-task coherently, because they were working on something else while waiting. The raw numbers showed a swing towards speedup, around eighteen per cent for returning participants, with a confidence interval running from a substantial gain to a small loss. METR's own reading was that the true effect is probably positive and probably larger than they measured, and that their data was too contaminated by selection to say by how much. They are redesigning the study.
I would encourage anyone building an AI adoption business case to sit with that for a moment. A well-funded research organisation, with a clean design, paid participants and screen recordings, could not construct a credible counterfactual eighteen months into broad adoption. The reason is not incompetence. The reason is that once a practice becomes ambient, the unexposed arm stops being a control and starts being a distortion.
Every large engineering organisation is now in that position, and most of them have not noticed. The window in which a defensible before-and-after baseline could have been captured has, for practical purposes, closed for tools already in the estate. What remains is the class of measurement that METR falsified directly: asking people how much faster they are. That number is still the number in most board packs.
The counter-narrative is not more reliable
There is a symmetrical failure worth naming, because a sceptical audience is at least as vulnerable to it, and because credibility here is bidirectional.
The most-cited statistic in enterprise AI is that ninety-five per cent of generative AI pilots fail. It comes from a July 2025 working paper out of an MIT Media Lab initiative, and it is weak in ways that are checkable rather than arguable. The document describes itself as preliminary findings, states that the views are the authors' rather than their institutions', and was not peer-reviewed. Its sample, as printed in the PDF, was fifty-two interviews and a survey of 153 leaders; the mainstream coverage that made it famous reported 150 interviews and a survey of 350 employees, and I have not found a reconciliation of the two. The failure threshold was measurable P&L impact within roughly six months, which is a window that would have condemned most successful data platform programmes I have worked on. And the paper concludes by recommending, as the path across the divide it has just described, the class of agentic infrastructure that the authoring initiative builds.
None of that makes the underlying observation wrong. Pilots do stall at the production boundary, for reasons that are familiar and mostly unglamorous. But an institution that governs its AI programme by that number is governing by a press cycle, and it will find the number difficult to defend when someone asks where it came from. The discipline that applies to vendor productivity claims applies equally to the statistics that flatter our scepticism.
What actually replicates
Set both narratives aside and a smaller, duller, more consistent picture remains, and it is the one worth building on.
The 2025 DORA research programme surveyed close to five thousand technology professionals and found that around ninety per cent now use AI at work. Its headline result is that AI functions as an amplifier: it magnifies whatever the organisation already was. The specific finding that matters for anyone with a control function is the shape of the correlation. Between 2024 and 2025 the relationship between AI adoption and software delivery throughput reversed from negative to positive — teams genuinely are shipping more. The relationship between AI adoption and software delivery instability did not reverse. Higher adoption continues to correlate with more change failures, more rework, and longer recovery. DORA's own language is that AI adoption not only fails to fix instability, it is currently associated with increasing it. Their explanation, drawn from the qualitative work, is that time saved in creation is being reallocated to auditing and verification, and that the systems around the developer have not adapted to the volume.
The structural signals point the same way. GitClear's 2026 analysis, tracking seven code-quality indicators across a large commit corpus, reports duplicated code blocks up eighty-one per cent against 2023 and at the highest level in their record; properly moved or refactored lines falling from around a fifth of changed lines in 2022 to under four per cent in the first half of 2026; cross-file function calls, a rough proxy for genuine reuse, down thirty-five per cent; work on legacy maintenance down roughly three-quarters; and constructs that mask errors rather than handle them up by nearly half. These are correlational, they come from a vendor with a measurement product to sell, and the causal arrow is contestable — some of this is a hiring-mix effect, some of it is a market that has been rewarding feature velocity since well before AI arrived. But the direction is consistent across independently produced datasets, and it is consistent with what the DORA instability finding would predict.
Developer sentiment has moved in the same direction as the structural data, which is not always the case and is therefore worth something. Stack Overflow's 2025 survey of roughly forty-nine thousand developers found adoption at eighty-four per cent using or planning to use AI tools, with trust in output accuracy falling to about a third from over forty per cent the year before, and active distrust at forty-six per cent. Just over three per cent reported high trust, and the figure was lower among experienced developers than among learners. The leading frustration, cited by two-thirds, was output that is almost right but not quite — which is precisely the failure mode that is expensive to catch, because it survives a fast review.
The honest summary is this. Generation capacity has increased, probably substantially, and the increase is real rather than imagined. The organisational capacity to verify, integrate, and take responsibility for the output has not increased at anything like the same rate, and in several measurable respects has degraded. That is not a claim about model capability. It is a claim about the shape of an organisation, and it is the claim that supervisors will eventually make in their own vocabulary.
The governance problem underneath
FINMA got there early, and in language worth quoting to a board.
Guidance 08/2024, published in December 2024 and drawn from ongoing supervisory observation rather than from theory, sets out expectations for governance and risk management where supervised institutions use AI. Switzerland has no AI-specific statute; the guidance rests on the existing technology-neutral framework, which is exactly why it is useful — it is a description of how current obligations bind, not a new regime to be phased in. The expectations are unremarkable on their face: a centralised inventory of AI applications, clearly assigned roles and responsibilities, due diligence and explicit contractual allocation of liability for outsourced solutions, and the requirement that institutions understand how their applications function rather than merely checking outputs.
Two of FINMA's supervisory observations are the ones I would put in front of an engineering leadership team. The first is that institutions frequently could not explain or critically assess the results their models produced, which undermines any assurance about robustness or accuracy. The second is that assigning clear responsibility when an AI system errs is genuinely difficult, and that FINMA regards this as a supervisory challenge rather than a solved problem.
That second observation is the whole essay in one line from a regulator. The accountability model that engineering governance is built on assumes an attributable author. Where the author is partly a model, the attribution has to be reconstructed from evidence, and the evidence has to have been captured at the time. It cannot be assembled afterwards, because the model that wrote the change has been deprecated, the prompt was not retained, the reviewer approved forty other changes that week, and nobody wrote down which parts they actually read.
Why It Matters
The clock is running, and the deferral is not a reprieve
The EU AI Act timeline moved this year, and moved in a way that is easy to misread as breathing room.
The Digital Omnibus on AI, proposed by the Commission in November 2025, reached provisional political agreement on 7 May 2026, was endorsed by Parliament on 16 June and given final approval by the Council on 29 June. The effect is a staggered deferral of the high-risk obligations: stand-alone Annex III systems move from 2 August 2026 to 2 December 2027, and AI embedded in regulated products under Annex I moves to 2 August 2028. Certain transparency obligations, including watermarking for systems already on the market, shift to December 2026 rather than August. The architecture of the Act — the risk tiers, conformity assessment, the GPAI track, the AI Office — did not move.
Sixteen months is a substantial deferral and it was granted because implementation was visibly off track, not because the obligations were reconsidered. The institutions that will be comfortable in December 2027 are the ones treating the interval as a build window for technical documentation, conformity assessment evidence and human-oversight design. The ones that will not be comfortable are the ones that have quietly rebaselined the programme and released the team. I have watched this pattern play out through three regulatory deferrals in two decades and the outcome has been the same each time: the deferral is consumed by the organisation's ordinary appetite for other work, and the eventual compliance effort is compressed, expensive, and staffed by whoever is available rather than whoever is competent.
AI vendors are already inside the third-party risk perimeter
The Digital Operational Resilience Act has applied since 17 January 2025 and it does not have an AI carve-out, which is the point. Providers of AI services are ICT third-party service providers. They belong in the Register of Information alongside everything else, they are subject to the same contractual requirements, and under Article 28(8) the exit strategy for any arrangement supporting a critical or important function must be documented, kept current, and tested.
In November 2025 the European Supervisory Authorities designated the first nineteen critical ICT third-party providers, drawn from the registers that financial entities had submitted, and brought them under direct oversight including inspections. The list is dominated by hyperscale cloud and platform providers — which is to say, by the same small set of organisations that also supply the frontier model capacity most institutions are building on. Concentration risk in this domain is no longer an architectural opinion; it is a named supervisory object with an oversight framework attached to it, and the institution retains its own responsibility for managing the dependency regardless of what the ESAs do with the provider.
There is a practical consequence that surfaces in every register review I have been near. The entry that reads "AI coding assistant, developer productivity tooling" and the entry that reads "code generation within the change path for systems supporting a critical or important function" describe the same contract and imply completely different obligations. Only one of them survives contact with a supervisor who asks what the tool actually touches. Most institutions have filed the first.
The supervisory logic is not new, which is the good news
It is worth remembering that the industry has already solved a structurally similar problem once, under duress.
Quantitative model governance was built after a decade in which institutions discovered they could not explain their own numbers. The answer was not to ban models. It was inventory, tiering by materiality, independent validation proportionate to that tiering, documented limitations, ongoing monitoring, and clear ownership. BCBS 239 and the RDARR programme applied the same instinct to data: a G-SIB should be able to reconstruct how a reported figure was produced, through documented lineage, with the golden record identifiable and the transformations traceable. Nobody enjoyed implementing either. Both are now unremarkable furniture.
The uncomfortable observation is that AI coding assistants are currently governed almost everywhere as tooling — a procurement decision, a licence line, an IT security review — when the shape of the risk is closer to a model risk problem sitting inside a change-control problem. They influence the content of production changes. They are versioned, they drift, their behaviour is not reproducible across versions, and their failure mode is confident plausibility rather than obvious breakage. An institution that already runs a model inventory with materiality tiering has most of the conceptual machinery. It has simply not pointed it at the software development lifecycle, largely because the tools arrived through a channel that was never designed to catch this class of thing.
The competence dimension is a supervisory question, not an HR complaint
The labour market data has firmed up enough to be worth acting on, with appropriate care about what it does and does not establish.
Stanford's Digital Economy Lab, working with ADP payroll records across millions of workers, reports that employment for software developers aged twenty-two to twenty-five fell by close to twenty per cent from its late-2022 peak through mid-2025, while employment for developers aged thirty and over at the same firms grew somewhere between six and twelve per cent over the same period. The researchers have been careful, and their February 2026 update noted that under the broadest set of controls the decline becomes statistically significant only from 2024, with earlier movement plausibly attributable to non-AI factors. They describe the work as measuring correlation between AI exposure and employment, not a clean causal effect. Take it, then, as a strong signal rather than a proof.
The reason this belongs in an operational resilience discussion rather than a workforce planning one is that the tasks being displaced are the tasks through which reviewers were historically produced. Boilerplate, test scaffolding, well-specified features, bug triage — tedious, and also the apprenticeship. The capacity to look at a plausible-looking change and notice that it is subtly wrong is not taught in a curriculum. It is accumulated by having written a great deal of code badly under supervision. An institution that removes the accumulation mechanism while increasing the volume of output requiring judgement has created a succession risk with a five-to-ten-year fuse, and succession and key-person concentration are categories that supervisors already have vocabulary for.
The complementary finding is the one I would watch internally. Controlled work on learning outcomes suggests the mode of use determines the effect: delegating implementation to a model correlates with materially lower comprehension of the resulting system, while using the same model for conceptual inquiry does not. Two engineers with identical AI usage statistics can therefore be on opposite trajectories, and no dashboard currently in production distinguishes them.
A Practical Model
What follows is ordered as a sequence rather than a menu, because the layers depend on each other and because the first one has a closing window.
1. Baseline before you scale, and accept where the window has already shut
The METR retreat is a warning about your own measurement programme. Once a tool is ambient across the estate, you cannot construct the counterfactual, and any subsequent claim about its effect will rest on self-report, which has been directly falsified in the one setting where it was rigorously tested.
For anything not yet broadly deployed — agentic tooling with commit or pipeline authority is the current example in most institutions — capture the baseline now, at team level, on outcome measures that generation volume cannot inflate: change failure rate, time to restore, incidents per change merged, rework within fourteen days, proportion of changes merged without substantive review, defect escape to production. For tools already ambient, be honest in the papers that a controlled baseline is no longer available and say what you are using instead. That honesty costs less than it appears to. An institution that says "we cannot isolate this effect, so we are governing to the outcome metrics" is in a stronger supervisory position than one presenting a confident percentage with no defensible derivation.
Self-reported productivity uplift should not appear in a board pack. If it must appear, it should be labelled as sentiment.
2. Provenance by default, captured at the moment of change
Every change entering the estate should carry, as structured metadata rather than convention: the identity and version of any assistive system involved, the degree of assistance, the named human who is accountable for the change, and the evidence of what was verified before merge. This is the lineage discipline that BCBS 239 forced onto data, applied to the artefact that produces the data.
Two design cautions, both learned the expensive way. First, do not make the prompt the artefact of record. Prompts are not reproducible — the model version has moved, the context window differed, the temperature was not yours to control — and an audit trail built on them creates an appearance of reconstructability that will not survive being tested. Capture the decision and the verification, not the incantation. Second, this must not be built or presented as attribution-for-blame. Engineers will degrade the data quality of any system they read as surveillance, and they are quite capable of doing so without technically violating anything. The purpose is reconstruction, and it should be stated as such in the policy, repeatedly.
3. Tier by consequence, not by enthusiasm
Autonomy in the change path should be granted against blast radius, not against team appetite or pilot success. The tiering that has held up in practice runs roughly: systems where a defect is recoverable within the sprint and visible immediately; systems where a defect reaches a client or a report but is correctable; systems supporting a critical or important function under DORA, or feeding regulatory reporting, or touching client money or client data at scale.
Machine-authored change can land in the first tier with lightweight verification. In the second it requires substantive human review with recorded evidence of what was examined. In the third the question is not the depth of review but whether a human authored the change, with the model used for analysis, test generation and adversarial critique rather than for the change itself. That last boundary will feel conservative in 2026 and may well be relaxed by 2028. It should be relaxed on evidence produced by your own tiering data, not on a vendor benchmark.
Set these trust boundaries in the platform, not in policy documents. A boundary that depends on an engineer remembering which tier they are in is not a boundary.
4. Fund verification as a capacity, not as a residue
This is the layer institutions consistently underfund, and the data on why is unambiguous. Review is now the binding constraint, and it has been treated as a free byproduct of having engineers.
If generation capacity rises materially, verification capacity must be planned rather than assumed, and it should appear in the operating plan as its own line with its own headcount. The practical instruments are unglamorous: a hard ceiling on change size, because review quality collapses against large diffs and the diffs are getting larger; measurement of review depth rather than review completion, since approval events are trivially gamed; explicit review budgets in team capacity planning; and automated verification — property tests, contract tests, mutation testing, static analysis tuned for the specific failure modes of generated code, particularly duplication and swallowed exceptions — treated as the primary defence rather than as a supplement to human attention that is now spread far thinner.
The signal to watch is the proportion of changes merging without substantive review. Where that is rising, throughput gains are being converted into deferred instability, and the conversion is not visible in any velocity metric.
5. Portability and exit discipline
Keep the model a replaceable component. Vendor-specific semantics should not accumulate in the control plane of the development lifecycle, and the abstraction boundary should be maintained deliberately rather than emerging by accident from whichever SDK was convenient. This is ordinary portability discipline and it is well covered elsewhere; the point specific to this argument is that DORA Article 28(8) does not ask whether you have written an exit strategy. It asks whether you have tested one. For an assistive tool embedded in the change path of a critical function, testing the exit means demonstrating that the organisation can continue to produce and verify changes at an acceptable rate without it — which is an operational rehearsal, not a document.
Institutions that have run that rehearsal report the same thing: the dependency was deeper than the register suggested, and the discovery was worth the disruption.
6. Maintain the competence you are consuming
Reviewers are produced by doing the work that is now being automated. If that pathway is not deliberately reconstructed, the institution is drawing down a reserve it cannot quickly rebuild.
The measures that seem to work are specific rather than cultural. Protect a defined portion of early-career work as unassisted, framed as training rather than restriction, and be candid about why. Require, in review, that the accountable engineer can explain the change without reference to the tool that produced it — this single norm does more than any policy, and it is enforceable without instrumentation. Rotate engineers through incident response and legacy maintenance, which are the two activities where the gap between understanding a system and generating code for it becomes immediately apparent. Assess on comprehension and judgement rather than throughput, since throughput is now a measure of the tool. And separate, in whatever internal telemetry exists, delegation from inquiry, because the productivity numbers do not distinguish an engineer who is compounding their capability from one who is spending it.
The Checklist
Framed as the questions an internal audit function or a supervisor would ask. An institution that can answer these with evidence is ready in the sense that matters; one that can answer them with intentions is not.
Inventory and attribution
- Is there a single inventory of AI systems in use across the engineering estate, including tools adopted through individual or team-level procurement?
- For each entry, is the accountable owner named, and is the materiality tier recorded?
- Can you identify, for any production change in the last twelve months, whether an assistive system was involved and which version?
- Does the Register of Information under DORA describe what these tools actually touch, or does it describe them as generic productivity tooling?
- Is a named human accountable for every change, irrespective of how it was authored?
Evidence and reconstruction
- For a change in a critical or important function, can you reconstruct what was verified before merge — not that it was approved, but what was examined?
- Is verification evidence captured at the time of change rather than assembled on request?
- Would that evidence survive the deprecation of the model that produced the change?
- Can you explain the behaviour of AI-influenced components to a supervisor, in the sense FINMA 08/2024 uses, rather than only demonstrating that outputs were checked?
Boundaries
- Are systems tiered by consequence, and is the tiering enforced in the platform rather than in policy?
- Is there a defined boundary beyond which machine-authored change may not land without human authorship?
- Who can move that boundary, and what evidence is required to move it?
- Are agentic tools with commit, pipeline or infrastructure authority governed under change control, or under tooling procurement?
Capacity
- Does verification capacity appear as a funded line in the operating plan?
- Is the proportion of changes merged without substantive review measured, and is it trending?
- Is change size capped, and has the cap held over the last two quarters?
- Are change failure rate, time to restore, and rework within fourteen days trending against a baseline you can defend?
- Does any productivity figure presented to the board rest on self-report?
Third-party and concentration
- Are AI providers assessed as ICT third-party providers, with contractual liability allocation explicit?
- Has the exit strategy for any AI dependency supporting a critical or important function been tested, not merely documented?
- Is concentration across model providers, cloud providers and tooling providers assessed jointly, given the overlap in the designated CTPP population?
- If a provider becomes unavailable at short notice, what is the demonstrated fallback for the change path?
Competence
- Is there a deliberate mechanism producing engineers capable of reviewing machine-authored change five years from now?
- Can accountable engineers explain their changes without reference to the tool, and is this tested in review rather than assumed?
- Has the reduction in early-career hiring been assessed as a key-person and succession risk, with an owner?
Regulatory readiness
- Has the AI Act deferral to December 2027 been treated as a build window with retained resourcing, or absorbed as relief?
- Is technical documentation for in-scope systems being produced now, at the pace required to be complete before the deadline rather than at it?
Closing
None of this is exotic, and that is deliberate. The instruments are inventory, tiering by materiality, independent verification proportionate to consequence, documented lineage, tested exit, and a succession pipeline — the same instruments the industry built for quantitative models after 2008 and for risk data aggregation after BCBS 239, applied now to a new and unusually prolific class of author. The novelty is not in the controls. It is in the recognition that a control framework resting on attributable human authorship needs re-anchoring when authorship becomes partly synthetic, and that the re-anchoring has to happen while the volume is rising rather than after.
The organisations that will look prepared in 2028 are the ones that were unfashionably boring in 2026: capturing provenance nobody was asking for yet, funding review capacity that did not show up in a velocity metric, protecting apprenticeship work that looked inefficient, and testing an exit from a vendor everybody assumed would still be there. None of that will feel like readiness at the time. It will feel like operational friction, and there will be internal pressure to remove it.
That pressure is worth resisting, and it is worth being specific about why. The engineering organisations that got into trouble in previous technology transitions were rarely the ones that adopted too slowly. They were the ones that adopted at a rate their evidence-production capacity could not match, and discovered the gap at the moment they were asked to explain something. The tooling question resolves itself; vendors are competent and the capability is genuine and improving. The evidentiary question does not resolve itself, and it is the one being asked.
References
- METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — Randomised controlled trial of experienced open-source developers, including measured slowdown and the gap between measured and perceived productivity.
- METR: We Are Changing Our Developer Productivity Experiment Design — Follow-up documenting participation, task-selection, concurrent-agent, and time-measurement limitations as AI adoption became ambient.
- DORA: State of AI-assisted Software Development 2025 — Research framing AI as an amplifier of existing organisational strengths and weaknesses.
- DORA: Balancing AI tensions—moving from adoption to effective SDLC use — DORA analysis connecting higher AI adoption with both increased delivery throughput and increased delivery instability.
- GitClear: Software engineering research studies — Vendor-produced code-quality research treated as correlational and directional rather than causal evidence.
- Stack Overflow: 2025 Developer Survey—AI — Survey evidence on AI adoption, trust, distrust, and the verification concerns reported by developers.
- FINMA Guidance 08/2024: Governance and risk management when using AI — Swiss supervisory expectations for AI inventory, governance, accountability, explainability, outsourcing, and risk management.
- FINMA: Survey on AI use at Swiss financial institutions — Survey evidence on adoption and planned AI use across supervised Swiss financial institutions.
- European Parliament: AI Act simplification measures approved — Official record of Parliament's 16 June 2026 approval and the revised application dates for high-risk systems.
- Council of the European Union: Final approval of the Digital Omnibus on AI — Official Council document recording final approval on 29 June 2026.
- EUR-Lex: Regulation (EU) 2022/2554—Digital Operational Resilience Act — Primary legal text, including Article 28 requirements for documented, sufficiently tested, and periodically reviewed exit plans.
- European Supervisory Authorities: Critical ICT third-party providers designated under DORA — Official 18 November 2025 designation announcement and explanation of the Register of Information-based assessment.
- Basel Committee: Principles for effective risk data aggregation and risk reporting — Primary BCBS 239 source for traceability, governance, and reconstruction principles used as an institutional analogy.
- Stanford Digital Economy Lab: Canaries in the Coal Mine? — Payroll-based analysis of employment changes among early-career workers in AI-exposed occupations, with the authors' causal caveats.
- Stanford HAI: 2026 AI Index—Economy — AI Index summary of labour-market signals, including changes affecting software developers aged 22 to 25.
- Shen and Tamkin: How AI Impacts Skill Formation — Controlled study of AI-assisted coding, comprehension, and the differences between delegation and conceptual-inquiry usage patterns.
- MIT NANDA: The GenAI Divide—State of AI in Business 2025 — Archived preliminary working paper cited critically to examine the sourcing and definition behind the widely repeated 95 percent claim.
Author
Géza Kuti is a senior Data and AI executive based in Bülach (ZH), Switzerland, focused on data strategy, enterprise architecture, AI governance, hybrid cloud, and regulated delivery.
Related expertise
Related articles
From Zero to Scale: Seven Principles for Building a High-Performing Data & AI Organisation
Seven practical principles for building and scaling a Data & AI organisation: start with the business mandate, hire for judgement, develop leaders early, make accountability explicit, use governance to accelerate delivery, design for multicultural work, and measure organisational value.
Hiring for the Stack, Paying for the Context
Swiss employers are pricing tools, locations, and engineering seats while underpricing the institutional context, specification work, and accountability that regulated data delivery requires.
The Record You Cannot Refactor
Why AI lifecycle evidence must be captured while systems are designed, tested, approved, released, and monitored—before models, prompts, suppliers, and people move on.