Skip to main content

AI SDLC

What AI Did Not Make Cheaper

The cost of producing software has fallen sharply. The cost of demonstrating that it is correct, controlled, and defensible has not.

··16 min read

Executive brief

AI-assisted delivery increases the volume of software faster than it improves assurance. In regulated institutions, the durable advantage comes from treating evaluation as a versioned asset, capturing provenance at the point of change, planning review capacity explicitly, and giving every production agent a uniquely attributable identity and accountable owner.

Key takeaways

  • Software generation and verification have decoupled: output volume is rising faster than the proportion that can be shipped safely.
  • Review queues and evidence packs reveal AI-assisted delivery constraints more honestly than coding-speed metrics.
  • Self-reported productivity gains require a defensible counterfactual, which is becoming harder to observe as AI adoption becomes universal.
  • Regulated institutions still owe complete inventories, risk classification, named accountability, testing evidence, limitations, and fallback plans.
  • Evaluation datasets and policy suites should be versioned, owned, and deployed as production assets.
  • Provenance should be captured when a change is made rather than reconstructed after an incident or audit request.
  • Production agents need unique identities, delegated authority, auditable context, and named accountable owners.

The cost of producing software has fallen sharply. The cost of demonstrating that the software is correct, controlled, and defensible has not. In regulated institutions, the second cost was always the real one.

TL;DR

  • Generation and verification have visibly decoupled. Longitudinal testing of code-generating models shows benchmark coding ability improving while the proportion of generated code free of known security flaws has stayed close to flat.
  • The instruments we use to measure AI's effect on delivery are themselves degrading. The best-known randomised trial in this area could not be cleanly repeated, because developers now decline to work without AI.
  • Self-reported gains and measured gains continue to diverge, and the gap is not closing so much as becoming harder to observe.
  • Where AI-assisted delivery lands hardest in a large institution is the review queue and the evidence pack, not the IDE.
  • Supervisory expectations already describe what an institution owes: an inventory, a risk classification, named accountability, and documentation covering purpose, data, model choice, limitations, testing and fallback. None of that got cheaper.
  • The EU's high-risk deadlines moved in June 2026. The obligations did not. Calendar relief is not scope relief.

Two curves diverging over time: model coding capability rising steadily, while the proportion of generated code passing security checks remains roughly flat.

Where it actually shows up

The first place AI-assisted delivery becomes visible inside a large institution is not the development environment. It is the pull request queue, and shortly afterwards the change advisory process, and eventually the evidence pack that someone in second line has to assemble before a release can be defended.

This is not how the business cases were written. The business cases were written around engineering hours, because engineering hours are the part of the system that finance already knows how to price. What the last two years suggest is that the hours were never the binding constraint, and that the activities which did constrain delivery — review, verification, attribution, sign-off — absorbed the additional volume rather than shrinking with it.

The telemetry now available on this is uncomfortable, and worth reading with the usual scepticism reserved for vendor datasets. Faros, which instruments delivery pipelines commercially and therefore has an interest in the finding, reported across roughly 22,000 developers that epics completed per developer rose by around 66 percent, while median time in pull request review rose by several hundred percent against their prior year's dataset, and the share of pull requests merging with no review at all increased by roughly a third. Those numbers should be treated as directional rather than precise. The direction is the point: throughput improved, and the assurance step either queued or was skipped.

A platform lead at a mid-sized Swiss institution described their review backlog through 2025 as the first honest measurement anyone had produced of the tooling rollout, and noted that it had appeared in no version of the business case. That is a familiar shape. It is the same shape as every capacity problem that arrives through a door nobody was watching.

A short institutional memory

It is worth remembering that this is not the first time the industry has reduced the cost of producing software artifacts and expected delivery economics to follow.

CASE tools in the late 1980s promised generated code from formal models. Fourth-generation languages promised applications from specifications. Model Driven Architecture, a decade later, promised that the model would become the system. Offshore delivery, running in parallel through the 2000s, reduced the unit cost of an engineering hour by a factor that dwarfs anything currently claimed for AI assistance. Each of these did something real. None of them produced the delivery transformation that was underwritten at the outset, and in most cases the cost migrated into integration, coordination, and verification — the parts of the work that resist being made cheaper because they are about establishing agreement rather than producing output.

Brooks made the distinction in 1986 and it has aged well. Accidental complexity is the labour of expressing a solution in a machine-readable form. Essential complexity is the labour of working out what the solution should be and establishing that it is correct. Current tooling is a genuine and substantial attack on the accidental part. It is a much weaker attack on the essential part, and in regulated delivery the essential part is where most of the cost lives.

I do not think this makes the current wave equivalent to the previous ones. The capability is different in kind, and the adoption curve is steeper than anything comparable. But the structural lesson holds: when you make one input dramatically cheaper without changing the system it feeds, the system's existing weaknesses stop being tolerable.

The measurement problem is getting worse, not better

This is the part of the picture that has changed most since 2024, and it deserves more attention than it gets.

In July 2025, METR published a randomised controlled trial in which sixteen experienced open-source developers worked through 246 real tasks in repositories they knew well, with AI access randomised task by task. The measured result was that tasks took around 19 percent longer with AI available. The participants had forecast a 24 percent speedup beforehand and, having completed the work, still estimated a 20 percent speedup afterwards. The confidence interval was wide and METR were explicit about the limits of the setting. The finding worth carrying forward was never the headline number. It was the roughly forty-point gap between what was measured and what the people doing the work believed had happened.

What happened next matters more. METR attempted to repeat the study with a larger and more diverse cohort through late 2025, and in February 2026 published a note explaining that the resulting data could not be interpreted reliably. Developers were declining to participate because they did not want to work without AI. Of those who did participate, between 30 and 50 percent reported withholding specific tasks from the study because they did not want those tasks assigned to the no-AI arm. Some developers were running multiple agents concurrently, which made self-reported task time close to meaningless. METR's own reading is that developers are probably faster now than in early 2025, that their raw numbers show some evidence of speedup, and that selection effects are severe enough that the estimate should be treated as a lower bound of unknown tightness. They are redesigning the study.

Read that as a risk professional rather than as a research result. The control arm is disappearing. The counterfactual — what this work would have cost without the tool — is becoming unobservable, and it is becoming unobservable precisely because adoption has succeeded. Within two or three years, an institution that wants to know what its AI-assisted delivery actually bought it may find that no clean comparison is available anywhere, inside or outside the firm.

That leaves surveys, which come with their own well-documented problems. METR's own survey work in early 2026, covering 349 technical workers, found a median self-reported speed gain of around 3x and a median self-reported value gain of between 1.4x and 2x. The distinction between the two is the most useful thing to come out of that work. Speed measures how much faster the tasks you did went. Value measures how much more your output was worth. They diverge when the tool changes which tasks you choose to do, and cheap tasks that were previously not worth prioritising start getting done. Some proportion of every reported productivity gain is substitution into work that is now cheap rather than work that is now valuable.

Two details from that survey are worth holding onto. The researchers' own staff reported the lowest gains of any subgroup, which the authors themselves suggest may reflect familiarity with the perception gap. And on qualitative review of the highest self-reported multipliers, the authors could not find corresponding evidence in the participants' visible output.

None of this says the tools do not work. It says that anyone presenting a confident productivity number to a board in 2026 should be asked how it was derived, and that "we surveyed the engineers" is not a sufficient answer.

Two curves that have come apart

The clearest signal in the current evidence is not about speed at all.

Veracode has been running the same longitudinal exercise for over a year: put a large number of models through security-sensitive coding tasks where the requested functionality can be implemented securely or insecurely, then run static analysis over what comes out. Across more than 150 models evaluated, roughly 55 percent of generation tasks produce code without a known security flaw. Their spring 2026 update found that figure essentially unchanged over the testing period — during which coding capability on standard benchmarks improved considerably. Models built for extended reasoning did better, in the region of 70 to 72 percent, which is a real improvement and still means that close to a third of output carries a known defect.

Their explanation is structural rather than incidental, and I find it persuasive. Models learn secure and insecure patterns from the same public corpus, and that corpus is heavily weighted toward twenty years of accumulated bad practice. String-concatenated SQL is far better represented than parameterised queries. Nothing about improving a model's reasoning on a coding benchmark causes it to prefer the less common pattern.

Set that alongside Veracode's broader application security data for 2026, drawn from roughly 1.6 million applications: security debt now affects around 82 percent of organisations, up from 74 percent, and high-risk findings have risen from 8.3 to 11.3 percent of the total. Remediation capacity has not moved.

So there are two curves. One is capability, which is rising steeply and is the curve everyone is watching. The other is assurance — the proportion of output that can be shipped without additional verification work — which is close to flat. Everything difficult about the current moment follows from the distance between them. Faster generation against a flat assurance rate does not produce faster delivery. It produces more verification work per unit of time, arriving in a function that was sized against the old rate.

This also explains a finding that confused people when it first appeared. DORA's 2024 research — the DevOps research programme, not the Digital Operational Resilience Act, a collision this industry will apparently live with indefinitely — found that a 25 percent increase in AI adoption was associated with a 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability. By 2025, across nearly 5,000 respondents, the throughput penalty had reversed while the instability remained. The delivery system adapted to the volume. It has not yet adapted to the defect rate. DORA's own framing is that AI is an amplifier, magnifying whatever the organisation already was, and their subsequent ROI work models value realisation as a J-curve with a genuine productivity dip in it. That is a more honest shape than anything in the vendor material.

The evidence obligation did not move

For a supervised institution, none of the above is primarily an engineering problem. It is a problem about what you can produce when asked.

FINMA set out its expectations in Guidance 08/2024 in December 2024, and the substance is unglamorous and entirely clear. Institutions are expected to maintain a comprehensive AI inventory, apply consistent risk classification, and establish named accountability across the lifecycle. For material applications, documentation is expected to cover the purpose of the application, the selection and preparation of data, the choice of model, performance measures, assumptions, limitations, tests and controls, and fallback arrangements. FINMA's survey of around 400 institutions, run over the turn of that year, found roughly half already using AI in some form and a further quarter intending to within three years.

Read that documentation list again with the delivery evidence in mind. Every item on it is a verification artifact. Not one of them is produced faster by generating code faster. Several of them get materially harder when the volume of change rises and the review step queues.

Institutions with a model risk function will recognise the shape of the problem, and will also recognise where the analogy breaks. Supervisory model risk management, in the lineage running from SR 11-7 onward, was built around models with stable specifications, defined input domains, periodic revalidation, and an owner who could be identified without ambiguity. A probabilistic component embedded in a delivery pipeline, updated by a third party on their schedule, invoked through prompts that change weekly, satisfies almost none of those assumptions. A head of model validation at a European bank put it to me in terms of intake: the queue had not changed shape, only length, and the function had been staffed against the old shape. That is optimistic. The shape has changed too.

The same applies to data lineage. Institutions that did the BCBS 239 work know what it costs to demonstrate where a number came from. Retrieval-augmented systems reintroduce that question at a point in the architecture where most firms have no lineage discipline at all, and they reintroduce it non-deterministically.

On timing, the European position shifted in June 2026 and has been widely misread. The Digital Omnibus on AI was adopted by Parliament on 16 June and by Council on 29 June, deferring high-risk obligations for standalone Annex III systems — which include creditworthiness assessment — from 2 August 2026 to 2 December 2027, and for AI embedded in Annex I regulated products to 2 August 2028. Marking obligations under Article 50(2) move to 2 December 2026. General-purpose model obligations under Articles 51 to 55, applicable since August 2025, are untouched.

Sixteen months of calendar relief is real and useful. It is not scope relief. Inventory, classification, conformity assessment and human oversight design do not become easier for having been postponed, and an institution that treats the deferral as permission to pause will arrive at December 2027 with the same work and less time. The firms I would expect to be comfortable are the ones spending the deferral on the parts that take longest to build: the inventory that is actually complete, and the evidence pipeline that produces documentation as a by-product of delivery rather than as a project.

Agents force the identity question early

The agentic shift makes one previously theoretical question operational, and it is worth stating plainly because it is easy to defer.

NIST's work through 2026, and the associated NCCoE material, frames the current state accurately: agents in enterprise environments are commonly deployed as generic service accounts, without dedicated identity, authorisation, or accountability controls. Singapore's IMDA framework, published in January 2026, went the other way and requires that each agent carry a verifiable identity together with an audit trail recording which agent acted under whose authorisation. That is the more demanding position, and I would expect it to become the general one.

The architectural consequence is not subtle. If an agent acts through a shared service account, there is no answer to the question of who authorised a given action, and no answer means no defensible control. Survey evidence on this is thin and vendor-flavoured, but the direction is consistent: a minority of organisations currently monitor AI traffic end-to-end across prompts, tool calls and outputs, and a small minority monitor agent-to-agent interaction at all.

The useful design constraint is that an audit trail for an agent must answer four things rather than one. Who requested the outcome. Which agent acted, under whose delegated authority. What context informed the action. What changed as a result. Conventional access logging answers the last of those and gestures at the second. Retiring an agent has a matching problem, because removing the primary account leaves tokens, delegated credentials, scheduled jobs and sub-agent permissions in place.

None of this is exotic engineering. It is identity and access management applied to a principal class that most IAM estates were not designed for, and it is considerably easier to establish before autonomy is granted than after.

What I would change in the architecture

I am wary of prescriptive lists in this area, because the honest answer is that the practices are still forming and anyone claiming otherwise is selling something. Four things, though, seem robust enough to commit to.

Evaluation should be treated as a deployable asset with an owner, a version, and a change history, rather than as a testing activity. The golden datasets, the adversarial cases, the retrieval relevance checks and the policy compliance suites are the artifacts that make a behavioural claim defensible, and they need the same lifecycle discipline as the systems they assess. In most estates today they live in notebooks.

Provenance should be attached at the point of change, not reconstructed afterwards. Knowing which changes were AI-assisted, at what level of autonomy, and who reviewed them, is the difference between a defensible position and an argument. It is cheap to capture at commit time and expensive to establish later.

Review capacity should be planned as a constraint rather than discovered as one. If generation rises and the assurance rate is flat, the review function needs either more capacity or better tooling before the volume arrives, and the evidence suggests the alternative is that review quietly stops happening.

Agent identity should gate production access rather than follow it. An agent without a uniquely attributable identity and a named accountable owner should not reach an environment where it can act on anything that matters.

Closing

The most durable observation I can make about the last two years is that the industry priced the wrong thing.

Producing a working artifact was treated as the expensive step, because it was the visible one and the one with a headcount attached. In a supervised institution it was never the expensive step. The expensive step was establishing that the artifact does what it is supposed to do, that the data behind it is what it claims to be, that someone specific is answerable for it, and that all of this can be shown to a third party months later without a special project. That work is not made cheaper by generating code faster. On the current evidence, some of it is made more expensive.

This is not an argument against adoption, and I would not want it read as one. The capability is real, the gains in the accidental part of the work are substantial, and institutions that opt out will not be rewarded for their caution. It is an argument about where the returns actually come from. They come from closing the distance between the two curves — from raising the proportion of output that can be shipped without additional verification, and from building the evidence pipeline that makes verification a by-product of delivery rather than a separate cost centre attached to it.

The organisations I would bet on are not the ones with the highest AI code share. They are the ones that can still answer, precisely and quickly, how a given decision was reached and who is accountable for it. That was the harder problem before any of this started. It remains the harder problem, and it is now arriving faster.


Part of the SDOP series on enterprise AI adoption, AI governance, software engineering, platform architecture, and continuous modernization.

References

  1. METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — Randomised controlled trial of experienced open-source developers, including measured slowdown and participant speed estimates.
  2. METR: We Are Changing Our Developer Productivity Experiment Design — Follow-up explaining participation, task-selection, concurrent-agent, and time-measurement limitations.
  3. METR: Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity — Survey evidence separating self-reported speed gains from self-reported value gains.
  4. DORA: Accelerate State of DevOps Report 2024 — Research on the relationship between AI adoption, delivery throughput, and delivery stability.
  5. DORA: State of AI-assisted Software Development 2025 — DORA research framing AI as an amplifier of the surrounding delivery system.
  6. DORA: ROI of AI-assisted Software Development — Research material on value realisation and the productivity J-curve.
  7. Veracode: Spring 2026 GenAI Code Security Update — Longitudinal testing of code-generation capability and security pass rates.
  8. Veracode: AI Coding Tools Are Creating a Security Gap — Analysis of security flaws in AI-generated code and the persistence of insecure patterns.
  9. Faros AI: Key Takeaways from the DORA Report 2025 — Commercial delivery telemetry discussed as directional evidence, with the vendor context made explicit in the article.
  10. FINMA Guidance 08/2024: Governance and risk management when using AI — Swiss supervisory expectations for AI inventory, classification, accountability, documentation, testing, and fallback arrangements.
  11. FINMA: Survey on AI use at Swiss financial institutions — Survey evidence on adoption and planned use of AI across Swiss financial institutions.
  12. Freshfields: The final Digital Omnibus on AI — Legal analysis of the June 2026 changes to the EU AI Act implementation timetable.
  13. Cloud Security Alliance: AI Agent Governance Framework Gap — Research note on identity, accountability, authorisation, and governance gaps for enterprise AI agents.
  14. NIST: AI Agent Standards Initiative — Standards initiative addressing secure, interoperable, and accountable AI agents.
  15. IMDA: Model AI Governance Framework for Agentic AI — Governance framework for accountable agent deployment, including identity and auditability.
  16. Frederick P. Brooks Jr.: No Silver Bullet—Essence and Accidents of Software Engineering — The classic distinction between essential and accidental complexity in software engineering.

Author

is a senior Data and AI executive based in Bülach (ZH), Switzerland, focused on data strategy, enterprise architecture, AI governance, hybrid cloud, and regulated delivery.

Data & AI Leadership Signal Brief

A concise executive note on data strategy, AI governance, platform transformation, and regulated delivery. No generic AI-news noise.

Topic preferences

Your email and topic preferences may be stored in Supabase and synced to Kit for newsletter delivery. You can unsubscribe from future emails.

Related articles

Engineering Management··16 min read

From Zero to Scale: Seven Principles for Building a High-Performing Data & AI Organisation

Seven practical principles for building and scaling a Data & AI organisation: start with the business mandate, hire for judgement, develop leaders early, make accountability explicit, use governance to accelerate delivery, design for multicultural work, and measure organisational value.