The case for AI coding tools is easy to make in a demonstration. Give an agent a task, watch it produce code in seconds, and imagine what a whole engineering team could achieve with that speed. The harder conversation starts when the CFO asks the CTO what the investment has delivered
MIT’s’s “State of AI in Business 2025” report showed that a staggering 95% of enterprise AI pilots delivered zero return on investment.
More code is being written, but more code does not automatically mean more valuable software. Some of the apparent gain disappears in review, testing and rework. A feature may work in isolation yet conflict with the architecture around it or miss a business rule that was never included in the prompt. Engineers then spend the time saved on the first draft making the result usable.
We call the gap between what a business intended to build, what its teams produced and what the finished software actually delivers engineering drift.It can exist without a failed project or a dramatic incident. Often, it shows up as a growing number of small corrections, inconsistent implementations, and handoffs where the original requirement is interpreted differently.
This is the question DesignVerse is taking to San Francisco Tech Week. At a session for CTOs and senior engineering leaders on 7 October, we will examine where AI creates enterprise value, where it gets lost and how teams can measure the difference. The goal is a working discussion grounded in what leaders are seeing across their own organisations.
The problem grows with scale. A small team may have an experienced engineer who knows why a particular integration works the way it does and can spot a deviation immediately. Across dozens of teams, that knowledge is dispersed. It lives in documentation, repositories, business procedures and the judgment of people who have maintained a system for years. AI can make the initial implementation faster while multiplying the number of places where that knowledge is missed.
Consider a request to update a customer-facing application. An agent might produce a clean solution that passes its immediate tests. But if it uses a different integration pattern from the rest of the estate, the senior reviewer must explain and correct it. If it misses a rule buried in an operational workflow, QA may find the problem later. If neither catches it, the difference can emerge when the software reaches production. Each stage adds a cost that a measure of coding speed alone will overlook.
There are practical ways to see that cost. Look at recent features and ask how much review time went into correcting architecture or pattern mismatches. Count the changes that were reopened after they were marked complete. Compare how different teams implemented the same requirement. Track the time from the initial request to a validated release, rather than the time it took to generate the first pull request. Those measures show whether AI is making delivery more effective or moving work downstream.
Buying a more capable model will help with many tasks. It cannot, by itself, supply a company’s undocumented rules or explain why an established system makes a particular trade-off. That knowledge must be made available to the tools doing the work. It also has to be maintained: a static document pasted into a prompt will become less useful as the software and the business change.
At DesignVerse, we build an enterprise context layer from a customer’s code, architecture, engineering documentation and business information. It sits between the request and the chosen AI model, supplying relevant rules and context so that generated code can fit the organization’s existing architecture and ways of working. The aim is to reduce the corrections that accumulate between a convincing first draft and software the business can actually use.
This matters most where an application has a long history or a mistake has serious consequences. DesignVerse’s work with EUROCONTROL – which governs the busy skies over Europe – began with internal application modernization and has extended towards software used in its operations. In environments like these, the useful question is not simply whether an agent can write code. It is whether the organization can understand, review and maintain what the agent helps to build.
There is a related procurement question. If AI is going to become part of critical software delivery, leaders need to know who controls the information it uses, which models they can choose, and how their own rules govern the output. Control over the workflow is as important to the investment case as control over the data. Without it, an organization can pay for faster code generation while absorbing the hidden cost of drift.
San Francisco Tech Week will be full of examples of what AI can do. Enterprise leaders should be just as curious about what happens after the demonstration. The most valuable coding tool is the one that helps a team carry its intent through review, testing and release. That is where the return on AI investment is won or lost.