Essay · 2 October 2026
Nobody uses an AI feature they cannot check

Only 25% of workers use AI regularly as part of their job, while 86% of chief executives think their people are ready for it. Both figures come from the same IBM study of more than 2,000 chief executives, published in May. If you have shipped an AI feature and the usage line is flat, the standard explanations are that the model needs work or the users need training. Usually it is neither. The feature is competing with doing the task by hand, and it is losing on the cost of checking.
Every AI output arrives with a second task attached: deciding whether it is right. That decision is work, and somebody has to do it. If that decision takes nearly as long as the task would have taken, nothing has been saved. The effort has moved from doing to verifying, and verifying is the duller half. People stop opening the feature, and nobody files a bug, because there is no bug to file.
Almost right is the expensive kind of wrong
Stack Overflow put this to 49,000 developers. The most cited frustration, named by 66% of them, was "AI solutions that are almost right, but not quite". Another 45% said debugging AI-generated code took more time than they expected. On accuracy, 46% actively distrusted what the tools produced, against 33% who trusted it. Even so, 84% were using the tools or planning to. The figures are in the 2025 developer survey.
Hold those two facts together, because the combination is the whole problem. Adoption high, trust low. That is a large population of people using something and then checking all of it.
Obviously wrong is cheap. It announces itself, you throw it away, you move on. Nearly right is expensive, because it survives a glance and fails in one detail, and the only way to find which detail is to go through the lot. Developers feel this sharply, since code eventually fails out loud. In a dashboard or a report, a nearly right number just sits there looking like a number.
People do not abandon an AI feature because they distrust it. They abandon it because checking it costs more than doing the work.
The design effort goes in the wrong place
In most products the design attention has been spent on the input. The prompt box, the placeholder text, the suggested starting points, the animation while it thinks. That is the easy half, and it is the half that demonstrates well in a sales call.
The expensive half starts the moment an answer appears. Cheap checking has a shape, and it is mostly provenance and reversal. Show where the answer came from at the level of the specific row, document or line, so a reader can spot-check one claim instead of re-deriving all of them. Say what was not looked at, because a confident answer drawn from half the data is the dangerous kind. Let the output be edited in place rather than accepted or rejected whole, so a nearly right answer can be corrected in ten seconds instead of discarded. Make undo complete and obvious, because people take fewer chances in a system they cannot reverse. And when the model is unsure, let that change what the interface does, rather than adding a hedge to the sentence.
None of that is model work. All of it is interface work, and it is the work that decides whether a feature gets opened a second time.
Showing the working
We do not sell AI design as a separate service, because it is not one. It is a context that changes brand, website and product work rather than a fourth thing sitting beside them. The hard part is not the interaction, it is showing the working. An agent acting on someone's behalf, a summary that replaces reading the source, a suggested value in a field a human will put their name to: each of those needs a visible chain back to something a sceptical person can inspect in seconds.
That is the honest test of whether a feature is finished. Not whether it gives a good answer on a good day, but whether a cautious user can satisfy themselves it is right faster than they could have done the job unaided. If they cannot, the usage graph is not being unfair. It is reporting the arithmetic correctly.
It is also why a disclosure, on its own, does very little. Telling someone they are reading machine output is now a legal requirement in a lot of cases, and it is the right thing to do, but a label answers a different question. It tells a reader to be careful. It does not help them be careful quickly.
Nobody owns the verification surface
Where this goes wrong is usually organisational rather than technical. The model belongs to engineering. The feature and the prompt belong to product. The question of how a person knows the answer is right belongs to nobody in particular, so it gets handled late by whoever has capacity, which is how it ends up as a tooltip.
Give that question an owner early, and the next version of the feature probably does not need a better model at all.