The demo went well. Everyone remembers the demo going well. The tool summarised, generated, forecast, whatever it was bought to do, and for a fortnight the Slack channel was busy with screenshots. Then the pilot ended, the renewal conversation started, and nobody could quite say what had changed. The output wasn’t wrong. It was fine. Fine is the most expensive word in an AI budget.
Somewhere between 70 and 85 per cent of AI projects fail to deliver what was promised, depending on whose research you read. That number gets quoted constantly, usually as an argument for more careful technology selection. Sit with the failures for a while, though, and a different pattern emerges. The tool rarely underperforms the demo. It performs exactly as advertised. What underperforms is the instruction it was given.
Generic brief, generic output
Hire the best freelance copywriter in the world and brief them with “make it better”, and you will get something polished, professional and generic. Not because they lack talent. Because you gave them nothing to work with. AI behaves the same way. The quality of the output is decided by the quality of the context, and most businesses hand the machine a brief they would be embarrassed to give a junior: vague objectives, numbers nobody fully trusts, no definition of what good looks like.
I built an AI trading intelligence platform for a consumer brand last year. The hardest part wasn’t the AI. It was the commercial work that came before it: agreeing which numbers the business actually trusted, what a good week looked like, which decisions the tool existed to inform. Once that was settled, the AI part moved quickly, and the platform now sits inside the weekly trading rhythm rather than beside it. Run the sequence the other way and you get something different: a very fast way of producing reports nobody believes.
The next model won’t save you
The counterargument is that the technology is still maturing, and the next generation of models will be capable enough to work things out for themselves. The models will certainly improve. But no model can retrieve a decision your board hasn’t made. If the business cannot articulate what it is optimising for, the machine cannot either. It fills the vacuum with the average of everyone else’s answer, which is precisely what generic output is. The sound of a vacuum being filled.
A pilot is not really a test of the technology. It is a test of the brief, and the brief is your commercial clarity written down.
This is why the brands winning with AI didn’t start with AI. They started with clarity about who they serve, what they are optimising for, and what good looks like. Then the same tools everyone else bought started compounding, because they had something to compound. The gap you see between those brands and the rest isn’t a technology gap. It never was.
The technology gets auditioned. The brief doesn’t. The 70 per cent are not unlucky; they automated before they clarified. The fix is to run the sequence the other way round, and it costs a great deal less than the pilot did.