Every AI development company can show you a working demo. That is not a useful signal, because a demo is the easiest artefact in this discipline to produce: one path, one curated dataset, one person driving who knows exactly which question to ask, and no consequence attached to being wrong.
The question that decides whether a programme succeeds is what happens on the first Tuesday after launch, when the system is running unattended against data nobody cleaned, for a user who does not know what it can and cannot do. Almost no evaluation tests for that, because it is harder to stage and much less satisfying to watch.
If you are choosing between firms, here is what actually separates them.
Ask what they have running, not what they have built
"Built" is a weak verb. It includes projects that shipped and were quietly switched off, pilots that never left the innovation team, and prototypes that impressed a steering committee and then stopped.
Ask instead: what is in production today, who uses it, and how long has it been running without the original team attached to it? A system that has survived two years of staff turnover, data drift and changing requirements is evidence of something a demo cannot demonstrate. A company that cannot answer this quickly is telling you something.
The follow-up question is better still: what have you had to fix since launch? Anyone with real production experience has a list and will talk about it without defensiveness. Anyone who says nothing has gone wrong has not been in production.
Ask what they refused to build
This is the most revealing question on the list, and the one most likely to produce an uncomfortable pause.
A firm whose answer is "nothing" is either very new or selling whatever you are buying. AI is the wrong answer to a great many problems: where the process is stable and rules express it perfectly, where the data needed does not exist, where the cost of being wrong is high and the accuracy achievable is not high enough, and where the real problem is that three systems disagree about what a customer is.
A partner worth having has talked at least one client out of an AI project, and can tell you why. That is not modesty. It is the same judgement that decides where AI genuinely pays, and you want it pointed at your problem.
Make them explain a wrong answer
Ask to see the system fail. Not a staged edge case — ask what it does when it does not know.
There are only a few honest behaviours. It can say it does not know. It can return what it found and let a person judge. It can escalate to a human. What it must not do is produce a confident, fluent answer with no basis, which is the default behaviour of a language model asked a question outside what it can support.
Then ask how you would find out afterwards. If a decision the system made turns out to be wrong six weeks later, can you reconstruct what it saw, what it did and why? In a regulated business that is not a nice-to-have. An answer you cannot audit cannot be defended, and a system whose reasoning is not recorded is one you will eventually have to switch off rather than explain.
Ask where the data goes
For a great many organisations this is the question that decides the architecture, and it should be asked early rather than discovered late.
If your contracts, patient records or transaction history cannot leave your infrastructure, say so on the first call. Open-weight models have made self-hosting genuinely practical, and the capability gap against hosted frontier models is now small enough that it rarely decides the outcome. We have run a full agentic platform for a hospital network entirely on that network's own GPUs, with no patient data leaving the premises, and the clinical results held up — specialist misrouting fell by 59%.
What matters is that the firm treats this as an architectural choice rather than an obstacle. If the answer is that self-hosting is impossible, what they mean is that they have not done it.
Ask who is going to do the work
Not who is in the room. Who writes the code.
The gap between the people at the pitch and the people on the project is the oldest complaint in this industry and it has not gone away. Ask for the names, the seniority and the time allocation of the engineers who will actually be assigned, and ask what happens when one of them leaves.
Ask, too, what the firm did before AI. A company that has been delivering software since 2011 has fifteen years of experience with the unglamorous parts — integration, data quality, migration, operations, the things that decide whether an AI feature survives — and those parts are most of the work. A firm founded in 2023 has enthusiasm and a model provider.
The uncomfortable summary
Almost everything that distinguishes AI vendors is invisible in a demo and obvious after eighteen months. You are not buying a model. Models are a commodity and they change every few months. You are buying the engineering judgement that decides what to build, the discipline to say when not to, and the willingness to still be there when the system does something surprising.
Ask the questions above and you will learn more in an hour than in any proof of concept.
Thinking about this for your own business?
We have been building and running enterprise systems since 2011. Talk to a solutions lead about where agents pay off first.
Talk to a solutions lead
