Starting a generative AI pilot has never been easier. A small team, a capable model and a few weeks are enough for a convincing demonstration: a sales assistant that drafts account plans, a service agent that answers policy questions, a finance copilot that writes variance commentary. Leadership sees the demonstration, approves the next step, and the pilot then waits in a holding pattern for months.
When pilots stall, the model is rarely the reason. They stall because nobody owns the outcome, nobody can show that the answers are good enough, nobody knows the cost at full volume, nobody is ready to support the tool, and the people meant to use it were never asked to change how they work.
Where pilots stall
Each of these gaps is ordinary, and each has an ordinary fix. Together they explain why so many organizations run pilot after pilot and have nothing in production. The way out is to close the gaps properly for one use case, then reuse what was built for the next ones. The first use case carries most of the effort; the ones that follow inherit it.
Name an owner who is measured on the result
A production use case needs a business owner: a leader whose team will use it and who will be measured on the result. That owner defines what good looks like, accepts the remaining risk and makes room for the change in how the work is done. The owner also decides when the pilot has succeeded, which prevents the common case of a pilot that runs on because nobody has the authority to end it. Technology teams own the platform and the build. They cannot own the outcome.
If no business leader is willing to own a pilot, that is useful information. It usually means the use case solves a problem nobody is paying for, and the kindest decision is to stop it early.
Prove quality before you scale
A demonstration shows that the assistant can be right. Production requires evidence that it is right often enough, on the cases that matter. Build an evaluation set of real cases with the expected answers, agree on the quality bar with the owner, and run the set after every change to the model, the instructions or the data. We describe how to build such a set in our article on getting data ready for AI agents.
Keep evaluating after launch. Review a sample of production answers each week, track the share that passes, and send uncertain or high-risk cases to a person. Evaluation is not a phase that ends; it is how the use case stays trustworthy as the model, the data and the business change around it.
Know the cost per answer
Pilot costs are small, and nobody watches them. At production volume, the cost of every request becomes a line in someone’s budget. Measure the cost per answer, per document or per conversation from the first day of the pilot, and project it at full volume before approving the rollout. In every review, cost should sit next to value, for the same period and in the same currency.
Most of the levers are design choices: use the smallest model that clears the quality bar, send only the context a request needs, reuse results for repeated questions, and set spending alerts for each use case. A use case whose value does not cover its running cost at scale should not reach production, however impressive the demonstration.
Plan support and change before go-live
Users need to know where to report a wrong answer and how quickly it will be corrected. Someone has to own the instructions, the evaluation set and the data connections after the project team moves on. Write that support model before go-live, as you would for any business application, and decide how a new version is released: who approves it, how it is tested and how users learn what changed.
Change is the larger task. People adopt a new way of working when it fits the tools they already use, when their managers expect it and when the old way is retired. Training sessions alone rarely achieve that; redesigning the process around the assistant usually does.
From one use case to a portfolio
The second use case should cost far less than the first, because it reuses what the first one built: identity and access, logging, the evaluation set and its tooling, cost monitoring and the support model. Treat these as a shared platform, not as parts of a single project. A shared platform also prevents the opposite problem, in which each department buys its own assistant, with its own controls, for the same need.
Choose the next use cases on value and feasibility, and sequence them in waves, with data foundations and governance running underneath. A portfolio view also lets leaders compare use cases on the same measures and move investment to the ones that pay back.
Governance that speeds things up
Good governance makes scaling faster, because every team knows the bar before it starts. A short set of gates is enough:
- a named business owner and an agreed measure of value;
- an evaluation set, and a quality bar that has been met;
- a known cost per use, projected at full volume;
- a support model and a plan for the change in how work is done;
- a review of the risks, with people approving any action that has consequences.
Review the portfolio every quarter against these gates and the value delivered. Scale what works, fix what is close and stop what is not. That rhythm, more than any single technology choice, is what turns a collection of pilots into a lasting capability.