Most AI programs in analytics report activity: accounts created, questions asked, prompts written. Activity shows that people tried the tools. It does not show whether answers arrive faster, whether the monthly pack became cheaper to produce, or whether decisions improved. Leaders who are asked to fund the next phase need those answers, in the same plain terms they use for any other investment.
Five measures cover most of the value. Each is simple to define, and each needs a baseline taken before the assistant goes live, because without the “before”, the “after” is only an opinion. Together they also show what to fix: an assistant that is accurate but unused needs different work from one that is popular but often wrong.
Time to answer
Measure the time from a question being asked to an answer the person trusts enough to act on. Before launch, that path usually runs through a request to an analyst, a queue, a few clarifying messages and a spreadsheet. Afterwards, routine questions should be answered in minutes, and the harder ones should still reach an analyst, now with a head start.
Track the median rather than the average, and include the questions that still end up with an analyst. A fast assistant that sends many questions back to the queue has not reduced the time to answer. It has added a step. The simplest method is to record two moments for a sample of requests: when the question arrived, and when the person who asked it confirmed the answer.
Report production hours removed
Much of an analytics team’s month goes into producing the same packs: refreshing extracts, rebuilding slides, writing commentary and reconciling one version against another. AI can draft the commentary and answer the follow-up questions that used to require a new report. The value lies in the hours that actually stop.
Count hours removed, not hours saved on paper. A pack that is still produced in full, with some AI help on the side, has removed nothing. The honest measure lists the packs retired, the steps no longer performed and the hours moved to analysis. Ask the people who build the packs, not only their managers; they know which steps disappeared and which ones quietly moved elsewhere.
Adoption by role
A single usage total hides the story. Executives may ask a few high-value questions a week while finance analysts ask far more, and one number says little about either group. Measure weekly active use by role, and how many people return after their first month. Repeat use matters more than first use: a role that comes back every week has found value, while a role that tries once and leaves has not.
Low adoption in a role is usually a design signal rather than a training problem. It means the assistant cannot answer the questions that role asks, because the data, the measures or the vocabulary are missing. The questions people tried and then abandoned show what to add next. Close those gaps first, then measure again.
Accuracy against certified numbers
Value counts only if the answers are right. Keep a set of real business questions with their certified answers, agreed with the finance or data owner, and run it after every change to the data, the semantic model or the assistant’s instructions. Report the share answered correctly, and set a floor below which a new version is not released.
Measure accuracy against the same certified figures that appear in the monthly pack. If the assistant and the pack disagree, leaders will trust neither, and every other measure on this list stops mattering. When a question fails, record why: a missing measure, an ambiguous name or a gap in the data. Those causes tell the data team what to fix first.
Decisions made faster
This is the measure that matters most and the one most often skipped. Choose a handful of recurring decisions where speed has value: replenishing stock, approving a price exception, signing off a forecast, moving budget between campaigns. Record how long each takes today, from the moment the signal arrives to the moment the decision is made, and measure the same cycle after launch. Keep the definitions strict, so that the signal, the decision and the decision maker are the same before and after.
Faster decisions are where saved hours and better accuracy turn into financial results. They also make the clearest story for a board: the same decision, of the same quality, made days earlier. A few such decisions, timed honestly, are more convincing than any usage chart.
Take the baseline before launch
The weeks before launch are the only time a clean baseline can be taken. Use them to:
- sample real requests to the analytics team, and record when each was asked and when it was answered;
- log the hours spent on each recurring pack, step by step;
- write the test set of questions with their certified answers;
- time the recurring decisions you selected, from signal to decision.
At the same time, agree on a target for each measure with the business sponsor, so that success is defined before anyone sees the results.
Report value on one page
Put the five measures on one page, every month, next to their baseline and target, with a short note on what changed. Then treat it like any other investment review: scale the use cases that move the measures, fix the ones that come close, and stop the ones that cannot. That page is also the strongest case for the next phase, because it shows value in numbers the organization already trusts. It is one of the first things we set up in our data and AI strategy work.