AI assistants that answer business questions in plain language are moving from pilots into the hands of executives. The appeal is speed: a question asked and answered in the same minute, without waiting for an analyst. But speed is not what leaders remember. They remember the day the assistant gave a revenue figure that did not match the board pack.

One wrong number can undo months of adoption work, because it puts every other answer in doubt. The remedy is not to slow the assistant down. It is to prove, before it reaches leaders and after every change, that its answers match the certified figures.

Start from the certified number

A certified number is a figure the organization has agreed to stand behind: reconciled to the source, built on a definition someone owns and published in the reports leaders already use. That number, not the assistant’s own arithmetic, is the reference. The assistant is right only when it reproduces the certified figure for the same definition, period and scope.

When no certified figure exists for a question, the assistant should say so, and say what it used instead. An honest “this is not a certified measure” protects trust far better than a confident guess.

Build a test set from real questions

The test set is a list of questions with their expected answers; few things in an AI analytics project are worth more. Build it from what leaders actually ask: questions sent to analysts, raised in review meetings or typed into the assistant during the pilot. Each test case records four things.

  • The question, worded the way people ask it, with a few rephrasings and the natural follow-up.
  • The expected answer, taken from the certified report, with its period, scope and unit.
  • The source the answer must cite, so that a correct number from the wrong place still fails.
  • The owner who confirms the expected answer and updates it when a definition changes.

Include questions the assistant should decline: periods that are not yet closed, data outside its scope and information the person asking is not allowed to see. Declining correctly counts as a pass. A few dozen well-chosen questions per topic make a better start than hundreds of generic ones.

Exhibit: how an AI agent handles routine finance questions, and its accuracy on a test set by question type, against a release bar.

Run it on every change

A test set used once, at launch, is a demonstration. Its value comes from running automatically whenever something changes: the semantic model, a measure, the assistant’s instructions, the underlying AI model, even the name of a column. Changes that look harmless are often the ones that shift answers, because the assistant reads names and descriptions to decide what a question means.

Set a release bar with the business owner and treat it like any other quality gate. A change that pushes accuracy below the bar does not reach users until it is fixed. Run the full set before each release and a shorter set every day, because the data changes too. Keep every run, so you can show how accuracy has moved over time.

Exhibit: test-set accuracy after each change; the one change that fell below the release bar was stopped before release.

Agree the tolerances first

Not every difference is an error, and arguing about it after a failure wastes time. Agree tolerances with finance before the first run. Counts and transactions should match exactly. Amounts may differ only by rounding, at the precision shown in the certified report. Ratios must use the certified definition, not a recalculation from rounded figures.

Wording may vary; numbers, periods and definitions may not. An answer that quotes gross revenue when the question was about net revenue is wrong, however well it reads. When a question is ambiguous, the assistant should state its assumption or ask, and the test set should check that it did.

When the assistant is wrong

Treat a wrong answer the way you would treat an error in a published report: as a defect with a cause, an owner and a fix. The causes are usually familiar: two measures with similar names, a missing description, a filter the assistant did not apply or a question outside the data it can see.

Fix the cause in the model or the instructions rather than patching the single answer, then add the failing question to the test set so the same mistake cannot return unnoticed. While the fix is under way, narrow the assistant’s scope: route questions on that topic to an analyst and tell users what changed. People forgive a known limit far more readily than a silent error.

Show the source in every answer

Every answer should carry its provenance: the certified report or model it came from, the measure used, the period and filters applied and when the data was last refreshed, with a link to the page where the number can be checked. Answers built on data that is not certified should say so plainly.

Showing the source does more than build confidence. It makes errors visible quickly, because readers notice at once when the period or the scope is not what they asked for. And the assistant should see only what the person asking is allowed to see, so that every answer follows the same access rules as the report it came from.

The bottom line

Speed is the promise of AI in analytics; trust is the condition for using it. Organizations that treat expected answers as a governed asset, owned jointly by finance and the analytics team and run on every change, can extend an assistant topic by topic, with evidence at each step. Those that skip the test set find out from their leaders, which is the most expensive way to learn.