When an AI assistant gives a wrong answer, the first instinct is to blame the AI. More often the data was wrong first: a load that ran late, a region that never arrived, a customer counted twice. An experienced analyst would have noticed that the number looked odd and checked before sending it. An assistant does not hesitate. It answers from whatever it finds, in the same confident tone as always.
The checks that good analysts run by instinct therefore have to become automatic, and they have to run before the assistant reads the data, not after a leader questions the answer. Five checks cover most of what goes wrong.
1. Completeness: did everything arrive?
Check that every expected source, entity, region and day is present, and that the fields answers depend on, such as customer, product, date and amount, are filled in. Compare record counts with the source and with the usual range for that day. A partial load is the most dangerous kind, because its totals look plausible. Watch for silence as well as errors: a region with no sales yesterday may be a failed extract, and an assistant will report it as a collapse in demand.
Where it runs: in the pipeline, right after each load. Who is alerted: the data engineering team. If the check fails, the new data is not published to the model the assistant reads.
2. Freshness: is the data current?
Every data set has a promised time by which it should be refreshed. Check the time of the last successful load and the date of the latest transaction, and compare both with that promise. Freshness must also be visible to the assistant: an answer should say how current its data is, and when data is late the assistant should say so or decline questions about today.
Where it runs: on a schedule, before the business day starts. Who is alerted: data engineering first, then the data owner if the delay will affect the morning’s decisions.
3. Reconciliation: does it tie to the source of record?
Totals in the analytics platform should tie to the system of record: revenue to the general ledger, headcount to the HR system, units shipped to the ERP. Agree the tolerance with the owner, run the comparison after each load and again at period close, and do not certify a period until it ties.
Where it runs: after each load and at every close. Who is alerted: the business data owner, usually in finance or operations, with the analytics team investigating.
4. Duplicates and orphans: does every record count once, in the right place?
Duplicate customers, invoices or shipments inflate totals. Orphan records, such as a sale with a product code missing from the product list, fall out of every breakdown, so the regions no longer add up to the total. AI assistants are especially exposed, because they slice data in combinations no report designer anticipated.
Where it runs: when the model is built or refreshed, by checking unique keys and the relationships between tables. Who is alerted: data engineering, and the owner of the master data that needs correcting. Fix duplicates at the source where possible; filtering them out in the model hides the problem from everyone who still works in the source system.
5. Definition drift: does each measure still mean what it says?
This is the quietest failure. A source system adds a new order type or status code, a product moves to another category or someone edits a measure, and net revenue silently changes meaning. Nothing breaks, so nothing raises an alert. Watch for new values in the fields that definitions depend on, compare certified measures with a reference set of known answers and require approval for any change to a certified measure.
Where it runs: on every change to the model and whenever new codes appear. Who is alerted: the measure owner in the business and the analytics lead. When a definition changes on purpose, record the date and the reason, and tell users: an assistant comparing this year with last year needs to know that the rule changed in between.
Where the checks run and who is alerted
Two principles matter more than the tools. First, checks run before publication: a failure stops the data from reaching the model, or marks it as not certified, rather than letting the assistant answer from it. Second, every check has a named owner and an alert route. If only the analytics team hears about problems, problems wait.
Start with the tables behind the questions leaders ask most often. Set thresholds from recent history rather than guesswork, and tune them in the first weeks so that alerts stay rare enough to be acted on. A check that fires every day is soon ignored.
Publish the status of each check where the assistant can read it, so it can add a caveat or decline. Keep the history as well. A record of checks passed day after day is the evidence that lets leaders rely on the answers, and it shows where the data pipelines need investment.
The bottom line
None of these checks is new. What AI changes is the cost of skipping them: a wrong number no longer waits in a report for an analyst to spot it; it goes straight to a leader in a fluent sentence. Organizations that build the five checks into their pipelines, with owners and alerts, give their assistants data worth reading.