Finance teams adopted generative AI the way they once adopted spreadsheets: one person at a time. An analyst finds a phrasing that produces a good variance summary, saves it in a notes file and passes it to a colleague, who changes a few words. Months later the team has many versions of the same request, and the CFO receives two different explanations of the same margin movement from two people who both used AI.

The problem is rarely the model. It is that the request, the data it reads and the expected answer were never written down. A prompt library fixes that: a shared, tested set of prompts for the questions finance answers every month, each with a named owner.

Why ad-hoc prompts drift

Language models respond to wording and context. Ask for “revenue versus plan” one day and “sales against budget” the next, and the assistant may choose different definitions, periods and rounding. Paste in an extract that is a day older than the pack, or let the vendor update the underlying model, and the same question returns a different answer with the same confidence.

None of this is visible to the reader. The answer arrives well written and confidently worded, with no sign of the choices made along the way. In finance, an answer that cannot be reproduced cannot be reviewed, and an answer that cannot be reviewed should not reach a leader.

What a library entry contains

An entry is more than the text of a prompt. It is a small, documented product with five parts.

  • The question it answers, in the words people actually use, with the usual variants: why margin moved, the margin bridge, what drove gross margin.
  • The data it reads: the certified model, the measures, the period and the access rules. A library prompt never depends on whatever happens to be pasted into the chat.
  • The expected format: length, order, units, sign conventions and rounding, so every answer looks the same and can go into a pack without editing.
  • Test questions with expected answers, checked against the certified figures and kept with the entry.
  • An owner and a version. One person is accountable for the entry, approves changes and decides when it is retired.

Writing an entry is a small piece of work, not a project. Most of the effort goes into the test questions and their expected answers, which is also where most of the value lies: they turn a good phrasing into something the team can rely on.

Exhibit: one entry in a finance prompt library, with its data, format, owner and test record.

Start with the questions finance answers every month

The best first entries are the recurring questions: revenue and margin against plan, operating expenses by cost centre, cash and working capital, headcount and labour cost. They are asked at every close, their correct answers are known, and their definitions already exist in the reporting model. A handful of well-built entries covers much of what leaders ask, and the questions that fall outside the library show where to build next.

Resist the urge to catalogue every prompt anyone has ever written. A library is valuable because each entry has been tested and has an owner; a long list of untested phrasings is just a larger notes file.

Exhibit: how often repeated runs return the certified answer, ad-hoc prompts against library prompts.

Test prompts the way you test reports

A prompt is code written in plain language, and it deserves the same discipline. Each entry carries its test questions, and the whole library is re-run whenever a prompt, the data model or the AI model changes. An entry that fails goes back to its owner before anyone relies on it. Over time, the test results become the evidence that the library can be trusted, and the record that auditors and controllers will ask to see.

The tests should check more than the headline number. A good answer uses the right period, the right comparison and the right sign convention, and says so; an answer with the correct total and the wrong period is still wrong.

Make every answer reviewable

Repeatable is not enough; an answer must also show its work. Library prompts should require the assistant to state the measures, filters and period it used and when the data was last refreshed, so a reviewer can trace every number to the certified model. Answers that leave the finance team, such as commentary for the board pack, are reviewed by a person and stored with the version of the prompt that produced them.

A fixed order helps the reviewer: the conclusion first, then the figures behind it, then the source and its refresh time. When every answer follows the same order, a missing piece is obvious at a glance.

Put the library where people already work

An entry that has to be copied from a document into a chat window will be edited on the way, and the edits will not be tested. Make the library available inside the tools the team already uses, such as the assistant in the reporting tool or an agent built for the finance team in Copilot Studio, so that people choose a tested entry rather than retype it. The easier the tested path, the less ad-hoc prompting survives.

Run the library like a product

Someone has to own the library, usually the FP&A or reporting lead, with a small group of contributors. New entries come from questions analysts answer again and again. A change to an entry follows a light version of change control: a proposal, a test run, the owner’s approval and a new version number. Retired entries are kept, with the reason, so that past answers remain explainable. Report on the library as you would on any product: which entries are used most, which fail their tests, and which questions arrive that it cannot yet answer. Access follows the same rules as the data, so a prompt cannot reveal what its user is not allowed to see; the rules that secure the data apply to every AI request made through the library.

The payoff is quieter than a new tool launch, and more durable. The same question gets the same answer, whoever asks it and whenever, and when an answer changes, there is a reason on record.