Policies and procedures are written to be followed, but they are rarely easy to find. Staff search intranet sites, open long manuals and, in the end, ask the colleague who seems to know. Answers vary, policy owners are interrupted, and nobody can tell which version of a rule people are actually following.

Generative AI offers a better route: ask a question in plain language and get a short answer drawn from the approved documents, with a link to the exact passage. Done well, it saves time on both sides of the question. Done carelessly, it gives confident answers from outdated or restricted material. The difference lies in four pieces of work that have little to do with the language model itself.

The pattern in plain words

The assistant does not know your policies. When a question arrives, it searches an index of approved documents for the passages most relevant to it, and the language model writes an answer using only those passages. Every answer cites its sources: the document, the section and the effective date. If no passage answers the question, the assistant says so and points to the policy owner instead of improvising.

This pattern, often called retrieval-augmented generation, has a useful property: the quality of the answers depends mostly on the quality of the documents and of the search, and the organization controls both. It also makes every answer checkable, because anyone can follow the citation and read the passage for themselves.

Be clear about what the assistant is for. It explains what a policy says and where to find it; it does not grant exceptions, approve requests or give legal advice. Those boundaries belong in its instructions and in the answers it gives.

Prepare the documents first

Most wrong answers trace back to the documents rather than the model: two versions of the same procedure, a regional variation that is not labelled as one, a scanned manual with no searchable text, a table split across pages. Preparing the collection is the largest part of the work:

  • One approved source per topic. Retire copies, drafts and superseded versions from the index, even if they stay in the archive.
  • Sections that stand on their own. The assistant retrieves sections, not whole manuals, so each heading must make sense without the pages around it.
  • Labels that matter for answers. Owner, effective date, and who the document applies to: role, region and business unit.
  • Readable content. Scanned pages converted to text, tables checked, and abbreviations spelled out at least once.

The effect shows up in testing. When the same questions are asked again after each round of clean-up, more answers come back right, with no change to the model. The work pays back even without AI: staff and auditors both benefit from a policy library with one current version of everything.

Exhibit: answers to the same test questions after each round of document clean-up.

Access by role

The assistant must use only the documents the person asking is allowed to read. The simplest way to guarantee it is to inherit permissions from the document repository rather than build a separate index with looser access. An investigation procedure in human resources, a security runbook or a board policy should never surface in an answer to someone who could not open the original.

Test this explicitly, with people in different roles asking the same questions. An answer that reveals restricted content is a more serious failure than an answer that is wrong, and it should stop the release in the same way. Keep personal information out of the collection altogether unless a use case truly needs it: employee files, case notes and health information belong in their own systems, not in a policy index.

Test before staff rely on it

Build a test set from questions staff actually ask, collected from help desk tickets, emails to policy owners and the questions new employees raise. For each one, record the correct answer and the passage it should cite. Include questions the assistant should decline, questions whose answer depends on region or role, and questions about policies that changed recently, because that is where errors hide.

Score each answer on three counts: whether it is right, whether it cites the right passage, and whether it admits what it does not know. An answer that is correct but cites the wrong source still fails, because staff will follow the link and read the wrong rule.

Policy owners review the results and decide when the assistant is ready. Set the release bar for each type of question, not only for the average: an overall score can pass while the questions that matter most still fail. Run the same test again whenever documents, settings or the model change, and keep the results so that every release can be compared with the last.

Exhibit: correct, sourced answers by type of question against the release bar.

Keep the answers current

A policy assistant is only as current as its index. Tie the index to the publishing process, so that a new or revised policy is searchable the day it takes effect and a retired one disappears the same day. Show the effective date in every answer, so staff can see how recent the source is, and flag answers drawn from a document that is past its review date.

Log the questions the assistant could not answer and the answers staff marked as unhelpful, and send them to the policy owners each month. They are the most direct evidence available of where the documentation is unclear, missing or contradictory. Give each owner a review date and a short report: how often their documents were cited, which questions failed and which answers drew complaints.

Start with one domain

Start with one collection and one audience: human resources policies for managers, IT procedures for the service desk, or operating procedures for frontline staff. Prove the answers there, publish the test results, then add the next collection. Expect the first weeks to surface more problems in the documents than in the technology, and plan the policy owners’ time accordingly. Our work on generative AI assistants follows the same order: documents, access, testing, then scale.