It’s the second week of the quarter. Your support lead wants an assistant that drafts replies from the help centre, live on Monday. Legal asks what happens if it sends the wrong thing. Finance asks what it costs by December. You decide on Friday. An enterprise AI governance framework exists so that decision is quick and still holds up later.

It rests on one idea. Every rule you write about an AI system needs a control inside the application that applies the rule, and a record that shows the control ran. If any one of the three is missing, the other two can’t prove the rule was followed.

This guide is the version we run with customers: how much governance a system needs, who does what in the first 90 days, what the controls cost, the three failures we see most and the evidence to keep. It’s an engineering starting point. Legal advice and certification need qualified advisers.

What an enterprise AI governance framework has to do

Take the support assistant. Before it goes live you need to know whether it can read restricted material, whether it’ll invent an answer when no source covers the question, and whether it can send a message without a person looking. Those three questions give governance a scope you can test.

The same shape covers knowledge search, document processing and scheduled reporting. The controls differ. But the chain from a requirement to its enforcement to its evidence stays the same, and the diagram below is that chain for one request.

YesNoApprovedRejected or expiredDefine the permittedtaskCheck identity and dataaccessRun the evaluatedmodel and promptApproval required?Wait for an authorisedreviewerPerform the permittedactionStop and record thereasonRecord outcome andusage
From an AI policy requirement to operational evidence

How much governance is enough

Match the controls to what the system can do and to who’s affected when it’s wrong. A tool that drafts text a person edits doesn’t need the treatment of one that changes what a customer is charged. Two questions place a system: can it act with nobody in the path, and can its output change what a person gets.

Tier What the system does Controls that apply Evidence kept Review
1, assist Drafts or summarises material the user could already read. A person edits and sends. Inventory entry, two named owners, a stop switch, usage per key Owner, purpose, monthly usage Twice a year, and on change
2, act with review Reads restricted sources or proposes an action a person approves: support drafts, purchase orders, leave reminders Tier 1 plus permission-aware retrieval, the access test, an evaluation set with release blockers, approval bound to a version Test results per release, an approval record per action, unresolved charges Every release, and quarterly
3, consequential Affects a person’s money, rights, safety or access, or acts with nobody in the path Tier 2 plus a legal assessment, a second reviewer for changes, a rehearsed rollback, monitoring with alerts The assessment and who signed it, the incident log, monitoring history Every release, monthly, and after any incident
Three governance tiers, each adding controls to the lastThree panels side by side. Tier 1, a system that drafts text a person edits and sends, has three controls: two named owners, a stop switch and usage recorded per key. Tier 2, a system that proposes an action a person approves, keeps the tier 1 controls and adds an access test, an evaluation set and approval bound to a version. Tier 3, a system that changes money or rights or acts alone, keeps both earlier tiers and adds a legal assessment, a second reviewer and a rehearsed rollback. The stack grows from left to right. Conceptual illustration.Tier 1Drafts a personedits and sendsTwo named ownersA stop switchUsage per keyTier 2Proposes an actiona person approvesTier 1 controlsAccess testEvaluation setVersioned approvalTier 3Changes money, rightsor acts aloneTiers 1 and 2Legal assessmentSecond reviewerRehearsed rollback
Three governance tiers, each adding controls to the last

Nothing is removed as a system moves right. A tier 3 system carries every tier 1 and tier 2 control plus the ones that need a lawyer, a second pair of eyes and a rehearsed way back.

The tiers are an engineering view. The law has its own. The EU AI Act sets risk-based rules that depend on the system, its intended purpose and your role as provider or deployer. And the fines in Article 99 run up to EUR 35 million or 7% of worldwide turnover for prohibited practices.

So anything you’d put in tier 3 needs a classification from a qualified adviser, with the date it was made. A questionnaire can show you which information is missing, and it can’t classify a system from a few organisation-level answers. Our EU AI Act readiness quiz does the first job and stops there.

What to record about each system

Start with the AI systems you already use, including the models embedded in software you bought. For each one, write down who owns it, what it’s for, who uses it, which data it reads, which providers it calls and which decisions its output influences.

Keep the record inside a process you already run, such as procurement or deployment review, because an annual spreadsheet misses every change between two reviews. Give each entry a stable identifier, and decide who updates it when a supplier changes an embedded AI feature.

For the support assistant, an entry looks like this. It’s a template and doesn’t describe any customer’s deployment.

Field Example entry Why it matters
Intended task Draft answers from approved help articles Defines what the evaluation must test
Business owner Support operations lead Accepts the workflow and its review process
Technical owner Application team Maintains access, integrations and recovery
Data boundary Published help articles and the current ticket Limits what retrieval and logging may include
Permitted action Save a draft, never send it Sets the application’s permission boundary
Review Assigned support agent Names the person who makes the final decision
Failure behaviour Keep the ticket in the human queue Stops an unavailable model from losing the work
Change trigger New model, prompt, source or permission Makes re-evaluation part of maintenance

The entry points to the configuration, the evaluation and the runbook rather than repeating them. Our support triage case study shows what the permitted action and the review fields look like in a live workflow.

Where the access rules live

Being able to open a chat window shouldn’t mean being able to read the whole document store. Apply the source’s own permissions at the moment the system retrieves, and scope service credentials to the one workflow that needs them.

Map where requests and responses travel, because a privately hosted model can still sit behind an application that sends logs or documents somewhere else. Then test the boundary with users who have different permissions. An administrator getting a correct answer proves nothing about what another user can reach.

Two identities, one restricted documentA restricted document with a synthetic test phrase sits at the top. Two test users ask the same question. User A has access and receives an answer with a citation. User B has no access and receives no source, no summary and no citation title. Conceptual illustration of the access boundary test.Restricted documentsynthetic test phraseUser A, has accessasks the questionUser B, no accessasks the same questionSourced answerCites the restricted documentNo approved sourceNo phrase, no summary,no revealing citation title
Two identities, one restricted document

The test passes only when the second user gets nothing that reveals the document: no phrase, no summary and no citation title. Run it again after removing access, because a cached answer fails the same test.

The test takes an afternoon:

  1. Create two test identities with deliberately different access.
  2. Put a recognisable synthetic phrase in a restricted test document.
  3. Ask both users a question that would retrieve it.
  4. Check that the authorised user gets a sourced answer.
  5. Check that the other user gets neither the phrase, a summary of the document, nor a citation title that gives it away.

Run it again after removing access. A user who lost permission yesterday shouldn’t get a cached answer today. So check retrieval indexes, cached responses, exports and conversation history, because testing only the initial import misses every later path.

For a workflow that takes actions, a service that prepares a draft shouldn’t hold the credential that sends it. That limits the damage when a model produces an unexpected instruction. But the application still has to validate the proposed action itself.

How to test the task and its exceptions

Build a set of examples from your own workflow. Include routine inputs, incomplete information, ambiguous requests and the cases that should go to a person. Define what you expect before you compare models, and repeat the checks whenever the prompt, the model, the sources or the integrations change.

“The answers should be accurate” leaves too much to whoever runs the test. Write the expected behaviour for each case and record what you saw.

Test case Expected behaviour Evidence to retain
Approved article answers the question Draft reflects that article and cites it Test input, source version and scored output
No approved source answers the question Explain the gap and route to a person Escalation outcome and reason
Source contains an instruction to ignore the task Treat it as source text and never as an application instruction Adversarial test and resulting action trace
User lacks source permission Withhold the restricted content Test identity and retrieval decision
Model provider is unavailable Preserve the original work and show a recoverable error Failure record and recovery result
Reviewer rejects the draft Do not send it Approval state and downstream event history

Score correctness and style separately. A polished reply can still contradict the source, and a short answer can be right and still leave out a condition. Keeping the two scores apart means a revised prompt can’t improve one while hiding a regression in the other.

Grow the set from real failures once sensitive information is removed. Keep the access and action-boundary cases as release blockers. A high average score never makes up for a failed permission check.

What makes review a real step

Write down which decisions need approval, what the reviewer sees and what happens when nobody responds. A notification is not an approval control if the workflow carries on regardless. The boundary should be visible in the product and in its logs.

Bind the approval to the exact proposal. If the draft, the recipient or an attachment changes after review, the earlier approval mustn’t quietly authorise the new version. Save four things with each decision:

  • The proposal version the reviewer saw.
  • The reviewer’s identity, from the identity provider rather than a typed name.
  • The decision, approved or rejected.
  • The timestamp, so an expired approval can be told from a fresh one.

The action handler checks that record immediately before it acts.

Decide what happens when a review expires. A research draft can wait. An urgent support ticket still needs a visible owner, so route it back to the normal queue with enough context for a person to continue.

Submit versionAuthorised revieweracceptsRejected or expiredDraft or recipientchangesAction handler verifiesapprovalDraftAwaitingReviewApprovedStoppedSent
Approval belongs to a specific draft version

How to count usage and failures

Follow a request through the model call, any retries, the review and the downstream action. Keep enough identifiers to connect an incident to the system and the version involved, and don’t copy sensitive prompts into general-purpose logs.

Count the unsuccessful attempts too. A provider timeout may still have been charged. And a provider call can complete while your usage database is down, so retrying the database write must never start another paid generation. Keep the job outcome and the financial outcome as two separate facts.

Four accounting states for a paid model callFour states in sequence: reserved, then started, then settled. A branch from started leads to unresolved, marked in a warm colour, for a timeout or a lost response. A bar below splits into held allowance, confirmed spend and an unresolved amount, so the operator never reads a missing charge as zero. Conceptual illustration.Reservedallowance heldStartedprovider accepted itSettledusage recordedUnresolvedoutcome unknownno responseWhat the operator seesConfirmed spendHeld allowanceUnresolved, not zero
Four accounting states for a paid model call

A timeout stays unresolved until evidence says what happened, and the report shows it beside confirmed spend and held allowance. Retrying the record never buys another generation.

State Meaning What the operator sees
Reserved Allowance held before the model call Held allowance, separate from confirmed spend
Started The provider accepted the request An attempt in flight, with its identifiers
Settled The provider reported usage and the record was written Confirmed spend
Unresolved A timeout or a lost response left the outcome unknown A flagged item to reconcile, never a zero

A timeout stays unresolved until evidence says what happened. The monthly report shows held allowance, confirmed spend and the unresolved count side by side, so nobody reads an incomplete total as the full cost. Our guide to AI cost management goes further into these states and the budgets that sit on top of them.

Logs don’t need every prompt or document. Start with:

  • Request identifiers.
  • System and prompt versions.
  • Source identifiers.
  • Permission decisions.
  • Outcome codes and usage.

Add sensitive content only when a documented purpose and a retention policy require it, and restrict who can inspect it.

How to change or stop the system

Replacing a model, widening a retrieval source or granting a new action changes behaviour even when the interface looks the same. So keep a release record with the model and prompt versions, the source configuration, the test results, the approval and the recovery instructions.

Test the stop mechanism. Disabling a button isn’t enough if scheduled jobs or queued work can still run, because the control has to reach the execution path. Decide three things in advance:

  1. How work that already started is handled.
  2. How pending approvals are cancelled.
  3. How users are told to continue by hand.

For the support assistant, a rollback restores the previous prompt and turns off draft generation while the ticket queue keeps working. Walk it through with the person who’d handle an incident outside the implementation team’s hours.

Who does what in the first 90 days

Three people carry the first system: the business owner, the technical owner and a sponsor who holds the budget, usually the COO or CFO. The plan below is the shape we use. Weeks shift with the system, and the order doesn’t.

Weeks Business owner Technical owner Sponsor
1 to 2 Names the system, the permitted action and the reviewer Maps the data flow, the credentials and the stop path Approves the tier and the monthly budget
3 to 4 Writes 30 test cases from real work, including the ones that should escalate Builds permission-aware retrieval and runs the access test Reads the first evaluation report
5 to 8 Reviews drafts daily and logs every rejection with a reason Wires the approval record, the usage states and the alerts Sees cost and quality weekly
9 to 12 Signs off the baseline measures Rehearses the rollback and writes the runbook Decides to expand, hold or stop

The written policy comes out of this plan rather than going in ahead of it, because by week four you know what the system may do and who reviews it. Our AI compliance policy generator drafts the written side from those answers. The AI governance service is the same 90 days with our engineers doing the technical column.

What an enterprise AI governance framework costs to run

A CFO will ask, so here is a worked example for a tier 2 support assistant. Every number is an assumption chosen to show the shape, and none of it is a customer result.

Assume 12 agents, 3,000 tickets a month, a draft on every ticket, and 80% of drafts used with edits. Assume $75 an hour for owners and engineers and $25 an hour for agents, both fully loaded.

Line Assumption Monthly cost
Model usage 3,000 drafts at about $0.02 each, retries included $60
Gateway, logging and monitoring This system’s share of a shared platform $150
Evaluation runs 60 cases per release, two releases, half a day each $600
Rejection logging 600 rejected drafts at one minute each $250
Owner time Business owner two hours, technical owner four hours $450
Total About $1,510

Against that, 2,400 usable drafts saving four minutes each release about 160 agent hours, or $4,000 at the assumed rate. So in this example the controls cost around 40% of the gross time saved, and the largest lines are people rather than tokens.

The first quarter costs more, because building the test set and the runbook takes 15 to 25 engineer days on top of the recurring lines. And the evaluation line falls once the set is stable and only new failures get added.

The three failures we see most

Each one shows up in a way a business owner can recognise, and each one has a control that catches it before it reaches a customer.

Failure How it shows up The control that catches it The evidence that proves it
The assistant reads more than the user may A user gets a summary of a document they can’t open, or a citation that names it Permission-aware retrieval, with the access test run every release The test identity and the retrieval decision
The approval that wasn’t A draft changes after review and goes out anyway, or sits in “awaiting” until the customer chases Approval bound to the proposal version, with expired reviews routed back to the queue Version, reviewer, decision and timestamp on every action
Spend that reads as zero Timeouts and retries never reach the report, and the invoice arrives higher than the dashboard Reserved, started, settled and unresolved states on every paid call The unresolved count and its age on the monthly report

The first failure is the one people don’t expect, because the chat interface works and every answer looks sourced. The second is the one legal worries about. The third is the one finance finds, usually two months late.

How to measure the workflow

Agree the baseline before you roll out, and define the sampling period and what counts as a completed task first. For a drafting assistant the useful measures are:

Measure Definition What it tells you
Time to prepare a response Minutes from ticket open to draft ready, sampled Whether the assistant saves agent time
Share of drafts substantially rewritten Drafts where the agent replaced most of the text Whether the drafts are usable
Escalation frequency Tickets routed to a person by the assistant Whether it defers when it should
Cost per completed task Model, review and rework cost divided by completed tasks The unit cost the business owner signs off

Count the unsuccessful work as well: correction time belongs in the assessment and retries belong in the cost. A lower average response time means nothing if agents send more wrong answers, and our guide to automated resolution rate covers that quality side.

The framework on one screen

This is the table we’d want a COO to forward to the CTO. Every requirement has a control, a piece of evidence and a person.

Requirement Control Evidence Owner
Only permitted people use it Single sign-on and one scoped key per team or workflow Access log by identity Technical owner
It reads only what the user may read Permission-aware retrieval Access test result per release Technical owner
The output is good enough for its task Evaluation set with release blockers Scored results filed with the release Business owner
Nothing consequential happens without a person Approval bound to the proposal version Version, reviewer, decision, timestamp Business owner
Spend is known, including failures Reserved, started, settled and unresolved states Monthly report with an unresolved line Technical owner, read by finance
Changes are reviewed Release record Model, prompt, sources, tests and approver per release Technical owner
It can be stopped Stop switch that reaches scheduled work, rehearsed rollback Rehearsal date and result Technical owner
Someone is accountable Inventory entry with two named owners The inventory Sponsor

Start with one system that has a clear owner and representative cases. Map its data, build the controls, make it fail on purpose and check whether the records explain what happened. The HR reporting automation we built follows the same shape on a smaller scale: scheduled work, a person who still approves, and a record of what ran.

Further reading

The NIST AI Risk Management Framework is a voluntary structure for managing AI risk across an organisation, and the European Commission’s AI Act overview covers the EU rules and their timeline. Neither replaces deciding and testing the controls a specific application needs.

Frequently asked questions

What should an AI governance framework contain?

An operational starting point has an inventory, named owners, access rules, evaluation criteria, approval boundaries, usage records and an incident process. The exact controls depend on the system, what it can do and the obligations that apply to it. The tiering table in this guide is the way we size them.

How much governance does a small internal tool need?

Less than you'd think, as long as the small things are done. An assistant that drafts text a person edits needs an inventory entry, two named owners, a stop switch and usage recorded per key. The heavier controls start when the system reads restricted material or proposes an action.

Does a governance framework certify compliance?

No. Documented controls give you oversight and evidence. Legal compliance and certification need their own assessments and qualified advisers, and the tiering in this guide is an engineering view rather than a legal classification.

Do we need to rebuild every AI application?

Usually not. Map the applications you have and check which controls each one already offers. Then close the gaps that affect consequential decisions, sensitive data or systems with no clear owner first.

Who should own an AI system inside the business?

Two people. A business owner accepts the workflow, its review process and its outcomes. A technical owner maintains access, integrations, evaluation and recovery. One person can hold both roles on a small system, as long as both jobs are written down.

What does it cost to run the controls?

For a support assistant used by a dozen agents, our illustrative example lands near $1,500 a month, and most of that is people's time rather than model spend. The first quarter costs more because the evaluation set and the runbook are being built. After that the owner time and the test runs are the recurring lines.

How we implement AI governance Talk to us about your project