<!-- Canonical: https://www.aigentcy.com/blog/enterprise-ai-governance-framework-guide/ -->

# An AI governance framework your team can operate

A decision guide for the people who sign off AI systems: how much governance each one needs, who does what in the first 90 days, what it costs to run, and the evidence that proves the rules were followed.

Aigentcy 15 April 2026 16 min read Updated 7 September 2026

## Key takeaways

-   Give every AI system a business owner, a technical owner and a written limit on what it may do, and size the rest of the controls to that limit.
-   Put access and approval rules inside the application. A policy on its own can't stop an unauthorised action.
-   Test ordinary work, missing sources, access failures and recovery before you widen the rollout, and keep the access test as a release blocker.
-   Record paid attempts separately from job outcomes. A timeout leaves an uncertain charge, and the report must show it as unresolved rather than zero.
-   Keep one release record that ties the model, the prompt, the sources, the test results and the approval decision together.

**A requirement, its enforcement and its evidence**

A policy sentence becomes a control in the application and a record the operator can read. If any of the three is missing, the other two cannot prove the rule was followed.

It’s the second week of the quarter. Your support lead wants an assistant that drafts replies from the help centre, live on Monday. Legal asks what happens if it sends the wrong thing. Finance asks what it costs by December. You decide on Friday. An enterprise AI governance framework exists so that decision is quick and still holds up later.

It rests on one idea. Every rule you write about an AI system needs a control inside the application that applies the rule, and a record that shows the control ran. If any one of the three is missing, the other two can’t prove the rule was followed.

This guide is the version we run with customers: how much governance a system needs, who does what in the first 90 days, what the controls cost, the three failures we see most and the evidence to keep. It’s an engineering starting point. Legal advice and certification need qualified advisers.

## What an enterprise AI governance framework has to do

Take the support assistant. Before it goes live you need to know whether it can read restricted material, whether it’ll invent an answer when no source covers the question, and whether it can send a message without a person looking. Those three questions give governance a scope you can test.

The same shape covers knowledge search, document processing and scheduled reporting. The controls differ. But the chain from a requirement to its enforcement to its evidence stays the same, and the diagram below is that chain for one request.

From an AI policy requirement to operational evidence

```mermaid
flowchart TD
accTitle: From an AI policy requirement to operational evidence
  A[Define the permitted task] --> B[Check identity and data access]
  B --> C[Run the evaluated model and prompt]
  C --> D{Approval required?}
  D -->|Yes| E[Wait for an authorised reviewer]
  D -->|No| F[Perform the permitted action]
  E -->|Approved| F
  E -->|Rejected or expired| G[Stop and record the reason]
  F --> H[Record outcome and usage]
```

## How much governance is enough

Match the controls to what the system can do and to who’s affected when it’s wrong. A tool that drafts text a person edits doesn’t need the treatment of one that changes what a customer is charged. Two questions place a system: can it act with nobody in the path, and can its output change what a person gets.

| Tier | What the system does | Controls that apply | Evidence kept | Review |
| --- | --- | --- | --- | --- |
| 1, assist | Drafts or summarises material the user could already read. A person edits and sends. | Inventory entry, two named owners, a stop switch, usage per key | Owner, purpose, monthly usage | Twice a year, and on change |
| 2, act with review | Reads restricted sources or proposes an action a person approves: support drafts, purchase orders, leave reminders | Tier 1 plus permission-aware retrieval, the access test, an evaluation set with release blockers, approval bound to a version | Test results per release, an approval record per action, unresolved charges | Every release, and quarterly |
| 3, consequential | Affects a person’s money, rights, safety or access, or acts with nobody in the path | Tier 2 plus a legal assessment, a second reviewer for changes, a rehearsed rollback, monitoring with alerts | The assessment and who signed it, the incident log, monitoring history | Every release, monthly, and after any incident |

**Three governance tiers, each adding controls to the last**

Nothing is removed as a system moves right. A tier 3 system carries every tier 1 and tier 2 control plus the ones that need a lawyer, a second pair of eyes and a rehearsed way back.

The tiers are an engineering view. The law has its own. The EU AI Act sets [risk-based rules](https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai) that depend on the system, its intended purpose and your role as provider or deployer. And the fines in [Article 99](https://artificialintelligenceact.eu/article/99/) run up to EUR 35 million or 7% of worldwide turnover for prohibited practices.

So anything you’d put in tier 3 needs a classification from a qualified adviser, with the date it was made. A questionnaire can show you which information is missing, and it can’t classify a system from a few organisation-level answers. Our [EU AI Act readiness quiz](https://www.aigentcy.com/tools/eu-ai-act-readiness-quiz/) does the first job and stops there.

## What to record about each system

Start with the AI systems you already use, including the models embedded in software you bought. For each one, write down who owns it, what it’s for, who uses it, which data it reads, which providers it calls and which decisions its output influences.

Keep the record inside a process you already run, such as procurement or deployment review, because an annual spreadsheet misses every change between two reviews. Give each entry a stable identifier, and decide who updates it when a supplier changes an embedded AI feature.

For the support assistant, an entry looks like this. It’s a template and doesn’t describe any customer’s deployment.

| Field | Example entry | Why it matters |
| --- | --- | --- |
| Intended task | Draft answers from approved help articles | Defines what the evaluation must test |
| Business owner | Support operations lead | Accepts the workflow and its review process |
| Technical owner | Application team | Maintains access, integrations and recovery |
| Data boundary | Published help articles and the current ticket | Limits what retrieval and logging may include |
| Permitted action | Save a draft, never send it | Sets the application’s permission boundary |
| Review | Assigned support agent | Names the person who makes the final decision |
| Failure behaviour | Keep the ticket in the human queue | Stops an unavailable model from losing the work |
| Change trigger | New model, prompt, source or permission | Makes re-evaluation part of maintenance |

The entry points to the configuration, the evaluation and the runbook rather than repeating them. Our [support triage case study](https://www.aigentcy.com/case-studies/customer-support-automation/) shows what the permitted action and the review fields look like in a live workflow.

## Where the access rules live

Being able to open a chat window shouldn’t mean being able to read the whole document store. Apply the source’s own permissions at the moment the system retrieves, and scope service credentials to the one workflow that needs them.

Map where requests and responses travel, because a privately hosted model can still sit behind an application that sends logs or documents somewhere else. Then test the boundary with users who have different permissions. An administrator getting a correct answer proves nothing about what another user can reach.

**Two identities, one restricted document**

The test passes only when the second user gets nothing that reveals the document: no phrase, no summary and no citation title. Run it again after removing access, because a cached answer fails the same test.

The test takes an afternoon:

1.  Create two test identities with deliberately different access.
2.  Put a recognisable synthetic phrase in a restricted test document.
3.  Ask both users a question that would retrieve it.
4.  Check that the authorised user gets a sourced answer.
5.  Check that the other user gets neither the phrase, a summary of the document, nor a citation title that gives it away.

Run it again after removing access. A user who lost permission yesterday shouldn’t get a cached answer today. So check retrieval indexes, cached responses, exports and conversation history, because testing only the initial import misses every later path.

For a workflow that takes actions, a service that prepares a draft shouldn’t hold the credential that sends it. That limits the damage when a model produces an unexpected instruction. But the application still has to validate the proposed action itself.

## How to test the task and its exceptions

Build a set of examples from your own workflow. Include routine inputs, incomplete information, ambiguous requests and the cases that should go to a person. Define what you expect before you compare models, and repeat the checks whenever the prompt, the model, the sources or the integrations change.

“The answers should be accurate” leaves too much to whoever runs the test. Write the expected behaviour for each case and record what you saw.

| Test case | Expected behaviour | Evidence to retain |
| --- | --- | --- |
| Approved article answers the question | Draft reflects that article and cites it | Test input, source version and scored output |
| No approved source answers the question | Explain the gap and route to a person | Escalation outcome and reason |
| Source contains an instruction to ignore the task | Treat it as source text and never as an application instruction | Adversarial test and resulting action trace |
| User lacks source permission | Withhold the restricted content | Test identity and retrieval decision |
| Model provider is unavailable | Preserve the original work and show a recoverable error | Failure record and recovery result |
| Reviewer rejects the draft | Do not send it | Approval state and downstream event history |

Score correctness and style separately. A polished reply can still contradict the source, and a short answer can be right and still leave out a condition. Keeping the two scores apart means a revised prompt can’t improve one while hiding a regression in the other.

Grow the set from real failures once sensitive information is removed. Keep the access and action-boundary cases as release blockers. A high average score never makes up for a failed permission check.

## What makes review a real step

Write down which decisions need approval, what the reviewer sees and what happens when nobody responds. A notification is not an approval control if the workflow carries on regardless. The boundary should be visible in the product and in its logs.

Bind the approval to the exact proposal. If the draft, the recipient or an attachment changes after review, the earlier approval mustn’t quietly authorise the new version. Save four things with each decision:

-   **The proposal version** the reviewer saw.
-   **The reviewer’s identity**, from the identity provider rather than a typed name.
-   **The decision**, approved or rejected.
-   **The timestamp**, so an expired approval can be told from a fresh one.

The action handler checks that record immediately before it acts.

Decide what happens when a review expires. A research draft can wait. An urgent support ticket still needs a visible owner, so route it back to the normal queue with enough context for a person to continue.

Approval belongs to a specific draft version

```mermaid
stateDiagram-v2
accTitle: Approval belongs to a specific draft version
  [*] --> Draft
  Draft --> AwaitingReview: Submit version
  AwaitingReview --> Approved: Authorised reviewer accepts
  AwaitingReview --> Stopped: Rejected or expired
  Approved --> AwaitingReview: Draft or recipient changes
  Approved --> Sent: Action handler verifies approval
  Sent --> [*]
  Stopped --> [*]
```

## How to count usage and failures

Follow a request through the model call, any retries, the review and the downstream action. Keep enough identifiers to connect an incident to the system and the version involved, and don’t copy sensitive prompts into general-purpose logs.

Count the unsuccessful attempts too. A provider timeout may still have been charged. And a provider call can complete while your usage database is down, so retrying the database write must never start another paid generation. Keep the job outcome and the financial outcome as two separate facts.

**Four accounting states for a paid model call**

A timeout stays unresolved until evidence says what happened, and the report shows it beside confirmed spend and held allowance. Retrying the record never buys another generation.

| State | Meaning | What the operator sees |
| --- | --- | --- |
| Reserved | Allowance held before the model call | Held allowance, separate from confirmed spend |
| Started | The provider accepted the request | An attempt in flight, with its identifiers |
| Settled | The provider reported usage and the record was written | Confirmed spend |
| Unresolved | A timeout or a lost response left the outcome unknown | A flagged item to reconcile, never a zero |

A timeout stays unresolved until evidence says what happened. The monthly report shows held allowance, confirmed spend and the unresolved count side by side, so nobody reads an incomplete total as the full cost. [Our guide to AI cost management](https://www.aigentcy.com/blog/enterprise-ai-cost-control/) goes further into these states and the budgets that sit on top of them.

Logs don’t need every prompt or document. Start with:

-   Request identifiers.
-   System and prompt versions.
-   Source identifiers.
-   Permission decisions.
-   Outcome codes and usage.

Add sensitive content only when a documented purpose and a retention policy require it, and restrict who can inspect it.

## How to change or stop the system

Replacing a model, widening a retrieval source or granting a new action changes behaviour even when the interface looks the same. So keep a release record with the model and prompt versions, the source configuration, the test results, the approval and the recovery instructions.

Test the stop mechanism. Disabling a button isn’t enough if scheduled jobs or queued work can still run, because the control has to reach the execution path. Decide three things in advance:

1.  How work that already started is handled.
2.  How pending approvals are cancelled.
3.  How users are told to continue by hand.

For the support assistant, a rollback restores the previous prompt and turns off draft generation while the ticket queue keeps working. Walk it through with the person who’d handle an incident outside the implementation team’s hours.

## Who does what in the first 90 days

Three people carry the first system: the business owner, the technical owner and a sponsor who holds the budget, usually the COO or CFO. The plan below is the shape we use. Weeks shift with the system, and the order doesn’t.

| Weeks | Business owner | Technical owner | Sponsor |
| --- | --- | --- | --- |
| 1 to 2 | Names the system, the permitted action and the reviewer | Maps the data flow, the credentials and the stop path | Approves the tier and the monthly budget |
| 3 to 4 | Writes 30 test cases from real work, including the ones that should escalate | Builds permission-aware retrieval and runs the access test | Reads the first evaluation report |
| 5 to 8 | Reviews drafts daily and logs every rejection with a reason | Wires the approval record, the usage states and the alerts | Sees cost and quality weekly |
| 9 to 12 | Signs off the baseline measures | Rehearses the rollback and writes the runbook | Decides to expand, hold or stop |

The written policy comes out of this plan rather than going in ahead of it, because by week four you know what the system may do and who reviews it. Our [AI compliance policy generator](https://www.aigentcy.com/tools/ai-compliance-policy-generator/) drafts the written side from those answers. The [AI governance service](https://www.aigentcy.com/services/ai-governance/) is the same 90 days with our engineers doing the technical column.

## What an enterprise AI governance framework costs to run

A CFO will ask, so here is a worked example for a tier 2 support assistant. Every number is an assumption chosen to show the shape, and none of it is a customer result.

Assume 12 agents, 3,000 tickets a month, a draft on every ticket, and 80% of drafts used with edits. Assume $75 an hour for owners and engineers and $25 an hour for agents, both fully loaded.

| Line | Assumption | Monthly cost |
| --- | --- | --- |
| Model usage | 3,000 drafts at about $0.02 each, retries included | $60 |
| Gateway, logging and monitoring | This system’s share of a shared platform | $150 |
| Evaluation runs | 60 cases per release, two releases, half a day each | $600 |
| Rejection logging | 600 rejected drafts at one minute each | $250 |
| Owner time | Business owner two hours, technical owner four hours | $450 |
| Total |  | About $1,510 |

Against that, 2,400 usable drafts saving four minutes each release about 160 agent hours, or $4,000 at the assumed rate. So in this example the controls cost around 40% of the gross time saved, and the largest lines are people rather than tokens.

The first quarter costs more, because building the test set and the runbook takes 15 to 25 engineer days on top of the recurring lines. And the evaluation line falls once the set is stable and only new failures get added.

## The three failures we see most

Each one shows up in a way a business owner can recognise, and each one has a control that catches it before it reaches a customer.

| Failure | How it shows up | The control that catches it | The evidence that proves it |
| --- | --- | --- | --- |
| The assistant reads more than the user may | A user gets a summary of a document they can’t open, or a citation that names it | Permission-aware retrieval, with the access test run every release | The test identity and the retrieval decision |
| The approval that wasn’t | A draft changes after review and goes out anyway, or sits in “awaiting” until the customer chases | Approval bound to the proposal version, with expired reviews routed back to the queue | Version, reviewer, decision and timestamp on every action |
| Spend that reads as zero | Timeouts and retries never reach the report, and the invoice arrives higher than the dashboard | Reserved, started, settled and unresolved states on every paid call | The unresolved count and its age on the monthly report |

The first failure is the one people don’t expect, because the chat interface works and every answer looks sourced. The second is the one legal worries about. The third is the one finance finds, usually two months late.

## How to measure the workflow

Agree the baseline before you roll out, and define the sampling period and what counts as a completed task first. For a drafting assistant the useful measures are:

| Measure | Definition | What it tells you |
| --- | --- | --- |
| Time to prepare a response | Minutes from ticket open to draft ready, sampled | Whether the assistant saves agent time |
| Share of drafts substantially rewritten | Drafts where the agent replaced most of the text | Whether the drafts are usable |
| Escalation frequency | Tickets routed to a person by the assistant | Whether it defers when it should |
| Cost per completed task | Model, review and rework cost divided by completed tasks | The unit cost the business owner signs off |

Count the unsuccessful work as well: correction time belongs in the assessment and retries belong in the cost. A lower average response time means nothing if agents send more wrong answers, and [our guide to automated resolution rate](https://www.aigentcy.com/blog/support-ai-resolution-metrics/) covers that quality side.

## The framework on one screen

This is the table we’d want a COO to forward to the CTO. Every requirement has a control, a piece of evidence and a person.

| Requirement | Control | Evidence | Owner |
| --- | --- | --- | --- |
| Only permitted people use it | Single sign-on and one scoped key per team or workflow | Access log by identity | Technical owner |
| It reads only what the user may read | Permission-aware retrieval | Access test result per release | Technical owner |
| The output is good enough for its task | Evaluation set with release blockers | Scored results filed with the release | Business owner |
| Nothing consequential happens without a person | Approval bound to the proposal version | Version, reviewer, decision, timestamp | Business owner |
| Spend is known, including failures | Reserved, started, settled and unresolved states | Monthly report with an unresolved line | Technical owner, read by finance |
| Changes are reviewed | Release record | Model, prompt, sources, tests and approver per release | Technical owner |
| It can be stopped | Stop switch that reaches scheduled work, rehearsed rollback | Rehearsal date and result | Technical owner |
| Someone is accountable | Inventory entry with two named owners | The inventory | Sponsor |

Start with one system that has a clear owner and representative cases. Map its data, build the controls, make it fail on purpose and check whether the records explain what happened. The [HR reporting automation](https://www.aigentcy.com/case-studies/hr-automation/) we built follows the same shape on a smaller scale: scheduled work, a person who still approves, and a record of what ran.

### Further reading

The [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework) is a voluntary structure for managing AI risk across an organisation, and the [European Commission’s AI Act overview](https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai) covers the EU rules and their timeline. Neither replaces deciding and testing the controls a specific application needs.

## Frequently asked questions

What should an AI governance framework contain?

An operational starting point has an inventory, named owners, access rules, evaluation criteria, approval boundaries, usage records and an incident process. The exact controls depend on the system, what it can do and the obligations that apply to it. The tiering table in this guide is the way we size them.

How much governance does a small internal tool need?

Less than you'd think, as long as the small things are done. An assistant that drafts text a person edits needs an inventory entry, two named owners, a stop switch and usage recorded per key. The heavier controls start when the system reads restricted material or proposes an action.

Does a governance framework certify compliance?

No. Documented controls give you oversight and evidence. Legal compliance and certification need their own assessments and qualified advisers, and the tiering in this guide is an engineering view rather than a legal classification.

Do we need to rebuild every AI application?

Usually not. Map the applications you have and check which controls each one already offers. Then close the gaps that affect consequential decisions, sensitive data or systems with no clear owner first.

Who should own an AI system inside the business?

Two people. A business owner accepts the workflow, its review process and its outcomes. A technical owner maintains access, integrations, evaluation and recovery. One person can hold both roles on a small system, as long as both jobs are written down.

What does it cost to run the controls?

For a support assistant used by a dozen agents, our illustrative example lands near $1,500 a month, and most of that is people's time rather than model spend. The first quarter costs more because the evaluation set and the runbook are being built. After that the owner time and the test runs are the recurring lines.

## Related reading

-   [AI cost management: what an accepted result costs](https://www.aigentcy.com/blog/enterprise-ai-cost-control/)How to report AI spend the way finance reports everything else: by team, by workflow and by accepted result, with unknown charges kept separate from zero.
-   [Enterprise AI enablement: governed access for the whole company](https://www.aigentcy.com/blog/enterprise-ai-enablement/)Give each team governed access to a few approved models through one gateway, and measure adoption by the work that gets finished.
-   [Automated resolution rate: measuring support AI by outcomes](https://www.aigentcy.com/blog/support-ai-resolution-metrics/)A bot can answer more conversations while your agents get harder work. Measure verified resolution, safe escalation, repeat contact and the cost per resolved case before you trust the headline rate.
-   [Support triage with replies ready for review](https://www.aigentcy.com/case-studies/customer-support-automation/)Case study
-   [HR reporting and reminders in BambooHR](https://www.aigentcy.com/case-studies/hr-automation/)Case study

[How we implement AI governance](https://www.aigentcy.com/services/ai-governance/) [Talk to us about your project](https://www.aigentcy.com/contact/)
