How Insurers Can Implement an AI Governance Framework for NAIC Exhibit B
A practical guide to implementing an AI governance framework for NAIC Exhibit B, from AIS programs to vendor oversight and consumer evidence.
For Governance leads and the CRO, GC, and compliance officers who have to sign what they produce, at insurers preparing for NAIC AI examinations.
Read if You have an AI governance policy and need to turn it into evidence that can answer NAIC Exhibit B during an exam.
NAIC Exhibit B is the part of the AI Systems Evaluation Tool that tests a governance framework in operation. It asks for evidence that a program is written, adopted, reviewed, and wired into the rest of the company. An elegant but dormant document will not satisfy that inquiry.
The NAIC AI Systems Evaluation Tool offers two ways to respond to Exhibit B: a narrative and a checklist. The narrative is a series of open questions about governance framework, material AI uses, vendor oversight, professional service provider oversight, and additional framework design. The checklist is more direct: a written AIS Program, board or management involvement, and fourteen governance elements labeled 3a through 3n, covering everything from unfair trade practice risk to consumer notification 1. Regulators are looking for a coherent picture of a program that is actually operating.
Read the exhibit’s own preamble before you read the fourteen elements as a to-do list. Exhibit B states that its references to and questions about a company’s AI Systems Governance Risk Assessments “are not intended to be interpreted as creating new requirements” for those assessments, and it instructs regulators to customize the tool to limit information requested to more targeted inquiries 1. They map the ground an examiner may walk over, which is how this article reads them throughout.
This article is a practical guide to making that picture real. It walks the five domains that sit behind the fourteen checklist elements, with questions you can use to pressure-test your own program before an examiner does.
Start with the AIS Program as a document that has a history
The first checklist question is deceptively simple: Has the company adopted a written AIS Program? If yes, when was it adopted, and what is the frequency of review? The question is simple, but the answer carries a lot of weight. A written, dated program adopted with board or management involvement is the minimum entry fee; the adoption date and review cadence reveal whether the document has a life cycle and predates the scrutiny.
A regulator-ready AIS Program should be able to show its own provenance. When was the first version approved? What triggered each revision? Did the committee minutes reference the program, or did it sit untouched while the company deployed new AI systems? The review cadence matters because AI systems change faster than most insurance policies. A program that is reviewed annually may already be stale if the company released four new models in the meantime.
Board or management involvement is the second element. The checklist asks whether the board or management was involved in adopting the AIS Program, and then, as item 2a, what their role is in the AI Systems Governance Framework. The tool’s preface is explicit that its governance questions are not there to create new requirements 1, and nothing in it says the board must read every model validation. It wants an answer, from whichever body owns one. The answers that hold up come with meeting minutes, escalation triggers, and a record that oversight was exercised rather than merely assigned. The tool asks for none of that. Minutes and triggers are just how a company shows the answer is true. A governance framework that never reaches the board deck remains a draft.
Wire governance into the organization
Exhibit B is fundamentally an organizational test, and it assumes the wider program our AI governance in insurance guide maps out already exists. The narrative asks who maintains the framework, how it is integrated throughout the company, how responsibility is assigned, and how consistency is ensured. The checklist reinforces this through elements 3f through 3h, which ask whether AI risk is considered within ERM, ORSA, and the software development lifecycle 1.
The goal is to make AI governance indistinguishable from how the company already manages risk. If your AI risk committee is a separate track that reports to compliance once a quarter, it will look peripheral. If AI risk feeds into the same ERM process that owns underwriting, catastrophe, and operational risk, it will look like part of the fabric.
Three structural details usually separate programs that pass this test from ones that do not. First, each AI system has a business-line owner who understands the consumer impact alongside the technical model owner. Second, the governance committee can pause deployments or acquisitions before approval. Third, AI risk appears in the ORSA and ERM reporting that senior management already receives.
Build a risk and materiality loop that can explain its own decisions
You do not have to call every model high-risk under Exhibit B. You have to show how you decide. Element 3 asks you to reference the processes and procedures in your governance framework. Three of its items approach tiering from different directions: residual unfair-trade-practice risk (3a), adverse-consumer-outcome risk (3c), and quantified AI system risk levels (3k) 1. Notice that you are asked to reference something. If the framework does not say it, there is nothing to point at. The narrative adds a pointed suggested question: How does the insurance company assess autonomy, reversibility, and reporting impact risk of AI systems?
This is where implementation becomes harder than documentation. A risk-tiering scheme must be written, reviewed, and applied. A model that affects underwriting, pricing, claims, or health utilization should have a clear rationale for its tier, documented when the tier is assigned. The same applies to autonomy. A supportive tool that ranks internal work queues carries a different risk profile than an automated system that denies coverage with no human review. The framework must be able to explain why the two are treated differently.
Adverse consumer outcome tracking is the element most often underestimated. Element 3m asks whether consumer complaints resulting from AI systems are identified, tracked, and addressed 1. The complaint log becomes useful when it operates as a feedback loop. If a model produces a pattern of complaints, the framework should be able to show how that pattern was detected, escalated, and fed back into testing, revalidation, or a change in human-review thresholds. Without that loop, adverse outcomes remain disconnected anecdotes.
Connect vendor oversight without rebuilding it here
Exhibit B asks how the governance framework covers vendors and professional service providers 1. In the implementation sequence, the important step is the connection: procurement identifies the system and its intended use, the governance process assigns a risk tier, contract limitations become recorded gaps, and material vendor changes trigger review.
The detailed diligence questions, contract terms, audit rights, and ongoing review cadence belong in the AI vendor risk assessment. Exhibit B should point to that operating record and show how its findings reach the governance owner. It should not contain a second vendor program written for the examination.
Connect the decision to the consumer-facing record
Exhibit B asks about privacy, consumer awareness, notification, and procedures that address adverse consumer outcomes 1. The framework therefore needs an owner and an escalation path for notices, corrections, complaints, appeals, and accommodations. Product and state rules determine the exact language and timing.
Do not turn Exhibit B into a nationwide notice survey. Record the common control and link each state-specific obligation to its owner; the Colorado analysis and New York guide carry their own scopes.
One completed business decision is enough to test that: the decision evidence pack connects its actual notice and later appeal back to the model-assisted action.
Choose the format that matches your evidence
The narrative and checklist formats are both acceptable, and the choice matters less than the substance. Our NAIC AI Evaluation Tool guide covers the tactics of that choice; the short version is that the checklist rewards a program that fits the fourteen boxes, and the narrative rewards one that needs context. Where the decision gets more interesting is when neither fits alone: recent acquisitions, legacy systems, or governance that differs by entity.
Some companies may want to submit both: a checklist for the main answers and a narrative appendix for context. Before doing that, check with your domestic regulator (contact details by state). The tool instructs companies to work with the regulator if governance differs by entity, line of business, or state 1. The same principle applies to format. A concise, traceable response makes the examiner’s job easier.
A self-audit worksheet before the exam notice arrives
Here is a set of questions you can run now. They do not replace the exhibit, but they will tell you where you are already solid and where you are still assembling a story.
The download has each question with a status column and a tab mapping the five governance domains behind the fourteen checklist elements. Fill it as a living record and date each entry. A healthy in-progress row looks like “ERM integration: partial, AI risk added to the ORSA risk taxonomy in the Q2 review, owner CRO, due Q4.” This is an illustrative entry. A bare “yes” cannot show progress or ownership.
- Can you produce the current AIS Program, the date of adoption, the dates of each review, and the committee minutes that approved them?
- Does the system register produce counts that match the systems your AIS Program claims to govern? Exhibit A requests those counts; the register is the source behind them.
- Is there a board or management committee with a defined AI oversight role, and has it met in the last quarter?
- Can you show where AI risk appears inside your ERM or ORSA risk taxonomy?
- For each high-risk AI system, can you explain the tiering rationale, the testing performed, and the human-review threshold?
- Do you have a written process for identifying, tracking, and addressing AI-related consumer complaints?
- Do your vendor contracts include AI-specific audit rights, data disclosure, and update-notification clauses?
- Have you documented your consumer disclosure and appeals process for AI-influenced decisions?
If the answer to any of these is “we are working on it,” that is fine. The point is to know which answers are still in progress before an examiner asks.
What Exhibit B is really testing
Exhibit B is an implementation test. The regulator wants to see that a written program has owners, a review history, a place in ERM and ORSA, a risk and materiality loop, vendor controls, and consumer-facing accountability. The narrative format and the checklist format are just two ways of telling that same story.
If you start with the self-audit worksheet above, the gaps will surface quickly. Most carriers will find that their policy is stronger than their evidence. Closing that distance is the work, and it goes faster with no examiner in the building. Start with the systems you have already counted in Exhibit A, because every governance answer in Exhibit B will eventually be checked against that inventory. If that count has never been reconciled against what business lines actually run, the shadow AI analysis is the prior step.
One thing this does not settle. Exhibit B asks whether a program is documented and operating; it has no view on whether the program is any good. A carrier can evidence all fourteen elements and still have drawn its materiality line in the wrong place, and nothing in the exhibit will catch that. Passing it is a floor.
Footnotes
The Bottom Line
- Exhibit B tests governance in operation. A dated AIS Program adopted with board or management involvement is the minimum entry fee.
- The fourteen checklist elements map onto five domains: program documentation, organizational wiring, risk and materiality, vendor controls, and consumer-facing evidence.
- Narrative and checklist formats are both acceptable; choose the one that lets you evidence your strongest controls without exposing blank boxes.
- Third-party AI is still your AI under Exhibit B. Vendor contracts, audit rights, and ongoing monitoring must live inside your governance framework, not beside it.
- The fastest way to fail is to answer Exhibit B before Exhibit A is honest. Your governance narrative must match the systems you actually count.
AI Model Monitoring After the Model Goes Live
A playbook for insurers on AI model monitoring, validation, drift detection, and retesting records that satisfy NAIC Model Bulletin and Exhibit C expectations.
Continue →
Simon Li · Founding Editor
Much of his time goes into reading NAIC meeting papers, state bulletins, bills, court filings, and public comments. He also keeps the site's 51-jurisdiction tracker up to date.
Free · Weekly
Track these developments weekly
Get the InsureAI Wire dispatch in your inbox. Free, sourced, no spam.
Free weekly · No spam · Unsubscribe anytime
Related reading
The NAIC AI Systems Evaluation Tool, Exhibit by Exhibit
What the NAIC AI Systems Evaluation Tool is, what Exhibits A through D ask, how regulators select them, and where each evidence task belongs.
How AI Governance Works in Insurance
A plain-language map of AI governance in insurance: the regulatory stack, the program it expects, and where to go for rules, implementation, and evidence.
How Insurers Assess AI Vendor Risk
A practical NAIC-aligned checklist for AI vendor risk assessment: due-diligence questions, contract clauses, and the ongoing monitoring that stays with the insurer.
Inside the NAIC AI Model Bulletin
What the NAIC AI Model Bulletin is, how adoption works, what belongs in a written AIS Program, and which implementation guide to use next.
Information aggregation and analysis, not legal advice. See our disclaimer.