The NAIC AI Systems Evaluation Tool, Exhibit by Exhibit
The NAIC AI Systems Evaluation Tool is now the AI Risk Evaluation Supplement in the record. What Exhibits A through D ask, and how a regulator picks among them.
In this article
For Compliance, examination, model-risk, actuarial, and legal teams preparing to understand or answer the NAIC's optional supplemental exhibits.
Read if You need to know what each exhibit does and how they connect, without turning the tool into a general governance manual.
Twelve states are taking part in a pilot in which regulators are applying the supplement across market conduct exams, financial exams, financial analyses, and more general regulatory inquiries,1 and the instrument has acquired a second name. The official summary of the Big Data and Artificial Intelligence (H) Working Group’s August 13, 2026 session calls it the AI Risk Evaluation Supplement.1 The file posted for public download still calls itself the AI Systems Evaluation Tool, and every page of it is watermarked DRAFT.2 Both descriptions are current. A team that may have to answer the thing should work from what the exhibits ask, because the exhibits are what arrives.
The supplement sits inside the broader insurance AI governance map.
Status, verified August 15, 2026. The supplement has not been adopted. In the working group’s August 13 summary, the only formal action recorded is the adoption of the group’s July 22 minutes; for the instrument itself the summary records a pilot update and no adoption or exposure action.1 The pilot began in March, involves twelve participating states, and continues through September.1
Two names and one unfinished document
The August 13 summary is worth reading for its own vocabulary. Recording what happened at the June 1 session, it uses the old name: an update on the “Artificial Intelligence (AI) Systems Evaluation Tool pilot process.” Recording the July 22 session, it switches, with the change flagged once and only once: an update on the “AI Risk Evaluation Supplement (formerly known as the AI Systems Evaluation Tool) pilot process.” Describing its own August 13 business, it drops the parenthetical and simply reports an update on the AI Risk Evaluation Supplement pilot process.1 Our reading is that the record now treats the new name as settled and the old one as history. The summary was written after the fact, so it fixes the working group’s current usage rather than the date on which the name actually changed.
The posted document has not caught up. Its cover reads “Artificial Intelligence Systems Evaluation: Optional Supplemental Exhibits for State Regulators,” its page footers read “AI Systems Evaluation Tool,” and the version number lives outside the document, in the committee page’s file label for the clean and tracked changes copies.23 The draft status is visible in smaller ways too. The cover lists Exhibit D as “AI Systems Model Data Details” while the exhibit’s own header reads “AI Systems Data Details,” and the definitions section contains an unfinished line offering an alternate definition of neural network models, attributed to an abbreviation the document never expands.2
None of that changes what an insurer would be asked. It does tell you how to cite the thing: name the exhibit and the question, because the cover, the footers, and the official record each give the document a different name.
What the supplement is, and what it is not
The document is addressed to state regulators and describes itself as optional supplemental exhibits.2 It supports market conduct, financial analysis, and financial examination review of AI systems, and it says regulators should continue to treat existing NAIC resources as authoritative while drawing on it to understand a company’s use of AI. Inquiries under it are to be coordinated consistently with the Market Regulation Handbook, the Financial Condition Examiners Handbook, and the Financial Analysis Handbook.2
Three consequences follow directly from the text.
Requests will be trimmed. Every exhibit carries the same regulator instruction to customize the tool and limit what is requested for a targeted or limited scope inquiry.2 An insurer should answer the request it receives rather than the public file, and should not present a carrier-built worksheet, including ours, as an official regulator form.
Prior work can count. If information sought through the tool has already been provided to that department or any other state insurance department, the company’s response should say so, and regulators may accept the earlier submission if it remains current and applicable.2 That is permission, not reciprocity, which is precisely why answers given in one state need to survive being read next to answers given in another.
Answers shape scope. The document tells regulators that responses will be considered when identifying the insurer’s inherent risks and should affect the planned examination or inquiry approach and the nature, timing, and extent of further procedures.2 The exhibits are a filter as much as a questionnaire. Separately, the confidentiality section instructs regulators to cite examination or other authority when requesting the information and to cite the relevant confidentiality statutes or protections.2 The obligation to respond, and the protection for what is filed, come from that cited authority rather than from the exhibits.
How a regulator chooses which exhibit to send
Before the exhibits, the document sets out a short routing table headed “Which Exhibit to Use?” that pairs a regulatory purpose with the exhibits that serve it.2 It is addressed to the other side of the correspondence, which is why insurers tend to skip it, and it is the most useful page in the file for anyone trying to read intent from an incoming request.
| Risk identification or assessment | Exhibits marked |
|---|---|
| Identify reputational risk | B, checklist form |
| Review company practices related to consumer complaints | B |
| Assess company financial risk, number of models implemented recently | A, and B in checklist form |
| Identify adverse consumer outcomes, AI systems and data use by operational area | A, B, C, D |
| Evaluate actions taken against the company’s use of high-risk AI systems, as defined by the company | C |
| Evaluate robustness of AI controls | B, C |
| Determine the types of data used by operational area | D |
One purpose in that table calls for all four exhibits, and it is the adverse consumer outcome inquiry. Our reading is that the shape of a request carries information: a lone Exhibit D asks a data question, an Exhibit B and C pairing asks a controls question, and the full set signals the consumer outcome track. That is an inference from the document’s own routing logic, not a rule, and a regulator remains free to build a request the table does not describe.
Exhibit A: counts, and the definition hiding inside them
Exhibit A quantifies the regulated entity’s use of AI systems. Its purpose line explains the sequencing plainly: based on the responses, regulators may ask for governance in Exhibit B, high-risk models in Exhibit C, and data types in Exhibit D where there is risk of adverse consumer outcomes or material adverse financial impact. The instructions add that a company’s answers may show AI use so limited or so low in inherent risk that no further inquiry is warranted.2
The grid has fifteen rows of operational or program areas, from marketing and premium quotes through underwriting, ratemaking, claims, customer service, utilization management, fraud, investments, legal, producer services, reserves, catastrophe triage, and reinsurance, ending in a catch-all row for other insurance practices. For each row it asks four counts and the use cases: models currently in use, models with direct consumer impact, models with material financial impact, and models implemented in the past twelve months.2
The definition underneath those counts matters more than the counts themselves. Exhibit A instructs that models which augment or automate decision making related to consumers are considered to have direct consumer impact.2 The appendix then defines augmentation as suggesting an answer or advising a human decision maker, automation as operating without human intervention, and support as providing information without suggesting a decision or action.2 Read together, the second column is a function of how the insurer classifies autonomy, and a system genuinely confined to support falls outside it. Two carriers with identical technology can report different numbers here without either misreporting, which is a good reason to write down the classification rule before writing down the number.
The header block also asks for the line of business covered and an “as of” date, and the instructions tell companies whose governance differs by entity, line of business, or state to work with their domestic regulator on whether multiple submissions are needed.2 The AI inventory playbook owns the register that makes these counts reproducible. Exhibit A itself stops well short of that register.
Exhibit B: one framework, two ways to show it
Exhibit B exists in a narrative version and a checklist version, and the choice is real rather than cosmetic.
The narrative version asks the company to provide its AI governance framework and then discuss who maintains it, the governance structure and board reporting frequency, its integration and remediation across the organization, and where responsibility is assigned.2 It also asks how the effectiveness of the framework and of individual models is assessed and modified, and how AI systems enter Own Risk and Solvency Assessment (ORSA) and enterprise risk management work. Further questions cover AI systems carrying material financial or consumer impact, vendor policy, model design, and the validation and testing of internally and externally developed models, and oversight of professional service providers such as actuarial, claim, MGA, and audit services.2 The last of the numbered questions turns to other aspects of framework design and assessment, including the units responsible and the approach and frequency of that assessment.2 One further item is marked as a suggested additional question: how the company assesses autonomy, reversibility, and reporting impact risk.2
The checklist version asks whether a written AIS Program has been adopted, when, and how often it is reviewed, whether the board or management was involved, and what their role is. Then it runs fourteen sub-items, from residual unfair trade practice risk at one end to consumer awareness through disclosure at the other, with ORSA and the software development lifecycle among the elements in between.2
The checklist puts a page number column next to each sub-item and a space to explain what is not specified in the governance documents. That is the test worth rehearsing. A narrative can carry a well-run program compactly, but the checklist asks where in the file each element lives, and an item nobody can point to reads as a gap regardless of how well the practice actually works. Note also the document’s own limit: the questions about governance risk assessment elements are not intended to create new requirements.2 The framework implementation guide turns these subject areas into an operating sequence, and the Model Bulletin guide explains the authority and AIS Program expectations behind them.
Exhibit C: the systems the insurer itself calls high risk
Exhibit C collects details on high-risk AI system models, for example those making automated decisions that could cause adverse consumer outcomes, material financial impact, or material financial reporting impact. The scope rule is explicit: AI system risk criteria are set by the insurance company. The exhibit adds that to identify which models to ask about, regulators may request the company’s risk assessment and a model inventory if those have not already been provided.2 Which of a carrier’s own systems fall inside that company-set scope is a separate question, and screening the inventory is where it gets answered.
Thirteen requests follow for each model.2 Four identify it: name and version number, model type, implementation date, and whether it was built in-house or by a named vendor. Three ask how the company classifies it: risk classification, known risks and limitations, and whether the system automates, augments, or supports. Four ask for evidence: output testing for drift, accuracy, unfair trade practices, unfair discrimination, and performance degradation, plus pre-deployment validation and ongoing monitoring; the last test date; use cases and purpose; the effect on financial statements, risk assessment or controls. The remaining two are legal: the compliance review against applicable law including unfair trade practices and unfair claims settlement statutes, and, where law permits, any action taken against the company over the model, from informal agreements and voluntary compliance plans to consent orders.2
Two boundaries matter for preparation. The instructions acknowledge that the template refers to both AI systems and models and that a company may need to discuss with regulators which level a given answer describes.2 And the listed fields contain no dedicated human-override item, so meaningful review sits in the testing and AI type answers rather than in a field of its own. What that looks like in a claims operation is examined in Agentic AI in Claims, while deployed-model thresholds belong to model monitoring.
Exhibit D: twenty-five data categories and their sources
Exhibit D’s purpose line asks for detailed information on “the source(s) and type(s) of data used in AI System(s),” in order “to identify risk of adverse consumer impact, unfair trade practices, material financial impact, or material financial reporting impact.”2 Its grid lists twenty-five data element categories. Twenty-four of them are named and run in alphabetical order, and the twenty-fifth is the open row. The range runs from traditional underwriting variables such as loss experience, income, and driving behavior to non-traditional ones such as voice analysis, online social media, and reasonable accommodations granted or requested, before that open row catches anything the named categories miss.2
For each category the company can be asked for the type of AI system involved, how the data is used across insurance operations including operational practices by line of insurance, the internal data source, and the third-party vendor name.2 The trigger is development: the instruction covers data elements used in the training or test data as part of developing the models, and tells companies to leave a row blank where the source is not used. A separate response is contemplated for each line of business.2
The exhibit maps sources. It stops before proxy testing, data quality thresholds, audit certification, and update frequency, which may arise under other obligations and should be labeled as such rather than folded into an Exhibit D answer. The Exhibit D playbook owns the field-by-field preparation work.
The seams a reviewer can test
The exhibits describe one program from four angles, and their answers have to reconcile.
| Seam | A question that exposes a mismatch |
|---|---|
| A to B | Does the AIS Program govern every operational area that Exhibit A says uses AI? |
| A to C | Can every selected high-risk model be found in the register supporting the Exhibit A counts? |
| C to D | Do the data sources in Exhibit D match the inputs and vendors described for the model in Exhibit C? |
| B to C | Does the risk classification follow the method the governance response says the insurer uses? |
| B to D | Do vendor and data controls cover the external sources actually listed? |
| A to D | Does the autonomy rule behind the direct consumer impact count match the systems whose data you described? |
These checks compare the exhibits without adding new ones. They are worth running because a polished answer still fails when another department produces a record that contradicts it.
Where the pilot stands, and what it is not
Three status words get used interchangeably here, and they carry different consequences. Pilot use means live regulators are asking these questions now under whatever examination authority they cite. Exposure would mean a draft formally out for comment. Adoption would mean the working group and its parent committee approved a version. None of the three would make the exhibits mandatory on their own, since the document remains optional supplemental material whose force in any given request comes from the authority the regulator cites.
Only the first of those three is on the record.1 The participating state regulators are determining which companies to include and coordinating regulatory communication, and public updates are promised throughout the pilot.1 The 2026 Fall National Meeting follows in Dallas, November 14 through 17.4 The pilot summary’s working timeline is what connects the pilot to that date: September and October to update the instrument on pilot feedback and reissue it for review, then November to consider the updated version for adoption at the Fall meeting.5 Our reading is that the timeline records an intention rather than a fixed agenda, since the step it names for November is consideration.
Preparing in the order the document implies
The register comes before the exhibits. Fix the classification rules first, because the autonomy rule decides the Exhibit A consumer impact column and the company’s own risk criteria decide who appears in Exhibit C. Then confirm the counts against the register and date them. Then use the actual request, and the routing table above, to work out which of B, C, and D the regulator is really pursuing. Prose comes last, after the numbers agree.
Where the company cannot name an owner, a version, a data source, or a last test date, record the gap and route it to the guide that owns the underlying work rather than building another general checklist here. The exhibits are answerable when one system register can support all four views without contradiction, and organization-wide exam packaging remains a separate, scheduled guide.
Footnotes
-
National Association of Insurance Commissioners, Big Data and Artificial Intelligence (H) Working Group meeting summary report, August 13, 2026, 2026 Summer National Meeting, Columbus, Ohio, circulated as Attachment Two to the Innovation, Cybersecurity, and Technology (H) Committee, August 14, 2026. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7
-
National Association of Insurance Commissioners, Artificial Intelligence Systems Evaluation: Optional Supplemental Exhibits for State Regulators, posted by the Big Data and Artificial Intelligence (H) Working Group as the clean version 4.0. Page footers read “AI Systems Evaluation Tool” and every page carries a DRAFT watermark. Retrieved August 15, 2026. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17 ↩18 ↩19 ↩20 ↩21 ↩22 ↩23 ↩24 ↩25 ↩26 ↩27 ↩28 ↩29
-
National Association of Insurance Commissioners, Big Data and Artificial Intelligence (H) Working Group, listing the working group’s 2026 charges and posted documents, including the clean and tracked changes copies of AI System Evaluation Tool 4.0 and the Pilot Project Summary. Retrieved August 15, 2026. ↩
-
National Association of Insurance Commissioners, Events, listing the 2026 Fall National Meeting in Dallas, November 14 through 17, 2026. Retrieved August 15, 2026. ↩
-
National Association of Insurance Commissioners, AI Systems Evaluation Tool Pilot: Pilot Project Background, whose working timeline schedules a September and October update of the tool on pilot feedback and then reads “NOVEMBER: Consider the updated Tool for adoption at the Fall National Meeting.” The document carries no printed date. Retrieved August 1, 2026. ↩
The Bottom Line
- The official record now calls the instrument the AI Risk Evaluation Supplement. The file posted for download still calls itself the AI Systems Evaluation Tool and is watermarked DRAFT on every page.
- Nothing has been adopted. The August 13, 2026 working group summary records a pilot update and no adoption or exposure action for the instrument.
- The document routes: Exhibit A counts come first, and the regulator's stated purpose decides whether B, C, or D follows.
- Exhibit C covers only the systems the insurer itself classifies as high risk, and the insurer sets the criteria.
- The four exhibits describe one AI footprint from four angles. The seams between them are where preparation usually fails.
How Shadow AI Undermines NAIC Exhibit A
Shadow AI corrupts the counts reported in NAIC Exhibit A. What ungoverned AI use means for insurers, and a practical plan to close the gap.
Continue →
Simon Li · Founding Editor
I write InsureAI Wire and maintain its 51-jurisdiction tracker. Most of the work is reading: NAIC working group papers, state bulletins, bills, court filings, and public comment letters. Every claim on the site carries the document it came from, so you never have to take my word for it.
Free · Weekly
Track these developments weekly
Get the InsureAI Wire dispatch in your inbox. Free, sourced, no spam.
Free weekly · No spam · Unsubscribe anytime
Related reading
AI Governance Documents to Prepare for a Market Conduct Exam
The insurance AI exam documentation to have ready for a market conduct exam: an insurer-built readiness file of eight evidence categories, and how to record a gap.
How to Identify the High-Risk AI Systems Exhibit C Asks About
Five screening lines that turn an existing AI inventory into a defensible list of the high-risk systems NAIC Exhibit C asks about, and the record behind it.
Who Owns the Evidence in Insurance AI Governance
Insurance AI governance roles as an ownership matrix: which function prepares each piece of NAIC evidence, which one signs it, and who answers for it in an exam.
AI Model Monitoring After the Model Goes Live
A playbook for insurers on AI model monitoring, validation, drift detection, and retesting records that satisfy NAIC Model Bulletin and Exhibit C expectations.
Information aggregation and analysis, not legal advice. See our disclaimer.