PLAYBOOK Steps, templates, and checklists you can run as written

Prove AI Governance

Model Data Documentation for NAIC Exhibit D

A playbook for insurers preparing NAIC Exhibit D responses: document the 25 data elements, show internal and third-party sources, and close the data gap before the exam.

For Data scientists, compliance officers, and model-risk leads at insurers preparing for an NAIC AI Systems Evaluation Tool request.

Read if You need to turn Exhibit D's data questions into a file list your team can assemble today.

By Simon Li · Updated JUL 31, 2026 · 11 min read

InsureAI Wire seal IAW 2026 Primary sources Methodology →

Engraved cover illustration: Model Data Documentation for NAIC Exhibit D
Ask AI

Exhibit D is where most NAIC AI Systems Evaluation Tool submissions fall apart for a boring reason. The model works, the governance memo exists, but nobody wrote down where the data came from. By the time a regulator asks for the source of a geodemographic score or the vendor behind a telematics feature, the data science team is already three models past the one in question.

The NAIC’s Exhibit D, labeled AI Systems Model Data Details, asks insurers to document the source and type of data used in AI System models. The purpose is not to evaluate the model’s accuracy. It is to identify risk of adverse consumer impact, unfair trade practices, material financial impact, or material financial reporting impact 1. This playbook turns the exhibit’s 25 data elements into a file-and-field checklist that teams can fill before the request arrives.

The document you are preparing for calls itself Optional Supplemental Exhibits for State Regulators, and it tells every regulator picking it up to cut the questions down to the inquiry at hand. Exhibit D also sits downstream: a regulator works from Exhibit A’s counts and asks for data types only where those counts point at consumer or financial risk 1. So filling in all 25 rows is readiness work. Nobody is obliged to come and ask you for it.

What Exhibit D asks for

The exhibit uses a table with one row per data element and five columns: the type of data element, the type of AI system, how the company uses it across insurance operations, the internal source, and the third-party source or vendor name 1. The instructions are direct: if a data element is used in training or testing, say whether it is sourced internally or from a third party, and if it is third-party, name the vendor. If a data element is not used, leave it blank 1.

The data elements read like a map of modern insurance data: aerial imagery, geocoding, telematics, medical and biometric information, online social media, and traditional underwriting variables such as loss experience, vehicle data, and weather. The list runs alphabetically to 25 rows, the last of which is an open “Other: Non-Traditional Data Elements” line asking for examples 1. The newest addition in Version 4.0 is Reasonable Accommodations or Policy Modifications, which captures accommodations granted or requested 12.

The 25 data elements, organized for preparation

The fastest way to prepare is to group the data elements by function, then assign each to the owner who actually knows the answer. The NAIC table is alphabetical and transactional; a carrier should translate it into a work plan.

Exhibit D data categories, likely owners, and evidence to gather
NAIC categoryExamplesLikely ownerEvidence to gather
Location and propertyAerial imagery, geocoding, geo-demographics, natural catastrophe hazard, weatherUnderwriting, actuarial, GISData purchase agreements, feature dictionaries, vendor methodology documents
Vehicle and drivingDriving behavior, telematics, vehicle-specific dataPersonal lines, UBI teamDevice or app data flows, consent records, vendor score documentation
Demographic and financialAge, gender, ethnicity/race, income, job history, education level, household composition, personal financial informationMarketing, underwriting, data acquisitionInternal data dictionaries, third-party vendor lists, proxy-screening results
Health and biometricMedical, biometrics, genetic information, pre-existing conditions, diagnostic data, facial or body detection, voice analysisLife, health, claimsHIPAA and privacy assessments, vendor contracts, data use restrictions
Behavioral and alternativeOnline social media, image/video analysis, criminal convictions, crime statistics, consumer or other risk scoresMarketing, fraud, underwritingVendor names, scoring logic disclosures, model cards
Core insuranceLoss experienceActuarial, claimsData lineage, claim-system extracts, reserving methodology
v4.0 additionReasonable Accommodations or Policy ModificationsCompliance, claims, customer serviceAccommodation logs, policy modification records, training data annotations
Open rowOther: Non-Traditional Data Elements (examples requested)Data science, model riskFeature dictionary entries for anything the 24 named rows do not cover

This grouping lets you route each row to the right team. The GIS team owns geocoding. The UBI team owns telematics. The compliance team owns the new accommodations field. A single data-science owner should coordinate the final table, but the answers must come from the people who bought, built, or configured the data.

How to document each data element

Exhibit D asks five questions per data element. The answer should be factual and short. Regulators are not looking for a narrative; they are looking for provenance.

1. Type of data element. Use the NAIC’s category name or a close plain-language equivalent. If the model uses a derived variable, such as a credit score or a geodemographic cluster, map it back to the source category. A score built from ZIP code demographics is still geo-demographic data.

2. Type of AI system. The column offers machine learning versus generative AI as its example contrast, and the same data element can feed more than one of them 1. List the system or model where the data is used, and note whether it is a predictive model, a generative assistant, or a rules-based system with an AI component. Those three labels are a practical shorthand the exhibit never publishes, so an accurate plain description beats a borrowed term.

3. How the company uses the data. This is the operational description. Be specific by line of business and use case. “Used in homeowners pricing to map property-level risk” is more useful than “used in underwriting.” The NAIC asks for operational practices by line of insurance 1.

4. Internal source. If the data comes from your own systems, name the system and the table or field where feasible. For loss experience, that might be the claims data warehouse. For policy modifications, that might be the policy administration system. The goal is to show you can trace the data to an internal record.

5. Third-party source or vendor name. If the data comes from outside, name the vendor or data provider. Do not write “external vendor” or “third-party score.” The NAIC instructions explicitly ask for the vendor name 1. If the same data flows through multiple vendors, list the originating source and any intermediaries in the description column.

Exhibit D asks five questions for every data element an AI system uses: the type of data element, the type of AI system, how the company uses the data, the internal source, and the third-party source or vendor name. A vendor score without a vendor name is a missing answer. 1 TYPE OF DATA ELEMENT 2 TYPE OF AI SYSTEM 3 HOW THE COMPANY USES THE DATA 4 INTERNAL SOURCE 5 THIRD-PARTY SOURCE OR VENDOR NAME : NO VENDOR NAME = MISSING ANSWER
FIG. 1: THE FIVE QUESTIONS EXHIBIT D ASKS PER DATA ELEMENT

The most common gaps

Carriers tend to have one of three problems when they first fill in Exhibit D.

The vendor score without a name. A model uses a “consumer risk score” but the contract is held by another department, or the vendor was acquired and renamed. The answer in Exhibit D must be the current vendor name and the data element it provides. If the score is proprietary, the contract and any model card or methodology disclosure should be attached.

The derived feature that forgot its origin. A model uses a feature engineered from multiple sources, such as a weather-adjusted geocoding score. The model team may know the feature; the data team may know the sources. Exhibit D requires both the source and the use case. Document the feature dictionary at the same time as the source table.

The reasonable-accommodations blind spot. Version 4.0 added Reasonable Accommodations or Policy Modifications to the data list 2. This is not a traditional underwriting variable. It asks whether the carrier tracks accommodations granted or requested as part of the model data 1. The likely owners are compliance, claims, and customer service, and the pricing team has no obvious claim on it. If this field has not been collected before, it is a gap that needs an owner and a date.

How to handle third-party data

Third-party data is the highest-risk area in Exhibit D because the carrier does not control the source. Under the NAIC Model Bulletin, governance obligations follow the data in from outside: buying it from a vendor changes nothing about what the insurer has to be able to show 3. That means naming the vendor, describing the data, and demonstrating oversight.

For each third-party data element, keep a file that contains:

  • The vendor name and the product or data element purchased.
  • The contract or data use agreement.
  • The vendor’s methodology disclosure or model card, if available.
  • The date the data was last validated or audited for suitability.
  • Any proxy-screening or bias analysis performed on the data.
  • The internal model or system where the data feeds in.

If a vendor will not disclose the data source or methodology, the absence is itself an answer. Exhibit D asks for vendor names; the follow-up question will be whether the carrier can demonstrate oversight. For a deeper look at vendor governance, see our AI vendor risk assessment checklist. Whether that oversight is documented anywhere a regulator can find it is an Exhibit B question, and our AI governance framework implementation guide takes that one.

The proxy-discrimination connection

Exhibit D sits next to the fairness questions in Exhibit C and the governance questions in Exhibit B. A data element that is neutral on its own can become a proxy for a protected class when combined with other variables. The NAIC’s list includes several elements that have already drawn regulatory attention: geocoding, geo-demographics, age, gender, ethnicity, income, education, and medical information.

Two states have written that concern into supervision, and neither in the shape Exhibit D takes. In New York, Circular Letter No. 7 says a carrier should not put external consumer data to work in underwriting or pricing until it has worked out how far that data correlates with protected class membership, and asked whether a legitimate business necessity stands behind keeping it 4. Colorado’s adopted Regulation 10-1-1 reaches only individually issued life, private passenger auto, and health benefit plan insurers, and requires a governance framework with documented discrimination testing behind it; what counts as adequate testing is still open, because the regulation meant to set that standard has never been adopted 5. Exhibit D is the data-element checklist those exercises run on. If you cannot produce that list, you cannot show what you assessed. For the state-level detail, read our NYDFS Circular Letter No. 7 analysis and the disparate impact in AI pricing guide.

A one-page checklist for data teams

Use this checklist before an examiner asks for the Exhibit D table. A “no” is not a failure; it is a gap with a name and a deadline.

  • Have we listed every AI system from Exhibit A that uses third-party or structured data?
  • For each system, can we name the data elements used in training, testing, or production?
  • For each data element, can we identify the internal source or the third-party vendor?
  • Have we mapped the data elements to the NAIC’s 25 categories, including the new reasonable-accommodations field?
  • Do we have a feature dictionary or data dictionary that explains derived variables?
  • For third-party data, do we have the contract, methodology disclosure, and last validation date?
  • Have we documented proxy-screening or bias analysis for data elements that could stand in for protected class status?
  • Do the data elements in Exhibit D match the model descriptions in Exhibit C?
  • Do the data sources in Exhibit D match the governance answers in Exhibit B?
  • Can the entire table be produced in one file without a last-minute scramble?

If the answer is yes to most of these, the Exhibit D response is in shape. If the answer is no, start with the high-risk models flagged in Exhibit C and work outward.

The same checklist ships as a worksheet you can hand to the data team: a data-element log with the five Exhibit D fields, the 25-category reference, and these checks on their own tab.

How Exhibit D fits into the A-D sequence

The four exhibits are not independent. Exhibit A counts the systems. Exhibit B asks how they are governed. Exhibit C asks for the details of high-risk systems. Exhibit D asks for the data behind those systems. If the high-risk model in Exhibit C uses a geocoding score, the vendor name must appear in Exhibit D. Exhibit D itself has no validation column. But if Exhibit B says third-party data is validated, the vendors named in Exhibit D are the list that claim will be checked against.

A carrier can have a polished Exhibit B narrative and still fail if Exhibit D cannot name the vendor behind the third-party score. For the exhibit-by-exhibit view of the tool, and why the cross-checks between exhibits decide more exams than any single answer, see our NAIC AI Evaluation Tool: Exhibits A-D guide.

Make the data traceable

Strip away the table formatting and Exhibit D asks one thing: can you trace every data element an AI system uses to a named source? The NAIC wants a table that names the data, names the AI system, describes the use, names the internal source, and names the third-party vendor. The reasonable-accommodations row added in Version 4.0 means one more category a carrier should be able to speak to: accommodations granted or requested, sourced like any other.

Start with the 25 categories. Route each one to the owner who knows the source. Close the gaps on third-party vendor names and derived feature dictionaries. Then cross-check the answers against Exhibit A, B, and C. None of this requires the data to be clean. It requires the data to be traceable.

Traceability is a floor. A fully sourced Exhibit D can still name a data element that has no business being in the model, and identifying the vendor behind a geodemographic score answers nothing about whether that score belongs in a rating plan. That question lives in Exhibit C and in the state proxy rules, not in this table.

Footnotes

  1. NAIC, “AI Systems Evaluation Tool 4.0,” 2026: https://content.naic.org/sites/default/files/inline-files/AI%20Systems%20Evaluation%20Tool%204.0%20%28Clean%29.pdf 2 3 4 5 6 7 8 9 10

  2. Fenwick, “NAIC Expands AI Systems Evaluation Tool Pilot Program to 12 States,” 2026: https://www.fenwick.com/insights/publications/naic-expands-ai-systems-evaluation-tool-pilot-program-to-12-states-key-updates-for-insurers-and-ai-vendors-supporting-insurers 2

  3. NAIC Model Bulletin, “Use of Artificial Intelligence Systems by Insurers,” adopted December 4, 2023: https://content.naic.org/sites/default/files/inline-files/2023-12-4%20Model%20Bulletin_Adopted_0.pdf

  4. New York State Department of Financial Services, Insurance Circular Letter No. 7 (2024), “Use of Artificial Intelligence Systems and External Consumer Data and Information Sources in Insurance Underwriting and Pricing,” July 11, 2024: https://www.dfs.ny.gov/industry-guidance/circular-letters/cl2024-07

  5. Colorado Division of Insurance, “Notice of Adoption, Amended Regulation 10-1-1, Governance and Risk Management Framework,” effective October 15, 2025 (scope: individually issued life, private passenger automobile, and health benefit plan insurers; documented discrimination testing and an annual officer-signed narrative report). The quantitative testing regulation remains a 2023 draft covering life underwriting alone: https://doi.colorado.gov/announcements/notice-of-adoption-amended-regulation-10-1-1-governance-and-risk-management-framework

The Bottom Line

  • Exhibit D asks where the data came from, not whether the model performs. That question belongs to Exhibit C, which a regulator reaches first.
  • The NAIC lists 25 data element categories in Exhibit D. Every category used in training or testing needs an internal source or a named third-party vendor; categories you do not use are left blank.
  • Version 4.0 added Reasonable Accommodations or Policy Modifications to the list. Carriers should now track accommodations granted or requested alongside the model data.
  • Third-party data must be named. A vendor score without a vendor name is a missing answer in Exhibit D.
  • Exhibit D sits at the end of a sequence. Its data answers should reconcile to the system register behind Exhibit A's counts, the governance in Exhibit B, and the high-risk model details in Exhibit C.

Path complete

You have reached the evidence stage.

Return to the reading map to choose another route or business line.

Choose another reading path →
Engraved portrait of Simon Li

Written by

Simon Li · Founding Editor

Much of his time goes into reading NAIC meeting papers, state bulletins, bills, court filings, and public comments. He also keeps the site's 51-jurisdiction tracker up to date.

Contact or report a correction →

Related reading

AI Model Monitoring After the Model Goes Live

AI Governance · Prove PLAYBOOK

AI Model Monitoring After the Model Goes Live

A playbook for insurers on AI model monitoring, validation, drift detection, and retesting records that satisfy NAIC Model Bulletin and Exhibit C expectations.

JUL 31, 2026 · 12 min read

How Insurers Assess AI Vendor Risk

AI Governance · Apply PLAYBOOK

How Insurers Assess AI Vendor Risk

A practical NAIC-aligned checklist for AI vendor risk assessment: due-diligence questions, contract clauses, and the ongoing monitoring that stays with the insurer.

AUG 7, 2026 · 13 min read

Information aggregation and analysis, not legal advice. See our disclaimer.