PLAYBOOK Steps, templates, and checklists you can run as written

Prove Business Lines

What to Keep in an Insurance AI Decision Evidence Pack

Insurance AI decision documentation for reconstructing one underwriting or claims outcome, including human review, notice, appeal, and model version.

For Business, compliance, claims, underwriting, operations, and model-risk teams that need to reconstruct one AI-assisted insurance decision.

Read if You can name the AI system but cannot yet replay what happened in one transaction, which version ran, who reviewed it, what the customer received, and how the matter ended.

By Simon Li · Updated AUG 7, 2026 · 8 min read

InsureAI Wire seal IAW 2026 Primary sources Methodology →

Engraved cover illustration: What to Keep in an Insurance AI Decision Evidence Pack
Ask AI

The practical test for insurance AI decision documentation is whether a second person can reconstruct one AI-assisted transaction without relying on oral history. Suppose a reviewer picks a decision and asks six ordinary questions. Which transaction is this, and how did it end? What information entered the process? Which system and version ran, and what did it output? Who reviewed it, and what did they do? What did the customer receive? What happened after a complaint or appeal? Every earlier control still has to connect to that one transaction.

Many organizations can answer each question somewhere. The problem is that the answers live in different systems and do not reliably point to one another. The claim file has the outcome. The model platform has a score. The vendor portal has a version. The correspondence system has the notice. The appeal log has the later correction. A decision evidence pack is the small connective record that lets a team reconstruct one transaction. Which facts have to be in it depends on the line of business, mapped in our guide to AI by business line.

Diagram: five system labels stacked down a red spine. The claim file, model platform, vendor portal, correspondence system, and appeal log each hold one piece of an AI-assisted decision. The spine is the stable transaction identifier, the key every one of them has to carry before their separate records can be read as one decision. CLAIM FILE has the outcome MODEL PLATFORM has a score VENDOR PORTAL has a version CORRESPONDENCE SYSTEM has the notice APPEAL LOG has the later correction STABLE TRANSACTION IDENTIFIER Without it, each system's record stands alone.
FIG. 1: FIVE SYSTEMS, ONE TRANSACTION

What this pack is, and what it is not

The AI inventory covers which systems the company runs; the model monitoring plan covers how one of those systems behaves over time. Both sit above the transaction.

The market conduct readiness file sits beside it, organizing what gets handed over once a department asks. This pack works at the transaction itself: one AI-assisted business decision, replayed.

The workbook is an InsureAI Wire practice tool. It is not an NAIC exhibit, a uniform state form, or a claim that every field is legally required in every jurisdiction. The last tab marks the basis for each field using three labels:

  • Explicit source requirement means a cited primary source directly asks for that information or record.
  • Derived from another obligation means the field is a practical way to show compliance with a broader duty, but the source does not prescribe this exact column.
  • InsureAI Wire practice recommendation means the field helps reconstruction and control testing, without being presented as regulator-written language.

That distinction matters. A useful internal record becomes misleading if an organization labels its own template as a regulator’s checklist.

The six linked records

1. Decision record

Start with a stable transaction identifier and the decision being reconstructed. Record the line of business, jurisdiction, policy or account reference, decision date, outcome, and the source material that entered the process. Do not copy unnecessary personal data into a new spreadsheet. Use controlled references to the systems that lawfully hold it.

The decision record should distinguish model output from business action. A risk score is not a declination, and a damage estimate is not a payment. Someone converted one into the other, and that conversion is the regulated act. Preserving both values shows where the system ended and the regulated decision began.

2. Human review and override

Record who reviewed the recommendation, their role and authority, what they could see, what they did, and why. “Human approved” is too thin when the business relies on human involvement as a control. A useful reason can be short, but it should identify the fact, policy provision, guideline, or professional judgment that explains the action.

The record should also allow “no human review” when that is the actual design. Hiding automation behind a mandatory name field produces a false record. Where meaningful review is the question, use the tests in Agentic AI in Claims.

3. Notice, complaint, and appeal

Connect the decision to the communication the consumer actually received. Preserve the notice template version, delivery date, stated reason, and available correction or appeal path. If a complaint or appeal followed, link it to the original decision, the new evidence, the reviewer, and the final outcome.

Warehousing correspondence in the workbook is not the point. The chain breaks in two familiar places: the business can find the decision but cannot show which explanation went out, or it can find the appeal but not the model-assisted action that appeal challenged.

4. Model and vendor version

Record the system name, model or ruleset version, vendor release where relevant, configuration, and the lookup location for validation or change records. A product name alone is rarely enough. Two decisions processed under different model versions may look comparable while resting on different features, thresholds, or vendor data.

Vendor responsibility does not replace insurer responsibility. The NAIC Model Bulletin says an insurer should address its acquisition and use of third-party AI within its AIS Program and remains responsible for decisions and actions that affect consumers.1 Detailed diligence, contract, and continuing-oversight methods belong in the vendor risk playbook. This pack only identifies the vendor version involved in the transaction.

5. Outcome-monitoring handoff

A transaction record should feed portfolio review without becoming the monitoring system itself. Record the cohort or monitoring group, any complaint or appeal outcome, later correction, override category, and the date the record entered monitoring. Those fields let model-risk and business owners find patterns without repeatedly reconstructing raw files.

For example, a single supplemental payment after an image estimate may be ordinary. A sustained increase for one damage type may show that the model, parts data, or review threshold needs attention. The monitoring playbook owns those thresholds and response decisions.

6. Instructions and source map

The last tab defines the fields and gives each one a source-status label. Keep this tab with the workbook. If the organization changes a field, it should update the definition and basis at the same time. That prevents a locally useful column from slowly acquiring a false regulatory pedigree.

Minimum viable reconstruction

When records are fragmented, start with a small set that can answer the reviewer’s questions:

Minimum record links needed to reconstruct an AI-assisted insurance decision
QuestionMinimum link
What happened?Transaction ID, decision type, date, outcome
What information drove it?Controlled source references and material inputs
What did the system do?System/version, output, confidence or threshold where used
What did a person do?Reviewer, authority, action, reason, time
What did the customer receive?Notice version, reason, delivery, recourse
What happened later?Complaint or appeal, new evidence, final outcome, monitoring handoff

This minimum is intentionally smaller than a full model file. It does not reproduce training data, validation studies, vendor contracts, committee minutes, or every policy document. It points to those materials when they are relevant.

How to use the workbook without creating a shadow database

The downloadable workbook is a design and sampling tool. For a live production process, the stable fields should usually be implemented in governed systems rather than maintained indefinitely in a desktop file.

  1. Choose one decision type, such as a homeowners repair estimate, life underwriting recommendation, fraud escalation, producer recommendation, or health claim action.
  2. Complete one real sample using controlled references instead of copying sensitive data.
  3. Mark every missing field as a gap, with the system owner who can close it.
  4. Test whether another qualified reviewer can reconstruct the decision without oral history from the original operator.
  5. Move the approved field design into the systems of record and preserve access controls, retention rules, and change history there.
  6. Send the cohort and outcome fields to the monitoring owner.

Do not backfill a missing contemporaneous fact as though it had been recorded at the time. Label a reconstructed explanation, identify who reconstructed it, and preserve the basis. A candid gap has less evidentiary value than a contemporaneous record, but more integrity than a polished fiction.

How the pack changes by business line

The six record groups stay stable, while the material facts change.

  • Underwriting: application data, external data, eligibility or class recommendation, exception, notice, and final action.
  • Claims: loss facts, estimate or recommendation, adjuster action, payment or denial reason, supplement, complaint, and appeal.
  • Health operations: request, clinical or administrative criteria, recommendation, qualified review, notice, reconsideration, and final disposition.
  • Homeowners pricing: property observations, derived features, filed or approved rating treatment, premium effect, correction request, and final result.
  • Fraud: alert, supporting indicators, escalation, investigative action, service delay, clearance, and adverse-outcome monitoring.
  • Reinsurance: submission and exposure data, model version, treaty recommendation, underwriting judgment, pricing or capacity action, and later performance review.
  • Distribution: customer needs, recommendation inputs, channel or producer action, disclosure, correction, and complaint outcome.

These examples do not make every transaction legally identical. They show why one generic “AI decision” field is not enough.

Sources and limits

The NAIC AI Systems Evaluation Tool asks regulators to begin with counts by operational area and, where selected, request governance, high-risk model, and data information.2 Its Exhibit C includes model identity, version, implementation, risk classification, use, testing discussion, last test date, and compliance review, every one of them recorded at the system level. The transaction record underneath it, including each field in this workbook, is the insurer’s own design.

The Model Bulletin describes governance, risk management, controls, documentation, testing, and third-party considerations for insurer AI systems.1 It is a model bulletin. A state must adopt or otherwise use it for it to carry that state’s stated supervisory effect, and state-specific law remains controlling.

Claims-handling, notice, privacy, retention, and appeal duties vary by product and jurisdiction. The NAIC’s Unfair Claims Settlement Practices Act illustrates established claims-handling subjects, but model laws are not uniform state law.3 Legal and records teams should map the implemented record to the requirements that actually apply; our state-by-state tracker carries the current insurance-AI guidance and the official source for each jurisdiction.

Completion test

The pack is working when a second person can select one transaction and reproduce the sequence from source facts to system output, human action, customer communication, later challenge, and final outcome. If the reconstruction depends on the memory of the employee who handled it, the record is not finished.

The vendor decides where this stops. The model version that scored the transaction belongs on the record; the feature values behind that score usually stay with the vendor, which treats them as proprietary. That is a named gap, and it moves to whoever owns the vendor contract.

Footnotes

  1. National Association of Insurance Commissioners, Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted December 4, 2023. 2

  2. National Association of Insurance Commissioners, AI Systems Evaluation Tool 4.0, 2026.

  3. National Association of Insurance Commissioners, Unfair Claims Settlement Practices Act, Model 900. State enactments and wording vary.

The Bottom Line

  • This is a transaction record, not an organization-wide AI inventory or an examiner's standard request list.
  • Keep the source facts, model and vendor version, recommendation, human action, notice, appeal, and outcome connected by stable identifiers.
  • The workbook distinguishes explicit source requirements, reasonable inferences from other duties, and InsureAI Wire practice recommendations.
  • A blank field should become a named gap, not a reconstructed fact presented as contemporaneous evidence.

Path complete

You have reached the evidence stage.

Return to the reading map to choose another route or business line.

Choose another reading path →
Engraved portrait of Simon Li

Written by

Simon Li · Founding Editor

Much of his time goes into reading NAIC meeting papers, state bulletins, bills, court filings, and public comments. He also keeps the site's 51-jurisdiction tracker up to date.

Contact or report a correction →

Related reading

AI Model Monitoring After the Model Goes Live

AI Governance · Prove PLAYBOOK

AI Model Monitoring After the Model Goes Live

A playbook for insurers on AI model monitoring, validation, drift detection, and retesting records that satisfy NAIC Model Bulletin and Exhibit C expectations.

JUL 31, 2026 · 12 min read

Information aggregation and analysis, not legal advice. See our disclaimer.