# What to Keep in an Insurance AI Decision Evidence Pack

> Insurance AI decision documentation for reconstructing one underwriting or claims outcome, including human review, notice, appeal, and model version.

- Source: https://insureaiwire.com/insurance-ai-decision-evidence-pack/
- Publication: InsureAI Wire
- Author: Simon Li
- Updated: 2026-08-07

---
The practical test for insurance AI decision documentation is whether a second person can reconstruct one AI-assisted transaction without relying on oral history. Suppose a reviewer picks a decision and asks six ordinary questions. Which transaction is this, and how did it end? What information entered the process? Which system and version ran, and what did it output? Who reviewed it, and what did they do? What did the customer receive? What happened after a complaint or appeal? Every earlier control still has to connect to that one transaction.

Many organizations can answer each question somewhere. The problem is that the answers live in different systems and do not reliably point to one another. The claim file has the outcome. The model platform has a score. The vendor portal has a version. The correspondence system has the notice. The appeal log has the later correction. A decision evidence pack is the small connective record that lets a team reconstruct one transaction. Which facts have to be in it depends on the line of business, mapped in our guide to [AI by business line](/ai-by-business-line/).

<div><a class="download-cta" href="/downloads/insurance-ai-decision-evidence-pack.xlsx" download><span class="dl-label">Download the decision evidence pack</span><span class="dl-ext">XLSX</span></a></div>

<figure class="figure">
<svg viewBox="0 0 460 360" width="460" role="img">
<title>Diagram: five system labels stacked down a red spine. The claim file, model platform, vendor portal, correspondence system, and appeal log each hold one piece of an AI-assisted decision. The spine is the stable transaction identifier, the key every one of them has to carry before their separate records can be read as one decision.</title>
<line x1="26" y1="18" x2="26" y2="302" class="s-red" stroke-width="2"/>
<rect x="22" y="26" width="8" height="8" class="f-ink"/>
<text x="46" y="35" class="t-label f-ink" font-size="14">CLAIM FILE</text>
<text x="46" y="55" class="t-note f-soft" font-size="14">has the outcome</text>
<rect x="22" y="82" width="8" height="8" class="f-ink"/>
<text x="46" y="91" class="t-label f-ink" font-size="14">MODEL PLATFORM</text>
<text x="46" y="111" class="t-note f-soft" font-size="14">has a score</text>
<rect x="22" y="138" width="8" height="8" class="f-ink"/>
<text x="46" y="147" class="t-label f-ink" font-size="14">VENDOR PORTAL</text>
<text x="46" y="167" class="t-note f-soft" font-size="14">has a version</text>
<rect x="22" y="194" width="8" height="8" class="f-ink"/>
<text x="46" y="203" class="t-label f-ink" font-size="14">CORRESPONDENCE SYSTEM</text>
<text x="46" y="223" class="t-note f-soft" font-size="14">has the notice</text>
<rect x="22" y="250" width="8" height="8" class="f-ink"/>
<text x="46" y="259" class="t-label f-ink" font-size="14">APPEAL LOG</text>
<text x="46" y="279" class="t-note f-soft" font-size="14">has the later correction</text>
<text x="14" y="324" class="t-label f-red" font-size="15">STABLE TRANSACTION IDENTIFIER</text>
<text x="14" y="350" class="t-note f-soft" font-size="14">Without it, each system's record stands alone.</text>
</svg>
<figcaption>FIG. 1: FIVE SYSTEMS, ONE TRANSACTION</figcaption>
</figure>

## What this pack is, and what it is not

The [AI inventory](/ai-inventory-by-line-of-business/) covers which systems the company runs; the [model monitoring plan](/ai-model-monitoring-insurance/) covers how one of those systems behaves over time. Both sit above the transaction.

The [market conduct readiness file](/market-conduct-exam-ai-docs-ready/) sits beside it, organizing what gets handed over once a department asks. This pack works at the transaction itself: one AI-assisted business decision, replayed.

The workbook is an InsureAI Wire practice tool. It is not an NAIC exhibit, a uniform state form, or a claim that every field is legally required in every jurisdiction. The last tab marks the basis for each field using three labels:

- **Explicit source requirement** means a cited primary source directly asks for that information or record.
- **Derived from another obligation** means the field is a practical way to show compliance with a broader duty, but the source does not prescribe this exact column.
- **InsureAI Wire practice recommendation** means the field helps reconstruction and control testing, without being presented as regulator-written language.

That distinction matters. A useful internal record becomes misleading if an organization labels its own template as a regulator's checklist.

## The six linked records

### 1. Decision record

Start with a stable transaction identifier and the decision being reconstructed. Record the line of business, jurisdiction, policy or account reference, decision date, outcome, and the source material that entered the process. Do not copy unnecessary personal data into a new spreadsheet. Use controlled references to the systems that lawfully hold it.

The decision record should distinguish model output from business action. A risk score is not a declination, and a damage estimate is not a payment. Someone converted one into the other, and that conversion is the regulated act. Preserving both values shows where the system ended and the regulated decision began.

### 2. Human review and override

Record who reviewed the recommendation, their role and authority, what they could see, what they did, and why. “Human approved” is too thin when the business relies on human involvement as a control. A useful reason can be short, but it should identify the fact, policy provision, guideline, or professional judgment that explains the action.

The record should also allow “no human review” when that is the actual design. Hiding automation behind a mandatory name field produces a false record. Where meaningful review is the question, use the tests in [Agentic AI in Claims](/agentic-ai-in-claims/).

### 3. Notice, complaint, and appeal

Connect the decision to the communication the consumer actually received. Preserve the notice template version, delivery date, stated reason, and available correction or appeal path. If a complaint or appeal followed, link it to the original decision, the new evidence, the reviewer, and the final outcome.

Warehousing correspondence in the workbook is not the point. The chain breaks in two familiar places: the business can find the decision but cannot show which explanation went out, or it can find the appeal but not the model-assisted action that appeal challenged.

### 4. Model and vendor version

Record the system name, model or ruleset version, vendor release where relevant, configuration, and the lookup location for validation or change records. A product name alone is rarely enough. Two decisions processed under different model versions may look comparable while resting on different features, thresholds, or vendor data.

Vendor responsibility does not replace insurer responsibility. The NAIC Model Bulletin says an insurer should address its acquisition and use of third-party AI within its AIS Program and remains responsible for decisions and actions that affect consumers.[^1] Detailed diligence, contract, and continuing-oversight methods belong in the [vendor risk playbook](/ai-vendor-risk-assessment/). This pack only identifies the vendor version involved in the transaction.

### 5. Outcome-monitoring handoff

A transaction record should feed portfolio review without becoming the monitoring system itself. Record the cohort or monitoring group, any complaint or appeal outcome, later correction, override category, and the date the record entered monitoring. Those fields let model-risk and business owners find patterns without repeatedly reconstructing raw files.

For example, a single supplemental payment after an image estimate may be ordinary. A sustained increase for one damage type may show that the model, parts data, or review threshold needs attention. The monitoring playbook owns those thresholds and response decisions.

### 6. Instructions and source map

The last tab defines the fields and gives each one a source-status label. Keep this tab with the workbook. If the organization changes a field, it should update the definition and basis at the same time. That prevents a locally useful column from slowly acquiring a false regulatory pedigree.

## Minimum viable reconstruction

When records are fragmented, start with a small set that can answer the reviewer's questions:

**Table [row-headers]:** Minimum record links needed to reconstruct an AI-assisted insurance decision

| Question | Minimum link |
|---|---|
| What happened? | Transaction ID, decision type, date, outcome |
| What information drove it? | Controlled source references and material inputs |
| What did the system do? | System/version, output, confidence or threshold where used |
| What did a person do? | Reviewer, authority, action, reason, time |
| What did the customer receive? | Notice version, reason, delivery, recourse |
| What happened later? | Complaint or appeal, new evidence, final outcome, monitoring handoff |

This minimum is intentionally smaller than a full model file. It does not reproduce training data, validation studies, vendor contracts, committee minutes, or every policy document. It points to those materials when they are relevant.

## How to use the workbook without creating a shadow database

The downloadable workbook is a design and sampling tool. For a live production process, the stable fields should usually be implemented in governed systems rather than maintained indefinitely in a desktop file.

1. Choose one decision type, such as a homeowners repair estimate, life underwriting recommendation, fraud escalation, producer recommendation, or health claim action.
2. Complete one real sample using controlled references instead of copying sensitive data.
3. Mark every missing field as a gap, with the system owner who can close it.
4. Test whether another qualified reviewer can reconstruct the decision without oral history from the original operator.
5. Move the approved field design into the systems of record and preserve access controls, retention rules, and change history there.
6. Send the cohort and outcome fields to the monitoring owner.

Do not backfill a missing contemporaneous fact as though it had been recorded at the time. Label a reconstructed explanation, identify who reconstructed it, and preserve the basis. A candid gap has less evidentiary value than a contemporaneous record, but more integrity than a polished fiction.

## How the pack changes by business line

The six record groups stay stable, while the material facts change.

- **Underwriting:** application data, external data, eligibility or class recommendation, exception, notice, and final action.
- **Claims:** loss facts, estimate or recommendation, adjuster action, payment or denial reason, supplement, complaint, and appeal.
- **Health operations:** request, clinical or administrative criteria, recommendation, qualified review, notice, reconsideration, and final disposition.
- **Homeowners pricing:** property observations, derived features, filed or approved rating treatment, premium effect, correction request, and final result.
- **Fraud:** alert, supporting indicators, escalation, investigative action, service delay, clearance, and adverse-outcome monitoring.
- **Reinsurance:** submission and exposure data, model version, treaty recommendation, underwriting judgment, pricing or capacity action, and later performance review.
- **Distribution:** customer needs, recommendation inputs, channel or [producer](/glossary/producer/) action, disclosure, correction, and complaint outcome.

These examples do not make every transaction legally identical. They show why one generic “AI decision” field is not enough.

## Sources and limits

The NAIC AI Systems Evaluation Tool asks regulators to begin with counts by operational area and, where selected, request governance, high-risk model, and data information.[^2] Its Exhibit C includes model identity, version, implementation, risk classification, use, testing discussion, last test date, and compliance review, every one of them recorded at the system level. The transaction record underneath it, including each field in this workbook, is the insurer's own design.

The Model Bulletin describes governance, risk management, controls, documentation, testing, and third-party considerations for insurer AI systems.[^1] It is a model bulletin. A state must adopt or otherwise use it for it to carry that state's stated supervisory effect, and state-specific law remains controlling.

Claims-handling, notice, privacy, retention, and appeal duties vary by product and jurisdiction. The NAIC's [Unfair Claims Settlement Practices Act](/glossary/unfair-claims-settlement-practices-act/) illustrates established claims-handling subjects, but model laws are not uniform state law.[^3] Legal and records teams should map the implemented record to the requirements that actually apply; our [state-by-state tracker](/states/) carries the current insurance-AI guidance and the official source for each jurisdiction.

## Completion test

The pack is working when a second person can select one transaction and reproduce the sequence from source facts to system output, human action, customer communication, later challenge, and final outcome. If the reconstruction depends on the memory of the employee who handled it, the record is not finished.

The vendor decides where this stops. The model version that scored the transaction belongs on the record; the feature values behind that score usually stay with the vendor, which treats them as proprietary. That is a named gap, and it moves to whoever owns the vendor contract.

[^1]: National Association of Insurance Commissioners, [Model Bulletin on the Use of Artificial Intelligence Systems by Insurers](https://content.naic.org/sites/default/files/inline-files/2023-12-4%20Model%20Bulletin_Adopted_0.pdf), adopted December 4, 2023.
[^2]: National Association of Insurance Commissioners, [AI Systems Evaluation Tool 4.0](https://content.naic.org/sites/default/files/inline-files/AI%20Systems%20Evaluation%20Tool%204.0%20%28Clean%29.pdf), 2026.
[^3]: National Association of Insurance Commissioners, [Unfair Claims Settlement Practices Act, Model 900](https://content.naic.org/sites/default/files/model-law-900.pdf). State enactments and wording vary.