Agentic AI in Claims and the Human Review Test
How agentic AI in insurance changes claims decisions, where Exhibit C records the risk, and how carriers can test whether human review is meaningful.
In this article
For Claims, risk, and legal owners at carriers where AI already touches claim outcomes.
Read if Your claims operation runs AI somewhere (triage, fraud flags, damage estimates, settlement recommendations) and you want to know what it takes to show the human review over it is real.
Claims AI can move from support to action inside a single workflow. A model scores the file and suggests a denial; an adjuster reads the summary, agrees, and clicks approve. Ask the carrier who made that decision and the answer is the adjuster, because there was a human in the loop. The better control question is narrower. Across the last quarter, how often did that person disagree with the model, and what record shows why? If disagreement is vanishingly rare and the exceptions leave no trace, the review is functioning as a rubber stamp.
That gap, between assigning human review and proving that it changes outcomes, can surface when a carrier explains a high-risk system in Exhibit C of the NAIC AI Systems Evaluation Tool. Claims is a useful place to test it because recommendation, approval, payment, and escalation can all happen in the same operational chain. The business-line map shows where that chain sits beside underwriting, fraud, and distribution.
The full claims guide maps intake, triage, estimation, coverage, payment, notice, and appeal. This article takes one narrower question from that chain: what evidence makes human review credible once a claims system begins recommending or taking action. The count begins in Exhibit A, and the insurer’s own risk classification decides whether the deeper Exhibit C record follows.
The autonomy trap
“Agentic AI” is the phrase doing the rounds for systems that act with some degree of autonomy to reach a goal, anywhere from quietly suggesting a next step to deciding and acting on their own. The label is less useful than it sounds, because it invites the comforting line carriers like to draw. Our AI does not make the final call, so it is not high-risk. It helps to know where the tool draws its own line, because the two lines are not in the same place.
The tool sorts AI by autonomy into three bands, and defines them more tightly than the words suggest. Support provides information without suggesting a decision or action. Augmentation recommends an answer to the person making the decision. Automation proceeds without human intervention 1. Many claims tools described as supportive fit the tool’s augmentation definition.
The band is one of the fields Exhibit C asks about, and Exhibit C only reaches models the insurer has already tiered as high risk. Exhibit A uses the bands for a different purpose: augmented and automated consumer decision-making count as direct consumer impact 1. That count does not by itself send a system to Exhibit C. The insurer’s risk criteria do. The Model Bulletin asks carriers to size controls using the nature of the decision, potential harm, human involvement, explainability, and reliance on third parties 2. A person at the end of the workflow is one factor in that assessment. The quality of the review matters more than its label.
Run the opening scenario through the tool’s instruction. A model that recommends a denial augments a consumer decision and belongs in Exhibit A’s direct-impact count 1. Whether it goes any further depends on the tier the insurer assigns it, and the records behind that human sign-off are part of what defends the tier.
So the practical scope is wider than most carriers assume. Triage tools, fraud flags, damage estimators, settlement recommenders: if any of them shape what happens to a claim, they belong in scope. The useful question for your own inventory is not whether a system is autonomous. It is whether the system changes what you end up doing to a claim.
Why claims creates a sharp oversight test
A caveat first, because it changes who is on the hook. The tool does not designate claims, or any other area, as high risk. It lists fifteen operational and program areas, claims and adjudication among them, and leaves the risk criteria to the company 1. So the case for treating claims AI as the hard one is an argument a carrier has to make, or fail to make, for itself. Here is the argument, in three parts, and it also tells you where your own exposure sits.
The first is the timing of the harm. An underwriting model shapes coverage a person might draw on at some later point. A claims decision moves money the policyholder is already owed. A denial, an underpayment, or a check that arrives three weeks late while the roof is still leaking does damage right away, and it is easy for a regulator to point at.
The second is how quickly claims AI grows up. These systems rarely show up as a finished model. A triage tool that began by sorting files picks up a settlement-suggestion feature in some later release. A fraud filter that used to flag anomalies for a human starts auto-routing small claims to denial. Each of those changes adds autonomy, and almost none of them arrives with a governance review attached.
The third is the documentation gap, and this is the one that does the real damage. Underwriting and pricing models live in a part of the company that already generates paper: actuarial memos, rate filings, validation reports. Claims AI grew up in operations rather than actuarial, and a lot of it has none of that. Some of it was never even logged as a model. The fraud-scoring tool the SIU team has run for years does not sit under data science, no one ever filed it, and so it never lands on the AI inventory. That holds right up until an examiner asks how claims get flagged, and someone has to account for a system that exists nowhere on paper.
If your claims AI governance is thinner than your underwriting AI governance, that does not make you an outlier. It makes you a fairly typical carrier, and Exhibit C is where that thinness has to be written down.
Where Exhibit C stops
For a system the insurer has classified as high risk, Exhibit C records identity, version, purpose, limitations, autonomy, testing, monitoring, and compliance review.1 The Evaluation Tool guide owns the complete field-by-field explanation.
The form does not contain a human-override field. Human involvement appears in the Model Bulletin as one factor in deciding how much control a system needs.2 A claims team therefore has to keep the operational evidence beside the model record: what the reviewer saw, what authority they held, whether they could disagree, what action they took, and why.
That gap defines this article’s task. Exhibit C can show that a carrier documented a model. The claim file has to show that human involvement changed, or was capable of changing, the decision.
The governance debt nobody booked
There is a pattern worth naming, because carriers rarely recognize it from the inside. A claims AI system tends to evolve along roughly this path:
- A simple rules engine sorts incoming claims, with little documentation because it only moves work.
- A machine-learning model takes over fraud flagging, and the available record is usually technical.
- A recommendation system starts proposing settlement amounts, while the documentation still describes a routing tool.
- Low-complexity claims get approved or denied automatically within set limits, but nobody has updated the review record to match.
At every step the system got more autonomous, and at no point did anyone schedule a governance update to match. What you end up with is a kind of debt. The system is making or shaping real financial decisions while sitting on documentation written for a smaller, simpler version of itself. The capability moved on and the paperwork stayed behind.
When the system is reviewed, that debt becomes visible. The question concerns the automated approval logic from step four, while the document the carrier can produce still describes step two.
Run the three-claim override test
The three-claim test measures whether the assigned reviewer could change an outcome and whether the record proves it.
Confirm the current authority. Classify what the system does today as supportive, augmented, or automated. A triage tool that now suggests settlement amounts should be evaluated for its current authority, even if its original approval described simple routing.
Stress-test the override story. Pick three recent claims where the AI recommended one thing and a human chose another. Produce the original recommendation, the information available to the reviewer, the final decision, and the recorded reason for the difference. A missing link identifies the exact point where review stopped being reconstructable.
Measure the review conditions. Record whether the reviewer could see the model’s reason and source facts, how much time they had, whether escalation was available, and how often recommendations were changed. A low override rate may reflect a strong model, a narrow review mandate, or a rubber stamp. Sampled claim records are what distinguish those explanations.
The result is a focused control test, not a substitute for inventory, fairness analysis, or portfolio monitoring. It answers one question: did a qualified person have a real opportunity to change this claims outcome, and can the carrier prove what happened?
Footnotes
-
NAIC, “AI Systems Evaluation Tool 4.0,” 2026: https://content.naic.org/sites/default/files/inline-files/AI%20Systems%20Evaluation%20Tool%204.0%20%28Clean%29.pdf ↩ ↩2 ↩3 ↩4 ↩5
-
NAIC Model Bulletin, “Use of Artificial Intelligence Systems by Insurers,” adopted December 4, 2023: https://content.naic.org/sites/default/files/inline-files/2023-12-4%20Model%20Bulletin_Adopted_0.pdf ↩ ↩2
The Bottom Line
- A human signing off does not remove an augmented claims system from Exhibit A's direct-consumer-impact count. Exhibit C applies only after the insurer's own risk classification puts the system there.
- Map your claims AI by autonomy band (supportive / augmented / automated), and watch the tools that quietly drifted upward.
- Stress-test the override story: pull three claims where the AI recommended one thing and a human chose another. Can you produce the recommendation, the decision, and the reason?
- Screen demographic, geographic, or credit-adjacent inputs for protected-class proxies. Exhibit D makes you name the data; Exhibit C asks what testing you ran on the outputs. "We never checked" is what falls out of the gap between the two.
When Is an Insurance AI Pilot Ready to Scale?
Insurance AI pilots often fail at scale because organizations focus on the model rather than outcomes, workflow integration, and oversight. A readiness checklist.
Continue →
Simon Li · Founding Editor
I write InsureAI Wire and maintain its 51-jurisdiction tracker. Most of the work is reading: NAIC working group papers, state bulletins, bills, court filings, and public comment letters. Every claim on the site carries the document it came from, so you never have to take my word for it.
Free · Weekly
Track these developments weekly
Get the InsureAI Wire dispatch in your inbox. Free, sourced, no spam.
Free weekly · No spam · Unsubscribe anytime
Related reading
Conversational AI in Insurance and Where the Rules Reach
What conversational AI and chatbots actually do across insurance, from quotes to claims, and the point where a customer-facing bot becomes a compliance question.
AI in Insurance Claims
AI in insurance claims, step by step from intake to appeal: what each system decides, where it can go wrong, and what record makes the step reviewable.
What to Keep in an Insurance AI Decision Evidence Pack
Insurance AI decision documentation for reconstructing one underwriting or claims outcome, including human review, notice, appeal, and model version.
Will AI Replace Insurance Agents? The Work Is Splitting
Will AI replace insurance agents? What current employment projections can show, which tasks are changing, and where agency AI use creates compliance exposure.
Information aggregation and analysis, not legal advice. See our disclaimer.