AI in Insurance Fraud Detection and the Governance of Suspicion
How insurers use AI to detect fraud, how generative AI changes the fight, and what governance is needed to avoid false positives and bias.
For Insurance fraud investigators, SIU leaders, claims executives, compliance officers, and regulators evaluating AI-enabled fraud detection.
Read if Your SIU runs on model scores, and nobody has yet written down what happens to a policyholder the model flags by mistake.
Fraud detection starts with suspicion, but a policyholder feels the actions that follow: investigation, delayed service, a claim decision, or an external referral. Insurers have long used rules, flags, and scoring models to identify suspicious claims. The scale of the data, the sophistication of the models, and the ease with which fraud itself can now be generated have raised the stakes. A useful signal can protect consumers and control loss. A false positive can delay a legitimate claim and create bad-faith exposure.
Insurance fraud is expensive. The Coalition Against Insurance Fraud estimates that at least $308.6 billion is stolen from American consumers every year, and puts fraud in about 10% of property-casualty insurance losses 1. The losses are passed on to policyholders, which makes fraud detection a consumer-protection issue as well as a profitability issue. The site’s fraud branch places this workflow beside the claims decisions it can influence.
How AI detects fraud
Traditional fraud detection relies on rules. If a claim is filed within a certain window after a policy is issued, or if the same provider appears in multiple unrelated claims, the system flags it for review. Rules are transparent and easy to audit, but they are brittle. Fraudsters learn the rules and route around them.
Machine learning changes the approach. Instead of looking for known patterns, an ML model looks for statistical anomalies: a claim that is unusual given the policyholder’s history, a provider whose billing pattern differs from peers, or a repair estimate that is out of line with local market rates. Health insurers have been among the earliest adopters. The NAIC’s 2025 health AI/ML survey puts fraud detection near the top of the table 2. Among insurers that use or plan to use AI in a given operational area, 70% already have fraud detection in production 2. That ties utilization management and sales and marketing, and trails only strategic operations. Prior authorization and product pricing sit at the other end of the same table, at 29% and 37%. The same claims-adjudication dynamics are covered in AI in health insurance claims, prior authorization, and risk adjustment.
P&C carriers are following. In a recent Deloitte survey of insurance executives, 35% ranked fraud detection among the top five areas for developing or implementing generative AI applications within the next twelve months 3. Deloitte also puts a number on the prize: $80 billion to $160 billion in P&C savings by 2032 3. Two conditions travel with that number. The forecast assumes carriers integrate real-time analysis across text, image, audio, and video; fraud detection deployed on its own sits outside it. And the range is built up from the projected size of the fraud-detection technology market and an assumed 20% to 40% savings rate, so it estimates what the technology could be worth 3. No carrier has banked it.
The methods are becoming more diverse. Insurers now use:
- Network analysis to identify rings of providers, attorneys, or claimants.
- Natural language processing to read claim notes, police reports, and medical records for inconsistencies.
- Computer vision to compare damage photos against repair estimates or to detect manipulated images.
- Behavioral analytics to flag claimants whose behavior differs from baseline patterns.
- Anomaly detection to find outliers in billing, treatment, or repair costs.
Generative AI works both sides of the claim
Generative AI is changing fraud in both directions. The NAIC Model Bulletin defines generative AI as a class of systems that produce data, text, images, sounds, or video similar to but not copied from pre-existing content 4. That definition doubles as an inventory of what a claim file holds: the narrative, the receipt, the repair estimate, the damage photo, the recorded call. Anything that can generate those five can fabricate them, and the number of people holding the capability is no longer small. The eMarketer forecast published in June 2024 projected 100.1 million people in the US using generative AI at least once a month that year, up from 7.8 million in 2022, an increase it puts at nearly 900% 5. No one has published a count of claim documents actually fabricated this way, and any figure offered as one is an estimate.
On the defense side, generative AI is being used to summarize claims files, identify inconsistencies across documents, and suggest follow-up questions for investigators. A model can read thousands of pages of medical records or police reports in minutes and flag contradictions that a human reviewer might miss. The same technology that creates fake documents can also detect them.
The challenge is that generative AI is not a reliable fact-checker. It can hallucinate, misinterpret context, or produce confident-sounding but incorrect conclusions. A fraud detection program that relies on generative AI without human verification is exposed to two errors: missing real fraud and flagging innocent claims.
Four ways the suspicion machine misfires
The risks in fraud detection AI are concentrated in four areas.
False positives. A model that flags too many legitimate claims creates consumer harm and operational cost. The harm is not just delay. A wrongfully flagged claim can lead to denied payment, litigation, and bad-faith damages. The model must be calibrated so that the cost of investigating false positives does not exceed the savings from detecting fraud.
Bias and proxy discrimination. Fraud models trained on historical data may inherit the biases of past investigations. If a zip code, occupation, or provider type was over-investigated in the past, the model may learn to flag it as suspicious even when the underlying behavior is not fraudulent. This can produce unfair outcomes that look like proxy discrimination.
Opacity and due process. Policyholders and providers need to know why a claim is delayed or denied. A fraud model that produces a score without explaining the reason takes that away. The Model Bulletin lists lack of transparency and explainability among the risks insurers should act to minimize, and makes the explainability of an outcome to the affected consumer one of five factors that set how tight the controls around a system have to be 4. The NAIC AI Principles the bulletin incorporates put the same idea in terms of recourse: stakeholders, consumers included, should have a way to inquire about, review, and seek recourse for AI-driven insurance decisions, in language that describes the factors behind the decision 6.
Data misuse. Fraud detection often depends on third-party data: public records, credit data, social media, and device data. Each source carries privacy and accuracy risks. A model that uses stale or inaccurate data can flag legitimate claims and expose the insurer to fair-credit-reporting and privacy violations.
Four transitions that decide what the flag becomes
A fraud score begins as a signal. Its consequences depend on the operational step that follows.
Signal to work queue. Record the model or ruleset version, the threshold crossed, and the indicators shown to the investigator. The model monitoring playbook owns threshold changes and portfolio performance; the fraud file needs enough detail to explain why this claim entered the queue.
Work queue to investigation. The investigator should be able to distinguish the model’s suspicion from corroborating facts. Record what they reviewed, what evidence changed the assessment, and who authorized the investigation. If the investigator sees only a score and almost never disagrees, use the meaningful human-review test to examine whether review exists in practice.
Investigation to service restriction. Extra document requests, payment holds, special handling, or delayed communication affect the policyholder before a formal denial occurs. The record should identify the restriction, its start and end, the reason, and the claim-handling clock that continued to apply. Coverage, payment, notice, and appeal remain part of the full claims lifecycle.
Suspicion to adverse action or external referral. A denial, cancellation, law-enforcement referral, or industry-database report requires a named decision maker and evidence beyond the score itself. Preserve the predicate facts, the authorized action, the notice or communication, and the later outcome. The decision evidence pack provides the transaction spine.
Fairness analysis and vendor diligence still apply, but their full methods live elsewhere. Use the proxy and outcome-testing guide when historical investigations may have shaped the model’s outcomes, and the vendor assessment when an outside provider controls the score, data, or release. Fraud’s own job is to show how a suspicion moved, or did not move, into action.
The Coalition Against Insurance Fraud reports that 78% of consumers are concerned about insurance fraud 1. That supports investigation while leaving the decision boundary intact: public concern about fraud does not turn a model score into proof.
The claims cluster, with a criminal-justice edge
Fraud owns the false-positive and escalation problem. Intake, estimation, coverage, payment, notice, and appeal belong to the full claims lifecycle. When a flag changes a particular transaction, the decision evidence pack connects the alert, investigation, business action, notice, and later correction. Keeping those scopes separate prevents a fraud flag from being treated as though it were already a claim decision.
On the org chart, fraud detection belongs to claims. In a governance program, it borrows from everywhere: underwriting’s model validation, pricing’s fairness testing, vendor management’s audit rights. Its own contribution is the cost of being wrong. A false positive in pricing produces a bad quote; a false positive here produces an accusation, sometimes a referral, against a policyholder who did nothing, and that is a file the carrier will be asked to explain long after the model has moved on.
Footnotes
-
Coalition Against Insurance Fraud, “Fraud Stats” (Impact panel: “Insurance fraud steals at least $308.6B every year from American consumers”; “Fraud occurs in about 10% of property-casualty insurance losses”; By Category, Consumer Attitudes: “78% say they are concerned about insurance fraud”). The page is a rolling summary and carries no date on these three figures; the dated studies listed beneath them run from 2019 to 2021: https://insurancefraud.org/fraud-stats/ ↩ ↩2
-
NAIC, “Health Insurance Artificial Intelligence/Machine Learning Survey Report,” May 2025, Survey Analysis Summary table, p.2, read against Table 5, p.14-15 (the percentages are the share already in production among companies that use or plan to use AI/ML in that operational area, not a share of all respondents; fraud detection 42/60 = 70%, product pricing 22/59 = 37%, prior authorization 18/62 = 29%): https://content.naic.org/sites/default/files/inline-files/Health%20Survey%20Report%20-%20FINAL%205.9.25.pdf ↩ ↩2
-
Deloitte, “Property and casualty carriers can win the fight against insurance fraud,” FSI Predictions 2025, April 2025 (the $80-160 billion range is conditioned on integrating “real-time analysis from multiple modalities”; per “About this prediction” it is derived from the fraud-detection technology market growing at a 25% CAGR from US$4 billion in 2023 to US$32 billion by 2032, with an assumed 20% to 40% savings rate): https://www.deloitte.com/us/en/insights/industry/financial-services/financial-services-industry-predictions/2025/ai-to-fight-insurance-fraud.html ↩ ↩2 ↩3
-
NAIC, “Use of Artificial Intelligence Systems by Insurers,” Model Bulletin adopted December 4, 2023 (Definitions, Generative AI; Section 1, risks including “lack of transparency and explainability”; the Section 3 preamble, clause (iv), explainability as one of five factors setting the strength of controls; Section 2, UTPA and UCSPA “regardless of the methods the Insurer used”; Section 3, the written AIS Program; §3.2, “bias analysis and minimization”): https://content.naic.org/sites/default/files/inline-files/2023-12-4%20Model%20Bulletin_Adopted_0.pdf ↩ ↩2
-
eMarketer, “Generative AI hits 100 million user milestone in US,” article by Sara Lebow, July 30, 2024, reporting eMarketer’s June 2024 forecast: “This year 100.1 million people in the US will use generative AI at least once per month, up nearly 900% from 7.8 million in 2022.” The figure is a forecast, not a count: https://www.emarketer.com/content/generative-ai-hits-100-million-users-milestone-us ↩
-
NAIC, “Principles on Artificial Intelligence (AI),” adopted by the Innovation and Technology (EX) Task Force August 7, 2020, by the Executive (EX) Committee August 13, 2020, and by the Executive (EX) Committee and Plenary August 14, 2020 (five headings: Fair and Ethical, Accountable, Compliant, Transparent, Secure, Safe and Robust; “These principles are guidance and do not carry the weight of law or impose any legal liability”; Transparent (b) on inquiry, review, and recourse): https://content.naic.org/sites/default/files/inline-files/AI%20principles%20as%20Adopted%20by%20the%20TF_0807.pdf ↩
The Bottom Line
- Fraud detection is a longstanding and closely scrutinized use of AI in insurance. It sits at the intersection of financial crime, consumer protection, and model risk.
- Machine learning can detect patterns that rules-based systems miss, but it can also generate false positives that delay legitimate claims and create bad faith exposure.
- Generative AI plays offense and defense at once. The same class of system that can produce text, images, and audio can also read documents, summarize claims, and flag inconsistencies.
- The critical boundary is where suspicion turns into action: investigation, service delay, a claim decision, or an external referral.
- A fraud record should preserve the signal, the investigator's basis, the action taken, and what happened when the suspicion was cleared or challenged.
When Is an Insurance AI Pilot Ready to Scale?
Insurance AI pilots often fail at scale because organizations focus on the model rather than outcomes, workflow integration, and oversight. A readiness checklist.
Continue →
Simon Li · Founding Editor
Much of his time goes into reading NAIC meeting papers, state bulletins, bills, court filings, and public comments. He also keeps the site's 51-jurisdiction tracker up to date.
Free · Weekly
Track these developments weekly
Get the InsureAI Wire dispatch in your inbox. Free, sourced, no spam.
Free weekly · No spam · Unsubscribe anytime
Related reading
AI in Insurance Claims
AI in insurance claims, step by step from intake to appeal: what each system decides, where it can go wrong, and what record makes the step reviewable.
What to Keep in an Insurance AI Decision Evidence Pack
Insurance AI decision documentation for reconstructing one underwriting or claims outcome, including human review, notice, appeal, and model version.
Will AI Replace Insurance Agents? The Work Is Splitting
Will AI replace insurance agents? What current employment projections can show, which tasks are changing, and where agency AI use creates compliance exposure.
Conversational AI in Insurance and Where the Rules Reach
What conversational AI and chatbots actually do across insurance, from quotes to claims, and the point where a customer-facing bot becomes a compliance question.
Information aggregation and analysis, not legal advice. See our disclaimer.