How Disparate Impact Shapes AI Pricing in Insurance
How proxy variables and outcome differences appear in insurance AI pricing, what testing can establish, and where state-specific procedures belong.
For Compliance officers, actuaries, model-risk teams, and counsel interpreting fairness results in insurance pricing and underwriting.
Read if You need the concepts behind proxy and outcome testing before applying a particular state's procedure.
The governance map separates general control design from the legal and analytical concepts used in particular decisions. In pricing, a model can be blind to race, sex, disability, or another protected characteristic and still produce uneven results. The model may use geography, credit, education, occupation, household information, or behavior that correlates with group status. Many weak signals can also combine into a stronger proxy no analyst intended to create.
This article owns the concepts used to examine that problem: disparate treatment, disparate impact, proxy variables, outcome differences, business necessity, and less discriminatory alternatives. It does not reproduce New York’s numbered procedure or Colorado’s filing regime. Those belong to the New York NYDFS Circular Letter 7 and Colorado guides.
Treatment, impact, and proxy are different ideas
Disparate treatment concerns intentional different treatment because of a protected characteristic. Disparate impact describes a neutral practice that has a disproportionate adverse effect on a protected group. The familiar burden-shifting structure comes from federal employment law: a challenger shows the effect, the decision-maker offers business necessity, and a less discriminatory alternative may still matter.1
Insurance statutes and regulatory documents do not always use the phrase “disparate impact.” They may refer to unfair discrimination, disproportionate adverse effects, bias, or protected-class correlation. The vocabulary and legal consequence depend on the jurisdiction and line. The analytical structure remains useful because it separates three questions:
- Was protected-class status used directly or was someone treated differently because of it?
- Does a neutral input or derived feature correlate with protected-class status?
- Does the complete model produce materially different outcomes across groups after relevant risk differences are considered?
A proxy is a mechanism, not a legal conclusion. A variable can correlate with a protected group and still have a legitimate insurance relationship. A model can also show outcome differences even when no single input has a strong correlation. That is why input and output analysis should not be collapsed.
How proxy effects enter a pricing model
Pricing models work by finding patterns. The same capability that improves prediction can reconstruct social and economic differences from ordinary-looking data.
| Data family | Why it can matter | What to examine |
|---|---|---|
| Geography | Historical segregation, infrastructure, hazard, service access, and local cost can overlap | Direct risk relationship, granularity, neighborhood error, alternatives |
| Credit and financial data | Payment behavior and financial stress may correlate with loss and with protected groups | Permitted use, stability, causal story, incremental lift, group effect |
| Occupation and education | Work and educational opportunity are unevenly distributed | Insurance relevance, categorical treatment, missingness, alternatives |
| Telematics and behavior | Collection and opportunity to change behavior vary | Exposure normalization, device bias, access, discount and surcharge effects |
| Property imagery | Image quality and housing stock differ by area and property type | Error by cohort, inspection correction, filed treatment, nonrenewal effect |
| Third-party scores | A single score can hide many sources and transformations | Source disclosure, version, validation population, limitation, audit evidence |
Exhibit D of the NAIC AI Systems Evaluation Tool helps identify categories of data and their sources.2 It does not ask whether a data element is a proxy. Exhibit C asks about testing outputs for matters that include accuracy, unfair trade practices, and unfair discrimination.2 The Evaluation Tool guide owns the form. The fairness team brings the proxy and outcome questions to the data and model record.
Input testing and outcome testing answer different questions
An input screen asks how strongly a feature, source, or derived score tracks protected-class status. It can find an obvious proxy and help prioritize review. It cannot show what the whole model does after interactions, transformations, and offsets.
An outcome analysis compares decisions or prices across groups, with careful attention to legitimate risk and product differences. It can reveal a model-level disparity. It cannot by itself identify the responsible feature or settle whether the difference is legally prohibited.
Both analyses depend on design choices:
- which population and time period are included;
- how protected-class information is obtained or estimated;
- which outcome is measured, such as quote, declination, tier, premium, renewal, or surcharge;
- what constitutes a similarly situated comparison;
- how small samples and missing data are handled;
- which thresholds trigger investigation rather than an automatic verdict.
The method, limitation, and decision made from the result belong in the record. A single metric without those elements is easy to overread.
Business necessity is a documented argument
When a feature or model produces concern, the insurer should be able to explain what business need it serves and why that need is legitimate, lawful, and fair under the governing regime. Predictive lift is relevant, but “the model performs better” is incomplete. Performance may be concentrated in one segment, depend on unstable data, or add little value after a less problematic variable is used.
A less discriminatory alternative asks whether the business objective can be met with another feature, transformation, threshold, model, or workflow that reduces the disparity. Useful comparisons preserve the business target and report the tradeoff. An alternative that destroys the intended risk signal is not equivalent; an alternative rejected because it is unfamiliar has not been seriously tested.
The analysis should preserve candidates considered, evaluation method, performance and group results, the decision, and the person authorized to make it. New York specifies where and how this search fits its process; the Circular Letter guide carries those details.
Three regulatory contexts, kept separate
| Instrument | What it contributes | What it does not do |
|---|---|---|
| NAIC Model Bulletin | Program expectations for detecting and addressing unfair discrimination and managing AI risk | Prescribe one statistical test or create a new statute3 |
| NAIC Evaluation Tool | Data-source and selected-model questions a regulator can use | Use the term proxy or define the insurer’s high-risk criteria2 |
| New York Circular Letter 7 | State-specific expectations for ECDIS, AIS, outcome assessment, alternatives, governance, transparency, and notice | Create a generic procedure for every state or every insurance function4 |
| Colorado SB 21-169 and Regulation 10-1-1 | Colorado insurance authority, governance framework, records, and reporting for the covered scope | Copy New York’s language or establish one nationwide metric5 |
This comparison prevents a common documentation error: writing a blended “regulator requirement” that no regulator actually issued.
A review that can be defended
Start with one pricing or underwriting model and create a compact review file:
- Identify the decision, population, product, jurisdictions, model version, and material data sources.
- Screen likely proxy features and record the method and limitations.
- Compare relevant outcomes across groups, controlling for legitimate differences with actuarial and legal review.
- Investigate the features, interactions, thresholds, or workflow steps that may explain a material difference.
- Evaluate business necessity and credible alternatives under the applicable standard.
- Record the decision, conditions, owner, monitoring handoff, and state-specific filing or notice work.
The model monitoring playbook owns production thresholds and drift response. The vendor playbook owns third-party diligence. Keeping those controls separate prevents this fairness review from becoming another copy of the entire governance program.
What a result can and cannot prove
A correlation is a signal to investigate. It is not a verdict. A clean aggregate result may hide a subgroup or interaction problem. A statistically significant result may be operationally small, while a large result may be too imprecise to interpret. Estimated protected-class data introduces its own uncertainty.
The defensible conclusion states the limits plainly: what was tested, what was not, which legal and actuarial judgments were applied, and what happens next. Whether a particular model violates law is a fact- and jurisdiction-specific legal question. Testing supplies evidence for that judgment; it does not replace it.
Footnotes
-
Civil Rights Act of 1964, Title VII, 42 U.S.C. § 2000e-2(k). ↩
-
National Association of Insurance Commissioners, AI Systems Evaluation Tool 4.0, 2026. ↩ ↩2 ↩3
-
National Association of Insurance Commissioners, Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted December 4, 2023. ↩
-
New York State Department of Financial Services, Insurance Circular Letter No. 7 (2024), July 11, 2024. ↩
-
Colorado Division of Insurance, SB21-169 insurance AI resources and Regulation 10-1-1. ↩
The Bottom Line
- A protected-class field can be absent while a model still reconstructs group differences through correlated variables.
- Input correlation and outcome difference are separate tests. Neither result, standing alone, proves that a model is lawful or unlawful.
- Business necessity and the search for a less discriminatory alternative are part of the reasoning, not a license to skip measurement.
- New York and Colorado use different legal instruments and scopes; their detailed procedures should not be blended into a national checklist.
How Insurers Assess AI Vendor Risk
A practical NAIC-aligned checklist for AI vendor risk assessment: due-diligence questions, contract clauses, and the ongoing monitoring that stays with the insurer.
Continue →
Simon Li · Founding Editor
Much of his time goes into reading NAIC meeting papers, state bulletins, bills, court filings, and public comments. He also keeps the site's 51-jurisdiction tracker up to date.
Free · Weekly
Track these developments weekly
Get the InsureAI Wire dispatch in your inbox. Free, sourced, no spam.
Free weekly · No spam · Unsubscribe anytime
Related reading
AI Governance Documents to Prepare for a Market Conduct Exam
The insurance AI exam documentation to have ready for a market conduct exam: an insurer-built readiness file of eight evidence categories, and how to record a gap.
AI Model Monitoring After the Model Goes Live
A playbook for insurers on AI model monitoring, validation, drift detection, and retesting records that satisfy NAIC Model Bulletin and Exhibit C expectations.
Does AI Insurance Regulation Apply to You? A Plain-Language Self-Check
Five questions that flag when the AI regulations for insurance companies may reach your business, and what to review before treating the answer as settled.
The NAIC Insurance Data Security Model Law, Explained
NAIC Model Law 668 binds where a state enacts it: a written security program, third-party oversight, breach investigation, and notice to the insurance commissioner.
Information aggregation and analysis, not legal advice. See our disclaimer.